munotes®

Reinforcement Learning: The Third Form

Get access to whole semester resourcesSemester Pass

Chapter Forty-Six

Syllabus topic Module 2, "reinforcement"

Pages 250 to 253 of 591

In one line

Reinforcement learning is learning from a score rather than from an answer: the agent is told how well it did, not what it should have done.

In the wording a student can write in an examination: in reinforcement learning an agent interacts with an environment, taking actions in states and receiving a numerical reward. It is not told the correct action; it must discover which actions yield the most reward by trying them. Two features distinguish it: the reward may be delayed, so an action's value depends on consequences far in the future, which is the credit assignment problem; and the agent's actions determine what data it sees, so it must balance exploration against exploitation.

What the feedback is

The difference from the other two forms is entirely in what the critic says, and it is worth stating three times over.

The agent does something, and the critic saysThe form
"the right answer was 7"supervised
nothing at allunsupervised
"that was worth 3 points"reinforcement

A reward says how good the outcome was, not what should have been done instead. That is a much weaker signal than a label. Told that a chess move scored badly, the agent does not learn which move was better; it learns only that this one was poor, and must try others to find out.

The two things that make it a different problem

The reward is delayed. A move in chess is followed by fifty more before the game is won or lost, and only then does a reward arrive. Which of the fifty deserves the credit? That is the credit assignment problem, and it does not arise in supervised learning at all, where the label arrives with the example.

The agent generates its own data. A supervised learner is handed a training set. A reinforcement learner sees only the consequences of the actions it chose, so a good action never tried is never learned about. It must sometimes take an action it believes is worse, purely to find out. That is the exploration against exploitation problem from The Learning Agent, and Q-Learning is where it becomes a concrete rule.

Where it fits, and where it does not

A paper can ask when reinforcement learning is the right choice, and the answer has conditions.

It suits a problem where the right action is not known but the outcome can be scored, where the agent can act repeatedly and cheaply, and where actions have consequences that unfold over time. Games, control, scheduling, resource allocation.

It does not suit a problem where acting is expensive or dangerous, because it learns by trying and its early attempts are bad. A reinforcement learner cannot be let loose on a real patient, a real vehicle or a real power station to find out what happens. The standard answer is a simulator, and the standard failure is that the simulator differs from reality in some way that matters.

munotes.in250

Reinforcement Learning: The Third Form

The three forms, in one table

This is the answer to MU's Forms of learning (supervised, unsupervised, reinforcement) label, and it is the table to reproduce.

SupervisedUnsupervisedReinforcement
The critic givesthe correct outputnothinga numerical reward
Datalabelled pairsunlabelled instancesexperience the agent generated
Feedback timingimmediatenonepossibly much delayed
The agent chooses its own datanonoyes
Central difficultyobtaining labelsno way to score a resultcredit assignment, and exploration
Goalpredict the labelfind structuremaximise total reward over time
MU's methodsk-NN, trees, naive Bayes, SVM, networks, ensemblesclustering, association rulesMDPs, Q-learning

And the one-line version worth memorising: supervised learning is told the answer, unsupervised learning is told nothing, and reinforcement learning is told the score.

Its relation to Module 1

The cross-module link, and a likely Q.3.

The Utility-Based Agent maximised expected utility over outcomes whose probabilities it was given. Reinforcement learning is that agent when nobody gave it the probabilities or the utilities: it must estimate both from experience. Markov Decision Processes sets out the problem when the model IS known, and Q-Learning solves it when it is not.

So reinforcement learning is not a third kind of prediction. It is a decision problem, and it is the only part of Module 2 that inherits directly from Module 1's agent framework rather than from its data.

Distinctions

A rewardA label
Sayshow good the outcome waswhat the correct output was
Tells you the right actionnoyes
Can arrive lateyesno
Fromthe environmentan annotator
Reinforcement learningSupervised learning
Who chose the datathe agent, by actingsomebody else, in advance
A good action never tried isnever learnedirrelevant, the data is fixed
Needs explorationyesno
Credit assignmentExploration against exploitation
The problemwhich of many past actions earned this rewardtake the best known action, or an informative one
Becausethe reward is delayedthe agent generates its own data
Solved in this book bythe value of a state, Markov Decision Processesepsilon-greedy, Q-Learning

What it does not mean

A reward is not a label. It scores an outcome and does not name a correct action.

Reinforcement learning is not trial and error without structure. It estimates the value of states and actions, and the estimate is what improves.

Delayed reward is not a minor complication. It is the defining difficulty, and the reason a whole row of MU's syllabus is given to it.

munotes.in251

Reinforcement Learning: The Third Form

Exploration is not noise. It is a deliberate, sometimes costly choice to act suboptimally in order to learn.

It is not always applicable. Where acting is expensive or dangerous it cannot learn by trying, and a simulator is needed, with all the risk of the simulator being wrong.

Quick revision

  • Reinforcement learning: an agent acts in states, receives a numerical reward, and must discover which actions yield the most reward over time.
  • The reward says how good, not what was right. A much weaker signal than a label.
  • Two defining difficulties: credit assignment, because the reward is delayed; and exploration against exploitation, because the agent generates its own data.
  • Supervised is told the answer, unsupervised is told nothing, reinforcement is told the score. Three kinds of critic.
  • Suits problems where the right action is unknown but the outcome can be scored and acting is cheap and repeatable. Unsuitable where acting is dangerous or expensive; a simulator is then used, and may differ from reality.
  • It is The Utility-Based Agent with the probabilities and utilities unknown. Markov Decision Processes is the known-model case, Q-Learning the unknown one.
  • It is a decision problem, not a prediction problem, and the only part of Module 2 inheriting from Module 1's agent framework.

Test yourself

1. Define reinforcement learning and say what the agent receives. An agent takes actions in states of an environment and receives a numerical reward. It is not told the correct action; it must discover by trying which actions yield the greatest total reward over time.

2. How does a reward differ from a label? A label states the correct output for an example. A reward states how good an outcome was, without saying which action would have been better, and it may arrive long after the action that earned it.

3. What is the credit assignment problem, and why does it not arise in supervised learning? When a reward arrives after a long sequence of actions, it is not clear which of them deserves the credit or blame. It does not arise in supervised learning because the correct output arrives with each example, so the feedback is attached to the case it concerns.

4. Why must a reinforcement learner explore? Because it sees only the consequences of the actions it takes, so an action it never tries is never learned about. Occasionally acting against its current belief is the only way to discover something better.

5. Give the three forms of learning in terms of the critic. Supervised learning has a critic that supplies the correct answer. Unsupervised learning has no critic at all. Reinforcement learning has a critic that supplies a numerical reward, which says how good the outcome was and nothing more.

munotes.in252

Reinforcement Learning: The Third Form

6. When is reinforcement learning unsuitable, and what is the usual workaround? When acting is expensive or dangerous, because the method learns by trying and its early attempts are poor. The usual workaround is to train in a simulator, with the standing risk that the simulator differs from reality in a way that matters.

7. How does reinforcement learning relate to the utility-based agent of Module 1? The utility-based agent maximises expected utility using probabilities and utilities it was given. Reinforcement learning is the same objective when neither has been given, so both must be estimated from experience. Markov decision processes handle the case where the model is known, and Q-learning the case where it is not.

munotes.in253

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!