The Model-Based Reflex Agent
Chapter Six
Syllabus topic Module 1, "model-based"
Pages 24 to 28 of 591
In one line
A model-based agent keeps a picture of the world in its head, updates it after every percept and every action, and decides from the picture instead of from the percept.
In the wording a student can write in an examination: a model-based reflex agent maintains internal state, a representation of the aspects of the world that its percepts do not reveal. The state is updated using two pieces of knowledge: a transition model, which says how the world changes, including as a result of the agent's own actions; and a sensor model, which says how the state of the world is reflected in the percepts. The agent then applies condition-action rules to the updated state rather than to the raw percept.
Why internal state is the fix
The previous chapter established the problem exactly: two situations can produce the same percept and need different actions. If the information is not in the percept, the only other place it can come from is the past, and using the past means keeping something.
The agent does not keep the whole percept history. That is the table-driven agent and it is unbuildable. It keeps a summary of the history that is sufficient for the decision, and working out what that summary must contain is the design problem.
For the vacuum world the summary is small: what do I believe about the square I am not standing in. Two extra facts, and the agent becomes able to finish.
The two models
The state is updated from two sources and both have a name. Getting the names the right way round is worth marks.
| Transition model | Sensor model | |
|---|---|---|
| Answers | how does the world change | how does the world show up in a percept |
| Covers | the effects of my own actions, and changes I did not cause | what my sensors report, given the state |
| In the vacuum world | sucking makes this square clean; nothing else changes on its own | I see the square I am in and its dirt, and nothing about the other square |
| Later in this book | the transition function of an MDP, chapter 79 | the emission matrix of an HMM, chapter 72 |
The update loop
Four steps, run in this order, every time round.
- Predict: apply the transition model for the action just taken, to get the state the world should now be in.
- Update from the percept: apply the sensor model in reverse, using what was actually perceived to correct or sharpen the predicted state.
- Decide: apply the rules to the state.
- Act, and remember what action was taken, because step 1 next time needs it.
The Model-Based Reflex Agent
It running, on the same world
Both squares dirty, agent starts in A, one point per clean square per step, minus one per move. Identical to the previous chapter in every respect except the agent.
# A model-based reflex agent: it keeps a MODEL of the parts it cannot see.
class ModelBasedAgent:
def __init__(self):
self.model = {"A": "unknown", "B": "unknown"}
def __call__(self, percept):
here, dirt = percept
self.model[here] = dirt # update state from percept
if dirt == "Dirty":
self.model[here] = "Clean" # and from the action's effect
return "Suck"
other = "B" if here == "A" else "A"
if self.model[other] == "Clean":
return "NoOp" # nothing left to do: STOP
return "Right" if here == "A" else "Left"
agent = ModelBasedAgent()
world, at, score = {"A": "Dirty", "B": "Dirty"}, "A", 0
for step in range(1, 9):
percept = (at, world[at])
action = agent(percept)
if action == "Suck":
world[at] = "Clean"
elif action == "Right":
at = "B"
elif action == "Left":
at = "A"
score += sum(1 for v in world.values() if v == "Clean")
if action in ("Left", "Right"):
score -= 1
print("step %d %-16s %-6s model %-36s score %d"
% (step, str(percept), action, str(agent.model), score))step 1 ('A', 'Dirty') Suck model {'A': 'Clean', 'B': 'unknown'} score 1
step 2 ('A', 'Clean') Right model {'A': 'Clean', 'B': 'unknown'} score 1
step 3 ('B', 'Dirty') Suck model {'A': 'Clean', 'B': 'Clean'} score 3
step 4 ('B', 'Clean') NoOp model {'A': 'Clean', 'B': 'Clean'} score 5
step 5 ('B', 'Clean') NoOp model {'A': 'Clean', 'B': 'Clean'} score 7
step 6 ('B', 'Clean') NoOp model {'A': 'Clean', 'B': 'Clean'} score 9
step 7 ('B', 'Clean') NoOp model {'A': 'Clean', 'B': 'Clean'} score 11
step 8 ('B', 'Clean') NoOp model {'A': 'Clean', 'B': 'Clean'} score 13Watch the model column, which is the agent's belief and not the world.
- Step 1: it is in A and sees Dirty. It records Dirty, decides to Suck, and immediately records Clean because it knows what sucking does. That is the transition model in one line.
- Step 2: the percept (A, Clean) is now not enough on its own, and it does not need to be. The model says B is unknown, so there is work to do and it moves right.
- Step 3: B is dirty. Suck, and record it.
- Step 4: the percept is (B, Clean), the identical percept that made the simple reflex agent move. This agent looks at its model, sees A recorded as Clean, and returns NoOp. It has worked out that it is finished.
- Steps 5 to 8: NoOp, costing nothing, scoring 2 every step.
Final score 13, against the simple reflex agent's 8 on the same world. The whole difference is that it can stop, and it can stop because it remembered.
The Model-Based Reflex Agent
What "unknown" is doing in the model
The initial model is not "both clean" or "both dirty". It is unknown, and that is the honest representation of an agent that has just been switched on. It matters twice.
It stops the agent concluding it is finished before it has looked. If the model started at Clean for both, step 1 would have sucked and step 2 would have returned NoOp with B still filthy.
And it is the first appearance of a theme that runs to the end of the book: an agent's state is a belief, not a fact, and the honest thing to store is what it actually knows. In Module 1's fourth row that belief becomes a probability distribution; in Module 2 it becomes a fitted model. The structure is the same.
Where the model can be wrong
A model-based agent acts on its belief, so a wrong belief produces a wrong action, confidently.
- The world changed and the agent did not see it. Somebody drops dirt in A at step 5. The model still says Clean and the agent sits doing NoOp on a dirty floor. This is what makes a dynamic environment hard.
- The action did not do what the model says. Sucking fails one time in five. The model records Clean, the square is dirty. This is a stochastic environment, and the fix is not a better model of this kind but a probabilistic one.
- The sensors were wrong. The dirt sensor misreports. Then even the directly perceived part of the model is unreliable.
None of these is an argument against keeping a model. They are the argument for keeping a model with probabilities in it, which is Module 1's fourth row, and for learning the model instead of being given it, which is Module 2.
The vocabulary, because two pairs get confused
| Word | Means | Not to be confused with |
|---|---|---|
| Internal state | what the agent has stored about the world | the state of the environment, which is the real thing |
| Belief state | the agent's internal state when it may be uncertain | the true state |
| Transition model | how the world changes | the sensor model |
| Sensor model | how the world shows in percepts | the transition model |
Distinctions
| Simple reflex | Model-based reflex | |
|---|---|---|
| Decides from | the current percept | internal state, updated from the percept |
| Internal state | none | yes |
| Needs | condition-action rules | rules, a transition model and a sensor model |
| Partially observable | fails | works |
| Can it tell that it is finished | no | yes |
| Score on this chapter's world | 8 | 13 |
| Model-based reflex | Goal-based | |
|---|---|---|
| Knows where it is | yes | yes |
| Knows where it wants to be | no, only what to do next | yes, explicitly |
| If the goal changes | the rules must be rewritten | the goal is changed and the rest stands |
| Considers action sequences | no | yes, which is search |
The Model-Based Reflex Agent
What it does not mean
The model is not the world. It is what the agent believes. A chapter that blurs the two makes every later error in this book invisible.
It does not store the percept history. It stores a summary sufficient for the decision. Storing the history would be the table-driven agent.
It is still a reflex agent. The final step is still a rule applied to a situation. What changed is that the situation is now the internal state and not the raw percept, which is why the name keeps the word "reflex".
A model does not make the agent correct. It makes it correct as long as the model is. A stale or wrong model produces confident wrong actions.
"Unknown" is not the same as "clean". Initialising a model to a definite value the agent has not observed is a bug, and it is the specific bug that would make this agent stop before it started.
Quick revision
- Model-based reflex agent: keeps internal state, updates it from a transition model (how the world changes, including by its own actions) and a sensor model (how the world appears in percepts), then applies rules to the state.
- The update loop: predict from the last action, correct from the percept, decide, act and remember.
- It solves the simple reflex agent's first failure, because the missing information now comes from the past instead of from the percept.
- On the same vacuum world it scores 13 against the simple reflex agent's 8, and the difference is that it returns NoOp at step 4.
- The state must start as unknown, not as a guessed value.
- It fails when the world changes unseen, when an action does not do what the model says, or when the sensors lie. Those three failures are the reason for probability in Module 1 and learning in Module 2.
- It is still reflex: a rule applied to a situation. Only the situation changed.
Test yourself
1. Define a model-based reflex agent. An agent that maintains internal state representing the parts of the world its percepts do not reveal, updates that state using a transition model and a sensor model, and then applies condition-action rules to the state rather than to the percept.
2. What is the transition model, and what is the sensor model? The transition model says how the world changes, including as a result of the agent's own actions. The sensor model says how the state of the world is reflected in what the sensors report.
The Model-Based Reflex Agent
3. In this chapter's run, why does step 4 return NoOp when the simple reflex agent moved Left? Both see the percept (B, Clean). The model-based agent also has A recorded as Clean in its internal state, so it can conclude that nothing is left to do. The simple reflex agent has only the percept, which does not contain that fact.
4. Why is the model initialised to "unknown" rather than to "clean"? Because the agent has observed nothing yet. If it began by believing both squares clean, it would return NoOp after cleaning the first square and leave the second dirty.
5. Give three ways the model can become wrong, and name the environment property responsible in each case. The world changes without the agent seeing it, which is a dynamic environment. An action fails to have its expected effect, which is a stochastic environment. The sensors misreport, which is noise and makes even the perceived part unreliable.
6. Is a model-based agent still a reflex agent? Justify. Yes. Its final step is a condition-action rule applied to a situation, with no consideration of goals or of action sequences. What changed is that the situation is the internal state, not the raw percept.
7. Distinguish internal state from the state of the environment. The state of the environment is how the world actually is. Internal state is what the agent has stored about it, which may be incomplete, stale or wrong. The agent acts on the second and is judged on the first.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.