munotes®

The Model-Based Reflex Agent

Get access to whole semester resourcesSemester Pass

Chapter Six

Syllabus topic Module 1, "model-based"

Pages 24 to 28 of 591

In one line

A model-based agent keeps a picture of the world in its head, updates it after every percept and every action, and decides from the picture instead of from the percept.

In the wording a student can write in an examination: a model-based reflex agent maintains internal state, a representation of the aspects of the world that its percepts do not reveal. The state is updated using two pieces of knowledge: a transition model, which says how the world changes, including as a result of the agent's own actions; and a sensor model, which says how the state of the world is reflected in the percepts. The agent then applies condition-action rules to the updated state rather than to the raw percept.

Why internal state is the fix

The previous chapter established the problem exactly: two situations can produce the same percept and need different actions. If the information is not in the percept, the only other place it can come from is the past, and using the past means keeping something.

The agent does not keep the whole percept history. That is the table-driven agent and it is unbuildable. It keeps a summary of the history that is sufficient for the decision, and working out what that summary must contain is the design problem.

For the vacuum world the summary is small: what do I believe about the square I am not standing in. Two extra facts, and the agent becomes able to finish.

The two models

The state is updated from two sources and both have a name. Getting the names the right way round is worth marks.

Transition modelSensor model
Answershow does the world changehow does the world show up in a percept
Coversthe effects of my own actions, and changes I did not causewhat my sensors report, given the state
In the vacuum worldsucking makes this square clean; nothing else changes on its ownI see the square I am in and its dirt, and nothing about the other square
Later in this bookthe transition function of an MDP, chapter 79the emission matrix of an HMM, chapter 72

The update loop

Four steps, run in this order, every time round.

  1. Predict: apply the transition model for the action just taken, to get the state the world should now be in.
  2. Update from the percept: apply the sensor model in reverse, using what was actually perceived to correct or sharpen the predicted state.
  3. Decide: apply the rules to the state.
  4. Act, and remember what action was taken, because step 1 next time needs it.
munotes.in24

The Model-Based Reflex Agent

It running, on the same world

Both squares dirty, agent starts in A, one point per clean square per step, minus one per move. Identical to the previous chapter in every respect except the agent.

# A model-based reflex agent: it keeps a MODEL of the parts it cannot see.
class ModelBasedAgent:
    def __init__(self):
        self.model = {"A": "unknown", "B": "unknown"}

    def __call__(self, percept):
        here, dirt = percept
        self.model[here] = dirt                       # update state from percept
        if dirt == "Dirty":
            self.model[here] = "Clean"                # and from the action's effect
            return "Suck"
        other = "B" if here == "A" else "A"
        if self.model[other] == "Clean":
            return "NoOp"                             # nothing left to do: STOP
        return "Right" if here == "A" else "Left"

agent = ModelBasedAgent()
world, at, score = {"A": "Dirty", "B": "Dirty"}, "A", 0
for step in range(1, 9):
    percept = (at, world[at])
    action = agent(percept)
    if action == "Suck":
        world[at] = "Clean"
    elif action == "Right":
        at = "B"
    elif action == "Left":
        at = "A"
    score += sum(1 for v in world.values() if v == "Clean")
    if action in ("Left", "Right"):
        score -= 1
    print("step %d  %-16s %-6s  model %-36s  score %d"
          % (step, str(percept), action, str(agent.model), score))
step 1  ('A', 'Dirty')   Suck    model {'A': 'Clean', 'B': 'unknown'}        score 1
step 2  ('A', 'Clean')   Right   model {'A': 'Clean', 'B': 'unknown'}        score 1
step 3  ('B', 'Dirty')   Suck    model {'A': 'Clean', 'B': 'Clean'}          score 3
step 4  ('B', 'Clean')   NoOp    model {'A': 'Clean', 'B': 'Clean'}          score 5
step 5  ('B', 'Clean')   NoOp    model {'A': 'Clean', 'B': 'Clean'}          score 7
step 6  ('B', 'Clean')   NoOp    model {'A': 'Clean', 'B': 'Clean'}          score 9
step 7  ('B', 'Clean')   NoOp    model {'A': 'Clean', 'B': 'Clean'}          score 11
step 8  ('B', 'Clean')   NoOp    model {'A': 'Clean', 'B': 'Clean'}          score 13

Watch the model column, which is the agent's belief and not the world.

  • Step 1: it is in A and sees Dirty. It records Dirty, decides to Suck, and immediately records Clean because it knows what sucking does. That is the transition model in one line.
  • Step 2: the percept (A, Clean) is now not enough on its own, and it does not need to be. The model says B is unknown, so there is work to do and it moves right.
  • Step 3: B is dirty. Suck, and record it.
  • Step 4: the percept is (B, Clean), the identical percept that made the simple reflex agent move. This agent looks at its model, sees A recorded as Clean, and returns NoOp. It has worked out that it is finished.
  • Steps 5 to 8: NoOp, costing nothing, scoring 2 every step.

Final score 13, against the simple reflex agent's 8 on the same world. The whole difference is that it can stop, and it can stop because it remembered.

munotes.in25

The Model-Based Reflex Agent

What "unknown" is doing in the model

The initial model is not "both clean" or "both dirty". It is unknown, and that is the honest representation of an agent that has just been switched on. It matters twice.

It stops the agent concluding it is finished before it has looked. If the model started at Clean for both, step 1 would have sucked and step 2 would have returned NoOp with B still filthy.

And it is the first appearance of a theme that runs to the end of the book: an agent's state is a belief, not a fact, and the honest thing to store is what it actually knows. In Module 1's fourth row that belief becomes a probability distribution; in Module 2 it becomes a fitted model. The structure is the same.

Where the model can be wrong

A model-based agent acts on its belief, so a wrong belief produces a wrong action, confidently.

  • The world changed and the agent did not see it. Somebody drops dirt in A at step 5. The model still says Clean and the agent sits doing NoOp on a dirty floor. This is what makes a dynamic environment hard.
  • The action did not do what the model says. Sucking fails one time in five. The model records Clean, the square is dirty. This is a stochastic environment, and the fix is not a better model of this kind but a probabilistic one.
  • The sensors were wrong. The dirt sensor misreports. Then even the directly perceived part of the model is unreliable.

None of these is an argument against keeping a model. They are the argument for keeping a model with probabilities in it, which is Module 1's fourth row, and for learning the model instead of being given it, which is Module 2.

The vocabulary, because two pairs get confused

WordMeansNot to be confused with
Internal statewhat the agent has stored about the worldthe state of the environment, which is the real thing
Belief statethe agent's internal state when it may be uncertainthe true state
Transition modelhow the world changesthe sensor model
Sensor modelhow the world shows in perceptsthe transition model

Distinctions

Simple reflexModel-based reflex
Decides fromthe current perceptinternal state, updated from the percept
Internal statenoneyes
Needscondition-action rulesrules, a transition model and a sensor model
Partially observablefailsworks
Can it tell that it is finishednoyes
Score on this chapter's world813
Model-based reflexGoal-based
Knows where it isyesyes
Knows where it wants to beno, only what to do nextyes, explicitly
If the goal changesthe rules must be rewrittenthe goal is changed and the rest stands
Considers action sequencesnoyes, which is search
munotes.in26

The Model-Based Reflex Agent

What it does not mean

The model is not the world. It is what the agent believes. A chapter that blurs the two makes every later error in this book invisible.

It does not store the percept history. It stores a summary sufficient for the decision. Storing the history would be the table-driven agent.

It is still a reflex agent. The final step is still a rule applied to a situation. What changed is that the situation is now the internal state and not the raw percept, which is why the name keeps the word "reflex".

A model does not make the agent correct. It makes it correct as long as the model is. A stale or wrong model produces confident wrong actions.

"Unknown" is not the same as "clean". Initialising a model to a definite value the agent has not observed is a bug, and it is the specific bug that would make this agent stop before it started.

Quick revision

  • Model-based reflex agent: keeps internal state, updates it from a transition model (how the world changes, including by its own actions) and a sensor model (how the world appears in percepts), then applies rules to the state.
  • The update loop: predict from the last action, correct from the percept, decide, act and remember.
  • It solves the simple reflex agent's first failure, because the missing information now comes from the past instead of from the percept.
  • On the same vacuum world it scores 13 against the simple reflex agent's 8, and the difference is that it returns NoOp at step 4.
  • The state must start as unknown, not as a guessed value.
  • It fails when the world changes unseen, when an action does not do what the model says, or when the sensors lie. Those three failures are the reason for probability in Module 1 and learning in Module 2.
  • It is still reflex: a rule applied to a situation. Only the situation changed.

Test yourself

1. Define a model-based reflex agent. An agent that maintains internal state representing the parts of the world its percepts do not reveal, updates that state using a transition model and a sensor model, and then applies condition-action rules to the state rather than to the percept.

2. What is the transition model, and what is the sensor model? The transition model says how the world changes, including as a result of the agent's own actions. The sensor model says how the state of the world is reflected in what the sensors report.

munotes.in27

The Model-Based Reflex Agent

3. In this chapter's run, why does step 4 return NoOp when the simple reflex agent moved Left? Both see the percept (B, Clean). The model-based agent also has A recorded as Clean in its internal state, so it can conclude that nothing is left to do. The simple reflex agent has only the percept, which does not contain that fact.

4. Why is the model initialised to "unknown" rather than to "clean"? Because the agent has observed nothing yet. If it began by believing both squares clean, it would return NoOp after cleaning the first square and leave the second dirty.

5. Give three ways the model can become wrong, and name the environment property responsible in each case. The world changes without the agent seeing it, which is a dynamic environment. An action fails to have its expected effect, which is a stochastic environment. The sensors misreport, which is noise and makes even the perceived part unreliable.

6. Is a model-based agent still a reflex agent? Justify. Yes. Its final step is a condition-action rule applied to a situation, with no consideration of goals or of action sequences. What changed is that the situation is the internal state, not the raw percept.

7. Distinguish internal state from the state of the environment. The state of the environment is how the world actually is. Internal state is what the agent has stored about it, which may be incomplete, stale or wrong. The agent acts on the second and is judged on the first.

munotes.in28

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!