munotes®

Types of Environment

Get access to whole semester resourcesSemester Pass

Chapter Four

Syllabus topic Module 1, "Types of environments"

Pages 15 to 19 of 591

In one line

Environments differ in ways that change what an agent has to be able to do, and there are seven standard questions to ask about any one of them.

In the wording a student can write in an examination: a task environment is classified along seven dimensions: fully or partially observable, single agent or multi agent, deterministic or stochastic, episodic or sequential, static or dynamic, discrete or continuous, and known or unknown. The classification determines which agent architecture is adequate, because each dimension rules out a simpler design.

Why classify at all

An agent design is only adequate relative to an environment. A design that works perfectly in one is useless in another, and the seven questions are how you find out which you are in before writing any code. The hardest case on every dimension is the real world, which is why the automated taxi is the standing example of a difficult task.

1. Fully observable or partially observable

Fully observable: the agent's sensors give it the complete state of the environment at each instant, so far as anything relevant to its choice of action is concerned.

Partially observable: they do not. Something relevant is hidden, either because the sensors are limited, because they are noisy, or because part of the state is somewhere else.

  • Fully observable: chess with a board in view, a crossword, the vacuum world if the agent can see both squares.
  • Partially observable: the vacuum world with a local dirt sensor only, poker, driving (you cannot see what is behind the lorry), medical diagnosis.
  • Unobservable: a special case with no sensors at all. Such an agent can still act sensibly if the world is predictable enough.

Consequence for the agent. A partially observable environment forces the agent to keep internal state, because the current percept does not determine the right action. That is exactly the step from the simple reflex agent to the model-based one, two chapters from now.

2. Single agent or multi agent

Single agent: the agent is the only decision maker. Everything else is part of the environment.

Multi agent: another entity is also choosing, and its choices depend on what our agent does.

The test is not whether other things move; it is whether their behaviour is best described as maximising a performance measure that depends on ours. A falling stone is part of the environment. A taxi in the next lane is another agent, because it will brake if you pull in front of it.

Multi agent splits in two:

  • Competitive: one agent's gain is another's loss. Chess. This is where Module 1's Adversarial search row comes from.
  • Cooperative: the measures agree at least partly. Two taxis both wanting to avoid a collision.
munotes.in15

Types of Environment

Consequence for the agent. In a competitive setting the agent must reason about what the other will do, which is minimax. It may also become rational to behave randomly, to stop the opponent predicting you, which never pays in a single agent setting.

3. Deterministic or stochastic

Deterministic: the next state is fixed completely by the current state and the action taken.

Stochastic: it is not. The same action in the same state can lead to different next states, and the agent can at best have probabilities.

  • Deterministic: chess, the 8-puzzle, the vacuum world as defined in the last chapter.
  • Stochastic: driving (a tyre may burst), a robot's wheels slipping, any system with real sensors.

"Nondeterministic" is a third word and is not a synonym for stochastic. A nondeterministic environment has several possible outcomes but attaches no probabilities to them, so the agent has to succeed whatever happens. A stochastic one attaches probabilities, so the agent can maximise an expectation.

Consequence for the agent. Determinism is what makes a plan a sequence of actions. Once the environment is stochastic, a plan has to be a policy: what to do in each state that might arise. That is the whole of Module 2's Markov Decision Processes.

4. Episodic or sequential

Episodic: the agent's experience divides into independent episodes. The action taken in one episode has no effect on the next.

Sequential: the current decision affects all future decisions.

  • Episodic: classifying a part on a conveyor belt; answering an exam question that stands alone; a spam filter judging one message.
  • Sequential: chess, driving, a game of any kind, a course of medical treatment.

Consequence for the agent. In an episodic environment the agent need not think ahead at all, which makes it much easier. Almost all of Module 2's supervised learning is episodic: each prediction stands alone. Almost all of Module 1 is sequential, which is why it is full of search.

5. Static or dynamic

Static: the environment does not change while the agent is deciding.

Dynamic: it does, so time spent thinking is time in which the world moved.

Semidynamic: the environment is static but the agent's SCORE changes with time. Chess with a clock is the standard example: the board does not move while you think, but your remaining time does.

  • Static: a crossword, an offline puzzle.
  • Dynamic: driving, a robot in a room with people in it.
  • Semidynamic: timed chess.

Consequence for the agent. In a dynamic environment the agent must act on incomplete deliberation, so an anytime algorithm that can give its best answer so far matters more than an optimal one that takes too long. This is the practical reason the search chapters report time complexity rather than only correctness.

munotes.in16

Types of Environment

6. Discrete or continuous

The distinction applies separately to the state, to time, and to the percepts and actions, and a question may ask about any of them.

  • Discrete state and time: chess. Finitely many board positions, moves at distinct instants.
  • Continuous state and time: driving. Position and speed are real numbers changing smoothly, and steering is a continuous action.
  • Mixed: a camera is a discrete sensor sampling a continuous world, and what it delivers is a grid of integers.

Consequence for the agent. Every search algorithm in Module 1 assumes a discrete state space with a finite set of actions at each state. A continuous problem has to be discretised before those algorithms apply, and the discretisation is a design decision that can make the problem easy or impossible.

7. Known or unknown

This is not the same as observable, and mixing the two is the commonest mistake on this topic.

Known: the agent knows the rules. It knows what its actions do, and in a stochastic environment it knows the probabilities.

Unknown: it does not, and it must learn them.

Observability is about the state; knownness is about the laws. All four combinations exist:

KnownUnknown
Fully observablechess: you see the board and know the rulesa new video game with the board on screen and no manual
Partially observablepoker: hidden cards, known rulesdriving in an unfamiliar country at night

Consequence for the agent. An unknown environment is the reason learning exists. It is also the difference between Module 2's Markov Decision Processes, where the model is known and the answer can be computed, and Q-Learning, where it is not and the agent has to find out by acting.

The classification, done

This table is the answer to "classify the following environment", and the six columns are the ones papers use.

Chess with a clockThe 8-puzzleDriving a taxiMedical diagnosisSpam filterPart-picking robot
Observablefullyfullypartiallypartiallyfullypartially
Agentsmultisinglemultisinglesinglesingle
Deterministicdeterministicdeterministicstochasticstochasticdeterministicstochastic
Episodicsequentialsequentialsequentialsequentialepisodicepisodic
Staticsemidynamicstaticdynamicdynamicstaticdynamic
Discretediscretediscretecontinuouscontinuousdiscretecontinuous
Knownknownknownknownunknownunknownknown

Two entries in that table are worth defending, because a marker may disagree and a student should be able to argue.

Spam filter, fully observable. The message is entirely available to the filter; nothing about it is hidden. What is hidden is the sender's INTENTION, but intention is not part of the environment state the filter acts on. Some texts call it partially observable for that reason; say which you mean.

munotes.in17

Types of Environment

Part-picking robot, episodic. Each part is judged on its own and the bin it goes in does not change how the next part should be judged. If the bins can overflow, it becomes sequential.

The hardest case

The worst case on every dimension at once is partially observable, multi agent, stochastic, sequential, dynamic, continuous and unknown. That is driving in a city you have never visited. It is why the automated taxi is used as the running example of a hard task, and why no algorithm in this paper solves it on its own.

Distinctions

Partially observableUnknown
What is missingpart of the current statethe rules of the environment
The fixkeep internal state, and reason under uncertaintylearn from experience
Examplepokera game whose rules you have not been told
Where on this syllabusmodel-based agents, Bayesian networksthe whole of Module 2
StochasticNondeterministic
Outcomesseveral, with probabilitiesseveral, with no probabilities
The agent aims atthe best expectationsuccess whatever happens
Needsa probability modela plan for every contingency
EpisodicSequential
This decision affects later onesnoyes
Lookahead needednoneyes
Typical ofclassificationgames, planning, control

What it does not mean

Fully observable does not mean simple. Chess is fully observable and unsolved.

Deterministic does not mean predictable by the agent. A deterministic environment the agent does not understand is effectively unpredictable to it, which is why "known" is a separate dimension.

Multi agent does not mean many moving objects. It means another entity is choosing, and its choice responds to yours.

Dynamic does not mean fast. It means the state changes while the agent deliberates, however slowly.

Continuous does not mean infinite state only. Time can be continuous while the state is discrete, and the dimensions are asked separately.

Quick revision

  • Seven dimensions: observable (fully or partially), agents (single or multi), deterministic or stochastic, episodic or sequential, static or dynamic, discrete or continuous, known or unknown.
  • Partially observable forces internal state. Stochastic forces a policy instead of a plan. Sequential forces lookahead. Dynamic forces acting on incomplete deliberation. Continuous forces discretisation. Unknown forces learning. Multi agent forces reasoning about the other, and can make randomising rational.
  • Semidynamic: the world is still but the score moves, as in timed chess.
  • Nondeterministic is not stochastic: outcomes without probabilities, so the agent must cover every case.
  • Observable is about the state; known is about the rules. All four combinations occur.
  • Hardest case, and the standing example: partially observable, multi agent, stochastic, sequential, dynamic, continuous, unknown.
munotes.in18

Types of Environment

Test yourself

1. List the seven dimensions on which a task environment is classified. Fully or partially observable; single or multi agent; deterministic or stochastic; episodic or sequential; static or dynamic; discrete or continuous; known or unknown.

2. Classify the environment of a taxi driving in Mumbai. Partially observable, multi agent, stochastic, sequential, dynamic, continuous and, for an unfamiliar city, unknown. It is the hardest case on every dimension.

3. Distinguish a partially observable environment from an unknown one. Partially observable means part of the current state is hidden from the sensors; unknown means the agent does not know the laws by which the environment behaves. One is fixed by keeping state and reasoning under uncertainty, the other by learning.

4. What is a semidynamic environment? Give an example. One in which the environment itself does not change while the agent deliberates, but the agent's performance score does. Chess played with a clock: the board is still, your remaining time is not.

5. Why does a partially observable environment force an agent to keep internal state? Because the current percept no longer determines the right action, so two situations needing different actions can look identical. The only way to tell them apart is to remember what came before.

6. Is a stone falling towards a robot another agent? No. Its behaviour is not usefully described as maximising a measure that depends on the robot's choices, so it is part of the environment. A second robot that would swerve to avoid a collision is another agent.

7. Which dimension is the reason the whole of Module 2 exists, and why? Known or unknown. If the agent already knew the environment's laws and the right answers, it would need no experience. Learning is the response to not knowing.

munotes.in19

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!