Types of Environment
Chapter Four
Syllabus topic Module 1, "Types of environments"
Pages 15 to 19 of 591
In one line
Environments differ in ways that change what an agent has to be able to do, and there are seven standard questions to ask about any one of them.
In the wording a student can write in an examination: a task environment is classified along seven dimensions: fully or partially observable, single agent or multi agent, deterministic or stochastic, episodic or sequential, static or dynamic, discrete or continuous, and known or unknown. The classification determines which agent architecture is adequate, because each dimension rules out a simpler design.
Why classify at all
An agent design is only adequate relative to an environment. A design that works perfectly in one is useless in another, and the seven questions are how you find out which you are in before writing any code. The hardest case on every dimension is the real world, which is why the automated taxi is the standing example of a difficult task.
1. Fully observable or partially observable
Fully observable: the agent's sensors give it the complete state of the environment at each instant, so far as anything relevant to its choice of action is concerned.
Partially observable: they do not. Something relevant is hidden, either because the sensors are limited, because they are noisy, or because part of the state is somewhere else.
- Fully observable: chess with a board in view, a crossword, the vacuum world if the agent can see both squares.
- Partially observable: the vacuum world with a local dirt sensor only, poker, driving (you cannot see what is behind the lorry), medical diagnosis.
- Unobservable: a special case with no sensors at all. Such an agent can still act sensibly if the world is predictable enough.
Consequence for the agent. A partially observable environment forces the agent to keep internal state, because the current percept does not determine the right action. That is exactly the step from the simple reflex agent to the model-based one, two chapters from now.
2. Single agent or multi agent
Single agent: the agent is the only decision maker. Everything else is part of the environment.
Multi agent: another entity is also choosing, and its choices depend on what our agent does.
The test is not whether other things move; it is whether their behaviour is best described as maximising a performance measure that depends on ours. A falling stone is part of the environment. A taxi in the next lane is another agent, because it will brake if you pull in front of it.
Multi agent splits in two:
- Competitive: one agent's gain is another's loss. Chess. This is where Module 1's
Adversarial searchrow comes from. - Cooperative: the measures agree at least partly. Two taxis both wanting to avoid a collision.
Types of Environment
Consequence for the agent. In a competitive setting the agent must reason about what the other will do, which is minimax. It may also become rational to behave randomly, to stop the opponent predicting you, which never pays in a single agent setting.
3. Deterministic or stochastic
Deterministic: the next state is fixed completely by the current state and the action taken.
Stochastic: it is not. The same action in the same state can lead to different next states, and the agent can at best have probabilities.
- Deterministic: chess, the 8-puzzle, the vacuum world as defined in the last chapter.
- Stochastic: driving (a tyre may burst), a robot's wheels slipping, any system with real sensors.
"Nondeterministic" is a third word and is not a synonym for stochastic. A nondeterministic environment has several possible outcomes but attaches no probabilities to them, so the agent has to succeed whatever happens. A stochastic one attaches probabilities, so the agent can maximise an expectation.
Consequence for the agent. Determinism is what makes a plan a sequence of actions. Once the environment is stochastic, a plan has to be a policy: what to do in each state that might arise. That is the whole of Module 2's Markov Decision Processes.
4. Episodic or sequential
Episodic: the agent's experience divides into independent episodes. The action taken in one episode has no effect on the next.
Sequential: the current decision affects all future decisions.
- Episodic: classifying a part on a conveyor belt; answering an exam question that stands alone; a spam filter judging one message.
- Sequential: chess, driving, a game of any kind, a course of medical treatment.
Consequence for the agent. In an episodic environment the agent need not think ahead at all, which makes it much easier. Almost all of Module 2's supervised learning is episodic: each prediction stands alone. Almost all of Module 1 is sequential, which is why it is full of search.
5. Static or dynamic
Static: the environment does not change while the agent is deciding.
Dynamic: it does, so time spent thinking is time in which the world moved.
Semidynamic: the environment is static but the agent's SCORE changes with time. Chess with a clock is the standard example: the board does not move while you think, but your remaining time does.
- Static: a crossword, an offline puzzle.
- Dynamic: driving, a robot in a room with people in it.
- Semidynamic: timed chess.
Consequence for the agent. In a dynamic environment the agent must act on incomplete deliberation, so an anytime algorithm that can give its best answer so far matters more than an optimal one that takes too long. This is the practical reason the search chapters report time complexity rather than only correctness.
Types of Environment
6. Discrete or continuous
The distinction applies separately to the state, to time, and to the percepts and actions, and a question may ask about any of them.
- Discrete state and time: chess. Finitely many board positions, moves at distinct instants.
- Continuous state and time: driving. Position and speed are real numbers changing smoothly, and steering is a continuous action.
- Mixed: a camera is a discrete sensor sampling a continuous world, and what it delivers is a grid of integers.
Consequence for the agent. Every search algorithm in Module 1 assumes a discrete state space with a finite set of actions at each state. A continuous problem has to be discretised before those algorithms apply, and the discretisation is a design decision that can make the problem easy or impossible.
7. Known or unknown
This is not the same as observable, and mixing the two is the commonest mistake on this topic.
Known: the agent knows the rules. It knows what its actions do, and in a stochastic environment it knows the probabilities.
Unknown: it does not, and it must learn them.
Observability is about the state; knownness is about the laws. All four combinations exist:
| Known | Unknown | |
|---|---|---|
| Fully observable | chess: you see the board and know the rules | a new video game with the board on screen and no manual |
| Partially observable | poker: hidden cards, known rules | driving in an unfamiliar country at night |
Consequence for the agent. An unknown environment is the reason learning exists. It is also the difference between Module 2's Markov Decision Processes, where the model is known and the answer can be computed, and Q-Learning, where it is not and the agent has to find out by acting.
The classification, done
This table is the answer to "classify the following environment", and the six columns are the ones papers use.
| Chess with a clock | The 8-puzzle | Driving a taxi | Medical diagnosis | Spam filter | Part-picking robot | |
|---|---|---|---|---|---|---|
| Observable | fully | fully | partially | partially | fully | partially |
| Agents | multi | single | multi | single | single | single |
| Deterministic | deterministic | deterministic | stochastic | stochastic | deterministic | stochastic |
| Episodic | sequential | sequential | sequential | sequential | episodic | episodic |
| Static | semidynamic | static | dynamic | dynamic | static | dynamic |
| Discrete | discrete | discrete | continuous | continuous | discrete | continuous |
| Known | known | known | known | unknown | unknown | known |
Two entries in that table are worth defending, because a marker may disagree and a student should be able to argue.
Spam filter, fully observable. The message is entirely available to the filter; nothing about it is hidden. What is hidden is the sender's INTENTION, but intention is not part of the environment state the filter acts on. Some texts call it partially observable for that reason; say which you mean.
Types of Environment
Part-picking robot, episodic. Each part is judged on its own and the bin it goes in does not change how the next part should be judged. If the bins can overflow, it becomes sequential.
The hardest case
The worst case on every dimension at once is partially observable, multi agent, stochastic, sequential, dynamic, continuous and unknown. That is driving in a city you have never visited. It is why the automated taxi is used as the running example of a hard task, and why no algorithm in this paper solves it on its own.
Distinctions
| Partially observable | Unknown | |
|---|---|---|
| What is missing | part of the current state | the rules of the environment |
| The fix | keep internal state, and reason under uncertainty | learn from experience |
| Example | poker | a game whose rules you have not been told |
| Where on this syllabus | model-based agents, Bayesian networks | the whole of Module 2 |
| Stochastic | Nondeterministic | |
|---|---|---|
| Outcomes | several, with probabilities | several, with no probabilities |
| The agent aims at | the best expectation | success whatever happens |
| Needs | a probability model | a plan for every contingency |
| Episodic | Sequential | |
|---|---|---|
| This decision affects later ones | no | yes |
| Lookahead needed | none | yes |
| Typical of | classification | games, planning, control |
What it does not mean
Fully observable does not mean simple. Chess is fully observable and unsolved.
Deterministic does not mean predictable by the agent. A deterministic environment the agent does not understand is effectively unpredictable to it, which is why "known" is a separate dimension.
Multi agent does not mean many moving objects. It means another entity is choosing, and its choice responds to yours.
Dynamic does not mean fast. It means the state changes while the agent deliberates, however slowly.
Continuous does not mean infinite state only. Time can be continuous while the state is discrete, and the dimensions are asked separately.
Quick revision
- Seven dimensions: observable (fully or partially), agents (single or multi), deterministic or stochastic, episodic or sequential, static or dynamic, discrete or continuous, known or unknown.
- Partially observable forces internal state. Stochastic forces a policy instead of a plan. Sequential forces lookahead. Dynamic forces acting on incomplete deliberation. Continuous forces discretisation. Unknown forces learning. Multi agent forces reasoning about the other, and can make randomising rational.
- Semidynamic: the world is still but the score moves, as in timed chess.
- Nondeterministic is not stochastic: outcomes without probabilities, so the agent must cover every case.
- Observable is about the state; known is about the rules. All four combinations occur.
- Hardest case, and the standing example: partially observable, multi agent, stochastic, sequential, dynamic, continuous, unknown.
Types of Environment
Test yourself
1. List the seven dimensions on which a task environment is classified. Fully or partially observable; single or multi agent; deterministic or stochastic; episodic or sequential; static or dynamic; discrete or continuous; known or unknown.
2. Classify the environment of a taxi driving in Mumbai. Partially observable, multi agent, stochastic, sequential, dynamic, continuous and, for an unfamiliar city, unknown. It is the hardest case on every dimension.
3. Distinguish a partially observable environment from an unknown one. Partially observable means part of the current state is hidden from the sensors; unknown means the agent does not know the laws by which the environment behaves. One is fixed by keeping state and reasoning under uncertainty, the other by learning.
4. What is a semidynamic environment? Give an example. One in which the environment itself does not change while the agent deliberates, but the agent's performance score does. Chess played with a clock: the board is still, your remaining time is not.
5. Why does a partially observable environment force an agent to keep internal state? Because the current percept no longer determines the right action, so two situations needing different actions can look identical. The only way to tell them apart is to remember what came before.
6. Is a stone falling towards a robot another agent? No. Its behaviour is not usefully described as maximising a measure that depends on the robot's choices, so it is part of the environment. A second robot that would swerve to avoid a collision is another agent.
7. Which dimension is the reason the whole of Module 2 exists, and why? Known or unknown. If the agent already knew the environment's laws and the right answers, it would need no experience. Learning is the response to not knowing.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.