The Utility-Based Agent
Chapter Eight
Syllabus topic Module 1, "utility-"
Pages 34 to 38 of 591
In one line
A utility-based agent puts a number on how good each outcome is, so it can pick the best of several plans that all work, and can weigh a small chance of a great outcome against a large chance of a fair one.
In the wording a student can write in an examination: a utility-based agent uses a utility function, a mapping from a state (or a sequence of states) to a real number measuring how desirable it is. Where outcomes are uncertain, the agent computes the expected utility of each action, the sum over possible outcomes of their probability times their utility, and selects the action of maximum expected utility. This is the principle of maximum expected utility, and it is the formal statement of rational behaviour used throughout this paper.
Why a goal is not enough
The previous chapter ended with three things a binary goal cannot do. Each becomes possible the moment states carry numbers.
Ranking two successes. Two plans clean the house. One takes three actions, the other four. A goal test says "both fine". A utility function that subtracts the cost of each move says which is better, and by how much.
Trading off conflicting objectives. Fast and safe disagree. With numbers, the question becomes an arithmetic one: is the two minutes saved worth the extra risk, at these values.
Preferring a better chance. If sucking works four times in five, no plan is certain. Utility lets the agent compare a plan that succeeds 92 times in 100 with one that succeeds 64 times in 100, and to notice that the safer plan also costs more electricity.
The two things you must be given
A utility-based agent cannot be built out of thin air. It needs two numerical inputs, and both are design decisions that a marker may ask you to justify.
- A utility function on outcomes. Here: a clean square at the end is worth 5. That is a choice; 5 is not discovered.
- Probabilities for the uncertain effects of actions. Here: sucking works with probability 0.8. Where those come from is the subject of Module 1's fourth row, and how they are learned from data is Module 2.
Add the costs, which are negative utility: every move costs 1 and every suck costs 0.2 in electricity.
Expected utility, defined and then computed
For an action with possible outcomes numbered 1 to n:
EU(action) = sum over i of P(outcome i) * U(outcome i)
Read it in words: take each thing that could happen, multiply how good it is by how likely it is, and add them up. Nothing more.
For one Suck on a dirty square in A, with the other square dirty:
The Utility-Based Agent
EU(Suck) = 0.8 U(A clean, B dirty) + 0.2 U(A dirty, B dirty) - 0.2
= 0.8 5 + 0.2 0 - 0.2
= 4.0 - 0.2
= 3.8
The 0.2 subtracted at the end is the electricity. That figure, 3.8, appears in the table below as the expected utility of the one-action plan Suck, so the hand arithmetic and the program agree.
Comparing whole plans
The program pushes a probability distribution over states through each action of a plan, accumulates the cost as it goes, and then computes the expected value of the squares that are clean at the end.
# A utility-based agent. It puts a NUMBER on every outcome, so it can compare two
# plans that BOTH reach the goal, which a goal-based agent cannot.
# Sucking works with probability 0.8. A clean square at the end is worth 5.
# Every move costs 1 and every suck costs 0.2 in electricity.
P_WORKS, MOVE_COST, SUCK_COST, CLEAN_VALUE = 0.8, 1.0, 0.2, 5.0
def step(dist, action):
"""Push a probability distribution over states through one action."""
out, cost = {}, 0.0
for (at, a, b), p in dist.items():
if action == "Suck":
cost += p * SUCK_COST
here = a if at == "A" else b
if here == "Dirty":
done = (at, "Clean", b) if at == "A" else (at, a, "Clean")
out[done] = out.get(done, 0.0) + p * P_WORKS
out[(at, a, b)] = out.get((at, a, b), 0.0) + p * (1 - P_WORKS)
else:
out[(at, a, b)] = out.get((at, a, b), 0.0) + p
elif action in ("Left", "Right"):
cost += p * MOVE_COST
nxt = ("A" if action == "Left" else "B", a, b)
out[nxt] = out.get(nxt, 0.0) + p
else:
out[(at, a, b)] = out.get((at, a, b), 0.0) + p
return out, cost
def expected_utility(plan, start):
dist, spent = {start: 1.0}, 0.0
for action in plan:
dist, cost = step(dist, action)
spent += cost
value = sum(p * CLEAN_VALUE * sum(1 for s in st[1:] if s == "Clean")
for st, p in dist.items())
return value - spent
start = ("A", "Dirty", "Dirty")
PLANS = [("Suck", "Right", "Suck"),
("Suck", "Suck", "Right", "Suck", "Suck"),
("Right", "Suck", "Left", "Suck"),
("Suck",),
()]
print("plan P(goal) expected utility")
for plan in PLANS:
d = {start: 1.0}
for action in plan:
d, _ = step(d, action)
pgoal = sum(p for st, p in d.items() if st[1] == "Clean" and st[2] == "Clean")
label = " then ".join(plan) if plan else "do nothing"
print("%-44s %7.3f %15.3f" % (label, pgoal, expected_utility(plan, start)))
print()
best = max(PLANS, key=lambda pl: expected_utility(pl, start))
print("chooses: %s" % (" then ".join(best) if best else "do nothing"))plan P(goal) expected utility
Suck then Right then Suck 0.640 6.600
Suck then Suck then Right then Suck then Suck 0.922 7.800
Right then Suck then Left then Suck 0.640 5.600
Suck 0.000 3.800
do nothing 0.000 0.000
chooses: Suck then Suck then Right then Suck then SuckThe Utility-Based Agent
Read the table row by row, because every row settles one of the three things a goal could not do.
Row 1 and row 3 both reach the goal with probability 0.640, and they are exactly the two plans the previous chapter could not choose between. Utility separates them: 6.600 against 5.600. The difference is one extra move, costing 1.
Row 2 is the agent's choice, and it is the interesting one. It sucks twice in each square. That is pointless if sucking always works, and it is the right thing when sucking works four times in five: the probability of finishing rises from 0.640 to 0.922, and the two extra sucks cost only 0.4 between them. A goal-based agent would never have considered it, because one suck already "reaches the goal" in its model.
Row 4 is the single action worked by hand above, 3.800, and it confirms the arithmetic.
Row 5 is doing nothing, worth 0.000, which is the baseline every other row is measured against.
Why utility is the definition of rationality used from here on
The definition in Acting Rationally, or Thinking Like a Human said an agent is rational when it selects the action expected to maximise its performance measure. This chapter is where that sentence becomes computable. Expected utility IS the expected performance measure, and maximum expected utility is the rule that implements the definition.
That is why the same quantity reappears three more times in this book under three more names:
| Where | What it is called | What it is |
|---|---|---|
| Reasoning under uncertainty | expected value | probability times value, summed |
| Markov decision processes | the value of a state under a policy | expected discounted utility of what follows |
| Evaluating a model | expected loss, or risk | expected utility with the sign turned round |
Utility against money, and why the function is not linear
One subtlety that is worth 5 marks and is usually missed. Utility is not the same as the quantity being measured.
Offered a certain 500 rupees or a coin flip for 1,000, most people take the 500, although the expected rupees are identical. That is not irrational. It means their utility for money is not linear: the first 500 rupees is worth more to them than the second 500. An agent whose utility function bends this way is called risk averse; one whose function bends the other way is risk seeking; a straight line is risk neutral.
The Utility-Based Agent
So a utility function is where an agent's attitude to risk lives. Two agents with the same probabilities and the same money can rationally make opposite choices, because their utility functions differ. There is no single correct utility function to be discovered; there is a designer's choice to be justified.
Distinctions
| Goal-based | Utility-based | |
|---|---|---|
| Value of a state | yes or no | a real number |
| Ranks two successful plans | no | yes |
| Conflicting objectives | cannot trade off | trades off numerically |
| Uncertain outcomes | no preference between chances | maximises expected utility |
| Needs | a goal test | a utility function and probabilities |
| Utility | Performance measure | |
|---|---|---|
| Whose it is | internal to the agent, it uses it to choose | external, the designer judges the agent by it |
| Purpose | to decide | to evaluate |
| When they agree | the agent is well designed |
| Risk averse | Risk neutral | Risk seeking | |
|---|---|---|---|
| Utility of money | bends down | straight | bends up |
| Prefers | the certain amount | indifferent | the gamble |
| Rational | yes | yes | yes |
What it does not mean
Utility is not the performance measure. The performance measure is how the designer judges the agent from outside. The utility function is inside the agent and is what it uses to choose. A well-designed agent has one that leads to a good score on the other, and confusing them makes it impossible to say what went wrong when they diverge.
Maximum expected utility does not mean the best outcome happens. It means the best average over what might happen. An agent that takes the highest expected utility and gets an unlucky result was still rational.
Utility numbers are not probabilities. They are not between 0 and 1 and they do not add to 1. Any scale will do, because only the ordering and the ratios of differences matter.
A higher probability of success is not automatically better. Row 2 of the table wins because the extra certainty is worth more than the extra electricity. Change the electricity cost to 2 per suck and row 1 wins instead. The numbers decide, not the story.
Risk aversion is not a bias to be corrected. It is a shape of utility function, and it is perfectly rational.
Quick revision
- Utility-based agent: uses a utility function from outcomes to real numbers, and where outcomes are uncertain chooses the action of maximum expected utility.
- Expected utility: sum over outcomes of probability times utility. Written out:
EU(a) = sum P(outcome) * U(outcome). - It needs two given things: a utility function and probabilities. Both are design decisions.
- It can do the three things a goal cannot: rank two successes, trade off conflicting objectives, and prefer a better chance.
- On this chapter's world it chooses Suck, Suck, Right, Suck, Suck, which raises the chance of finishing from 0.640 to 0.922 for 0.4 in extra electricity. A goal-based agent could not have considered it.
- Utility is the agent's internal measure; the performance measure is the designer's external one.
- Utility is not linear in the underlying quantity. A curve that bends down is risk averse, a straight line risk neutral, one that bends up risk seeking, and all three are rational.
The Utility-Based Agent
Test yourself
1. Define a utility-based agent and state the principle it acts on. An agent that maps outcomes to real numbers with a utility function and, where outcomes are uncertain, selects the action whose expected utility is greatest. The principle is maximum expected utility.
2. Write the formula for expected utility and explain each part. EU(a) is the sum over possible outcomes of P(outcome given a) times U(outcome). P is how likely the outcome is, U is how desirable it is, and the sum is the average desirability weighted by likelihood.
3. Compute the expected utility of one Suck on a dirty square, given a success probability of 0.8, a clean square worth 5, and a suck costing 0.2. 0.8 times 5 plus 0.2 times 0, minus 0.2, which is 4.0 minus 0.2, that is 3.8.
4. Two plans both clean the house. Why can a goal-based agent not choose between them, and how does a utility-based agent do so? A goal test returns only yes or no, and both plans end in a goal state, so it is indifferent. A utility function assigns a number that includes the cost of each move, so the shorter plan scores higher, 6.600 against 5.600 in this chapter's run.
5. Why does the agent in this chapter choose to suck twice in the same square? Because sucking succeeds only four times in five. Sucking twice raises the probability of finishing the job from 0.640 to 0.922, and the two extra sucks cost 0.4 in total, which is far less than the extra expected value.
6. Distinguish the utility function from the performance measure. The utility function is internal to the agent and is what it uses to choose actions. The performance measure is external, chosen by the designer, and is what the agent is judged by. A good design makes maximising the first produce a high score on the second.
7. A person prefers a certain 500 rupees to a fair coin flip for 1,000. Is this irrational? Explain in terms of utility. No. The expected number of rupees is the same, but the person's utility for money is not linear: the second 500 rupees adds less utility than the first, so the certain amount has the higher expected utility for them. This shape of utility function is called risk aversion.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.