Bayes Theorem
Chapter Thirty-Nine
Syllabus topic Module 1, "Bayes theorem"
Pages 209 to 214 of 591
In one line
Bayes theorem turns the probability of a symptom given a disease, which is what a laboratory can measure, into the probability of the disease given the symptom, which is what a patient wants to know.
In the wording a student can write in an examination: Bayes theorem states that for propositions a and b with P(b) greater than 0,
P(a | b) = P(b | a) * P(a) / P(b)
Here P(a) is the prior, P(b | a) the likelihood, P(b) the evidence or normalising constant, and P(a | b) the posterior. The theorem is used when the likelihood is easy to obtain and the posterior is what is wanted, which is the usual situation in diagnosis.
The derivation, in two lines
It is worth being able to produce, because a paper can ask for it and it takes twenty seconds.
The product rule from Why an Agent Needs Probability gives the same quantity two ways:
P(a and b) = P(a | b) * P(b)
P(a and b) = P(b | a) * P(a)
So the right-hand sides are equal. Divide both by P(b):
P(a | b) P(b) = P(b | a) P(a)
P(a | b) = P(b | a) * P(a) / P(b)
That is the whole proof. Bayes theorem is not a new assumption; it is the product rule written twice and rearranged, which is why it cannot fail and why anyone who rejects it has rejected the axioms.
The four names, and why the theorem is worth having
| Term | Name | In a diagnosis |
|---|---|---|
P(a) | the prior | how common the disease is, the base rate |
P(b given a) | the likelihood | how often the test is positive in people who have it |
P(b) | the evidence, or normalising constant | how often the test is positive at all |
P(a given b) | the posterior | the answer: does this patient have it |
The point of the theorem is the direction. A laboratory can measure P(positive | disease) by testing people known to have the disease. Nobody can directly measure P(disease | positive), because that depends on how common the disease is in the population being tested. Bayes theorem is the bridge, and the base rate is the toll.
The denominator is usually computed by cases rather than looked up:
P(b) = P(b | a) P(a) + P(b | not a) P(not a)
This is the law of total probability, and it says: the test comes out positive either because the patient has the disease and it was detected, or because they do not and it misfired. Adding the two is the only way to get P(b) from what a laboratory can measure.
Bayes Theorem
Three worked problems
# Bayes theorem, worked on three problems. The BASE RATE TRAP is the third and it
# is the one every examiner sets.
def bayes(prior, sens, spec, name, positive="positive"):
"""prior = P(cause); sens = P(test+ | cause); spec = P(test- | not cause)."""
fp = 1 - spec # false positive rate
joint_yes = prior * sens # has it AND tests positive
joint_no = (1 - prior) * fp # has it not AND tests positive
evidence = joint_yes + joint_no
post = joint_yes / evidence
print(name)
print(" %-34s = %.4f" % ("P(cause)", prior))
print(" %-34s = %.3f the sensitivity" % ("P(%s | cause)" % positive, sens))
print(" %-34s = %.3f 1 - specificity" % ("P(%s | no cause)" % positive, fp))
print(" %-34s = %.6f" % ("P(cause) * P(+ | cause)", joint_yes))
print(" %-34s = %.6f" % ("P(no cause) * P(+ | no cause)", joint_no))
print(" %-34s = %.6f" % ("P(%s), the total" % positive, evidence))
print(" %-34s = %.6f %.2f%%" % ("P(cause | %s)" % positive, post, 100 * post))
print()
return post
bayes(0.001, 0.99, 0.99,
"A DISEASE TEST. 1 person in 1000 has it. The test is 99% accurate both ways.")
bayes(0.02, 0.95, 0.90,
"A MACHINE FAULT. 2% of parts are faulty. The scanner catches 95% and\n"
"wrongly flags 10% of good parts.", "flagged")
bayes(0.30, 0.98, 0.95,
"A SPAM FILTER. 30% of mail is spam. The filter catches 98% and wrongly\n"
"flags 5% of good mail.", "marked spam")
print("THE BASE RATE TRAP, in one sentence:")
print(" the first test is 99% accurate and a positive result still leaves the")
print(" patient more likely NOT to have the disease than to have it, because")
print(" the healthy are a thousand times more numerous. Out of 100,000 people:")
n = 100000
ill = int(n * 0.001)
well = n - ill
tp = round(ill * 0.99)
fp = round(well * 0.01)
print(" %6d have it, of whom %4d test positive" % (ill, tp))
print(" %6d do not, of whom %4d test positive" % (well, fp))
print(" so %d positives in all, and only %d of them are ill: %.2f%%"
% (tp + fp, tp, 100.0 * tp / (tp + fp)))A DISEASE TEST. 1 person in 1000 has it. The test is 99% accurate both ways.
P(cause) = 0.0010
P(positive | cause) = 0.990 the sensitivity
P(positive | no cause) = 0.010 1 - specificity
P(cause) * P(+ | cause) = 0.000990
P(no cause) * P(+ | no cause) = 0.009990
P(positive), the total = 0.010980
P(cause | positive) = 0.090164 9.02%
A MACHINE FAULT. 2% of parts are faulty. The scanner catches 95% and
wrongly flags 10% of good parts.
P(cause) = 0.0200
P(flagged | cause) = 0.950 the sensitivity
P(flagged | no cause) = 0.100 1 - specificity
P(cause) * P(+ | cause) = 0.019000
P(no cause) * P(+ | no cause) = 0.098000
P(flagged), the total = 0.117000
P(cause | flagged) = 0.162393 16.24%
A SPAM FILTER. 30% of mail is spam. The filter catches 98% and wrongly
flags 5% of good mail.
P(cause) = 0.3000
P(marked spam | cause) = 0.980 the sensitivity
P(marked spam | no cause) = 0.050 1 - specificity
P(cause) * P(+ | cause) = 0.294000
P(no cause) * P(+ | no cause) = 0.035000
P(marked spam), the total = 0.329000
P(cause | marked spam) = 0.893617 89.36%
THE BASE RATE TRAP, in one sentence:
the first test is 99% accurate and a positive result still leaves the
patient more likely NOT to have the disease than to have it, because
the healthy are a thousand times more numerous. Out of 100,000 people:
100 have it, of whom 99 test positive
99900 do not, of whom 999 test positive
so 1098 positives in all, and only 99 of them are ill: 9.02%Bayes Theorem
The base rate trap, which is the examinable part
Read the first block. The test is 99 per cent accurate in both directions, the patient has tested positive, and the probability that they have the disease is 9.02 per cent. They are about ten times more likely to be well than ill.
This is not a paradox and the test is not bad. The arithmetic at the foot of the output is the whole explanation, and it is the form to use when explaining it to anybody:
Out of 100,000 people, 100 have the disease and 99,900 do not. Of the 100, the test correctly flags 99. Of the 99,900, the test wrongly flags one per cent, which is 999. So there are 1,098 positive results and only 99 of them are true. 99 out of 1,098 is 9.02 per cent.
The false positives outnumber the true positives eleven to one, because the healthy group is a thousand times larger. A one per cent error rate on a group a thousand times bigger produces ten times more errors than the whole of the small group.
The name of the mistake. Ignoring the prior and answering 99 per cent is called the base rate fallacy. It is the commonest error on this topic, it is made by doctors and lawyers as well as students, and the way to avoid it is to compute the denominator by cases as the law of total probability requires.
And note the comparison across the three problems. The same structure with a base rate of 0.001 gives 9 per cent, with 0.02 gives 16 per cent, and with 0.30 gives 89 per cent. The prior does most of the work, and the third case is why spam filters are usable while rare-disease screening is hard.
Bayes Theorem
Two more things the theorem gives
The odds form, which is quicker for a sequence of evidence and is worth knowing.
posterior odds = prior odds * likelihood ratio
where the likelihood ratio is P(b | a) / P(b | not a)
For the first problem: prior odds are 1 to 999, the likelihood ratio is 0.99 divided by 0.01, that is 99, so the posterior odds are 99 to 999, which is 99 out of 1,098. The same answer with no division, and the normalising constant disappears entirely. This is also why a second independent positive test helps so much: multiply by 99 again.
Sequential updating. The posterior after one piece of evidence is the prior for the next. Beliefs are revised piece by piece, and the order does not matter. It requires the pieces of evidence to be conditionally independent given the cause, which is the next chapter and is the assumption The Naive Bayes Classifier is named for.
Where the priors come from
An honest section, because a paper can ask and because the answer is sometimes uncomfortable.
From data, when there is data: the prevalence of a disease, the proportion of faulty parts, the fraction of mail that is spam. This is the usual case and the numbers are then defensible.
From an expert's judgement, when there is not. This is legitimate under the Bayesian reading in Why an Agent Needs Probability and it is also where a system can be criticised.
From a deliberately uninformative choice, such as assigning equal probability to each alternative, when nothing is known. That is itself an assumption and not a neutral position: "equally likely" is a claim.
And the practical reassurance: with enough evidence the posterior becomes insensitive to the prior, because the likelihood ratios accumulate. With little evidence the prior dominates, which is exactly the rare-disease case.
Distinctions
| Prior | Likelihood | Posterior | |
|---|---|---|---|
| Symbol | P(a) | P(b given a) | P(a given b) |
| Measured by | how common the cause is | testing known cases | Bayes theorem |
| In the first problem | 0.001 | 0.99 | 0.0902 |
P(positive given disease) | P(disease given positive) | |
|---|---|---|
| Is | the sensitivity, a property of the test | the answer the patient wants |
| Depends on the base rate | no | yes |
| In the first problem | 0.99 | 0.0902 |
| Confusing them is | the base rate fallacy |
| Sensitivity | Specificity | |
|---|---|---|
| Is | P(positive given disease) | P(negative given no disease) |
| Measures | catching the ill | not alarming the well |
| In the first problem | 0.99 | 0.99 |
Bayes Theorem
What it does not mean
Bayes theorem is not an extra assumption. It is the product rule written two ways and rearranged, so it follows from the axioms alone.
P(a | b) is not P(b | a). Swapping them is the base rate fallacy, and in the first problem the two differ by a factor of eleven.
A 99 per cent accurate test does not give a 99 per cent answer. It gives 9 per cent when the disease affects one person in a thousand.
The result is not evidence that the test is useless. It multiplied the belief from 0.001 to 0.090, a factor of 90. For a rare disease that is exactly why screening is followed by a second, different test.
The denominator is not usually looked up. It is computed by cases, using the law of total probability.
A prior is not optional. Leaving it out is not neutrality; it is assuming a base rate of one half, which in the first problem is wrong by a factor of 500.
Quick revision
- Bayes theorem:
P(a | b) = P(b | a) * P(a) / P(b). Derived in two lines from the product rule written both ways. - Names:
P(a)prior,P(b | a)likelihood,P(b)evidence or normalising constant,P(a | b)posterior. - Law of total probability for the denominator:
P(b) = P(b|a)P(a) + P(b|not a)P(not a). - The base rate trap. A 99 per cent accurate test, a positive result, a disease affecting 1 in 1,000: the answer is 9.02 per cent. Out of 100,000 people, 99 true positives against 999 false ones.
- The base rate fallacy is answering 99 per cent, that is confusing
P(disease | positive)withP(positive | disease). - The same structure at base rates 0.001, 0.02 and 0.30 gives 9 per cent, 16 per cent and 89 per cent. The prior does most of the work.
- Odds form: posterior odds equal prior odds times the likelihood ratio
P(b|a) / P(b|not a). 1 to 999 times 99 gives 99 to 999. No division, and no normalising constant. - Sequential updating: today's posterior is tomorrow's prior, provided the pieces of evidence are conditionally independent given the cause.
- Priors come from data, from expert judgement, or from a deliberately uninformative choice, which is itself an assumption.
Test yourself
1. State Bayes theorem and derive it. P(a | b) equals P(b | a) times P(a), divided by P(b). The product rule gives P(a and b) as both P(a|b)P(b) and P(b|a)P(a); equating those and dividing by P(b) gives the theorem.
2. Name the four quantities in the theorem. P(a) is the prior, P(b | a) the likelihood, P(b) the evidence or normalising constant, and P(a | b) the posterior.
Bayes Theorem
3. A disease affects 1 person in 1,000. A test detects 99 per cent of cases and wrongly flags 1 per cent of healthy people. A patient tests positive. What is the probability they have the disease? The numerator is 0.001 times 0.99, which is 0.00099. The other term is 0.999 times 0.01, which is 0.00999. The total is 0.01098, so the posterior is 0.00099 divided by 0.01098, that is 0.0902, about 9 per cent.
4. Explain that answer by counting people. In 100,000 people, 100 have the disease and 99 of them test positive. Of the 99,900 who do not, one per cent, that is 999, test positive. So 1,098 people test positive and only 99 of them are ill, which is 9.02 per cent. The healthy group is a thousand times larger, so its small error rate produces ten times more positives than the entire ill group.
5. What is the base rate fallacy? Answering the question P(disease | positive) with the value of P(positive | disease), that is ignoring how common the disease is. In the example above it gives 99 per cent instead of 9 per cent.
6. Give the odds form of the theorem and apply it to the disease problem. Posterior odds equal prior odds times the likelihood ratio, the likelihood ratio being P(b|a) divided by P(b | not a). Here the prior odds are 1 to 999 and the ratio is 0.99 over 0.01, that is 99, so the posterior odds are 99 to 999, giving 99 out of 1,098.
7. The same test structure at base rates of 0.001, 0.02 and 0.30 gives posteriors of 9, 16 and 89 per cent. What does that show? That the prior does most of the work. A test of fixed quality is decisive when the condition is common and nearly uninformative when it is rare, which is why spam filtering works well and why screening for a rare disease needs a second, independent test.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.