What Machine Learning Is
Chapter Forty-Three
Syllabus topic Module 2, "Introduction of Machine Learning"
Pages 234 to 239 of 591
In one line
Machine learning is writing a program that gets better at a task by looking at examples, instead of being told the rule.
In the wording a student can write in an examination: machine learning is the study of algorithms that improve their performance at a task through experience. The standard definition, due to Tom Mitchell, is that a program learns from experience E with respect to a task T and a performance measure P if its performance at T, as measured by P, improves with E. A model is the thing learned; training is the process of fitting it to data; and generalisation, performing well on data not seen during training, is the only thing that matters.
Where this module sits
The Learning Agent gave the frame and it is worth restating, because Module 2 is one box of it.
| Component | What it does | Which of MU's labels |
|---|---|---|
| Performance element | chooses actions | all of Module 1 |
| Critic | says how well it did, against a fixed standard | the three forms of learning are three kinds of critic |
| Learning element | changes the performance element | all the rest of Module 2 |
| Problem generator | suggests informative actions | exploration, in Q-Learning |
So machine learning is not a separate subject bolted on. It is the answer to Module 1's last environment property, unknown: if the agent does not know the laws it is operating under, it has to find them out, and finding them out from experience is learning.
Mitchell's definition, and why it is worth the three letters
a program learns from experience E with respect to task T and measure P
if its performance at T, measured by P, improves with E
The value of the definition is that it forces three things to be named before anything is built, and a paper often asks for them on a given example.
| Example | T, the task | P, the measure | E, the experience |
|---|---|---|---|
| Spam filter | classify a message as spam or not | fraction classified correctly, with a heavy penalty for losing wanted mail | a mailbox of messages already labelled |
| Handwriting recognition | read a digit from an image | fraction of digits read correctly | images of digits with their true values |
| Predicting marks | predict a mark from hours studied | average squared error of the prediction | past students' hours and marks |
| Playing draughts | choose a move | fraction of games won | games played against itself |
The measure is not optional and it is not obvious. A spam filter measured on plain accuracy will learn to mark everything as not-spam if 97 per cent of mail is legitimate, and score 97 per cent. Evaluating a Model is where this is taken properly, and it is licensed by MU's own Course Outcome 3.
What Machine Learning Is
How it differs from Module 1
| Module 1 | Module 2 | |
|---|---|---|
| The model is | given by the designer | learned from data |
| The question is | what should I do | what is the pattern |
| Correct means | the answer follows from the model | it generalises to unseen data |
| Fails when | the model is wrong | the data is unrepresentative, or the model memorises |
| Needs | a specification | examples |
The single sentence that joins the modules: Module 1 assumes somebody supplied the model, and Module 2 is how the model is obtained from data. Search needs a transition model; logic needs rules; a Bayesian network needs its tables. Learning with Complete Data fills in those tables by counting.
When to use it, and when not to
A paper can ask this, and the honest answer is a rule with three conditions.
Use machine learning when the rule is not known, or is too complicated to write down, or changes over time; and examples are available; and some errors are tolerable.
Do not use it when the rule is known and simple. Nobody should learn a model to decide whether a number is even. Nobody should learn the rules of chess. A program that can be written in ten lines should be written in ten lines, and a learned model in its place is slower, larger, less reliable and impossible to explain.
And do not use it when errors are not tolerable and cannot be checked. A learned model is right most of the time and gives no warning when it is not, which is the substance of MU's Responsible AI row.
The smallest honest learner
Marks against hours studied, for eight students. Fit the straight line that minimises the squared error, then use it.
# The smallest honest learner: fit a straight line to marks against hours studied,
# by least squares, and then predict. Twenty lines, no library.
DATA = [(2, 32), (3, 41), (4, 48), (5, 56), (6, 61), (7, 72), (8, 77), (9, 85)]
def fit(points):
"""Least squares: the line y = a + b*x that minimises the squared error."""
n = len(points)
mx = sum(x for x, _ in points) / n
my = sum(y for _, y in points) / n
sxy = sum((x - mx) * (y - my) for x, y in points)
sxx = sum((x - mx) ** 2 for x, _ in points)
b = sxy / sxx
a = my - b * mx
return a, b
a, b = fit(DATA)
print("training data: hours studied against marks out of 100")
for x, y in DATA:
print(" %d hours -> %d marks" % (x, y))
print()
print("the fitted line: marks = %.3f + %.3f * hours" % (a, b))
print()
print("hours | actual | predicted | error")
total = 0.0
for x, y in DATA:
pred = a + b * x
err = y - pred
total += err * err
print(" %3d | %6d | %9.2f | %+6.2f" % (x, y, pred, err))
print(" mean squared error on the training data: %.3f" % (total / len(DATA)))
print()
print("and the point of it: a prediction for hours nobody studied")
for x in (1, 10, 12):
print(" %2d hours -> %.1f marks" % (x, a + b * x))
print()
print("the last one is EXTRAPOLATION beyond the range of the data, and the")
print("model cannot know that marks stop at 100: at 12 hours it predicts %.1f." % (a + b * 12))What Machine Learning Is
training data: hours studied against marks out of 100
2 hours -> 32 marks
3 hours -> 41 marks
4 hours -> 48 marks
5 hours -> 56 marks
6 hours -> 61 marks
7 hours -> 72 marks
8 hours -> 77 marks
9 hours -> 85 marks
the fitted line: marks = 17.881 + 7.476 * hours
hours | actual | predicted | error
2 | 32 | 32.83 | -0.83
3 | 41 | 40.31 | +0.69
4 | 48 | 47.79 | +0.21
5 | 56 | 55.26 | +0.74
6 | 61 | 62.74 | -1.74
7 | 72 | 70.21 | +1.79
8 | 77 | 77.69 | -0.69
9 | 85 | 85.17 | -0.17
mean squared error on the training data: 1.060
and the point of it: a prediction for hours nobody studied
1 hours -> 25.4 marks
10 hours -> 92.6 marks
12 hours -> 107.6 marks
the last one is EXTRAPOLATION beyond the range of the data, and the
model cannot know that marks stop at 100: at 12 hours it predicts 107.6.Everything Module 2 is about is visible in twenty lines.
The model is the two numbers, 17.881 and 7.476. Not the data: once fitted, the eight rows can be thrown away and the line still predicts.
Nothing was told to the program about studying. It does not know that hours cause marks, or which way round the relation should go. It found the line that fits.
The point is the last block. Nobody studied for 1, 10 or 12 hours, and the model answers anyway. Generalisation is the whole purpose: a model that only repeats its training data is a lookup table and has learned nothing.
And the last row is the honest warning. At 12 hours it predicts 107.6 marks. The model has no idea that a mark cannot exceed 100, because nobody told it and no training example was near 12 hours. Extrapolating beyond the range of the data is where a learned model fails silently, and it does so confidently.
What Machine Learning Is
The three ingredients of any learning method
Every method in this module is these three choices, and naming them makes the module a single subject rather than a list of algorithms.
| Ingredient | What it is | In the example above |
|---|---|---|
| The hypothesis space | the set of models the method is willing to consider | all straight lines |
| The loss function | how badly a particular model fits | squared error |
| The optimisation | how the best model in the space is found | the closed-form least-squares formula |
A method is not better than another in general; it makes different choices here. A decision tree's hypothesis space is all trees; a neural network's is all settings of its weights. The Statistical Learning Framework states this properly, and it is worth knowing from the first chapter so the methods can be compared rather than merely collected.
Vocabulary
Defined once, and used unchanged for the rest of the module.
- Instance or example: one row of data.
- Feature or attribute: one measured quantity of an instance. Hours studied here.
- Label or target: the answer being learned, where there is one. Marks here.
- Training set: the examples the model is fitted to. Test set: examples held back to measure generalisation.
- Model or hypothesis: what the learning produces.
- Parameter: a number the model learns, such as 7.476. Hyperparameter: a number the designer chooses before learning, such as the degree of a polynomial.
- Inference or prediction: using a fitted model on a new instance. Not the same "inference" as Module 1's logical inference, and the clash is unfortunate and standard.
Distinctions
| Machine learning | A program written by hand | |
|---|---|---|
| The rule comes from | data | a person |
| Suits | a rule too complex to write, or unknown, or changing | a known, simple, stable rule |
| Fails | silently, on data unlike the training set | visibly, as a bug |
| Explains itself | only with extra work | yes, it is the code |
| Parameter | Hyperparameter | |
|---|---|---|
| Chosen by | the learning algorithm, from data | the designer, before learning |
| Example | the slope 7.476 | the degree of the polynomial |
| Tuned using | the training set | a validation set, never the test set |
| Training error | Generalisation | |
|---|---|---|
| Measured on | the data fitted to | data never seen |
| Can be made zero | usually, yes | no |
| What matters | not this | this |
What it does not mean
Machine learning is not the whole of AI. Module 1 contains no learning at all, and a minimax chess program learns nothing.
The model is not the data. It is a summary of it, and the data can be discarded after training.
What Machine Learning Is
A low training error is not success. A model can fit its training data perfectly and be useless, which is Overfitting and Underfitting.
A learned model does not know when it is out of its depth. At 12 hours the line predicts 107.6 marks with no hesitation.
Learning does not discover causes. The line does not say hours cause marks. A model fitted to marks against shoe size would fit just as willingly.
"Inference" here is not Module 1's inference. Here it means using a fitted model to predict; there it meant deriving what follows from a knowledge base.
Quick revision
- Machine learning: algorithms that improve at a task with experience. Mitchell: a program learns from experience
Ewith respect to taskTand measurePif its performance atT, measured byP, improves withE. - It is the learning element of
The Learning Agent, and the answer to Module 1's unknown environment. - Module 1 assumes the model was supplied; Module 2 is how the model comes from data.
- Use it when the rule is unknown, too complex, or changing; examples exist; and some errors are tolerable. Do not use it for a rule you can write in ten lines.
- The worked line:
marks = 17.881 + 7.476 * hours. The model is the two numbers; the data can be discarded. - Generalisation is the point. At 12 hours it predicts 107.6 marks, because nobody told it a mark stops at 100 and no example was near 12 hours. Extrapolation is where a learned model fails silently.
- Three ingredients of any method: the hypothesis space, the loss function, the optimisation.
- Vocabulary: instance, feature, label, training and test set, model, parameter (learned) against hyperparameter (chosen), prediction.
Test yourself
1. Give Mitchell's definition of learning and apply it to a spam filter. A program learns from experience E with respect to a task T and a performance measure P if its performance at T, as measured by P, improves with E. For a spam filter: the task is classifying a message as spam or not; the measure is the fraction classified correctly with a heavy penalty for losing wanted mail; and the experience is a mailbox of messages already labelled.
2. How does Module 2 relate to Module 1? It is the learning element of the learning agent introduced in Module 1. Module 1 assumes a model was supplied, whether a transition model, a set of rules or a set of probability tables, and Module 2 is how such a model is obtained from data. It is the response to an environment being unknown.
What Machine Learning Is
3. When should machine learning not be used? When the rule is known and simple enough to write directly, since a learned model is then slower, larger, less reliable and harder to explain. And when errors cannot be tolerated and cannot be checked, since a learned model fails silently.
4. In the worked example, what exactly is the model, and what can be discarded? The model is the two numbers 17.881 and 7.476, the intercept and the slope. Once they are fitted the eight training rows can be discarded and the model still predicts.
5. The model predicts 107.6 marks for 12 hours of study. What has gone wrong, and what is the general lesson? Nothing has gone wrong inside the model; it has extrapolated beyond the range of its data. It was never told that a mark cannot exceed 100 and no training example was near 12 hours. The lesson is that a learned model fails silently and confidently outside the region its data covered.
6. Name the three ingredients of any learning method, with the example's values. The hypothesis space, here all straight lines; the loss function, here squared error; and the optimisation procedure, here the closed-form least-squares formula.
7. Distinguish a parameter from a hyperparameter, and say which data is used to set each. A parameter is a number the learning algorithm fits from the training data, such as the slope. A hyperparameter is chosen by the designer before learning, such as the degree of a polynomial, and is tuned on a validation set held out from the training data, never on the test set.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.