munotes®

What Machine Learning Is

Get access to whole semester resourcesSemester Pass

Chapter Forty-Three

Syllabus topic Module 2, "Introduction of Machine Learning"

Pages 234 to 239 of 591

In one line

Machine learning is writing a program that gets better at a task by looking at examples, instead of being told the rule.

In the wording a student can write in an examination: machine learning is the study of algorithms that improve their performance at a task through experience. The standard definition, due to Tom Mitchell, is that a program learns from experience E with respect to a task T and a performance measure P if its performance at T, as measured by P, improves with E. A model is the thing learned; training is the process of fitting it to data; and generalisation, performing well on data not seen during training, is the only thing that matters.

Where this module sits

The Learning Agent gave the frame and it is worth restating, because Module 2 is one box of it.

ComponentWhat it doesWhich of MU's labels
Performance elementchooses actionsall of Module 1
Criticsays how well it did, against a fixed standardthe three forms of learning are three kinds of critic
Learning elementchanges the performance elementall the rest of Module 2
Problem generatorsuggests informative actionsexploration, in Q-Learning

So machine learning is not a separate subject bolted on. It is the answer to Module 1's last environment property, unknown: if the agent does not know the laws it is operating under, it has to find them out, and finding them out from experience is learning.

Mitchell's definition, and why it is worth the three letters

a program learns from experience E with respect to task T and measure P

if its performance at T, measured by P, improves with E

The value of the definition is that it forces three things to be named before anything is built, and a paper often asks for them on a given example.

ExampleT, the taskP, the measureE, the experience
Spam filterclassify a message as spam or notfraction classified correctly, with a heavy penalty for losing wanted maila mailbox of messages already labelled
Handwriting recognitionread a digit from an imagefraction of digits read correctlyimages of digits with their true values
Predicting markspredict a mark from hours studiedaverage squared error of the predictionpast students' hours and marks
Playing draughtschoose a movefraction of games wongames played against itself

The measure is not optional and it is not obvious. A spam filter measured on plain accuracy will learn to mark everything as not-spam if 97 per cent of mail is legitimate, and score 97 per cent. Evaluating a Model is where this is taken properly, and it is licensed by MU's own Course Outcome 3.

munotes.in234

What Machine Learning Is

How it differs from Module 1

Module 1Module 2
The model isgiven by the designerlearned from data
The question iswhat should I dowhat is the pattern
Correct meansthe answer follows from the modelit generalises to unseen data
Fails whenthe model is wrongthe data is unrepresentative, or the model memorises
Needsa specificationexamples

The single sentence that joins the modules: Module 1 assumes somebody supplied the model, and Module 2 is how the model is obtained from data. Search needs a transition model; logic needs rules; a Bayesian network needs its tables. Learning with Complete Data fills in those tables by counting.

When to use it, and when not to

A paper can ask this, and the honest answer is a rule with three conditions.

Use machine learning when the rule is not known, or is too complicated to write down, or changes over time; and examples are available; and some errors are tolerable.

Do not use it when the rule is known and simple. Nobody should learn a model to decide whether a number is even. Nobody should learn the rules of chess. A program that can be written in ten lines should be written in ten lines, and a learned model in its place is slower, larger, less reliable and impossible to explain.

And do not use it when errors are not tolerable and cannot be checked. A learned model is right most of the time and gives no warning when it is not, which is the substance of MU's Responsible AI row.

The smallest honest learner

Marks against hours studied, for eight students. Fit the straight line that minimises the squared error, then use it.

# The smallest honest learner: fit a straight line to marks against hours studied,
# by least squares, and then predict. Twenty lines, no library.
DATA = [(2, 32), (3, 41), (4, 48), (5, 56), (6, 61), (7, 72), (8, 77), (9, 85)]

def fit(points):
    """Least squares: the line y = a + b*x that minimises the squared error."""
    n = len(points)
    mx = sum(x for x, _ in points) / n
    my = sum(y for _, y in points) / n
    sxy = sum((x - mx) * (y - my) for x, y in points)
    sxx = sum((x - mx) ** 2 for x, _ in points)
    b = sxy / sxx
    a = my - b * mx
    return a, b

a, b = fit(DATA)
print("training data: hours studied against marks out of 100")
for x, y in DATA:
    print("   %d hours -> %d marks" % (x, y))
print()
print("the fitted line: marks = %.3f + %.3f * hours" % (a, b))
print()
print("hours | actual | predicted | error")
total = 0.0
for x, y in DATA:
    pred = a + b * x
    err = y - pred
    total += err * err
    print("  %3d | %6d | %9.2f | %+6.2f" % (x, y, pred, err))
print("  mean squared error on the training data: %.3f" % (total / len(DATA)))
print()
print("and the point of it: a prediction for hours nobody studied")
for x in (1, 10, 12):
    print("   %2d hours -> %.1f marks" % (x, a + b * x))
print()
print("the last one is EXTRAPOLATION beyond the range of the data, and the")
print("model cannot know that marks stop at 100: at 12 hours it predicts %.1f." % (a + b * 12))
munotes.in235

What Machine Learning Is

training data: hours studied against marks out of 100
   2 hours -> 32 marks
   3 hours -> 41 marks
   4 hours -> 48 marks
   5 hours -> 56 marks
   6 hours -> 61 marks
   7 hours -> 72 marks
   8 hours -> 77 marks
   9 hours -> 85 marks

the fitted line: marks = 17.881 + 7.476 * hours

hours | actual | predicted | error
    2 |     32 |     32.83 |  -0.83
    3 |     41 |     40.31 |  +0.69
    4 |     48 |     47.79 |  +0.21
    5 |     56 |     55.26 |  +0.74
    6 |     61 |     62.74 |  -1.74
    7 |     72 |     70.21 |  +1.79
    8 |     77 |     77.69 |  -0.69
    9 |     85 |     85.17 |  -0.17
  mean squared error on the training data: 1.060

and the point of it: a prediction for hours nobody studied
    1 hours -> 25.4 marks
   10 hours -> 92.6 marks
   12 hours -> 107.6 marks

the last one is EXTRAPOLATION beyond the range of the data, and the
model cannot know that marks stop at 100: at 12 hours it predicts 107.6.

Everything Module 2 is about is visible in twenty lines.

The model is the two numbers, 17.881 and 7.476. Not the data: once fitted, the eight rows can be thrown away and the line still predicts.

Nothing was told to the program about studying. It does not know that hours cause marks, or which way round the relation should go. It found the line that fits.

The point is the last block. Nobody studied for 1, 10 or 12 hours, and the model answers anyway. Generalisation is the whole purpose: a model that only repeats its training data is a lookup table and has learned nothing.

And the last row is the honest warning. At 12 hours it predicts 107.6 marks. The model has no idea that a mark cannot exceed 100, because nobody told it and no training example was near 12 hours. Extrapolating beyond the range of the data is where a learned model fails silently, and it does so confidently.

munotes.in236

What Machine Learning Is

The three ingredients of any learning method

Every method in this module is these three choices, and naming them makes the module a single subject rather than a list of algorithms.

IngredientWhat it isIn the example above
The hypothesis spacethe set of models the method is willing to considerall straight lines
The loss functionhow badly a particular model fitssquared error
The optimisationhow the best model in the space is foundthe closed-form least-squares formula

A method is not better than another in general; it makes different choices here. A decision tree's hypothesis space is all trees; a neural network's is all settings of its weights. The Statistical Learning Framework states this properly, and it is worth knowing from the first chapter so the methods can be compared rather than merely collected.

Vocabulary

Defined once, and used unchanged for the rest of the module.

  • Instance or example: one row of data.
  • Feature or attribute: one measured quantity of an instance. Hours studied here.
  • Label or target: the answer being learned, where there is one. Marks here.
  • Training set: the examples the model is fitted to. Test set: examples held back to measure generalisation.
  • Model or hypothesis: what the learning produces.
  • Parameter: a number the model learns, such as 7.476. Hyperparameter: a number the designer chooses before learning, such as the degree of a polynomial.
  • Inference or prediction: using a fitted model on a new instance. Not the same "inference" as Module 1's logical inference, and the clash is unfortunate and standard.

Distinctions

Machine learningA program written by hand
The rule comes fromdataa person
Suitsa rule too complex to write, or unknown, or changinga known, simple, stable rule
Failssilently, on data unlike the training setvisibly, as a bug
Explains itselfonly with extra workyes, it is the code
ParameterHyperparameter
Chosen bythe learning algorithm, from datathe designer, before learning
Examplethe slope 7.476the degree of the polynomial
Tuned usingthe training seta validation set, never the test set
Training errorGeneralisation
Measured onthe data fitted todata never seen
Can be made zerousually, yesno
What mattersnot thisthis

What it does not mean

Machine learning is not the whole of AI. Module 1 contains no learning at all, and a minimax chess program learns nothing.

The model is not the data. It is a summary of it, and the data can be discarded after training.

munotes.in237

What Machine Learning Is

A low training error is not success. A model can fit its training data perfectly and be useless, which is Overfitting and Underfitting.

A learned model does not know when it is out of its depth. At 12 hours the line predicts 107.6 marks with no hesitation.

Learning does not discover causes. The line does not say hours cause marks. A model fitted to marks against shoe size would fit just as willingly.

"Inference" here is not Module 1's inference. Here it means using a fitted model to predict; there it meant deriving what follows from a knowledge base.

Quick revision

  • Machine learning: algorithms that improve at a task with experience. Mitchell: a program learns from experience E with respect to task T and measure P if its performance at T, measured by P, improves with E.
  • It is the learning element of The Learning Agent, and the answer to Module 1's unknown environment.
  • Module 1 assumes the model was supplied; Module 2 is how the model comes from data.
  • Use it when the rule is unknown, too complex, or changing; examples exist; and some errors are tolerable. Do not use it for a rule you can write in ten lines.
  • The worked line: marks = 17.881 + 7.476 * hours. The model is the two numbers; the data can be discarded.
  • Generalisation is the point. At 12 hours it predicts 107.6 marks, because nobody told it a mark stops at 100 and no example was near 12 hours. Extrapolation is where a learned model fails silently.
  • Three ingredients of any method: the hypothesis space, the loss function, the optimisation.
  • Vocabulary: instance, feature, label, training and test set, model, parameter (learned) against hyperparameter (chosen), prediction.

Test yourself

1. Give Mitchell's definition of learning and apply it to a spam filter. A program learns from experience E with respect to a task T and a performance measure P if its performance at T, as measured by P, improves with E. For a spam filter: the task is classifying a message as spam or not; the measure is the fraction classified correctly with a heavy penalty for losing wanted mail; and the experience is a mailbox of messages already labelled.

2. How does Module 2 relate to Module 1? It is the learning element of the learning agent introduced in Module 1. Module 1 assumes a model was supplied, whether a transition model, a set of rules or a set of probability tables, and Module 2 is how such a model is obtained from data. It is the response to an environment being unknown.

munotes.in238

What Machine Learning Is

3. When should machine learning not be used? When the rule is known and simple enough to write directly, since a learned model is then slower, larger, less reliable and harder to explain. And when errors cannot be tolerated and cannot be checked, since a learned model fails silently.

4. In the worked example, what exactly is the model, and what can be discarded? The model is the two numbers 17.881 and 7.476, the intercept and the slope. Once they are fitted the eight training rows can be discarded and the model still predicts.

5. The model predicts 107.6 marks for 12 hours of study. What has gone wrong, and what is the general lesson? Nothing has gone wrong inside the model; it has extrapolated beyond the range of its data. It was never told that a mark cannot exceed 100 and no training example was near 12 hours. The lesson is that a learned model fails silently and confidently outside the region its data covered.

6. Name the three ingredients of any learning method, with the example's values. The hypothesis space, here all straight lines; the loss function, here squared error; and the optimisation procedure, here the closed-form least-squares formula.

7. Distinguish a parameter from a hyperparameter, and say which data is used to set each. A parameter is a number the learning algorithm fits from the training data, such as the slope. A hyperparameter is chosen by the designer before learning, such as the degree of a polynomial, and is tuned on a validation set held out from the training data, never on the test set.

munotes.in239

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!