munotes®

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

Get access to whole semester resourcesSemester Pass

Chapter Three

Syllabus topic Module 1, the evaluation every learning practical asks for: "Evaluate the accuracy and effectiveness of the decision tree on test data.", "Evaluate the performance of the trained network on test data.", "Evaluate the performance of the SVM model on test data and analyze the results.", "Train the ensemble model on a given dataset and evaluate its performance.", "Evaluate the accuracy of the model on test data and analyze the results.", "Evaluate the accuracy or error of the predictions and analyze the results.", and MU's own "performance evaluation".

Pages 11 to 19 of 206

Aim

To learn the one thing six of the ten practicals in this module ask for and none of them explains: how a model is measured. The training set and the test set, the confusion matrix, accuracy, precision, recall and F1, the error of a prediction that is a number, and the baseline a model has to beat.

Why this chapter exists

Read MU's Module 1 and count the sentences that say "evaluate". There are six of them. "Evaluate the accuracy and effectiveness of the decision tree on test data." "Evaluate the performance of the trained network on test data." And so on through the SVM, the ensemble, Naive Bayes and K-NN.

Not one of those sentences says what accuracy is, what test data is, or how the data becomes test data. Six practicals ask for a measurement the syllabus never defines, and a student who has not been taught it either prints a number with no name or, much worse, prints a number measured on the wrong data.

That is what this chapter is for. Learn it once here, use it in six chapters.

The dataset this module uses

Most of this module runs on one small dataset of thirty students. Two columns are what we know about a student, the features, and one is what we are trying to predict, the label.

name,attendance,practice,result
Aarav,50,22,Fail
Isha,48,16,Fail
Rohan,88,38,Pass
Sanya,97,39,Pass
Vikram,90,30,Fail
Meera,35,41,Fail
Farhan,45,9,Fail
Nikita,71,8,Fail
Omkar,92,2,Fail
Pooja,97,45,Pass
Rahul,75,15,Fail
Sneha,85,18,Pass
Tejas,79,24,Pass
Urmila,83,34,Pass
Varun,44,23,Fail
Yash,46,37,Pass
Zoya,72,20,Fail
Amit,44,9,Pass
Bhavna,74,3,Fail
Chirag,82,25,Pass
Deepa,94,29,Pass
Eshan,46,27,Fail
Gauri,98,9,Fail
Harsh,89,34,Pass
Ira,97,27,Pass
Jatin,68,28,Pass
Kavya,96,34,Pass
Lalit,38,38,Fail
Manav,63,10,Fail
Neha,41,35,Fail

attendance is a percentage. practice is the number of laboratory hours the student put in. result is Pass or Fail. Thirty rows is small on purpose: every intermediate number a program computes can be checked by hand, which is exactly what an examiner may ask you to do.

import csv

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))

data = [([int(r["attendance"]), int(r["practice"])], r["result"]) for r in rows]

passes = sum(1 for _, y in data if y == "Pass")
print("rows     :", len(data))
print("features :", 2, "(attendance, practice)")
print("Pass     :", passes)
print("Fail     :", len(data) - passes)
print("first row:", data[0])
rows     : 30
features : 2 (attendance, practice)
Pass     : 14
Fail     : 16
first row: ([50, 22], 'Fail')

Fourteen Pass and sixteen Fail. Keep those two numbers in mind: they are what the baseline below is built from.

The training set and the test set, and why the split is the whole idea

A model learns from data and is then asked about data. If you ask it about the same rows it learned from, you learn nothing at all, because a model is allowed to remember. A model that memorises all thirty rows answers all thirty correctly and knows nothing.

munotes.in11

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

So the rows are split in two:

  • the training set, which the model learns from, and
  • the test set, which it has never seen, and which is the only honest place to measure it.

The usual split is 70 or 80 per cent for training and the rest for testing. Two rules matter.

Shuffle before you split. Take the last six rows of our file as the test set and you get five Fail and one Pass, because the file happens to end that way. A test set that does not look like the data tells you nothing about the data.

Seed the shuffle. Otherwise your journal and the examiner's re-run disagree, for the reason the previous chapter gives.

import csv
import random

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
data = [([int(r["attendance"]), int(r["practice"])], r["result"]) for r in rows]

def split(data, seed, frac=0.7):
    order = list(range(len(data)))
    random.seed(seed)
    random.shuffle(order)
    cut = int(frac * len(data))
    return [data[i] for i in order[:cut]], [data[i] for i in order[cut:]]

def balance(rows):
    p = sum(1 for _, y in rows if y == "Pass")
    return "%d rows, %d Pass, %d Fail" % (len(rows), p, len(rows) - p)

train, test = split(data, 1)
print("training :", balance(train))
print("test     :", balance(test))
training : 21 rows, 9 Pass, 12 Fail
test     : 9 rows, 5 Pass, 4 Fail

Twenty-one rows to learn from, nine to be tested on, chosen by a seeded shuffle. That is the split every later chapter in this module uses.

How much the split itself moves the answer

The whole file is 14 Pass to 16 Fail, which is 47 per cent Pass. Seed 1 put 5 Pass into the nine test rows, which is 56 per cent. Change the seed and the test set changes with it:

import csv
import random

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
data = [([int(r["attendance"]), int(r["practice"])], r["result"]) for r in rows]

def split(data, seed, frac=0.7):
    order = list(range(len(data)))
    random.seed(seed)
    random.shuffle(order)
    cut = int(frac * len(data))
    return [data[i] for i in order[:cut]], [data[i] for i in order[cut:]]

print("%-6s %12s %12s" % ("seed", "test Pass", "test Fail"))
for seed in range(1, 9):
    _, test = split(data, seed)
    p = sum(1 for _, y in test if y == "Pass")
    print("%-6d %12d %12d" % (seed, p, len(test) - p))
seed      test Pass    test Fail
1                 5            4
2                 4            5
3                 5            4
4                 6            3
5                 5            4
6                 5            4
7                 6            3
8                 4            5

Four to six Pass out of nine, purely from the seed. On nine test rows one row is 0.111 of the accuracy, so two models that differ by 0.05 on a test set this small have not been distinguished at all, and a write-up that ranks them is reading noise. Say the size of your test set whenever you quote a score.

munotes.in12

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

The fix for the balance, which scikit-learn calls stratify, keeps the proportion of each class the same in both halves. The fix for the noise is a bigger test set, or cross validation: split the data five ways, train five times, test on the part that was held out each time, and average the five scores. Cross validation is out of this syllabus, but the sentence "on nine test rows this difference is not measurable" is always worth writing.

The baseline: the number your model has to beat

Before measuring any model, measure the stupidest possible one: the model that ignores its input and answers whichever class was commonest in the training data.

import csv
import random
from collections import Counter

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
data = [([int(r["attendance"]), int(r["practice"])], r["result"]) for r in rows]

def split(data, seed, frac=0.7):
    order = list(range(len(data)))
    random.seed(seed)
    random.shuffle(order)
    cut = int(frac * len(data))
    return [data[i] for i in order[:cut]], [data[i] for i in order[cut:]]

train, test = split(data, 1)
commonest = Counter(y for _, y in train).most_common(1)[0][0]
right = sum(1 for _, y in test if y == commonest)

print("commonest class in training :", commonest)
print("it answers                  :", commonest, "every time")
print("baseline on the test set    : %d of %d right, accuracy %.3f"
      % (right, len(test), right / len(test)))
commonest class in training : Fail
it answers                  : Fail every time
baseline on the test set    : 4 of 9 right, accuracy 0.444

Read that carefully, because it is not the answer most students expect. The commonest class in the training set is Fail, 12 of 21. But the test set has 5 Pass and 4 Fail, so answering Fail every time is right only 4 times out of 9: an accuracy of 0.444, which is worse than a coin.

Two things follow, and both belong in a write-up.

The baseline is decided on training data and measured on test data, like everything else. You are not allowed to look at the test labels to choose what your stupid model says, because you are not allowed to look at them at all.

On a small test set a baseline can come out below 0.5. That is not a bug. It is what a nine-row test set does, and it is another reason to say how big the test set was.

Compute the baseline anyway, every time. It is four lines, and it is the difference between a number and a result.

The confusion matrix

Accuracy on its own hides which mistakes a model makes, and the two mistakes are not equally serious. Telling a student who will fail that they will pass is not the same error as telling a student who will pass that they will fail.

munotes.in13

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

Pick one class as positive. Here, Pass. Then every prediction falls into one of four boxes:

Predicted PassPredicted Fail
Actually PassTrue Positive (TP)False Negative (FN)
Actually FailFalse Positive (FP)True Negative (TN)

Read the names from right to left: a False Positive is a prediction of Positive that is false.

def confusion(actual, predicted, positive):
    tp = fp = fn = tn = 0
    for a, p in zip(actual, predicted):
        if p == positive and a == positive:
            tp += 1
        elif p == positive and a != positive:
            fp += 1
        elif p != positive and a == positive:
            fn += 1
        else:
            tn += 1
    return tp, fp, fn, tn

actual = ["Pass", "Fail", "Pass", "Fail", "Fail", "Pass", "Fail", "Pass", "Fail"]
guess = ["Pass", "Fail", "Fail", "Fail", "Pass", "Pass", "Fail", "Pass", "Fail"]

tp, fp, fn, tn = confusion(actual, guess, "Pass")
print("             predicted Pass  predicted Fail")
print("actual Pass  %14d  %14d" % (tp, fn))
print("actual Fail  %14d  %14d" % (fp, tn))
print()
print("TP %d  FP %d  FN %d  TN %d  total %d" % (tp, fp, fn, tn, tp + fp + fn + tn))
             predicted Pass  predicted Fail
actual Pass               3               1
actual Fail               1               4

TP 3  FP 1  FN 1  TN 4  total 9

The four numbers add up to the number of test rows, always. If they do not, the counting is wrong.

Accuracy, precision, recall and F1

All four come out of those same four numbers.

Accuracy is how often the model was right, over everything.

accuracy = (TP + TN) / (TP + FP + FN + TN)

Precision is how often it was right when it said Pass. Of the students you predicted would pass, what fraction did?

precision = TP / (TP + FP)

Recall is how many of the real passes it found. Of the students who actually passed, what fraction did you catch?

recall = TP / (TP + FN)

Precision and recall pull against each other. Predict Pass for everybody and recall is 1.00 and precision is poor. Predict Pass only for the one student you are certain of and precision is 1.00 and recall is terrible.

F1 is the single number that refuses to let you cheat either way. It is the harmonic mean of precision and recall, which is low whenever either of them is low.

F1 = 2 precision recall / (precision + recall)

def confusion(actual, predicted, positive):
    tp = fp = fn = tn = 0
    for a, p in zip(actual, predicted):
        if p == positive and a == positive:
            tp += 1
        elif p == positive:
            fp += 1
        elif a == positive:
            fn += 1
        else:
            tn += 1
    return tp, fp, fn, tn

def scores(actual, predicted, positive):
    tp, fp, fn, tn = confusion(actual, predicted, positive)
    total = tp + fp + fn + tn
    acc = (tp + tn) / total
    prec = tp / (tp + fp) if tp + fp else 0.0
    rec = tp / (tp + fn) if tp + fn else 0.0
    f1 = 2 * prec * rec / (prec + rec) if prec + rec else 0.0
    return acc, prec, rec, f1

actual = ["Pass", "Fail", "Pass", "Fail", "Fail", "Pass", "Fail", "Pass", "Fail"]
models = [("the model", ["Pass", "Fail", "Fail", "Fail", "Pass", "Pass", "Fail", "Pass", "Fail"]),
          ("always Fail", ["Fail"] * 9),
          ("always Pass", ["Pass"] * 9)]

print("%-12s %9s %10s %7s %6s" % ("model", "accuracy", "precision", "recall", "F1"))
for name, guess in models:
    acc, prec, rec, f1 = scores(actual, guess, "Pass")
    print("%-12s %9.3f %10.3f %7.3f %6.3f" % (name, acc, prec, rec, f1))
munotes.in14

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

model         accuracy  precision  recall     F1
the model        0.778      0.750   0.750  0.750
always Fail      0.556      0.000   0.000  0.000
always Pass      0.444      0.444   1.000  0.615

Read the bottom row, because it is the whole reason F1 exists. "Always Pass" has perfect recall: it catches every student who really passed, because it says Pass about everybody. Its precision is 0.444, because more than half the students it named are wrong. And its F1 is 0.615, not the 0.722 an ordinary average of 0.444 and 1.000 would give, because the harmonic mean is dragged towards the smaller number.

That is the point. A model cannot buy a good F1 by making one of the two perfect at the other's cost, and a model that answers the same thing every time is caught by every column except the one it cheated.

Print these as a table with a heading row, as the program above does. Five numbers on one long line with the labels squeezed between them is how a figure gets read against the wrong heading, and that mistake in a journal costs marks that the program had already earned.

Note one more thing before leaving this table. "Always Fail" scores 0.556 here, and on the real seed-1 test set earlier in this chapter the same stupid model scored 0.444. Both are right: this table's nine rows are four Pass and five Fail, and the seed-1 split's nine rows are five Pass and four Fail. A baseline belongs to one test set. Quote the two together or quote neither.

munotes.in15

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

When the answer is a number, not a class

K-NN, which MU sets as "classification or regression", can predict a number instead of a class: not Pass or Fail but a mark out of 100. Accuracy makes no sense then, because a prediction of 71 when the truth is 72 is not "wrong".

Two measures do the work.

Mean absolute error, the average size of the mistake, in the same units as the answer.

MAE = (1/n) * sum of |actual - predicted|

Root mean squared error, which squares each mistake before averaging, so one big miss counts for much more than several small ones.

RMSE = square root of ((1/n) * sum of (actual - predicted)^2)

import math

actual = [72, 55, 88, 41, 66, 79]
guess = [70, 58, 84, 45, 65, 95]

n = len(actual)
errors = [a - p for a, p in zip(actual, guess)]
mae = sum(abs(e) for e in errors) / n
rmse = math.sqrt(sum(e * e for e in errors) / n)

print("%-8s %8s %8s %8s" % ("actual", "guess", "error", "squared"))
for a, p, e in zip(actual, guess, errors):
    print("%-8d %8d %8d %8d" % (a, p, e, e * e))
print()
print("MAE  : %.3f" % mae)
print("RMSE : %.3f" % rmse)
actual      guess    error  squared
72             70        2        4
55             58       -3        9
88             84        4       16
41             45       -4       16
66             65        1        1
79             95      -16      256

MAE  : 5.000
RMSE : 7.095

Five of the six predictions are within four marks. The sixth is out by sixteen, and the two measures treat that very differently: MAE is 5.000 and RMSE is 7.095, because squaring turned that one miss into 256 of the 302 total. RMSE is the one to quote when a big mistake matters more than several small ones, which for a mark prediction it usually does.

The same measurements from scikit-learn

Everything above is four lines of scikit-learn. Write your own once, so you know what the four lines do, then use theirs.

from sklearn.metrics import (accuracy_score, precision_score, recall_score,
                             f1_score, confusion_matrix)

actual = ["Pass", "Fail", "Pass", "Fail", "Fail", "Pass", "Fail", "Pass", "Fail"]
guess = ["Pass", "Fail", "Fail", "Fail", "Pass", "Pass", "Fail", "Pass", "Fail"]

print("accuracy :", accuracy_score(actual, guess))
print("precision:", precision_score(actual, guess, pos_label="Pass"))
print("recall   :", recall_score(actual, guess, pos_label="Pass"))
print("F1       :", f1_score(actual, guess, pos_label="Pass"))
print("matrix, labels [Fail, Pass]:")
print(confusion_matrix(actual, guess, labels=["Fail", "Pass"]))
accuracy : 0.7777777777777778
precision: 0.75
recall   : 0.75
F1       : 0.75
matrix, labels [Fail, Pass]:
[[4 1]
 [1 3]]

The four figures are the ones our own program computed, which is the check that our own program is right.

confusion_matrix lays the matrix out the other way round from the table earlier in this chapter: rows are the true label in the order you give in labels, columns are the prediction, so with labels=["Fail", "Pass"] the top-left cell is TN, not TP. Always pass labels and always say in your write-up which corner is which, because half the marks lost on a confusion matrix are lost to reading it upside down.

munotes.in16

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

And the split, in one line:

import csv
from sklearn.model_selection import train_test_split

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
X = [[int(r["attendance"]), int(r["practice"])] for r in rows]
y = [r["result"] for r in rows]

Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=1,
                                      stratify=y)
print("train:", len(Xtr), "rows,", ytr.count("Pass"), "Pass")
print("test :", len(Xte), "rows,", yte.count("Pass"), "Pass")
train: 21 rows, 10 Pass
test : 9 rows, 4 Pass

stratify=y is what keeps the proportion of Pass the same in both halves: 10 of 21 and 4 of 9, against the file's 14 of 30. Pass random_state or the split changes on every run and so does every figure in your journal.

Procedure

  1. Save students.csv and read it, converting the two feature columns to whole numbers.
  2. Count the rows and the two classes.
  3. Shuffle with a seed, split 70 to 30, and print the balance of both halves.
  4. Compute the baseline: the accuracy of always answering the commoner class on your test set.
  5. Count TP, FP, FN and TN for a set of predictions and print the confusion matrix with its headings.
  6. Compute accuracy, precision, recall and F1 from those four counts, and print them as a table with a heading row.
  7. Compute MAE and RMSE for a set of numeric predictions, with the per-row error and squared error shown.
  8. Repeat steps 5 to 7 with sklearn.metrics and confirm the figures agree with your own.

Observations

MeasuredValue
Rows in the dataset30
Class balance14 Pass, 16 Fail
Split at 70 per cent, seed 121 training, 9 test
Test balance at seed 15 Pass, 4 Fail
Test balance over seeds 1 to 8between 4 and 6 Pass out of 9
Test balance, stratified4 Pass, 5 Fail
Baseline on the seed-1 test set0.444, because training says Fail and the test set has more Pass
The example modelaccuracy 0.778, precision 0.750, recall 0.750, F1 0.750
Always Passaccuracy 0.444, precision 0.444, recall 1.000, F1 0.615
Always Failaccuracy 0.556, precision 0.000, recall 0.000, F1 0.000
Numeric predictionsMAE 5.000, RMSE 7.095 (one error of 16 contributed 256 of the 302)
scikit-learn agreementaccuracy, precision, recall and F1 identical to our own

Result

A measurement kit for this whole module was built and checked: a seeded 70 to 30 split, a baseline, a confusion matrix, accuracy, precision, recall and F1 from the four counts, and MAE and RMSE for numeric answers. Every figure our own code produced was reproduced by sklearn.metrics on the same inputs.

munotes.in17

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

Where marks are lost

Measuring on the training set. The commonest and the most serious. A model tested on what it learned from can score 1.00 and be useless. Say in your journal which rows the score was measured on.

No baseline. 0.70 accuracy sounds like a result until the reader learns that answering Fail every time scores 0.667.

Quoting accuracy alone on an unbalanced dataset. Report the confusion matrix too, or precision and recall.

Reading the confusion matrix upside down. sklearn's rows are the truth and its columns are the prediction, and the order of the classes is whatever labels says. State which corner is TP.

Not seeding the split. Then no number in the journal can be reproduced.

Reading a number off a long line against the wrong label. Print a table with a heading row. This chapter contains a worked example of getting that wrong.

Using accuracy for a numeric prediction. Use MAE or RMSE, and say which.

For the journal

Aim; the dataset with its row count and class balance; the split program with the sizes and balance of both halves; the baseline and its figure; the confusion matrix with headings and the four counts; the metrics table with its heading row; the MAE and RMSE table with the per-row squared error; the sklearn.metrics run confirming the same figures; the observation table above; the result.

Quick revision

  • Train on the training set. Measure on the test set. Never the other way round.
  • Shuffle before splitting, and seed the shuffle.
  • stratify=y keeps the class proportions in both halves.
  • Baseline: always answer the commoner class. Our stratified test set gives 0.556.
  • TP, FP, FN, TN. A False Positive is a prediction of Positive that is false.
  • accuracy = (TP + TN) / total.
  • precision = TP / (TP + FP): how often "yes" was right.
  • recall = TP / (TP + FN): how much of the real "yes" was found.
  • F1 = 2 precision recall / (precision + recall), the harmonic mean, low if either is low.
  • MAE is the average size of the error. RMSE squares first, so one big miss dominates.
  • sklearn.metrics has all of them; confusion_matrix has truth as rows and prediction as columns.

Questions you must be able to answer

1. Why must a model be tested on data it has not seen? Because a model is allowed to remember. One that memorises the training rows answers all of them correctly and has learned nothing, so a score on the training set measures memory, not learning.

2. What is the baseline, and why compute it? The score of the stupidest model, which is to answer the commonest class every time. It is what any real model has to beat, and without it a number like 0.70 cannot be judged.

munotes.in18

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

3. Define precision and recall in one sentence each. Precision is the fraction of the rows you predicted positive that really were positive. Recall is the fraction of the really positive rows that you predicted positive.

4. Which do you care about more when predicting that a student will fail? Recall, if the purpose is to reach every student who is in trouble: missing one is worse than warning one extra. Precision, if a warning is expensive. Say which you chose and why.

5. Why is F1 the harmonic mean and not the ordinary average? Because the harmonic mean is dragged down by the smaller of the two, so a model cannot score well by making precision or recall perfect at the other's expense. Precision 0.444 with recall 1.000 gives F1 0.615, not 0.722.

6. What is in each of the four cells of a confusion matrix? TP, correct positives. FP, predicted positive but actually negative. FN, predicted negative but actually positive. TN, correct negatives. They add to the number of test rows.

7. When would you report RMSE instead of accuracy? When the prediction is a number rather than a class, as in K-NN regression, and especially when one large error matters more than several small ones.

8. What does random_state do to train_test_split, and why does it matter here? It fixes the shuffle, so the same rows go into the test set on every run, so the figure in your journal is the figure the examiner will see.

9. Our test set has nine rows. What is the smallest change in accuracy you can measure? One row in nine, which is 0.111. A difference of 0.05 between two models on nine test rows is not a difference at all, and a write-up should say so rather than rank them.

munotes.in19

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Issue
Done!