munotes®

Multi-Class Classification, and How It Is Scored

Get access to whole semester resourcesSemester Pass

Chapter Eighty-Four

Syllabus topic Module 2, "Multi-class classification"

Pages 300 to 302 of 378

In one line

With more than two classes there is no single "positive" class, so accuracy alone stops being a report and the confusion matrix starts.

In the wording you can write in an examination: multi-class classification assigns each object one of more than two labels. A binary classifier can be extended to it by one-against-rest, training one classifier per class to distinguish it from all the others, or by one-against-one, training a classifier for every pair. Performance is reported by a confusion matrix, from which per-class precision and recall are computed, since a single accuracy figure hides which classes are confused with which.

The two ways to extend a binary method

One-against-rest. For k classes, build k classifiers. The i-th distinguishes class i from everything else. To classify, run all k and take the one that answers most strongly.

One-against-one. For k classes, build k times k minus 1, over 2 classifiers, one for each pair. To classify, run all of them and take the class that wins most pairwise contests.

One-against-restOne-against-one
Number of classifierskk(k-1)/2
For k equal to 333
For k equal to 7721
Each classifier seesall the dataonly two classes' data
Training setsimbalanced, one class against allbalanced, if the classes are
Tiespossible, resolved by the strength of the answerspossible, resolved by a rule

For the three doṣa the two methods need the same number of classifiers, which is a coincidence of k equal to three. For the seven labels of [Prakṛti: Constitution as a Class Label] they need seven and twenty-one.

And the multi-label reading of that chapter needs neither. Three independent yes-or-no decisions is three classifiers, which is one-against-rest without the requirement that exactly one wins.

The confusion matrix

What it is. A table with one row per true class and one column per predicted class. The cell at row i, column j is the number of cases whose true class is i and whose predicted class is j.

What the diagonal is. Correct predictions.

What everything else is. Errors, and the position says which error.

Worked. The example below is invented for this chapter and not the result of any experiment, because no labelled data for this scheme exists.

predicted windpredicted bilepredicted phlegm
true wind4064
true bile8284
true phlegm226

Read three things off it before computing anything.

The totals. 50 true wind, 40 true bile, 10 true phlegm. The classes are imbalanced, five to one between the commonest and the rarest.

The diagonal. 40 plus 28 plus 6 is 74 correct out of 100, so accuracy is 74 per cent.

munotes.in300

Multi-Class Classification, and How It Is Scored

And the errors are not symmetric. Eight true bile were called wind; six true wind were called bile. The classifier leans toward wind, which is the commonest class, and an accuracy figure says nothing about that.

Precision and recall, per class

Precision for a class. Of the cases predicted to be that class, how many were? It is the diagonal cell divided by the COLUMN total.

Recall for a class. Of the cases that truly were that class, how many were found? It is the diagonal cell divided by the ROW total.

Worked, for all three.

ClassColumn totalPrecisionRow totalRecall
wind40 + 8 + 2 = 5040 / 50 = 0.8040 + 6 + 4 = 5040 / 50 = 0.80
bile6 + 28 + 2 = 3628 / 36 = 0.788 + 28 + 4 = 4028 / 40 = 0.70
phlegm4 + 4 + 6 = 146 / 14 = 0.432 + 2 + 6 = 106 / 10 = 0.60

And now the report says something the accuracy figure did not. Overall accuracy is 74 per cent and phlegm's precision is 43 per cent: fewer than half the cases called phlegm were phlegm. A user acting on a prediction of phlegm is wrong more often than right, and nothing in "74 per cent accurate" warned them.

Check the arithmetic. The column totals add to 100 and the row totals add to 100, and both must, because every case is counted once by its true class and once by its predicted class. That is a free check and it catches a mistranscribed cell at once.

Combining the per-class figures

Three ways, and a question may ask the difference.

Macro average. The mean of the per-class figures, each class counting equally. Macro precision here is the mean of 0.80, 0.78 and 0.43, which is 0.67.

Weighted average. The mean weighted by how common each class is. Phlegm is only a tenth of the cases, so it barely moves the figure.

Micro average. Pool all the cases and compute one figure. For single-label multi-class classification the micro average equals the accuracy.

Which to use. Macro if the rare classes matter as much as the common ones. Weighted if they do not. Reporting only one of them is how an unflattering result is hidden, and reporting which one was used is the minimum.

The baseline you must beat

Always answering the commonest class. Here that is wind, 50 of 100, so a classifier that always says wind is 50 per cent accurate.

So 74 per cent is 24 points above the baseline, not 74 points above nothing. Quoting accuracy without the baseline is quoting half a number, and with seven classes and one of them dominant the baseline can be very high indeed.

munotes.in301

Multi-Class Classification, and How It Is Scored

What this does NOT apply to in this book

No figure in this chapter comes from an experiment. The matrix above is made up to be worked. The programs in this block are not evaluated for accuracy, because no ground truth exists, and [Āyurveda as a Śāstra, and What This Chapter Does Not Claim] says so.

And a submission that reported an accuracy for such a program would be reporting a number it had invented. That is worth repeating at the point where the machinery for reporting accuracy has just been taught.

Quick revision

  • One-against-rest: k classifiers, imbalanced training sets. One-against-one: k(k-1)/2 classifiers, balanced pairs. For k equal to 3 both need three.
  • Confusion matrix: rows are true classes, columns are predicted; the diagonal is correct and the position of an error says which error.
  • Precision is the diagonal cell over its COLUMN total; recall is the diagonal cell over its ROW total.
  • In the worked matrix accuracy is 74 per cent while phlegm's precision is 43 per cent, so a prediction of phlegm is wrong more often than right.
  • Row totals and column totals must each add to the number of cases, which is a free check.
  • Macro, weighted and micro averages differ, and micro equals accuracy in this setting. Say which you used.
  • Quote the majority-class baseline beside any accuracy figure.

Test yourself

1. Compute precision and recall for a class from a confusion matrix.

Precision is the diagonal cell for that class divided by the total of its column, the cases predicted to be it. Recall is the same cell divided by the total of its row, the cases that truly were it.

2. In the worked matrix, why is 74 per cent accuracy a misleading summary?

Because it hides that only 43 per cent of the cases predicted to be phlegm actually were, so a prediction of phlegm is wrong more often than right, and because the majority-class baseline is already 50 per cent.

3. Distinguish macro from weighted averaging and say when each is right.

Macro averages the per-class figures with every class counting equally; weighted averages them by how common each class is. Macro is right when rare classes matter as much as common ones; weighted when the overall case load is what matters.

4. Why does this book report no accuracy figure for its own Āyurvedic classifier?

Because no labelled data with a ground truth exists for the scheme, so there is nothing to measure against. A reported accuracy would be a number the submission had invented.

munotes.in302

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Issue
Done!