munotes®

Unsupervised Learning

Get access to whole semester resourcesSemester Pass

Chapter Forty-Five

Syllabus topic Module 2, "unsupervised"

Pages 246 to 249 of 591

In one line

Unsupervised learning is looking for structure in data that nobody has labelled, so there is no right answer to be checked against.

In the wording a student can write in an examination: unsupervised learning infers structure from a training set of unlabelled instances. There is no target output and no critic, so success cannot be measured against a correct answer; the learner is instead judged by whether the structure it finds is useful or meaningful. The principal tasks are clustering, association rule mining, dimensionality reduction and anomaly detection.

What it means to have no labels

The definition is easiest to see as what is missing from Supervised Learning.

SupervisedUnsupervised
Each example carriesan input and its correct outputan input only
The learner is askedpredict the outputfind structure
Error isthe difference from the labelnot defined
Success ismeasurable, exactlya judgement, usually
The criticsupplies the answerdoes not exist

The third and fourth rows are the whole difficulty. With no label there is no error to minimise, so there is no obvious objective and no obvious way to tell a good result from a bad one. Every unsupervised method therefore has to invent an objective, and different inventions give different answers on the same data. Hierarchical Clustering, and Judging a Clustering is where that is taken seriously.

The four things it is used for

A paper asking what unsupervised learning does wants these, and an example of each.

Clustering. Group the instances so that those in a group are more like each other than like those in other groups. Segmenting students by study pattern; grouping news articles by topic. Clustering and k-Means is MU's label for it.

Association rule mining. Find which things occur together. Which items are bought in the same basket. Support, Confidence and Lift and The Apriori Algorithm are MU's labels.

Dimensionality reduction. Describe each instance with fewer numbers while losing as little as possible. A hundred measurements of a student reduced to three that capture most of the variation, which makes everything downstream cheaper and often works better. Not on MU's label; named here because it is one of the four and a paper may ask for the list.

Anomaly detection. Find the instances unlike the rest. A transaction unlike any the cardholder has made before. It is unsupervised precisely because nobody has a labelled set of every kind of fraud that might be invented tomorrow.

Why it is worth doing when it cannot be scored

Three honest reasons, and they are why the row exists at all.

The data is there and the labels are not. Almost all data is unlabelled. A shop's till records exist whether or not anyone has classified the baskets. Using them requires a method that does not need labels.

munotes.in246

Unsupervised Learning

It is a step before supervised learning. Clustering can suggest what the classes should be; dimensionality reduction can shrink the input to something a supervised model can learn from with the data available. Much practical unsupervised work is preparation rather than an end in itself.

The structure itself is the answer. A shop wanting to know which products sell together is not predicting anything. Support, Confidence and Lift produces the answer directly.

What makes it hard, stated plainly

Three difficulties, and a paper can ask for any of them.

There is no correct answer to compare with. Two clusterings of the same customers, one by spending and one by region, can both be defensible. Which is better depends on what the clustering is for, and the algorithm was not told.

A structure is always found. Run k-means asking for four clusters on data with no groups in it at all, and it returns four clusters. The method cannot report that there was nothing there. That is the single most dangerous property of unsupervised learning and it is why Hierarchical Clustering, and Judging a Clustering insists on asking whether a clustering means anything before using it.

The number of groups is usually a hyperparameter. k-means must be told k. Choosing it is a judgement, and the standard devices, the elbow and the silhouette, are measurements rather than proofs.

Semi-supervised and self-supervised, named once

Neither is on MU's label and both come up, so one line each prevents confusion.

Semi-supervised learning uses a small labelled set together with a large unlabelled one. The unlabelled data reveals the shape of the input distribution, and the few labels say which part of it is which. It suits exactly the common situation where labels are expensive and raw data is free.

Self-supervised learning manufactures labels from the data itself: hide a word in a sentence and predict it, hide part of an image and predict that. The data was unlabelled and the task is then supervised, so it is a way of turning an unsupervised situation into a supervised one. It is how modern language models are trained, and it is worth knowing that the distinction between supervised and unsupervised is about the data, not about the machinery.

Distinctions

ClusteringClassification
Classes arediscoveredgiven in advance by the labels
Dataunlabelledlabelled
Right answernone to check againstthe label
Resultgroups with no namesa prediction of a known class
UnsupervisedSemi-supervisedSelf-supervised
Labelsnonea fewmade from the data itself
Then trained asunsupervisedboth togethersupervised
Anomaly detectionClassification of fraud
Needs examples of the bad classnoyes
Handles a kind never seen beforeyesno
Defines abnormal asunlike the restlike the labelled bad examples
munotes.in247

Unsupervised Learning

What it does not mean

Unsupervised does not mean unguided. The method's objective, its distance measure and its number of groups are all chosen by the designer, and they decide the answer.

No labels does not mean no assumptions. k-means assumes clusters are roughly round and of similar size. Change the assumption and the groups change.

Finding clusters does not mean clusters exist. The algorithm returns k groups whether or not the data has any.

It is not a weaker form of supervised learning. It answers different questions. "Which products sell together" has no label to predict.

Self-supervised learning is not unsupervised. The data is unlabelled and the task manufactured from it is supervised.

Quick revision

  • Unsupervised learning finds structure in unlabelled data. No critic, no target, no error to minimise, so each method must invent an objective.
  • Four tasks: clustering, association rule mining, dimensionality reduction, anomaly detection.
  • Why it is used: most data is unlabelled; it is a preparation step for supervised learning; and sometimes the structure is the answer.
  • Three difficulties: no correct answer to compare with; a structure is always found, even in data with none; and the number of groups is a hyperparameter.
  • Semi-supervised: a few labels with much unlabelled data. Self-supervised: labels manufactured from the data, after which the task is supervised.
  • Clustering discovers the classes; classification is given them.
  • Anomaly detection needs no examples of the bad class, which is why it handles kinds of fraud nobody has seen.

Test yourself

1. Define unsupervised learning and say what is missing compared with supervised learning. Learning structure from a training set of unlabelled instances. What is missing is the label on each example, so there is no target output, no critic, and no error that can be computed against a correct answer.

2. Name the four principal unsupervised tasks with an example of each. Clustering, such as grouping students by study pattern. Association rule mining, such as finding which products are bought together. Dimensionality reduction, such as describing each student with three numbers instead of a hundred. And anomaly detection, such as flagging a transaction unlike any the cardholder has made.

3. Why is unsupervised learning hard to evaluate? Because there is no correct answer to compare against. Two different groupings of the same data can both be defensible, and which is better depends on the purpose, which the algorithm was never told.

4. What is the most dangerous property of clustering, and what follows from it? That a structure is always found. Asked for four clusters in data with no groups at all, the algorithm returns four clusters and cannot report that there was nothing there. It follows that a clustering must be judged before it is used, not merely computed.

munotes.in248

Unsupervised Learning

5. Give three reasons unsupervised learning is worth doing despite the difficulty of scoring it. Most data is unlabelled, so a method that needs no labels is the only one available. It often serves as preparation for supervised learning, by suggesting classes or by reducing the input. And sometimes the structure found is itself the answer, as when a shop wants to know which products sell together.

6. Distinguish clustering from classification. Clustering discovers groups in unlabelled data and the groups have no names. Classification predicts one of a set of classes fixed in advance by the labels in the training data, and its accuracy can be measured against those labels.

7. Is self-supervised learning a form of unsupervised learning? Explain. The data is unlabelled, so the situation is unsupervised, but a label is manufactured from the data itself, such as hiding a word and predicting it. The task that is then trained is supervised. The distinction between supervised and unsupervised concerns whether labels were supplied with the data, not the machinery used afterwards.

munotes.in249

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!