munotes®

What Deep Learning Is

Get access to whole semester resourcesSemester Pass

Chapter Sixty-Three

Syllabus topic Module 2, "concept of deep learning"

Pages 357 to 361 of 591

In one line

Deep learning is a neural network with many layers, in which each layer learns features from the layer below, so the features are learned instead of designed.

In the wording a student can write in an examination: deep learning is machine learning with artificial neural networks of many hidden layers. Its defining property is representation learning: rather than being given features designed by a person, the network learns a hierarchy of representations, each layer composing the features of the one beneath into something more abstract. The same backpropagation trains it; what changed to make deep networks practical was more data, more computation, and a set of specific technical remedies.

What depth actually buys

The Multilayer Network and Backpropagation recorded the universal approximation theorem: one hidden layer of enough units can approximate any continuous function. So depth is not required for representability, and a paper asking why depth matters is asking for a different answer.

Depth buys efficiency of representation. Certain functions need exponentially many units in one layer and only polynomially many when the layers are stacked. A deep network expresses a composition of simple steps; a shallow one must enumerate the whole result.

And it buys a hierarchy. In a network trained on images, the early layers respond to edges, the middle layers to parts, the later layers to whole objects. Nobody designed that progression: each layer's features are composed from those below it because that is the cheapest way for the network to reduce its loss.

That is the claim worth being careful about. The hierarchy is a well-attested property of trained image networks; this book has not measured it and does not assert it as a general law. What is safe to say, and what a paper wants, is the principle: the features are learned rather than designed.

The change that actually happened

A paper often asks why deep learning succeeded when neural networks had been known since the 1950s. The honest answer has four parts and none of them is a new idea about networks.

What changedWhy it mattered
Dataa network with millions of weights needs a great many examples, and digital data became abundant
Computationthe graphics processor turned out to do exactly the matrix arithmetic a network needs, thousands of times in parallel
ActivationsReLU in place of the sigmoid removed the vanishing gradient of The Multilayer Network and Backpropagation
Techniquebetter initialisation, normalisation between layers, dropout, and optimisers such as Adam made deep networks trainable in practice

Backpropagation did not change. It is the algorithm of 1986, applied to larger networks on larger data with better activations. Saying that a new learning algorithm was discovered is a common and wrong answer.

munotes.in357

What Deep Learning Is

What was given up

Every advantage in the previous section was paid for, and naming the costs is what separates an informed answer from an enthusiastic one.

Data. Deep networks need far more labelled data than any other method in MU's list. The Naive Bayes Classifier works on a few hundred rows; a deep network on a few hundred rows will simply memorise them.

Computation. Training a large network costs a great deal of electricity and time, and that cost is concentrated in the few organisations that can afford it.

Opacity. This is the one that matters for the rest of this syllabus. A decision tree's path is a rule a person can read; Reading, Drawing and Pruning a Decision Tree turned one into five sentences. A network's decision lives in millions of weights and cannot be read at all. Transparency and Explainability is where MU takes that up, and it is the direct cost of this chapter's subject.

Brittleness. A network can be extremely accurate on data like its training data and fail strangely on data slightly unlike it, with no warning and with high confidence. That is the same failure What Machine Learning Is showed with a straight line predicting 107.6 marks, and it does not become milder with depth.

Hyperparameters. The number of layers, their widths, the activations, the learning rate, the regularization and the initialisation are all choices, and the results depend on them.

The architectures, named once

MU sets the concept, not the architectures, so each gets one line and a paper wanting more than that is asking beyond the label.

  • Convolutional network (CNN): units look at a small patch of the input and the same weights are reused across the whole image, so the number of weights does not grow with the image and a feature learned in one place is available everywhere. For images and anything with a grid structure.
  • Recurrent network (RNN), and LSTM: the network has a loop, so its state carries information from one step of a sequence to the next. For sequences, and the direct relative of Hidden Markov Models, which does the same job with an explicit probabilistic state.
  • Transformer: processes a whole sequence at once and learns which positions to attend to. It is the architecture behind current language models.
  • Autoencoder: trained to reproduce its own input through a narrow middle layer, which forces that layer to be a compressed representation. It is unsupervised, and it is the bridge to Unsupervised Learning.

Where it fits among MU's models

The comparison a Q.3 would ask for.

Deep networkDecision treeNaive BayesSVMk-NN
Data neededa great dealmoderatevery littlemoderatemoderate
Features designed bythe networka persona persona person, or a kernela person
Interpretablenoyespartlypartlyby example
Training costhighlownegligiblemoderate to highnone
Best whereimages, speech, language, very large datarules matter and must be readtext, small datamany features, few rowslow dimensions
munotes.in358

What Deep Learning Is

A deep network is not the default choice. On a table of a few hundred rows with a dozen features, which is what most real problems look like, a tree or an ensemble of trees will usually match or beat it, train in a second, and be readable. Deep learning earns its cost where the input is raw and high-dimensional and the data is abundant, which is exactly the case for images, audio and text.

What this book can and cannot show

An honest limit, stated rather than concealed. Every program in this book is plain Python run on three interpreters, and a deep network large enough to demonstrate representation learning cannot be trained that way in a chapter. What this book does show is the whole mechanism at a scale a reader can check: The Multilayer Network and Backpropagation computes one weight update by hand to eight decimal places and then trains a network to do what one neuron provably cannot. A deep network is that, repeated, with more layers and more data. Nothing conceptual is missing; scale is.

And Using an AI Library Responsibly is where the reader is shown what a library call does in terms of these chapters, with the same honesty about what is not installed on the machine this book was checked on.

Distinctions

Shallow networkDeep network
Hidden layersonemany
Can approximate any continuous functionyes, in theoryyes
Units needed for some functionsexponentially manypolynomially many
Featuresstill a single transformationa hierarchy
Feature engineeringRepresentation learning
Features come froma person's knowledge of the domainthe network, from data
Needs domain expertiseyesless
Needs datalessmuch more
Transfers to a new problemrarelyoften, by reusing early layers
What changed by 1986What changed by 2012
The algorithmbackpropagation was publishedunchanged
Datascarceabundant
Computationscarcethe graphics processor
ActivationsigmoidReLU

What it does not mean

Deep does not mean better. On small tabular data a tree or an ensemble usually wins, trains in a second and can be read.

Deep learning is not a new learning algorithm. It is backpropagation, from 1986, on more data with better activations.

The universal approximation theorem does not make depth pointless. One layer can represent the function; it may need exponentially many units and training may not find it.

munotes.in359

What Deep Learning Is

Learned features are not interpretable features. They are effective and they are not readable, which is the cost Transparency and Explainability addresses.

More layers are not automatically better. Depth adds capacity, and capacity without data is overfitting.

Accuracy on a benchmark is not reliability in use. A network can be excellent on data like its training data and fail without warning on data slightly unlike it.

Quick revision

  • Deep learning: neural networks with many hidden layers, whose defining property is representation learning, a hierarchy of features learned rather than designed.
  • Depth is not needed for representability, by the universal approximation theorem. It buys efficiency: some functions need exponentially many units in one layer and polynomially many when stacked.
  • Four things changed, and none of them is the algorithm: data, computation (the graphics processor), ReLU in place of the sigmoid, and technique (initialisation, normalisation, dropout, Adam).
  • The costs: much more data, much more computation, opacity, brittleness, and many hyperparameters.
  • Architectures, one line each: CNN (shared weights over patches, for grids), RNN and LSTM (a loop carrying state, for sequences, the relative of HMMs), transformer (attends over a whole sequence), autoencoder (reproduces its input through a narrow layer, unsupervised).
  • Not the default choice. On a few hundred rows with a dozen features a tree or an ensemble usually wins. Deep learning earns its cost on raw, high-dimensional input with abundant data.
  • This book shows the mechanism at a checkable scale and says plainly that it cannot show the scale.

Test yourself

1. Define deep learning and name its defining property. Machine learning with neural networks of many hidden layers. Its defining property is representation learning: the network learns a hierarchy of features from the data instead of being given features designed by a person.

2. If one hidden layer can approximate any continuous function, why does depth matter? Because the theorem concerns representability, not efficiency or trainability. Some functions require exponentially many units in a single layer and only polynomially many when the computation is composed across several, and a deep network's features are built from those beneath them rather than enumerated.

3. Why did deep learning succeed when neural networks had been known for decades? Because of abundant data, the graphics processor providing the parallel matrix arithmetic networks need, the ReLU activation removing the vanishing gradient, and practical technique such as better initialisation, normalisation, dropout and modern optimisers. The learning algorithm itself, backpropagation, did not change.

4. Give four costs of deep learning. It needs far more labelled data than any other method in this syllabus; training is computationally expensive; the resulting model is opaque and cannot be read as a rule; it is brittle, failing confidently on data slightly unlike its training data; and it has many hyperparameters whose values change the result.

munotes.in360

What Deep Learning Is

5. Name three deep architectures and say what each suits. A convolutional network, which reuses the same weights across small patches and suits images and other grid data. A recurrent network or LSTM, which carries state along a sequence and suits sequential data, playing the role that hidden Markov models play probabilistically. And a transformer, which attends across a whole sequence at once and underlies current language models.

6. When would you not use a deep network? On a modest table of a few hundred rows with a dozen designed features, where a decision tree or an ensemble of trees will usually match or beat it, train almost instantly, and be readable. Deep learning earns its cost when the input is raw and high-dimensional, such as images, audio or text, and the data is abundant.

7. What does the opacity of a deep network cost, and where does this syllabus address it? It costs the ability to explain a decision. A decision tree's path is a rule a person can read, while a network's decision is distributed across millions of weights. MU addresses this in the Responsible AI row, under transparency and explainability, and it is the direct consequence of choosing a learned representation over a designed one.

munotes.in361

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!