What Deep Learning Is
Chapter Sixty-Three
Syllabus topic Module 2, "concept of deep learning"
Pages 357 to 361 of 591
In one line
Deep learning is a neural network with many layers, in which each layer learns features from the layer below, so the features are learned instead of designed.
In the wording a student can write in an examination: deep learning is machine learning with artificial neural networks of many hidden layers. Its defining property is representation learning: rather than being given features designed by a person, the network learns a hierarchy of representations, each layer composing the features of the one beneath into something more abstract. The same backpropagation trains it; what changed to make deep networks practical was more data, more computation, and a set of specific technical remedies.
What depth actually buys
The Multilayer Network and Backpropagation recorded the universal approximation theorem: one hidden layer of enough units can approximate any continuous function. So depth is not required for representability, and a paper asking why depth matters is asking for a different answer.
Depth buys efficiency of representation. Certain functions need exponentially many units in one layer and only polynomially many when the layers are stacked. A deep network expresses a composition of simple steps; a shallow one must enumerate the whole result.
And it buys a hierarchy. In a network trained on images, the early layers respond to edges, the middle layers to parts, the later layers to whole objects. Nobody designed that progression: each layer's features are composed from those below it because that is the cheapest way for the network to reduce its loss.
That is the claim worth being careful about. The hierarchy is a well-attested property of trained image networks; this book has not measured it and does not assert it as a general law. What is safe to say, and what a paper wants, is the principle: the features are learned rather than designed.
The change that actually happened
A paper often asks why deep learning succeeded when neural networks had been known since the 1950s. The honest answer has four parts and none of them is a new idea about networks.
| What changed | Why it mattered |
|---|---|
| Data | a network with millions of weights needs a great many examples, and digital data became abundant |
| Computation | the graphics processor turned out to do exactly the matrix arithmetic a network needs, thousands of times in parallel |
| Activations | ReLU in place of the sigmoid removed the vanishing gradient of The Multilayer Network and Backpropagation |
| Technique | better initialisation, normalisation between layers, dropout, and optimisers such as Adam made deep networks trainable in practice |
Backpropagation did not change. It is the algorithm of 1986, applied to larger networks on larger data with better activations. Saying that a new learning algorithm was discovered is a common and wrong answer.
What Deep Learning Is
What was given up
Every advantage in the previous section was paid for, and naming the costs is what separates an informed answer from an enthusiastic one.
Data. Deep networks need far more labelled data than any other method in MU's list. The Naive Bayes Classifier works on a few hundred rows; a deep network on a few hundred rows will simply memorise them.
Computation. Training a large network costs a great deal of electricity and time, and that cost is concentrated in the few organisations that can afford it.
Opacity. This is the one that matters for the rest of this syllabus. A decision tree's path is a rule a person can read; Reading, Drawing and Pruning a Decision Tree turned one into five sentences. A network's decision lives in millions of weights and cannot be read at all. Transparency and Explainability is where MU takes that up, and it is the direct cost of this chapter's subject.
Brittleness. A network can be extremely accurate on data like its training data and fail strangely on data slightly unlike it, with no warning and with high confidence. That is the same failure What Machine Learning Is showed with a straight line predicting 107.6 marks, and it does not become milder with depth.
Hyperparameters. The number of layers, their widths, the activations, the learning rate, the regularization and the initialisation are all choices, and the results depend on them.
The architectures, named once
MU sets the concept, not the architectures, so each gets one line and a paper wanting more than that is asking beyond the label.
- Convolutional network (CNN): units look at a small patch of the input and the same weights are reused across the whole image, so the number of weights does not grow with the image and a feature learned in one place is available everywhere. For images and anything with a grid structure.
- Recurrent network (RNN), and LSTM: the network has a loop, so its state carries information from one step of a sequence to the next. For sequences, and the direct relative of
Hidden Markov Models, which does the same job with an explicit probabilistic state. - Transformer: processes a whole sequence at once and learns which positions to attend to. It is the architecture behind current language models.
- Autoencoder: trained to reproduce its own input through a narrow middle layer, which forces that layer to be a compressed representation. It is unsupervised, and it is the bridge to
Unsupervised Learning.
Where it fits among MU's models
The comparison a Q.3 would ask for.
| Deep network | Decision tree | Naive Bayes | SVM | k-NN | |
|---|---|---|---|---|---|
| Data needed | a great deal | moderate | very little | moderate | moderate |
| Features designed by | the network | a person | a person | a person, or a kernel | a person |
| Interpretable | no | yes | partly | partly | by example |
| Training cost | high | low | negligible | moderate to high | none |
| Best where | images, speech, language, very large data | rules matter and must be read | text, small data | many features, few rows | low dimensions |
What Deep Learning Is
A deep network is not the default choice. On a table of a few hundred rows with a dozen features, which is what most real problems look like, a tree or an ensemble of trees will usually match or beat it, train in a second, and be readable. Deep learning earns its cost where the input is raw and high-dimensional and the data is abundant, which is exactly the case for images, audio and text.
What this book can and cannot show
An honest limit, stated rather than concealed. Every program in this book is plain Python run on three interpreters, and a deep network large enough to demonstrate representation learning cannot be trained that way in a chapter. What this book does show is the whole mechanism at a scale a reader can check: The Multilayer Network and Backpropagation computes one weight update by hand to eight decimal places and then trains a network to do what one neuron provably cannot. A deep network is that, repeated, with more layers and more data. Nothing conceptual is missing; scale is.
And Using an AI Library Responsibly is where the reader is shown what a library call does in terms of these chapters, with the same honesty about what is not installed on the machine this book was checked on.
Distinctions
| Shallow network | Deep network | |
|---|---|---|
| Hidden layers | one | many |
| Can approximate any continuous function | yes, in theory | yes |
| Units needed for some functions | exponentially many | polynomially many |
| Features | still a single transformation | a hierarchy |
| Feature engineering | Representation learning | |
|---|---|---|
| Features come from | a person's knowledge of the domain | the network, from data |
| Needs domain expertise | yes | less |
| Needs data | less | much more |
| Transfers to a new problem | rarely | often, by reusing early layers |
| What changed by 1986 | What changed by 2012 | |
|---|---|---|
| The algorithm | backpropagation was published | unchanged |
| Data | scarce | abundant |
| Computation | scarce | the graphics processor |
| Activation | sigmoid | ReLU |
What it does not mean
Deep does not mean better. On small tabular data a tree or an ensemble usually wins, trains in a second and can be read.
Deep learning is not a new learning algorithm. It is backpropagation, from 1986, on more data with better activations.
The universal approximation theorem does not make depth pointless. One layer can represent the function; it may need exponentially many units and training may not find it.
What Deep Learning Is
Learned features are not interpretable features. They are effective and they are not readable, which is the cost Transparency and Explainability addresses.
More layers are not automatically better. Depth adds capacity, and capacity without data is overfitting.
Accuracy on a benchmark is not reliability in use. A network can be excellent on data like its training data and fail without warning on data slightly unlike it.
Quick revision
- Deep learning: neural networks with many hidden layers, whose defining property is representation learning, a hierarchy of features learned rather than designed.
- Depth is not needed for representability, by the universal approximation theorem. It buys efficiency: some functions need exponentially many units in one layer and polynomially many when stacked.
- Four things changed, and none of them is the algorithm: data, computation (the graphics processor), ReLU in place of the sigmoid, and technique (initialisation, normalisation, dropout, Adam).
- The costs: much more data, much more computation, opacity, brittleness, and many hyperparameters.
- Architectures, one line each: CNN (shared weights over patches, for grids), RNN and LSTM (a loop carrying state, for sequences, the relative of HMMs), transformer (attends over a whole sequence), autoencoder (reproduces its input through a narrow layer, unsupervised).
- Not the default choice. On a few hundred rows with a dozen features a tree or an ensemble usually wins. Deep learning earns its cost on raw, high-dimensional input with abundant data.
- This book shows the mechanism at a checkable scale and says plainly that it cannot show the scale.
Test yourself
1. Define deep learning and name its defining property. Machine learning with neural networks of many hidden layers. Its defining property is representation learning: the network learns a hierarchy of features from the data instead of being given features designed by a person.
2. If one hidden layer can approximate any continuous function, why does depth matter? Because the theorem concerns representability, not efficiency or trainability. Some functions require exponentially many units in a single layer and only polynomially many when the computation is composed across several, and a deep network's features are built from those beneath them rather than enumerated.
3. Why did deep learning succeed when neural networks had been known for decades? Because of abundant data, the graphics processor providing the parallel matrix arithmetic networks need, the ReLU activation removing the vanishing gradient, and practical technique such as better initialisation, normalisation, dropout and modern optimisers. The learning algorithm itself, backpropagation, did not change.
4. Give four costs of deep learning. It needs far more labelled data than any other method in this syllabus; training is computationally expensive; the resulting model is opaque and cannot be read as a rule; it is brittle, failing confidently on data slightly unlike its training data; and it has many hyperparameters whose values change the result.
What Deep Learning Is
5. Name three deep architectures and say what each suits. A convolutional network, which reuses the same weights across small patches and suits images and other grid data. A recurrent network or LSTM, which carries state along a sequence and suits sequential data, playing the role that hidden Markov models play probabilistically. And a transformer, which attends across a whole sequence at once and underlies current language models.
6. When would you not use a deep network? On a modest table of a few hundred rows with a dozen designed features, where a decision tree or an ensemble of trees will usually match or beat it, train almost instantly, and be readable. Deep learning earns its cost when the input is raw and high-dimensional, such as images, audio or text, and the data is abundant.
7. What does the opacity of a deep network cost, and where does this syllabus address it? It costs the ability to explain a decision. A decision tree's path is a rule a person can read, while a network's decision is distributed across millions of weights. MU addresses this in the Responsible AI row, under transparency and explainability, and it is the direct consequence of choosing a learned representation over a designed one.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.