Parametric and Nonparametric Models
Chapter Forty-Seven
Syllabus topic Module 2, "Parametric vs Nonparametric"
Pages 254 to 258 of 591
In one line
A parametric model boils the data down to a fixed handful of numbers and then throws the data away; a nonparametric model keeps the data and consults it every time.
In the wording a student can write in an examination: a parametric model summarises the training data in a fixed, finite number of parameters chosen before the data is seen; the size of the model does not grow with the amount of data. A nonparametric model has a number of parameters that grows with the training set, so it retains more of the data and can represent a wider range of functions. The distinction is about how model complexity relates to data size, not about whether parameters exist.
The two on one dataset
Eight students, hours studied against marks.
The parametric model is the straight line of What Machine Learning Is: marks = 17.881 + 7.476 * hours. Two numbers. Fit them, discard the eight rows, and the model still predicts.
The nonparametric model is nearest neighbour: to predict for a new number of hours, find the closest student in the training set and return their mark. No numbers are fitted at all, and the whole training set is the model.
# Parametric against nonparametric, on the SAME eight students, with the memory
# and the prediction cost of each COUNTED rather than described.
DATA = [(2, 32), (3, 41), (4, 48), (5, 56), (6, 61), (7, 72), (8, 77), (9, 85)]
def fit_line(points):
n = len(points)
mx = sum(x for x, _ in points) / n
my = sum(y for _, y in points) / n
b = (sum((x - mx) * (y - my) for x, y in points)
/ sum((x - mx) ** 2 for x, _ in points))
return my - b * mx, b
a, b = fit_line(DATA)
def parametric(h):
"""Two numbers. The data is not consulted."""
return a + b * h, 0 # (prediction, rows examined)
def nonparametric(h):
"""The whole training set IS the model. Every row is examined."""
best = min(DATA, key=lambda row: abs(row[0] - h))
return float(best[1]), len(DATA)
print("the parametric model is two numbers: %.3f and %.3f" % (a, b))
print("the nonparametric model is all %d rows of the training data" % len(DATA))
print()
print(" hours | parametric | rows read | nonparametric | rows read")
for h in (2.0, 4.5, 6.5, 9.0, 12.0):
p, pr = parametric(h)
q, qr = nonparametric(h)
print(" %4.1f | %10.1f | %9d | %13.1f | %9d" % (h, p, pr, q, qr))
print()
print("what each costs, as the training set grows:")
print(" n rows | parametric: numbers stored | nonparametric: numbers stored")
for n in (8, 100, 10000, 1000000):
print(" %7d | %26d | %29d" % (n, 2, 2 * n))
print()
print("and the cost of ONE prediction:")
print(" parametric: 2 multiplications, whatever n is")
print(" nonparametric: n comparisons, growing with the data")
print()
print("at 12 hours the parametric model says %.1f, which is impossible," % parametric(12.0)[0])
print("and the nonparametric model says %.1f, the nearest student it has."
% nonparametric(12.0)[0])
print("neither is right. they are wrong in DIFFERENT ways, and that is the trade.")Parametric and Nonparametric Models
the parametric model is two numbers: 17.881 and 7.476
the nonparametric model is all 8 rows of the training data
hours | parametric | rows read | nonparametric | rows read
2.0 | 32.8 | 0 | 32.0 | 8
4.5 | 51.5 | 0 | 48.0 | 8
6.5 | 66.5 | 0 | 61.0 | 8
9.0 | 85.2 | 0 | 85.0 | 8
12.0 | 107.6 | 0 | 85.0 | 8
what each costs, as the training set grows:
n rows | parametric: numbers stored | nonparametric: numbers stored
8 | 2 | 16
100 | 2 | 200
10000 | 2 | 20000
1000000 | 2 | 2000000
and the cost of ONE prediction:
parametric: 2 multiplications, whatever n is
nonparametric: n comparisons, growing with the data
at 12 hours the parametric model says 107.6, which is impossible,
and the nonparametric model says 85.0, the nearest student it has.
neither is right. they are wrong in DIFFERENT ways, and that is the trade.What each is wrong about
Read the last two lines of the output, because they are the honest comparison.
At 12 hours, beyond anything in the data, the parametric model extrapolates its line and predicts a mark above 100, which is impossible. The nonparametric model returns the mark of the nearest student it has, 85 at 9 hours, and refuses to extrapolate at all.
Neither is right, and they are wrong in opposite directions. The line assumes the relation continues; the neighbour assumes nothing beyond its data and therefore says nothing new. That is the trade in one example, and it is worth more than any list of properties.
The trade, set out
| Parametric | Nonparametric | |
|---|---|---|
| Number of parameters | fixed, chosen in advance | grows with the data |
| Training data after fitting | can be discarded | must be kept |
| Memory | constant | proportional to n |
| Training cost | usually higher, an optimisation | often nothing at all |
| Prediction cost | constant and small | grows with n |
| Assumes a form for the function | yes, and it may be wrong | no |
| With little data | better: the assumption substitutes for data | worse: too few neighbours to trust |
| With a great deal of data | limited by its own form | better: it can represent anything |
| Bias | high, if the form is wrong | low |
| Variance | low | high |
| Extrapolates | yes, confidently and often wrongly | no |
| Examples in MU's list | linear models, naive Bayes, a neural network of fixed size | k-NN, a decision tree grown to fit, kernel SVM |
Parametric and Nonparametric Models
The two rows on bias and variance are the connection to the next chapter. A parametric model's fixed form is exactly a restriction on its hypothesis space, so it has high bias and low variance. A nonparametric model has almost no restriction, so low bias and high variance. Bias and Variance measures both.
The misconception in the name
Nonparametric does not mean "has no parameters". k-nearest neighbours has a parameter, k. A decision tree has one parameter per split, and often many. A kernel SVM has one weight per support vector.
What the word actually means is that the number of parameters is not fixed in advance: it grows with the training set, so the model's capacity grows as more data arrives. The precise reading of the term is "not characterised by a fixed, finite set of parameters", and a paper that asks you to define it wants exactly that.
A second, related misreading: nonparametric does not mean assumption-free. k-NN assumes that nearby inputs have similar outputs, which is a strong assumption and is false wherever the function jumps. Every method has an inductive bias, as Supervised Learning established.
Which to choose
A paper asking "when would you use each" expects the conditions, not a preference.
Choose parametric when the data is limited, when the form of the relation is known or can be assumed with confidence, when prediction must be fast or must run on a small device, when the model must be inspected or explained, and when the fitted model has to be transported without the data.
Choose nonparametric when there is plenty of data, when the form of the relation is unknown or clearly not simple, when training must be cheap or the model must absorb new examples continuously, and when accuracy matters more than prediction speed.
And the honest middle. Most of MU's models can be pushed either way. A decision tree is nonparametric when grown freely and effectively parametric when its depth is capped. A neural network of fixed architecture is parametric; adding layers as data arrives makes it nonparametric. The distinction is about how capacity relates to data, so it is a property of how a method is used as much as of the method.
Distinctions
| Parametric | Nonparametric | |
|---|---|---|
| Model size | fixed | grows with n |
| Keeps the training data | no | yes |
| Assumes a functional form | yes | no |
| Bias and variance | high bias, low variance | low bias, high variance |
Parametric and Nonparametric Models
| "Nonparametric" as commonly misread | What it means | |
|---|---|---|
| Has no parameters | not fixed in NUMBER in advance | |
| Makes no assumptions | still has an inductive bias, such as nearby means similar |
| Instance-based, or lazy | Model-based, or eager | |
|---|---|---|
| Training does | almost nothing | the work |
| Prediction does | the work | almost nothing |
| Example | k-NN | a fitted line |
| Usually | nonparametric | parametric |
What it does not mean
Nonparametric does not mean without parameters. It means the number of them is not fixed in advance.
Nonparametric does not mean without assumptions. k-NN assumes nearby inputs have similar labels.
Parametric does not mean simple. A neural network with a billion weights is parametric, because the count was fixed before training.
Neither is better. With little data the parametric model's assumption is an advantage; with a great deal of data it becomes the limit.
The distinction is not fixed per algorithm. A depth-capped decision tree behaves parametrically; a freely grown one does not.
Keeping the data is not only a memory cost. It is also a privacy and a deployment cost: a nonparametric model cannot be shipped without shipping the training data with it.
Quick revision
- Parametric: a fixed number of parameters, chosen before seeing the data; the data can be discarded after fitting. Nonparametric: the number of parameters grows with the training set, so the data must be kept.
- The distinction is about capacity against data size, not about whether parameters exist.
- On the eight students: the line is two numbers; nearest neighbour is all eight rows, and reads every one for each prediction.
- At 12 hours the line predicts an impossible mark and the neighbour returns its nearest student's mark. Wrong in opposite directions: one extrapolates, the other refuses to.
- Parametric: high bias, low variance, better with little data, constant memory and fast prediction. Nonparametric: low bias, high variance, better with much data, memory and prediction cost growing with
n. - Nonparametric does not mean no parameters (k-NN has
k, a tree has one per split) and does not mean no assumptions (k-NN assumes nearby means similar). - MU's examples: parametric are linear models, naive Bayes, a fixed neural network; nonparametric are k-NN, a freely grown decision tree, a kernel SVM.
- Lazy or instance-based methods do nothing at training and the work at prediction; eager methods do the reverse.
Test yourself
1. Define parametric and nonparametric models. A parametric model has a fixed, finite number of parameters chosen before the data is seen, so its size does not grow with the data. A nonparametric model has a number of parameters that grows with the training set, so its capacity increases as more data arrives.
2. On the eight-student example, what is each model, and what does each cost? The parametric model is two numbers, the intercept and the slope, and predicting costs two multiplications whatever the data size. The nonparametric model is all eight rows, it stores two numbers per row, and predicting requires examining every row.
Parametric and Nonparametric Models
3. At 12 hours of study the two models disagree. What does each do, and what does that show? The line extrapolates and predicts a mark above 100, which is impossible. The nearest neighbour returns the mark of the closest student it has, 85 at nine hours, and cannot go beyond its data. Both are wrong, in opposite ways: one assumes the relation continues, the other assumes nothing at all.
4. Correct the statement "nonparametric models have no parameters". They have parameters, often many: k-nearest neighbours has k, a decision tree has one per split, a kernel support vector machine has a weight per support vector. What is not fixed in advance is the number of them, which grows with the training set.
5. Relate the distinction to bias and variance. A parametric model's fixed form restricts the hypothesis space, giving high bias and low variance. A nonparametric model imposes little restriction, giving low bias and high variance. That is why parametric models are better with little data and nonparametric ones with a great deal.
6. When would you choose a parametric model? When data is limited, when the form of the relation is known or safely assumed, when prediction must be fast or run on a small device, when the model must be explained, or when it must be deployed without shipping the training data.
7. Is a decision tree parametric or nonparametric? Justify. It depends on how it is used. Grown freely until the leaves are pure, it adds parameters as the data grows and is nonparametric. With its depth capped in advance, its size is bounded regardless of the data and it behaves parametrically. The distinction concerns how capacity relates to data size, which is partly a matter of how a method is used.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.