Practical 4: The Feed Forward Backpropagation Neural Network
Chapter Seven
Syllabus topic Module 1, "Feed Forward Backpropagation Neural Network: Implement the Feed Forward Backpropagation algorithm to train a neural network. Use a given dataset to train the neural network for a specific task. Evaluate the performance of the trained network on test data."
Pages 51 to 61 of 206
Aim
To implement the feed forward backpropagation algorithm, to train a neural network on a dataset, and to measure the trained network on data it has not seen.
What you need to know before you start
A neuron does two things. It adds up its inputs, each multiplied by a weight, plus one number of its own called the bias. Then it passes that sum through a squashing function.
z = b + w1x1 + w2x2 + ... + wn*xn
a = sigmoid(z) = 1 / (1 + e^(-z))
The sigmoid turns any number into one between 0 and 1: a large negative z gives nearly 0, a large positive z nearly 1, and z of 0 gives exactly 0.5. That is what lets the network's answer be read as "yes" or "no".
It has one more property, and it is the reason it is used here rather than anything simpler: its slope can be written in terms of its own output.
d/dz sigmoid(z) = sigmoid(z) (1 - sigmoid(z)) = a (1 - a)
Every a * (1 - a) in the program below is that line.
Feed forward means the values only ever travel one way: inputs to hidden layer, hidden layer to output. Backpropagation is how the blame for a wrong answer travels the other way: the output's error is shared out among the hidden neurons in proportion to the weights that carried their values forward.
The network in this chapter:
Figure 7.1 The 2-4-1 network. W1 holds the eight weights on the left and W2 the four on the right, and each neuron also has a bias of its own.
Step 1: one forward pass and one backward pass, with every number shown
Before any loop, do it once by hand. Two inputs, two hidden neurons, one output, weights chosen to be easy to read, the input 1 0 and the target 1.
"""One forward pass and one backward pass, with every number shown."""
import math
def sigmoid(z):
return 1 / (1 + math.exp(-z))
# a 2-2-1 network with weights chosen so the arithmetic can be checked by hand
w1 = [[0.10, 0.40], # input 1 to hidden 1, hidden 2
[0.20, 0.30]] # input 2 to hidden 1, hidden 2
b1 = [0.05, -0.10]
w2 = [0.50, -0.60] # hidden 1, hidden 2 to output
b2 = 0.15
x = [1.0, 0.0]
target = 1.0
rate = 0.5
print("FORWARD")
z1 = []
a1 = []
for j in range(2):
z = b1[j] + sum(x[i] * w1[i][j] for i in range(2))
z1.append(z)
a1.append(sigmoid(z))
print(" hidden %d: z = %.4f + %.2f*%.2f + %.2f*%.2f = %.4f a = sigmoid(z) = %.6f"
% (j + 1, b1[j], x[0], w1[0][j], x[1], w1[1][j], z, a1[j]))
z_out = b2 + sum(a1[j] * w2[j] for j in range(2))
a_out = sigmoid(z_out)
print(" output : z = %.4f + %.6f*%.2f + %.6f*%.2f = %.6f" % (b2, a1[0], w2[0], a1[1], w2[1], z_out))
print(" output : a = sigmoid(z) = %.6f" % a_out)
print()
err = 0.5 * (target - a_out) ** 2
print("ERROR")
print(" target %.1f, output %.6f" % (target, a_out))
print(" E = 0.5 * (target - output)^2 = %.6f" % err)
print()
print("BACKWARD")
d_out = (a_out - target) * a_out * (1 - a_out)
print(" delta at the output = (a - t) * a * (1 - a)")
print(" = (%.6f - %.1f) * %.6f * %.6f = %.6f"
% (a_out, target, a_out, 1 - a_out, d_out))
d_hid = []
for j in range(2):
d = d_out * w2[j] * a1[j] * (1 - a1[j])
d_hid.append(d)
print(" delta at hidden %d = %.6f * %.2f * %.6f * %.6f = %.6f"
% (j + 1, d_out, w2[j], a1[j], 1 - a1[j], d))
print()
print("UPDATE, learning rate %.1f" % rate)
for j in range(2):
new = w2[j] - rate * d_out * a1[j]
print(" w2[%d]: %.4f - %.1f * %.6f * %.6f = %.6f" % (j, w2[j], rate, d_out, a1[j], new))
new_b2 = b2 - rate * d_out
print(" b2 : %.4f - %.1f * %.6f = %.6f" % (b2, rate, d_out, new_b2))
for i in range(2):
for j in range(2):
new = w1[i][j] - rate * d_hid[j] * x[i]
print(" w1[%d][%d]: %.4f - %.1f * %.6f * %.1f = %.6f"
% (i, j, w1[i][j], rate, d_hid[j], x[i], new))Practical 4: The Feed Forward Backpropagation Neural Network
FORWARD
hidden 1: z = 0.0500 + 1.00*0.10 + 0.00*0.20 = 0.1500 a = sigmoid(z) = 0.537430
hidden 2: z = -0.1000 + 1.00*0.40 + 0.00*0.30 = 0.3000 a = sigmoid(z) = 0.574443
output : z = 0.1500 + 0.537430*0.50 + 0.574443*-0.60 = 0.074049
output : a = sigmoid(z) = 0.518504
ERROR
target 1.0, output 0.518504
E = 0.5 * (target - output)^2 = 0.115919
BACKWARD
delta at the output = (a - t) * a * (1 - a)
= (0.518504 - 1.0) * 0.518504 * 0.481496 = -0.120209
delta at hidden 1 = -0.120209 * 0.50 * 0.537430 * 0.462570 = -0.014942
delta at hidden 2 = -0.120209 * -0.60 * 0.574443 * 0.425557 = 0.017632
UPDATE, learning rate 0.5
w2[0]: 0.5000 - 0.5 * -0.120209 * 0.537430 = 0.532302
w2[1]: -0.6000 - 0.5 * -0.120209 * 0.574443 = -0.565473
b2 : 0.1500 - 0.5 * -0.120209 = 0.210105
w1[0][0]: 0.1000 - 0.5 * -0.014942 * 1.0 = 0.107471
w1[0][1]: 0.4000 - 0.5 * 0.017632 * 1.0 = 0.391184
w1[1][0]: 0.2000 - 0.5 * -0.014942 * 0.0 = 0.200000
w1[1][1]: 0.3000 - 0.5 * 0.017632 * 0.0 = 0.300000Practical 4: The Feed Forward Backpropagation Neural Network
Work down that output with a calculator once and backpropagation is no longer mysterious. Four things in it are worth saying in the journal.
The error before training is 0.1159 and the output is 0.5185 when the target is 1. The untrained network is guessing, which is what random weights do.
The delta at the output is (a - t) a (1 - a). The first factor is how wrong the answer is; the second and third are the slope of the sigmoid there. Multiplying by the slope is what makes this gradient descent rather than a guess: a neuron that is already saturated, output near 0 or near 1, has almost no slope and so moves almost not at all.
The delta at a hidden neuron is the output's delta, carried back through the weight that joined them, times that neuron's own slope. Hidden 1 gets a negative delta and hidden 2 a positive one, because their weights to the output have opposite signs. That single line is the whole of "backpropagation".
The two weights from input 2 did not move at all. Look at the last two lines: w1[1][0] and w1[1][1] come out exactly as they went in, because input 2 was 0 and every update is multiplied by the input. A weight from an input that is zero carries no blame, and it learns nothing from that example.
Step 2: the network, and a task it can only do with a hidden layer
XOR: output 1 when exactly one input is 1. It is the standard first task for a network because no straight line separates its two classes, so it cannot be done without a hidden layer, and demonstrating that is half the practical.
"""Practical 4: a feed forward network trained by backpropagation, on XOR."""
import math
import random
def sigmoid(z):
return 1 / (1 + math.exp(-z))
class Network:
"""One hidden layer. Weights are lists of lists; no library is used."""
def __init__(self, n_in, n_hidden, n_out, seed=1):
rng = random.Random(seed)
span = 1.0
self.w1 = [[rng.uniform(-span, span) for _ in range(n_hidden)] for _ in range(n_in)]
self.b1 = [rng.uniform(-span, span) for _ in range(n_hidden)]
self.w2 = [[rng.uniform(-span, span) for _ in range(n_out)] for _ in range(n_hidden)]
self.b2 = [rng.uniform(-span, span) for _ in range(n_out)]
def forward(self, x):
self.a1 = [sigmoid(self.b1[j] + sum(x[i] * self.w1[i][j] for i in range(len(x))))
for j in range(len(self.b1))]
self.a2 = [sigmoid(self.b2[k] + sum(self.a1[j] * self.w2[j][k] for j in range(len(self.a1))))
for k in range(len(self.b2))]
return self.a2
def backward(self, x, target, rate):
out = self.a2
d_out = [(out[k] - target[k]) * out[k] * (1 - out[k]) for k in range(len(out))]
d_hid = [sum(d_out[k] * self.w2[j][k] for k in range(len(out)))
* self.a1[j] * (1 - self.a1[j]) for j in range(len(self.a1))]
for j in range(len(self.a1)):
for k in range(len(out)):
self.w2[j][k] -= rate * d_out[k] * self.a1[j]
for k in range(len(out)):
self.b2[k] -= rate * d_out[k]
for i in range(len(x)):
for j in range(len(self.a1)):
self.w1[i][j] -= rate * d_hid[j] * x[i]
for j in range(len(self.a1)):
self.b1[j] -= rate * d_hid[j]
def train(self, data, epochs, rate, report_every=0):
curve = []
for epoch in range(1, epochs + 1):
total = 0.0
for x, t in data:
out = self.forward(x)
total += sum(0.5 * (t[k] - out[k]) ** 2 for k in range(len(t)))
self.backward(x, t, rate)
curve.append(total)
if report_every and (epoch == 1 or epoch % report_every == 0):
print(" epoch %6d total error %.6f" % (epoch, total))
return curve
XOR = [([0.0, 0.0], [0.0]),
([0.0, 1.0], [1.0]),
([1.0, 0.0], [1.0]),
([1.0, 1.0], [0.0])]
net = Network(2, 4, 1, seed=1)
print("training on XOR, 2 inputs, 4 hidden, 1 output, learning rate 0.5")
net.train(XOR, 20000, 0.5, report_every=4000)
print()
print("%-8s %10s %10s %8s" % ("input", "target", "output", "rounded"))
right = 0
for x, t in XOR:
out = net.forward(x)[0]
got = 1 if out >= 0.5 else 0
right += (got == int(t[0]))
print("%-8s %10.0f %10.6f %8d" % ("%d %d" % (x[0], x[1]), t[0], out, got))
print()
print("correct: %d of %d" % (right, len(XOR)))Practical 4: The Feed Forward Backpropagation Neural Network
training on XOR, 2 inputs, 4 hidden, 1 output, learning rate 0.5
epoch 1 total error 0.570178
epoch 4000 total error 0.001644
epoch 8000 total error 0.000655
epoch 12000 total error 0.000401
epoch 16000 total error 0.000287
epoch 20000 total error 0.000223
input target output rounded
0 0 0 0.010836 0
0 1 1 0.989533 1
1 0 1 0.989801 1
1 1 0 0.010678 0
correct: 4 of 4The error falls from 0.570 to 0.000223 and all four cases come out right. Note that the outputs are 0.0108 and 0.9895, not 0 and 1: a sigmoid never quite reaches either end, so the answer is rounded at 0.5. Print the raw output as well as the rounded one, because the raw number is the network's confidence and an examiner may ask for it.
Note also where the error falls. Between epoch 1 and epoch 4,000 it drops from 0.570 to 0.0016; over the next 16,000 epochs it only reaches 0.00022. Almost all the learning happens early, and training ten times longer buys very little. That is the shape of every learning curve you will draw in this subject.
Step 3: three things that decide whether it learns at all
"""Three things that decide whether a network learns anything at all."""
import math, random
def sigmoid(z): return 1 / (1 + math.exp(-z))
class Network:
def __init__(self, n_in, n_hidden, n_out, seed=1):
rng = random.Random(seed)
self.w1 = [[rng.uniform(-1, 1) for _ in range(n_hidden)] for _ in range(n_in)]
self.b1 = [rng.uniform(-1, 1) for _ in range(n_hidden)]
self.w2 = [[rng.uniform(-1, 1) for _ in range(n_out)] for _ in range(n_hidden)]
self.b2 = [rng.uniform(-1, 1) for _ in range(n_out)]
def forward(self, x):
self.a1 = [sigmoid(self.b1[j] + sum(x[i]*self.w1[i][j] for i in range(len(x))))
for j in range(len(self.b1))]
self.a2 = [sigmoid(self.b2[k] + sum(self.a1[j]*self.w2[j][k] for j in range(len(self.a1))))
for k in range(len(self.b2))]
return self.a2
def backward(self, x, t, rate):
out = self.a2
d_out = [(out[k]-t[k])*out[k]*(1-out[k]) for k in range(len(out))]
d_hid = [sum(d_out[k]*self.w2[j][k] for k in range(len(out)))*self.a1[j]*(1-self.a1[j])
for j in range(len(self.a1))]
for j in range(len(self.a1)):
for k in range(len(out)): self.w2[j][k] -= rate*d_out[k]*self.a1[j]
for k in range(len(out)): self.b2[k] -= rate*d_out[k]
for i in range(len(x)):
for j in range(len(self.a1)): self.w1[i][j] -= rate*d_hid[j]*x[i]
for j in range(len(self.a1)): self.b1[j] -= rate*d_hid[j]
def train(self, data, epochs, rate):
for _ in range(epochs):
total = 0.0
for x, t in data:
o = self.forward(x)
total += sum(0.5*(t[k]-o[k])**2 for k in range(len(t)))
self.backward(x, t, rate)
return total
XOR = [([0.,0.],[0.]), ([0.,1.],[1.]), ([1.,0.],[1.]), ([1.,1.],[0.])]
def score(net):
return sum(1 for x, t in XOR if (1 if net.forward(x)[0] >= 0.5 else 0) == int(t[0]))
print("A. how many hidden neurons XOR needs")
print("%-10s %14s %10s" % ("hidden", "final error", "correct"))
for n in (0, 1, 2, 3, 4):
if n == 0:
# no hidden layer at all: a single sigmoid unit on the two inputs
rng = random.Random(1)
w = [rng.uniform(-1, 1) for _ in range(2)]; b = rng.uniform(-1, 1)
for _ in range(20000):
total = 0.0
for x, t in XOR:
o = sigmoid(b + w[0]*x[0] + w[1]*x[1])
total += 0.5*(t[0]-o)**2
d = (o-t[0])*o*(1-o)
w[0] -= 0.5*d*x[0]; w[1] -= 0.5*d*x[1]; b -= 0.5*d
got = sum(1 for x, t in XOR
if (1 if sigmoid(b+w[0]*x[0]+w[1]*x[1]) >= 0.5 else 0) == int(t[0]))
print("%-10s %14.6f %10d" % ("none", total, got))
continue
net = Network(2, n, 1, seed=1)
err = net.train(XOR, 20000, 0.5)
print("%-10d %14.6f %10d" % (n, err, score(net)))
print()
print("B. the learning rate, 4 hidden neurons, 20000 epochs")
print("%-14s %14s %10s" % ("learning rate", "final error", "correct"))
for rate in (0.01, 0.1, 0.5, 2.0, 10.0, 50.0, 200.0):
net = Network(2, 4, 1, seed=1)
err = net.train(XOR, 20000, rate)
print("%-14.2f %14.6f %10d" % (rate, err, score(net)))
print()
print("C. the starting weights, 4 hidden neurons, rate 0.5")
print("%-10s %14s %10s" % ("seed", "final error", "correct"))
for seed in (1, 2, 3, 4, 5):
net = Network(2, 4, 1, seed=seed)
err = net.train(XOR, 20000, 0.5)
print("%-10d %14.6f %10d" % (seed, err, score(net)))
net = Network(2, 4, 1, seed=1)
net.w1 = [[0.0]*4 for _ in range(2)]; net.b1 = [0.0]*4
net.w2 = [[0.0] for _ in range(4)]; net.b2 = [0.0]
err = net.train(XOR, 20000, 0.5)
print("%-10s %14.6f %10d" % ("all zero", err, score(net)))Practical 4: The Feed Forward Backpropagation Neural Network
A. how many hidden neurons XOR needs
hidden final error correct
none 0.532731 2
1 0.350102 3
2 0.000262 4
3 0.000231 4
4 0.000223 4
B. the learning rate, 4 hidden neurons, 20000 epochs
learning rate final error correct
0.01 0.479992 2
0.10 0.001411 4
0.50 0.000223 4
2.00 0.000050 4
10.00 0.000010 4
50.00 0.499976 3
200.00 1.000000 2
C. the starting weights, 4 hidden neurons, rate 0.5
seed final error correct
1 0.000223 4
2 0.000183 4
3 0.000184 4
4 0.000259 4
5 0.000190 4
all zero 0.342703 3Practical 4: The Feed Forward Backpropagation Neural Network
Three separate lessons, each measured rather than asserted.
A. XOR needs a hidden layer, and needs at least two neurons in it. With no hidden layer the single unit trained for 20,000 epochs and got 2 of 4, which is what guessing gets. With one hidden neuron, 3 of 4. With two or more, 4 of 4. That is the oldest result in the subject: a single layer of weights can only carve the input space with one straight line, and no straight line separates XOR.
B. The learning rate has a floor and a ceiling. At 0.01 the network was still at error 0.48 after 20,000 epochs: it was learning, just far too slowly. From 0.1 to 10 it worked, and on this very small problem the larger rates reached a lower error. At 50 it broke, and at 200 the error sat at exactly 1.000, which is the worst it can be: the steps are so large that each one overshoots and the weights are thrown further out than they started. Too small and it never arrives; too large and it never settles. On real data the useful range is much narrower than it is here, and 0.01 to 0.5 is where to start.
C. The starting weights must be random and different. Five different seeds all reached 4 of 4. Setting every weight to zero gave 3 of 4 and an error of 0.343. The reason is worth a sentence in the journal: if all the hidden neurons start identical, they receive identical deltas, so they stay identical for ever, and a layer of four identical neurons can do exactly what one can. Random starting weights are what break that symmetry.
Step 4: a real dataset, and the mistake that ruins it
MU says "use a given dataset to train the neural network for a specific task". The task: predict Pass or Fail from attendance and practice hours, on the thirty students of chapter 3.
There is one thing to get right first. Attendance runs from 35 to 98 and practice from 2 to 45. Feed those numbers straight into a sigmoid and the sum z is in the hundreds before training starts, the sigmoid is flat out at 1.000, a * (1 - a) is almost zero, and every delta is almost zero, so nothing learns. That is called saturation, and the cure is to squeeze every feature into the same small range first.
Practical 4: The Feed Forward Backpropagation Neural Network
scaled = (value - smallest) / (largest - smallest)
name,attendance,practice,result
Aarav,50,22,Fail
Isha,48,16,Fail
Rohan,88,38,Pass
Sanya,97,39,Pass
Vikram,90,30,Fail
Meera,35,41,Fail
Farhan,45,9,Fail
Nikita,71,8,Fail
Omkar,92,2,Fail
Pooja,97,45,Pass
Rahul,75,15,Fail
Sneha,85,18,Pass
Tejas,79,24,Pass
Urmila,83,34,Pass
Varun,44,23,Fail
Yash,46,37,Pass
Zoya,72,20,Fail
Amit,44,9,Pass
Bhavna,74,3,Fail
Chirag,82,25,Pass
Deepa,94,29,Pass
Eshan,46,27,Fail
Gauri,98,9,Fail
Harsh,89,34,Pass
Ira,97,27,Pass
Jatin,68,28,Pass
Kavya,96,34,Pass
Lalit,38,38,Fail
Manav,63,10,Fail
Neha,41,35,Fail"""The same network on a real dataset, and what scaling does to it."""
import csv, math, random
def sigmoid(z):
if z < -60: return 0.0
if z > 60: return 1.0
return 1 / (1 + math.exp(-z))
class Network:
def __init__(self, n_in, n_hidden, n_out, seed=1):
rng = random.Random(seed)
self.w1 = [[rng.uniform(-1, 1) for _ in range(n_hidden)] for _ in range(n_in)]
self.b1 = [rng.uniform(-1, 1) for _ in range(n_hidden)]
self.w2 = [[rng.uniform(-1, 1) for _ in range(n_out)] for _ in range(n_hidden)]
self.b2 = [rng.uniform(-1, 1) for _ in range(n_out)]
def forward(self, x):
self.a1 = [sigmoid(self.b1[j] + sum(x[i]*self.w1[i][j] for i in range(len(x))))
for j in range(len(self.b1))]
self.a2 = [sigmoid(self.b2[k] + sum(self.a1[j]*self.w2[j][k] for j in range(len(self.a1))))
for k in range(len(self.b2))]
return self.a2
def backward(self, x, t, rate):
out = self.a2
d_out = [(out[k]-t[k])*out[k]*(1-out[k]) for k in range(len(out))]
d_hid = [sum(d_out[k]*self.w2[j][k] for k in range(len(out)))*self.a1[j]*(1-self.a1[j])
for j in range(len(self.a1))]
for j in range(len(self.a1)):
for k in range(len(out)): self.w2[j][k] -= rate*d_out[k]*self.a1[j]
for k in range(len(out)): self.b2[k] -= rate*d_out[k]
for i in range(len(x)):
for j in range(len(self.a1)): self.w1[i][j] -= rate*d_hid[j]*x[i]
for j in range(len(self.a1)): self.b1[j] -= rate*d_hid[j]
with open("students.csv", newline="") as fh:
rows = list(csv.DictReader(fh))
raw = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [1.0 if r["result"] == "Pass" else 0.0 for r in rows]
random.seed(1)
order = list(range(len(rows))); random.shuffle(order)
cut = int(0.7 * len(rows))
tr_i, te_i = order[:cut], order[cut:]
def train_and_score(X, epochs=4000, rate=0.5, seed=1):
net = Network(2, 4, 1, seed=seed)
data = [(X[i], [y[i]]) for i in tr_i]
for _ in range(epochs):
for x, t in data:
net.forward(x); net.backward(x, t, rate)
def acc(idx):
ok = sum(1 for i in idx if (1.0 if net.forward(X[i])[0] >= 0.5 else 0.0) == y[i])
return ok / len(idx)
return acc(tr_i), acc(te_i)
lo = [min(r[k] for r in raw) for k in range(2)]
hi = [max(r[k] for r in raw) for k in range(2)]
scaled = [[(r[k] - lo[k]) / (hi[k] - lo[k]) for k in range(2)] for r in raw]
print("attendance ranges %.0f to %.0f, practice %.0f to %.0f" % (lo[0], hi[0], lo[1], hi[1]))
print("first row raw :", raw[0])
print("first row scaled: [%.4f, %.4f]" % (scaled[0][0], scaled[0][1]))
print()
print("%-22s %10s %9s" % ("features", "train acc", "test acc"))
a, b = train_and_score(raw)
print("%-22s %10.3f %9.3f" % ("raw, 35 to 98", a, b))
a, b = train_and_score(scaled)
print("%-22s %10.3f %9.3f" % ("scaled to 0 and 1", a, b))Practical 4: The Feed Forward Backpropagation Neural Network
attendance ranges 35 to 98, practice 2 to 45
first row raw : [50.0, 22.0]
first row scaled: [0.2381, 0.4651]
features train acc test acc
raw, 35 to 98 0.571 0.444
scaled to 0 and 1 0.952 0.7780.444 on the test set with raw features, 0.778 with scaled ones. The raw run did not merely do worse, it did worse than the baseline of chapter 3. Scaling is not a refinement here; without it the network learned almost nothing.
Note one honest point about the test figure: nine test rows, so 0.778 is 7 of 9 and one row is worth 0.111. The training figure, 0.952, is 20 of 21.
Step 5: the same network from scikit-learn
"""The same job from scikit-learn: MLPClassifier, and scaling again."""
import csv
from sklearn.neural_network import MLPClassifier
from sklearn.preprocessing import MinMaxScaler
from sklearn.model_selection import train_test_split
with open("students.csv", newline="") as fh:
rows = list(csv.DictReader(fh))
X = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [r["result"] for r in rows]
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=1, stratify=y)
def run(name, Atr, Ate):
net = MLPClassifier(hidden_layer_sizes=(4,), activation="logistic",
solver="lbfgs", max_iter=4000, random_state=1)
net.fit(Atr, ytr)
print("%-22s %10.3f %9.3f" % (name, net.score(Atr, ytr), net.score(Ate, yte)))
print("%-22s %10s %9s" % ("features", "train acc", "test acc"))
run("raw", Xtr, Xte)
scaler = MinMaxScaler().fit(Xtr)
run("scaled to 0 and 1", scaler.transform(Xtr), scaler.transform(Xte))features train acc test acc
raw 0.857 0.778
scaled to 0 and 1 0.952 0.778MLPClassifier with hidden_layer_sizes=(4,) and activation="logistic" is the same network: one hidden layer of four sigmoid neurons. Its scaled result, 0.952 training and 0.778 test, is exactly what our own program reached, and that agreement is the best evidence available that the implementation above is correct. Put both in the journal side by side and say so.
Two differences to be honest about. solver="lbfgs" is a cleverer optimiser than the plain gradient descent we wrote, which is why scikit-learn reaches the same answer in 4,000 iterations of a different kind. And its raw-feature run scored 0.857 rather than our 0.571, because lbfgs copes with badly scaled inputs better than plain gradient descent does; it still gained from scaling.
Procedure
- Write
sigmoidand confirm sigmoid(0) is 0.5. - Work one forward pass by hand on a 2-2-1 network with fixed weights, then one backward pass, and check every number against the program.
- Write the
Networkclass: random weights,forward,backward,train. - Train it on XOR and record the error at intervals and the four final outputs.
- Repeat with 0, 1, 2, 3 and 4 hidden neurons and record how many of the four cases each gets right.
- Repeat with learning rates from 0.01 to 200 and record the final error.
- Repeat with five different random seeds, and once with every weight set to zero.
- Scale the student dataset to the range 0 to 1, split it, train, and measure on the held-out rows. Then do it again without scaling and record both.
- Run
MLPClassifieron the same split and compare.
Practical 4: The Feed Forward Backpropagation Neural Network
Observations
| Measured | Value |
|---|---|
| Output of the untrained 2-2-1 network on input 1 0 | 0.518504, target 1 |
| Error before training | 0.115919 |
| Delta at the output | -0.120209 |
| Weights from an input of 0 | unchanged, exactly |
| XOR, 4 hidden, rate 0.5, 20,000 epochs | error 0.000223, 4 of 4 right |
| XOR error after 1 epoch, and after 4,000 | 0.570178, then 0.001644 |
| Hidden neurons | Final error | Correct of 4 |
|---|---|---|
| none | 0.532731 | 2 |
| 1 | 0.350102 | 3 |
| 2 | 0.000262 | 4 |
| 4 | 0.000223 | 4 |
| Learning rate | Final error | Correct of 4 |
|---|---|---|
| 0.01 | 0.479992 | 2 |
| 0.10 | 0.001411 | 4 |
| 0.50 | 0.000223 | 4 |
| 10.00 | 0.000010 | 4 |
| 50.00 | 0.499976 | 3 |
| 200.00 | 1.000000 | 2 |
| Starting weights | Final error | Correct of 4 |
|---|---|---|
| five different random seeds | 0.000183 to 0.000259 | 4 each |
| every weight zero | 0.342703 | 3 |
| Student dataset, 21 training and 9 test rows | train | test |
|---|---|---|
| our network, raw features | 0.571 | 0.444 |
| our network, scaled to 0 and 1 | 0.952 | 0.778 |
| MLPClassifier, raw features | 0.857 | 0.778 |
| MLPClassifier, scaled to 0 and 1 | 0.952 | 0.778 |
Result
A feed forward network trained by backpropagation was implemented from nothing and used on two tasks. One forward and one backward pass were worked with every intermediate value printed and checked. On XOR the network reached an error of 0.000223 and classified all four cases correctly, while the same network with no hidden layer managed 2 of 4 and with all weights initialised to zero managed 3 of 4. The learning rate worked between 0.1 and 10 and failed outside it in both directions. On the thirty-student dataset, scaling the features to the range 0 to 1 raised test accuracy from 0.444 to 0.778 and training accuracy from 0.571 to 0.952, and MLPClassifier on the same scaled split reached the same 0.952 and 0.778.
Where marks are lost
No hidden layer. XOR cannot be learned without one, and the program will happily train for 20,000 epochs and get half of them right.
Unscaled inputs. The sigmoid saturates, the slope goes to zero, and the network learns nothing. This cost 0.33 of test accuracy above.
All weights initialised to zero. Every hidden neuron stays identical to every other one for ever.
Reporting the rounded output only. Print the raw sigmoid value too; it is the confidence.
Practical 4: The Feed Forward Backpropagation Neural Network
Forgetting the bias. A neuron with no bias is a line through the origin, and there is no reason the answer should pass through the origin.
Using the same learning rate everywhere. It has a floor and a ceiling and both were found above. Report the one you used.
Not seeding. Random starting weights mean a different answer every run, and then the journal and the re-run disagree.
Measuring on the training rows. Chapter 3, and it applies here as much as anywhere.
Updating the weights while still computing the deltas. Compute all the deltas from the old weights, then update. Changing w2 before d_hid is computed uses the new weights to apportion the old blame, and the network trains slowly or not at all.
For the journal
Aim; the neuron equation and the sigmoid with its derivative; the network diagram; the hand-worked forward and backward pass with every number; the Network class; the XOR training run with the error at intervals and the four outputs; the three experiment tables for hidden size, learning rate and initialisation; the scaling formula; the student dataset run with and without scaling; the MLPClassifier comparison; all the observation tables; the result.
Quick revision
- A neuron computes z = bias + sum of weight times input, then a = sigmoid(z).
- sigmoid(z) = 1 / (1 + e to the minus z); its slope is
a * (1 - a). - Feed forward: inputs to hidden to output. Backpropagation: the error travels back through the same weights.
- Delta at the output =
(output - target) output (1 - output). - Delta at a hidden neuron = sum over outputs of (their delta times the joining weight), times its own
a * (1 - a). - Update:
weight = weight - rate delta input, where the input is the one that weight carried. - Compute every delta before updating any weight.
- XOR needs a hidden layer: without one, 2 of 4.
- Zero initial weights leave every hidden neuron identical for ever.
- Scale every input to a small range or the sigmoid saturates: 0.444 against 0.778 here.
- Almost all the learning happens in the first few thousand epochs.
- Round the output at 0.5, but report the raw value too.
Questions you must be able to answer
1. Why a sigmoid and not a step function? Because backpropagation needs a slope to work with, and a step has none anywhere except one point where it is undefined. The sigmoid has a smooth slope everywhere, and it can be written as a * (1 - a).
2. Why can a network with no hidden layer not learn XOR? Because one layer of weights draws one straight boundary, and the two classes of XOR cannot be separated by a straight line. Measured here: 2 of 4 after 20,000 epochs.
Practical 4: The Feed Forward Backpropagation Neural Network
3. What is the delta at the output, and what are its parts? (output - target) output (1 - output). The first factor is how wrong the answer is and the rest is the slope of the sigmoid there, so a saturated neuron barely moves.
4. How is the error shared with the hidden layer? Each hidden neuron takes the output's delta multiplied by the weight joining them, and multiplies by its own slope. That is backpropagation in one sentence.
5. What happens if every weight starts at zero? Every hidden neuron computes the same thing, receives the same delta, and is updated identically, so they stay identical for ever. The four-neuron layer behaves like one neuron: 3 of 4 here.
6. Why must the features be scaled? Because a raw sum of large inputs drives the sigmoid to 0 or 1, where its slope is almost zero, so the deltas are almost zero and nothing learns. Scaling raised test accuracy from 0.444 to 0.778.
7. What does the learning rate do, and what are the symptoms of getting it wrong? It sets how far each weight moves per step. Too small and training is still far from the answer after any number of epochs; too large and it overshoots, and the error stops falling or rises. Both were produced above.
8. Your network outputs 0.9895 for a target of 1. Is it wrong? No. A sigmoid never reaches 1. Round at 0.5 to read the class, and quote the raw value as the confidence.
9. Why compute all the deltas before changing any weight? Because the hidden deltas are defined in terms of the weights that carried the values forward. Change those weights first and the blame is apportioned with the wrong numbers.
10. How do you know your own implementation is right? By making it agree with a known one on the same data. Ours and MLPClassifier both reached 0.952 training and 0.778 test on the same scaled split.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.