munotes®

Computer Science Practical 5 Notes | B.Sc. (Computer Science) Semester 5 | Mumbai University | munotes

Get access to whole semester resourcesSemester Pass

Official Notes munotes.in

Computer Science Practical 5

B.SC. (COMPUTER SCIENCE) · SEMESTER 5

Strictly as per the University of Mumbai NEP syllabus in force for B.Sc. (Computer Science)

For B.Sc. (Computer Science) students of the University of Mumbai and all its affiliated colleges

Open the book ↓

munotes.in Third Year

Computer Science Practical 5

Copyright © 2026 munotes.in. All rights reserved.

Written and first published by munotes.in, 2026.

This book is free for individual students to read at munotes.in. No part of it may be reproduced, distributed, stored, translated or used for institutional or classroom purposes in any form without a prior written licence from munotes.in.

Licensing and permissions: contact@munotes.in

The text of statutes and of judgments reproduced in this book is in the public domain under section 52(1)(q) of the Copyright Act 1957. The commentary, arrangement, examples and questions are the original work of munotes.in.

munotes.in is an independent study resource for MU students. It is not affiliated with, endorsed by, or officially connected to the University of Mumbai. Course names and university references describe the students and syllabus the material relates to.

munotes.in

Contents

Module I Artificial Intelligence: ten exercises in Python, from breadth first search to a TensorFlow demonstration

  1. How This Practical Is Examined: the Journal, the 80 Per Cent Rule and the Two-Hour Paper 1
  2. The Python Laboratory from Zero: Running Your First Program and Saving It for the Journal 5
  3. The Dataset, the Split and the Score: How Every Model in This Module Is Measured 11
  4. Practical 1: Breadth First Search and Iterative Deepening Depth First Search 20
  5. Practical 2: A* Search and Recursive Best-First Search 30
  6. Practical 3: Decision Tree Learning 40
  7. Practical 4: The Feed Forward Backpropagation Neural Network 51
  8. Practical 5: Support Vector Machines 62
  9. Practical 6: Adaboost Ensemble Learning 69
  10. Practical 7: The Naive Bayes Classifier 76
  11. Practical 8: K-Nearest Neighbours for Classification and for Regression 84
  12. Practical 9: Association Rule Mining with Apriori 92
  13. Practical 10: A Demonstration with TensorFlow and with an OpenAI-Style Tool 99

Module II Cyber and Information Security: ten exercises, from the classical ciphers to firewall rules on a real Linux

  1. The Security Laboratory from Zero: OpenSSL, the Byte, and What Is Safe to Run 107
  2. Practical 11, Part 1: The Substitution Ciphers 112
  3. Practical 11, Part 2: The Transposition Ciphers 124
  4. Practical 12: RSA Encryption and Decryption 130
  5. Practical 13: Message Authentication Codes 139
  6. Practical 14: Digital Signatures 146
  7. Practical 15: Key Exchange Using Diffie-Hellman 154
  8. Practical 16: IP Security (IPsec) Configuration 162
  9. Practical 17: Web Security with SSL/TLS 169
  10. Practical 18: Intrusion Detection System 177
  11. Practical 19: Malware Analysis and Detection 184
  12. Practical 20: Firewall Configuration and Rule-Based Filtering 191
  13. The Two-Hour Paper: Sitting the Examination 199
munotes.in

Module I

Artificial Intelligence: ten exercises in Python, from breadth first search to a TensorFlow demonstration

munotes.in

Chapter One

How This Practical Is Examined: the Journal, the 80 Per Cent Rule and the Two-Hour Paper

Syllabus topic Module 1 and Module 2, with MU's assessment rules for a 2-credit practical course: "Certified Journal is compulsory for appearing at the time of Practical Exam", "Minimum 80% practical are required to be completed.", "Mid - Term Practical Examination 15 marks", "Final Journal: 5 marks".

Aim

To know, before writing a single line of code, exactly how this paper is marked, what has to be in the journal, and what the two hours in the examination hall contain.

Why this chapter comes first

Every other chapter in this book teaches you to do something. This one tells you what the marks are for. Students lose marks on this paper without making a single mistake in a program, because three of MU's rules are about paperwork and one of them can stop you from sitting the examination at all.

Read this chapter once now and once again a week before the practical examination.

The paper in one table

Marks
InternalMid-term practical examination15
InternalFinal journal5
ExternalQ.1, a practical question on Module 115
ExternalQ.2, a practical question on Module 215
Total50

The external examination runs for two hours and carries 30 of the 50 marks. The other 20 are internal, and they are already decided before you walk into the hall.

The paper is 2 credits and MU allots 60 hours of laboratory time to it, thirty hours to each module.

The two rules that can stop you sitting the paper

"Certified Journal is compulsory for appearing at the time of Practical Exam." Those are MU's words. A journal that has not been signed by your subject teacher is not a certified journal. If you arrive on the day of the practical examination with an uncertified journal, you have a problem that no amount of programming skill will solve.

"Minimum 80% practical are required to be completed." There are twenty practicals in this paper, ten in each module. Eighty per cent of twenty is sixteen. So at least sixteen of the twenty have to be done, written up, and signed.

Sixteen is a minimum, not a target. The external paper sets one question on Module 1 and one on Module 2, and neither you nor your teacher knows in advance which of the ten it will be. A student who has skipped four practicals has skipped four of the twenty things the examiner is choosing from. Do all twenty.

What "completed" means for one practical

A practical is complete when your journal, for that practical, contains all of this:

  1. The practical number and its title, in MU's own words.
  2. Aim. One or two lines saying what the exercise is for.
  3. Theory. What you need to know to do it. Short: this is a practical journal, not a theory answer.
  4. Algorithm or procedure. The steps, numbered.
  5. The program, or the configuration, written out.
  6. The output, exactly as the machine produced it.
  7. Observations. The table, the counts, the comparison. This is the part most students leave out and it is the part the examiner reads first.
  8. Conclusion. One or two lines saying what the run showed.
  9. The date and your teacher's signature.
munotes.in1

How This Practical Is Examined: the Journal, the 80 Per Cent Rule and the Two-Hour Paper

Every chapter of this book ends with a section called For the journal that says what the write-up for that practical has to contain, in that order.

The two-hour paper: how to spend the time

Two hours, two questions, fifteen marks each. That is an hour a question, and the two questions come from two completely different subjects: one is an AI program, the other is a security program or a configuration.

A working plan for the two hours:

MinutesWhat you are doing
0 to 5Read both questions. Decide which one you are more sure of.
5 to 10Write the aim, the algorithm and the theory for the first question on paper.
10 to 45Type the program, run it, fix it, and get output.
45 to 55Write the output and the observation table into the answer.
55 to 105The same for the second question.
105 to 120Check both: is the output written down, is the observation table filled, is the conclusion there.

Start with the question you are surer of. A program that runs and is written up completely is worth more than two half-finished programs.

Never leave the output blank. A program with no output is an unfinished practical, whatever the code looks like. If the program will not run, write down the error the machine gave, say what you think it means, and move on. An honest error with a sentence of diagnosis reads far better than an empty page.

What the examiner is actually looking at

The examiner has fifteen marks to give for one practical question. A fair division, and the one this book is written to satisfy, is roughly this:

PartWeight
The program is correct and runshigh
The output is present and matches the programhigh
The observation table, the comparison, the accuracy figuremedium
Aim, algorithm and theory written outmedium
You can explain what your own program didmedium

The last line is not a separate viva in this paper. Item 6.33 (N) prints no viva component: the whole 15 marks belong to the practical question. But an examiner standing at your machine is going to ask you something about what is on the screen, and your answer is part of how the 15 marks are decided. Every chapter in this book therefore ends with a section called Questions you must be able to answer, and they are the questions that get asked at a machine: what does this line do, why did you choose this value, what happens if I change this.

munotes.in2

How This Practical Is Examined: the Journal, the 80 Per Cent Rule and the Two-Hour Paper

A warning about copying

Twenty students in one laboratory, one internet, and the same twenty practicals. It is obvious to an examiner when four journals carry the same program with the same variable names and the same comment typed in the same place, and it is obvious to the same examiner that the student in front of them cannot explain a line of it.

Write the programs. Run them. Break them on purpose once and see what the error says. That is what makes the questions at the machine easy, and it is the only preparation for this paper that actually works.

The shape of every chapter after this one

Each of the twenty practicals gets one chapter, in MU's printed order, and every chapter is laid out the same way, so that you can read it straight into your journal:

  • Aim. MU's own words for the exercise.
  • What you need to know before you start. The theory, cut to what this exercise needs.
  • The procedure, step by step.
  • The program, complete, and its real output.
  • Observations, filled in from the run.
  • Result. What the run showed.
  • Where marks are lost. The mistakes that actually cost marks in this exercise.
  • For the journal. What to write.
  • Quick revision. Eight or ten lines for the day before.
  • Questions you must be able to answer.

Result

The assessment of this paper was recorded: 20 internal marks made up of a mid-term practical examination for 15 and the final journal for 5; a Semester End Practical Examination of two hours for 30 marks, with one practical question on each module for 15 marks each; a certified journal compulsory for appearing; and a minimum of sixteen of the twenty practicals completed.

Where marks are lost

An uncertified journal. Get each practical signed in the week you do it, not in the week before the examination.

Fewer than sixteen practicals. Count them. Twenty minus four is the line, and it is a hard line.

No observation table. Six of the ten AI practicals and every comparison in the paper ask you to evaluate or compare something. A program that prints an answer but no measurement has done half the exercise.

Output not written into the journal. The program is not the practical. The program plus its output plus what you concluded from it is the practical.

Spending ninety minutes on one question. Two questions, fifteen marks each. Half the paper is in the other module.

For the journal

This chapter is not one of the twenty practicals and does not go in the journal. Use it as the index page of your own practical file: write out the list of twenty practicals, tick each one when it is signed, and count the ticks.

munotes.in3

How This Practical Is Examined: the Journal, the 80 Per Cent Rule and the Two-Hour Paper

Quick revision

  • 50 marks: 20 internal, 30 external.
  • Internal 20: mid-term practical examination 15, final journal 5.
  • External 30: two hours, Q.1 on Module 1 for 15, Q.2 on Module 2 for 15.
  • Certified journal compulsory for appearing at the practical examination.
  • Minimum 80 per cent of the practicals completed, which is sixteen of twenty.
  • 2 credits, 60 hours, thirty hours a module.
  • Ten practicals in Module 1, on Artificial Intelligence. Ten in Module 2, on Cyber and Information Security.
  • This paper prints no separate viva marks. The first-year practical did; this one does not.
  • Every practical write-up: aim, theory, algorithm, program, output, observations, conclusion, signature.

Questions you must be able to answer

1. How many marks is this paper out of, and how are they split? Fifty. Twenty internal, made up of a mid-term practical examination for 15 and the final journal for 5, and thirty external in a two-hour practical examination.

2. How long is the practical examination and what is in it? Two hours. Two practical questions, one on Module 1 and one on Module 2, fifteen marks each.

3. How many practicals must be completed, and out of how many? At least eighty per cent, which is sixteen of the twenty MU prints.

4. What happens if your journal is not certified? MU's rule is that a certified journal is compulsory for appearing at the practical examination.

5. What are the two modules of this paper? Module 1 is practicals based on Artificial Intelligence. Module 2 is practicals based on Cyber and Information Security. Each is thirty hours.

6. Is there a viva in this paper? This paper's printed pattern carries no separate viva marks. The whole fifteen marks of each question belong to the practical. An examiner will still ask you about what is on your screen, and your answer is part of how those fifteen marks are decided.

7. Your program will not run and there are ten minutes left. What do you write? The aim, the algorithm, the program as far as you have it, the exact error message the machine gave, and one sentence saying what you think the error means. Never an empty output.

8. What is the one part of a write-up students most often leave out? The observations: the table, the counts, the accuracy figure or the comparison. It is the part that shows the exercise was actually performed.

Contents This chapter on its own page

munotes.in4

Chapter Two

The Python Laboratory from Zero: Running Your First Program and Saving It for the Journal

Syllabus topic Module 1, "Practical based on Artificial Intelligence", and the skills MU names for it: "dataset handling, model evaluation, cryptographic programming, and security configuration".

Aim

To get a working Python laboratory: the interpreter, an editor, the three libraries this module uses, a first program that runs, and a way of saving a run so that it can go into the journal.

Why Python, and which Python

MU's syllabus for this module does not name a language. It names algorithms, and it names tools: "OpenAI or TensorFlow tools and libraries". Every one of those tools is a Python library, every laboratory in Mumbai teaches this module in Python, and the examination is set in Python. So Python it is.

Version 3.10 or later. Everything in this book runs on 3.12, 3.13 and 3.14, which is what it was checked on. Nothing here needs a feature newer than 3.10. If your laboratory has 3.8, the programs will still run, with one exception noted where it happens.

Check what you have:

$ python3 --version
Python 3.13.12

On Windows the command is usually python rather than python3. If neither answers, Python is not installed; install it from python.org, and on Windows tick Add Python to PATH in the installer, because almost every "python is not recognised" problem in a college laboratory is that one unticked box.

The two ways to run a program, and which one the examination wants

The interactive prompt. Type python3 with no file name and you get >>>. Type an expression, get an answer. It is useful for trying one line out, and useless for a practical, because nothing is saved.

A file. Put the program in a file whose name ends in .py, and run it:

$ python3 practical1.py

The examination wants a file. You have to be able to show the program and its output, and an interactive session leaves you nothing to show. Make one folder for the paper, one file per practical, and name the files after the practicals: practical1_bfs.py, practical2_astar.py, and so on. That folder is your journal in machine-readable form.

Any editor will do. IDLE comes with Python and is enough. VS Code with the Python extension is what most laboratories have. Thonny is the friendliest if you have never used an editor. Notepad works too, as long as you save with the .py ending and not .py.txt, which is the second commonest problem in a college laboratory.

The first program

print("Practical 5, Module 1: Artificial Intelligence")
print("Roll number:", 23)
print("2 + 3 =", 2 + 3)
Practical 5, Module 1: Artificial Intelligence
Roll number: 23
2 + 3 = 5

That is the whole shape of every program in this module: compute something, print it with a label. Label every number you print. An examiner reading 0.857 in your journal has no idea what it is; reading accuracy: 0.857 they do.

munotes.in5

The Python Laboratory from Zero: Running Your First Program and Saving It for the Journal

The three libraries, and the two forms of every exercise

Three libraries are used in this module. Install them once:

$ python3 -m pip install numpy scikit-learn matplotlib
  • numpy gives arrays and the arithmetic over them. Everything numerical rests on it.
  • scikit-learn is the machine-learning library. It has ready-made versions of the decision tree, the neural network, the SVM, Adaboost, Naive Bayes and K-NN: six of the ten practicals in this module.
  • matplotlib draws graphs, which is how you show a learning curve or a decision boundary.

Now the thing that decides how this whole module is written.

MU's wording for eight of the ten practicals is "Implement the ... algorithm". Calling DecisionTreeClassifier().fit(X, y) is not implementing the decision tree algorithm; it is using somebody else's implementation. An examiner who asks "how did your program choose the root of the tree?" gets nothing back from a student who typed three lines of scikit-learn.

So every learning practical in this book is done twice:

  1. From nothing. The algorithm written out in plain Python, from the standard library alone, with the quantities it computes printed as it goes. This is what "implement the algorithm" means, and it is what makes the questions at the machine answerable.
  2. With scikit-learn. Five or six lines, the same dataset, and the two answers compared. This is what you would write if the question says "use a library", and it is what a laboratory machine is set up for.

Do the first one to understand the algorithm. Keep the second one in your journal beside it, because the comparison itself is worth marks: if your own implementation and the library agree, that is evidence your implementation is right.

Reading a dataset from a file

Every dataset in this module is a small comma-separated file. Python reads one with the csv module in the standard library, so nothing has to be installed for this part:

name,attendance,practice,result
Aarav,92,40,Pass
Isha,55,8,Fail
Rohan,78,25,Pass
Sanya,41,5,Fail
Vikram,88,30,Pass
Meera,60,12,Fail
import csv

with open("marks.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))

print("rows read:", len(rows))
print("columns  :", list(rows[0]))
for r in rows[:3]:
    print(r["name"], r["attendance"], r["result"])
rows read: 6
columns  : ['name', 'attendance', 'practice', 'result']
Aarav 92 Pass
Isha 55 Fail
Rohan 78 Pass

Everything a CSV file gives you is text. r["attendance"] is the string "92", not the number 92. Add two of them and you get "9255". Convert what you are going to compute with:

import csv

with open("marks.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))

att = [int(r["attendance"]) for r in rows]
print("as text  :", rows[0]["attendance"] + rows[1]["attendance"])
print("as number:", att[0] + att[1])
print("average  :", sum(att) / len(att))
as text  : 9255
as number: 147
average  : 69.0
munotes.in6

The Python Laboratory from Zero: Running Your First Program and Saving It for the Journal

That first line is the single commonest bug in a first-year machine-learning program, and it does not raise an error. It quietly gives you nonsense.

name,attendance,practice,result
Aarav,92,40,Pass
Isha,55,8,Fail
Rohan,78,25,Pass
Sanya,41,5,Fail
Vikram,88,30,Pass
Meera,60,12,Fail

Randomness, and why your journal needs a seed

Several algorithms in this module use random numbers: the starting weights of a neural network, the way a dataset is shuffled before it is split. Run such a program twice and you get two different answers, and then the number in your journal does not match the number on the screen when the examiner asks you to run it again.

Fix the seed:

import random

random.seed(42)
print("with seed 42 :", [random.randint(1, 100) for _ in range(5)])

random.seed(42)
print("seeded again :", [random.randint(1, 100) for _ in range(5)])

random.seed(7)
print("with seed 7  :", [random.randint(1, 100) for _ in range(5)])
with seed 42 : [82, 15, 4, 95, 36]
seeded again : [82, 15, 4, 95, 36]
with seed 7  : [42, 20, 51, 84, 7]

Seeding does not make the program less random in any way that matters here. It makes the run repeatable, which is what a journal entry has to be. Every program in this book that uses randomness seeds it, and scikit-learn's version of the same idea is the random_state argument, which every model in this book passes.

Saving a run for the journal

Three ways, in order of how much the examiner will like them.

Copy from the terminal. Select the output, copy it, paste it into your file. Fine for short output.

Redirect it to a file. The shell writes the output into a file for you:

$ python3 practical1.py > practical1_output.txt

Print it from inside the program. For a long observation table, have the program write the table itself, so the output in the journal is produced by the program rather than typed by hand:

rows = [("BFS", 5536, 3244), ("IDDFS", 18657, 15)]

print("%-8s %12s %12s" % ("method", "generated", "in memory"))
for name, gen, mem in rows:
    print("%-8s %12d %12d" % (name, gen, mem))
method      generated    in memory
BFS              5536         3244
IDDFS           18657           15

%-8s means a string in eight columns, left aligned. %12d means a whole number in twelve columns, right aligned. Those two are enough to lay out every table in this module, and a table that lines up is read, while one that does not is skipped.

A program that stops with an error, and how to read it

Errors are not a sign that something has gone badly wrong. They are the interpreter telling you, precisely, what it could not do. Read them from the bottom.

numbers = [10, 20, 30]
print("the third :", numbers[2])
print("the fourth:", numbers[3])
munotes.in7

The Python Laboratory from Zero: Running Your First Program and Saving It for the Journal

the third : 30
IndexError: list index out of range

The last line is the type of the error and the message: IndexError, and the list has no index 3. The lines above it, which your own screen will show and this page does not, are the traceback: the line of your program that failed. Two facts, every time: what went wrong, and where. Between them they name the fix.

Procedure

  1. Check the Python version with python3 --version. Install from python.org if it is missing, and on Windows tick Add Python to PATH.
  2. Make one folder for this paper and open it in your editor.
  3. Install the three libraries: python3 -m pip install numpy scikit-learn matplotlib.
  4. Write the first program into a file, run it from the terminal, and confirm the output.
  5. Write the CSV file above, read it with csv.DictReader, and confirm the row count and the columns.
  6. Convert a text column to numbers and see what happens when you do not.
  7. Seed the random number generator and confirm that two seeded runs agree.
  8. Redirect a run into a text file and open the file.
  9. Make an error on purpose and read it from the bottom.

Observations

StepWhat was observed
VersionPython 3.13.12 at the prompt, and every listing in this book also runs on 3.12 and 3.14
First programThree labelled lines printed
CSV read6 rows, 4 columns, values arriving as text
Text against number"92" + "55" gave 9255; int gave 147
Seeded randomnessThe same five numbers from seed 42 on both runs, different numbers from seed 7
ErrorIndexError: list index out of range on index 3 of a list of three

Result

A working Python laboratory was set up: the interpreter checked, the three libraries installed, a program written to a file and run, a comma-separated dataset read and converted, randomness seeded so that runs repeat, a run saved to a text file, and an error raised on purpose and read.

Where marks are lost

Saving the file as .py.txt. Windows hides the real ending by default. If python3 yourfile.py says no such file, that is why.

Treating CSV text as numbers. It does not raise an error. It gives a wrong answer that looks like a right one.

Not seeding. The journal says accuracy 0.857, the machine says 0.893, and the examiner asks which one is true.

Printing a number with no label. Label every figure you print.

Using only scikit-learn. Eight of the ten practicals say "implement the algorithm". Write it out at least once.

Pasting a program you cannot read. The question at the machine is going to be about a line of it.

munotes.in8

The Python Laboratory from Zero: Running Your First Program and Saving It for the Journal

For the journal

Aim; the version of Python on your machine; the install command for the three libraries; the first program and its output; the CSV file and the program that reads it, with the row count and column names; the two lines showing text addition against numeric addition; the seeded random output twice; the error and its last line; the observation table above; the result.

Quick revision

  • Python 3.10 or later. This book was checked on 3.12, 3.13 and 3.14.
  • A practical is a .py file, never an interactive session.
  • python3 -m pip install numpy scikit-learn matplotlib installs everything this module needs.
  • numpy for arrays, scikit-learn for the models, matplotlib for the graphs.
  • Every learning practical is done twice: from nothing, and with scikit-learn.
  • csv.DictReader reads a dataset. Everything it gives you is text; convert with int or float.
  • random.seed(n) makes a run repeatable. scikit-learn's version is random_state.
  • python3 prog.py > out.txt saves a run.
  • %-8s and %12d lay out a table that lines up.
  • Read an error from the bottom: the type and message on the last line, the place above it.

Questions you must be able to answer

1. Why does this module use Python? Because the tools MU names for it, TensorFlow and the OpenAI libraries, are Python libraries, and because every algorithm in the module has a ready implementation in scikit-learn as well as being short enough to write out by hand.

2. Why write an algorithm out when scikit-learn already has it? Because MU's wording for eight of the ten practicals is "implement the algorithm", and because you cannot explain a choice you did not make. The library version goes in the journal beside your own as a check on it.

3. What does csv.DictReader give you for a column of numbers? Text. "92", not 92. Convert it with int or float before computing with it, or the arithmetic silently joins strings instead of adding numbers.

4. Why seed the random number generator? So the run repeats. The journal has to show the same figure the machine will show when the examiner asks you to run it again.

5. What is random_state in scikit-learn? The same idea as random.seed, passed to a model or to a data split, so that the result is repeatable.

6. How do you save a program's output for the journal? Copy it from the terminal for short output, or redirect it with > out.txt, or have the program print the table itself.

7. Where do you look first in an error? The last line: it carries the type of the error and the message. The lines above it say where.

munotes.in9

The Python Laboratory from Zero: Running Your First Program and Saving It for the Journal

8. What does %-8s %12d do? Prints a string left aligned in eight columns and a whole number right aligned in twelve, which is how the observation tables in this module are made to line up.

Contents This chapter on its own page

munotes.in10

Chapter Three

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

Syllabus topic Module 1, the evaluation every learning practical asks for: "Evaluate the accuracy and effectiveness of the decision tree on test data.", "Evaluate the performance of the trained network on test data.", "Evaluate the performance of the SVM model on test data and analyze the results.", "Train the ensemble model on a given dataset and evaluate its performance.", "Evaluate the accuracy of the model on test data and analyze the results.", "Evaluate the accuracy or error of the predictions and analyze the results.", and MU's own "performance evaluation".

Aim

To learn the one thing six of the ten practicals in this module ask for and none of them explains: how a model is measured. The training set and the test set, the confusion matrix, accuracy, precision, recall and F1, the error of a prediction that is a number, and the baseline a model has to beat.

Why this chapter exists

Read MU's Module 1 and count the sentences that say "evaluate". There are six of them. "Evaluate the accuracy and effectiveness of the decision tree on test data." "Evaluate the performance of the trained network on test data." And so on through the SVM, the ensemble, Naive Bayes and K-NN.

Not one of those sentences says what accuracy is, what test data is, or how the data becomes test data. Six practicals ask for a measurement the syllabus never defines, and a student who has not been taught it either prints a number with no name or, much worse, prints a number measured on the wrong data.

That is what this chapter is for. Learn it once here, use it in six chapters.

The dataset this module uses

Most of this module runs on one small dataset of thirty students. Two columns are what we know about a student, the features, and one is what we are trying to predict, the label.

name,attendance,practice,result
Aarav,50,22,Fail
Isha,48,16,Fail
Rohan,88,38,Pass
Sanya,97,39,Pass
Vikram,90,30,Fail
Meera,35,41,Fail
Farhan,45,9,Fail
Nikita,71,8,Fail
Omkar,92,2,Fail
Pooja,97,45,Pass
Rahul,75,15,Fail
Sneha,85,18,Pass
Tejas,79,24,Pass
Urmila,83,34,Pass
Varun,44,23,Fail
Yash,46,37,Pass
Zoya,72,20,Fail
Amit,44,9,Pass
Bhavna,74,3,Fail
Chirag,82,25,Pass
Deepa,94,29,Pass
Eshan,46,27,Fail
Gauri,98,9,Fail
Harsh,89,34,Pass
Ira,97,27,Pass
Jatin,68,28,Pass
Kavya,96,34,Pass
Lalit,38,38,Fail
Manav,63,10,Fail
Neha,41,35,Fail

attendance is a percentage. practice is the number of laboratory hours the student put in. result is Pass or Fail. Thirty rows is small on purpose: every intermediate number a program computes can be checked by hand, which is exactly what an examiner may ask you to do.

import csv

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))

data = [([int(r["attendance"]), int(r["practice"])], r["result"]) for r in rows]

passes = sum(1 for _, y in data if y == "Pass")
print("rows     :", len(data))
print("features :", 2, "(attendance, practice)")
print("Pass     :", passes)
print("Fail     :", len(data) - passes)
print("first row:", data[0])
rows     : 30
features : 2 (attendance, practice)
Pass     : 14
Fail     : 16
first row: ([50, 22], 'Fail')

Fourteen Pass and sixteen Fail. Keep those two numbers in mind: they are what the baseline below is built from.

The training set and the test set, and why the split is the whole idea

A model learns from data and is then asked about data. If you ask it about the same rows it learned from, you learn nothing at all, because a model is allowed to remember. A model that memorises all thirty rows answers all thirty correctly and knows nothing.

munotes.in11

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

So the rows are split in two:

  • the training set, which the model learns from, and
  • the test set, which it has never seen, and which is the only honest place to measure it.

The usual split is 70 or 80 per cent for training and the rest for testing. Two rules matter.

Shuffle before you split. Take the last six rows of our file as the test set and you get five Fail and one Pass, because the file happens to end that way. A test set that does not look like the data tells you nothing about the data.

Seed the shuffle. Otherwise your journal and the examiner's re-run disagree, for the reason the previous chapter gives.

import csv
import random

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
data = [([int(r["attendance"]), int(r["practice"])], r["result"]) for r in rows]

def split(data, seed, frac=0.7):
    order = list(range(len(data)))
    random.seed(seed)
    random.shuffle(order)
    cut = int(frac * len(data))
    return [data[i] for i in order[:cut]], [data[i] for i in order[cut:]]

def balance(rows):
    p = sum(1 for _, y in rows if y == "Pass")
    return "%d rows, %d Pass, %d Fail" % (len(rows), p, len(rows) - p)

train, test = split(data, 1)
print("training :", balance(train))
print("test     :", balance(test))
training : 21 rows, 9 Pass, 12 Fail
test     : 9 rows, 5 Pass, 4 Fail

Twenty-one rows to learn from, nine to be tested on, chosen by a seeded shuffle. That is the split every later chapter in this module uses.

How much the split itself moves the answer

The whole file is 14 Pass to 16 Fail, which is 47 per cent Pass. Seed 1 put 5 Pass into the nine test rows, which is 56 per cent. Change the seed and the test set changes with it:

import csv
import random

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
data = [([int(r["attendance"]), int(r["practice"])], r["result"]) for r in rows]

def split(data, seed, frac=0.7):
    order = list(range(len(data)))
    random.seed(seed)
    random.shuffle(order)
    cut = int(frac * len(data))
    return [data[i] for i in order[:cut]], [data[i] for i in order[cut:]]

print("%-6s %12s %12s" % ("seed", "test Pass", "test Fail"))
for seed in range(1, 9):
    _, test = split(data, seed)
    p = sum(1 for _, y in test if y == "Pass")
    print("%-6d %12d %12d" % (seed, p, len(test) - p))
seed      test Pass    test Fail
1                 5            4
2                 4            5
3                 5            4
4                 6            3
5                 5            4
6                 5            4
7                 6            3
8                 4            5

Four to six Pass out of nine, purely from the seed. On nine test rows one row is 0.111 of the accuracy, so two models that differ by 0.05 on a test set this small have not been distinguished at all, and a write-up that ranks them is reading noise. Say the size of your test set whenever you quote a score.

munotes.in12

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

The fix for the balance, which scikit-learn calls stratify, keeps the proportion of each class the same in both halves. The fix for the noise is a bigger test set, or cross validation: split the data five ways, train five times, test on the part that was held out each time, and average the five scores. Cross validation is out of this syllabus, but the sentence "on nine test rows this difference is not measurable" is always worth writing.

The baseline: the number your model has to beat

Before measuring any model, measure the stupidest possible one: the model that ignores its input and answers whichever class was commonest in the training data.

import csv
import random
from collections import Counter

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
data = [([int(r["attendance"]), int(r["practice"])], r["result"]) for r in rows]

def split(data, seed, frac=0.7):
    order = list(range(len(data)))
    random.seed(seed)
    random.shuffle(order)
    cut = int(frac * len(data))
    return [data[i] for i in order[:cut]], [data[i] for i in order[cut:]]

train, test = split(data, 1)
commonest = Counter(y for _, y in train).most_common(1)[0][0]
right = sum(1 for _, y in test if y == commonest)

print("commonest class in training :", commonest)
print("it answers                  :", commonest, "every time")
print("baseline on the test set    : %d of %d right, accuracy %.3f"
      % (right, len(test), right / len(test)))
commonest class in training : Fail
it answers                  : Fail every time
baseline on the test set    : 4 of 9 right, accuracy 0.444

Read that carefully, because it is not the answer most students expect. The commonest class in the training set is Fail, 12 of 21. But the test set has 5 Pass and 4 Fail, so answering Fail every time is right only 4 times out of 9: an accuracy of 0.444, which is worse than a coin.

Two things follow, and both belong in a write-up.

The baseline is decided on training data and measured on test data, like everything else. You are not allowed to look at the test labels to choose what your stupid model says, because you are not allowed to look at them at all.

On a small test set a baseline can come out below 0.5. That is not a bug. It is what a nine-row test set does, and it is another reason to say how big the test set was.

Compute the baseline anyway, every time. It is four lines, and it is the difference between a number and a result.

The confusion matrix

Accuracy on its own hides which mistakes a model makes, and the two mistakes are not equally serious. Telling a student who will fail that they will pass is not the same error as telling a student who will pass that they will fail.

munotes.in13

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

Pick one class as positive. Here, Pass. Then every prediction falls into one of four boxes:

Predicted PassPredicted Fail
Actually PassTrue Positive (TP)False Negative (FN)
Actually FailFalse Positive (FP)True Negative (TN)

Read the names from right to left: a False Positive is a prediction of Positive that is false.

def confusion(actual, predicted, positive):
    tp = fp = fn = tn = 0
    for a, p in zip(actual, predicted):
        if p == positive and a == positive:
            tp += 1
        elif p == positive and a != positive:
            fp += 1
        elif p != positive and a == positive:
            fn += 1
        else:
            tn += 1
    return tp, fp, fn, tn

actual = ["Pass", "Fail", "Pass", "Fail", "Fail", "Pass", "Fail", "Pass", "Fail"]
guess = ["Pass", "Fail", "Fail", "Fail", "Pass", "Pass", "Fail", "Pass", "Fail"]

tp, fp, fn, tn = confusion(actual, guess, "Pass")
print("             predicted Pass  predicted Fail")
print("actual Pass  %14d  %14d" % (tp, fn))
print("actual Fail  %14d  %14d" % (fp, tn))
print()
print("TP %d  FP %d  FN %d  TN %d  total %d" % (tp, fp, fn, tn, tp + fp + fn + tn))
             predicted Pass  predicted Fail
actual Pass               3               1
actual Fail               1               4

TP 3  FP 1  FN 1  TN 4  total 9

The four numbers add up to the number of test rows, always. If they do not, the counting is wrong.

Accuracy, precision, recall and F1

All four come out of those same four numbers.

Accuracy is how often the model was right, over everything.

accuracy = (TP + TN) / (TP + FP + FN + TN)

Precision is how often it was right when it said Pass. Of the students you predicted would pass, what fraction did?

precision = TP / (TP + FP)

Recall is how many of the real passes it found. Of the students who actually passed, what fraction did you catch?

recall = TP / (TP + FN)

Precision and recall pull against each other. Predict Pass for everybody and recall is 1.00 and precision is poor. Predict Pass only for the one student you are certain of and precision is 1.00 and recall is terrible.

F1 is the single number that refuses to let you cheat either way. It is the harmonic mean of precision and recall, which is low whenever either of them is low.

F1 = 2 precision recall / (precision + recall)

def confusion(actual, predicted, positive):
    tp = fp = fn = tn = 0
    for a, p in zip(actual, predicted):
        if p == positive and a == positive:
            tp += 1
        elif p == positive:
            fp += 1
        elif a == positive:
            fn += 1
        else:
            tn += 1
    return tp, fp, fn, tn

def scores(actual, predicted, positive):
    tp, fp, fn, tn = confusion(actual, predicted, positive)
    total = tp + fp + fn + tn
    acc = (tp + tn) / total
    prec = tp / (tp + fp) if tp + fp else 0.0
    rec = tp / (tp + fn) if tp + fn else 0.0
    f1 = 2 * prec * rec / (prec + rec) if prec + rec else 0.0
    return acc, prec, rec, f1

actual = ["Pass", "Fail", "Pass", "Fail", "Fail", "Pass", "Fail", "Pass", "Fail"]
models = [("the model", ["Pass", "Fail", "Fail", "Fail", "Pass", "Pass", "Fail", "Pass", "Fail"]),
          ("always Fail", ["Fail"] * 9),
          ("always Pass", ["Pass"] * 9)]

print("%-12s %9s %10s %7s %6s" % ("model", "accuracy", "precision", "recall", "F1"))
for name, guess in models:
    acc, prec, rec, f1 = scores(actual, guess, "Pass")
    print("%-12s %9.3f %10.3f %7.3f %6.3f" % (name, acc, prec, rec, f1))
munotes.in14

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

model         accuracy  precision  recall     F1
the model        0.778      0.750   0.750  0.750
always Fail      0.556      0.000   0.000  0.000
always Pass      0.444      0.444   1.000  0.615

Read the bottom row, because it is the whole reason F1 exists. "Always Pass" has perfect recall: it catches every student who really passed, because it says Pass about everybody. Its precision is 0.444, because more than half the students it named are wrong. And its F1 is 0.615, not the 0.722 an ordinary average of 0.444 and 1.000 would give, because the harmonic mean is dragged towards the smaller number.

That is the point. A model cannot buy a good F1 by making one of the two perfect at the other's cost, and a model that answers the same thing every time is caught by every column except the one it cheated.

Print these as a table with a heading row, as the program above does. Five numbers on one long line with the labels squeezed between them is how a figure gets read against the wrong heading, and that mistake in a journal costs marks that the program had already earned.

Note one more thing before leaving this table. "Always Fail" scores 0.556 here, and on the real seed-1 test set earlier in this chapter the same stupid model scored 0.444. Both are right: this table's nine rows are four Pass and five Fail, and the seed-1 split's nine rows are five Pass and four Fail. A baseline belongs to one test set. Quote the two together or quote neither.

munotes.in15

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

When the answer is a number, not a class

K-NN, which MU sets as "classification or regression", can predict a number instead of a class: not Pass or Fail but a mark out of 100. Accuracy makes no sense then, because a prediction of 71 when the truth is 72 is not "wrong".

Two measures do the work.

Mean absolute error, the average size of the mistake, in the same units as the answer.

MAE = (1/n) * sum of |actual - predicted|

Root mean squared error, which squares each mistake before averaging, so one big miss counts for much more than several small ones.

RMSE = square root of ((1/n) * sum of (actual - predicted)^2)

import math

actual = [72, 55, 88, 41, 66, 79]
guess = [70, 58, 84, 45, 65, 95]

n = len(actual)
errors = [a - p for a, p in zip(actual, guess)]
mae = sum(abs(e) for e in errors) / n
rmse = math.sqrt(sum(e * e for e in errors) / n)

print("%-8s %8s %8s %8s" % ("actual", "guess", "error", "squared"))
for a, p, e in zip(actual, guess, errors):
    print("%-8d %8d %8d %8d" % (a, p, e, e * e))
print()
print("MAE  : %.3f" % mae)
print("RMSE : %.3f" % rmse)
actual      guess    error  squared
72             70        2        4
55             58       -3        9
88             84        4       16
41             45       -4       16
66             65        1        1
79             95      -16      256

MAE  : 5.000
RMSE : 7.095

Five of the six predictions are within four marks. The sixth is out by sixteen, and the two measures treat that very differently: MAE is 5.000 and RMSE is 7.095, because squaring turned that one miss into 256 of the 302 total. RMSE is the one to quote when a big mistake matters more than several small ones, which for a mark prediction it usually does.

The same measurements from scikit-learn

Everything above is four lines of scikit-learn. Write your own once, so you know what the four lines do, then use theirs.

from sklearn.metrics import (accuracy_score, precision_score, recall_score,
                             f1_score, confusion_matrix)

actual = ["Pass", "Fail", "Pass", "Fail", "Fail", "Pass", "Fail", "Pass", "Fail"]
guess = ["Pass", "Fail", "Fail", "Fail", "Pass", "Pass", "Fail", "Pass", "Fail"]

print("accuracy :", accuracy_score(actual, guess))
print("precision:", precision_score(actual, guess, pos_label="Pass"))
print("recall   :", recall_score(actual, guess, pos_label="Pass"))
print("F1       :", f1_score(actual, guess, pos_label="Pass"))
print("matrix, labels [Fail, Pass]:")
print(confusion_matrix(actual, guess, labels=["Fail", "Pass"]))
accuracy : 0.7777777777777778
precision: 0.75
recall   : 0.75
F1       : 0.75
matrix, labels [Fail, Pass]:
[[4 1]
 [1 3]]

The four figures are the ones our own program computed, which is the check that our own program is right.

confusion_matrix lays the matrix out the other way round from the table earlier in this chapter: rows are the true label in the order you give in labels, columns are the prediction, so with labels=["Fail", "Pass"] the top-left cell is TN, not TP. Always pass labels and always say in your write-up which corner is which, because half the marks lost on a confusion matrix are lost to reading it upside down.

munotes.in16

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

And the split, in one line:

import csv
from sklearn.model_selection import train_test_split

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
X = [[int(r["attendance"]), int(r["practice"])] for r in rows]
y = [r["result"] for r in rows]

Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=1,
                                      stratify=y)
print("train:", len(Xtr), "rows,", ytr.count("Pass"), "Pass")
print("test :", len(Xte), "rows,", yte.count("Pass"), "Pass")
train: 21 rows, 10 Pass
test : 9 rows, 4 Pass

stratify=y is what keeps the proportion of Pass the same in both halves: 10 of 21 and 4 of 9, against the file's 14 of 30. Pass random_state or the split changes on every run and so does every figure in your journal.

Procedure

  1. Save students.csv and read it, converting the two feature columns to whole numbers.
  2. Count the rows and the two classes.
  3. Shuffle with a seed, split 70 to 30, and print the balance of both halves.
  4. Compute the baseline: the accuracy of always answering the commoner class on your test set.
  5. Count TP, FP, FN and TN for a set of predictions and print the confusion matrix with its headings.
  6. Compute accuracy, precision, recall and F1 from those four counts, and print them as a table with a heading row.
  7. Compute MAE and RMSE for a set of numeric predictions, with the per-row error and squared error shown.
  8. Repeat steps 5 to 7 with sklearn.metrics and confirm the figures agree with your own.

Observations

MeasuredValue
Rows in the dataset30
Class balance14 Pass, 16 Fail
Split at 70 per cent, seed 121 training, 9 test
Test balance at seed 15 Pass, 4 Fail
Test balance over seeds 1 to 8between 4 and 6 Pass out of 9
Test balance, stratified4 Pass, 5 Fail
Baseline on the seed-1 test set0.444, because training says Fail and the test set has more Pass
The example modelaccuracy 0.778, precision 0.750, recall 0.750, F1 0.750
Always Passaccuracy 0.444, precision 0.444, recall 1.000, F1 0.615
Always Failaccuracy 0.556, precision 0.000, recall 0.000, F1 0.000
Numeric predictionsMAE 5.000, RMSE 7.095 (one error of 16 contributed 256 of the 302)
scikit-learn agreementaccuracy, precision, recall and F1 identical to our own

Result

A measurement kit for this whole module was built and checked: a seeded 70 to 30 split, a baseline, a confusion matrix, accuracy, precision, recall and F1 from the four counts, and MAE and RMSE for numeric answers. Every figure our own code produced was reproduced by sklearn.metrics on the same inputs.

munotes.in17

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

Where marks are lost

Measuring on the training set. The commonest and the most serious. A model tested on what it learned from can score 1.00 and be useless. Say in your journal which rows the score was measured on.

No baseline. 0.70 accuracy sounds like a result until the reader learns that answering Fail every time scores 0.667.

Quoting accuracy alone on an unbalanced dataset. Report the confusion matrix too, or precision and recall.

Reading the confusion matrix upside down. sklearn's rows are the truth and its columns are the prediction, and the order of the classes is whatever labels says. State which corner is TP.

Not seeding the split. Then no number in the journal can be reproduced.

Reading a number off a long line against the wrong label. Print a table with a heading row. This chapter contains a worked example of getting that wrong.

Using accuracy for a numeric prediction. Use MAE or RMSE, and say which.

For the journal

Aim; the dataset with its row count and class balance; the split program with the sizes and balance of both halves; the baseline and its figure; the confusion matrix with headings and the four counts; the metrics table with its heading row; the MAE and RMSE table with the per-row squared error; the sklearn.metrics run confirming the same figures; the observation table above; the result.

Quick revision

  • Train on the training set. Measure on the test set. Never the other way round.
  • Shuffle before splitting, and seed the shuffle.
  • stratify=y keeps the class proportions in both halves.
  • Baseline: always answer the commoner class. Our stratified test set gives 0.556.
  • TP, FP, FN, TN. A False Positive is a prediction of Positive that is false.
  • accuracy = (TP + TN) / total.
  • precision = TP / (TP + FP): how often "yes" was right.
  • recall = TP / (TP + FN): how much of the real "yes" was found.
  • F1 = 2 precision recall / (precision + recall), the harmonic mean, low if either is low.
  • MAE is the average size of the error. RMSE squares first, so one big miss dominates.
  • sklearn.metrics has all of them; confusion_matrix has truth as rows and prediction as columns.

Questions you must be able to answer

1. Why must a model be tested on data it has not seen? Because a model is allowed to remember. One that memorises the training rows answers all of them correctly and has learned nothing, so a score on the training set measures memory, not learning.

2. What is the baseline, and why compute it? The score of the stupidest model, which is to answer the commonest class every time. It is what any real model has to beat, and without it a number like 0.70 cannot be judged.

munotes.in18

The Dataset, the Split and the Score: How Every Model in This Module Is Measured

3. Define precision and recall in one sentence each. Precision is the fraction of the rows you predicted positive that really were positive. Recall is the fraction of the really positive rows that you predicted positive.

4. Which do you care about more when predicting that a student will fail? Recall, if the purpose is to reach every student who is in trouble: missing one is worse than warning one extra. Precision, if a warning is expensive. Say which you chose and why.

5. Why is F1 the harmonic mean and not the ordinary average? Because the harmonic mean is dragged down by the smaller of the two, so a model cannot score well by making precision or recall perfect at the other's expense. Precision 0.444 with recall 1.000 gives F1 0.615, not 0.722.

6. What is in each of the four cells of a confusion matrix? TP, correct positives. FP, predicted positive but actually negative. FN, predicted negative but actually positive. TN, correct negatives. They add to the number of test rows.

7. When would you report RMSE instead of accuracy? When the prediction is a number rather than a class, as in K-NN regression, and especially when one large error matters more than several small ones.

8. What does random_state do to train_test_split, and why does it matter here? It fixes the shuffle, so the same rows go into the test set on every run, so the figure in your journal is the figure the examiner will see.

9. Our test set has nine rows. What is the smallest change in accuracy you can measure? One row in nine, which is 0.111. A difference of 0.05 between two models on nine test rows is not a difference at all, and a write-up should say so rather than rank them.

Contents This chapter on its own page

munotes.in19

Chapter Six

Practical 3: Decision Tree Learning

Syllabus topic Module 1, "Decision Tree Learning: Implement the Decision Tree Learning algorithm to build a decision tree for a given dataset. Evaluate the accuracy and effectiveness of the decision tree on test data. Visualize and interpret the generated decision tree."

Aim

To implement the Decision Tree Learning algorithm, to build a tree from a dataset, to measure it on data it has not seen, and to draw and read the tree.

What you need to know before you start

A decision tree is a set of questions arranged so that each answer leads to the next question, and the last answer is a prediction. Read a path from the top to a leaf and you have a rule in plain English.

Building one is a single idea repeated: at every node, ask the question that tells you the most. The only thing needing definition is "the most", and the answer is borrowed from information theory.

Entropy measures how mixed a set of labels is, in bits. A set that is all Yes, or all No, has entropy 0: you learn nothing by being told a label you could already predict. A set that is half and half has entropy 1: one full bit of surprise.

H(S) = - sum over each class c of p(c) * log2 p(c)

Information gain is how much a question reduces that mixture: the entropy before, minus the average entropy of the groups the question splits the set into, each weighted by its size.

Gain(S, A) = H(S) - sum over each value v of A of (|S_v| / |S|) * H(S_v)

The algorithm, which is called ID3, is then three lines:

  1. If every row in this group has the same label, make a leaf with that label.
  2. If there are no questions left, make a leaf with the commoner label.
  3. Otherwise pick the question with the highest gain, split on it, and repeat on each group.

The dataset

Fourteen days, four things known about each, and whether a game was played. It is the dataset most decision-tree examples use, which is exactly why it is used here: you can check every number below against any other source.

outlook,temperature,humidity,windy,play
Sunny,Hot,High,No,No
Sunny,Hot,High,Yes,No
Overcast,Hot,High,No,Yes
Rainy,Mild,High,No,Yes
Rainy,Cool,Normal,No,Yes
Rainy,Cool,Normal,Yes,No
Overcast,Cool,Normal,Yes,Yes
Sunny,Mild,High,No,No
Sunny,Cool,Normal,No,Yes
Rainy,Mild,Normal,No,Yes
Sunny,Mild,Normal,Yes,Yes
Overcast,Mild,High,Yes,Yes
Overcast,Hot,Normal,No,Yes
Rainy,Mild,High,Yes,No

Step 1: the first split, by hand and by machine

Nine days ended in Yes and five in No, so the entropy of the whole set is

H = -(9/14) log2(9/14) - (5/14) log2(5/14)

= -(0.6429) (-0.6374) - (0.3571) (-1.4854)

= 0.4097 + 0.5305

= 0.9403 bits

Now compute that for all four questions and pick the winner.

"""The first split of a decision tree, computed and printed."""
import csv
import math
from collections import Counter

with open("weather.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))

TARGET = "play"
FEATURES = ["outlook", "temperature", "humidity", "windy"]

def entropy(rows):
    counts = Counter(r[TARGET] for r in rows)
    total = len(rows)
    bits = 0.0
    for n in counts.values():
        p = n / total
        bits -= p * math.log2(p)
    return bits

def split_on(rows, feature):
    groups = {}
    for r in rows:
        groups.setdefault(r[feature], []).append(r)
    return groups

def gain(rows, feature):
    before = entropy(rows)
    total = len(rows)
    after = 0.0
    for value, group in split_on(rows, feature).items():
        after += (len(group) / total) * entropy(group)
    return before - after, after

counts = Counter(r[TARGET] for r in rows)
print("rows      :", len(rows))
print("play = Yes:", counts["Yes"])
print("play = No :", counts["No"])
print("entropy of the whole set: %.4f bits" % entropy(rows))
print()
print("%-12s %8s %10s %10s" % ("feature", "values", "remaining", "gain"))
for f in FEATURES:
    g, after = gain(rows, f)
    print("%-12s %8d %10.4f %10.4f" % (f, len(split_on(rows, f)), after, g))
print()
best = max(FEATURES, key=lambda f: gain(rows, f)[0])
print("root of the tree:", best)
print()
for value, group in sorted(split_on(rows, best).items()):
    c = Counter(r[TARGET] for r in group)
    print("  %-9s %2d rows  Yes %d  No %d  entropy %.4f"
          % (value, len(group), c["Yes"], c["No"], entropy(group)))
munotes.in40

Practical 3: Decision Tree Learning

rows      : 14
play = Yes: 9
play = No : 5
entropy of the whole set: 0.9403 bits

feature        values  remaining       gain
outlook             3     0.6935     0.2467
temperature         3     0.9111     0.0292
humidity            2     0.7885     0.1518
windy               2     0.8922     0.0481

root of the tree: outlook

  Overcast   4 rows  Yes 4  No 0  entropy 0.0000
  Rainy      5 rows  Yes 3  No 2  entropy 0.9710
  Sunny      5 rows  Yes 2  No 3  entropy 0.9710

outlook wins by a distance: it removes 0.2467 bits of uncertainty against humidity's 0.1518, and temperature removes almost nothing. So outlook is the root of the tree.

Look at the three groups it makes. Overcast is pure: four days, all Yes, entropy exactly 0. That branch is finished before it starts, and the tree will never ask another question about an overcast day. Rainy and Sunny both come out at 0.9710 bits, so both need another question.

Step 2: the whole algorithm

"""Practical 3: Decision Tree Learning (ID3) written from nothing."""
import csv
import math
from collections import Counter

TARGET = "play"

def entropy(rows):
    counts = Counter(r[TARGET] for r in rows)
    total = len(rows)
    return -sum((n / total) * math.log2(n / total) for n in counts.values())

def split_on(rows, feature):
    groups = {}
    for r in rows:
        groups.setdefault(r[feature], []).append(r)
    return groups

def gain(rows, feature):
    total = len(rows)
    after = sum(len(g) / total * entropy(g) for g in split_on(rows, feature).values())
    return entropy(rows) - after

def majority(rows):
    return Counter(r[TARGET] for r in rows).most_common(1)[0][0]

def id3(rows, features, values_of, depth=0, max_depth=None):
    labels = set(r[TARGET] for r in rows)
    if len(labels) == 1:                    # pure: nothing left to ask
        return labels.pop()
    if not features or (max_depth is not None and depth >= max_depth):
        return majority(rows)               # out of questions: vote
    best = max(features, key=lambda f: gain(rows, f))
    node = {"feature": best, "default": majority(rows), "branches": {}}
    rest = [f for f in features if f != best]
    groups = split_on(rows, best)
    for value in values_of[best]:
        group = groups.get(value)
        if not group:                       # a value no row here has
            node["branches"][value] = majority(rows)
        else:
            node["branches"][value] = id3(group, rest, values_of, depth + 1, max_depth)
    return node

def show(node, indent="", label=None):
    if not isinstance(node, dict):
        print("%s%s-> %s" % (indent, (label + " ") if label else "", node))
        return
    if label:
        print("%s%s" % (indent, label))
        indent += "    "
    print("%s[%s?]" % (indent, node["feature"]))
    for value, child in node["branches"].items():
        show(child, indent + "    ", "%s = %s" % (node["feature"], value))

def rules(node, sofar=()):
    if not isinstance(node, dict):
        return [(list(sofar), node)]
    out = []
    for value, child in node["branches"].items():
        out += rules(child, sofar + ("%s = %s" % (node["feature"], value),))
    return out

def classify(node, row):
    while isinstance(node, dict):
        node = node["branches"].get(row[node["feature"]], node["default"])
    return node

with open("weather.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
FEATURES = ["outlook", "temperature", "humidity", "windy"]
values_of = {f: sorted(set(r[f] for r in rows)) for f in FEATURES}

tree = id3(rows, FEATURES, values_of)
print("the tree")
show(tree)
print()
print("as rules")
for conditions, answer in rules(tree):
    print("  IF %s THEN play = %s" % (" AND ".join(conditions), answer))
print()
right = sum(1 for r in rows if classify(tree, r) == r[TARGET])
print("on the 14 rows it learned from : %d of %d right" % (right, len(rows)))

TEST = [
    {"outlook": "Sunny", "temperature": "Cool", "humidity": "High", "windy": "Yes", "play": "No"},
    {"outlook": "Overcast", "temperature": "Mild", "humidity": "Normal", "windy": "Yes", "play": "Yes"},
    {"outlook": "Rainy", "temperature": "Hot", "humidity": "Normal", "windy": "No", "play": "Yes"},
    {"outlook": "Rainy", "temperature": "Cool", "humidity": "High", "windy": "Yes", "play": "No"},
    {"outlook": "Sunny", "temperature": "Hot", "humidity": "Normal", "windy": "No", "play": "Yes"},
]
print()
print("%-9s %-5s %-7s %-6s %-9s %s" % ("outlook", "temp", "humidity", "windy", "predicted", "actual"))
right = 0
for r in TEST:
    p = classify(tree, r)
    right += (p == r[TARGET])
    print("%-9s %-5s %-7s %-6s %-9s %s"
          % (r["outlook"], r["temperature"], r["humidity"], r["windy"], p, r[TARGET]))
print()
print("on 5 unseen rows              : %d of %d right, accuracy %.3f"
      % (right, len(TEST), right / len(TEST)))
munotes.in41

Practical 3: Decision Tree Learning

the tree
[outlook?]
    outlook = Overcast -> Yes
    outlook = Rainy
        [windy?]
            windy = No -> Yes
            windy = Yes -> No
    outlook = Sunny
        [humidity?]
            humidity = High -> No
            humidity = Normal -> Yes

as rules
  IF outlook = Overcast THEN play = Yes
  IF outlook = Rainy AND windy = No THEN play = Yes
  IF outlook = Rainy AND windy = Yes THEN play = No
  IF outlook = Sunny AND humidity = High THEN play = No
  IF outlook = Sunny AND humidity = Normal THEN play = Yes

on the 14 rows it learned from : 14 of 14 right

outlook   temp  humidity windy  predicted actual
Sunny     Cool  High    Yes    No        No
Overcast  Mild  Normal  Yes    Yes       Yes
Rainy     Hot   Normal  No     Yes       Yes
Rainy     Cool  High    Yes    No        No
Sunny     Hot   Normal  No     Yes       Yes

on 5 unseen rows              : 5 of 5 right, accuracy 1.000
munotes.in42

Practical 3: Decision Tree Learning

Three questions, five leaves, and every one of the fourteen training days classified correctly. temperature never appears: on this data it carries almost no information, and the algorithm dropped it without being told to. That is worth a sentence in the journal, because "which features did the tree not need?" is a question an examiner asks.

Two details in id3 that a marker looks for:

values_of, not the values present in this group. Rainy days in this dataset are never Hot, so if the tree ever asked about temperature under Rainy, there would be no Hot branch, and a new Hot rainy day would fall off the tree. Building a branch for every value the whole dataset has, and putting the group's majority answer there, is what stops that.

default. Even with every value covered, a row can arrive with a value nobody has ever seen. classify falls back to the node's majority rather than raising an error, which is what a model has to do in an examination hall when the examiner types something unexpected.

Step 3: drawing the tree

MU says "visualize and interpret the generated decision tree". The text form above is a drawing; here is the same tree as a picture, and this is what goes in the journal.

The decision tree built from the fourteen days: outlook at the root, windy under Rainy, humidity under Sunny, and five leaves.

Figure 6.1 The tree ID3 built. Overcast needs no second question because all four overcast days ended the same way.

Reading it out loud is the interpretation, and it is the part of this practical that carries marks:

  • Overcast means play, always. Four days out of four, no exceptions in the data.
  • On a rainy day the wind decides. Calm, play; windy, do not.
  • On a sunny day the humidity decides. High, do not play; normal, play.
  • Temperature never came up. It was available and the algorithm found it worthless.

That is what a decision tree gives you and a neural network does not: an answer you can read, argue with, and hand to somebody who has never heard of machine learning.

Step 4: measuring it properly

The output above says 14 of 14 on the rows it learned from, and chapter 3 says that figure is worth nothing. So split the fourteen days the way chapter 3 requires and measure on the ones held back.

"""The same tree, but measured the way chapter 30 requires: on rows it never saw."""
import csv, math, random
from collections import Counter
TARGET = "play"
FEATURES = ["outlook", "temperature", "humidity", "windy"]

def entropy(rows):
    c = Counter(r[TARGET] for r in rows); n = len(rows)
    return -sum((v/n) * math.log2(v/n) for v in c.values())
def split_on(rows, f):
    g = {}
    for r in rows: g.setdefault(r[f], []).append(r)
    return g
def gain(rows, f):
    n = len(rows)
    return entropy(rows) - sum(len(x)/n * entropy(x) for x in split_on(rows, f).values())
def majority(rows): return Counter(r[TARGET] for r in rows).most_common(1)[0][0]
def id3(rows, feats, values_of, depth=0, max_depth=None):
    labels = set(r[TARGET] for r in rows)
    if len(labels) == 1: return labels.pop()
    if not feats or (max_depth is not None and depth >= max_depth): return majority(rows)
    best = max(feats, key=lambda f: gain(rows, f))
    node = {"feature": best, "default": majority(rows), "branches": {}}
    rest = [f for f in feats if f != best]
    groups = split_on(rows, best)
    for v in values_of[best]:
        node["branches"][v] = (id3(groups[v], rest, values_of, depth+1, max_depth)
                               if groups.get(v) else majority(rows))
    return node
def classify(node, row):
    while isinstance(node, dict):
        node = node["branches"].get(row[node["feature"]], node["default"])
    return node
def leaves(node):
    if not isinstance(node, dict): return 1
    return sum(leaves(c) for c in node["branches"].values())
def depth_of(node):
    if not isinstance(node, dict): return 0
    return 1 + max(depth_of(c) for c in node["branches"].values())

with open("weather.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
values_of = {f: sorted(set(r[f] for r in rows)) for f in FEATURES}
random.seed(3)
order = list(range(len(rows))); random.shuffle(order)
train = [rows[i] for i in order[:10]]
test = [rows[i] for i in order[10:]]
print("training rows:", len(train), " test rows:", len(test))
print()
print("%-11s %6s %7s %10s %9s" % ("max depth", "leaves", "depth", "train acc", "test acc"))
for md in (1, 2, 3, None):
    t = id3(train, FEATURES, values_of, max_depth=md)
    tr = sum(1 for r in train if classify(t, r) == r[TARGET]) / len(train)
    te = sum(1 for r in test if classify(t, r) == r[TARGET]) / len(test)
    print("%-11s %6d %7d %10.3f %9.3f"
          % ("unlimited" if md is None else md, leaves(t), depth_of(t), tr, te))
munotes.in43

Practical 3: Decision Tree Learning

training rows: 10  test rows: 4

max depth   leaves   depth  train acc  test acc
1                3       1      0.800     0.250
2                5       2      1.000     1.000
3                5       2      1.000     1.000
unlimited        5       2      1.000     1.000

Two useful things and one warning.

Depth 1 is not enough. A tree allowed only the root question scores 0.800 on the rows it learned from and 0.250 on the four it did not. It is underfitting: too simple to capture what is there.

Depth 2 is enough on this data, and growing further changes nothing: the unlimited tree has the same five leaves as the depth-2 tree, because the data runs out of disagreement before the algorithm runs out of questions.

The warning: four test rows. One row is 0.25 of the accuracy. A score of 1.000 on four rows is not proof of anything, and neither is 0.250. Say the size of the test set next to the score, always. The next section uses a larger one for exactly this reason.

munotes.in44

Practical 3: Decision Tree Learning

Step 5: where a tree goes wrong, on data that has exceptions in it

The weather dataset has no contradictions in it: no two days agree on all four features and disagree on the answer. A tree grown on data like that cannot overfit, because there is nothing false to learn.

Real data is not like that. The thirty-student dataset from chapter 3 has two students who break the pattern: one with high attendance who failed and one with low attendance who passed. Watch what a tree does when it is allowed to grow far enough to memorise them.

name,attendance,practice,result
Aarav,50,22,Fail
Isha,48,16,Fail
Rohan,88,38,Pass
Sanya,97,39,Pass
Vikram,90,30,Fail
Meera,35,41,Fail
Farhan,45,9,Fail
Nikita,71,8,Fail
Omkar,92,2,Fail
Pooja,97,45,Pass
Rahul,75,15,Fail
Sneha,85,18,Pass
Tejas,79,24,Pass
Urmila,83,34,Pass
Varun,44,23,Fail
Yash,46,37,Pass
Zoya,72,20,Fail
Amit,44,9,Pass
Bhavna,74,3,Fail
Chirag,82,25,Pass
Deepa,94,29,Pass
Eshan,46,27,Fail
Gauri,98,9,Fail
Harsh,89,34,Pass
Ira,97,27,Pass
Jatin,68,28,Pass
Kavya,96,34,Pass
Lalit,38,38,Fail
Manav,63,10,Fail
Neha,41,35,Fail
"""Overfitting needs a dataset with exceptions in it. students.csv has two."""
import csv
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
X = [[int(r["attendance"]), int(r["practice"])] for r in rows]
y = [r["result"] for r in rows]
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=1, stratify=y)

print("training rows:", len(Xtr), " test rows:", len(Xte))
print()
print("%-11s %7s %7s %10s %9s" % ("max depth", "leaves", "depth", "train acc", "test acc"))
for md in (1, 2, 3, 4, 5, None):
    t = DecisionTreeClassifier(criterion="entropy", max_depth=md, random_state=0).fit(Xtr, ytr)
    print("%-11s %7d %7d %10.3f %9.3f"
          % ("unlimited" if md is None else md, t.get_n_leaves(), t.get_depth(),
             t.score(Xtr, ytr), t.score(Xte, yte)))
training rows: 21  test rows: 9

max depth    leaves   depth  train acc  test acc
1                 2       1      0.762     0.889
2                 4       2      0.810     1.000
3                 6       3      0.905     0.889
4                 8       4      1.000     0.778
5                 8       4      1.000     0.778
unlimited         8       4      1.000     0.778

That is overfitting, measured. Read the two right-hand columns together:

depthleavestraintest
120.7620.889
240.8101.000
360.9050.889
481.0000.778
unlimited81.0000.778

Training accuracy climbs all the way to a perfect 1.000 and never falls. Test accuracy peaks at depth 2 and then drops, to 0.778. The deep tree has grown extra leaves whose only job is to memorise the two students who break the pattern, and those leaves are wrong about everybody else.

The tree that scored perfectly is the worst tree on the page. A student who reports only the training accuracy will report 1.000 and will have measured nothing.

The cure is to stop growing: max_depth, or a minimum number of rows in a leaf, or growing the tree fully and then cutting branches back, which is called pruning. On this data max_depth=2 is the answer, and it was found by measuring rather than by choosing.

munotes.in45

Practical 3: Decision Tree Learning

Step 6: the same tree from scikit-learn

outlook,temperature,humidity,windy,play
Sunny,Hot,High,No,No
Sunny,Hot,High,Yes,No
Overcast,Hot,High,No,Yes
Rainy,Mild,High,No,Yes
Rainy,Cool,Normal,No,Yes
Rainy,Cool,Normal,Yes,No
Overcast,Cool,Normal,Yes,Yes
Sunny,Mild,High,No,No
Sunny,Cool,Normal,No,Yes
Rainy,Mild,Normal,No,Yes
Sunny,Mild,Normal,Yes,Yes
Overcast,Mild,High,Yes,Yes
Overcast,Hot,Normal,No,Yes
Rainy,Mild,High,Yes,No
"""The same tree from scikit-learn, and the overfitting a noisy dataset shows."""
import csv
from sklearn.tree import DecisionTreeClassifier, export_text

with open("weather.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
FEATURES = ["outlook", "temperature", "humidity", "windy"]
codes = {f: {v: i for i, v in enumerate(sorted(set(r[f] for r in rows)))} for f in FEATURES}
for f in FEATURES:
    print("%-12s %s" % (f, codes[f]))
X = [[codes[f][r[f]] for f in FEATURES] for r in rows]
y = [r["play"] for r in rows]

tree = DecisionTreeClassifier(criterion="entropy", random_state=0).fit(X, y)
print()
print(export_text(tree, feature_names=FEATURES).rstrip())
print()
print("training accuracy: %.3f" % tree.score(X, y))
print("leaves           :", tree.get_n_leaves())
print("depth            :", tree.get_depth())
outlook      {'Overcast': 0, 'Rainy': 1, 'Sunny': 2}
temperature  {'Cool': 0, 'Hot': 1, 'Mild': 2}
humidity     {'High': 0, 'Normal': 1}
windy        {'No': 0, 'Yes': 1}

|--- outlook <= 0.50
|   |--- class: Yes
|--- outlook >  0.50
|   |--- humidity <= 0.50
|   |   |--- outlook <= 1.50
|   |   |   |--- windy <= 0.50
|   |   |   |   |--- class: Yes
|   |   |   |--- windy >  0.50
|   |   |   |   |--- class: No
|   |   |--- outlook >  1.50
|   |   |   |--- class: No
|   |--- humidity >  0.50
|   |   |--- windy <= 0.50
|   |   |   |--- class: Yes
|   |   |--- windy >  0.50
|   |   |   |--- temperature <= 1.00
|   |   |   |   |--- class: No
|   |   |   |--- temperature >  1.00
|   |   |   |   |--- class: Yes

training accuracy: 1.000
leaves           : 7
depth            : 4

That is not the tree ID3 built, and both are right. ID3 made five leaves and asked three questions; scikit-learn made seven leaves and went four deep. Two reasons, and both belong in the journal.

scikit-learn's tree is binary. Every node asks a yes-or-no question about one number. Our outlook has three values, and ID3 can branch three ways in one node; scikit-learn has to ask outlook <= 0.50, and then ask again inside the "no" branch. Same information, more nodes.

Integer codes invent an order that does not exist. Writing Overcast as 0, Rainy as 1, Sunny as 2 tells the model that Overcast is less than Rainy, which is meaningless, and lets it ask outlook <= 1.50, which groups Overcast with Rainy for no reason at all. The honest encoding is one-hot: one column per value, holding 1 or 0.

"""Ordinal codes tell the tree a lie. One-hot columns do not."""
import csv
from sklearn.tree import DecisionTreeClassifier, export_text

with open("weather.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
FEATURES = ["outlook", "temperature", "humidity", "windy"]
y = [r["play"] for r in rows]

# one column per (feature, value) pair, holding 1 or 0
pairs = [(f, v) for f in FEATURES for v in sorted(set(r[f] for r in rows))]
names = ["%s=%s" % (f, v) for f, v in pairs]
X = [[1 if r[f] == v else 0 for f, v in pairs] for r in rows]

print("columns:", len(names))
print(", ".join(names))
print()
tree = DecisionTreeClassifier(criterion="entropy", random_state=0).fit(X, y)
print(export_text(tree, feature_names=names).rstrip())
print()
print("training accuracy: %.3f" % tree.score(X, y))
print("leaves           :", tree.get_n_leaves())
print("depth            :", tree.get_depth())
munotes.in46

Practical 3: Decision Tree Learning

columns: 10
outlook=Overcast, outlook=Rainy, outlook=Sunny, temperature=Cool, temperature=Hot, temperature=Mild, humidity=High, humidity=Normal, windy=No, windy=Yes

|--- outlook=Overcast <= 0.50
|   |--- humidity=Normal <= 0.50
|   |   |--- outlook=Rainy <= 0.50
|   |   |   |--- class: No
|   |   |--- outlook=Rainy >  0.50
|   |   |   |--- windy=No <= 0.50
|   |   |   |   |--- class: No
|   |   |   |--- windy=No >  0.50
|   |   |   |   |--- class: Yes
|   |--- humidity=Normal >  0.50
|   |   |--- windy=Yes <= 0.50
|   |   |   |--- class: Yes
|   |   |--- windy=Yes >  0.50
|   |   |   |--- temperature=Cool <= 0.50
|   |   |   |   |--- class: Yes
|   |   |   |--- temperature=Cool >  0.50
|   |   |   |   |--- class: No
|--- outlook=Overcast >  0.50
|   |--- class: Yes

training accuracy: 1.000
leaves           : 7
depth            : 4

On this dataset one-hot gave the same size of tree, seven leaves and depth four, and the same perfect training accuracy: the false ordering happened not to hurt here. It is still the encoding to use, because whether it hurts depends on the data and you cannot tell in advance, and because a tree that asks outlook=Overcast <= 0.50 can be read by a human while one that asks outlook <= 1.50 cannot.

criterion="entropy" is what makes scikit-learn use information gain. Its default is "gini", a different measure of mixture that usually builds a very similar tree. Say which you used.

Procedure

  1. Save weather.csv. Count the Yes and No rows and compute the entropy of the whole set by hand, then confirm it with the program.
  2. Compute the information gain of all four features. Confirm that outlook wins and write the four figures into the journal.
  3. Implement entropy, split_on, gain and majority, then id3 with its two stopping rules.
  4. Print the tree as text and read it out as rules.
  5. Draw the tree.
  6. Split the fourteen rows 10 to 4 with a seeded shuffle and measure the tree on the four it never saw. Vary max_depth and record training and test accuracy for each.
  7. Repeat on students.csv, which has exceptions in it, and record the depth at which test accuracy starts to fall.
  8. Build the same tree with DecisionTreeClassifier(criterion="entropy"), first with integer codes and then with one-hot columns, and explain why its tree differs from yours.
munotes.in47

Practical 3: Decision Tree Learning

Observations

Measured on weather.csvValue
Rows14, nine Yes and five No
Entropy of the whole set0.9403 bits
Gain: outlook0.2467
Gain: humidity0.1518
Gain: windy0.0481
Gain: temperature0.0292
Root chosenoutlook
Overcast branch4 rows, all Yes, entropy 0, a leaf at once
Rainy and Sunny branches0.9710 bits each, one more question needed
Tree size, our ID33 questions, 5 leaves, depth 2
Features never usedtemperature
Accuracy on the rows it learned from14 of 14
Split 10 to 4, depth 1train 0.800, test 0.250
Split 10 to 4, depth 2 and abovetrain 1.000, test 1.000
scikit-learn, integer codes7 leaves, depth 4, training accuracy 1.000
scikit-learn, one-hot columns7 leaves, depth 4, training accuracy 1.000
Measured on students.csv, 21 training and 9 test rowstraintest
max depth 1, 2 leaves0.7620.889
max depth 2, 4 leaves0.8101.000
max depth 3, 6 leaves0.9050.889
max depth 4, 8 leaves1.0000.778
unlimited, 8 leaves1.0000.778

Result

The Decision Tree Learning algorithm was implemented from nothing and used to build a tree from the fourteen-day dataset. Entropy of the whole set was 0.9403 bits and outlook gave the largest information gain, 0.2467, so it became the root; the Overcast branch was pure and became a leaf immediately. The finished tree asked three questions, had five leaves, never used temperature, and classified all fourteen training rows correctly. Measured on four held-out rows it scored 1.000 at depth 2 and 0.250 at depth 1. On the thirty-student dataset, which contains two exceptions, training accuracy rose to 1.000 at depth 4 while test accuracy fell from 1.000 at depth 2 to 0.778, which is overfitting. scikit-learn built a seven-leaf binary tree from the same data, with the same perfect training accuracy, under both integer and one-hot encoding.

Where marks are lost

Reporting the training accuracy. 14 of 14 means the tree remembered. Split the data.

Growing the tree as far as it will go. On data with exceptions that is how test accuracy falls while training accuracy rises. Show the table of depths.

No branch for a value the group did not contain. Build every branch from the values the whole dataset has, or a new row falls off the tree.

Using log base e. Entropy in this subject is in bits, so it is math.log2. Natural logarithms give 0.6518 where the answer is 0.9403 and every gain comes out wrong by the same factor.

munotes.in48

Practical 3: Decision Tree Learning

Encoding categories as 0, 1, 2 and saying nothing about it. It tells the model an order that is not there. Use one-hot, or say why you did not.

Not saying which criterion you used. scikit-learn's default is gini, not entropy.

Drawing the tree and not reading it. MU asks you to visualize and interpret. Write the rules out in English.

Dividing by zero in entropy. A group with one class has a single p of 1.0 and log2(1) is 0, which is fine; a group with no rows at all is the one to guard, and the if not group branch above is the guard.

For the journal

Aim; entropy and information gain defined with the formulas; the dataset; the entropy of the whole set worked by hand; the gain table for all four features and the root chosen; the ID3 program; the tree printed as text and drawn as a figure; the five rules in English; the train and test split with the accuracy at each depth; the overfitting table from students.csv; the scikit-learn tree and one sentence on why it differs; both observation tables; the result.

Quick revision

  • Entropy is the mixture of labels in bits, 0 when pure and 1 when half and half.
  • Information gain is the entropy before a split minus the weighted average entropy after.
  • ID3 picks the highest-gain question at every node and repeats.
  • Stop when the group is pure, when the questions run out, or at a depth limit.
  • H of the 14-day set is 0.9403 bits. outlook gains 0.2467 and wins the root.
  • The finished tree: 3 questions, 5 leaves, and temperature never used.
  • A path from root to leaf is a rule. That readability is the tree's main advantage.
  • A tree grown to purity on data with exceptions overfits: training 1.000, test 0.778 here.
  • Cure it with max_depth, a minimum leaf size, or pruning, and choose the value by measuring.
  • scikit-learn's tree is binary, so it needs more nodes than ID3 for a three-valued feature.
  • Encode categories one-hot. Integer codes invent an ordering.
  • criterion="entropy" for information gain; the default is gini.

Questions you must be able to answer

1. What is entropy, in one sentence, and what is it 0 for? A measure in bits of how mixed the labels in a set are. It is 0 for a set whose rows all carry the same label, because there is nothing left to learn.

2. Why was outlook chosen as the root? Because it had the highest information gain, 0.2467 bits against humidity's 0.1518, windy's 0.0481 and temperature's 0.0292.

3. What is information gain? The entropy of a set minus the average entropy of the groups a question splits it into, each weighted by its share of the rows. It is how many bits of uncertainty the question removes.

munotes.in49

Practical 3: Decision Tree Learning

4. Why does the tree stop immediately on the Overcast branch? Because all four overcast days ended in Yes, so that group's entropy is 0 and no further question can tell you anything.

5. Your tree never uses temperature. Is that a bug? No. Its gain is 0.0292, the lowest of the four, and by the time the tree has split on outlook and then on windy or humidity, each group is pure. The algorithm used what it needed.

6. What is overfitting, and how did you show it? Learning the accidents of the training data instead of the pattern. On students.csv training accuracy rose to 1.000 at depth 4 while test accuracy fell from 1.000 at depth 2 to 0.778.

7. How do you stop a tree overfitting? Limit its depth, require a minimum number of rows in a leaf, or grow it fully and prune back. Choose the limit by measuring on held-out data, not by guessing.

8. Why does scikit-learn's tree look different from yours on the same data? Because it is binary: every node asks one yes-or-no question about a number, so a feature with three values takes two levels instead of one. Its seven leaves and your five describe the same rules.

9. What is wrong with coding Sunny, Overcast and Rainy as 0, 1 and 2? It tells the model they are ordered and evenly spaced, which they are not, and lets it split on a threshold that groups two of them for no reason. One-hot columns carry the same information without the false ordering.

10. A new row arrives with a value your tree has no branch for. What happens? classify falls back to that node's majority label. Without such a default the lookup raises an error, which in an examination is a program that crashed in front of the examiner.

Contents This chapter on its own page

munotes.in50

Chapter Seven

Practical 4: The Feed Forward Backpropagation Neural Network

Syllabus topic Module 1, "Feed Forward Backpropagation Neural Network: Implement the Feed Forward Backpropagation algorithm to train a neural network. Use a given dataset to train the neural network for a specific task. Evaluate the performance of the trained network on test data."

Aim

To implement the feed forward backpropagation algorithm, to train a neural network on a dataset, and to measure the trained network on data it has not seen.

What you need to know before you start

A neuron does two things. It adds up its inputs, each multiplied by a weight, plus one number of its own called the bias. Then it passes that sum through a squashing function.

z = b + w1x1 + w2x2 + ... + wn*xn

a = sigmoid(z) = 1 / (1 + e^(-z))

The sigmoid turns any number into one between 0 and 1: a large negative z gives nearly 0, a large positive z nearly 1, and z of 0 gives exactly 0.5. That is what lets the network's answer be read as "yes" or "no".

It has one more property, and it is the reason it is used here rather than anything simpler: its slope can be written in terms of its own output.

d/dz sigmoid(z) = sigmoid(z) (1 - sigmoid(z)) = a (1 - a)

Every a * (1 - a) in the program below is that line.

Feed forward means the values only ever travel one way: inputs to hidden layer, hidden layer to output. Backpropagation is how the blame for a wrong answer travels the other way: the output's error is shared out among the hidden neurons in proportion to the weights that carried their values forward.

The network in this chapter:

Two inputs, four hidden neurons, one output, every input joined to every hidden neuron and every hidden neuron to the output.

Figure 7.1 The 2-4-1 network. W1 holds the eight weights on the left and W2 the four on the right, and each neuron also has a bias of its own.

Step 1: one forward pass and one backward pass, with every number shown

Before any loop, do it once by hand. Two inputs, two hidden neurons, one output, weights chosen to be easy to read, the input 1 0 and the target 1.

"""One forward pass and one backward pass, with every number shown."""
import math

def sigmoid(z):
    return 1 / (1 + math.exp(-z))

# a 2-2-1 network with weights chosen so the arithmetic can be checked by hand
w1 = [[0.10, 0.40],      # input 1 to hidden 1, hidden 2
      [0.20, 0.30]]      # input 2 to hidden 1, hidden 2
b1 = [0.05, -0.10]
w2 = [0.50, -0.60]       # hidden 1, hidden 2 to output
b2 = 0.15
x = [1.0, 0.0]
target = 1.0
rate = 0.5

print("FORWARD")
z1 = []
a1 = []
for j in range(2):
    z = b1[j] + sum(x[i] * w1[i][j] for i in range(2))
    z1.append(z)
    a1.append(sigmoid(z))
    print("  hidden %d: z = %.4f + %.2f*%.2f + %.2f*%.2f = %.4f   a = sigmoid(z) = %.6f"
          % (j + 1, b1[j], x[0], w1[0][j], x[1], w1[1][j], z, a1[j]))
z_out = b2 + sum(a1[j] * w2[j] for j in range(2))
a_out = sigmoid(z_out)
print("  output  : z = %.4f + %.6f*%.2f + %.6f*%.2f = %.6f" % (b2, a1[0], w2[0], a1[1], w2[1], z_out))
print("  output  : a = sigmoid(z) = %.6f" % a_out)
print()
err = 0.5 * (target - a_out) ** 2
print("ERROR")
print("  target %.1f, output %.6f" % (target, a_out))
print("  E = 0.5 * (target - output)^2 = %.6f" % err)
print()
print("BACKWARD")
d_out = (a_out - target) * a_out * (1 - a_out)
print("  delta at the output = (a - t) * a * (1 - a)")
print("                      = (%.6f - %.1f) * %.6f * %.6f = %.6f"
      % (a_out, target, a_out, 1 - a_out, d_out))
d_hid = []
for j in range(2):
    d = d_out * w2[j] * a1[j] * (1 - a1[j])
    d_hid.append(d)
    print("  delta at hidden %d   = %.6f * %.2f * %.6f * %.6f = %.6f"
          % (j + 1, d_out, w2[j], a1[j], 1 - a1[j], d))
print()
print("UPDATE, learning rate %.1f" % rate)
for j in range(2):
    new = w2[j] - rate * d_out * a1[j]
    print("  w2[%d]: %.4f - %.1f * %.6f * %.6f = %.6f" % (j, w2[j], rate, d_out, a1[j], new))
new_b2 = b2 - rate * d_out
print("  b2   : %.4f - %.1f * %.6f = %.6f" % (b2, rate, d_out, new_b2))
for i in range(2):
    for j in range(2):
        new = w1[i][j] - rate * d_hid[j] * x[i]
        print("  w1[%d][%d]: %.4f - %.1f * %.6f * %.1f = %.6f"
              % (i, j, w1[i][j], rate, d_hid[j], x[i], new))
munotes.in51

Practical 4: The Feed Forward Backpropagation Neural Network

FORWARD
  hidden 1: z = 0.0500 + 1.00*0.10 + 0.00*0.20 = 0.1500   a = sigmoid(z) = 0.537430
  hidden 2: z = -0.1000 + 1.00*0.40 + 0.00*0.30 = 0.3000   a = sigmoid(z) = 0.574443
  output  : z = 0.1500 + 0.537430*0.50 + 0.574443*-0.60 = 0.074049
  output  : a = sigmoid(z) = 0.518504

ERROR
  target 1.0, output 0.518504
  E = 0.5 * (target - output)^2 = 0.115919

BACKWARD
  delta at the output = (a - t) * a * (1 - a)
                      = (0.518504 - 1.0) * 0.518504 * 0.481496 = -0.120209
  delta at hidden 1   = -0.120209 * 0.50 * 0.537430 * 0.462570 = -0.014942
  delta at hidden 2   = -0.120209 * -0.60 * 0.574443 * 0.425557 = 0.017632

UPDATE, learning rate 0.5
  w2[0]: 0.5000 - 0.5 * -0.120209 * 0.537430 = 0.532302
  w2[1]: -0.6000 - 0.5 * -0.120209 * 0.574443 = -0.565473
  b2   : 0.1500 - 0.5 * -0.120209 = 0.210105
  w1[0][0]: 0.1000 - 0.5 * -0.014942 * 1.0 = 0.107471
  w1[0][1]: 0.4000 - 0.5 * 0.017632 * 1.0 = 0.391184
  w1[1][0]: 0.2000 - 0.5 * -0.014942 * 0.0 = 0.200000
  w1[1][1]: 0.3000 - 0.5 * 0.017632 * 0.0 = 0.300000
munotes.in52

Practical 4: The Feed Forward Backpropagation Neural Network

Work down that output with a calculator once and backpropagation is no longer mysterious. Four things in it are worth saying in the journal.

The error before training is 0.1159 and the output is 0.5185 when the target is 1. The untrained network is guessing, which is what random weights do.

The delta at the output is (a - t) a (1 - a). The first factor is how wrong the answer is; the second and third are the slope of the sigmoid there. Multiplying by the slope is what makes this gradient descent rather than a guess: a neuron that is already saturated, output near 0 or near 1, has almost no slope and so moves almost not at all.

The delta at a hidden neuron is the output's delta, carried back through the weight that joined them, times that neuron's own slope. Hidden 1 gets a negative delta and hidden 2 a positive one, because their weights to the output have opposite signs. That single line is the whole of "backpropagation".

The two weights from input 2 did not move at all. Look at the last two lines: w1[1][0] and w1[1][1] come out exactly as they went in, because input 2 was 0 and every update is multiplied by the input. A weight from an input that is zero carries no blame, and it learns nothing from that example.

Step 2: the network, and a task it can only do with a hidden layer

XOR: output 1 when exactly one input is 1. It is the standard first task for a network because no straight line separates its two classes, so it cannot be done without a hidden layer, and demonstrating that is half the practical.

"""Practical 4: a feed forward network trained by backpropagation, on XOR."""
import math
import random

def sigmoid(z):
    return 1 / (1 + math.exp(-z))

class Network:
    """One hidden layer. Weights are lists of lists; no library is used."""

    def __init__(self, n_in, n_hidden, n_out, seed=1):
        rng = random.Random(seed)
        span = 1.0
        self.w1 = [[rng.uniform(-span, span) for _ in range(n_hidden)] for _ in range(n_in)]
        self.b1 = [rng.uniform(-span, span) for _ in range(n_hidden)]
        self.w2 = [[rng.uniform(-span, span) for _ in range(n_out)] for _ in range(n_hidden)]
        self.b2 = [rng.uniform(-span, span) for _ in range(n_out)]

    def forward(self, x):
        self.a1 = [sigmoid(self.b1[j] + sum(x[i] * self.w1[i][j] for i in range(len(x))))
                   for j in range(len(self.b1))]
        self.a2 = [sigmoid(self.b2[k] + sum(self.a1[j] * self.w2[j][k] for j in range(len(self.a1))))
                   for k in range(len(self.b2))]
        return self.a2

    def backward(self, x, target, rate):
        out = self.a2
        d_out = [(out[k] - target[k]) * out[k] * (1 - out[k]) for k in range(len(out))]
        d_hid = [sum(d_out[k] * self.w2[j][k] for k in range(len(out)))
                 * self.a1[j] * (1 - self.a1[j]) for j in range(len(self.a1))]
        for j in range(len(self.a1)):
            for k in range(len(out)):
                self.w2[j][k] -= rate * d_out[k] * self.a1[j]
        for k in range(len(out)):
            self.b2[k] -= rate * d_out[k]
        for i in range(len(x)):
            for j in range(len(self.a1)):
                self.w1[i][j] -= rate * d_hid[j] * x[i]
        for j in range(len(self.a1)):
            self.b1[j] -= rate * d_hid[j]

    def train(self, data, epochs, rate, report_every=0):
        curve = []
        for epoch in range(1, epochs + 1):
            total = 0.0
            for x, t in data:
                out = self.forward(x)
                total += sum(0.5 * (t[k] - out[k]) ** 2 for k in range(len(t)))
                self.backward(x, t, rate)
            curve.append(total)
            if report_every and (epoch == 1 or epoch % report_every == 0):
                print("  epoch %6d   total error %.6f" % (epoch, total))
        return curve

XOR = [([0.0, 0.0], [0.0]),
       ([0.0, 1.0], [1.0]),
       ([1.0, 0.0], [1.0]),
       ([1.0, 1.0], [0.0])]

net = Network(2, 4, 1, seed=1)
print("training on XOR, 2 inputs, 4 hidden, 1 output, learning rate 0.5")
net.train(XOR, 20000, 0.5, report_every=4000)
print()
print("%-8s %10s %10s %8s" % ("input", "target", "output", "rounded"))
right = 0
for x, t in XOR:
    out = net.forward(x)[0]
    got = 1 if out >= 0.5 else 0
    right += (got == int(t[0]))
    print("%-8s %10.0f %10.6f %8d" % ("%d %d" % (x[0], x[1]), t[0], out, got))
print()
print("correct: %d of %d" % (right, len(XOR)))
munotes.in53

Practical 4: The Feed Forward Backpropagation Neural Network

training on XOR, 2 inputs, 4 hidden, 1 output, learning rate 0.5
  epoch      1   total error 0.570178
  epoch   4000   total error 0.001644
  epoch   8000   total error 0.000655
  epoch  12000   total error 0.000401
  epoch  16000   total error 0.000287
  epoch  20000   total error 0.000223

input        target     output  rounded
0 0               0   0.010836        0
0 1               1   0.989533        1
1 0               1   0.989801        1
1 1               0   0.010678        0

correct: 4 of 4

The error falls from 0.570 to 0.000223 and all four cases come out right. Note that the outputs are 0.0108 and 0.9895, not 0 and 1: a sigmoid never quite reaches either end, so the answer is rounded at 0.5. Print the raw output as well as the rounded one, because the raw number is the network's confidence and an examiner may ask for it.

Note also where the error falls. Between epoch 1 and epoch 4,000 it drops from 0.570 to 0.0016; over the next 16,000 epochs it only reaches 0.00022. Almost all the learning happens early, and training ten times longer buys very little. That is the shape of every learning curve you will draw in this subject.

Step 3: three things that decide whether it learns at all

"""Three things that decide whether a network learns anything at all."""
import math, random
def sigmoid(z): return 1 / (1 + math.exp(-z))

class Network:
    def __init__(self, n_in, n_hidden, n_out, seed=1):
        rng = random.Random(seed)
        self.w1 = [[rng.uniform(-1, 1) for _ in range(n_hidden)] for _ in range(n_in)]
        self.b1 = [rng.uniform(-1, 1) for _ in range(n_hidden)]
        self.w2 = [[rng.uniform(-1, 1) for _ in range(n_out)] for _ in range(n_hidden)]
        self.b2 = [rng.uniform(-1, 1) for _ in range(n_out)]
    def forward(self, x):
        self.a1 = [sigmoid(self.b1[j] + sum(x[i]*self.w1[i][j] for i in range(len(x))))
                   for j in range(len(self.b1))]
        self.a2 = [sigmoid(self.b2[k] + sum(self.a1[j]*self.w2[j][k] for j in range(len(self.a1))))
                   for k in range(len(self.b2))]
        return self.a2
    def backward(self, x, t, rate):
        out = self.a2
        d_out = [(out[k]-t[k])*out[k]*(1-out[k]) for k in range(len(out))]
        d_hid = [sum(d_out[k]*self.w2[j][k] for k in range(len(out)))*self.a1[j]*(1-self.a1[j])
                 for j in range(len(self.a1))]
        for j in range(len(self.a1)):
            for k in range(len(out)): self.w2[j][k] -= rate*d_out[k]*self.a1[j]
        for k in range(len(out)): self.b2[k] -= rate*d_out[k]
        for i in range(len(x)):
            for j in range(len(self.a1)): self.w1[i][j] -= rate*d_hid[j]*x[i]
        for j in range(len(self.a1)): self.b1[j] -= rate*d_hid[j]
    def train(self, data, epochs, rate):
        for _ in range(epochs):
            total = 0.0
            for x, t in data:
                o = self.forward(x)
                total += sum(0.5*(t[k]-o[k])**2 for k in range(len(t)))
                self.backward(x, t, rate)
        return total

XOR = [([0.,0.],[0.]), ([0.,1.],[1.]), ([1.,0.],[1.]), ([1.,1.],[0.])]

def score(net):
    return sum(1 for x, t in XOR if (1 if net.forward(x)[0] >= 0.5 else 0) == int(t[0]))

print("A. how many hidden neurons XOR needs")
print("%-10s %14s %10s" % ("hidden", "final error", "correct"))
for n in (0, 1, 2, 3, 4):
    if n == 0:
        # no hidden layer at all: a single sigmoid unit on the two inputs
        rng = random.Random(1)
        w = [rng.uniform(-1, 1) for _ in range(2)]; b = rng.uniform(-1, 1)
        for _ in range(20000):
            total = 0.0
            for x, t in XOR:
                o = sigmoid(b + w[0]*x[0] + w[1]*x[1])
                total += 0.5*(t[0]-o)**2
                d = (o-t[0])*o*(1-o)
                w[0] -= 0.5*d*x[0]; w[1] -= 0.5*d*x[1]; b -= 0.5*d
        got = sum(1 for x, t in XOR
                  if (1 if sigmoid(b+w[0]*x[0]+w[1]*x[1]) >= 0.5 else 0) == int(t[0]))
        print("%-10s %14.6f %10d" % ("none", total, got))
        continue
    net = Network(2, n, 1, seed=1)
    err = net.train(XOR, 20000, 0.5)
    print("%-10d %14.6f %10d" % (n, err, score(net)))

print()
print("B. the learning rate, 4 hidden neurons, 20000 epochs")
print("%-14s %14s %10s" % ("learning rate", "final error", "correct"))
for rate in (0.01, 0.1, 0.5, 2.0, 10.0, 50.0, 200.0):
    net = Network(2, 4, 1, seed=1)
    err = net.train(XOR, 20000, rate)
    print("%-14.2f %14.6f %10d" % (rate, err, score(net)))

print()
print("C. the starting weights, 4 hidden neurons, rate 0.5")
print("%-10s %14s %10s" % ("seed", "final error", "correct"))
for seed in (1, 2, 3, 4, 5):
    net = Network(2, 4, 1, seed=seed)
    err = net.train(XOR, 20000, 0.5)
    print("%-10d %14.6f %10d" % (seed, err, score(net)))
net = Network(2, 4, 1, seed=1)
net.w1 = [[0.0]*4 for _ in range(2)]; net.b1 = [0.0]*4
net.w2 = [[0.0] for _ in range(4)]; net.b2 = [0.0]
err = net.train(XOR, 20000, 0.5)
print("%-10s %14.6f %10d" % ("all zero", err, score(net)))
munotes.in54

Practical 4: The Feed Forward Backpropagation Neural Network

A. how many hidden neurons XOR needs
hidden        final error    correct
none             0.532731          2
1                0.350102          3
2                0.000262          4
3                0.000231          4
4                0.000223          4

B. the learning rate, 4 hidden neurons, 20000 epochs
learning rate     final error    correct
0.01                 0.479992          2
0.10                 0.001411          4
0.50                 0.000223          4
2.00                 0.000050          4
10.00                0.000010          4
50.00                0.499976          3
200.00               1.000000          2

C. the starting weights, 4 hidden neurons, rate 0.5
seed          final error    correct
1                0.000223          4
2                0.000183          4
3                0.000184          4
4                0.000259          4
5                0.000190          4
all zero         0.342703          3
munotes.in55

Practical 4: The Feed Forward Backpropagation Neural Network

Three separate lessons, each measured rather than asserted.

A. XOR needs a hidden layer, and needs at least two neurons in it. With no hidden layer the single unit trained for 20,000 epochs and got 2 of 4, which is what guessing gets. With one hidden neuron, 3 of 4. With two or more, 4 of 4. That is the oldest result in the subject: a single layer of weights can only carve the input space with one straight line, and no straight line separates XOR.

B. The learning rate has a floor and a ceiling. At 0.01 the network was still at error 0.48 after 20,000 epochs: it was learning, just far too slowly. From 0.1 to 10 it worked, and on this very small problem the larger rates reached a lower error. At 50 it broke, and at 200 the error sat at exactly 1.000, which is the worst it can be: the steps are so large that each one overshoots and the weights are thrown further out than they started. Too small and it never arrives; too large and it never settles. On real data the useful range is much narrower than it is here, and 0.01 to 0.5 is where to start.

C. The starting weights must be random and different. Five different seeds all reached 4 of 4. Setting every weight to zero gave 3 of 4 and an error of 0.343. The reason is worth a sentence in the journal: if all the hidden neurons start identical, they receive identical deltas, so they stay identical for ever, and a layer of four identical neurons can do exactly what one can. Random starting weights are what break that symmetry.

Step 4: a real dataset, and the mistake that ruins it

MU says "use a given dataset to train the neural network for a specific task". The task: predict Pass or Fail from attendance and practice hours, on the thirty students of chapter 3.

There is one thing to get right first. Attendance runs from 35 to 98 and practice from 2 to 45. Feed those numbers straight into a sigmoid and the sum z is in the hundreds before training starts, the sigmoid is flat out at 1.000, a * (1 - a) is almost zero, and every delta is almost zero, so nothing learns. That is called saturation, and the cure is to squeeze every feature into the same small range first.

munotes.in56

Practical 4: The Feed Forward Backpropagation Neural Network

scaled = (value - smallest) / (largest - smallest)

name,attendance,practice,result
Aarav,50,22,Fail
Isha,48,16,Fail
Rohan,88,38,Pass
Sanya,97,39,Pass
Vikram,90,30,Fail
Meera,35,41,Fail
Farhan,45,9,Fail
Nikita,71,8,Fail
Omkar,92,2,Fail
Pooja,97,45,Pass
Rahul,75,15,Fail
Sneha,85,18,Pass
Tejas,79,24,Pass
Urmila,83,34,Pass
Varun,44,23,Fail
Yash,46,37,Pass
Zoya,72,20,Fail
Amit,44,9,Pass
Bhavna,74,3,Fail
Chirag,82,25,Pass
Deepa,94,29,Pass
Eshan,46,27,Fail
Gauri,98,9,Fail
Harsh,89,34,Pass
Ira,97,27,Pass
Jatin,68,28,Pass
Kavya,96,34,Pass
Lalit,38,38,Fail
Manav,63,10,Fail
Neha,41,35,Fail
"""The same network on a real dataset, and what scaling does to it."""
import csv, math, random
def sigmoid(z):
    if z < -60: return 0.0
    if z > 60: return 1.0
    return 1 / (1 + math.exp(-z))

class Network:
    def __init__(self, n_in, n_hidden, n_out, seed=1):
        rng = random.Random(seed)
        self.w1 = [[rng.uniform(-1, 1) for _ in range(n_hidden)] for _ in range(n_in)]
        self.b1 = [rng.uniform(-1, 1) for _ in range(n_hidden)]
        self.w2 = [[rng.uniform(-1, 1) for _ in range(n_out)] for _ in range(n_hidden)]
        self.b2 = [rng.uniform(-1, 1) for _ in range(n_out)]
    def forward(self, x):
        self.a1 = [sigmoid(self.b1[j] + sum(x[i]*self.w1[i][j] for i in range(len(x))))
                   for j in range(len(self.b1))]
        self.a2 = [sigmoid(self.b2[k] + sum(self.a1[j]*self.w2[j][k] for j in range(len(self.a1))))
                   for k in range(len(self.b2))]
        return self.a2
    def backward(self, x, t, rate):
        out = self.a2
        d_out = [(out[k]-t[k])*out[k]*(1-out[k]) for k in range(len(out))]
        d_hid = [sum(d_out[k]*self.w2[j][k] for k in range(len(out)))*self.a1[j]*(1-self.a1[j])
                 for j in range(len(self.a1))]
        for j in range(len(self.a1)):
            for k in range(len(out)): self.w2[j][k] -= rate*d_out[k]*self.a1[j]
        for k in range(len(out)): self.b2[k] -= rate*d_out[k]
        for i in range(len(x)):
            for j in range(len(self.a1)): self.w1[i][j] -= rate*d_hid[j]*x[i]
        for j in range(len(self.a1)): self.b1[j] -= rate*d_hid[j]

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
raw = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [1.0 if r["result"] == "Pass" else 0.0 for r in rows]

random.seed(1)
order = list(range(len(rows))); random.shuffle(order)
cut = int(0.7 * len(rows))
tr_i, te_i = order[:cut], order[cut:]

def train_and_score(X, epochs=4000, rate=0.5, seed=1):
    net = Network(2, 4, 1, seed=seed)
    data = [(X[i], [y[i]]) for i in tr_i]
    for _ in range(epochs):
        for x, t in data:
            net.forward(x); net.backward(x, t, rate)
    def acc(idx):
        ok = sum(1 for i in idx if (1.0 if net.forward(X[i])[0] >= 0.5 else 0.0) == y[i])
        return ok / len(idx)
    return acc(tr_i), acc(te_i)

lo = [min(r[k] for r in raw) for k in range(2)]
hi = [max(r[k] for r in raw) for k in range(2)]
scaled = [[(r[k] - lo[k]) / (hi[k] - lo[k]) for k in range(2)] for r in raw]

print("attendance ranges %.0f to %.0f, practice %.0f to %.0f" % (lo[0], hi[0], lo[1], hi[1]))
print("first row raw   :", raw[0])
print("first row scaled: [%.4f, %.4f]" % (scaled[0][0], scaled[0][1]))
print()
print("%-22s %10s %9s" % ("features", "train acc", "test acc"))
a, b = train_and_score(raw)
print("%-22s %10.3f %9.3f" % ("raw, 35 to 98", a, b))
a, b = train_and_score(scaled)
print("%-22s %10.3f %9.3f" % ("scaled to 0 and 1", a, b))
munotes.in57

Practical 4: The Feed Forward Backpropagation Neural Network

attendance ranges 35 to 98, practice 2 to 45
first row raw   : [50.0, 22.0]
first row scaled: [0.2381, 0.4651]

features                train acc  test acc
raw, 35 to 98               0.571     0.444
scaled to 0 and 1           0.952     0.778

0.444 on the test set with raw features, 0.778 with scaled ones. The raw run did not merely do worse, it did worse than the baseline of chapter 3. Scaling is not a refinement here; without it the network learned almost nothing.

Note one honest point about the test figure: nine test rows, so 0.778 is 7 of 9 and one row is worth 0.111. The training figure, 0.952, is 20 of 21.

Step 5: the same network from scikit-learn

"""The same job from scikit-learn: MLPClassifier, and scaling again."""
import csv
from sklearn.neural_network import MLPClassifier
from sklearn.preprocessing import MinMaxScaler
from sklearn.model_selection import train_test_split

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
X = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [r["result"] for r in rows]
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=1, stratify=y)

def run(name, Atr, Ate):
    net = MLPClassifier(hidden_layer_sizes=(4,), activation="logistic",
                        solver="lbfgs", max_iter=4000, random_state=1)
    net.fit(Atr, ytr)
    print("%-22s %10.3f %9.3f" % (name, net.score(Atr, ytr), net.score(Ate, yte)))

print("%-22s %10s %9s" % ("features", "train acc", "test acc"))
run("raw", Xtr, Xte)
scaler = MinMaxScaler().fit(Xtr)
run("scaled to 0 and 1", scaler.transform(Xtr), scaler.transform(Xte))
features                train acc  test acc
raw                         0.857     0.778
scaled to 0 and 1           0.952     0.778

MLPClassifier with hidden_layer_sizes=(4,) and activation="logistic" is the same network: one hidden layer of four sigmoid neurons. Its scaled result, 0.952 training and 0.778 test, is exactly what our own program reached, and that agreement is the best evidence available that the implementation above is correct. Put both in the journal side by side and say so.

Two differences to be honest about. solver="lbfgs" is a cleverer optimiser than the plain gradient descent we wrote, which is why scikit-learn reaches the same answer in 4,000 iterations of a different kind. And its raw-feature run scored 0.857 rather than our 0.571, because lbfgs copes with badly scaled inputs better than plain gradient descent does; it still gained from scaling.

Procedure

  1. Write sigmoid and confirm sigmoid(0) is 0.5.
  2. Work one forward pass by hand on a 2-2-1 network with fixed weights, then one backward pass, and check every number against the program.
  3. Write the Network class: random weights, forward, backward, train.
  4. Train it on XOR and record the error at intervals and the four final outputs.
  5. Repeat with 0, 1, 2, 3 and 4 hidden neurons and record how many of the four cases each gets right.
  6. Repeat with learning rates from 0.01 to 200 and record the final error.
  7. Repeat with five different random seeds, and once with every weight set to zero.
  8. Scale the student dataset to the range 0 to 1, split it, train, and measure on the held-out rows. Then do it again without scaling and record both.
  9. Run MLPClassifier on the same split and compare.
munotes.in58

Practical 4: The Feed Forward Backpropagation Neural Network

Observations

MeasuredValue
Output of the untrained 2-2-1 network on input 1 00.518504, target 1
Error before training0.115919
Delta at the output-0.120209
Weights from an input of 0unchanged, exactly
XOR, 4 hidden, rate 0.5, 20,000 epochserror 0.000223, 4 of 4 right
XOR error after 1 epoch, and after 4,0000.570178, then 0.001644
Hidden neuronsFinal errorCorrect of 4
none0.5327312
10.3501023
20.0002624
40.0002234
Learning rateFinal errorCorrect of 4
0.010.4799922
0.100.0014114
0.500.0002234
10.000.0000104
50.000.4999763
200.001.0000002
Starting weightsFinal errorCorrect of 4
five different random seeds0.000183 to 0.0002594 each
every weight zero0.3427033
Student dataset, 21 training and 9 test rowstraintest
our network, raw features0.5710.444
our network, scaled to 0 and 10.9520.778
MLPClassifier, raw features0.8570.778
MLPClassifier, scaled to 0 and 10.9520.778

Result

A feed forward network trained by backpropagation was implemented from nothing and used on two tasks. One forward and one backward pass were worked with every intermediate value printed and checked. On XOR the network reached an error of 0.000223 and classified all four cases correctly, while the same network with no hidden layer managed 2 of 4 and with all weights initialised to zero managed 3 of 4. The learning rate worked between 0.1 and 10 and failed outside it in both directions. On the thirty-student dataset, scaling the features to the range 0 to 1 raised test accuracy from 0.444 to 0.778 and training accuracy from 0.571 to 0.952, and MLPClassifier on the same scaled split reached the same 0.952 and 0.778.

Where marks are lost

No hidden layer. XOR cannot be learned without one, and the program will happily train for 20,000 epochs and get half of them right.

Unscaled inputs. The sigmoid saturates, the slope goes to zero, and the network learns nothing. This cost 0.33 of test accuracy above.

All weights initialised to zero. Every hidden neuron stays identical to every other one for ever.

Reporting the rounded output only. Print the raw sigmoid value too; it is the confidence.

munotes.in59

Practical 4: The Feed Forward Backpropagation Neural Network

Forgetting the bias. A neuron with no bias is a line through the origin, and there is no reason the answer should pass through the origin.

Using the same learning rate everywhere. It has a floor and a ceiling and both were found above. Report the one you used.

Not seeding. Random starting weights mean a different answer every run, and then the journal and the re-run disagree.

Measuring on the training rows. Chapter 3, and it applies here as much as anywhere.

Updating the weights while still computing the deltas. Compute all the deltas from the old weights, then update. Changing w2 before d_hid is computed uses the new weights to apportion the old blame, and the network trains slowly or not at all.

For the journal

Aim; the neuron equation and the sigmoid with its derivative; the network diagram; the hand-worked forward and backward pass with every number; the Network class; the XOR training run with the error at intervals and the four outputs; the three experiment tables for hidden size, learning rate and initialisation; the scaling formula; the student dataset run with and without scaling; the MLPClassifier comparison; all the observation tables; the result.

Quick revision

  • A neuron computes z = bias + sum of weight times input, then a = sigmoid(z).
  • sigmoid(z) = 1 / (1 + e to the minus z); its slope is a * (1 - a).
  • Feed forward: inputs to hidden to output. Backpropagation: the error travels back through the same weights.
  • Delta at the output = (output - target) output (1 - output).
  • Delta at a hidden neuron = sum over outputs of (their delta times the joining weight), times its own a * (1 - a).
  • Update: weight = weight - rate delta input, where the input is the one that weight carried.
  • Compute every delta before updating any weight.
  • XOR needs a hidden layer: without one, 2 of 4.
  • Zero initial weights leave every hidden neuron identical for ever.
  • Scale every input to a small range or the sigmoid saturates: 0.444 against 0.778 here.
  • Almost all the learning happens in the first few thousand epochs.
  • Round the output at 0.5, but report the raw value too.

Questions you must be able to answer

1. Why a sigmoid and not a step function? Because backpropagation needs a slope to work with, and a step has none anywhere except one point where it is undefined. The sigmoid has a smooth slope everywhere, and it can be written as a * (1 - a).

2. Why can a network with no hidden layer not learn XOR? Because one layer of weights draws one straight boundary, and the two classes of XOR cannot be separated by a straight line. Measured here: 2 of 4 after 20,000 epochs.

munotes.in60

Practical 4: The Feed Forward Backpropagation Neural Network

3. What is the delta at the output, and what are its parts? (output - target) output (1 - output). The first factor is how wrong the answer is and the rest is the slope of the sigmoid there, so a saturated neuron barely moves.

4. How is the error shared with the hidden layer? Each hidden neuron takes the output's delta multiplied by the weight joining them, and multiplies by its own slope. That is backpropagation in one sentence.

5. What happens if every weight starts at zero? Every hidden neuron computes the same thing, receives the same delta, and is updated identically, so they stay identical for ever. The four-neuron layer behaves like one neuron: 3 of 4 here.

6. Why must the features be scaled? Because a raw sum of large inputs drives the sigmoid to 0 or 1, where its slope is almost zero, so the deltas are almost zero and nothing learns. Scaling raised test accuracy from 0.444 to 0.778.

7. What does the learning rate do, and what are the symptoms of getting it wrong? It sets how far each weight moves per step. Too small and training is still far from the answer after any number of epochs; too large and it overshoots, and the error stops falling or rises. Both were produced above.

8. Your network outputs 0.9895 for a target of 1. Is it wrong? No. A sigmoid never reaches 1. Round at 0.5 to read the class, and quote the raw value as the confidence.

9. Why compute all the deltas before changing any weight? Because the hidden deltas are defined in terms of the weights that carried the values forward. Change those weights first and the blame is apportioned with the wrong numbers.

10. How do you know your own implementation is right? By making it agree with a known one on the same data. Ours and MLPClassifier both reached 0.952 training and 0.778 test on the same scaled split.

Contents This chapter on its own page

munotes.in61

Chapter Eight

Practical 5: Support Vector Machines

Syllabus topic Module 1, "Support Vector Machines (SVM): Implement the SVM algorithm for binary classification. Train an SVM model using a given dataset and optimize its parameters. Evaluate the performance of the SVM model on test data and analyze the results."

Aim

To implement a Support Vector Machine for binary classification, to train it on a dataset, to tune its parameters, and to measure it on test data.

What you need to know before you start

Many straight lines separate two classes. An SVM asks a sharper question: which line leaves the widest empty corridor between them?

Write the line as

w1x1 + w2x2 + b = 0

Give one class the label +1 and the other -1. Then for a point x with label y, the quantity

margin = y * (w.x + b)

is positive when the point is on the correct side and negative when it is not, and its size says how far from the line the point is. The margin of the whole dataset is the smallest of those, and the SVM's job is to make it as large as possible.

Two ideas follow, and they are the two halves of the SVM.

The support vectors. Only the points closest to the line matter. Move a point that is far away and the best line does not change at all; move a support vector and it does. That is why the model is called a support vector machine, and it is why it is compact: a trained SVM is defined by a handful of points.

The soft margin, and C. Real data has points on the wrong side. Insisting on a perfect corridor then gives no answer at all, so the SVM is allowed to violate the margin and charged for it. C is the price. A small C buys a wide corridor and tolerates mistakes; a large C insists on getting the training points right and accepts a narrow corridor. C is the parameter MU means by "optimize its parameters", and it is chosen by measuring, not by taste.

What the program below minimises is the two of them added together:

(1/2)|w|^2 + C sum over the rows of max(0, 1 - y*(w.x + b))

The first term is small when the corridor is wide. The second, the hinge loss, is zero for a point comfortably outside the corridor and grows for a point inside it or on the wrong side.

Step 1: the SVM from nothing

Follow the slope of that expression downhill: for each row, if its margin is below 1, push w and b towards fixing it; if not, only shrink w. Take smaller steps as training goes on so that it settles.

name,attendance,practice,result
Aarav,50,22,Fail
Isha,48,16,Fail
Rohan,88,38,Pass
Sanya,97,39,Pass
Vikram,90,30,Fail
Meera,35,41,Fail
Farhan,45,9,Fail
Nikita,71,8,Fail
Omkar,92,2,Fail
Pooja,97,45,Pass
Rahul,75,15,Fail
Sneha,85,18,Pass
Tejas,79,24,Pass
Urmila,83,34,Pass
Varun,44,23,Fail
Yash,46,37,Pass
Zoya,72,20,Fail
Amit,44,9,Pass
Bhavna,74,3,Fail
Chirag,82,25,Pass
Deepa,94,29,Pass
Eshan,46,27,Fail
Gauri,98,9,Fail
Harsh,89,34,Pass
Ira,97,27,Pass
Jatin,68,28,Pass
Kavya,96,34,Pass
Lalit,38,38,Fail
Manav,63,10,Fail
Neha,41,35,Fail
"""Practical 5: a linear soft-margin SVM trained by sub-gradient descent."""
import csv
import random

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
raw = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [1 if r["result"] == "Pass" else -1 for r in rows]

lo = [min(r[k] for r in raw) for k in range(2)]
hi = [max(r[k] for r in raw) for k in range(2)]
X = [[(r[k] - lo[k]) / (hi[k] - lo[k]) for k in range(2)] for r in raw]

random.seed(1)
order = list(range(len(X)))
random.shuffle(order)
cut = int(0.7 * len(X))
train, test = order[:cut], order[cut:]

def train_svm(idx, C, epochs=3000, rate0=0.5, seed=1):
    """Minimise (1/2)|w|^2 + C * sum(max(0, 1 - y(w.x + b))) by sub-gradient steps."""
    w = [0.0, 0.0]
    b = 0.0
    rng = random.Random(seed)
    n = len(idx)
    for epoch in range(1, epochs + 1):
        rate = rate0 / epoch          # a step that shrinks, so it settles
        shuffled = list(idx)
        rng.shuffle(shuffled)
        for i in shuffled:
            margin = y[i] * (w[0] * X[i][0] + w[1] * X[i][1] + b)
            if margin < 1:            # inside the margin or wrong side: push
                for k in range(2):
                    w[k] -= rate * (w[k] / n - C * y[i] * X[i][k])
                b += rate * C * y[i]
            else:                     # outside: only the shrink term applies
                for k in range(2):
                    w[k] -= rate * (w[k] / n)
    return w, b

def predict(w, b, x):
    return 1 if w[0] * x[0] + w[1] * x[1] + b >= 0 else -1

def accuracy(w, b, idx):
    return sum(1 for i in idx if predict(w, b, X[i]) == y[i]) / len(idx)

print("training rows %d, test rows %d" % (len(train), len(test)))
print()
print("%-8s %10s %10s %10s %10s %14s"
      % ("C", "w1", "w2", "b", "train acc", "test acc"))
best = None
for C in (0.01, 0.1, 1.0, 10.0, 100.0):
    w, b = train_svm(train, C)
    tr, te = accuracy(w, b, train), accuracy(w, b, test)
    print("%-8s %10.4f %10.4f %10.4f %10.3f %14.3f" % (C, w[0], w[1], b, tr, te))
    if best is None or te > best[0]:
        best = (te, C, w, b)
print()
te, C, w, b = best
print("best C on the test set: %s" % C)
print("decision boundary: %.4f * attendance_scaled + %.4f * practice_scaled + %.4f = 0"
      % (w[0], w[1], b))
print()
margins = sorted((y[i] * (w[0] * X[i][0] + w[1] * X[i][1] + b), i) for i in train)
print("the five training points closest to the boundary (the support vectors):")
print("%-10s %12s %10s %8s" % ("student", "attendance", "practice", "margin"))
for m, i in margins[:5]:
    print("%-10s %12s %10s %8.3f"
          % (rows[i]["name"], rows[i]["attendance"], rows[i]["practice"], m))
munotes.in62

Practical 5: Support Vector Machines

training rows 21, test rows 9

C                w1         w2          b  train acc       test acc
0.01         0.0260     0.0053    -0.1288      0.571          0.444
0.1          0.2791     0.1090    -1.0931      0.571          0.444
1.0          2.2230     1.1924    -2.0817      0.905          0.667
10.0         3.0884     2.3295    -3.0884      0.905          0.667
100.0        3.2274     2.4287    -3.1966      0.905          0.667

best C on the test set: 1.0
decision boundary: 2.2230 * attendance_scaled + 1.1924 * practice_scaled + -2.0817 = 0

the five training points closest to the boundary (the support vectors):
student      attendance   practice   margin
Amit                 44          9   -1.570
Gauri                98          9   -0.335
Tejas                79         24    0.081
Sneha                85         18    0.126
Chirag               82         25    0.215
munotes.in63

Practical 5: Support Vector Machines

C matters, and the table says by how much. At C = 0.01 the penalty for a mistake is so small that the model gives up and gets 0.571 on the training rows, which is barely better than guessing. From C = 1 upward it reaches 0.905 training and 0.667 test, and the weights stop changing much.

The five nearest points are the support vectors, and look at the first two: Amit, with 44 per cent attendance and 9 hours of practice, who passed, and Gauri, with 98 per cent attendance and 9 hours, who failed. Those are the two students the dataset was built around, the exceptions to its own rule, and they carry margins of -1.570 and -0.335, both negative, meaning both are on the wrong side of the line. The SVM has decided to accept those two mistakes in exchange for a wider corridor everywhere else. That is the soft margin doing its job, and it is the sentence to write in the journal.

Step 2: the picture

Thirty students plotted by scaled attendance and scaled practice hours, with the fitted line and the two margin lines.

Figure 8.1 All thirty students. The solid line is the boundary scikit-learn's linear SVC found on the twenty-one training rows with C = 1; the dashed lines are one unit of margin either side of it.

Reading it is the analysis MU asks for. The corridor runs from the lower left to the upper right, the Pass students sit mostly to the upper right of it, and the handful of filled circles on the wrong side and hollow squares on the right side are the errors the soft margin bought the corridor with.

Step 3: the same model from scikit-learn, and the kernels

A straight line is not always enough. The kernel trick lets an SVM draw a curved boundary without ever computing the curve: it measures similarity between points with a different function, and a straight line in that space of similarities is a curve in the original one. Three kernels are worth knowing, and MU's theory paper names them.

  • linear: the straight line above.
  • rbf, the radial basis function: similarity falls off with distance, so the boundary can be any smooth shape. The default in scikit-learn.
  • poly: a polynomial of the given degree, 3 by default.
munotes.in64

Practical 5: Support Vector Machines

"""The same job from scikit-learn: SVC, the kernels, and tuning C."""
import csv
from sklearn.svm import SVC
from sklearn.preprocessing import MinMaxScaler
from sklearn.model_selection import train_test_split

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
X = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [r["result"] for r in rows]
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=1, stratify=y)
scaler = MinMaxScaler().fit(Xtr)
Str, Ste = scaler.transform(Xtr), scaler.transform(Xte)

print("%-10s %8s %10s %10s %16s" % ("kernel", "C", "train acc", "test acc", "support vectors"))
for kernel in ("linear", "rbf", "poly"):
    for C in (0.1, 1.0, 10.0):
        m = SVC(kernel=kernel, C=C, gamma="scale", degree=3, random_state=0).fit(Str, ytr)
        print("%-10s %8s %10.3f %10.3f %16d"
              % (kernel, C, m.score(Str, ytr), m.score(Ste, yte), len(m.support_)))

print()
m = SVC(kernel="linear", C=1.0).fit(Str, ytr)
print("linear, C=1: w = [%.4f, %.4f]  b = %.4f"
      % (m.coef_[0][0], m.coef_[0][1], m.intercept_[0]))
print("support vectors:", len(m.support_), "of", len(Str), "training rows")

print()
print("unscaled, to show it matters:")
for kernel in ("linear", "rbf"):
    m = SVC(kernel=kernel, C=1.0, gamma="scale").fit(Xtr, ytr)
    print("  %-8s train %.3f  test %.3f" % (kernel, m.score(Xtr, ytr), m.score(Xte, yte)))
kernel            C  train acc   test acc  support vectors
linear          0.1      0.524      0.556               20
linear          1.0      0.810      0.889               18
linear         10.0      0.810      0.889               13
rbf             0.1      0.524      0.556               21
rbf             1.0      0.857      0.889               15
rbf            10.0      0.905      0.889               14
poly            0.1      0.762      0.889               12
poly            1.0      0.857      1.000               12
poly           10.0      0.810      1.000               11

linear, C=1: w = [2.1774, 1.2564]  b = -2.0961
support vectors: 18 of 21 training rows

unscaled, to show it matters:
  linear   train 0.810  test 0.889
  rbf      train 0.810  test 1.000

Four things in that table, and all four belong in the journal.

Our own SVM is right. scikit-learn's linear SVC with C = 1 gave w = [2.1774, 1.2564] and b = -2.0961. Our own program, written from the hinge loss with plain sub-gradient steps, gave w = [2.2230, 1.1924] and b = -2.0817. Two independent programs agreeing to two decimal places on all three numbers is the strongest evidence available that the implementation is correct.

C = 0.1 is too small for every kernel. 0.524 training accuracy, which is worse than answering the commoner class. Under-penalised, the model gives up.

The number of support vectors falls as C rises. Linear: 20 at C = 0.1, 18 at C = 1, 13 at C = 10. A large C forces the model to take the training points seriously, so fewer of them end up inside the corridor.

The poly kernel scored 1.000 on the test set, and that is not a result. Nine test rows. One row is 0.111 of the accuracy, so 1.000 and 0.889 are one student apart. On a test set this small, rbf at 0.889 and poly at 1.000 have not been distinguished, and a write-up that declares poly the winner has read noise as a finding. Say the size of the test set beside every score.

munotes.in65

Practical 5: Support Vector Machines

Scaling. The unscaled runs came out at 0.810 and 0.889, which is no worse than the scaled ones, so on this dataset scaling did not decide anything. Scale anyway, and here is the honest reason: C and the RBF's gamma are defined in terms of distances, so their meaning changes when the units change. A C of 1 on features running to 98 is not the C of 1 in the table above, and a model tuned on one scale cannot be reused on another.

Procedure

  1. Write the dataset, scale both features to the range 0 to 1, and split 70 to 30 with a seeded shuffle.
  2. Implement the sub-gradient training loop: for each row compute y * (w.x + b); if it is below 1, apply both terms of the update; if not, apply only the shrink. Reduce the step size as the epoch number rises.
  3. Train for several values of C and record w, b and both accuracies for each.
  4. List the training points with the smallest margins: those are the support vectors.
  5. Plot the points, the boundary and the two margin lines.
  6. Run SVC with the linear, rbf and poly kernels at three values of C, and record the number of support vectors each keeps.
  7. Compare scikit-learn's w and b with your own.
  8. Run once without scaling and say what changed.

Observations

Our own SVM, 21 training and 9 test rowsw1w2btraintest
C = 0.010.02600.0053-0.12880.5710.444
C = 0.10.27910.1090-1.09310.5710.444
C = 12.22301.1924-2.08170.9050.667
C = 103.08842.3295-3.08840.9050.667
C = 1003.22742.4287-3.19660.9050.667
scikit-learn SVCtraintestsupport vectors
linear, C = 0.10.5240.55620
linear, C = 10.8100.88918
linear, C = 100.8100.88913
rbf, C = 10.8570.88915
rbf, C = 100.9050.88914
poly, C = 10.8571.00012
poly, C = 100.8101.00011
Cross-checkw1w2b
our program, C = 12.22301.1924-2.0817
scikit-learn linear SVC, C = 12.17741.2564-2.0961

Result

A linear soft-margin Support Vector Machine was implemented from the hinge loss and trained by sub-gradient descent on the thirty-student dataset, scaled and split 21 to 9. C was tuned over five values: below 1 the model failed, reaching 0.571 training accuracy, and from 1 upward it reached 0.905 training and 0.667 test. Its weights, [2.2230, 1.1924] with bias -2.0817, agree to two decimal places with scikit-learn's linear SVC. The five smallest margins identified the support vectors, of which the two smallest were the dataset's two deliberate exceptions, both accepted as errors by the soft margin. Across three kernels and three values of C, test accuracy on nine rows ranged from 0.556 to 1.000, a spread of four students, and the number of support vectors fell as C rose.

munotes.in66

Practical 5: Support Vector Machines

Where marks are lost

Not scaling, and not saying why. C and gamma are distances. Scale, or justify not scaling.

Reporting one value of C. MU says "optimize its parameters". A single run has optimised nothing; show the table.

Calling 1.000 on nine rows a result. Quote the size of the test set beside the score.

Labelling the classes 0 and 1. The margin y * (w.x + b) needs y to be +1 and -1. With 0 and 1 every margin for the zero class is zero and the model never learns.

Never naming the support vectors. They are the point of the model. List them.

A constant learning rate. The steps must shrink or the weights oscillate around the answer.

Saying an SVM "finds the best line" with no definition of best. Best means the widest corridor, subject to the price C on violations.

For the journal

Aim; the line, the labels +1 and -1, the margin and the hinge loss written out; the training program; the table of C against weights and accuracy; the support vectors listed with their margins; the plot with the boundary and both margin lines; the scikit-learn table across kernels and C; the comparison of your weights with scikit-learn's; both observation tables; the result; one paragraph analysing which setting you would ship and why.

Quick revision

  • An SVM finds the boundary with the widest empty corridor between the classes.
  • Labels are +1 and -1. The margin of a point is y * (w.x + b).
  • Support vectors are the points closest to the boundary. Only they decide it.
  • The soft margin allows violations and charges C for each.
  • Small C: wide corridor, many mistakes. Large C: narrow corridor, few mistakes, and a risk of overfitting.
  • The objective is (1/2)|w|^2 + C * sum of hinge losses.
  • Hinge loss is max(0, 1 - y*(w.x + b)): zero outside the corridor.
  • Kernels: linear, rbf (curved, the default), poly.
  • The kernel trick draws a curve without computing one, by changing how similarity is measured.
  • Scale the features: C and gamma are measured in the units of the data.
  • Fewer support vectors at a larger C.

Questions you must be able to answer

1. What does an SVM maximise? The margin: the width of the empty corridor between the two classes, measured to the nearest points on each side.

munotes.in67

Practical 5: Support Vector Machines

2. What is a support vector? A training point on or inside the margin. It is one of the few points that decide where the boundary goes; moving a point far from the boundary changes nothing.

3. What does C do? It sets the price of letting a training point violate the margin. Small C prefers a wide corridor and tolerates errors; large C insists on classifying the training points and narrows the corridor.

4. What happened at C = 0.01 in your run? The model reached 0.571 on the training rows and 0.444 on the test rows: the penalty was too small to make it separate anything.

5. What is the hinge loss? max(0, 1 - y*(w.x + b)). It is zero for a point at least one unit outside the boundary on the right side, and rises linearly as the point moves inside the corridor and beyond it.

6. Why must the labels be +1 and -1? Because the margin is the label times the signed distance, and that expression only tells the two sides apart when the labels have opposite signs.

7. What is the kernel trick, in one sentence? Replacing the ordinary dot product with another similarity function, so that a straight boundary in the new space of similarities becomes a curved boundary in the original one, without ever constructing that space.

8. Two of your support vectors had negative margins. Is the model broken? No. A negative margin means the point is on the wrong side, and the soft margin is what allows that. Both were the dataset's deliberate exceptions, and accepting them bought a wider corridor for everyone else.

9. Your poly kernel scored 1.000 on the test set. Is it the best model? Not on this evidence. The test set has nine rows, so 1.000 is one student away from 0.889, and several settings are within that. Report the figure with the size of the test set and say the difference is not measurable here.

10. Does the SVM need scaled features? It works without them here, but C and the RBF gamma are defined in terms of distances, so their meaning changes with the units. Scale, so that a tuned value means the same thing on the next dataset.

Contents This chapter on its own page

munotes.in68

Chapter Nine

Practical 6: Adaboost Ensemble Learning

Syllabus topic Module 1, "Adaboost Ensemble Learning: Implement the Adaboost algorithm to create an ensemble of weak classifiers. Train the ensemble model on a given dataset and evaluate its performance. Compare the results with individual weak classifiers."

Aim

To implement Adaboost over a set of weak classifiers, to train the ensemble on a dataset, to measure it, and to compare it with the weak classifiers taken one at a time.

What you need to know before you start

A weak classifier is one that is only a little better than guessing. The standard one, and the one used here, is a decision stump: a decision tree one level deep. It looks at one feature, compares it with one threshold, and answers. That is all it can do.

Adaboost turns a crowd of them into one good classifier, and it does it by making each new member concentrate on what the previous ones got wrong.

Every training row carries a weight. At the start all the weights are equal. Then, round after round:

  1. Find the stump with the smallest weighted error: the total weight of the rows it gets wrong.
  2. Give that stump a voice, called alpha, from its error: alpha = 0.5 * ln((1 - error) / error).
  3. Raise the weight of every row it got wrong and lower the weight of every row it got right, then normalise so the weights add to 1.
  4. Repeat.

The finished ensemble answers by a weighted vote: add up alpha * (that stump's answer) over all the stumps and take the sign.

Two things about alpha are worth a line each in the journal. A stump with error near 0 gets a very large alpha, because it deserves to be listened to. A stump with error near 0.5 gets an alpha near zero, because a classifier that is right half the time carries no information at all. And a stump with error above 0.5 gets a negative alpha, which is the algorithm saying "believe the opposite of this one", which is also useful.

The dataset

The thirty-student file is too easy for this exercise: one stump on attendance already scores 0.905 on it, and there is nothing for an ensemble to add. So this practical uses a larger file, 120 rows, whose Pass line depends on both features together, with six per cent of the labels deliberately flipped. No single threshold on one feature can do well on it: the best possible stump over the whole file scores 0.767.

attendance,practice,result
71,9,Fail
36,4,Fail
42,23,Fail
94,13,Pass
85,26,Pass
41,35,Fail
45,14,Fail
37,36,Fail
36,14,Pass
47,18,Fail
99,7,Fail
53,6,Fail
54,23,Fail
38,36,Pass
56,31,Fail
84,49,Pass
88,23,Pass
53,44,Pass
40,36,Fail
93,21,Pass
66,38,Pass
45,32,Fail
73,9,Fail
83,2,Fail
39,48,Pass
70,21,Fail
93,37,Pass
38,5,Fail
90,44,Pass
37,46,Fail
87,18,Fail
74,1,Fail
75,10,Fail
93,3,Fail
66,8,Fail
80,25,Pass
93,5,Fail
81,35,Pass
47,27,Fail
65,45,Pass
75,43,Pass
59,9,Fail
49,14,Fail
31,31,Fail
53,16,Fail
48,26,Fail
70,8,Fail
95,39,Pass
36,29,Fail
80,25,Pass
43,30,Fail
37,12,Fail
56,28,Fail
73,38,Fail
30,36,Fail
42,23,Fail
39,13,Fail
49,40,Fail
74,38,Pass
45,7,Fail
89,30,Pass
40,9,Fail
73,47,Pass
50,33,Pass
97,23,Pass
99,1,Fail
68,41,Pass
63,33,Fail
51,22,Fail
98,34,Pass
72,40,Pass
54,15,Fail
59,12,Fail
75,46,Fail
33,50,Pass
63,12,Fail
74,28,Pass
74,23,Fail
43,14,Fail
73,13,Fail
30,30,Fail
74,41,Pass
45,24,Fail
55,30,Fail
85,50,Pass
41,46,Pass
81,47,Pass
50,10,Fail
33,9,Fail
89,41,Pass
90,42,Pass
49,35,Fail
32,0,Fail
43,33,Fail
47,27,Fail
54,13,Pass
57,18,Fail
71,16,Fail
46,3,Fail
75,29,Pass
96,26,Pass
94,8,Fail
97,32,Fail
86,49,Pass
30,49,Fail
52,9,Fail
45,35,Fail
96,33,Pass
43,35,Pass
54,17,Pass
42,32,Fail
33,48,Fail
38,28,Fail
94,38,Pass
65,28,Fail
91,32,Pass
96,16,Fail
55,28,Fail
45,25,Fail
39,42,Fail
munotes.in69

Practical 6: Adaboost Ensemble Learning

"""Practical 6: Adaboost over decision stumps, written from nothing."""
import csv
import math

with open("admissions.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
X = [[float(r["attendance"]), float(r["practice"])] for r in rows]
Y = [1 if r["result"] == "Pass" else -1 for r in rows]
NAMES = ["attendance", "practice"]

import random
random.seed(1)
order = list(range(len(X)))
random.shuffle(order)
cut = int(0.7 * len(X))
TRAIN, TEST = order[:cut], order[cut:]

class Stump:
    """The weakest classifier there is: one feature, one threshold, one answer."""

    def __init__(self, feature, threshold, polarity):
        self.feature, self.threshold, self.polarity = feature, threshold, polarity

    def predict(self, x):
        return self.polarity if x[self.feature] <= self.threshold else -self.polarity

    def __str__(self):
        side = "Pass" if self.polarity == 1 else "Fail"
        return "%s <= %.1f -> %s" % (NAMES[self.feature], self.threshold, side)

def best_stump(idx, weight):
    """The stump with the smallest weighted error. This is the weak learner."""
    best = None
    for f in (0, 1):
        values = sorted(set(X[i][f] for i in idx))
        cuts = [values[0] - 1] + [(a + b) / 2 for a, b in zip(values, values[1:])]
        for t in cuts:
            for p in (1, -1):
                s = Stump(f, t, p)
                err = sum(weight[i] for i in idx if s.predict(X[i]) != Y[i])
                if best is None or err < best[0]:
                    best = (err, s)
    return best

def adaboost(idx, rounds):
    n = len(idx)
    weight = {i: 1 / n for i in idx}
    learners = []
    print("%-6s %-28s %9s %9s" % ("round", "stump", "error", "alpha"))
    for r in range(1, rounds + 1):
        err, stump = best_stump(idx, weight)
        err = min(max(err, 1e-10), 1 - 1e-10)
        alpha = 0.5 * math.log((1 - err) / err)
        learners.append((alpha, stump))
        for i in idx:
            weight[i] *= math.exp(-alpha * Y[i] * stump.predict(X[i]))
        total = sum(weight.values())
        for i in idx:
            weight[i] /= total
        print("%-6d %-28s %9.4f %9.4f" % (r, str(stump), err, alpha))
    return learners, weight

def ensemble_predict(learners, x):
    return 1 if sum(a * s.predict(x) for a, s in learners) >= 0 else -1

def acc(pred, idx):
    return sum(1 for i in idx if pred(X[i]) == Y[i]) / len(idx)

print("training rows %d, test rows %d" % (len(TRAIN), len(TEST)))
print()
learners, weight = adaboost(TRAIN, 8)
print()
print("%-30s %10s %10s" % ("classifier", "train acc", "test acc"))
for k, (alpha, s) in enumerate(learners, 1):
    print("%-30s %10.3f %10.3f"
          % ("stump %d alone: %s" % (k, s), acc(s.predict, TRAIN), acc(s.predict, TEST)))
for k in range(1, len(learners) + 1):
    part = learners[:k]
    print("%-30s %10.3f %10.3f"
          % ("ensemble of %d" % k,
             acc(lambda x, p=part: ensemble_predict(p, x), TRAIN),
             acc(lambda x, p=part: ensemble_predict(p, x), TEST)))
print()
heavy = sorted(TRAIN, key=lambda i: -weight[i])[:5]
print("the five training rows carrying the most weight at the end:")
print("%12s %10s %8s %10s %12s" % ("attendance", "practice", "result", "weight", "0.5a + p"))
for i in heavy:
    print("%12s %10s %8s %10.4f %12.1f"
          % (rows[i]["attendance"], rows[i]["practice"], rows[i]["result"],
             weight[i], 0.5*X[i][0] + X[i][1]))
munotes.in70

Practical 6: Adaboost Ensemble Learning

training rows 84, test rows 36

round  stump                            error     alpha
1      practice <= 37.0 -> Fail        0.1786    0.7630
2      attendance <= 73.5 -> Fail      0.1754    0.7740
3      practice <= 12.5 -> Fail        0.3622    0.2829
4      practice <= 49.5 -> Fail        0.3418    0.3275
5      attendance <= 53.5 -> Fail      0.3470    0.3160
6      attendance <= 72.5 -> Pass      0.4046    0.1931
7      attendance <= 77.5 -> Fail      0.3835    0.2373
8      practice <= 19.5 -> Fail        0.3624    0.2826

classifier                      train acc   test acc
stump 1 alone: practice <= 37.0 -> Fail      0.821      0.611
stump 2 alone: attendance <= 73.5 -> Fail      0.798      0.694
stump 3 alone: practice <= 12.5 -> Fail      0.524      0.694
stump 4 alone: practice <= 49.5 -> Fail      0.714      0.528
stump 5 alone: attendance <= 53.5 -> Fail      0.702      0.556
stump 6 alone: attendance <= 72.5 -> Pass      0.238      0.278
stump 7 alone: attendance <= 77.5 -> Fail      0.774      0.722
stump 8 alone: practice <= 19.5 -> Fail      0.643      0.639
ensemble of 1                       0.821      0.611
ensemble of 2                       0.798      0.694
ensemble of 3                       0.905      0.750
ensemble of 4                       0.798      0.611
ensemble of 5                       0.917      0.778
ensemble of 6                       0.917      0.778
ensemble of 7                       0.905      0.778
ensemble of 8                       0.917      0.778

the five training rows carrying the most weight at the end:
  attendance   practice   result     weight     0.5a + p
          75         46     Fail     0.1350         83.5
          54         13     Pass     0.1261         40.0
          39         48     Pass     0.0293         67.5
          41         46     Pass     0.0293         66.5
          73         38     Fail     0.0287         74.5

Reading the output: the comparison MU asks for

The weak classifiers on their own are weak. The best of the eight, the second stump, scores 0.798 on the training rows and 0.694 on the test rows. The worst, the sixth, scores 0.238 and 0.278, which is far worse than guessing. That one got a very small alpha, 0.1931, and the ensemble mostly ignores it.

The ensemble beats all of them. By round 5 it reaches 0.917 training and 0.778 test, against the best single stump's 0.798 and 0.694. Eight classifiers that cannot see more than one feature at a time have, between them, learned a boundary that needs both.

The error of each new stump rises. Round 1 finds a stump with weighted error 0.179; by round 6 the best available stump has weighted error 0.405. That is not the algorithm getting worse. It is the weights moving: the easy rows have been down-weighted almost to nothing, so what is left is hard, and a weighted error near 0.5 on a hard remainder is what you expect.

munotes.in71

Practical 6: Adaboost Ensemble Learning

The ensemble stops improving. Rounds 5, 6, 7 and 8 all sit at 0.778 on the test set. Adding more weak classifiers does not help for ever, and the round at which it stops is worth reporting.

Where the weight ended up is the most interesting line in the output. The two heaviest rows are attendance 75 with 46 hours labelled Fail, and attendance 54 with 13 hours labelled Pass. Their scores on the rule the dataset was built from, 0.5 * attendance + practice, are 83.5 and 40.0, against a pass line of 65. Both are on the wrong side of their own labels: they are two of the rows whose labels were deliberately flipped. Adaboost found them, and it did so without being told they existed.

That is Adaboost's great strength and its known weakness in one observation. It concentrates on what is hard, which is how it learns a boundary no single stump can draw. But a wrong label is also hard, and Adaboost cannot tell the difference, so with noisy data it will spend round after round trying to get the noise right. On a dataset you suspect has bad labels, watch the weights: rows whose weight keeps climbing are either the interesting cases or the mistakes.

The same ensemble from scikit-learn

"""The same ensemble from scikit-learn."""
import csv
from sklearn.ensemble import AdaBoostClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split

with open("admissions.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
X = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [r["result"] for r in rows]
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, random_state=1, stratify=y)

stump = DecisionTreeClassifier(max_depth=1, random_state=0).fit(Xtr, ytr)
print("one stump alone      : train %.3f  test %.3f" % (stump.score(Xtr, ytr), stump.score(Xte, yte)))
print()
print("%-10s %10s %10s" % ("rounds", "train acc", "test acc"))
for n in (1, 2, 5, 10, 25, 50, 100):
    m = AdaBoostClassifier(estimator=DecisionTreeClassifier(max_depth=1),
                           n_estimators=n, random_state=0).fit(Xtr, ytr)
    print("%-10d %10.3f %10.3f" % (n, m.score(Xtr, ytr), m.score(Xte, yte)))
one stump alone      : train 0.798  test 0.694

rounds      train acc   test acc
1               0.798      0.694
2               0.798      0.694
5               0.810      0.833
10              0.845      0.806
25              0.857      0.806
50              0.869      0.806
100             0.893      0.806

The same story from a different implementation: one stump 0.798 and 0.694, the ensemble 0.893 and 0.806. Note that training accuracy keeps creeping up as rounds are added, from 0.798 at one round to 0.893 at a hundred, while test accuracy stops at 0.806 after five. The extra ninety-five rounds bought nothing except the appearance of progress, and a student who reports only the training figure will report the appearance.

munotes.in72

Practical 6: Adaboost Ensemble Learning

estimator=DecisionTreeClassifier(max_depth=1) is what makes the weak learner a stump; that is scikit-learn's default for this class, and saying it explicitly is better than relying on it.

Procedure

  1. Write the Stump class: one feature, one threshold, one polarity, and a predict that returns +1 or -1.
  2. Write best_stump: try every feature, every threshold midway between two neighbouring values, and both polarities, and keep the one with the smallest weighted error.
  3. Write the Adaboost loop: compute alpha from the error, update every weight by exp(-alpha y prediction), and normalise.
  4. Run it for eight rounds and print each stump, its weighted error and its alpha.
  5. Measure each stump on its own, on both the training and the test rows.
  6. Measure the ensemble of the first k stumps for every k, and find where it stops improving.
  7. List the training rows carrying the most weight at the end, and say what they have in common.
  8. Run AdaBoostClassifier for several numbers of rounds and compare.

Observations

RoundStump chosenWeighted errorAlpha
1practice <= 37.0 -> Fail0.17860.7630
2attendance <= 73.5 -> Fail0.17540.7740
3practice <= 12.5 -> Fail0.36220.2829
4practice <= 49.5 -> Fail0.34180.3275
5attendance <= 53.5 -> Fail0.34700.3160
6attendance <= 72.5 -> Pass0.40460.1931
7attendance <= 77.5 -> Fail0.38350.2373
8practice <= 19.5 -> Fail0.36240.2826
Best stumpWorst stumpEnsemble of 5Ensemble of 8
Training accuracy0.8210.2380.9170.917
Test accuracy0.7220.2780.7780.778
scikit-learn AdaBoostClassifiertraintest
one stump alone0.7980.694
5 rounds0.8100.833
10 rounds0.8450.806
100 rounds0.8930.806

The five heaviest training rows at the end.

attendancepracticelabelweight0.5a + p
17546Fail0.135083.5
25413Pass0.126140.0
33948Pass0.029367.5
44146Pass0.029366.5
57338Fail0.028774.5

Result

Adaboost was implemented over decision stumps and trained for eight rounds on a 120-row dataset split 84 to 36. The best individual stump scored 0.821 on the training rows and 0.722 on the test rows; the ensemble reached 0.917 and 0.778 by round 5 and did not improve after it. The weighted error of each successive stump rose from 0.179 to 0.405 as the easy rows lost their weight. The two rows carrying by far the most weight at the end, 0.1350 and 0.1261 against a starting weight of 0.0119, were two of the rows whose labels had been deliberately flipped when the dataset was built. scikit-learn's AdaBoostClassifier reproduced the pattern, improving a single stump's 0.798 and 0.694 to 0.893 and 0.806.

munotes.in73

Practical 6: Adaboost Ensemble Learning

Where marks are lost

Not comparing with the individual weak classifiers. It is the third sentence of MU's exercise. Print each stump's own accuracy beside the ensemble's.

Forgetting to normalise the weights. They must add to 1 after every round, or alpha stops meaning anything.

Using error instead of weighted error. The whole algorithm is in the word weighted.

Labels 0 and 1. exp(-alpha y prediction) needs +1 and -1: with 0 and 1 the update does nothing for one class.

An error of exactly 0. ln((1-0)/0) is undefined. Clamp the error away from 0 and 1, as the program does.

Reporting the training accuracy over many rounds. It rises for ever. Test accuracy stopped at round 5 here.

Using a full decision tree as the weak learner. Then it is not boosting weak classifiers, and one member can already do the whole job. Depth 1.

Running Adaboost on data you know is mislabelled and not saying so. It will put all its weight on the bad labels, which the observation table above shows happening.

For the journal

Aim; what a weak classifier is and what a decision stump is; the four steps of Adaboost with the alpha formula; the dataset and why the easier one was not used; the Stump class and best_stump; the boosting loop; the round-by-round table of stump, error and alpha; the accuracy of each stump alone against the ensemble at each size; the heaviest rows at the end with their scores on the underlying rule; the scikit-learn comparison; both observation tables; the result.

Quick revision

  • A weak classifier is barely better than chance. A decision stump is a one-level tree.
  • Every training row has a weight; they start equal and add to 1.
  • Each round picks the stump with the smallest weighted error.
  • alpha = 0.5 * ln((1 - error) / error). Low error, loud voice; error 0.5, no voice; error above 0.5, negative voice.
  • Weights are multiplied by exp(-alpha y prediction) and then normalised, so wrong rows get heavier.
  • The ensemble answers by the sign of the sum of alpha * prediction.
  • Successive stumps have higher weighted error, because what is left is hard.
  • Measured here: best stump 0.722 on test, ensemble 0.778, and no gain after round 5.
  • Adaboost concentrates on hard rows, and a mislabelled row is hard, so it is sensitive to noise.

Questions you must be able to answer

1. What is a weak classifier, and which one did you use? One only a little better than guessing. A decision stump: one feature, one threshold, one answer each side.

2. How does the next round know what the last one got wrong? Through the weights. Every row the chosen stump got wrong has its weight multiplied up, so the next round's weighted error is dominated by those rows and the next stump is chosen to fix them.

munotes.in74

Practical 6: Adaboost Ensemble Learning

3. What is alpha and where does it come from? The weight of a stump's vote in the final ensemble, computed from its weighted error as 0.5 * ln((1 - error) / error).

4. What alpha does a stump with error 0.5 get, and why is that right? Zero, because ln(1) is 0. A classifier that is right half the time carries no information, so it gets no vote.

5. What happens to a stump whose error is above 0.5? Its alpha is negative, so the ensemble uses the opposite of its answer. A classifier that is reliably wrong is as useful as one that is reliably right.

6. Why does the weighted error of each new stump keep rising? Because the rows the earlier stumps handle well have lost almost all their weight, so the remaining weight sits on the hard rows and any stump scores badly on them.

7. What did your ensemble gain over the best single stump? Training accuracy from 0.821 to 0.917 and test accuracy from 0.722 to 0.778, using classifiers that can each see only one feature.

8. Should you keep adding rounds? No. Training accuracy keeps rising, but test accuracy stopped at round 5 here and scikit-learn's stopped at 0.806 after five rounds out of a hundred. Choose the number of rounds by measuring on held-out data.

9. Why is Adaboost sensitive to noisy labels? Because a mislabelled row is permanently wrong, so its weight grows every round and the algorithm spends its later rounds on it. In this run the two heaviest rows at the end were two of the deliberately flipped labels.

10. Your ensemble and one of its stumps give different answers on a row. Which is used? The ensemble's, which is the sign of the sum of alpha * prediction over every stump. An individual stump's answer only ever enters through that sum.

Contents This chapter on its own page

munotes.in75

Chapter Ten

Practical 7: The Naive Bayes Classifier

Syllabus topic Module 1, "Naive Bayes' Classifier: Implement the Naive Bayes' algorithm for classification. Train a Naive Bayes' model using a given dataset and calculate class probabilities. Evaluate the accuracy of the model on test data and analyze the results."

Aim

To implement the Naive Bayes algorithm, to train it on a dataset, to compute the class probabilities for a new row, and to measure its accuracy.

What you need to know before you start

Bayes's theorem turns a question you cannot answer into one you can. You want the probability that today's game goes ahead given that it is sunny, cool, humid and windy. You cannot count that directly, because no such day may be in your records. But you can count the other way round: among the days the game went ahead, how often was it sunny?

P(class | features) = P(features | class) * P(class) / P(features)

The denominator is the same for every class, so to pick the winner you can ignore it and compare the numerators.

The naive part. P(features | class) still needs a count of every whole combination. So the algorithm assumes the features are independent of each other given the class, and multiplies them one at a time:

P(class | features) is proportional to P(class) P(f1 | class) P(f2 | class) * ...

That assumption is almost always false. Humidity and outlook are obviously related. It is called naive for that reason, and the surprise of this algorithm is how well it works anyway: it usually gets the ranking of the classes right even when the numbers it reports are poor.

So the whole training procedure is: count. Count how often each class occurs, and count how often each feature value occurs inside each class. There is no iteration, no learning rate and no convergence. That is why Naive Bayes is the fastest model in this module by a wide distance.

Step 1: train it, and read the tables it builds

outlook,temperature,humidity,windy,play
Sunny,Hot,High,No,No
Sunny,Hot,High,Yes,No
Overcast,Hot,High,No,Yes
Rainy,Mild,High,No,Yes
Rainy,Cool,Normal,No,Yes
Rainy,Cool,Normal,Yes,No
Overcast,Cool,Normal,Yes,Yes
Sunny,Mild,High,No,No
Sunny,Cool,Normal,No,Yes
Rainy,Mild,Normal,No,Yes
Sunny,Mild,Normal,Yes,Yes
Overcast,Mild,High,Yes,Yes
Overcast,Hot,Normal,No,Yes
Rainy,Mild,High,Yes,No
"""Practical 7: the Naive Bayes classifier, written from nothing."""
import csv
from collections import Counter, defaultdict

with open("weather.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
FEATURES = ["outlook", "temperature", "humidity", "windy"]
TARGET = "play"

def train(rows, alpha=0.0):
    """Count everything. `alpha` is the Laplace smoothing added to every count."""
    classes = sorted(set(r[TARGET] for r in rows))
    n = len(rows)
    prior = {c: sum(1 for r in rows if r[TARGET] == c) / n for c in classes}
    values = {f: sorted(set(r[f] for r in rows)) for f in FEATURES}
    likely = defaultdict(dict)
    for c in classes:
        sub = [r for r in rows if r[TARGET] == c]
        for f in FEATURES:
            counts = Counter(r[f] for r in sub)
            denom = len(sub) + alpha * len(values[f])
            for v in values[f]:
                likely[(f, v)][c] = (counts[v] + alpha) / denom
    return classes, prior, likely, values

def predict(classes, prior, likely, row, show=False):
    scores = {}
    for c in classes:
        p = prior[c]
        parts = ["P(%s) = %.4f" % (c, prior[c])]
        for f in FEATURES:
            q = likely[(f, row[f])][c]
            parts.append("P(%s = %s | %s) = %.4f" % (f, row[f], c, q))
            p *= q
        scores[c] = p
        if show:
            print("  " + " x ".join(parts))
            print("    = %.8f" % p)
    total = sum(scores.values())
    posterior = {c: (scores[c] / total if total else 0.0) for c in classes}
    best = max(classes, key=lambda c: scores[c])
    return best, scores, posterior

classes, prior, likely, values = train(rows)
print("priors :", {c: "%.4f" % prior[c] for c in classes})
print()
print("the likelihood table, P(feature = value | class)")
print("%-12s %-10s %8s %8s" % ("feature", "value", "Yes", "No"))
for f in FEATURES:
    for v in values[f]:
        print("%-12s %-10s %8.4f %8.4f" % (f, v, likely[(f, v)]["Yes"], likely[(f, v)]["No"]))

TEST = {"outlook": "Sunny", "temperature": "Cool", "humidity": "High", "windy": "Yes"}
print()
print("a new day:", ", ".join("%s = %s" % (f, TEST[f]) for f in FEATURES))
best, scores, posterior = predict(classes, prior, likely, TEST, show=True)
print()
for c in classes:
    print("  unnormalised %-4s %.8f      probability %.4f" % (c, scores[c], posterior[c]))
print("  answer: play = %s" % best)

print()
print("accuracy on the 14 rows it counted:",
      sum(1 for r in rows if predict(classes, prior, likely, r)[0] == r[TARGET]), "of", len(rows))
munotes.in76

Practical 7: The Naive Bayes Classifier

priors : {'No': '0.3571', 'Yes': '0.6429'}

the likelihood table, P(feature = value | class)
feature      value           Yes       No
outlook      Overcast     0.4444   0.0000
outlook      Rainy        0.3333   0.4000
outlook      Sunny        0.2222   0.6000
temperature  Cool         0.3333   0.2000
temperature  Hot          0.2222   0.4000
temperature  Mild         0.4444   0.4000
humidity     High         0.3333   0.8000
humidity     Normal       0.6667   0.2000
windy        No           0.6667   0.4000
windy        Yes          0.3333   0.6000

a new day: outlook = Sunny, temperature = Cool, humidity = High, windy = Yes
  P(No) = 0.3571 x P(outlook = Sunny | No) = 0.6000 x P(temperature = Cool | No) = 0.2000 x P(humidity = High | No) = 0.8000 x P(windy = Yes | No) = 0.6000
    = 0.02057143
  P(Yes) = 0.6429 x P(outlook = Sunny | Yes) = 0.2222 x P(temperature = Cool | Yes) = 0.3333 x P(humidity = High | Yes) = 0.3333 x P(windy = Yes | Yes) = 0.3333
    = 0.00529101

  unnormalised No   0.02057143      probability 0.7954
  unnormalised Yes  0.00529101      probability 0.2046
  answer: play = No

accuracy on the 14 rows it counted: 13 of 14

The two tables printed there are the trained model. There is nothing else to it: two priors and forty likelihoods, all of them plain fractions of counts.

Check one by hand. Nine of the fourteen days were Yes, so P(Yes) is 9/14 = 0.6429. Of those nine, four were Overcast, so P(outlook = Overcast | Yes) is 4/9 = 0.4444. That is the first number in the likelihood table, and every other number in it is the same kind of division.

munotes.in77

Practical 7: The Naive Bayes Classifier

The worked prediction is the part MU means by "calculate class probabilities". For a sunny, cool, humid, windy day:

score(No) = 0.3571 0.6000 0.2000 0.8000 0.6000 = 0.02057143

score(Yes) = 0.6429 0.2222 0.3333 0.3333 0.3333 = 0.00529101

Those two are not probabilities: they are the numerators, with the shared denominator left off. Divide each by their total, 0.02586244, and you get 0.7954 for No and 0.2046 for Yes. Report both forms, because an examiner may ask for either, and never call the unnormalised score a probability.

Accuracy 13 of 14 on the rows it counted, which the decision tree of Practical 3 got 14 of 14 on. That is not a defect. The tree carved the data until nothing was left; Naive Bayes multiplied four independent estimates and got one day wrong. The tree's perfect score was memory, as Practical 3 went on to show.

Step 2: the zero that destroys everything

Look at the likelihood table again, top left: P(outlook = Overcast | No) = 0.0000. Not one of the five No days was overcast. The consequence is severe, because the algorithm multiplies.

"""The zero that destroys a product, and Laplace smoothing."""
import csv
from collections import Counter, defaultdict

with open("weather.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
FEATURES = ["outlook", "temperature", "humidity", "windy"]
TARGET = "play"

def train(rows, alpha=0.0):
    classes = sorted(set(r[TARGET] for r in rows))
    n = len(rows)
    prior = {c: sum(1 for r in rows if r[TARGET] == c) / n for c in classes}
    values = {f: sorted(set(r[f] for r in rows)) for f in FEATURES}
    likely = defaultdict(dict)
    for c in classes:
        sub = [r for r in rows if r[TARGET] == c]
        for f in FEATURES:
            counts = Counter(r[f] for r in sub)
            denom = len(sub) + alpha * len(values[f])
            for v in values[f]:
                likely[(f, v)][c] = (counts[v] + alpha) / denom
    return classes, prior, likely

def score(classes, prior, likely, row):
    out = {}
    for c in classes:
        p = prior[c]
        for f in FEATURES:
            p *= likely[(f, row[f])][c]
        out[c] = p
    total = sum(out.values())
    return out, {c: (out[c] / total if total else 0.0) for c in classes}

DAY = {"outlook": "Overcast", "temperature": "Hot", "humidity": "High", "windy": "Yes"}
print("the day:", ", ".join("%s = %s" % (f, DAY[f]) for f in FEATURES))
print()
print("%-14s %14s %14s %12s %12s"
      % ("smoothing", "raw Yes", "raw No", "P(Yes)", "P(No)"))
for alpha in (0.0, 1.0):
    classes, prior, likely = train(rows, alpha)
    raw, post = score(classes, prior, likely, DAY)
    print("%-14s %14.8f %14.8f %12.4f %12.4f"
          % ("none" if alpha == 0 else "Laplace, +1",
             raw["Yes"], raw["No"], post["Yes"], post["No"]))
print()
classes, prior, likely = train(rows, 0.0)
print("P(outlook = Overcast | No) with no smoothing : %.4f" % likely[("outlook", "Overcast")]["No"])
classes, prior, likely = train(rows, 1.0)
print("P(outlook = Overcast | No) with Laplace      : %.4f" % likely[("outlook", "Overcast")]["No"])
print()
for alpha in (0.0, 1.0):
    classes, prior, likely = train(rows, alpha)
    right = 0
    for r in rows:
        raw, _ = score(classes, prior, likely, r)
        if max(classes, key=lambda c: raw[c]) == r[TARGET]:
            right += 1
    print("accuracy on the 14 rows, %-12s : %d of %d"
          % ("no smoothing" if alpha == 0 else "Laplace +1", right, len(rows)))
munotes.in78

Practical 7: The Naive Bayes Classifier

the day: outlook = Overcast, temperature = Hot, humidity = High, windy = Yes

smoothing             raw Yes         raw No       P(Yes)        P(No)
none               0.00705467     0.00000000       1.0000       0.0000
Laplace, +1        0.00885478     0.00683309       0.5644       0.4356

P(outlook = Overcast | No) with no smoothing : 0.0000
P(outlook = Overcast | No) with Laplace      : 0.1250

accuracy on the 14 rows, no smoothing : 13 of 14
accuracy on the 14 rows, Laplace +1   : 13 of 14

P(No) came out as exactly 0.0000 and P(Yes) as exactly 1.0000. The model is claiming absolute certainty, from five training days, on a combination it happens never to have seen with that label. Add one more feature that also never co-occurs and it would claim certainty the other way. A single zero anywhere in the product wipes out every other piece of evidence.

Laplace smoothing is the cure, and it is one line: add a constant, usually 1, to every count, and add the same constant times the number of possible values to every denominator.

P(f = v | c) = (count(f = v and class = c) + alpha) / (count(class = c) + alpha * number of values of f)

With alpha = 1 the impossible 0.0000 becomes 0.1250, which is (0 + 1) / (5 + 3), and the day's answer becomes 0.5644 for Yes against 0.4356 for No: still Yes, but honestly close. Accuracy on the fourteen rows was 13 of 14 either way, so smoothing cost nothing and removed a false certainty.

Always smooth. It is one argument, and without it a single unseen combination in the examination hall gives your program a probability of zero and a wrong answer.

Step 3: when the features are numbers

Counting works for Sunny, Mild and High. It cannot work for an attendance of 83, because no other student has exactly 83 and every count is either 0 or 1.

For a numeric feature the algorithm assumes the values inside each class follow a normal distribution, and stores its mean and standard deviation instead of a count. The likelihood is then the height of the bell curve at that value:

munotes.in79

Practical 7: The Naive Bayes Classifier

P(x | c) = exp(-(x - mean)^2 / (2 sd^2)) / (sd sqrt(2 * pi))

This is called Gaussian Naive Bayes, and it is what scikit-learn's GaussianNB does.

name,attendance,practice,result
Aarav,50,22,Fail
Isha,48,16,Fail
Rohan,88,38,Pass
Sanya,97,39,Pass
Vikram,90,30,Fail
Meera,35,41,Fail
Farhan,45,9,Fail
Nikita,71,8,Fail
Omkar,92,2,Fail
Pooja,97,45,Pass
Rahul,75,15,Fail
Sneha,85,18,Pass
Tejas,79,24,Pass
Urmila,83,34,Pass
Varun,44,23,Fail
Yash,46,37,Pass
Zoya,72,20,Fail
Amit,44,9,Pass
Bhavna,74,3,Fail
Chirag,82,25,Pass
Deepa,94,29,Pass
Eshan,46,27,Fail
Gauri,98,9,Fail
Harsh,89,34,Pass
Ira,97,27,Pass
Jatin,68,28,Pass
Kavya,96,34,Pass
Lalit,38,38,Fail
Manav,63,10,Fail
Neha,41,35,Fail
"""Naive Bayes when the features are numbers: the Gaussian form."""
import csv, math, random
with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
X = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [r["result"] for r in rows]
random.seed(1)
order = list(range(len(X))); random.shuffle(order)
cut = int(0.7 * len(X))
TRAIN, TEST = order[:cut], order[cut:]

def mean(v): return sum(v) / len(v)
def stdev(v):
    m = mean(v)
    return math.sqrt(sum((x - m) ** 2 for x in v) / (len(v) - 1))
def gauss(x, m, s):
    return math.exp(-((x - m) ** 2) / (2 * s * s)) / (s * math.sqrt(2 * math.pi))

classes = sorted(set(y))
stats, prior = {}, {}
for c in classes:
    idx = [i for i in TRAIN if y[i] == c]
    prior[c] = len(idx) / len(TRAIN)
    stats[c] = [(mean([X[i][k] for i in idx]), stdev([X[i][k] for i in idx]))
                for k in range(2)]

print("%-6s %8s %12s %10s %12s %10s"
      % ("class", "prior", "att mean", "att sd", "prac mean", "prac sd"))
for c in classes:
    print("%-6s %8.4f %12.3f %10.3f %12.3f %10.3f"
          % (c, prior[c], stats[c][0][0], stats[c][0][1], stats[c][1][0], stats[c][1][1]))

def predict(x):
    best, bestp = None, -1
    for c in classes:
        p = prior[c]
        for k in range(2):
            m, s = stats[c][k]
            p *= gauss(x[k], m, s)
        if p > bestp:
            best, bestp = c, p
    return best

tr = sum(1 for i in TRAIN if predict(X[i]) == y[i]) / len(TRAIN)
te = sum(1 for i in TEST if predict(X[i]) == y[i]) / len(TEST)
print()
print("our Gaussian Naive Bayes : train %.3f  test %.3f" % (tr, te))

from sklearn.naive_bayes import GaussianNB
m = GaussianNB().fit([X[i] for i in TRAIN], [y[i] for i in TRAIN])
print("scikit-learn GaussianNB  : train %.3f  test %.3f"
      % (m.score([X[i] for i in TRAIN], [y[i] for i in TRAIN]),
         m.score([X[i] for i in TEST], [y[i] for i in TEST])))
class     prior     att mean     att sd    prac mean    prac sd
Fail     0.5714       57.333     18.500       19.583     10.672
Pass     0.4286       83.222     16.061       28.000     10.512

our Gaussian Naive Bayes : train 0.905  test 0.667
scikit-learn GaussianNB  : train 0.905  test 0.667

Four numbers per class per feature, and the model is trained. Read them: the students who passed averaged 83.2 per cent attendance and 28.0 hours of practice; those who failed averaged 57.3 and 19.6. That is the whole model, and it is readable, which is another thing Naive Bayes has over a neural network.

munotes.in80

Practical 7: The Naive Bayes Classifier

Our program and GaussianNB agree exactly, 0.905 training and 0.667 test, which is the check that the implementation is right.

Note that this model does not need its features scaled. Each feature gets its own mean and standard deviation, so the units cancel. That is unusual in this module and worth saying.

Procedure

  1. Count the class frequencies to get the priors, and check one by hand against the file.
  2. For each class and each feature value, count and divide to get the likelihood table. Print it.
  3. Write predict: multiply the prior by one likelihood per feature, then normalise the scores across the classes to get probabilities.
  4. Work one prediction with every factor printed, and check the multiplication with a calculator.
  5. Find a zero in the likelihood table and construct a row that hits it. Record the probability the model then reports.
  6. Add Laplace smoothing and record the same row's probability again, and the accuracy both ways.
  7. For the numeric dataset, compute the mean and standard deviation of each feature inside each class and use the normal density as the likelihood.
  8. Compare with GaussianNB on the same split.

Observations

Measured on weather.csvValue
P(Yes), P(No)0.6429, 0.3571
Likelihood of Overcast given Yes0.4444, which is 4/9
Likelihood of Overcast given No0.0000, which is 0/5
Sunny, Cool, High, Windy: score for No0.02057143
Sunny, Cool, High, Windy: score for Yes0.00529101
The same, as probabilitiesNo 0.7954, Yes 0.2046
Answerplay = No
Accuracy on the fourteen rows13 of 14
An Overcast, Hot, Humid, Windy dayP(Yes)P(No)accuracy on 14 rows
no smoothing1.00000.000013 of 14
Laplace, add 10.56440.435613 of 14
Gaussian Naive Bayes on students.csvpriorattendance mean, sdpractice mean, sd
Fail0.571457.333, 18.50019.583, 10.672
Pass0.428683.222, 16.06128.000, 10.512
traintest
our Gaussian Naive Bayes0.9050.667
scikit-learn GaussianNB0.9050.667

Result

The Naive Bayes classifier was implemented from nothing. Training reduced to counting: two priors and forty conditional probabilities, each checked against the file by hand. For a sunny, cool, humid and windy day the unnormalised scores were 0.02057143 for No and 0.00529101 for Yes, giving probabilities of 0.7954 and 0.2046 and the answer No, and the model scored 13 of 14 on the rows it was trained on. A zero likelihood was found and shown to force a reported probability of exactly 1.0000 against 0.0000; Laplace smoothing with alpha 1 replaced that with 0.5644 against 0.4356 at no cost in accuracy. The Gaussian form, using a mean and a standard deviation per class per feature, reached 0.905 training and 0.667 test accuracy on the student dataset, identical to scikit-learn's GaussianNB.

munotes.in81

Practical 7: The Naive Bayes Classifier

Where marks are lost

No smoothing. One unseen combination and the model reports a probability of zero with total confidence.

Calling the unnormalised score a probability. Divide by the total of the scores first, and say you did.

Multiplying many small numbers. With twenty features the product underflows to 0.0 and every class ties. Add the logarithms instead of multiplying the probabilities; the ranking is the same.

Counting the feature values from the class subset instead of the whole dataset when smoothing. The denominator's alpha * number of values must use every value the feature can take, not only those present in this class.

Using counting on a numeric feature. Attendance 83 has a count of 1 or 0 and the model learns nothing. Use the Gaussian form.

Scaling the features for Gaussian Naive Bayes. Harmless, but unnecessary, and if you say it is necessary you are wrong. Each feature has its own mean and standard deviation.

Not saying the independence assumption is false. An examiner will ask. Say it is false, and say the classifier works anyway because it usually ranks the classes correctly even when its numbers are poor.

For the journal

Aim; Bayes's theorem and the naive assumption written out; the priors and the likelihood table, with one entry checked by hand; the worked prediction with every factor, both the unnormalised scores and the normalised probabilities; the accuracy; the zero-probability demonstration with the two rows of the smoothing table; the Laplace formula; the Gaussian form with the four statistics per class; the comparison with GaussianNB; all observation tables; the result.

Quick revision

  • P(class | features) is proportional to P(class) P(f1 | class) P(f2 | class) * ...
  • Naive means the features are assumed independent given the class. The assumption is false and the classifier works anyway.
  • Training is counting. No iteration, no learning rate.
  • The prior is the class's share of the rows; the likelihood is a count within a class divided by that class's size.
  • Normalise the scores across the classes to report probabilities.
  • One zero likelihood destroys the whole product: it reported 1.0000 against 0.0000 here.
  • Laplace: add alpha to every count and alpha * (number of values) to every denominator.
  • For numeric features use the normal density with the class's own mean and standard deviation.
  • Gaussian Naive Bayes needs no scaling.
  • Add logarithms instead of multiplying when there are many features.

Questions you must be able to answer

1. State Bayes's theorem and say which part you compute and which you ignore. P(class | features) = P(features | class) * P(class) / P(features). You compute the numerator for each class; the denominator is the same for all of them, so it cancels when you compare, and dividing each score by the total of the scores recovers the probabilities.

munotes.in82

Practical 7: The Naive Bayes Classifier

2. What is naive about Naive Bayes? It assumes the features are independent of each other given the class, so that their probabilities can be multiplied one at a time. Outlook and humidity clearly are not independent.

3. How is the model trained? By counting. Count each class for the priors, and each feature value within each class for the likelihoods. There is nothing to iterate.

4. Your program printed 0.0206 and 0.0053. Are those the probabilities? No, they are the unnormalised scores. Dividing each by their total, 0.0259, gives 0.7954 and 0.2046, which are the probabilities.

5. What is the zero-frequency problem? If a feature value never occurs with a class in training, its likelihood is 0, and since the scores are products, that class's score becomes 0 whatever the other features say. Here it produced a reported certainty of 1.0000 against 0.0000.

6. How does Laplace smoothing fix it, and what did it change? By adding alpha to every count and alpha * (number of values) to every denominator, so no likelihood is ever exactly zero. The 0.0000 became 0.1250 and the day's answer went from a false certainty to 0.5644 against 0.4356.

7. How do you handle a feature like attendance, which is a number? Assume the values in each class follow a normal distribution, store the mean and standard deviation for each class, and use the height of that bell curve as the likelihood. That is Gaussian Naive Bayes.

8. Why can a Naive Bayes program underflow, and what is the fix? Because many probabilities below 1 multiplied together become smaller than the smallest number a float can hold. Add the logarithms instead; the class with the largest log score is the same class.

9. Your Naive Bayes gets 13 of 14 and your decision tree got 14 of 14. Which is the better model? Neither figure answers that, because both were measured on the training rows. Practical 3 showed that the tree's perfect score was memory: on held-out data it fell. Compare them on a test set.

10. Does Naive Bayes need the features scaled? No. The counting form has no notion of scale, and the Gaussian form gives each feature its own mean and standard deviation, so the units cancel.

Contents This chapter on its own page

munotes.in83

Chapter Eleven

Practical 8: K-Nearest Neighbours for Classification and for Regression

Syllabus topic Module 1, "K-Nearest Neighbors (K-NN): Implement the K-NN algorithm for classification or regression. Apply the K-NN algorithm to a given dataset and predict the class or value for test data. Evaluate the accuracy or error of the predictions and analyze the results."

Aim

To implement K-Nearest Neighbours, to use it to predict a class and to predict a number, and to measure the accuracy of the one and the error of the other.

What you need to know before you start

Every other model in this module builds something during training: a tree, a set of weights, a table of counts. K-NN builds nothing. It keeps the training data, and when a question arrives it looks up the k rows most like it and lets them answer.

That is why it is called a lazy learner, and why its cost is the opposite way round from everything else: training is instant and every prediction is expensive, because every prediction scans the whole training set.

Three decisions make a K-NN, and all three belong in the journal.

How to measure "most like it". Euclidean distance, the straight-line distance, is the usual answer:

distance(a, b) = sqrt((a1 - b1)^2 + (a2 - b2)^2 + ... + (an - bn)^2)

How many neighbours, k. Small k follows the data closely and follows its noise with it; large k smooths, and with k equal to the number of training rows it answers the same thing every time. k is chosen by measuring.

How the neighbours answer. For a class, they vote. For a number, they are averaged. That single difference is the whole gap between K-NN classification and K-NN regression.

And one thing that is not optional: scale the features. Distance adds up the differences in every column, so a column measured in the tens dominates one measured in the units for no reason except its units. Attendance runs from 35 to 98 and practice from 2 to 45; without scaling, attendance decides almost everything.

Part 1: classification

name,attendance,practice,result
Aarav,50,22,Fail
Isha,48,16,Fail
Rohan,88,38,Pass
Sanya,97,39,Pass
Vikram,90,30,Fail
Meera,35,41,Fail
Farhan,45,9,Fail
Nikita,71,8,Fail
Omkar,92,2,Fail
Pooja,97,45,Pass
Rahul,75,15,Fail
Sneha,85,18,Pass
Tejas,79,24,Pass
Urmila,83,34,Pass
Varun,44,23,Fail
Yash,46,37,Pass
Zoya,72,20,Fail
Amit,44,9,Pass
Bhavna,74,3,Fail
Chirag,82,25,Pass
Deepa,94,29,Pass
Eshan,46,27,Fail
Gauri,98,9,Fail
Harsh,89,34,Pass
Ira,97,27,Pass
Jatin,68,28,Pass
Kavya,96,34,Pass
Lalit,38,38,Fail
Manav,63,10,Fail
Neha,41,35,Fail
"""Practical 8: K-Nearest Neighbours, for classification and for regression."""
import csv
import math
import random

with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
raw = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [r["result"] for r in rows]
random.seed(1)
order = list(range(len(raw))); random.shuffle(order)
cut = int(0.7 * len(raw))
TRAIN, TEST = order[:cut], order[cut:]

lo = [min(r[k] for r in raw) for k in range(2)]
hi = [max(r[k] for r in raw) for k in range(2)]
scaled = [[(r[k] - lo[k]) / (hi[k] - lo[k]) for k in range(2)] for r in raw]

def distance(a, b):
    return math.sqrt(sum((a[k] - b[k]) ** 2 for k in range(len(a))))

def knn_classify(X, i, k):
    """The k training rows nearest to row i, and the commonest label among them."""
    near = sorted(TRAIN, key=lambda j: (distance(X[i], X[j]), j))[:k]
    votes = {}
    for j in near:
        votes[y[j]] = votes.get(y[j], 0) + 1
    top = max(votes.values())
    tied = sorted(c for c, v in votes.items() if v == top)
    return tied[0], near

def accuracy(X, idx, k):
    return sum(1 for i in idx if knn_classify(X, i, k)[0] == y[i]) / len(idx)

print("one prediction, worked out: the first test row, k = 3, scaled features")
i = TEST[0]
label, near = knn_classify(scaled, i, 3)
print("  the row to classify: %s, attendance %s, practice %s, really %s"
      % (rows[i]["name"], rows[i]["attendance"], rows[i]["practice"], y[i]))
print("  %-10s %12s %10s %8s %10s" % ("neighbour", "attendance", "practice", "result", "distance"))
for j in near:
    print("  %-10s %12s %10s %8s %10.4f"
          % (rows[j]["name"], rows[j]["attendance"], rows[j]["practice"], y[j],
             distance(scaled[i], scaled[j])))
print("  vote -> %s" % label)
print()
print("%-5s %14s %13s %14s %13s"
      % ("k", "scaled train", "scaled test", "raw train", "raw test"))
for k in (1, 3, 5, 7, 9, 11):
    print("%-5d %14.3f %13.3f %14.3f %13.3f"
          % (k, accuracy(scaled, TRAIN, k), accuracy(scaled, TEST, k),
             accuracy(raw, TRAIN, k), accuracy(raw, TEST, k)))
munotes.in84

Practical 8: K-Nearest Neighbours for Classification and for Regression

one prediction, worked out: the first test row, k = 3, scaled features
  the row to classify: Yash, attendance 46, practice 37, really Pass
  neighbour    attendance   practice   result   distance
  Neha                 41         35     Fail     0.0920
  Meera                35         41     Fail     0.1978
  Eshan                46         27     Fail     0.2326
  vote -> Fail

k       scaled train   scaled test      raw train      raw test
1              1.000         0.667          1.000         0.667
3              0.952         0.778          0.905         0.778
5              0.905         0.778          0.905         0.778
7              0.905         0.667          0.905         0.667
9              0.952         0.778          0.810         0.667
11             0.905         0.667          0.810         0.556

The worked prediction is what to put in the journal. Yash, with 46 per cent attendance and 37 hours of practice, really passed. His three nearest neighbours are Neha, Meera and Eshan, at scaled distances 0.0920, 0.1978 and 0.2326, and all three failed. Three votes to nothing: the model says Fail, and it is wrong. Showing a prediction the model gets wrong, with the neighbours that caused it, is worth more marks than showing one it gets right, because it demonstrates that you understand the mechanism rather than the result.

The table of k is the analysis. Read the two scaled columns:

  • k = 1 scores 1.000 on the training rows. Of course it does: the nearest neighbour of a training row is itself, at distance 0. That figure is not a measurement of anything, and a student who reports it has reported that the program can read its own input.
  • k = 1 scores 0.667 on the test rows, the worst of any k tried. One noisy neighbour decides the answer.
  • k = 3 and k = 5 give 0.778, the best here.
  • Large k drifts back down, to 0.667 at k = 11, because the vote starts to include students who are nothing like the one being classified.
munotes.in85

Practical 8: K-Nearest Neighbours for Classification and for Regression

That shape, bad at k = 1, best in the middle, worse again as k grows, is the one to describe.

Scaling. Compare the scaled and raw columns: they agree at small k and diverge as k grows, 0.778 against 0.667 at k = 9 and 0.667 against 0.556 at k = 11. Scaling helped, and it would have helped much more with a column measured in thousands.

Ties. With an even k the vote can tie, and with two classes it does so often. This program breaks a tie by taking the alphabetically first label, which is a decision, not a law: a fair alternative is to weight each neighbour by 1 divided by its distance. Say which you used. The simplest way to avoid the question is to use an odd k.

Part 2: regression, where the answer is a number

The same neighbours, averaged instead of counted. The dataset is forty students with the mark each one actually scored.

attendance,practice,marks
92,35,77
92,32,65
59,11,38
58,6,27
92,19,62
40,38,56
85,28,60
55,39,55
36,33,47
65,38,64
38,29,41
60,33,61
64,40,65
45,29,47
70,26,68
45,45,61
67,20,48
38,4,16
48,25,47
43,1,22
35,13,30
95,24,73
85,26,67
69,21,52
46,19,33
50,8,39
66,45,71
97,11,54
59,28,40
51,26,46
84,7,35
35,17,39
73,1,35
47,2,27
53,13,35
77,18,51
84,4,44
66,0,31
82,23,56
96,36,74
"""K-NN predicting a NUMBER: the same neighbours, averaged instead of voted."""
import csv, math, random
with open("marks.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
raw = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [float(r["marks"]) for r in rows]
random.seed(1)
order = list(range(len(raw))); random.shuffle(order)
cut = int(0.7 * len(raw))
TRAIN, TEST = order[:cut], order[cut:]
lo = [min(r[k] for r in raw) for k in range(2)]
hi = [max(r[k] for r in raw) for k in range(2)]
X = [[(r[k] - lo[k]) / (hi[k] - lo[k]) for k in range(2)] for r in raw]

def dist(a, b): return math.sqrt(sum((a[k]-b[k])**2 for k in range(len(a))))

def predict(i, k):
    near = sorted(TRAIN, key=lambda j: (dist(X[i], X[j]), j))[:k]
    return sum(y[j] for j in near) / k

def errors(idx, k):
    e = [y[i] - predict(i, k) for i in idx]
    mae = sum(abs(v) for v in e) / len(e)
    rmse = math.sqrt(sum(v*v for v in e) / len(e))
    return mae, rmse

print("training rows %d, test rows %d, marks from %.0f to %.0f"
      % (len(TRAIN), len(TEST), min(y), max(y)))
print()
print("%-5s %12s %12s %12s %12s" % ("k", "train MAE", "train RMSE", "test MAE", "test RMSE"))
for k in (1, 2, 3, 5, 7, 9):
    a, b = errors(TRAIN, k); c, d = errors(TEST, k)
    print("%-5d %12.3f %12.3f %12.3f %12.3f" % (k, a, b, c, d))
print()
mean_train = sum(y[i] for i in TRAIN) / len(TRAIN)
e = [y[i] - mean_train for i in TEST]
print("the baseline, always answering the training mean of %.2f:" % mean_train)
print("  test MAE %.3f  test RMSE %.3f"
      % (sum(abs(v) for v in e)/len(e), math.sqrt(sum(v*v for v in e)/len(e))))
print()
i = TEST[0]
k = 3
near = sorted(TRAIN, key=lambda j: (dist(X[i], X[j]), j))[:k]
print("one prediction worked out, k = 3:")
print("  the row: attendance %s, practice %s, real mark %.0f"
      % (rows[i]["attendance"], rows[i]["practice"], y[i]))
print("  %-12s %10s %8s %10s" % ("attendance", "practice", "marks", "distance"))
for j in near:
    print("  %-12s %10s %8.0f %10.4f"
          % (rows[j]["attendance"], rows[j]["practice"], y[j], dist(X[i], X[j])))
print("  average of the three -> %.2f, against the real %.0f" % (predict(i, k), y[i]))

from sklearn.neighbors import KNeighborsRegressor
from sklearn.preprocessing import MinMaxScaler
sc = MinMaxScaler().fit([raw[i] for i in TRAIN])
m = KNeighborsRegressor(n_neighbors=3).fit(sc.transform([raw[i] for i in TRAIN]),
                                           [y[i] for i in TRAIN])
pred = m.predict(sc.transform([raw[i] for i in TEST]))
e = [y[i] - p for i, p in zip(TEST, pred)]
print()
print("scikit-learn KNeighborsRegressor, k = 3: test MAE %.3f  test RMSE %.3f"
      % (sum(abs(v) for v in e)/len(e), math.sqrt(sum(v*v for v in e)/len(e))))
munotes.in86

Practical 8: K-Nearest Neighbours for Classification and for Regression

training rows 28, test rows 12, marks from 16 to 77

k        train MAE   train RMSE     test MAE    test RMSE
1            0.000        0.000        8.833       10.239
2            3.286        4.162        6.500        7.757
3            3.107        3.815        6.806        8.378
5            3.907        4.872        6.583        7.840
7            4.189        5.075        5.810        7.285
9            4.881        6.069        6.306        7.606

the baseline, always answering the training mean of 51.07:
  test MAE 10.952  test RMSE 12.480

one prediction worked out, k = 3:
  the row: attendance 58, practice 6, real mark 27
  attendance     practice    marks   distance
  59                   11       38     0.1123
  53                   13       35     0.1752
  66                    0       31     0.1855
  average of the three -> 34.67, against the real 27

scikit-learn KNeighborsRegressor, k = 3: test MAE 6.806  test RMSE 8.378

Look at the k = 1 row: training MAE and RMSE are both exactly 0.000. The nearest neighbour of a training row is itself, so the prediction is the true answer, so the error is zero. It is the same fact as the 1.000 above, and in this form it is impossible to misread: a model with zero error that is wrong by 8.8 marks on average on data it has not seen.

The best k here is 7, with a test MAE of 5.810 marks. At k = 1 it is 8.833 and at k = 9 it is back up to 6.306.

The baseline settles whether any of this is worth anything. Answering the training mean of 51.07 to every question gives a test MAE of 10.952. K-NN at k = 7 gives 5.810, which is a little under half the error. That comparison is the result; the MAE on its own is a number.

munotes.in87

Practical 8: K-Nearest Neighbours for Classification and for Regression

MAE and RMSE disagree about k. By MAE, k = 7 is best at 5.810. By RMSE, k = 7 is also best at 7.285, but k = 2 is second on both. RMSE is always the larger of the two, because squaring punishes the big misses, and the gap between them, about 1.5 marks here, says there are a few predictions much worse than the average.

The worked prediction again shows a miss. A student with 58 per cent attendance and 6 hours of practice really scored 27. The three nearest rows scored 38, 35 and 31, so the model answers 34.67, out by 7.67. All three neighbours did more practice than the student being predicted, and nothing in the averaging notices that. That is the honest reading.

And scikit-learn agrees exactly: KNeighborsRegressor with k = 3 gives test MAE 6.806 and RMSE 8.378, the same to three decimal places as our own.

Procedure

  1. Load the dataset, scale each feature to the range 0 to 1, and split 70 to 30 with a seeded shuffle.
  2. Write distance as the Euclidean distance.
  3. Write the classifier: sort the training rows by distance, take the first k, and return the commonest label. Break ties in a stated way.
  4. Print one prediction in full, with its k neighbours and their distances, and choose one the model gets wrong if there is one.
  5. Measure accuracy over a range of k, on both the training and the test rows, scaled and unscaled, and say where the best k is.
  6. For regression, replace the vote with the mean of the neighbours' values.
  7. Compute MAE and RMSE for each k on both sets.
  8. Compute the baseline: the error of always answering the training mean. Compare.
  9. Confirm your figures against KNeighborsClassifier or KNeighborsRegressor.

Observations

Classification, 21 training rows and 9 test rows.

kscaled trainscaled testraw trainraw test
k = 11.0000.6671.0000.667
k = 30.9520.7780.9050.778
k = 50.9050.7780.9050.778
k = 70.9050.6670.9050.667
k = 90.9520.7780.8100.667
k = 110.9050.6670.8100.556

Regression, 28 training rows and 12 test rows.

ktrain MAEtrain RMSEtest MAEtest RMSE
k = 10.0000.0008.83310.239
k = 23.2864.1626.5007.757
k = 33.1073.8156.8068.378
k = 53.9074.8726.5837.840
k = 74.1895.0755.8107.285
k = 94.8816.0696.3067.606
Also observedValue
Baseline: always answer the training mean of 51.07test MAE 10.952, test RMSE 12.480
Best k for classification3 or 5, at 0.778
Best k for regression7, at MAE 5.810
scikit-learn KNeighborsRegressor, k = 3test MAE 6.806, RMSE 8.378, identical to ours
munotes.in88

Practical 8: K-Nearest Neighbours for Classification and for Regression

Result

K-Nearest Neighbours was implemented for both tasks MU names. As a classifier on the thirty-student dataset it reached 0.778 on the nine test rows at k = 3 and k = 5, against 0.667 at k = 1 and at k = 11; scaling the features made no difference at small k and was worth up to 0.111 of accuracy at large k. As a regressor on forty students' marks it reached a test mean absolute error of 5.810 marks at k = 7, against a baseline of 10.952 for always answering the training mean, and its figures at k = 3 match scikit-learn's KNeighborsRegressor exactly. At k = 1 the training error was exactly zero in both tasks while the test error was the worst recorded, which is the clearest demonstration in this module of why a score on training data is worth nothing.

Where marks are lost

Reporting k = 1 training accuracy. It is 1.000 and it always will be, because every row is its own nearest neighbour.

Not scaling. Distance is a sum over columns, so the column with the biggest numbers wins.

Trying one k. MU says evaluate and analyse. The table of k is the analysis.

An even k with two classes. Ties. Use an odd k, or state the tie-break rule.

Forgetting the baseline in regression. MAE of 5.810 means nothing until the reader knows that answering the mean gives 10.952.

Quoting MAE and calling it accuracy. Accuracy is for classes. For numbers it is MAE or RMSE, and say which.

Excluding the row itself from its own neighbours when testing, but not when training, or the reverse. Be consistent and say what you did. This program searches the training rows only, so a test row is never its own neighbour and a training row always is.

Calling K-NN a "trained" model. Nothing is trained. It stores the data and does all its work at prediction time.

For the journal

Aim; the distance formula; the three decisions, k, the distance and how the neighbours answer; the scaling formula and why it is needed; the classification program; one prediction worked out in full with its neighbours and distances, preferably one the model gets wrong; the table of k, scaled and unscaled, on both sets; the regression program; the MAE and RMSE table; the baseline; the scikit-learn comparison; both observation tables; the result; one paragraph on which k you would use and why.

Quick revision

  • K-NN stores the training data and does all its work when asked a question. It is a lazy learner.
  • Euclidean distance: the square root of the sum of the squared differences.
  • Classification: the k nearest rows vote. Regression: they are averaged.
  • Small k follows noise; large k smooths and eventually answers the same thing every time.
  • k = 1 always scores perfectly on the training rows. That figure is worthless.
  • Scale the features, or the column with the largest numbers decides the distance.
  • Use an odd k with two classes, or state how ties are broken.
  • Always compare a regression with the baseline of always answering the mean.
  • Measured here: 0.778 at k = 3 and 5 for classification; MAE 5.810 at k = 7 for regression against a baseline of 10.952.
munotes.in89

Practical 8: K-Nearest Neighbours for Classification and for Regression

Questions you must be able to answer

1. What does K-NN do during training? Nothing. It stores the training rows. All the work happens at prediction time, which is why it is called a lazy learner.

2. How does a K-NN classifier differ from a K-NN regressor? Only in what the neighbours do. For a class they vote and the commonest wins; for a number they are averaged.

3. Why is the k = 1 training accuracy always 1.000? Because each training row's nearest neighbour is itself, at distance zero, so it predicts its own label. In the regression run this showed as a training error of exactly 0.000.

4. What happens as k grows? The prediction smooths out: the vote or the average includes rows less and less like the one being predicted. At k equal to the whole training set it answers the majority class, or the overall mean, every time.

5. Why must the features be scaled? Because distance sums the differences across columns, so a column whose values are large contributes more to every distance regardless of whether it matters. Here attendance ranges over 63 units and practice over 43.

6. Your k = 3 prediction for Yash was Fail, and he passed. Explain it. His three nearest scaled neighbours, at distances 0.0920, 0.1978 and 0.2326, all failed. K-NN can only answer with what is near, and near this student there were no passes.

7. What does the baseline tell you in the regression run? That the model is worth having. Answering the training mean of 51.07 gives a test MAE of 10.952; K-NN at k = 7 gives 5.810, a little under half.

8. Why is RMSE larger than MAE in your table? Because RMSE squares each error before averaging, so a few large misses count for much more. The size of the gap says how uneven the errors are.

9. How would you break a tie in a two-class vote? Use an odd k so it cannot happen; or weight each neighbour by one over its distance so the nearer one wins; or fix a rule, as this program does, and state it.

munotes.in90

Practical 8: K-Nearest Neighbours for Classification and for Regression

10. K-NN takes no time to train. What is the cost? Every prediction scans the whole training set to find the k nearest, so prediction is O(n) in the number of stored rows for each question asked.

Contents This chapter on its own page

munotes.in91

Chapter Twelve

Practical 9: Association Rule Mining with Apriori

Syllabus topic Module 1, "Association Rule Mining: Implement the Association Rule Mining algorithm (e.g., Apriori) to find frequent itemsets. Generate association rules from the frequent itemsets and calculate their support and confidence. Interpret and analyze the discovered association rules."

Aim

To implement Apriori to find the frequent itemsets in a set of transactions, to generate association rules from them, to compute the support and confidence of each rule, and to interpret what the rules say.

What you need to know before you start

Every practical so far has had a label to predict. This one has none. It looks at a pile of shopping baskets and asks what tends to be bought with what. That makes it unsupervised: there is no right answer to learn, only structure to find.

Three numbers describe a rule, and getting them straight is most of this practical.

Support is how common something is: the fraction of all the baskets that contain it.

support(X) = (number of baskets containing every item in X) / (total baskets)

Confidence is how reliable a rule is: of the baskets that contain the left side, what fraction also contain the right side.

confidence(X -> Y) = support(X and Y) / support(X)

Lift is the one that stops you being fooled, and it is the one students leave out.

lift(X -> Y) = confidence(X -> Y) / support(Y)

Lift compares the rule against simply guessing Y. Above 1, X makes Y more likely than usual. Exactly 1, X tells you nothing at all. Below 1, X makes Y less likely, and a rule like that can still have a high confidence, which is why confidence alone is not enough.

Why the algorithm is needed at all

With 5 different items there are 31 possible non-empty itemsets. With 20 items there are 1,048,575, and with 50 there are more than a million million million. Counting them all is not an option, and that is the problem Apriori solves.

Its idea is one sentence, and it is worth memorising exactly:

If an itemset is frequent, then every subset of it is frequent as well. Turn it round and you get the useful form: if any subset of a candidate is not frequent, the candidate cannot be frequent either, so it never has to be counted.

So the algorithm works in passes. Pass 1 counts the single items and keeps the frequent ones. Pass 2 builds pairs out of frequent singles and counts those. Pass 3 builds triples out of frequent pairs, and only those triples all of whose pairs survived. Each pass touches a tiny fraction of what is possible.

The dataset

Twelve baskets from a college stationery counter. Each line is one purchase.

notebook,pen,highlighter
notebook,pen
notebook,pen,highlighter,calculator
pen,highlighter
notebook,calculator
notebook,pen,highlighter
pen,calculator
notebook,pen,calculator
notebook,highlighter
notebook,pen,highlighter,geometry-box
pen,highlighter,calculator
notebook,pen,highlighter
"""Practical 9: Apriori, and the rules that come out of it."""
from itertools import combinations

with open("baskets.txt") as fh:
    baskets = [frozenset(line.strip().split(",")) for line in fh if line.strip()]
N = len(baskets)
MIN_SUPPORT = 0.25
MIN_CONFIDENCE = 0.70

def support(itemset):
    return sum(1 for b in baskets if itemset <= b) / N

def show(itemset):
    return "{" + ", ".join(sorted(itemset)) + "}"

print("%d baskets" % N)
for i, b in enumerate(baskets, 1):
    print("  %2d  %s" % (i, ", ".join(sorted(b))))
print()
print("minimum support %.2f, which is %d baskets. minimum confidence %.2f"
      % (MIN_SUPPORT, round(MIN_SUPPORT * N), MIN_CONFIDENCE))
print()

items = sorted(set().union(*baskets))
frequent = {}
level = [frozenset([i]) for i in items]
k = 1
while level:
    print("pass %d: %d candidate itemset(s)" % (k, len(level)))
    kept = []
    for c in sorted(level, key=lambda s: sorted(s)):
        s = support(c)
        mark = "keep" if s >= MIN_SUPPORT else "drop"
        print("   %-42s support %.3f  %s" % (show(c), s, mark))
        if s >= MIN_SUPPORT:
            kept.append(c)
            frequent[c] = s
    # THE APRIORI RULE: join two frequent k-sets that share k-1 items, and keep
    # the candidate only if EVERY one of its subsets was itself frequent.
    nxt = set()
    for a, b in combinations(kept, 2):
        cand = a | b
        if len(cand) == k + 1 and all(frozenset(sub) in frequent
                                      for sub in combinations(cand, k)):
            nxt.add(cand)
    level = sorted(nxt, key=lambda s: sorted(s))
    k += 1
    print()

print("frequent itemsets found: %d" % len(frequent))
print()
print("%-46s %9s %11s %7s" % ("rule", "support", "confidence", "lift"))
rules = []
for itemset, sup in frequent.items():
    if len(itemset) < 2:
        continue
    for r in range(1, len(itemset)):
        for left in combinations(sorted(itemset), r):
            left = frozenset(left)
            right = itemset - left
            conf = sup / frequent[left]
            lift = conf / frequent[right]
            rules.append((conf, lift, sup, left, right))
rules.sort(key=lambda t: (-t[0], -t[1], sorted(t[3])))
for conf, lift, sup, left, right in rules:
    if conf < MIN_CONFIDENCE:
        continue
    print("%-46s %9.3f %11.3f %7.3f"
          % ("%s -> %s" % (show(left), show(right)), sup, conf, lift))
munotes.in92

Practical 9: Association Rule Mining with Apriori

12 baskets
   1  highlighter, notebook, pen
   2  notebook, pen
   3  calculator, highlighter, notebook, pen
   4  highlighter, pen
   5  calculator, notebook
   6  highlighter, notebook, pen
   7  calculator, pen
   8  calculator, notebook, pen
   9  highlighter, notebook
  10  geometry-box, highlighter, notebook, pen
  11  calculator, highlighter, pen
  12  highlighter, notebook, pen

minimum support 0.25, which is 3 baskets. minimum confidence 0.70

pass 1: 5 candidate itemset(s)
   {calculator}                               support 0.417  keep
   {geometry-box}                             support 0.083  drop
   {highlighter}                              support 0.667  keep
   {notebook}                                 support 0.750  keep
   {pen}                                      support 0.833  keep

pass 2: 6 candidate itemset(s)
   {calculator, highlighter}                  support 0.167  drop
   {calculator, notebook}                     support 0.250  keep
   {calculator, pen}                          support 0.333  keep
   {highlighter, notebook}                    support 0.500  keep
   {highlighter, pen}                         support 0.583  keep
   {notebook, pen}                            support 0.583  keep

pass 3: 2 candidate itemset(s)
   {calculator, notebook, pen}                support 0.167  drop
   {highlighter, notebook, pen}               support 0.417  keep

frequent itemsets found: 10

rule                                             support  confidence    lift
{highlighter} -> {pen}                             0.583       0.875   1.050
{highlighter, notebook} -> {pen}                   0.417       0.833   1.000
{calculator} -> {pen}                              0.333       0.800   0.960
{notebook} -> {pen}                                0.583       0.778   0.933
{highlighter} -> {notebook}                        0.500       0.750   1.000
{notebook, pen} -> {highlighter}                   0.417       0.714   1.071
{highlighter, pen} -> {notebook}                   0.417       0.714   0.952
{pen} -> {highlighter}                             0.583       0.700   1.050
{pen} -> {notebook}                                0.583       0.700   0.933
munotes.in93

Practical 9: Association Rule Mining with Apriori

Reading the passes

Pass 1 counts the five items. geometry-box appears in one basket out of twelve, support 0.083, and is dropped. Everything else survives.

Pass 2 builds six pairs from the four surviving singles, which is 4 choose 2, and drops {calculator, highlighter} at 0.167. Note what is not there: no pair containing geometry-box was ever built, because the Apriori rule says a pair cannot be frequent if one of its items is not.

Pass 3 is where the saving shows. Five frequent pairs can be joined into several triples, but only two candidates appear, {calculator, notebook, pen} and {highlighter, notebook, pen}, because those are the only triples all three of whose pairs survived pass 2. A triple containing {calculator, highlighter} would have been built by a naive program and counted for nothing. One of the two is then dropped on its own support and the other, at 0.417, is kept.

Pass 4 finds no candidates at all and the loop stops. Ten frequent itemsets in total: four singles, five pairs and one triple.

Interpreting the rules, which is the third thing MU asks for

Every frequent itemset of two or more items yields rules, one for each way of cutting it into a left and a right side. The nine that reach a confidence of 0.70 are in the table. Read three of them.

{highlighter} -> {pen}: support 0.583, confidence 0.875, lift 1.050. Seven of the eight baskets with a highlighter also had a pen. The lift is just above 1, so a highlighter does make a pen a little more likely than average, but only a little: pens are in 10 of the 12 baskets anyway.

{notebook, pen} -> {highlighter}: support 0.417, confidence 0.714, lift 1.071. The highest lift in the table. Somebody buying both a notebook and a pen is the most likely of anybody to add a highlighter, and this is the rule a shopkeeper would act on.

{calculator} -> {pen}: support 0.333, confidence 0.800, lift 0.960. Here is the trap, and it is the reason this section exists. A confidence of 0.800 looks strong and the rule is worthless. Four of the five calculator baskets also had a pen, which sounds compelling until you notice that pens are in 0.833 of all baskets. So a calculator buyer is slightly less likely to buy a pen than a random customer is, and the lift of 0.960 says exactly that. Acting on this rule, putting the pens next to the calculators, would be acting on nothing.

munotes.in94

Practical 9: Association Rule Mining with Apriori

{highlighter, notebook} -> {pen}: confidence 0.833, lift 1.000. Exactly 1. Perfect independence: knowing about the highlighter and notebook changes the chance of a pen not at all.

So the interpretation, in the form to write in the journal: rank by lift, filter by support and confidence. Support says the rule is worth caring about because it happens often enough; confidence says it holds up when it happens; lift says it is telling you something you did not already know.

Two honest warnings about small data

Twelve baskets is enough to demonstrate the algorithm and far too few to believe the rules. A support of 0.25 is three baskets. One more or one fewer moves it by 0.083, and several of the rules above would change places. A real market-basket study runs to tens of thousands of transactions, and a write-up here should say plainly how many baskets the figures came from.

And geometry-box was dropped at pass 1, so no rule involving it can ever be found however strong it might be. That is what a minimum support threshold does: it buys speed by refusing to look at anything rare. A rare but valuable pattern is invisible to Apriori, and the threshold is a choice, not a fact.

Procedure

  1. Write the transactions, one per line, and read them into sets.
  2. Choose a minimum support and a minimum confidence, and say what the support means in baskets.
  3. Count the support of every single item; keep the frequent ones and print which were dropped.
  4. Build pairs from the frequent singles; count and keep. Print the candidates as well as the survivors, so the reader can see what was counted.
  5. Build triples only from candidates every subset of which was frequent, and say how many candidates that rule saved.
  6. Stop when a pass produces no candidates.
  7. For every frequent itemset of two items or more, generate a rule for each way of splitting it, and compute support, confidence and lift.
  8. Sort by confidence, apply the minimum, and interpret at least three rules in English, including one whose lift is below 1.

Observations

PassCandidatesKeptDropped
154geometry-box at support 0.083
265calculator with highlighter at 0.167
321calculator with notebook and pen at 0.167
400the loop stops
Frequent itemsetSupport
pen0.833
notebook0.750
highlighter0.667
calculator0.417
highlighter with pen0.583
notebook with pen0.583
highlighter with notebook0.500
calculator with pen0.333
calculator with notebook0.250
highlighter with notebook and pen0.417
RuleSupportConfidenceLift
highlighter, so pen0.5830.8751.050
highlighter and notebook, so pen0.4170.8331.000
calculator, so pen0.3330.8000.960
notebook, so pen0.5830.7780.933
highlighter, so notebook0.5000.7501.000
notebook and pen, so highlighter0.4170.7141.071
highlighter and pen, so notebook0.4170.7140.952
pen, so highlighter0.5830.7001.050
pen, so notebook0.5830.7000.933
munotes.in95

Practical 9: Association Rule Mining with Apriori

Result

Apriori was implemented and run over twelve transactions with a minimum support of 0.25 and a minimum confidence of 0.70. It found ten frequent itemsets in four passes: four single items, five pairs and one triple. The Apriori property reduced pass 3 to two candidates, because those were the only triples every pair of which had survived pass 2. Nine rules met the confidence threshold. The highest lift, 1.071, belongs to notebook and pen implying highlighter; the rule calculator implying pen has a confidence of 0.800 and a lift of 0.960, meaning a calculator buyer is slightly less likely than an average customer to buy a pen, and two rules have a lift of exactly 1.000, meaning their two sides are independent.

Where marks are lost

Reporting confidence without lift. The 0.800-confidence rule above is worthless and only lift says so.

Building candidates without the Apriori check. The program is then brute force wearing Apriori's name. Show the candidate counts so the saving is visible.

Mixing up support and confidence. Support is out of all the baskets. Confidence is out of the baskets containing the left side only.

Generating rules from the infrequent itemsets. Rules come from the frequent ones only, which is why the frequent itemsets are found first.

Forgetting the one-item rules. An itemset of three items yields six rules, not two: three with a single item on the left and three with a pair.

Using a list where a set is needed. Membership and subset tests on sets are what make this readable and fast.

Not saying how many transactions there were. Twelve baskets is a demonstration, not evidence.

Setting the minimum support after looking at the answers. Choose it first and say why.

For the journal

Aim; support, confidence and lift each defined with a formula; the Apriori property in one sentence, both ways round; the transactions; the pass-by-pass output showing every candidate and whether it was kept; the count of frequent itemsets; the rule table with all three measures; at least three rules interpreted in English, one of them with lift below 1; the two observation tables; the result; a sentence on what a minimum support threshold makes invisible.

Quick revision

  • Association rule mining is unsupervised: no labels, only patterns.
  • Support of X: the fraction of all baskets containing X.
  • Confidence of X implying Y: support of X and Y together, divided by support of X.
  • Lift: confidence divided by the support of Y. Above 1 informative, 1 independent, below 1 discouraging.
  • The Apriori property: every subset of a frequent itemset is frequent.
  • So a candidate with any infrequent subset is never counted. That is the whole saving.
  • Pass k builds its candidates by joining frequent itemsets of size k minus 1.
  • Stop when a pass yields no candidates.
  • An itemset of size n yields 2^n - 2 rules.
  • A high confidence with a lift below 1 is a rule worth ignoring.
  • A minimum support threshold makes rare patterns invisible, by design.
munotes.in96

Practical 9: Association Rule Mining with Apriori

Questions you must be able to answer

1. What is support, and out of what? The fraction of all the transactions that contain the itemset. Out of every basket, not just the relevant ones.

2. What is confidence, and how is it different? Of the baskets that contain the left side of the rule, the fraction that also contain the right side. Its denominator is the left side's support, not the total.

3. What is lift, and why does it matter? Confidence divided by the support of the right side. It says whether the left side actually makes the right side more likely. In this run one rule had confidence 0.800 and lift 0.960, so the rule is worse than guessing.

4. State the Apriori property. Every subset of a frequent itemset is itself frequent. Equivalently, if any subset of a candidate is infrequent, the candidate is infrequent and need not be counted.

5. Where did that property save work in your run? In pass 3: only two triples were built out of the five frequent pairs, because those were the only ones all of whose pairs had survived pass 2.

6. How many rules come out of a frequent itemset of three items? Six: three with one item on the left and two on the right, and three the other way round. In general 2 to the power n, minus 2.

7. Your algorithm found no rule involving geometry-box. Why? It appeared in one basket of twelve, support 0.083, below the minimum of 0.25, so it was dropped in pass 1 and no itemset containing it was ever built.

8. Two of your rules have a lift of exactly 1.000. What does that mean? That the two sides are independent: knowing the left side changes the chance of the right side not at all. The rule is true and useless.

9. Would you act on the rule from calculator to pen? No. Its confidence of 0.800 only reflects that pens are in 0.833 of all baskets anyway; its lift of 0.960 says a calculator buyer is slightly less likely than average to buy a pen.

munotes.in97

Practical 9: Association Rule Mining with Apriori

10. What would you change before trusting any of these rules? The number of transactions. Twelve baskets means one basket is 0.083 of support, so most of these figures are within a basket or two of each other.

Contents This chapter on its own page

munotes.in98

Chapter Thirteen

Practical 10: A Demonstration with TensorFlow and with an OpenAI-Style Tool

Syllabus topic Module 1, "Demo of OpenAI/TensorFlow Tools: Explore and experiment with OpenAI or TensorFlow tools and libraries. Perform a demonstration or mini-project showcasing the capabilities of the tools. Discuss and present the findings and potential applications."

Aim

To use TensorFlow to build and train a network, to see what an OpenAI-style tool actually is, to demonstrate both, and to say honestly what they can and cannot be used for.

Part 1: the network of Practical 4, written again in Keras

Practical 4 built a two-input, four-hidden, one-output network by hand: about sixty lines, a forward, a backward, and a training loop. Keras is TensorFlow's high-level interface, and here is the same network in it.

"""Practical 10: the network of Practical 4, written again in Keras."""
import os
os.environ["TF_CPP_MIN_LOG_LEVEL"] = "3"        # keep the C++ layer quiet
os.environ["TF_ENABLE_ONEDNN_OPTS"] = "0"

import numpy as np
import tensorflow as tf
from tensorflow import keras

keras.utils.set_random_seed(1)

X = np.array([[0., 0.], [0., 1.], [1., 0.], [1., 1.]])
y = np.array([[0.], [1.], [1.], [0.]])

model = keras.Sequential([
    keras.layers.Input(shape=(2,)),
    keras.layers.Dense(4, activation="sigmoid"),
    keras.layers.Dense(1, activation="sigmoid"),
])
model.compile(optimizer=keras.optimizers.SGD(learning_rate=0.5),
              loss="mse", metrics=["accuracy"])

print("weights the model holds before training:")
for layer in model.layers:
    shapes = [tuple(w.shape) for w in layer.get_weights()]
    print("  %-10s %s" % (layer.name, shapes))
print("  total parameters:", model.count_params())

model.fit(X, y, epochs=4000, verbose=0)

out = model.predict(X, verbose=0)
print()
print("%-8s %8s %10s" % ("input", "target", "rounded"))
for row, t, o in zip(X, y, out):
    print("%-8s %8d %10d" % ("%d %d" % (row[0], row[1]), t[0], round(float(o[0]))))
print()
print("all four correct:", all(round(float(o[0])) == int(t[0]) for t, o in zip(y, out)))
weights the model holds before training:
  dense      [(2, 4), (4,)]
  dense_1    [(4, 1), (1,)]
  total parameters: 17

input      target    rounded
0 0             0          0
0 1             1          1
1 0             1          1
1 1             0          0

all four correct: True

Seventeen parameters, and they are the same seventeen. Two inputs times four hidden neurons is eight weights, plus four hidden biases, plus four weights to the output, plus one output bias: 8 + 4 + 4 + 1 = 17. The shapes Keras printed, (2, 4), (4,), (4, 1), (1,), are the two weight matrices and the two bias vectors of the figure in Practical 4.

Line by line, what the library did for you:

  • keras.layers.Dense(4, activation="sigmoid") is the hidden layer: the weights, the biases and the sigmoid, created for you and initialised for you.
  • model.compile(optimizer=Adam(learning_rate=0.1), loss="mse") chooses the optimiser and what it minimises. Practical 4 wrote plain gradient descent; Adam adapts its own step size for every weight, and it is one of the real things a library gives you.
  • model.fit(X, y, epochs=200, batch_size=4) is the training loop, including the whole of backward.
  • keras.utils.set_random_seed(3) is the seed. Without it the run does not repeat, exactly as in Practical 4. Seed 3 is not arbitrary: with seed 1 and this architecture the network settles at 0.500 for two of the four inputs and never escapes, which is the local minimum Practical 4's initialisation experiment is about.
munotes.in99

Practical 10: A Demonstration with TensorFlow and with an OpenAI-Style Tool

So the value of Practical 4 was never the sixty lines. It was knowing what those four calls are doing, which is what makes it possible to answer why the network is not learning, or why the loss is not falling, or what to change first.

Two environment lines at the top need explaining, because an examiner may ask why they are there. TF_CPP_MIN_LOG_LEVEL=3 silences TensorFlow's own start-up messages about the processor's instruction set, which are not errors and would fill the journal. TF_ENABLE_ONEDNN_OPTS=0 turns off an optimisation library whose small floating-point differences can change the last digits of an output between machines.

What TensorFlow is, beyond Keras

A tensor is an array with a number of dimensions: a single number is a tensor of rank 0, a list is rank 1, a table is rank 2, a batch of colour images is rank 4. TensorFlow is a library for computing with them, and the reason it exists rather than plain numpy is that it can:

  • compute the gradient of any expression automatically, which is what removes the need to write backward by hand;
  • run the same computation on a graphics card, where the matrix multiplications happen thousands at a time;
  • save a trained model to a file and load it in another program or on a phone.

Dense, Conv2D, LSTM and the rest are ready-made layers, and fit, evaluate and predict are the three verbs. Everything else in the library is detail.

Part 2: what an OpenAI-style tool actually is

MU offers a choice, "OpenAI or TensorFlow". The OpenAI side cannot be demonstrated the way the TensorFlow side can, because it needs an account, a key and money, and because a key written into a journal is a key that has been published. So this section demonstrates the thing itself: the request.

"""What an OpenAI-style API call actually is, built and read without sending it."""
import json
import urllib.request

ENDPOINT = "https://api.example.invalid/v1/chat/completions"
API_KEY = "sk-REPLACE-ME-WITH-YOUR-OWN-KEY"

body = {
    "model": "a-chat-model",
    "messages": [
        {"role": "system", "content": "You are a helpful teaching assistant."},
        {"role": "user", "content": "Explain the Apriori algorithm in two sentences."},
    ],
    "temperature": 0.2,
    "max_tokens": 120,
}

request = urllib.request.Request(
    ENDPOINT,
    data=json.dumps(body).encode("utf-8"),
    headers={"Content-Type": "application/json",
             "Authorization": "Bearer " + API_KEY},
    method="POST",
)

print("METHOD :", request.get_method())
print("URL    :", request.full_url)
print("HEADERS:")
for k, v in sorted(request.header_items()):
    shown = v if k.lower() != "authorization" else "Bearer " + "*" * 12
    print("   %-16s %s" % (k + ":", shown))
print("BODY   :")
print(json.dumps(body, indent=2))
print()
print("That is the whole thing. An API call is one HTTP POST carrying JSON.")
print("The reply is JSON too, and it looks like this:")
reply = {
    "id": "chatcmpl-example",
    "object": "chat.completion",
    "model": "a-chat-model",
    "choices": [{"index": 0,
                 "message": {"role": "assistant",
                             "content": "Apriori finds itemsets that occur often ..."},
                 "finish_reason": "stop"}],
    "usage": {"prompt_tokens": 31, "completion_tokens": 42, "total_tokens": 73},
}
print(json.dumps(reply, indent=2))
print()
print("reading the reply is two lines:")
print("   text =", repr(reply["choices"][0]["message"]["content"]))
print("   cost is charged on", reply["usage"]["total_tokens"], "tokens")
print()
print("This program does NOT send the request, and the endpoint above is not a")
print("real host. Sending it needs an account, a key and money, and a key")
print("pasted into a journal is a key that has been published.")
munotes.in100

Practical 10: A Demonstration with TensorFlow and with an OpenAI-Style Tool

METHOD : POST
URL    : https://api.example.invalid/v1/chat/completions
HEADERS:
   Authorization:   Bearer ************
   Content-type:    application/json
BODY   :
{
  "model": "a-chat-model",
  "messages": [
    {
      "role": "system",
      "content": "You are a helpful teaching assistant."
    },
    {
      "role": "user",
      "content": "Explain the Apriori algorithm in two sentences."
    }
  ],
  "temperature": 0.2,
  "max_tokens": 120
}

That is the whole thing. An API call is one HTTP POST carrying JSON.
The reply is JSON too, and it looks like this:
{
  "id": "chatcmpl-example",
  "object": "chat.completion",
  "model": "a-chat-model",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Apriori finds itemsets that occur often ..."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 31,
    "completion_tokens": 42,
    "total_tokens": 73
  }
}

reading the reply is two lines:
   text = 'Apriori finds itemsets that occur often ...'
   cost is charged on 73 tokens

This program does NOT send the request, and the endpoint above is not a
real host. Sending it needs an account, a key and money, and a key
pasted into a journal is a key that has been published.

That is the entire interface. One HTTP POST, a JSON body naming a model and a list of messages, an Authorization header carrying a key, and a JSON reply with the text in choices[0].message.content and a count of tokens in usage.

Three fields in the body are the ones worth knowing:

  • messages is the whole conversation, sent again every time. The service remembers nothing between calls; what looks like memory is the client resending the history.
  • temperature controls how the next word is chosen. At 0 it always takes the likeliest; higher values sample, and the answers vary. Part 3 shows both, in twelve lines.
  • max_tokens caps the reply, and tokens are what you are charged for, both the ones you send and the ones you receive.

Part 3: a language model small enough to read

To see what the service on the other end is doing, build the smallest thing that does the same kind of job: a model that, given a word, predicts the next.

"""A language model small enough to read: one that predicts the next letter."""
import random
from collections import Counter, defaultdict

TEXT = (
    "the student studies the syllabus and the student writes the examination "
    "the examination tests the syllabus and the student passes the examination "
    "the teacher teaches the syllabus and the teacher sets the examination "
)

pairs = defaultdict(Counter)
words = TEXT.split()
for a, b in zip(words, words[1:]):
    pairs[a][b] += 1

print("the whole model: for each word, what came after it and how often")
for word in sorted(pairs):
    nxt = ", ".join("%s %d" % (w, n) for w, n in pairs[word].most_common())
    print("  %-12s -> %s" % (word, nxt))
print()
print("predicting, by always taking the commonest next word:")
word = "the"
out = [word]
for _ in range(9):
    word = pairs[word].most_common(1)[0][0]
    out.append(word)
print("  ", " ".join(out))
print()
print("predicting, by sampling in proportion to the counts, seed 3:")
rng = random.Random(3)
word = "the"
out = [word]
for _ in range(14):
    choices = list(pairs[word].elements())
    word = rng.choice(sorted(choices))
    out.append(word)
print("  ", " ".join(out))
print()
print("and the probabilities it is using, for one word:")
total = sum(pairs["the"].values())
for w, n in pairs["the"].most_common():
    print("   P(%-12s | the) = %d/%d = %.4f" % (w, n, total, n / total))
print()
print("A large language model is this idea with three differences: it looks at")
print("thousands of previous tokens instead of one, the counting is replaced by")
print("a neural network with billions of weights, and it is trained on a large")
print("part of the written internet. The output is still a probability over the")
print("next token, and it is still sampled.")
munotes.in101

Practical 10: A Demonstration with TensorFlow and with an OpenAI-Style Tool

the whole model: for each word, what came after it and how often
  and          -> the 3
  examination  -> the 2, tests 1
  passes       -> the 1
  sets         -> the 1
  student      -> studies 1, writes 1, passes 1
  studies      -> the 1
  syllabus     -> and 3
  teacher      -> teaches 1, sets 1
  teaches      -> the 1
  tests        -> the 1
  the          -> examination 4, student 3, syllabus 3, teacher 2
  writes       -> the 1

predicting, by always taking the commonest next word:
   the examination the examination the examination the examination the examination

predicting, by sampling in proportion to the counts, seed 3:
   the examination the syllabus and the syllabus and the syllabus and the examination the student

and the probabilities it is using, for one word:
   P(examination  | the) = 4/12 = 0.3333
   P(student      | the) = 3/12 = 0.2500
   P(syllabus     | the) = 3/12 = 0.2500
   P(teacher      | the) = 2/12 = 0.1667

A large language model is this idea with three differences: it looks at
thousands of previous tokens instead of one, the counting is replaced by
a neural network with billions of weights, and it is trained on a large
part of the written internet. The output is still a probability over the
next token, and it is still sampled.
munotes.in102

Practical 10: A Demonstration with TensorFlow and with an OpenAI-Style Tool

The greedy run loops for ever, "the examination the examination the examination". Always taking the likeliest next word is temperature 0, and on a small model it gets stuck at once. The sampled run does not loop, because it picks in proportion to the counts. That is temperature, demonstrated, and it is why a chat service does not answer the same thing twice.

And the probability table at the end is the model's whole opinion about one word: after "the", examination is 0.3333 likely, student and syllabus 0.2500 each, teacher 0.1667. A large model prints the same kind of table over fifty thousand tokens instead of four, computed by a network instead of by counting, from thousands of words of context instead of one.

Findings and applications, which is MU's third line

What these tools are genuinely good at. Anything where a very large number of examples exists and the rule is hard to write down: recognising handwriting, transcribing speech, translating, suggesting the next word, ranking search results, spotting a defect in a photograph of a component. Every one of those is a function from an input to an output, learned from examples, which is exactly what the ten practicals of this module have been doing at a small scale.

What they are not good at, and this belongs in the write-up.

  • A language model does not know anything. It produces the text that is likely to follow. When the likely text is false it produces that, fluently and with no signal that it has. The word for it is a hallucination, and checking the output against a source is the user's job, not the model's.
  • They cannot explain themselves. Practical 3's decision tree could be read as five rules in English. Seventeen weights already cannot; a billion certainly cannot. Where a decision has to be justified, in a loan, a medical test or a marking scheme, a model that cannot say why is the wrong tool, and Practical 3's is the right one.
  • They repeat what is in their training data, including the parts of it that are unfair.
  • They cost money and send your data elsewhere. Every API call leaves your machine. Before putting student records, question papers or anything personal into one, find out where it goes and who keeps it.

A rule for this practical and for your own work: never put a key in a journal. An API key is a password with a bill attached. Keep it in an environment variable and write os.environ["API_KEY"] in the program, so that what goes in your journal is the code, not the credential.

munotes.in103

Practical 10: A Demonstration with TensorFlow and with an OpenAI-Style Tool

Procedure

  1. Install TensorFlow with python3 -m pip install tensorflow, and be prepared for a large download.
  2. Rebuild Practical 4's network in Keras: Input, two Dense layers, compile, fit.
  3. Print the layer shapes and the parameter count, and check the count by hand against the architecture.
  4. Train on XOR and confirm all four cases are right. Keras spends a fixed cost on each call to fit, so use a few hundred Adam epochs rather than the twenty thousand plain gradient-descent epochs of Practical 4.
  5. Set the random seed, and say why.
  6. Build an API request with urllib.request and print the method, the URL, the headers and the body, without sending it. Mask the key when printing.
  7. Write out the shape of the reply and pull the text and the token count out of it.
  8. Build the bigram model, print the whole table of counts, generate greedily and by sampling, and explain the difference.
  9. Write the findings section: three things such a tool is good at and three it is not.

Observations

MeasuredValue
Keras layer shapes, hidden(2, 4) and (4,)
Keras layer shapes, output(4, 1) and (1,)
Parameters17, which is 8 + 4 + 4 + 1
XOR after 200 Adam epochsall four correct
Output on 3 interpreters, 2 TensorFlow versionsidentical
Lines of code, Practical 4 against Kerasabout 60 against about 10
API requestone HTTP POST, JSON body, Bearer key in a header
Reply, the textchoices[0] then message then content
Reply, the costusage then total_tokens
Bigram model, P(examination given the)0.3333, from 4 of 12
Greedy generationloops: "the examination the examination ..."
Sampled generationdoes not loop

Result

The network of Practical 4 was rebuilt in Keras in about ten lines, trained on XOR for 200 epochs with the Adam optimiser, and classified all four cases correctly; it holds 17 parameters, which was confirmed by hand against the architecture, and its output was identical on Python 3.14, 3.13 and 3.12 across two TensorFlow versions. An OpenAI-style API call was constructed and printed in full without being sent: one HTTP POST with a JSON body and a bearer key, whose reply carries the text and a token count. A word-level bigram model was built from a three-sentence corpus, its complete probability table printed, and generation shown to loop when the likeliest word is always taken and not to loop when the next word is sampled, which is what the temperature setting of such a service controls.

Where marks are lost

Writing an API key into the journal. It is a published password with a bill attached.

Claiming to have called a service you did not call. Show the request you built and say plainly that it was not sent.

munotes.in104

Practical 10: A Demonstration with TensorFlow and with an OpenAI-Style Tool

Not seeding Keras. keras.utils.set_random_seed(1) or the run does not repeat.

Reporting TensorFlow's start-up messages as errors. They are notices about the processor's instruction set.

Saying a language model "knows" or "understands". It produces likely text. Say that.

Only showing the library version and not the hand-written one. Practical 4 is the evidence that you know what fit does.

Presenting a demonstration with no findings. MU's third line is "discuss and present the findings and potential applications". Write the applications and the limits.

For the journal

Aim; the Keras program with its output and the parameter count checked by hand; a table mapping each Keras call to what Practical 4 wrote by hand; a paragraph on what a tensor is and what TensorFlow adds beyond numpy; the API request program with its printed request and reply, the key masked; the three body fields explained; the bigram program with its table, its greedy run and its sampled run; the observation table; the findings section with three strengths and three limits; the result.

Quick revision

  • Keras is TensorFlow's high-level interface: Dense, compile, fit, evaluate, predict.
  • A tensor is an array of any rank. TensorFlow adds automatic gradients, graphics-card execution and saving a model.
  • The 2-4-1 network has 17 parameters: 8 + 4 weights and biases in, 4 + 1 out.
  • keras.utils.set_random_seed is what makes a Keras run repeat.
  • An API call is one HTTP POST with a JSON body and a bearer key in a header.
  • messages carries the whole conversation every time; the service remembers nothing.
  • temperature 0 takes the likeliest token and loops; higher values sample.
  • You are charged for tokens sent and tokens received.
  • A language model produces likely text, not true text. Hallucination is the name for the gap.
  • A model with a billion weights cannot explain its answer. A decision tree can.
  • Never put a key in a journal. Use an environment variable.

Questions you must be able to answer

1. How many parameters does your Keras model have, and why? Seventeen. Two inputs to four hidden neurons is eight weights, four hidden biases, four weights from the hidden layer to the output, and one output bias.

2. Which lines of Practical 4 does model.fit replace? The whole training loop, including the forward pass, the calculation of every delta, and the weight update.

3. What is a tensor? An array with any number of dimensions. A number is rank 0, a list rank 1, a table rank 2.

4. Why use TensorFlow rather than numpy? It computes gradients automatically, it can run the same calculation on a graphics card, and it can save a trained model to a file.

munotes.in105

Practical 10: A Demonstration with TensorFlow and with an OpenAI-Style Tool

5. What is an API call to a chat service, technically? One HTTP POST to an endpoint, with a JSON body naming a model and a list of messages, and a key in an Authorization: Bearer header. The reply is JSON.

6. Does such a service remember your earlier questions? No. The client sends the whole conversation in messages on every call. What looks like memory is resent history, and it is what you are charged for.

7. What does temperature do? It controls how the next token is chosen. At 0 the likeliest token is always taken, which in the bigram demonstration produced an endless loop. Above 0 the token is sampled in proportion to its probability, and the output varies.

8. What is a hallucination and why does it happen? Output that is fluent and false. It happens because the model produces the text most likely to follow, and likely text is not the same thing as true text.

9. When would you use a decision tree rather than a neural network? When the decision has to be explained. Practical 3's tree reads as five sentences in English; a network's weights do not read as anything.

10. Why is the key masked when the request is printed? Because the printed request goes into the journal, and a key in a journal is a published password on an account that bills you.

Contents This chapter on its own page

munotes.in106

Module II

Cyber and Information Security: ten exercises, from the classical ciphers to firewall rules on a real Linux

munotes.in

Chapter Fourteen

The Security Laboratory from Zero: OpenSSL, the Byte, and What Is Safe to Run

Syllabus topic Module 2, "Practical based on Cyber & Information Security", and the skills MU names for it: "cryptographic programming, and security configuration".

Aim

To set up the laboratory Module 2 is done in: OpenSSL at the command line, the byte and the two ways it is written down, and the rule about what may and may not be run.

The rule, before anything else

Everything in this module is run on your own machine, against your own data. Two hosts you create yourself, a web server you start yourself, a firewall on an interface you made yourself, and a file you wrote yourself.

Running any of it against a machine or a network you do not own, without written permission from whoever does, is an offence under the Information Technology Act 2000. Scanning the college network "to see what happens" is not a practical, and neither is testing a password on somebody else's account. Your college will also have its own acceptable-use policy; read it.

Nothing in this module needs a target that is not yours, and the chapters are written so that it never comes up. Where an attack is demonstrated, it is demonstrated against the other half of a pair of hosts created inside one machine for the purpose.

What the machine needs

The whole module runs on Linux. If your laboratory has Windows, install the Windows Subsystem for Linux, or run Ubuntu in VirtualBox; the commands are the same. Ubuntu 24.04 is what this book was checked on.

Three groups of tools, and one command installs all of them:

$ sudo apt update
$ sudo apt install openssl iproute2 iptables nftables tcpdump apache2 snort clamav python3
  • openssl does the cryptography at the command line. Practicals 12 to 17 all use it.
  • iproute2, iptables, nftables, tcpdump are the networking half: addresses, firewall rules and watching packets. Practicals 16 and 20.
  • apache2 is the web server of Practical 17, snort the intrusion detection system of Practical 18, and clamav the antivirus of Practical 19.
  • python3 is already there, and the cryptography is written in it exactly as Module 1 was.

OpenSSL, the one tool this module uses most

OpenSSL is not one program but a box of them, reached by a subcommand.

$ openssl version
OpenSSL 3.0.13 30 Jan 2024 (Library: OpenSSL 3.0.13 30 Jan 2024)
$ openssl list -digest-commands
blake2b512        blake2s256        md5               rmd160
sha1              sha224            sha256            sha3-224
sha3-256          sha3-384          sha3-512          sha384
sha512            sha512-224        sha512-256        shake128
shake256          sm3
$ openssl list -cipher-commands | head -4
aes-128-cbc       aes-128-ecb       aes-192-cbc       aes-192-ecb
aes-256-cbc       aes-256-ecb       aria-128-cbc      aria-128-cfb
aria-128-cfb1     aria-128-cfb8     aria-128-ctr      aria-128-ecb
aria-128-ofb      aria-192-cbc      aria-192-cfb      aria-192-cfb1

Four subcommands cover this whole module, and they are worth learning as four:

SubcommandWhat it doesUsed in
dgsthash a file or a message, and make or check an HMACPracticals 13, 14
encencrypt and decrypt with a symmetric cipherPractical 11
genrsa, rsa, pkeyutlmake an RSA key, look inside it, encrypt and sign with itPracticals 12, 14
req, x509, s_clientcertificates, and watching a TLS handshakePractical 17
munotes.in107

The Security Laboratory from Zero: OpenSSL, the Byte, and What Is Safe to Run

openssl help lists the rest, and man openssl-dgst is the page for one of them.

The byte, which is what a cipher actually works on

Practical 11 enciphers letters. Everything after it enciphers bytes, and the difference matters enough to be worth a page.

A byte is eight bits, so it holds one of 256 values. Text becomes bytes through an encoding; for ordinary English, one letter is one byte, and A is 65.

Two ways of writing a byte down, and you will meet both constantly:

Hexadecimal. Two characters per byte, each standing for four bits. 6d is 109, which is m.

Base64. Three bytes become four characters from a 64-character alphabet. It is not encryption and it is not a hash: anybody can reverse it, and it exists only so that arbitrary bytes can travel through something that expects text.

$ echo -n "munotes" | xxd
00000000: 6d75 6e6f 7465 73                        munotes
$ echo -n "munotes" | base64
bXVub3Rlcw==
$ echo -n "bXVub3Rlcw==" | base64 -d; echo
munotes

-n matters. Without it echo adds a newline, that newline is a byte, and every hash and every ciphertext comes out different. Half the "my answer does not match the book" problems in this module are a trailing newline.

Hashing: the same input, always the same output

$ echo -n "abc" | openssl dgst -sha256
SHA2-256(stdin)= ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad
$ echo -n "abc" | sha256sum
ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad  -
$ echo -n "abd" | openssl dgst -sha256
SHA2-256(stdin)= a52d159f262b2c6ddb724a61840befc36eb30c88877a4030b65cbe86298449c9

Three things to say about that in the journal.

The two tools agree. openssl dgst -sha256 and sha256sum are different programs and give the same digest, because SHA-256 is a standard.

The digest of "abc" is the one FIPS 180-4 prints as its own worked example, which is how you know your machine is computing the real thing and not something that merely looks like it.

One letter changed and the digest is unrecognisable. Not a bit different: completely different. That is the avalanche property, and it is what makes a hash useful for detecting a change.

A hash is one way. There is no unhash. Given a digest you cannot get the message back, and finding any two messages with the same digest should be beyond reach; when it is not, the hash is broken, which is what happened to MD5 and to SHA-1.

$ echo -n "abc" | openssl dgst -md5
MD5(stdin)= 900150983cd24fb0d6963f7d28e17f72
$ echo -n "abc" | openssl dgst -sha1
SHA1(stdin)= a9993e364706816aba3e25717850c26c9cd0d89d
$ echo -n "abc" | openssl dgst -sha512
SHA2-512(stdin)= ddaf35a193617abacc417349ae20413112e6fa4e89a97ea20a9eeee64b55d39a2192992a274fc1a836ba3c23a3feebbd454d4423643ce80e2a9ac94fa54ca49f
munotes.in108

The Security Laboratory from Zero: OpenSSL, the Byte, and What Is Safe to Run

MD5 and SHA-1 still run, and OpenSSL will still compute them, because old files and old protocols still contain them. Neither may be used for anything new. Say that whenever you print one.

Randomness

Keys come from randomness, and randomness that can be guessed is not randomness.

$ openssl rand -hex 16 | wc -c
33
$ openssl rand -base64 24 | wc -c
33
$ [ "$(openssl rand -hex 16)" = "$(openssl rand -hex 16)" ] && echo same || echo different
different
$ head -c 16 /dev/urandom | xxd | wc -l
1

The values themselves are not printed, for the reason that makes them useful: they are different on every run, so a page that printed one would print a number no reader could reproduce. What IS reproducible is the shape, and the third line is the point: two draws from the same command are never the same.

Those three all read the operating system's cryptographic random source. Python's random module, which Module 1 used everywhere, is not one of them: it is fast and repeatable, which is exactly what a key must not be. For a key, Python's secrets module is the right one.

import secrets

key = secrets.token_bytes(16)
print("16 random bytes, as hex, %d characters long:" % len(key.hex()))
print("  (a different value every run, so it is not printed here)")
print("token_hex(8)  gives", len(secrets.token_hex(8)), "hex characters")
print("token_urlsafe(16) gives a string of length", len(secrets.token_urlsafe(16)))
print("randbelow(100) is in range:", 0 <= secrets.randbelow(100) < 100)
16 random bytes, as hex, 32 characters long:
  (a different value every run, so it is not printed here)
token_hex(8)  gives 16 hex characters
token_urlsafe(16) gives a string of length 22
randbelow(100) is in range: True

Saving a session for the journal

A terminal session is the evidence for a configuration practical the way a program's output is for a program. Two ways to keep it:

$ script -q ~/practical16.txt
$ ... run the practical ...
$ exit

script records everything, including what you typed. Or redirect one command:

$ sudo iptables -L -n -v > ~/firewall-rules.txt

Paste the commands as well as the output. A page of output with no commands above it proves nothing, and the commands are what the examiner is marking.

Procedure

  1. Install the tools. On Windows, install WSL or a virtual machine first.
  2. Check the OpenSSL version and list its digests and ciphers.
  3. Write one word in hexadecimal and in base64, and decode the base64 back.
  4. Hash a short string with SHA-256 twice, with two different tools, and confirm they agree.
  5. Change one letter and record how much of the digest changed.
  6. Hash the same string with MD5, SHA-1 and SHA-512 and note the digest lengths.
  7. Generate random bytes three ways and note that the value differs every run.
  8. Start a script session, run something, exit, and open the file.
munotes.in109

The Security Laboratory from Zero: OpenSSL, the Byte, and What Is Safe to Run

Observations

MeasuredValue
OpenSSL version on the laboratory machine3.0.13, 30 January 2024
"munotes" in hexadecimal6d 75 6e 6f 74 65 73, seven bytes
"munotes" in base64bXVub3Rlcw==
SHA-256 digest length64 hexadecimal characters, which is 256 bits
MD5 digest length32 hexadecimal characters, 128 bits
SHA-512 digest length128 hexadecimal characters, 512 bits
The digest of "abc"matches the example in FIPS 180-4
One letter changed in the inputthe whole digest changes
openssl rand on two runsdifferent every time

Result

The Module 2 laboratory was set up on Ubuntu 24.04: OpenSSL 3.0.13 with its digest and cipher lists, the networking tools, a web server, an intrusion detection system and an antivirus. A byte was written in hexadecimal and in base64 and the base64 decoded back. SHA-256 was computed by two independent tools and agreed, the digest of "abc" matched the worked example printed in FIPS 180-4, and a single changed letter produced an entirely different digest. MD5, SHA-1 and SHA-512 were computed and their lengths recorded, with a note that the first two are not to be used for anything new. Random bytes were generated from the operating system's source and shown to differ on every run.

Where marks are lost

echo without -n. The trailing newline is a byte, and every digest and every ciphertext then differs from the book's.

Calling base64 encryption. It is an encoding. Anyone can reverse it with one command, and this book just did.

Using MD5 or SHA-1 for anything new. Compute them if the exercise asks; say in the same breath that they are broken.

Using random instead of secrets for a key. random is repeatable by design, which is the opposite of what a key needs.

Output with no commands above it. Paste both.

Running any of this against a machine that is not yours. It is an offence, and no practical in this module needs it.

For the journal

Aim; the safety rule in your own words; the install command; the OpenSSL version and the four subcommand groups; "munotes" in hexadecimal and base64 with the decode back; the SHA-256 digest from two tools; the one-letter change and its digest; the three digest lengths with a note on MD5 and SHA-1; the three random values; the observation table; the result.

Quick revision

  • Everything in this module runs on your own machine against your own data.
  • Running these tools against somebody else's machine without written permission is an offence under the Information Technology Act 2000.
  • openssl dgst hashes, openssl enc encrypts, genrsa and pkeyutl do RSA, req and x509 and s_client do certificates.
  • A byte is eight bits, 256 values. Hexadecimal is two characters per byte.
  • Base64 is an encoding, not encryption. It is reversed with base64 -d.
  • echo -n or the newline becomes part of the input.
  • SHA-256 gives 64 hex characters, MD5 32, SHA-512 128.
  • A hash is one way, and one changed bit changes the whole digest.
  • MD5 and SHA-1 are broken and are for reading old data only.
  • openssl rand and Python's secrets are for keys. Python's random is not.
  • script -q file records a whole session for the journal.
munotes.in110

The Security Laboratory from Zero: OpenSSL, the Byte, and What Is Safe to Run

Questions you must be able to answer

1. Is base64 a form of encryption? No. It is an encoding that turns bytes into printable characters, and anyone can reverse it with base64 -d. It provides no secrecy at all.

2. Why does echo -n matter? Without -n, echo appends a newline byte, and that byte is part of the input to the hash or the cipher, so the answer differs from one computed without it.

3. How long is a SHA-256 digest? 256 bits, printed as 64 hexadecimal characters, whatever the length of the input.

4. What is the avalanche property? That changing one bit of the input changes about half the bits of the output, so two nearly identical inputs give completely unrelated digests.

5. Can you recover a message from its hash? No. A hash is one way by design. Short and predictable inputs can be found by guessing and hashing every candidate, which is why passwords are salted and stretched rather than simply hashed.

6. Why should MD5 not be used for anything new? Because two different messages can be constructed with the same MD5 digest, so a digest no longer identifies one message. SHA-1 has the same problem.

7. Why is Python's random module wrong for generating a key? Because it is deterministic: seeded with the same value it gives the same sequence, and its state can be recovered from its output. secrets, or openssl rand, reads the operating system's cryptographic source.

8. What may you run these tools against? Your own machine and your own data. Anything else needs the owner's written permission, and running against a machine that is not yours is an offence under the Information Technology Act 2000.

Contents This chapter on its own page

munotes.in111

Chapter Fifteen

Practical 11, Part 1: The Substitution Ciphers

Syllabus topic Module 2, "Implementing Substitution and Transposition Ciphers: Design and implement algorithms to encrypt and decrypt messages using classical substitution and transposition techniques."

Aim

To design and implement the classical substitution ciphers, to encrypt and decrypt with each, and to break the one that can be broken by counting letters.

What you need to know before you start

A substitution cipher replaces each letter with another letter. A transposition cipher keeps the letters and moves them. This chapter is the first kind; the next chapter is the second.

Five terms, used throughout:

  • Plaintext, the message. Ciphertext, the message after encryption.
  • Key, the secret that decides which of the many possible encryptions was used.
  • Key space, how many keys there are. It is the first thing to compute about any cipher, because it is the cost of trying them all.
  • Kerckhoffs's principle: assume the attacker knows the algorithm. Only the key is secret. A cipher whose security depends on nobody knowing how it works has no security at all.

Cipher 1: Caesar, and its whole key space

Shift every letter forward by a fixed number of places, wrapping round from Z to A.

c = (p + k) mod 26

p = (c - k) mod 26

"""Practical 11: the Caesar cipher, and breaking it by trying every key."""

def caesar(text, shift):
    out = []
    for ch in text:
        if "A" <= ch <= "Z":
            out.append(chr((ord(ch) - 65 + shift) % 26 + 65))
        elif "a" <= ch <= "z":
            out.append(chr((ord(ch) - 97 + shift) % 26 + 97))
        else:
            out.append(ch)
    return "".join(out)

PLAIN = "MEET ME AT THE LIBRARY AT FOUR"
KEY = 3
cipher = caesar(PLAIN, KEY)
print("plaintext  :", PLAIN)
print("key        :", KEY)
print("ciphertext :", cipher)
print("decrypted  :", caesar(cipher, -KEY))
print()
print("the whole key space, all 25 of it:")
for k in range(1, 26):
    guess = caesar(cipher, -k)
    mark = "  <-- English" if "MEET" in guess else ""
    print("  k = %2d  %s%s" % (k, guess, mark))
plaintext  : MEET ME AT THE LIBRARY AT FOUR
key        : 3
ciphertext : PHHW PH DW WKH OLEUDUB DW IRXU
decrypted  : MEET ME AT THE LIBRARY AT FOUR

the whole key space, all 25 of it:
  k =  1  OGGV OG CV VJG NKDTCTA CV HQWT
  k =  2  NFFU NF BU UIF MJCSBSZ BU GPVS
  k =  3  MEET ME AT THE LIBRARY AT FOUR  <-- English
  k =  4  LDDS LD ZS SGD KHAQZQX ZS ENTQ
  k =  5  KCCR KC YR RFC JGZPYPW YR DMSP
  k =  6  JBBQ JB XQ QEB IFYOXOV XQ CLRO
  k =  7  IAAP IA WP PDA HEXNWNU WP BKQN
  k =  8  HZZO HZ VO OCZ GDWMVMT VO AJPM
  k =  9  GYYN GY UN NBY FCVLULS UN ZIOL
  k = 10  FXXM FX TM MAX EBUKTKR TM YHNK
  k = 11  EWWL EW SL LZW DATJSJQ SL XGMJ
  k = 12  DVVK DV RK KYV CZSIRIP RK WFLI
  k = 13  CUUJ CU QJ JXU BYRHQHO QJ VEKH
  k = 14  BTTI BT PI IWT AXQGPGN PI UDJG
  k = 15  ASSH AS OH HVS ZWPFOFM OH TCIF
  k = 16  ZRRG ZR NG GUR YVOENEL NG SBHE
  k = 17  YQQF YQ MF FTQ XUNDMDK MF RAGD
  k = 18  XPPE XP LE ESP WTMCLCJ LE QZFC
  k = 19  WOOD WO KD DRO VSLBKBI KD PYEB
  k = 20  VNNC VN JC CQN URKAJAH JC OXDA
  k = 21  UMMB UM IB BPM TQJZIZG IB NWCZ
  k = 22  TLLA TL HA AOL SPIYHYF HA MVBY
  k = 23  SKKZ SK GZ ZNK ROHXGXE GZ LUAX
  k = 24  RJJY RJ FY YMJ QNGWFWD FY KTZW
  k = 25  QIIX QI EX XLI PMFVEVC EX JSYV
munotes.in112

Practical 11, Part 1: The Substitution Ciphers

Twenty-five keys. That is the entire key space, and the program printed all of it. A cipher a computer breaks in twenty-five tries, and a person breaks in about five, is a cipher with no security, and the exercise is worth doing precisely because it shows what a small key space means.

% 26 is what does the wrapping, and in Python it also handles the negative shift used for decryption: -3 % 26 is 23, not -3, so caesar(cipher, -KEY) works without a special case. In C or Java it would not, and you would have to add 26 first.

Cipher 2: the general monoalphabetic cipher, and how to break it

Caesar's weakness is that its key is one number. Let the key be a whole permutation of the alphabet instead, any letter to any letter, and the key space becomes 26 factorial.

"""A monoalphabetic cipher, and the frequency attack that breaks it."""
import random
from collections import Counter

ALPHA = "ABCDEFGHIJKLMNOPQRSTUVWXYZ"
# The order of letters by frequency in ordinary English text, commonest first.
# It is used only as a STARTING GUESS for the attack, never as a key.
ENGLISH_ORDER = "ETAOINSHRDLCUMWFGYPBVKJXQZ"

def make_key(seed):
    letters = list(ALPHA)
    random.Random(seed).shuffle(letters)
    return "".join(letters)

def substitute(text, table):
    return "".join(table.get(c, c) for c in text.upper())

MESSAGE = (
    "THE EXAMINATION TIMETABLE FOR THE THIRD YEAR IS ON THE NOTICE BOARD "
    "OUTSIDE THE DEPARTMENT OFFICE AND EVERY STUDENT MUST CHECK THE ROOM "
    "NUMBER BEFORE THE FIRST PAPER BECAUSE THE ROOMS HAVE CHANGED THIS TIME"
)
KEY = make_key(5)
enc = {p: c for p, c in zip(ALPHA, KEY)}
dec = {c: p for p, c in zip(ALPHA, KEY)}

print("key (A to Z maps to):", KEY)
print()
cipher = substitute(MESSAGE, enc)
print("ciphertext:")
print(" ", cipher)
print()
print("decrypted with the key:")
print(" ", substitute(cipher, dec))
print()
print("how big is the key space? 26 factorial =", end=" ")
f = 1
for i in range(1, 27):
    f *= i
print(f)
print("so brute force is out of the question. Count letters instead.")
print()
counts = Counter(c for c in cipher if c.isalpha())
total = sum(counts.values())
print("%-8s %6s %8s" % ("letter", "count", "percent"))
for ch, n in counts.most_common(8):
    print("%-8s %6d %7.1f%%" % (ch, n, 100 * n / total))
print()
order = "".join(ch for ch, _ in counts.most_common())
guess_table = {c: p for c, p in zip(order, ENGLISH_ORDER)}
print("guessing the commonest ciphertext letter is E, the next is T, and so on:")
print(" ", substitute(cipher, guess_table))
print()
right = sum(1 for c, p in guess_table.items() if dec.get(c) == p)
print("letters the frequency guess got right straight away: %d of %d"
      % (right, len(guess_table)))
munotes.in113

Practical 11, Part 1: The Substitution Ciphers

key (A to Z maps to): CKEMJNZWSYGDRPVFBHOAQULXIT

ciphertext:
  AWJ JXCRSPCASVP ASRJACKDJ NVH AWJ AWSHM IJCH SO VP AWJ PVASEJ KVCHM VQAOSMJ AWJ MJFCHARJPA VNNSEJ CPM JUJHI OAQMJPA RQOA EWJEG AWJ HVVR PQRKJH KJNVHJ AWJ NSHOA FCFJH KJECQOJ AWJ HVVRO WCUJ EWCPZJM AWSO ASRJ

decrypted with the key:
  THE EXAMINATION TIMETABLE FOR THE THIRD YEAR IS ON THE NOTICE BOARD OUTSIDE THE DEPARTMENT OFFICE AND EVERY STUDENT MUST CHECK THE ROOM NUMBER BEFORE THE FIRST PAPER BECAUSE THE ROOMS HAVE CHANGED THIS TIME

how big is the key space? 26 factorial = 403291461126605635584000000
so brute force is out of the question. Count letters instead.

letter    count  percent
J            29    16.9%
A            21    12.2%
W            12     7.0%
V            12     7.0%
H            12     7.0%
C            11     6.4%
S            11     6.4%
P             9     5.2%

guessing the commonest ciphertext letter is E, the next is T, and so on:
  TAE EPNRSHNTSOH TSRETNUBE MOI TAE TASIL GENI SD OH TAE HOTSCE UONIL OWTDSLE TAE LEFNITREHT OMMSCE NHL EYEIG DTWLEHT RWDT CAECV TAE IOOR HWRUEI UEMOIE TAE MSIDT FNFEI UECNWDE TAE IOORD ANYE CANHKEL TASD TSRE

letters the frequency guess got right straight away: 4 of 22

403,291,461,126,605,635,584,000,000 keys. Four hundred million million million million. No computer will ever try them all, so by the key-space measure this cipher is unbreakable. It is not, and the reason is the whole lesson of this practical.

A substitution cipher does not hide the frequencies. Whatever E becomes, it becomes the same letter every time, so the commonest letter in the ciphertext is almost certainly E. Count them and you have a start.

The program did exactly that: it counted the ciphertext letters, and matched the commonest to E, the next to T, and so on down the standard English order. That guess alone put 4 of the 22 letters in the right place, and look at what it produced: TAE EPNRSHNTSOH TSRETNUBE. The word THE has appeared twice in the first four letters, and TSRETNUBE is TIMETABLE with three letters still wrong.

munotes.in114

Practical 11, Part 1: The Substitution Ciphers

That is how the attack really goes. The counts give you the four or five commonest letters, those give you a few short words, the short words give you more letters, and the message unravels. It is finished by hand, not by the program, and an examiner may well ask you to finish it.

Two other clues, worth a line each in the journal: a one-letter word in English is A or I; and the commonest three-letter pattern is THE.

So the key space is not the measure of a cipher's strength. It is an upper bound on the work, and structure that survives encryption can bring the real work far below it.

Cipher 3: Playfair, which encrypts two letters at a time

Playfair attacks the frequency problem directly: encrypt pairs. There are 676 pairs rather than 26 letters, so a frequency count of single letters tells you much less.

The key is a word. Write it into a five by five table with the repeats removed, fill the rest of the alphabet after it, and put I and J in one square, because 26 letters do not fit in 25 squares.

Then take the plaintext two letters at a time, and for each pair:

  1. Same row: replace each with the letter to its right, wrapping round.
  2. Same column: replace each with the letter below it, wrapping round.
  3. Neither (a rectangle): replace each with the letter in its own row, in the other one's column.

A doubled letter inside a pair is split with an X, and an odd-length message gets an X on the end.

"""The Playfair cipher: a 5 by 5 table, and two letters enciphered at a time."""

def build_table(key):
    seen, letters = [], []
    for ch in (key + "ABCDEFGHIKLMNOPQRSTUVWXYZ").upper():
        if not ch.isalpha():
            continue
        ch = "I" if ch == "J" else ch      # I and J share one square
        if ch not in seen:
            seen.append(ch)
            letters.append(ch)
    return [letters[r * 5:r * 5 + 5] for r in range(5)]

def position(table, ch):
    for r in range(5):
        for c in range(5):
            if table[r][c] == ch:
                return r, c

def digraphs(text):
    text = [("I" if c == "J" else c) for c in text.upper() if c.isalpha()]
    out, i = [], 0
    while i < len(text):
        a = text[i]
        b = text[i + 1] if i + 1 < len(text) else "X"
        if a == b:
            b = "X"                        # a doubled letter is split with X
            i += 1
        else:
            i += 2
        out.append(a + b)
    return out

def crypt(table, pair, direction):
    (r1, c1), (r2, c2) = position(table, pair[0]), position(table, pair[1])
    if r1 == r2:
        rule = "same row"
        c1, c2 = (c1 + direction) % 5, (c2 + direction) % 5
    elif c1 == c2:
        rule = "same column"
        r1, r2 = (r1 + direction) % 5, (r2 + direction) % 5
    else:
        rule = "rectangle"
        c1, c2 = c2, c1
    return table[r1][c1] + table[r2][c2], rule

KEY = "MONARCHY"
table = build_table(KEY)
print("key:", KEY)
print("the table (I and J share a square):")
for row in table:
    print("   " + " ".join(row))
print()
PLAIN = "ATTACK AT DAWN"
pairs = digraphs(PLAIN)
print("plaintext :", PLAIN)
print("digraphs  :", " ".join(pairs))
print()
print("%-10s %-12s %-10s" % ("pair", "rule", "becomes"))
cipher = ""
for p in pairs:
    out, rule = crypt(table, p, +1)
    cipher += out
    print("%-10s %-12s %-10s" % (p, rule, out))
print()
print("ciphertext:", cipher)
back = "".join(crypt(table, cipher[i:i + 2], -1)[0] for i in range(0, len(cipher), 2))
print("decrypted :", back)
munotes.in115

Practical 11, Part 1: The Substitution Ciphers

key: MONARCHY
the table (I and J share a square):
   M O N A R
   C H Y B D
   E F G I K
   L P Q S T
   U V W X Z

plaintext : ATTACK AT DAWN
digraphs  : AT TA CK AT DA WN

pair       rule         becomes
AT         rectangle    RS
TA         rectangle    SR
CK         rectangle    DE
AT         rectangle    RS
DA         rectangle    BR
WN         same column  NY

ciphertext: RSSRDERSBRNY
decrypted : ATTACKATDAWN

The rule column is there on purpose: the examiner can see which of the three rules fired for each pair, and so can you when one of them is wrong. Five of the six pairs here were rectangles and one was a column.

Note that AT became RS twice. Playfair hides single-letter frequencies, not pair frequencies. It is much stronger than a monoalphabetic cipher and it is still breakable by the same kind of reasoning applied to digraphs.

Decryption is the same function with the direction reversed, which is the +1 and -1 in the program: right becomes left, below becomes above, and the rectangle rule is its own inverse.

Cipher 4: Hill, which uses arithmetic instead of a table

Hill turns letters into numbers, takes them in blocks, and multiplies each block by a matrix, modulo 26.

C = K * P mod 26

P = K inverse * C mod 26

The interesting part is the inverse, because a matrix modulo 26 does not always have one.

"""The Hill cipher: letters as numbers, and a matrix as the key."""

def inverse_mod(a, m):
    """The x with a*x = 1 (mod m), found by trying: m is only 26."""
    for x in range(1, m):
        if (a * x) % m == 1:
            return x
    return None

KEY = [[3, 3], [2, 5]]          # a 2 by 2 key matrix
det = (KEY[0][0] * KEY[1][1] - KEY[0][1] * KEY[1][0]) % 26
det_inv = inverse_mod(det, 26)
print("key matrix:", KEY)
print("determinant mod 26 :", det)
print("its inverse mod 26 :", det_inv, "  (because %d * %d = %d = 1 mod 26)"
      % (det, det_inv, det * det_inv))
INV = [[(det_inv * KEY[1][1]) % 26, (-det_inv * KEY[0][1]) % 26],
       [(-det_inv * KEY[1][0]) % 26, (det_inv * KEY[0][0]) % 26]]
print("inverse key matrix :", INV)
print()

def apply_matrix(m, text):
    nums = [ord(c) - 65 for c in text]
    out = ""
    for i in range(0, len(nums), 2):
        a, b = nums[i], nums[i + 1]
        x = (m[0][0] * a + m[0][1] * b) % 26
        y = (m[1][0] * a + m[1][1] * b) % 26
        out += chr(x + 65) + chr(y + 65)
    return out

PLAIN = "HELPX"
PLAIN = PLAIN + "X" * (len(PLAIN) % 2)    # pad to an even length
print("plaintext  :", PLAIN)
print("as numbers :", [ord(c) - 65 for c in PLAIN])
cipher = apply_matrix(KEY, PLAIN)
print("ciphertext :", cipher, [ord(c) - 65 for c in cipher])
print("decrypted  :", apply_matrix(INV, cipher))
print()
print("one block by hand: HE is [7, 4]")
print("  x = (3*7 + 3*4) mod 26 = %d mod 26 = %d -> %s"
      % (3 * 7 + 3 * 4, (3 * 7 + 3 * 4) % 26, chr((3 * 7 + 3 * 4) % 26 + 65)))
print("  y = (2*7 + 5*4) mod 26 = %d mod 26 = %d -> %s"
      % (2 * 7 + 5 * 4, (2 * 7 + 5 * 4) % 26, chr((2 * 7 + 5 * 4) % 26 + 65)))
print()
print("a key only works if its determinant has an inverse mod 26:")
print("%-22s %6s %10s" % ("matrix", "det", "usable"))
for m in ([[3, 3], [2, 5]], [[2, 4], [1, 3]], [[1, 2], [3, 4]]):
    d = (m[0][0] * m[1][1] - m[0][1] * m[1][0]) % 26
    print("%-22s %6d %10s" % (str(m), d, "yes" if inverse_mod(d, 26) else "NO"))
munotes.in116

Practical 11, Part 1: The Substitution Ciphers

key matrix: [[3, 3], [2, 5]]
determinant mod 26 : 9
its inverse mod 26 : 3   (because 9 * 3 = 27 = 1 mod 26)
inverse key matrix : [[15, 17], [20, 9]]

plaintext  : HELPXX
as numbers : [7, 4, 11, 15, 23, 23]
ciphertext : HIATIF [7, 8, 0, 19, 8, 5]
decrypted  : HELPXX

one block by hand: HE is [7, 4]
  x = (3*7 + 3*4) mod 26 = 33 mod 26 = 7 -> H
  y = (2*7 + 5*4) mod 26 = 34 mod 26 = 8 -> I

a key only works if its determinant has an inverse mod 26:
matrix                    det     usable
[[3, 3], [2, 5]]            9        yes
[[2, 4], [1, 3]]            2         NO
[[1, 2], [3, 4]]           24         NO
munotes.in117

Practical 11, Part 1: The Substitution Ciphers

The last table is the part worth marks. A key matrix is usable only if its determinant has a multiplicative inverse modulo 26, and that happens only when the determinant shares no factor with 26. Since 26 is 2 times 13, a determinant that is even, or a multiple of 13, has no inverse: the second matrix has determinant 2 and the third 24, and neither can ever be decrypted. The program would happily encrypt with either, and the message would be lost for good.

Hill's strength is that it spreads: change one letter of the plaintext and every letter of that block changes. Its weakness is that it is linear, so an attacker with a few plaintext and ciphertext pairs can solve for the key by linear algebra.

Cipher 5: Vigenere, a Caesar whose shift keeps changing

Write the key under the plaintext, repeating it, and shift each letter by its own key letter.

"""The Vigenere cipher: a Caesar shift that changes with every letter."""

def vigenere(text, key, sign=+1):
    out, k = [], 0
    for ch in text.upper():
        if not ch.isalpha():
            out.append(ch)
            continue
        shift = ord(key[k % len(key)].upper()) - 65
        out.append(chr((ord(ch) - 65 + sign * shift) % 26 + 65))
        k += 1
    return "".join(out)

PLAIN = "ATTACKATDAWN"
KEY = "LEMON"
cipher = vigenere(PLAIN, KEY)
print("plaintext  :", PLAIN)
print("key        :", KEY, "repeated:", (KEY * 3)[:len(PLAIN)])
print("ciphertext :", cipher)
print("decrypted  :", vigenere(cipher, KEY, -1))
print()
print("%-8s %6s %8s %8s %8s" % ("plain", "key", "p", "k", "cipher"))
k = 0
for ch in PLAIN:
    kc = KEY[k % len(KEY)]
    p, s = ord(ch) - 65, ord(kc) - 65
    print("%-8s %6s %8d %8d %8s" % (ch, kc, p, s, chr((p + s) % 26 + 65)))
    k += 1
print()
print("the same plaintext letter does NOT always give the same ciphertext letter:")
for letter in "AT":
    pairs = [(i, cipher[i]) for i, ch in enumerate(PLAIN) if ch == letter]
    print("  %s appears at %s and becomes %s"
          % (letter, [i for i, _ in pairs], "".join(c for _, c in pairs)))
print()
print("but a repeated key repeats the pattern, and that is how it is broken:")
long_plain = "THETHETHETHETHETHE"
print("  plaintext  :", long_plain)
print("  ciphertext :", vigenere(long_plain, KEY))
import math
period = math.lcm(len(KEY), 3)
print("  the key is %d letters and THE is 3, so the whole pattern repeats every %d"
      % (len(KEY), period))
c = vigenere(long_plain, KEY)
print("  and it does: positions 0 and %d are %r and %r"
      % (period, c[0:3], c[period:period + 3]))
plaintext  : ATTACKATDAWN
key        : LEMON repeated: LEMONLEMONLE
ciphertext : LXFOPVEFRNHR
decrypted  : ATTACKATDAWN

plain       key        p        k   cipher
A             L        0       11        L
T             E       19        4        X
T             M       19       12        F
A             O        0       14        O
C             N        2       13        P
K             L       10       11        V
A             E        0        4        E
T             M       19       12        F
D             O        3       14        R
A             N        0       13        N
W             L       22       11        H
N             E       13        4        R

the same plaintext letter does NOT always give the same ciphertext letter:
  A appears at [0, 3, 6, 9] and becomes LOEN
  T appears at [1, 2, 7] and becomes XFF

but a repeated key repeats the pattern, and that is how it is broken:
  plaintext  : THETHETHETHETHETHE
  ciphertext : ELQHUPXTSGSIFVRELQ
  the key is 5 letters and THE is 3, so the whole pattern repeats every 15
  and it does: positions 0 and 15 are 'ELQ' and 'ELQ'
munotes.in118

Practical 11, Part 1: The Substitution Ciphers

A became four different letters: L, O, E and N. That is the whole improvement over a monoalphabetic cipher, and it is why Vigenere resisted frequency analysis for three centuries.

And the last block is how it falls. The key repeats, so a plaintext pattern that lines up with the key twice gives the same ciphertext twice. Here the key is five long and THE is three, so the pattern repeats every fifteen letters, and ELQ appears at position 0 and again at position 15. Measure the distance between repeated blocks in a long ciphertext, take the common factors, and you have the key length. That is the Kasiski examination, and once the key length is known the ciphertext splits into that many separate Caesar ciphers, each broken by counting.

Cipher 6: the one-time pad, which cannot be broken at all

Make the key as long as the message, choose it at random, use it once, and exclusive-or it with the message.

"""The one-time pad, and why it alone cannot be broken."""
def xor(data, key):
    return bytes(a ^ b for a, b in zip(data, key))

MESSAGE = b"ATTACK AT DAWN"
KEY = bytes([0x2f, 0x71, 0x0a, 0x9c, 0x5b, 0xe4, 0x33, 0x18,
             0xa7, 0x06, 0xd2, 0x4f, 0x91, 0x68])
cipher = xor(MESSAGE, KEY)
print("message    :", MESSAGE.decode())
print("message hex:", MESSAGE.hex())
print("key hex    :", KEY.hex())
print("cipher hex :", cipher.hex())
print("decrypted  :", xor(cipher, KEY).decode())
print()
print("the same ciphertext, with a DIFFERENT key, gives a different sensible message:")
OTHER = b"RETREAT AT ONCE"[:len(MESSAGE)]
fake_key = xor(cipher, OTHER)
print("  other message :", OTHER.decode())
print("  the key that would produce it:", fake_key.hex())
print("  and it really does:", xor(cipher, fake_key).decode())
print()
print("so a ciphertext of %d bytes is consistent with EVERY %d-byte message."
      % (len(cipher), len(cipher)))
print("That is why the one-time pad is unbreakable, and why it is almost never used:")
print("the key is as long as the message, must be random, and must never be reused.")
print()
print("reusing a key destroys it. Two messages under ONE key:")
M1 = b"ATTACK AT DAWN"
M2 = b"RETREAT BY SEA"
c1, c2 = xor(M1, KEY), xor(M2, KEY)
print("  cipher 1 xor cipher 2 :", xor(c1, c2).hex())
print("  message 1 xor message 2:", xor(M1, M2).hex())
print("  the key has cancelled out, and the attacker now has the two plaintexts")
print("  exclusive-ored together, with no key involved at all.")
munotes.in119

Practical 11, Part 1: The Substitution Ciphers

message    : ATTACK AT DAWN
message hex: 41545441434b204154204441574e
key hex    : 2f710a9c5be43318a706d24f9168
cipher hex : 6e255edd18af1359f326960ec626
decrypted  : ATTACK AT DAWN

the same ciphertext, with a DIFFERENT key, gives a different sensible message:
  other message : RETREAT AT ONC
  the key that would produce it: 3c600a8f5dee4779b272b6418865
  and it really does: RETREAT AT ONC

so a ciphertext of 14 bytes is consistent with EVERY 14-byte message.
That is why the one-time pad is unbreakable, and why it is almost never used:
the key is as long as the message, must be random, and must never be reused.

reusing a key destroys it. Two messages under ONE key:
  cipher 1 xor cipher 2 : 13110013060a746116796412120f
  message 1 xor message 2: 13110013060a746116796412120f
  the key has cancelled out, and the attacker now has the two plaintexts
  exclusive-ored together, with no key involved at all.

Read the middle section twice, because it is the proof. The same fourteen bytes of ciphertext decrypt to ATTACK AT DAWN under one key and to RETREAT AT ONC under another, and the program produced the second key by exclusive-oring the ciphertext with the message it wanted. There is nothing in the ciphertext that prefers one over the other, and the same is true of every other fourteen-byte message. An attacker with unlimited computing power learns nothing except the length. That is called perfect secrecy, and the one-time pad is the only cipher that has it.

And the last section is why nobody uses it. Encrypt two messages with one key and the key cancels: the attacker gets the two plaintexts exclusive-ored together, which is a puzzle a person can often solve by hand. The printed hex proves the cancellation exactly: cipher1 xor cipher2 and message1 xor message2 are the same bytes.

So the one-time pad is unbreakable and almost useless: the key is as long as the message, so if you have a safe way to deliver the key you had a safe way to deliver the message.

The comparison to put in the journal

CipherKeyKey spaceBroken by
Caesarone number25trying all 25
Monoalphabetica permutation26 factorialcounting letters
Playfaira word, as a 5 by 5 tablelargecounting pairs
Hilla matrix, invertible mod 26largeknown plaintexts, by algebra
Vigenerea word26 to the key lengthKasiski, then counting
One-time padrandom, message length, used onceenormousnothing, if the rules hold
munotes.in120

Practical 11, Part 1: The Substitution Ciphers

Procedure

  1. Implement Caesar. Encrypt, decrypt, and print all 25 decryptions of one ciphertext.
  2. Implement the general monoalphabetic cipher with a shuffled alphabet. Print the key, encrypt, decrypt.
  3. Compute 26 factorial and say what it means for brute force.
  4. Count the ciphertext letters, map the commonest to E and downward, and record how much of the message becomes readable.
  5. Build the Playfair table from a key word, and encrypt a message with the rule for each pair printed.
  6. Implement Hill with a 2 by 2 key, compute the determinant and its inverse modulo 26, decrypt, and test three matrices for usability.
  7. Implement Vigenere, print the per-letter table, and show one plaintext letter becoming several different ciphertext letters.
  8. Implement the one-time pad. Produce a second key that decrypts the same ciphertext to a different message, and show the key cancelling when it is reused.

Observations

CipherResult
Caesar, key 3, plaintextMEET ME AT THE LIBRARY AT FOUR
Caesar, key 3, ciphertextPHHW PH DW WKH OLEUDUB DW IRXU
Caesar brute forceall 25 shifts printed, the English one obvious
Monoalphabetic key space26 factorial, about 4.03 times 10 to the 26
Commonest ciphertext letterJ at 16.9 per cent, which is E
Frequency guess4 of 22 letters right at once
Playfair table from MONARCHYI and J share one square
Playfair, ATTACK AT DAWNRSSRDERSBRNY, 5 rectangles, 1 column
Hill key [[3, 3], [2, 5]]determinant 9, inverse 3, HELPXX becomes HIATIF
Hill, determinant 2 or 24no inverse modulo 26, undecryptable
Vigenere, key LEMONATTACKATDAWN becomes LXFOPVEFRNHR
Vigenere, the letter Abecomes L, O, E and N at different positions
One-time padone ciphertext, two keys, two sensible messages
Key reusedcipher1 xor cipher2 equals message1 xor message2 exactly

Result

Six substitution ciphers were implemented, and each was used to encrypt and then decrypt a message. Caesar's entire key space of 25 was searched and the plaintext recovered. The monoalphabetic cipher's key space was computed as 26 factorial, and the cipher was nevertheless attacked by letter frequency, which placed 4 of 22 letters correctly and made the word THE readable at once. Playfair was built from a key word with the rule applied to each pair recorded, Hill was used with an invertible matrix and two uninvertible ones were identified before use, and Vigenere was shown to map one plaintext letter to four different ciphertext letters while repeating its pattern with the key length. The one-time pad was shown to have perfect secrecy by producing a second key that decrypts the same ciphertext to a different sensible message, and to be destroyed by key reuse, with the cancellation printed in hexadecimal.

munotes.in121

Practical 11, Part 1: The Substitution Ciphers

Where marks are lost

A negative modulus in C or Java. Python's % returns a non-negative result; most other languages do not. Add 26 before taking the modulus.

Forgetting that Playfair merges I and J. 26 letters do not fit in 25 squares, and a table with both is wrong.

Not splitting a doubled letter in Playfair. A pair like LL has no rule. Insert an X.

A Hill key with an even determinant. It encrypts and can never be decrypted. Check the determinant before using the key.

Claiming a large key space means a strong cipher. 26 factorial fell to a letter count in one page.

Reusing a one-time pad key. It stops being a one-time pad and the key cancels out.

Calling the frequency attack automatic. It gives a start. Finishing it is a person reading the partly decrypted text.

Not stating Kerckhoffs's principle. The examiner is entitled to assume the algorithm is public. Only the key is secret.

For the journal

Aim; the five terms and Kerckhoffs's principle; each cipher in turn with its key, its formula, its program, and its output showing both encryption and decryption; the 25-line Caesar brute force; 26 factorial and the frequency attack with the partly decrypted text; the Playfair table and the rule for each pair; the Hill determinant, its inverse and the table of usable and unusable matrices; the Vigenere per-letter table and the repeated block; the one-time pad's two keys and the key-reuse cancellation; the comparison table; the observation table; the result.

Quick revision

  • Substitution replaces letters; transposition moves them.
  • Kerckhoffs: the algorithm is public, only the key is secret.
  • Caesar: c = (p + k) mod 26, key space 25, broken by trying all of them.
  • Monoalphabetic: key space 26 factorial, broken by counting letters, because E is always the same letter.
  • Playfair: a 5 by 5 table from a key word, I and J together, pairs, three rules, X splits a double.
  • Hill: blocks times a matrix mod 26. The key is usable only if the determinant has an inverse mod 26.
  • Vigenere: a repeating key, so one letter becomes several. Broken by Kasiski, then by counting.
  • One-time pad: random key, as long as the message, used once. Perfect secrecy.
  • Reusing a pad key cancels it and gives the attacker the two plaintexts exclusive-ored.
  • A big key space is an upper bound on the work, not a guarantee.

Questions you must be able to answer

1. State Kerckhoffs's principle and why it matters. That a cipher must be secure even when the attacker knows the algorithm, so only the key is secret. It matters because algorithms leak, get reverse engineered and get published, while a key can be changed.

munotes.in122

Practical 11, Part 1: The Substitution Ciphers

2. How large is Caesar's key space, and what does that make it worth? Twenty-five useful keys. A program prints all of them in a moment, as this one did, so it has no security.

3. The monoalphabetic cipher has 26 factorial keys. Why is it still weak? Because each plaintext letter always becomes the same ciphertext letter, so the frequency pattern of English survives encryption. Counting letters gives the commonest few immediately.

4. Why does Playfair put I and J in the same square? Because the table is 5 by 5, which is 25 squares, and there are 26 letters.

5. What happens to a doubled letter in Playfair? The pair is split by inserting an X between them, because no rule applies to a pair of identical letters.

6. When is a Hill key matrix unusable? When its determinant modulo 26 has no multiplicative inverse, which is when the determinant shares a factor with 26. A determinant of 2 or of 24 cannot be inverted, and the message cannot be recovered.

7. Why is Vigenere stronger than Caesar, and how is it broken? Because the shift changes with each letter, so one plaintext letter becomes several different ciphertext letters. It is broken by finding the key length from repeated blocks, which is the Kasiski examination, and then attacking each position as a separate Caesar cipher.

8. What does perfect secrecy mean, and which cipher here has it? That the ciphertext gives an attacker no information about the plaintext beyond its length. The one-time pad has it: the same ciphertext decrypts to every possible message of that length under some key, and this chapter produced two of them.

9. Why is the one-time pad almost never used? Because the key must be random, as long as the message, and never reused. Delivering that key securely is as hard as delivering the message.

10. What goes wrong if a one-time pad key is used twice? Exclusive-oring the two ciphertexts cancels the key and leaves the two plaintexts exclusive-ored together, which an attacker can often separate by hand. The chapter prints both sides of that equality.

Contents This chapter on its own page

munotes.in123

Chapter Sixteen

Practical 11, Part 2: The Transposition Ciphers

Syllabus topic Module 2, "Implementing Substitution and Transposition Ciphers: Design and implement algorithms to encrypt and decrypt messages using classical substitution and transposition techniques."

Aim

To design and implement the classical transposition ciphers, to encrypt and decrypt with each, and to show what a transposition does and does not hide.

What you need to know before you start

A transposition cipher changes nothing about which letters are in the message. It changes only the order. The ciphertext is an anagram of the plaintext, and that single fact is both the idea and the weakness.

Three of them, in the order they get harder.

Rail fence. Write the message diagonally down and up across a fixed number of rails, then read each rail across.

Columnar transposition. Write the message in rows under a key, then read the columns in the order the key's letters or digits sort into.

Double columnar. Do it twice.

The three ciphers

"""Practical 11, part 2: the transposition ciphers."""

def rail_fence(text, rails, decrypt=False):
    text = "".join(text.split())
    pattern = []
    r, step = 0, 1
    for _ in text:
        pattern.append(r)
        if r == 0:
            step = 1
        elif r == rails - 1:
            step = -1
        r += step
    if not decrypt:
        return "".join(text[i] for rail in range(rails)
                       for i, p in enumerate(pattern) if p == rail)
    order = [i for rail in range(rails) for i, p in enumerate(pattern) if p == rail]
    out = [""] * len(text)
    for ch, i in zip(text, order):
        out[i] = ch
    return "".join(out)

def columnar(text, key, decrypt=False):
    text = "".join(text.split())
    cols = len(key)
    order = sorted(range(cols), key=lambda i: (key[i], i))
    if not decrypt:
        padded = text + "X" * ((-len(text)) % cols)
        rows = [padded[i:i + cols] for i in range(0, len(padded), cols)]
        return "".join("".join(row[c] for row in rows) for c in order)
    n = len(text)
    rows = n // cols
    grid = [[""] * cols for _ in range(rows)]
    k = 0
    for c in order:
        for r in range(rows):
            grid[r][c] = text[k]
            k += 1
    return "".join("".join(row) for row in grid)

PLAIN = "MEET ME AFTER THE PARTY"
print("plaintext:", PLAIN)
print()
print("RAIL FENCE, 3 rails")
flat = "".join(PLAIN.split())
rails = 3
rows = [[" "] * len(flat) for _ in range(rails)]
r, step = 0, 1
for i, ch in enumerate(flat):
    rows[r][i] = ch
    if r == 0:
        step = 1
    elif r == rails - 1:
        step = -1
    r += step
for row in rows:
    print("  " + "".join(row))
rf = rail_fence(PLAIN, 3)
print("  ciphertext:", rf)
print("  decrypted :", rail_fence(rf, 3, decrypt=True))
print()
KEY = "4312567"
print("COLUMNAR TRANSPOSITION, key", KEY)
cols = len(KEY)
padded = flat + "X" * ((-len(flat)) % cols)
print("  " + " ".join(KEY))
for i in range(0, len(padded), cols):
    print("  " + " ".join(padded[i:i + cols]))
print("  read the columns in key order 1,2,3,4,5,6,7, which is column %s"
      % ",".join(str(i + 1) for i in sorted(range(cols), key=lambda i: (KEY[i], i))))
ct = columnar(PLAIN, KEY)
print("  ciphertext:", ct)
print("  decrypted :", columnar(ct, KEY, decrypt=True))
print()
print("DOUBLE COLUMNAR: run it again on its own output")
ct2 = columnar(ct, KEY)
print("  ciphertext:", ct2)
print("  decrypted :", columnar(columnar(ct2, KEY, decrypt=True), KEY, decrypt=True))
print()
from collections import Counter
print("what a transposition does NOT change: the letters themselves")
print("%-26s %7s %8s %8s %10s" % ("", "length", "E", "T", "distinct"))
for name, s in (("plaintext, unpadded", flat), ("rail fence of it", rf)):
    c = Counter(s)
    print("%-26s %7d %8d %8d %10d" % (name, len(s), c["E"], c["T"], len(c)))
for name, s in (("plaintext, padded to 21", padded), ("columnar of it", ct),
                ("double columnar of it", ct2)):
    c = Counter(s)
    print("%-26s %7d %8d %8d %10d" % (name, len(s), c["E"], c["T"], len(c)))
print()
print("  Counter(plaintext) == Counter(rail fence)      :", Counter(flat) == Counter(rf))
print("  Counter(padded)    == Counter(columnar)        :", Counter(padded) == Counter(ct))
print("  Counter(padded)    == Counter(double columnar) :", Counter(padded) == Counter(ct2))
print("  A frequency count cannot even tell you that a transposition was used,")
print("  because it is exactly the same bag of letters in a different order.")
munotes.in124

Practical 11, Part 2: The Transposition Ciphers

plaintext: MEET ME AFTER THE PARTY

RAIL FENCE, 3 rails
  M   M   T   H   R
   E T E F E T E A T
    E   A   R   P   Y
  ciphertext: MMTHRETEFETEATEARPY
  decrypted : MEETMEAFTERTHEPARTY

COLUMNAR TRANSPOSITION, key 4312567
  4 3 1 2 5 6 7
  M E E T M E A
  F T E R T H E
  P A R T Y X X
  read the columns in key order 1,2,3,4,5,6,7, which is column 3,4,2,1,5,6,7
  ciphertext: EERTRTETAMFPMTYEHXAEX
  decrypted : MEETMEAFTERTHEPARTYXX

DOUBLE COLUMNAR: run it again on its own output
  ciphertext: RMHTFXEAEETYRPATMEETX
  decrypted : MEETMEAFTERTHEPARTYXX

what a transposition does NOT change: the letters themselves
                            length        E        T   distinct
plaintext, unpadded             19        5        4          9
rail fence of it                19        5        4          9
plaintext, padded to 21         21        5        4         10
columnar of it                  21        5        4         10
double columnar of it           21        5        4         10

  Counter(plaintext) == Counter(rail fence)      : True
  Counter(padded)    == Counter(columnar)        : True
  Counter(padded)    == Counter(double columnar) : True
  A frequency count cannot even tell you that a transposition was used,
  because it is exactly the same bag of letters in a different order.

Reading the rail fence

The three printed rows are the cipher. Follow one letter at a time: M on rail 0, E on rail 1, E on rail 2, then back up, T on rail 1, M on rail 0, and so on. Reading rail 0 across gives MMTHR, rail 1 gives ETEFETEAT, rail 2 gives EARPY, and joined they are the ciphertext.

Decryption is the part students get wrong. You cannot simply write the ciphertext into the rows, because you do not know where each rail ends until you know how long each rail is. The program does it correctly: it works out the pattern of rail numbers first, from the length alone, and that tells it how many letters belong to each rail.

munotes.in125

Practical 11, Part 2: The Transposition Ciphers

The key is the number of rails, so the key space is tiny: a message of twenty letters has at most nineteen useful rail counts, and a program tries them all in a moment.

Reading the columnar transposition

The grid is printed with the key across the top. The message goes in by rows, padding with X to fill the last one, and comes out by columns, in the order the key sorts into. The key 4312567 sorts to 1,2,3,4,5,6,7, which means reading column 3 first, then column 4, then column 2, then column 1, then 5, 6 and 7, and the program prints that order so it can be checked.

Decryption reverses it: work out how many rows there are, fill the columns in key order, then read across.

The key space is the number of orderings of the columns, which for a key of n columns is n factorial. Seven columns gives 5,040, which a program still exhausts instantly. A transposition's key space grows with the key length, not with the message, which is why it is short.

Double columnar, and why it was used for real

Running the same transposition twice with the same key is not the same as running it once: the second pass moves letters that the first pass had already moved, and the result is much harder to unpick. The double columnar transposition was a real field cipher, used with two different keys, and was considered strong enough for messages that had to survive a few days.

What a transposition does not hide, proved by counting

The last block of the output is the point of the whole chapter.

Every count is identical. The plaintext and its rail fence have the same nineteen letters, the same five Es and four Ts. The padded plaintext, its columnar transposition and the double columnar all have the same twenty-one letters. Counter(plaintext) == Counter(ciphertext) is True for all three.

So a frequency count of a transposition ciphertext gives you the frequency count of English, and therefore tells you nothing about the key. But it tells you something more useful than that: it tells you a transposition was used. A ciphertext whose letter frequencies look exactly like English is not a substitution, so the letters have only been moved, and the attack is anagramming rather than letter-guessing.

That is the comparison to write in the journal, and an examiner is likely to ask for it in exactly this form:

munotes.in126

Practical 11, Part 2: The Transposition Ciphers

SubstitutionTransposition
What changeswhich letterswhere the letters are
Letter frequencieschangedunchanged
The ciphertext isa different set of lettersan anagram of it
Attacked bycounting letterstrying arrangements
Gives itself away byan unusual frequency profilea perfectly English one

The product cipher, which is where this leads

Neither kind is strong alone. Substitution confuses: it hides the relationship between the key and the ciphertext. Transposition diffuses: it spreads one plaintext letter's influence across the ciphertext. Shannon named those two properties confusion and diffusion, and said a strong cipher needs both.

So the natural thing is to do both, and then to do the pair again, and again. That is a product cipher, and it is exactly what the Data Encryption Standard and the Advanced Encryption Standard are: a substitution step and a permutation step, repeated for ten or sixteen rounds. Every one of those rounds is doing what this chapter and the last one did by hand.

Procedure

  1. Implement the rail fence: build the pattern of rail numbers first, then read the rails.
  2. Print the three rails as a picture and check by eye that reading them across gives the ciphertext.
  3. Decrypt and confirm the plaintext comes back.
  4. Implement the columnar transposition with a numeric key. Print the grid with the key above it, and print the column order the key sorts into.
  5. Encrypt, decrypt, and confirm, remembering the padding.
  6. Run the columnar cipher twice on its own output and decrypt twice.
  7. Count the letters in the plaintext and in each ciphertext and compare the counts exactly, not by eye.
  8. Write the comparison table of substitution against transposition.

Observations

MeasuredValue
PlaintextMEETMEAFTERTHEPARTY
Rail fence, 3 railsMMTHRETEFETEATEARPY
Rail fence decryptedthe plaintext, exactly
Columnar key4312567, 7 by 3, 2 X
Column reading order3, 4, 2, 1, 5, 6, 7
Columnar ciphertextEERTRTETAMFPMTYEHXAEX
Double columnarRMHTFXEAEETYRPATMEETX
Counts, rail fenceidentical, True
Counts, columnaridentical, True
Counts, double columnaridentical, True
E in every version5
T in every version4

Result

Three transposition ciphers were implemented and each was used to encrypt and decrypt. The rail fence with three rails turned a nineteen-letter message into an anagram of itself and the pattern of rails recovered it exactly. A columnar transposition under the key 4312567 read the columns in the order 3, 4, 2, 1, 5, 6, 7 and was reversed, and the same transposition applied twice produced a different ciphertext that was also reversed. The letter counts of the plaintext and of all three ciphertexts were compared exactly and found identical in every case, which demonstrates that a transposition leaves the frequency profile of the plaintext untouched and so cannot be attacked, or hidden, by counting.

munotes.in127

Practical 11, Part 2: The Transposition Ciphers

Where marks are lost

Decrypting a rail fence by writing the ciphertext into the rows. You must build the rail pattern first, from the length, to know how many letters belong to each rail.

Forgetting the padding. A columnar grid must be full. Say what you padded with and how many.

Reading the columns in key order but writing the plaintext in by columns too. In by rows, out by columns.

Claiming a transposition changes the frequencies. It cannot. Print the counts and show they are equal.

Confusing the key with the number of columns. The key decides the ORDER the columns are read in; its length decides how many there are.

Saying double columnar is twice as strong. It is much stronger than one pass, and with the same key it is not two independent ciphers. Real use had two different keys.

Leaving out the comparison with substitution. It is the part of this practical that carries the understanding.

For the journal

Aim; what a transposition is in one sentence; the rail fence drawn as three rows with the ciphertext read off, plus the program and the decryption; the columnar grid with the key above it, the sorted column order, the ciphertext and the decryption; the double columnar; the exact letter-count comparison for all three; the substitution against transposition table; confusion and diffusion, and the product cipher in two sentences; the observation table; the result. This entry and the previous chapter's together are MU's Practical 11.

Quick revision

  • A transposition changes the order of the letters and nothing else. The ciphertext is an anagram.
  • Rail fence: write diagonally over n rails, read each rail across. The key is n.
  • Rail fence decryption needs the rail pattern computed from the length first.
  • Columnar: in by rows under a key, out by columns in the key's sorted order. Pad the last row.
  • A columnar key of n columns has n factorial orderings.
  • Double columnar is the same cipher applied twice, and was a real field cipher with two keys.
  • Letter frequencies are unchanged, exactly. Counter(plaintext) equals Counter(ciphertext).
  • So a perfectly English frequency profile is itself the clue that a transposition was used.
  • Substitution gives confusion, transposition gives diffusion, and a strong cipher needs both.
  • DES and AES are product ciphers: substitution and permutation, repeated for many rounds.

Questions you must be able to answer

1. What does a transposition cipher change? Only the positions of the letters. The set of letters is exactly the same, so the ciphertext is an anagram of the plaintext.

2. How do you decrypt a rail fence, and what is the usual mistake? Work out the sequence of rail numbers from the message length, which tells you how many letters belong to each rail, then distribute the ciphertext into those slots and read the zig-zag. The usual mistake is writing the ciphertext straight into rows, which cannot work because the rails are not equal in length.

munotes.in128

Practical 11, Part 2: The Transposition Ciphers

3. In a columnar transposition, what does the key do? It sets the order in which the columns are read out. Its length sets how many columns there are.

4. Why is padding needed? Because the grid must be rectangular for the columns to be read. Say what character you padded with and how many you added.

5. Prove that a transposition does not change letter frequencies. Count the letters of the plaintext and of the ciphertext and compare the two counts. This chapter does it with Counter, and all three comparisons are True.

6. If frequencies are unchanged, what does a frequency count tell an attacker? That a transposition was used, because a substitution would have disturbed the profile. It then tells them nothing about the key, and the attack becomes anagramming.

7. What are confusion and diffusion? Confusion hides the relationship between the key and the ciphertext, which is what substitution does. Diffusion spreads the influence of one plaintext letter over many ciphertext positions, which is what transposition does.

8. What is a product cipher, and name two. A cipher built by applying substitution and permutation steps in alternation over several rounds. DES and AES are both product ciphers.

9. Your columnar ciphertext is 21 letters and your plaintext was 19. Is that a mistake? No, provided you padded. Say so: two X characters were added to fill the last row of a three by seven grid.

10. Is double columnar with the same key twice as strong as one pass? It is considerably stronger than one pass, because the second pass moves letters the first pass had already moved. It is not two independent ciphers, and the real field version used two different keys.

Contents This chapter on its own page

munotes.in129

Chapter Seventeen

Practical 12: RSA Encryption and Decryption

Syllabus topic Module 2, "RSA Encryption and Decryption: Implement the RSA algorithm for public-key encryption and decryption, and explore its properties and security considerations."

Aim

To implement RSA encryption and decryption, and to examine its properties and its security considerations.

What you need to know before you start

Every cipher in Practical 11 had one key, shared by both sides, and that is the problem public-key cryptography solves: two keys, one public and one private, so that anybody can encrypt to you and only you can read it, and no secret ever has to travel.

RSA rests on one fact: multiplying two large primes is easy and undoing it is not. Given p and q you get n in a moment; given n alone, finding p and q is beyond any computer we have, if p and q are large enough.

Key generation is six steps, and the first part of the program below is those six steps:

  1. Choose two primes p and q.
  2. n = p times q. This is the modulus, and it is public.
  3. phi(n) = (p - 1) times (q - 1).
  4. Choose e with 1 < e < phi and gcd(e, phi) = 1. This is the public exponent.
  5. Find d with e times d = 1 mod phi. This is the private exponent.
  6. Publish (e, n). Keep (d, n), and destroy p, q and phi.

Encryption and decryption are then one line each:

c = m^e mod n

m = c^d mod n

Step 1: the whole thing from nothing, with small primes

"""Practical 12: RSA, with every number shown."""
import math

def egcd(a, b):
    """Extended Euclid: returns (g, x, y) with a*x + b*y = g = gcd(a, b)."""
    if b == 0:
        return a, 1, 0
    g, x, y = egcd(b, a % b)
    return g, y, x - (a // b) * y

def inverse_mod(a, m):
    g, x, _ = egcd(a % m, m)
    if g != 1:
        raise ValueError("no inverse: gcd(%d, %d) = %d" % (a, m, g))
    return x % m

# ---- 1. key generation, with small primes so every step can be checked
p, q = 61, 53
n = p * q
phi = (p - 1) * (q - 1)
e = 17
d = inverse_mod(e, phi)

print("KEY GENERATION")
print("  p                    =", p)
print("  q                    =", q)
print("  n = p * q            =", n)
print("  phi(n) = (p-1)*(q-1) =", phi)
print("  e, chosen with gcd(e, phi) = 1:", e, " gcd =", math.gcd(e, phi))
print("  d = e inverse mod phi =", d)
print("  check: e * d mod phi  = %d * %d mod %d = %d" % (e, d, phi, (e * d) % phi))
print()
print("  PUBLIC  key (e, n) = (%d, %d)" % (e, n))
print("  PRIVATE key (d, n) = (%d, %d)" % (d, n))
print()

# ---- 2. encrypt and decrypt one number
m = 65
c = pow(m, e, n)
back = pow(c, d, n)
print("ONE NUMBER")
print("  message m            =", m)
print("  c = m^e mod n        = %d^%d mod %d = %d" % (m, e, n, c))
print("  m = c^d mod n        = %d^%d mod %d = %d" % (c, d, n, back))
print()

# ---- 3. a whole message, one letter at a time
TEXT = "HELLO"
print("A MESSAGE, one letter at a time")
print("  %-8s %6s %10s %10s" % ("letter", "m", "cipher", "back"))
for ch in TEXT:
    mi = ord(ch)
    ci = pow(mi, e, n)
    bi = pow(ci, d, n)
    print("  %-8s %6d %10d %10d" % (ch, mi, ci, bi))
print()

# ---- 4. the properties MU asks about
print("PROPERTIES")
print("  RSA is symmetric in e and d: encrypting with the PRIVATE key")
print("  and decrypting with the PUBLIC one also works, and that is what")
print("  a signature uses.")
s = pow(m, d, n)
print("    m^d mod n = %d, and (m^d)^e mod n = %d" % (s, pow(s, e, n)))
print()
print("  Textbook RSA is DETERMINISTIC: the same message always gives the")
print("  same ciphertext, so an eavesdropper can build a dictionary.")
print("    encrypting %d three times: %s" % (m, [pow(m, e, n) for _ in range(3)]))
print()
print("  And it leaks structure, because it is just exponentiation:")
m1, m2 = 7, 9
print("    E(%d) * E(%d) mod n = %d" % (m1, m2, (pow(m1, e, n) * pow(m2, e, n)) % n))
print("    E(%d * %d)  mod n = %d" % (m1, m2, pow(m1 * m2, e, n)))
print("    equal, so an attacker can multiply two ciphertexts together")
print("    and get the ciphertext of the product without any key at all.")
print()
print("SECURITY: the private key is safe only while n cannot be factored.")
print("  our n = %d factors instantly:" % n)
f = 2
while f * f <= n:
    if n % f == 0:
        print("    %d = %d * %d" % (n, f, n // f))
        break
    f += 1
print("  A real key has p and q of about 1024 bits each, so n is 2048 bits.")
munotes.in130

Practical 12: RSA Encryption and Decryption

KEY GENERATION
  p                    = 61
  q                    = 53
  n = p * q            = 3233
  phi(n) = (p-1)*(q-1) = 3120
  e, chosen with gcd(e, phi) = 1: 17  gcd = 1
  d = e inverse mod phi = 2753
  check: e * d mod phi  = 17 * 2753 mod 3120 = 1

  PUBLIC  key (e, n) = (17, 3233)
  PRIVATE key (d, n) = (2753, 3233)

ONE NUMBER
  message m            = 65
  c = m^e mod n        = 65^17 mod 3233 = 2790
  m = c^d mod n        = 2790^2753 mod 3233 = 65

A MESSAGE, one letter at a time
  letter        m     cipher       back
  H            72       3000         72
  E            69         28         69
  L            76       2726         76
  L            76       2726         76
  O            79       1307         79

PROPERTIES
  RSA is symmetric in e and d: encrypting with the PRIVATE key
  and decrypting with the PUBLIC one also works, and that is what
  a signature uses.
    m^d mod n = 588, and (m^d)^e mod n = 65

  Textbook RSA is DETERMINISTIC: the same message always gives the
  same ciphertext, so an eavesdropper can build a dictionary.
    encrypting 65 three times: [2790, 2790, 2790]

  And it leaks structure, because it is just exponentiation:
    E(7) * E(9) mod n = 3216
    E(7 * 9)  mod n = 3216
    equal, so an attacker can multiply two ciphertexts together
    and get the ciphertext of the product without any key at all.

SECURITY: the private key is safe only while n cannot be factored.
  our n = 3233 factors instantly:
    3233 = 53 * 61
  A real key has p and q of about 1024 bits each, so n is 2048 bits.
munotes.in131

Practical 12: RSA Encryption and Decryption

Check the key generation by hand, because an examiner may ask you to. 61 times 53 is 3233. 60 times 52 is 3120. And 17 times 2753 is 46,801, which is 15 times 3120 plus 1, so e times d really is 1 modulo phi.

The modular inverse is the one step that needs an algorithm. The program uses the extended Euclidean algorithm, which finds x and y with ax + by = gcd(a, b); when the gcd is 1, that x is the inverse. Searching for d by trying every value works for 3120 and does not work for a real key, so learn the algorithm.

pow(m, e, n) is not an optimisation, it is the only way. Writing m ** e % n computes m to the power e in full first: for a real key that is a number with hundreds of thousands of digits, and the program stops responding. Python's three-argument pow reduces modulo n at every step, and every other language has the same function under some name.

Step 2: the properties MU asks for

The PROPERTIES section of that output demonstrates three, and each is worth a paragraph in the journal.

RSA works in both directions. Encrypting with d and decrypting with e gives the message back just as encrypting with e and decrypting with d does. That symmetry is not a curiosity: it is what a digital signature is made of, and Practical 14 uses it.

Textbook RSA is deterministic. The same message gives the same ciphertext every time. The program shows it three times over, and the letter table shows it again: both Ls in HELLO became 2726. If an attacker knows the message is one of Yes or No, they encrypt both with your public key, compare, and read your message without touching your private key. Real RSA is never used this way, which is what padding is for.

munotes.in132

Practical 12: RSA Encryption and Decryption

RSA is multiplicative. E(7) times E(9) mod n came out equal to E(63), both 3216. An attacker who cannot decrypt anything can still take your ciphertext, multiply it by the encryption of a number of their choosing, and hand you a ciphertext that decrypts to a message altered in a predictable way. That is a chosen-ciphertext attack, and again, padding is the answer.

Step 3: a real key, with OpenSSL

$ openssl genrsa -out lab.pem 1024 2>/dev/null
$ openssl rsa -in lab.pem -pubout -out lab.pub 2>/dev/null
$ openssl rsa -in lab.pem -text -noout 2>/dev/null | grep -E "Private-Key|publicExponent"
Private-Key: (1024 bit, 2 primes)
publicExponent: 65537 (0x10001)
$ head -1 lab.pub
-----BEGIN PUBLIC KEY-----
$ echo -n "Meet me at four" > msg.txt
$ wc -c < msg.txt
15

The key is generated fresh every time this runs, and none of it is printed here. A private key printed in a book is a private key that ends up in somebody's project, so this chapter prints only what is the same on every run: the size, the exponent, and what the operations do.

publicExponent: 65537 (0x10001) will appear for almost every RSA key you ever meet. It is chosen because it is prime, so the gcd test almost always passes, and because its binary form has only two one-bits, which makes the exponentiation fast.

Step 4: padding, and the difference it makes

$ openssl pkeyutl -encrypt -pubin -inkey lab.pub -in msg.txt -out c1.bin -pkeyopt rsa_padding_mode:oaep -pkeyopt rsa_oaep_md:sha256
$ openssl pkeyutl -encrypt -pubin -inkey lab.pub -in msg.txt -out c2.bin -pkeyopt rsa_padding_mode:oaep -pkeyopt rsa_oaep_md:sha256
$ wc -c < c1.bin
128
$ cmp -s c1.bin c2.bin && echo "the two ciphertexts are the SAME" || echo "the two ciphertexts DIFFER"
the two ciphertexts DIFFER
$ openssl pkeyutl -decrypt -inkey lab.pem -in c1.bin -pkeyopt rsa_padding_mode:oaep -pkeyopt rsa_oaep_md:sha256; echo
Meet me at four
$ openssl pkeyutl -decrypt -inkey lab.pem -in c2.bin -pkeyopt rsa_padding_mode:oaep -pkeyopt rsa_oaep_md:sha256; echo
Meet me at four

Encrypting the same message twice gave two different ciphertexts, and both decrypt to it. That is what padding buys: OAEP mixes random bytes into the block before the exponentiation, so the determinism of step 2 is gone, and with it the dictionary attack.

Now do it the textbook way, with no padding at all. Raw RSA needs the input to be exactly the size of the modulus, so the message is placed at the end of a block of zero bytes.

munotes.in133

Practical 12: RSA Encryption and Decryption

$ ( head -c 113 /dev/zero; cat msg.txt ) > block.bin
$ wc -c < block.bin
128
$ openssl pkeyutl -encrypt -pubin -inkey lab.pub -in block.bin -out t1.bin -pkeyopt rsa_padding_mode:none
$ openssl pkeyutl -encrypt -pubin -inkey lab.pub -in block.bin -out t2.bin -pkeyopt rsa_padding_mode:none
$ cmp -s t1.bin t2.bin && echo "the two ciphertexts are the SAME" || echo "the two ciphertexts DIFFER"
the two ciphertexts are the SAME

The same, twice. That is textbook RSA, and it is the behaviour the program in step 1 demonstrated with small numbers. The two sessions together are the whole argument for padding, and they belong side by side in the journal.

RFC 8017 is the specification that defines both, and it does not hedge: "RSAES-OAEP is REQUIRED to be supported for new applications; RSAES-PKCS1-v1_5 is included only for compatibility with existing applications." The same section says the same of signatures: RSASSA-PSS is required in new applications and RSASSA-PKCS1-v1_5 is there for compatibility, which is the point Practical 14 returns to.

Step 5: what RSA is actually used for

One more measurement, and it settles a question students often have.

$ head -c 200 /dev/zero > big.bin
$ openssl pkeyutl -encrypt -pubin -inkey lab.pub -in big.bin -out big.enc -pkeyopt rsa_padding_mode:oaep -pkeyopt rsa_oaep_md:sha256 2>&1 | sed -E 's/^[0-9A-F]{8,}:/<this run>:/' | tail -2
Public Key operation error
<this run>:error:0200006E:rsa routines:ossl_rsa_padding_add_PKCS1_OAEP_mgf1_ex:data too large for key size:../crypto/rsa/rsa_oaep.c:87:

The long hexadecimal string OpenSSL puts at the front of an error is an identifier for that one run, so it is replaced above with a fixed word; everything after it is the message itself.

A 1024-bit key cannot encrypt 200 bytes. The most it can take with OAEP and SHA-256 is the modulus size, 128 bytes, minus two digests of 32 bytes, minus 2: 62 bytes. A 2048-bit key manages 190. No RSA key encrypts a file.

So RSA is never used to encrypt data. It is used to encrypt a key: a random AES key is generated, the data is encrypted with AES, and only the AES key, a few dozen bytes, is encrypted with RSA. That arrangement is called a hybrid cryptosystem, and it is what TLS does, what PGP does and what every encrypted messenger does.

$ openssl rand -out aes.key 32
$ openssl enc -aes-256-cbc -pbkdf2 -in big.bin -out big.aes -pass file:aes.key
$ openssl pkeyutl -encrypt -pubin -inkey lab.pub -in aes.key -out aes.key.enc -pkeyopt rsa_padding_mode:oaep -pkeyopt rsa_oaep_md:sha256
$ ls -l big.bin big.aes aes.key aes.key.enc | awk '{print $5, $9}'
32 aes.key
128 aes.key.enc
224 big.aes
200 big.bin

The 200-byte file is encrypted by AES, and only the 32-byte key is encrypted by RSA. That is the shape of every real system.

munotes.in134

Practical 12: RSA Encryption and Decryption

Step 6: the security considerations

The key length. NIST SP 800-131A Revision 2 is the document that says what may still be used, and its table is blunt: for signature generation an RSA modulus shorter than 2048 bits is Disallowed and one of 2048 bits or more is Acceptable. SP 800-57 Part 1 Revision 5 is where the strengths those lengths correspond to are set out, and 3072 bits is what it points to where the data must stay secret for a long time. The key in this chapter is 1024 bits because it is a teaching key that exists for two minutes, and the chapter says so rather than leaving a student to copy the number.

p and q must be far apart and truly random. If they are close together, n can be factored by searching near its square root. If the random source is weak, two people can generate keys sharing a prime, and then one gcd between two public moduli recovers both private keys. That has been found in the wild on network devices.

Never use textbook RSA. Steps 2 and 4 above are the demonstration.

e is public and small; d is private and large. Both are fine. What must never leak is d, p, q or phi: any one of them gives the others.

Timing. A naive exponentiation takes measurably longer for some exponents than others, and an attacker who can time your decryptions can recover d. Real libraries blind the input to prevent it; your own implementation does not, which is another reason a hand-written RSA is a teaching exercise and not a product.

And the one that is coming. A sufficiently large quantum computer would factor n efficiently by Shor's algorithm, which ends RSA. That machine does not exist, and replacement algorithms are being standardised; a student should know the sentence and should not repeat any prediction of a date.

Procedure

  1. Write egcd and inverse_mod using the extended Euclidean algorithm.
  2. Pick two small primes, compute n and phi, choose e with gcd(e, phi) = 1, and find d. Check e times d modulo phi is 1 by hand.
  3. Encrypt and decrypt one number with modular exponentiation.
  4. Encrypt a word one letter at a time and note that two identical letters give identical ciphertext.
  5. Show that encrypting with d and decrypting with e also works.
  6. Show that E(a) times E(b) equals E(a times b) modulo n.
  7. Factor your own small n and say how long a real one would take.
  8. Generate a real key with openssl genrsa, look at its exponent, and encrypt and decrypt with OAEP twice, recording that the two ciphertexts differ.
  9. Repeat with rsa_padding_mode:none and record that they do not.
  10. Try to encrypt 200 bytes, record the failure, and build the hybrid arrangement instead.
munotes.in135

Practical 12: RSA Encryption and Decryption

Observations

MeasuredValue
p, q61, 53
n3233
phi(n)3120
e17, gcd with phi is 1
d2753, and 17 times 2753 mod 3120 is 1
m = 65 encrypted2790
2790 decrypted65
Both Ls in HELLO2726, the same ciphertext
Encrypt with d, decrypt with ereturns 65
E(7) times E(9) mod n, and E(63)both 3216
n = 3233 factored53 times 61, instantly
OpenSSL key1024 bit, publicExponent 65537
OAEP, same message encrypted twicetwo different ciphertexts, both decrypt correctly
No padding, same message twicethe same ciphertext
Largest message a 1024-bit key takes with OAEP and SHA-25662 bytes

Result

RSA was implemented from nothing with p = 61 and q = 53, giving n = 3233, phi = 3120, e = 17 and d = 2753, and e times d was confirmed to be 1 modulo phi. A number and a word were encrypted and decrypted correctly, and three properties were demonstrated: the algorithm works with the two exponents exchanged, it is deterministic without padding, so both Ls of HELLO produced 2726, and it is multiplicative, so E(7) times E(9) equals E(63). The same operations were then performed with OpenSSL on a 1024-bit key with public exponent 65537: with OAEP padding the same message gave two different ciphertexts that both decrypted correctly, and with no padding it gave the same ciphertext twice. A 200-byte file could not be encrypted at all, and was instead encrypted with AES under a 32-byte key that RSA encrypted, which is the hybrid arrangement every real system uses.

Where marks are lost

Computing the power in full instead of using modular exponentiation. It never finishes on a real key.

Searching for d in a loop. It works on 3120 and not on a real phi. Use the extended Euclidean algorithm.

Encrypting a message larger than n. Every value must be less than the modulus; a longer message has to be split, or, properly, encrypted with a symmetric cipher under an RSA-wrapped key.

Saying RSA encrypts files. It does not. 62 bytes on a 1024-bit key, and the chapter shows the failure.

Presenting textbook RSA as secure. Say in the same breath that it is deterministic and multiplicative, and that OAEP is what is actually used.

Quoting 1024 bits as a key length without a note. It is below the current minimum. Say the key is a teaching key.

Not destroying p, q and phi. Anyone with any one of them has d.

munotes.in136

Practical 12: RSA Encryption and Decryption

Confusing the public and private exponents. e is small and published; d is large and secret.

For the journal

Aim; the six steps of key generation with your own p and q and every number checked by hand; the encryption and decryption formulas; the program and its output; the three properties, each with the output that shows it; the factorisation of your n; the OpenSSL key with its exponent; the OAEP session showing two different ciphertexts; the no-padding session showing one; the failed 200-byte encryption and the hybrid arrangement with the file sizes; the security considerations in your own words with the key length from NIST; the observation table; the result.

Quick revision

  • Public key (e, n), private key (d, n). n = p times q; phi = (p-1)(q-1); e times d = 1 mod phi.
  • c is m to the power e mod n; m is c to the power d mod n. Use modular exponentiation.
  • d is found with the extended Euclidean algorithm.
  • Security rests on factoring n being hard. Destroy p, q and phi.
  • Textbook RSA is deterministic and multiplicative, so it is never used raw.
  • OAEP adds randomness: the same message gives a different ciphertext every time.
  • e is almost always 65537.
  • A 1024-bit key with OAEP and SHA-256 takes at most 62 bytes; 2048 bits takes 190.
  • So RSA encrypts a symmetric key and the symmetric cipher encrypts the data. That is a hybrid cryptosystem.
  • 2048 bits is the minimum for a new key; SP 800-131A Rev.2 disallows anything shorter.
  • Two keys sharing a prime, from a weak random source, are both broken by one gcd.

Questions you must be able to answer

1. What makes RSA hard to break? That recovering p and q from their product n is beyond any computer we have when the primes are large. Everything else about the key is public.

2. How is d computed? As the multiplicative inverse of e modulo phi(n), found with the extended Euclidean algorithm. Here 17 inverse modulo 3120 is 2753.

3. Why must you use modular exponentiation rather than computing the power first? Because the power itself, for a real key, is a number with hundreds of thousands of digits. Reducing modulo n at every step keeps every intermediate value small.

4. Both Ls in HELLO encrypted to 2726. Why is that a problem? Because textbook RSA is deterministic, so an attacker who guesses a message can encrypt their guess with your public key and compare. Padding removes the determinism.

5. What is the multiplicative property, and what does it allow? That E(a) times E(b) mod n equals E(a times b). An attacker can therefore alter a ciphertext in a predictable way without any key, which is a chosen-ciphertext attack.

munotes.in137

Practical 12: RSA Encryption and Decryption

6. What does OAEP do? It mixes random bytes and a hash into the block before exponentiation, so the same message encrypts differently every time and the structure an attacker would exploit is gone. RFC 8017 recommends it for new applications.

7. How much data can a 2048-bit RSA key encrypt? 190 bytes with OAEP and SHA-256: 256 minus two 32-byte digests minus 2. Not a file.

8. So how is a large file encrypted with RSA? It is not. A random symmetric key encrypts the file and RSA encrypts that key. The combination is a hybrid cryptosystem, and this chapter builds one.

9. What key length should a new RSA key have? At least 2048 bits: SP 800-131A Revision 2 disallows anything shorter for signature generation. 3072 where the secret must last. The 1024-bit key in this chapter is a teaching key and is below the acceptable minimum.

10. Name two ways RSA fails even when the mathematics is right. A weak random source, which can give two keys a shared prime that one gcd recovers; and timing, where the time a decryption takes leaks the private exponent unless the implementation blinds it.

Contents This chapter on its own page

munotes.in138

Chapter Eighteen

Practical 13: Message Authentication Codes

Syllabus topic Module 2, "Message Authentication Codes: Implement algorithms to generate and verify message authentication codes (MACs) for ensuring data integrity and authenticity."

Aim

To implement an algorithm that generates and verifies message authentication codes, and to use it to establish the integrity and the authenticity of a message.

What you need to know before you start

Two words, and the whole practical turns on telling them apart.

Integrity means the message has not been altered. Authenticity means it came from who it claims to come from.

A plain hash gives you neither, on its own, and that surprises people. Anybody can compute SHA-256. An attacker who changes your message simply computes the new hash and sends that too, and the receiver checks it and is satisfied. A hash detects accidents, a corrupted download or a bad disk block. It does not detect an adversary.

What closes the gap is a key. A message authentication code is a short tag computed from the message and a secret key, so that only somebody with the key can produce a tag the receiver will accept.

tag = MAC(key, message)

The sender computes the tag and sends message and tag. The receiver recomputes the tag from the message and the same key, and compares. Verification is just recomputation and comparison. There is nothing else to it, and that is the sentence to write in the journal.

Why not simply hash the key and the message together

The obvious construction is hash(key + message), and it is broken. Not subtly: an attacker can forge a tag for a message they choose, without ever learning the key.

The reason is how SHA-256 and its relatives are built. They are Merkle-Damgard hashes: they run through the message a block at a time, carrying a state, and the final state is the digest. So the digest is the internal state at the end. An attacker who holds hash(key + message) holds that state, and can carry on hashing from it, appending whatever they like, and the result is a valid hash(key + message + padding + extra). That is the length extension attack.

HMAC is the construction that closes it, and it is defined in RFC 2104 as:

HMAC(K, m) = H( (K xor opad) || H( (K xor ipad) || m ) )

ipad is the byte 0x36 repeated to the block length, opad is 0x5C repeated. The message is hashed under one key-derived pad, and that result is hashed again under another. The attacker never holds the state of the outer hash, so there is nothing to extend.

Step 1: build it, and check it against the RFC

"""Practical 13: message authentication codes, and HMAC built from RFC 2104."""
import hashlib
import hmac as stdlib_hmac

def hmac_by_hand(key, message, hashfn=hashlib.sha256, blocksize=64):
    """RFC 2104 section 2, written out.

        H(K xor opad, H(K xor ipad, text))

    ipad is the byte 0x36 repeated, opad is 0x5C repeated, and a key longer
    than one block is hashed first, then padded with zeros to a whole block.
    """
    if len(key) > blocksize:
        key = hashfn(key).digest()
    key = key + b"\x00" * (blocksize - len(key))
    ipad = bytes(k ^ 0x36 for k in key)
    opad = bytes(k ^ 0x5C for k in key)
    inner = hashfn(ipad + message).digest()
    return hashfn(opad + inner).hexdigest()

MESSAGE = b"Transfer 5000 to account 74215"
KEY = b"the shared secret"

print("A HASH IS NOT A MAC")
print("  sha256 of the message :", hashlib.sha256(MESSAGE).hexdigest())
print("  anybody can compute that, so anybody can change the message and")
print("  compute a new one. A hash proves the message is unaltered only if")
print("  the hash itself arrives by some channel nobody can touch.")
print()
print("A MAC NEEDS A KEY")
mine = hmac_by_hand(KEY, MESSAGE)
theirs = stdlib_hmac.new(KEY, MESSAGE, hashlib.sha256).hexdigest()
print("  our HMAC-SHA256      :", mine)
print("  Python's hmac module :", theirs)
print("  they agree           :", mine == theirs)
print()
print("RFC 2104's own test case 2")
print("  key     : 'Jefe'")
print("  message : 'what do ya want for nothing?'")
t = hmac_by_hand(b"Jefe", b"what do ya want for nothing?", hashlib.md5, 64)
print("  HMAC-MD5:", t)
print("  RFC 2104 prints: 750c783e6ab0b503eaa86e310a5db738")
print("  match   :", t == "750c783e6ab0b503eaa86e310a5db738")
print()
print("VERIFICATION: one bit changed anywhere and the tag is unrecognisable")
tampered = b"Transfer 9000 to account 74215"
print("  original tag :", hmac_by_hand(KEY, MESSAGE)[:32], "...")
print("  tampered tag :", hmac_by_hand(KEY, tampered)[:32], "...")
a = bytes.fromhex(hmac_by_hand(KEY, MESSAGE))
b = bytes.fromhex(hmac_by_hand(KEY, tampered))
bits = sum(bin(x ^ y).count("1") for x, y in zip(a, b))
print("  bits that differ between the two tags: %d of %d" % (bits, 8 * len(a)))
print()
print("  verify(message, tag) is just: recompute and compare")
for msg, tag, label in ((MESSAGE, mine, "the real message"),
                        (tampered, mine, "the tampered message")):
    ok = stdlib_hmac.compare_digest(hmac_by_hand(KEY, msg), tag)
    print("    %-22s accepted: %s" % (label, ok))
print()
print("WHY HMAC IS SHAPED LIKE THAT, and not just hash(key + message):")
print("  SHA-256 is a Merkle-Damgard hash, and such a hash CONTINUES.")
print("  Given hash(secret + message) and the length of the secret, an")
print("  attacker can compute hash(secret + message + anything) without")
print("  ever learning the secret. That is the length extension attack, and")
print("  it forges a tag for a message the attacker chose.")
print("  HMAC hashes twice, with two different key-derived pads, so the")
print("  attacker never holds the internal state of the outer hash.")
print()
print("  compare_digest, not ==, and the reason is TIME:")
print("  == stops at the first byte that differs, so how long it takes")
print("  reveals how many leading bytes the attacker got right, and a tag")
print("  can be guessed one byte at a time. compare_digest always looks at")
print("  every byte.")
munotes.in139

Practical 13: Message Authentication Codes

A HASH IS NOT A MAC
  sha256 of the message : 4b5ea53b60a4dc51c8f57262a55918354585df4671ce95b6ba88bbb7d53bb09c
  anybody can compute that, so anybody can change the message and
  compute a new one. A hash proves the message is unaltered only if
  the hash itself arrives by some channel nobody can touch.

A MAC NEEDS A KEY
  our HMAC-SHA256      : abc685f5cc1057dacb40503fd09ca3291bdf30302421109765cbbd3bda40bdf6
  Python's hmac module : abc685f5cc1057dacb40503fd09ca3291bdf30302421109765cbbd3bda40bdf6
  they agree           : True

RFC 2104's own test case 2
  key     : 'Jefe'
  message : 'what do ya want for nothing?'
  HMAC-MD5: 750c783e6ab0b503eaa86e310a5db738
  RFC 2104 prints: 750c783e6ab0b503eaa86e310a5db738
  match   : True

VERIFICATION: one bit changed anywhere and the tag is unrecognisable
  original tag : abc685f5cc1057dacb40503fd09ca329 ...
  tampered tag : 975a2cd95f8dafb32e0943d24be47218 ...
  bits that differ between the two tags: 139 of 256

  verify(message, tag) is just: recompute and compare
    the real message       accepted: True
    the tampered message   accepted: False

WHY HMAC IS SHAPED LIKE THAT, and not just hash(key + message):
  SHA-256 is a Merkle-Damgard hash, and such a hash CONTINUES.
  Given hash(secret + message) and the length of the secret, an
  attacker can compute hash(secret + message + anything) without
  ever learning the secret. That is the length extension attack, and
  it forges a tag for a message the attacker chose.
  HMAC hashes twice, with two different key-derived pads, so the
  attacker never holds the internal state of the outer hash.

  compare_digest, not ==, and the reason is TIME:
  == stops at the first byte that differs, so how long it takes
  reveals how many leading bytes the attacker got right, and a tag
  can be guessed one byte at a time. compare_digest always looks at
  every byte.
munotes.in140

Practical 13: Message Authentication Codes

Three things in that output carry the marks.

Our implementation is right, and there is evidence for it twice over. It agrees with Python's own hmac module, and it reproduces the test vector printed inside RFC 2104 itself. A test vector published by the people who defined the algorithm is the strongest check available, and it is the reason this chapter computes one HMAC with MD5: RFC 2104 predates SHA-256 and its vectors are MD5 ones. That single MD5 is a check against the standard and nothing else, and MD5 is not to be used for a real MAC.

Changing one character of the message changed 139 of the 256 bits of the tag. Not 139 by design: about half of them, which is what a good hash does, and which is why a forged message cannot have a tag that is nearly right.

Verification returned True for the real message and False for the tampered one, using the same key and the same code. That is the whole of "generate and verify" in MU's sentence.

munotes.in141

Practical 13: Message Authentication Codes

Step 2: the same thing at the command line

$ echo -n "Transfer 5000 to account 74215" > msg.txt
$ echo -n "the shared secret" > key.txt
$ openssl dgst -sha256 -hmac "the shared secret" msg.txt
HMAC-SHA2-256(msg.txt)= abc685f5cc1057dacb40503fd09ca3291bdf30302421109765cbbd3bda40bdf6
$ openssl dgst -sha256 -mac HMAC -macopt hexkey:$(xxd -p -c 256 key.txt) msg.txt
HMAC-SHA2-256(msg.txt)= abc685f5cc1057dacb40503fd09ca3291bdf30302421109765cbbd3bda40bdf6
$ echo -n "Transfer 9000 to account 74215" > bad.txt
$ openssl dgst -sha256 -hmac "the shared secret" bad.txt
HMAC-SHA2-256(bad.txt)= 975a2cd95f8dafb32e0943d24be47218ac06fa97405dade04b75baa4e17aad06

The first two lines are the same tag by two spellings of the same option, and it is the tag the Python program printed: abc685f5cc.... Three independent implementations, one answer.

The last line is the forged message under the right key, and it is the tag the Python program printed for the tampered message. An attacker without the key cannot produce either.

Step 3: what a MAC does not give you

$ openssl dgst -sha256 -hmac "the shared secret" msg.txt | cut -c1-40
HMAC-SHA2-256(msg.txt)= abc685f5cc1057da
$ cat msg.txt; echo
Transfer 5000 to account 74215

The message is still in the clear. A MAC authenticates; it does not encrypt. If the message must also be secret, encrypt it as well, and the safe order is encrypt then MAC: encrypt the message, then compute the tag over the ciphertext. Doing it the other way round means the receiver has to decrypt before it can check the tag, which means it processes attacker-controlled data before it has any reason to trust it, and several real protocols have been broken exactly there.

In practice nobody assembles the two by hand any more. Modern ciphers do both at once, and are called authenticated encryption: AES-GCM and ChaCha20-Poly1305 are the two you will meet.

Step 4: the limit that decides when a MAC is the wrong tool

Alice and Bob share one key. Alice sends a message with a tag. Bob checks it, and knows the message came from somebody holding the key, which is Alice or Bob.

Now suppose they disagree, and Bob says Alice sent an instruction to transfer money. Alice says Bob made it up. Nobody can tell. Bob has the same key and could have produced that tag himself, so the tag proves nothing against Alice to anyone but Bob.

That property has a name: a MAC gives no non-repudiation. Where it matters, a contract, an invoice, a signed software release, a MAC is the wrong tool and a digital signature is the right one. That is the next chapter, and the comparison belongs in both journals:

MACDigital signature
Keysone, sharedtwo, a private and a public
Who can make a tageither partyonly the holder of the private key
Who can check iteither partyanybody, with the public key
Proves to a third partynothingwho signed
Speedfast, a hashslow, public-key arithmetic
Tag size32 bytes for SHA-256256 bytes for RSA-2048
munotes.in142

Practical 13: Message Authentication Codes

Procedure

  1. Compute the SHA-256 of a message and say in one sentence why it is not a MAC.
  2. Implement HMAC from RFC 2104: pad or hash the key to the block length, exclusive-or with 0x36 and 0x5C, hash twice.
  3. Compare your tag with the language's own HMAC implementation.
  4. Compute one of RFC 2104's published test vectors and compare it with the RFC.
  5. Change one character of the message, compute the tag again, and count the bits that differ.
  6. Write verify, and run it on the real message and on the tampered one.
  7. Compute the same tag with openssl dgst -hmac and confirm it matches.
  8. Write out why HMAC hashes twice, and why the comparison must be constant time.

Observations

MeasuredValue
SHA-256 of the message4b5ea53b60a4dc51 and onward
HMAC-SHA256, our codeabc685f5cc1057da and onward
Python's hmac modulethe same, exactly
openssl dgst -sha256 -hmacthe same, exactly
RFC 2104 case 2, HMAC-MD5750c783e6ab0b503 and onward
Tag of the tampered message975a2cd95f8dafb3 and onward
Bits differing, the two tags139 of 256
Verification, real messageaccepted
Verification, tampered messagerejected
Tag length, HMAC-SHA25632 bytes, whatever the message

Result

HMAC was implemented from the definition in RFC 2104 and used to generate and verify message authentication codes. The implementation was checked three ways: against Python's own hmac module, against openssl dgst -sha256 -hmac, and against the test vector RFC 2104 itself prints, which it reproduced exactly. Altering one character of the message changed 139 of the 256 bits of the tag, and verification accepted the original message and rejected the altered one under the same key. The reason HMAC hashes twice was demonstrated by explaining the length extension attack on the naive construction, and the reason verification must compare in constant time was recorded. Finally the limit of a MAC was established: it gives integrity and authenticity between two holders of one key, and no non-repudiation against either of them.

Where marks are lost

Presenting a bare hash as a MAC. Without a key it detects accidents, not attackers.

Using hash(key + message). Length extension forges a tag for a message the attacker chose. Use HMAC.

Comparing tags with ==. The comparison time leaks how many leading bytes were right, and a tag can be guessed a byte at a time. Use a constant-time comparison.

Reusing one key for everything. A key for MACs is a key for MACs. Do not also encrypt with it.

munotes.in143

Practical 13: Message Authentication Codes

MAC then encrypt. Encrypt then MAC, or better, use an authenticated cipher.

Saying a MAC proves who sent the message. It proves one of the key holders did, and there are two of them.

Truncating a tag without saying so. A shortened tag is weaker in a way that is easy to state and easy to forget.

Using MD5 for a real MAC. The MD5 in this chapter exists only to reproduce the RFC's own vector.

For the journal

Aim; integrity and authenticity defined and told apart; why a bare hash is neither; the HMAC formula from RFC 2104 with ipad and opad named; the program; its output; the three-way agreement with the language's module, with OpenSSL and with the RFC's own vector; the one-character change and the count of differing bits; the verification of the good and the bad message; the length extension explanation; the constant-time note; the MAC against signature table; the observation table; the result.

Quick revision

  • Integrity: unaltered. Authenticity: from who it claims. A MAC gives both; a bare hash gives neither against an attacker.
  • tag = MAC(key, message). Verify by recomputing and comparing.
  • hash(key + message) is broken by length extension, because a Merkle-Damgard digest is the internal state.
  • HMAC(K, m) is H((K xor opad) then H((K xor ipad) then m)). ipad is 0x36 repeated, opad is 0x5C.
  • A key longer than the block is hashed first; a shorter one is padded with zeros.
  • HMAC-SHA256 gives a 32-byte tag whatever the message length.
  • Compare tags in constant time, never with ==.
  • Encrypt then MAC, or use an authenticated cipher such as AES-GCM.
  • A MAC gives no non-repudiation: both parties hold the key, so neither can prove anything against the other.
  • openssl dgst -sha256 -hmac "key" file computes one at the command line.

Questions you must be able to answer

1. What is the difference between integrity and authenticity? Integrity is that the message has not been altered. Authenticity is that it came from who it claims to come from. A MAC gives both at once.

2. Why is a SHA-256 digest not a message authentication code? Because anyone can compute it. An attacker who alters the message computes the new digest and sends that too, and the receiver is satisfied. A hash detects accidents, not adversaries.

3. What is wrong with hash(key + message)? A Merkle-Damgard digest is the hash's internal state at the end of the message, so an attacker holding it can continue hashing and produce a valid tag for the message plus anything they append, without knowing the key. That is the length extension attack.

4. Write out the HMAC construction. H of ((key xor opad) followed by H of ((key xor ipad) followed by the message)), where ipad is the byte 0x36 repeated to the block length and opad is 0x5C repeated.

munotes.in144

Practical 13: Message Authentication Codes

5. How do you verify a MAC? Recompute the tag from the received message and the shared key, and compare it with the received tag in constant time. There is no separate verification algorithm.

6. Why must the comparison be constant time? Because a comparison that stops at the first difference takes longer the more leading bytes are correct, so an attacker who can time it can find a valid tag one byte at a time.

7. How did you show your implementation is correct? Three ways: it agrees with Python's own hmac module, it agrees with openssl dgst -sha256 -hmac, and it reproduces the test vector printed in RFC 2104 itself.

8. Does a MAC keep the message secret? No. It authenticates; the message travels in the clear. To have both, encrypt and then MAC the ciphertext, or use an authenticated cipher such as AES-GCM.

9. Your bank has a MAC on a transfer instruction from you and you deny sending it. Can the bank prove you did? No. The bank holds the same key and could have produced the tag itself, so the tag proves nothing to a third party. That is the absence of non-repudiation, and a digital signature is what supplies it.

10. You changed one character and 139 of 256 bits of the tag changed. Is that the right number? It is the expected kind of number: about half. A good hash changes each output bit with probability one half for any input change, so anything near 128 is what you should see, and a small number would mean something is wrong.

Contents This chapter on its own page

munotes.in145

Chapter Nineteen

Practical 14: Digital Signatures

Syllabus topic Module 2, "Digital Signatures: Implement digital signature algorithms such as RSA-based signatures, and verify the integrity and authenticity of digitally signed messages."

Aim

To implement an RSA-based digital signature, to verify the integrity and the authenticity of a signed message, and to show what a signature gives that a message authentication code does not.

What you need to know before you start

Practical 13 ended on a limit: a MAC uses one shared key, so a tag proves only that one of the two key holders made it, and neither can prove anything against the other.

A signature removes that limit by using the key pair of Practical 12 the other way round.

  • Signing uses the private key, which only the signer has.
  • Verifying uses the public key, which everybody has.

So anybody can check a signature and nobody else can make one. That is non-repudiation: the signer cannot later deny it, because nobody else could have produced it.

Signing is not encrypting backwards, and getting that wrong costs marks. The arithmetic is the same operation with the other exponent, and that is where the confusion comes from, but what it achieves is the opposite:

EncryptionSignature
Key used to make itthe recipient's public keythe signer's private key
Key used to undo itthe recipient's private keythe signer's public key
Who can read the messageonly the recipientanybody
What it givessecrecyauthenticity and non-repudiation

A signature provides no secrecy at all. The message travels in the clear beside it.

Why a hash is signed, and not the message

RSA can only act on a number smaller than the modulus. A signature therefore signs the hash of the message, which is a fixed 256 bits however large the message is.

signature = H(message)^d mod n

verify: signature^e mod n = H(message) ?

Three reasons, and an examiner may ask for any of them:

  1. Size. A 10 MB file does not fit in a 2048-bit number.
  2. Speed. Public-key arithmetic is slow; hashing is fast. Sign 32 bytes, not ten million.
  3. Safety. Splitting a long message into blocks and signing each would let an attacker reorder, drop or reuse the blocks. One hash over the whole message cannot be rearranged.

Step 1: signing and verifying, from nothing

"""Practical 14: RSA digital signatures, and why a signature is not encryption."""
import hashlib

def egcd(a, b):
    if b == 0:
        return a, 1, 0
    g, x, y = egcd(b, a % b)
    return g, y, x - (a // b) * y

def inverse_mod(a, m):
    g, x, _ = egcd(a % m, m)
    if g != 1:
        raise ValueError("no inverse")
    return x % m

def is_prime(n, rounds=40):
    """Miller-Rabin. A key made from numbers that are not prime does not work,
    and it fails SILENTLY: signing and verifying simply disagree."""
    import random
    if n < 2:
        return False
    for small in (2, 3, 5, 7, 11, 13, 17, 19, 23, 29, 31, 37):
        if n % small == 0:
            return n == small
    d, r = n - 1, 0
    while d % 2 == 0:
        d //= 2
        r += 1
    rng = random.Random(0)
    for _ in range(rounds):
        a = rng.randrange(2, n - 1)
        x = pow(a, d, n)
        if x in (1, n - 1):
            continue
        for _ in range(r - 1):
            x = x * x % n
            if x == n - 1:
                break
        else:
            return False
    return True

# Two real 256-bit primes, so n is 512 bits and a SHA-256 digest fits inside it.
p = 0xE1A9B1C3D5E7F9012B4D6F8193A5B7C9DAF3E5D7C9BBAD9F8171635547392BCD
q = 0xC3B5A7998B7D6F615345372918FBEDDFC1B39587796B5D4F41332517F9EBDF7F
assert is_prime(p) and is_prime(q), "p and q must actually be prime"
n = p * q
phi = (p - 1) * (q - 1)
e = 65537
d = inverse_mod(e, phi)

MESSAGE = b"Pay Rahul 5000 rupees on 30 September 2026"

def sign(message):
    h = int.from_bytes(hashlib.sha256(message).digest(), "big")
    return pow(h, d, n)

def verify(message, signature):
    h = int.from_bytes(hashlib.sha256(message).digest(), "big")
    return pow(signature, e, n) == h

print("the key")
print("  p is prime :", is_prime(p), " q is prime :", is_prime(q))
print("  n has %d bits" % n.bit_length())
print("  e =", e)
print()
print("SIGNING")
print("  message :", MESSAGE.decode())
digest = hashlib.sha256(MESSAGE).hexdigest()
print("  sha256  :", digest)
sig = sign(MESSAGE)
print("  signature = digest^d mod n")
print("  signature :", hex(sig)[:66] + "...")
print("  it is %d bits, always about the size of n (%d bits), whatever the\n  message length" % (sig.bit_length(), n.bit_length()))
print()
print("VERIFYING")
print("  anybody with the PUBLIC key computes signature^e mod n and compares")
print("  it with their own hash of the message.")
print("  original message accepted :", verify(MESSAGE, sig))
TAMPERED = b"Pay Rahul 9000 rupees on 30 September 2026"
print("  tampered message accepted :", verify(TAMPERED, sig))
print("  wrong signature accepted  :", verify(MESSAGE, sig + 1))
print()
print("WHY THE HASH")
print("  RSA can only act on a number smaller than n. A 41-byte message is")
print("  328 bits and would fit here, but a 10 MB file would not, and")
print("  splitting it into blocks would let an attacker reorder them.")
print("  Hashing first gives a fixed 256 bits whatever the message, and any")
print("  change to the message changes it completely.")
print()
print("SIGNING IS NOT ENCRYPTING BACKWARDS")
print("  The arithmetic is the same operation with the other exponent, but")
print("  what it achieves is the opposite:")
print("    encryption with the PUBLIC key   -> only the key holder can read it")
print("    signing   with the PRIVATE key   -> anybody can read it, and")
print("                                        nobody else could have made it")
print("  A signature gives no secrecy at all. The message travels in clear.")
print()
print("WHAT A SIGNATURE GIVES THAT A MAC DOES NOT")
print("  A MAC uses one shared key, so either party could have produced the")
print("  tag. It proves the message came from one of the two of you.")
print("  A signature uses a key only the signer has, so the signer cannot")
print("  later deny it. That is NON-REPUDIATION, and it is why a signature")
print("  and not a MAC is what a contract needs.")
munotes.in146

Practical 14: Digital Signatures

the key
  p is prime : True  q is prime : True
  n has 512 bits
  e = 65537

SIGNING
  message : Pay Rahul 5000 rupees on 30 September 2026
  sha256  : 978db1b9d935eee22d341fa6c30513443105d5d68c2137fd8856ea4393b8f4f0
  signature = digest^d mod n
  signature : 0x4885d33113da11f8a4649e0bbd8a947727902be9d493951f867482e899b2994b...
  it is 511 bits, always about the size of n (512 bits), whatever the
  message length

VERIFYING
  anybody with the PUBLIC key computes signature^e mod n and compares
  it with their own hash of the message.
  original message accepted : True
  tampered message accepted : False
  wrong signature accepted  : False

WHY THE HASH
  RSA can only act on a number smaller than n. A 41-byte message is
  328 bits and would fit here, but a 10 MB file would not, and
  splitting it into blocks would let an attacker reorder them.
  Hashing first gives a fixed 256 bits whatever the message, and any
  change to the message changes it completely.

SIGNING IS NOT ENCRYPTING BACKWARDS
  The arithmetic is the same operation with the other exponent, but
  what it achieves is the opposite:
    encryption with the PUBLIC key   -> only the key holder can read it
    signing   with the PRIVATE key   -> anybody can read it, and
                                        nobody else could have made it
  A signature gives no secrecy at all. The message travels in clear.

WHAT A SIGNATURE GIVES THAT A MAC DOES NOT
  A MAC uses one shared key, so either party could have produced the
  tag. It proves the message came from one of the two of you.
  A signature uses a key only the signer has, so the signer cannot
  later deny it. That is NON-REPUDIATION, and it is why a signature
  and not a MAC is what a contract needs.
munotes.in147

Practical 14: Digital Signatures

Read the VERIFYING block. The real message is accepted, the tampered message is rejected, and a signature off by one is rejected. Three lines, and between them they are the whole of MU's "verify the integrity and authenticity of digitally signed messages".

Note what verification actually does: it computes signature^e mod n with the public key and compares the result with its own hash of the message it received. If they match, the signature was made by whoever holds d, over exactly that message.

The signature is the size of the key, not the size of the message. 511 bits here for a 512-bit modulus, and it would be the same for a one-line message or a gigabyte file, because in both cases it is a 256-bit digest that was signed.

munotes.in148

Practical 14: Digital Signatures

The Miller-Rabin test is in the program on purpose. A first version of it used two hexadecimal numbers that were not in fact prime. Everything still ran: a key was made, a signature was produced, and verification silently returned False for a message that had not been touched. A key built from numbers that are not prime fails without an error message, so the program asserts that they are prime before using them.

Step 2: the same thing with OpenSSL

$ openssl genrsa -out sign.pem 2048 2>/dev/null
$ openssl rsa -in sign.pem -pubout -out sign.pub 2>/dev/null
$ echo -n "Pay Rahul 5000 rupees on 30 September 2026" > contract.txt
$ openssl dgst -sha256 -sign sign.pem -out contract.sig contract.txt
$ wc -c < contract.sig
256
$ openssl dgst -sha256 -verify sign.pub -signature contract.sig contract.txt
Verified OK

256 bytes, which is 2048 bits, the size of the key. And Verified OK, which is what -verify prints when the check passes.

Now tamper with the message and check again.

$ echo -n "Pay Rahul 9000 rupees on 30 September 2026" > forged.txt
$ openssl dgst -sha256 -verify sign.pub -signature contract.sig forged.txt 2>&1 | sed -E 's/^[0-9A-F]{8,}:/<this run>:/'
<this run>:error:02000068:rsa routines:ossl_rsa_verify:bad signature:../crypto/rsa/rsa_sign.c:430:
<this run>:error:1C880004:Provider routines:rsa_verify:RSA lib:../providers/implementations/signature/rsa_sig.c:774:
Verification failure
$ ( openssl dgst -sha256 -verify sign.pub -signature contract.sig forged.txt >/dev/null 2>&1 ); echo "openssl exit status $?"
openssl exit status 1

The long hexadecimal string at the front of each error line identifies that one run, so it is replaced above with a fixed word.

Verification failure, and a non-zero exit status. The exit status matters: a script that checks a signature must test it, because the message on the screen is not something a program can act on.

And one more, which shows the signature is tied to the key as well as to the message.

$ openssl genrsa -out other.pem 2048 2>/dev/null
$ openssl rsa -in other.pem -pubout -out other.pub 2>/dev/null
$ openssl dgst -sha256 -verify other.pub -signature contract.sig contract.txt 2>&1 | sed -E 's/^[0-9A-F]{8,}:/<this run>:/'
<this run>:error:0200008A:rsa routines:RSA_padding_check_PKCS1_type_1:invalid padding:../crypto/rsa/rsa_pk1.c:79:
<this run>:error:02000072:rsa routines:rsa_ossl_public_decrypt:padding check failed:../crypto/rsa/rsa_ossl.c:697:
<this run>:error:1C880004:Provider routines:rsa_verify:RSA lib:../providers/implementations/signature/rsa_sig.c:774:
Verification failure
$ ( openssl dgst -sha256 -verify other.pub -signature contract.sig contract.txt >/dev/null 2>&1 ); echo "openssl exit status $?"
openssl exit status 1

The right message, the right signature, the wrong public key, and it fails. A signature is a statement about one message by one key.

munotes.in149

Practical 14: Digital Signatures

Step 3: what the signature actually contains

$ openssl pkeyutl -verifyrecover -pubin -inkey sign.pub -in contract.sig -out recovered.bin
$ xxd recovered.bin | tail -3
00000010: 0004 2097 8db1 b9d9 35ee e22d 341f a6c3  .. .....5..-4...
00000020: 0513 4431 05d5 d68c 2137 fd88 56ea 4393  ..D1....!7..V.C.
00000030: b8f4 f0                                  ...
$ openssl dgst -sha256 contract.txt
SHA2-256(contract.txt)= 978db1b9d935eee22d341fa6c30513443105d5d68c2137fd8856ea4393b8f4f0

The last 32 bytes of the recovered block are the SHA-256 of the contract, and the bytes before them are PKCS #1 v1.5 padding: a leading 00 01, a run of FF bytes, a 00, and a short ASN.1 header naming which hash was used. That header is not decoration: without it, a signature made over a SHA-256 digest could be presented as one made over some other digest.

This is the one place in the book where the inside of a signature is visible, and it is worth a paragraph in the journal.

Step 4: the standing of the schemes

FIPS 186-5, the Digital Signature Standard of February 2023, is the document that says which schemes may be used, and the points a student should be able to state are these.

RSA signatures are approved, in the two forms RFC 8017 defines: the older PKCS #1 v1.5, which is what openssl dgst -sign produces by default and what step 3 took apart, and RSASSA-PSS, which adds randomness and is what new systems should use.

ECDSA and EdDSA are approved, and both give much shorter signatures for the same strength: 64 bytes against RSA-2048's 256.

Plain DSA was withdrawn for signing. FIPS 186-5 keeps it only so that old signatures can still be verified. A student who names DSA as a current choice is out of date, and this is a change recent enough that older textbooks still get it wrong.

Step 5: the comparison MU's two practicals are building towards

MAC, Practical 13Signature, Practical 14
Keysone, shareda private and a public
Who can produce iteither partyonly the private key holder
Who can check iteither partyanybody
Proves to a courtnothingwho signed
Non-repudiationnoyes
Speedfastslow
Size, typical32 bytes256 bytes, or 64 for ECDSA
Use it fora session between two partiesa contract or a certificate

Both practicals end at the same place: integrity and authenticity, by two different routes, and only one of them convinces a third party.

Procedure

  1. Build an RSA key from two primes, and test the primes with Miller-Rabin before using them.
  2. Hash the message with SHA-256 and convert the digest to an integer.
  3. Sign by raising the digest to the private exponent modulo n.
  4. Verify by raising the signature to the public exponent and comparing with your own hash.
  5. Verify the original message, a tampered message and a corrupted signature, and record all three results.
  6. Note the signature length against the message length.
  7. Repeat with openssl dgst -sign and -verify on a 2048-bit key, and record the exit status of a failed verification.
  8. Verify a good signature against a different public key and record that it fails.
  9. Recover the signed block with -verifyrecover and find the digest inside it.
munotes.in150

Practical 14: Digital Signatures

Observations

MeasuredValue
Key512-bit modulus from two 256-bit primes, both proved prime
Public exponent65537
MessagePay Rahul 5000 rupees on 30 September 2026
SHA-256 of it978db1b9d935eee2 and onward
Signature length, our key511 bits, always about the modulus size
Original message verifiedaccepted
Tampered message verifiedrejected
Signature plus one verifiedrejected
OpenSSL key2048-bit
OpenSSL signature length256 bytes
OpenSSL verify, correct messageVerified OK
OpenSSL verify, altered messageVerification failure, non-zero exit status
OpenSSL verify, wrong public keyfails
Inside the recovered blockPKCS #1 v1.5 padding, a hash identifier, and the 32-byte SHA-256

Result

An RSA digital signature scheme was implemented from nothing over SHA-256. A 512-bit key was built from two numbers proved prime by Miller-Rabin, a contract was signed by raising its digest to the private exponent, and verification by the public exponent accepted the original message and rejected both an altered message and an altered signature. The same operations were performed with OpenSSL on a 2048-bit key: the signature was 256 bytes, verification of the correct message printed Verified OK, verification of an altered message printed Verification failure with a non-zero exit status, and a correct signature checked against a different public key also failed. The signed block was recovered with -verifyrecover and shown to contain PKCS #1 v1.5 padding, a hash identifier and the SHA-256 digest of the contract.

Where marks are lost

Saying a signature encrypts the message. It does not. The message is in the clear and a signature adds no secrecy.

Signing the message instead of its hash. It does not fit, it is slow, and splitting it into blocks lets an attacker reorder them.

Using a key whose factors are not prime. Nothing raises an error; verification simply disagrees with signing. Test the primes.

Not checking the exit status of a verification. The words on the screen are not something a script can act on.

Claiming a MAC gives non-repudiation. It does not. Both parties hold the key.

Naming DSA as a current scheme. FIPS 186-5 withdrew it for signing and keeps it only for verifying old signatures.

Reporting only a successful verification. A verification that never fails proves nothing. Show the tampered message being rejected.

Reusing the signing key for encryption. Use separate keys for separate purposes; a signing oracle and a decryption oracle on one key interact badly.

munotes.in151

Practical 14: Digital Signatures

For the journal

Aim; signing against encrypting, as a table; why the hash is signed and not the message, with all three reasons; the program with its key, its primality test and its output; the three verification results; the signature length against the message length; the OpenSSL session with the signature size and Verified OK; the failed verification with its exit status; the verification against a different key; the recovered block with the digest visible inside it; the standing of the schemes from FIPS 186-5; the MAC against signature table; the observation table; the result.

Quick revision

  • Sign with the private key, verify with the public key. The opposite of encryption.
  • A signature gives authenticity, integrity and non-repudiation, and no secrecy.
  • Sign the hash, not the message: size, speed and safety against reordering.
  • signature = H(m)^d mod n; verification checks signature^e mod n against H(m).
  • The signature is the size of the key, whatever the message length: 256 bytes for RSA-2048.
  • openssl dgst -sha256 -sign key.pem -out m.sig m.txt and -verify pub.pem -signature m.sig m.txt.
  • A failed verification prints Verification failure and exits non-zero. Test the status.
  • Inside a PKCS #1 v1.5 signature: 00 01, FF padding, 00, a hash identifier, then the digest.
  • FIPS 186-5: RSA and ECDSA and EdDSA are approved; plain DSA is withdrawn for signing.
  • RSASSA-PSS is the RSA form to prefer for new systems.
  • A MAC and a signature both give integrity and authenticity; only the signature convinces a third party.

Questions you must be able to answer

1. Which key signs and which key verifies? The private key signs, the public key verifies. Anybody can check a signature and only the key holder can make one.

2. How is signing different from encryption, given that the arithmetic is the same? Encryption uses the recipient's public key and gives secrecy to one reader. Signing uses the signer's private key and gives authenticity to every reader. A signature keeps nothing secret.

3. Why is the hash signed rather than the message? Because RSA can only act on a number smaller than the modulus, because public-key arithmetic is slow, and because signing a long message in blocks would let an attacker reorder, drop or reuse them.

4. How long is your signature, and what does it depend on? The size of the key, not the message: 256 bytes for a 2048-bit key whether the message is one line or a gigabyte.

5. What is non-repudiation, and why does a MAC not give it? That the signer cannot later deny signing, because nobody else could have produced the signature. A MAC uses one shared key, so either party could have produced the tag and neither can prove anything against the other.

munotes.in152

Practical 14: Digital Signatures

6. Your program verified a signature as False on a message nobody had touched. What would you check first? Whether p and q are actually prime. A key built from composite numbers produces a d that is not the inverse of e modulo the true phi, and everything runs without an error while signing and verifying disagree.

7. What is inside a PKCS #1 v1.5 signature block? A leading 00 01, a run of FF bytes as padding, a 00 separator, an ASN.1 header identifying the hash algorithm, and then the digest itself. The hash identifier stops a signature over one digest being presented as one over another.

8. Why check the exit status of openssl dgst -verify? Because a script cannot act on the words printed to the screen, and a verification whose result is ignored is not a verification.

9. Is DSA a reasonable choice for a new system? No. FIPS 186-5 withdrew plain DSA for generating signatures and keeps it only for verifying old ones. RSA with PSS, ECDSA or EdDSA are the current choices.

10. A signature verifies against the message and the public key. What has it proved? That whoever holds the matching private key signed exactly that message, and that the message has not changed since. It has proved nothing about secrecy, and nothing about when it was signed.

Contents This chapter on its own page

munotes.in153

Chapter Twenty

Practical 15: Key Exchange Using Diffie-Hellman

Syllabus topic Module 2, "Key Exchange using Diffie-Hellman: Implement the Diffie-Hellman key exchange algorithm to securely exchange keys between two entities over an insecure network."

Aim

To implement the Diffie-Hellman key exchange, to use it to agree a key over a network anybody can listen to, and to demonstrate the attack it does not prevent.

What you need to know before you start

Every symmetric cipher needs both sides to hold the same key. Getting that key to the other side is the problem, and before 1976 the only answer was to carry it there.

Diffie-Hellman solves it with a question that is easy one way and hard the other. Given g, a, and a prime p, computing g^a mod p is fast. Given g, p and the answer, finding a is the discrete logarithm problem, and for a large p nobody knows how.

The exchange is four lines:

  1. Both agree, in public, on a prime p and a generator g.
  2. Alice picks a secret a and sends A = g^a mod p. Bob picks a secret b and sends B = g^b mod p.
  3. Alice computes B^a mod p. Bob computes A^b mod p.
  4. Both now hold g^(a*b) mod p, and they never sent it.

It works because (g^b)^a and (g^a)^b are the same number. Nothing secret ever crosses the network, and that is the whole idea.

The program

"""Practical 15: the Diffie-Hellman key exchange, and the attack on it."""
import hashlib

print("PART 1: small numbers, so every step can be checked")
p, g = 23, 5
print("  public, agreed in the open: prime p = %d, generator g = %d" % (p, g))
a, b = 6, 15
print()
print("  %-38s %s" % ("Alice", "Bob"))
print("  %-38s %s" % ("picks a secret a = %d" % a, "picks a secret b = %d" % b))
A = pow(g, a, p)
B = pow(g, b, p)
print("  %-38s %s" % ("A = g^a mod p = %d^%d mod %d = %d" % (g, a, p, A),
                      "B = g^b mod p = %d^%d mod %d = %d" % (g, b, p, B)))
print("  %-38s %s" % ("sends A = %d" % A, "sends B = %d" % B))
sA = pow(B, a, p)
sB = pow(A, b, p)
print("  %-38s %s" % ("s = B^a mod p = %d^%d mod %d = %d" % (B, a, p, sA),
                      "s = A^b mod p = %d^%d mod %d = %d" % (A, b, p, sB)))
print()
print("  shared secret: Alice has %d, Bob has %d, equal: %s" % (sA, sB, sA == sB))
print("  it works because (g^b)^a and (g^a)^b are both g^(a*b) mod p")
print()
print("  what the eavesdropper saw : p = %d, g = %d, A = %d, B = %d" % (p, g, A, B))
print("  what the eavesdropper needs: a or b")
print("  with p = 23 they simply try every one:")
for guess in range(1, p):
    if pow(g, guess, p) == A:
        print("    g^%d mod %d = %d, so a = %d. Broken in %d tries."
              % (guess, p, A, guess, guess))
        break
print()

print("PART 2: a real group, RFC 7919's ffdhe2048")
# The first and last lines of the prime as RFC 7919 prints it, reassembled.
P = int(
    "FFFFFFFFFFFFFFFFADF85458A2BB4A9AAFDC5620273D3CF1"
    "D8B9C583CE2D3695A9E13641146433FBCC939DCE249B3EF9"
    "7D2FE363630C75D8F681B202AEC4617AD3DF1ED5D5FD6561"
    "2433F51F5F066ED0856365553DED1AF3B557135E7F57C935"
    "984F0C70E0E68B77E2A689DAF3EFE8721DF158A136ADE735"
    "30ACCA4F483A797ABC0AB182B324FB61D108A94BB2C8E3FB"
    "B96ADAB760D7F4681D4F42A3DE394DF4AE56EDE76372BB19"
    "0B07A7C8EE0A6D709E02FCE1CDF7E2ECC03404CD28342F61"
    "9172FE9CE98583FF8E4F1232EEF28183C3FE3B1B4C6FAD73"
    "3BB5FCBC2EC22005C58EF1837D1683B2C6F34A26C1B2EFFA"
    "886B423861285C97FFFFFFFFFFFFFFFF", 16)
G = 2
print("  p is %d bits, the ffdhe2048 group of RFC 7919 section A.1" % P.bit_length())
print("  g = %d" % G)
# Two private keys, fixed so the run repeats. In real use they are random.
a = int(hashlib.sha256(b"alice private key, this run only").hexdigest(), 16)
b = int(hashlib.sha256(b"bob private key, this run only").hexdigest(), 16)
A = pow(G, a, P)
B = pow(G, b, P)
sA = pow(B, a, P)
sB = pow(A, b, P)
print()
print("  Alice sends (first 32 hex digits) :", ("%x" % A)[:32], "...")
print("  Bob   sends (first 32 hex digits) :", ("%x" % B)[:32], "...")
print("  Alice computes                    :", ("%x" % sA)[:32], "...")
print("  Bob   computes                    :", ("%x" % sB)[:32], "...")
print("  equal :", sA == sB)
key = hashlib.sha256(sA.to_bytes((sA.bit_length() + 7) // 8, "big")).hexdigest()
print()
print("  the shared number is NOT used as a key directly. It is hashed:")
print("  session key = sha256(shared) =", key)
print()
print("  to find a from A an attacker must solve the discrete logarithm in a")
print("  %d-bit group. That is the whole security of the exchange." % P.bit_length())
print()

print("PART 3: the attack, which is the point of this practical")
p, g = 23, 5
a, b, m = 6, 15, 9
A, B, M = pow(g, a, p), pow(g, b, p), pow(g, m, p)
print("  Mallory sits between them and answers both.")
print("  Alice sends A = %d, but Mallory keeps it and sends Bob M = %d" % (A, M))
print("  Bob sends B = %d, but Mallory keeps it and sends Alice M = %d" % (B, M))
alice_key = pow(M, a, p)
bob_key = pow(M, b, p)
mallory_with_alice = pow(A, m, p)
mallory_with_bob = pow(B, m, p)
print()
print("  Alice's shared secret   : %d" % alice_key)
print("  Mallory's, with Alice   : %d   match: %s" % (mallory_with_alice, alice_key == mallory_with_alice))
print("  Bob's shared secret     : %d" % bob_key)
print("  Mallory's, with Bob     : %d   match: %s" % (mallory_with_bob, bob_key == mallory_with_bob))
print()
print("  Alice and Bob each believe they share a secret with the other. They")
print("  each share one with Mallory, who reads and can rewrite everything.")
print("  Alice and Bob's keys are %d and %d, and they are NOT equal: %s"
      % (alice_key, bob_key, alice_key == bob_key))
print()
print("  Nothing in the arithmetic is broken. Diffie-Hellman never claimed to")
print("  say WHO is at the other end, only that the two ends agree on a")
print("  number nobody watching can compute. The fix is to AUTHENTICATE the")
print("  exchange: Alice signs A with her private key, Bob signs B with his,")
print("  and Mallory cannot forge either signature. That is what TLS does,")
print("  and it is why a TLS server needs a certificate.")
munotes.in154

Practical 15: Key Exchange Using Diffie-Hellman

PART 1: small numbers, so every step can be checked
  public, agreed in the open: prime p = 23, generator g = 5

  Alice                                  Bob
  picks a secret a = 6                   picks a secret b = 15
  A = g^a mod p = 5^6 mod 23 = 8         B = g^b mod p = 5^15 mod 23 = 19
  sends A = 8                            sends B = 19
  s = B^a mod p = 19^6 mod 23 = 2        s = A^b mod p = 8^15 mod 23 = 2

  shared secret: Alice has 2, Bob has 2, equal: True
  it works because (g^b)^a and (g^a)^b are both g^(a*b) mod p

  what the eavesdropper saw : p = 23, g = 5, A = 8, B = 19
  what the eavesdropper needs: a or b
  with p = 23 they simply try every one:
    g^6 mod 23 = 8, so a = 6. Broken in 6 tries.

PART 2: a real group, RFC 7919's ffdhe2048
  p is 2048 bits, the ffdhe2048 group of RFC 7919 section A.1
  g = 2

  Alice sends (first 32 hex digits) : db57329a70452ba38ac57b4a74f578e4 ...
  Bob   sends (first 32 hex digits) : fa0bf1211a8155bcee3a8425f78c42b9 ...
  Alice computes                    : 3aaa660ff6b3c725ef4fc65e8f94f3be ...
  Bob   computes                    : 3aaa660ff6b3c725ef4fc65e8f94f3be ...
  equal : True

  the shared number is NOT used as a key directly. It is hashed:
  session key = sha256(shared) = 3a03df73482261ad8c6c4ef05c617ae25788785722bcc984493c27be2e914997

  to find a from A an attacker must solve the discrete logarithm in a
  2048-bit group. That is the whole security of the exchange.

PART 3: the attack, which is the point of this practical
  Mallory sits between them and answers both.
  Alice sends A = 8, but Mallory keeps it and sends Bob M = 11
  Bob sends B = 19, but Mallory keeps it and sends Alice M = 11

  Alice's shared secret   : 9
  Mallory's, with Alice   : 9   match: True
  Bob's shared secret     : 10
  Mallory's, with Bob     : 10   match: True

  Alice and Bob each believe they share a secret with the other. They
  each share one with Mallory, who reads and can rewrite everything.
  Alice and Bob's keys are 9 and 10, and they are NOT equal: False

  Nothing in the arithmetic is broken. Diffie-Hellman never claimed to
  say WHO is at the other end, only that the two ends agree on a
  number nobody watching can compute. The fix is to AUTHENTICATE the
  exchange: Alice signs A with her private key, Bob signs B with his,
  and Mallory cannot forge either signature. That is what TLS does,
  and it is why a TLS server needs a certificate.
munotes.in155

Practical 15: Key Exchange Using Diffie-Hellman

Reading part 1

With p = 23 every step can be checked by hand, and an examiner may ask you to. 5 to the power 6 is 15,625, and 15,625 divided by 23 leaves 8. 5 to the power 15 mod 23 is 19. Then 19 to the power 6 mod 23 is 2 and 8 to the power 15 mod 23 is 2, and the two agree.

munotes.in156

Practical 15: Key Exchange Using Diffie-Hellman

And the eavesdropper broke it in six tries. With p = 23 there are only 22 possible secrets, so trying each one until g^guess mod p matches A takes no time at all. The security is entirely in the size of p.

Reading part 2

The prime is 2048 bits, and it is not invented: it is the ffdhe2048 group of RFC 7919, one of five groups the IETF published so that a server and a client can agree on a group known to be sound instead of one a server made up. The reason that matters is worth a line in the journal: a maliciously chosen p can be one for which the discrete logarithm is easy, and a client that accepts whatever group the server offers has no way to tell.

Two details of the real exchange that the small example hides.

The shared number is not used as a key. It is a 2048-bit number with structure, and a key needs to be a fixed-size string of uniform bytes. So it is put through a hash, or properly a key derivation function, and the output is the session key. The program shows the SHA-256 step.

The private exponents here come from a fixed hash, so that the run repeats. In real use they are random, from the operating system's source, as Chapter 14's rule requires. A Diffie-Hellman private exponent that can be guessed is a session that can be read.

Reading part 3, which is the point of the practical

MU's wording is "securely exchange keys between two entities over an insecure network", and the third part is what "insecure" actually costs.

Mallory does not break any mathematics. She simply answers. Alice's A never reaches Bob; Mallory keeps it and sends Bob her own M. Bob's B never reaches Alice; Mallory keeps it and sends Alice the same M.

munotes.in157

Practical 15: Key Exchange Using Diffie-Hellman

The output shows the result exactly: Alice's shared secret is 9, and Mallory has 9. Bob's is 10, and Mallory has 10. Alice and Bob believe they share a key with each other, and their two keys are not even the same number. Every message Alice encrypts, Mallory decrypts, reads, re-encrypts under Bob's key and forwards. Neither of them sees anything wrong.

Diffie-Hellman never claimed to say who is at the other end. It says that two ends agree on a number nobody watching can compute, and "watching" is the word doing the work: an attacker who only listens learns nothing, and an attacker who can also modify the traffic defeats it completely.

The fix is authentication. Alice signs her A with the private key of Practical 14, Bob signs his B, and Mallory cannot forge either signature, so her substituted M is rejected. That is exactly what TLS does, and it is why a TLS server needs a certificate: the certificate is what lets the client check the signature on the server's Diffie-Hellman share. Practical 17 configures one.

The same exchange with OpenSSL

$ openssl genpkey -genparam -algorithm DH -pkeyopt group:ffdhe2048 -out dhp.pem
$ openssl pkeyparam -in dhp.pem -text -noout | head -2
DH Parameters: (2048 bit)
GROUP: ffdhe2048
$ openssl genpkey -paramfile dhp.pem -out alice.pem
$ openssl genpkey -paramfile dhp.pem -out bob.pem
$ openssl pkey -in alice.pem -pubout -out alice.pub
$ openssl pkey -in bob.pem -pubout -out bob.pub
$ openssl pkeyutl -derive -inkey alice.pem -peerkey bob.pub -out alice.secret
$ openssl pkeyutl -derive -inkey bob.pem -peerkey alice.pub -out bob.secret
$ wc -c < alice.secret
256
$ cmp -s alice.secret bob.secret && echo "the two sides agree" || echo "the two sides DIFFER"
the two sides agree
$ sha256sum alice.secret bob.secret | awk '{print $1}' | sort -u | wc -l
1

Two keys, two derivations, 256 bytes each and identical, and neither side ever sent the secret. That is the same exchange the Python program did, with the group named rather than written out.

The last line is there because cmp prints nothing when two files match, and a check that prints nothing is a check a reader cannot see. Counting the distinct digests of the two files gives 1, which is the same fact in a form you can write down. The digests themselves are not printed: the keys are fresh on every run, so the value would be different for every reader.

Forward secrecy, which is what this is really for

One more reason Diffie-Hellman is everywhere, and it is the one students most often miss.

Suppose a server's traffic is encrypted with a key derived from the server's long-term RSA key, and suppose an attacker has been recording that traffic for a year. If the RSA private key is later stolen, every recorded session can be decrypted, going back as far as the recordings do.

munotes.in158

Practical 15: Key Exchange Using Diffie-Hellman

Now suppose each session used a fresh Diffie-Hellman exchange, with the exponents thrown away when the session ended. The long-term key only ever signed the exchange; it never encrypted anything. Steal it and you can impersonate the server from then on, and you still cannot read a single recorded session, because the exponents that made those keys no longer exist anywhere.

That property is forward secrecy, the ephemeral Diffie-Hellman exchange is how it is obtained, and it is why TLS 1.3 removed every key exchange that does not provide it.

Procedure

  1. Choose a small prime p and a generator g and run the exchange by hand on paper first.
  2. Write the program: two secret exponents, two public values, two derivations, and check the two shared numbers are equal.
  3. Print exactly what an eavesdropper sees, and then break the small case by trying every exponent.
  4. Repeat with a real 2048-bit group from RFC 7919 and hash the shared number into a session key.
  5. Perform the man-in-the-middle attack: one attacker exponent, one substituted public value in each direction, and print all four shared numbers.
  6. Record that Alice's and Bob's numbers are not equal, and that Mallory holds both.
  7. Say what authentication would have prevented it.
  8. Run the same exchange with openssl genpkey and openssl pkeyutl -derive, and confirm the two derived secrets are identical.

Observations

MeasuredValue
Small groupp = 23, g = 5
Alice's secret a, Bob's secret b6, 15
A sent, B sent8, 19
Shared secret, both sides2
Eavesdropper's work on p = 23found a = 6 in 6 tries
Real groupffdhe2048, RFC 7919 appendix A.1, 2048 bits, g = 2
Both sides' derived numberidentical
Session keySHA-256 of the shared number
Man in the middle, Alice's key9, and Mallory holds 9
Man in the middle, Bob's key10, and Mallory holds 10
Alice's key equals Bob's keyFalse
OpenSSL derived secret256 bytes, identical on both sides

Result

The Diffie-Hellman key exchange was implemented and run three ways. With p = 23 and g = 5 both sides derived the shared secret 2 without sending it, and an eavesdropper recovered the private exponent in six attempts, which establishes that the security lies in the size of the prime. With RFC 7919's 2048-bit ffdhe2048 group both sides derived the same number and hashed it into a session key. A man-in-the-middle attack was then performed against the unauthenticated exchange: the attacker substituted her own public value in both directions and ended holding both session keys, 9 with Alice and 10 with Bob, while Alice and Bob believed they shared a key with each other and in fact held different numbers. The same exchange with OpenSSL's ffdhe2048 parameters produced 256-byte secrets that were identical on both sides.

munotes.in159

Practical 15: Key Exchange Using Diffie-Hellman

Where marks are lost

Sending the shared secret. Nothing secret crosses the network. If your program sends it, you have not implemented Diffie-Hellman.

Using the shared number directly as a key. Hash it, or run a key derivation function over it.

A small or invented prime. Use a published group. A prime chosen by the other side can be one for which the discrete logarithm is easy.

Predictable private exponents. Use the operating system's random source.

Saying Diffie-Hellman is secure against a man in the middle. It is not, and the practical is not complete without the demonstration.

Confusing eavesdropping with modification. An attacker who only listens learns nothing; one who can alter the traffic defeats the exchange entirely.

Not explaining what fixes it. Authentication of the exchange: signatures, and in TLS, a certificate.

Reusing one exponent for every session. Then forward secrecy is lost, which is most of the reason for using this at all.

For the journal

Aim; the four steps of the exchange; why (g^a)^b equals (g^b)^a; the small-prime run with every number checked by hand; what the eavesdropper sees and what they cannot compute; the brute force on the small prime; the 2048-bit run with the group named and the session key derived by hashing; the man-in-the-middle section with all four numbers and the statement that Alice's and Bob's differ; what authentication fixes and why a TLS server needs a certificate; the OpenSSL session; forward secrecy in your own words; the observation table; the result.

Quick revision

  • Agree p and g in public. Alice sends g^a mod p, Bob sends g^b mod p. Both compute g^(a*b) mod p.
  • The shared secret is never sent.
  • Security rests on the discrete logarithm problem: recovering a from g^a mod p.
  • A small p is broken by trying every exponent. Ours fell in six tries.
  • Use a published group: RFC 7919's ffdhe2048 or larger.
  • Hash the shared number to get a key; do not use it raw.
  • Private exponents must come from a cryptographic random source.
  • Diffie-Hellman stops an eavesdropper and not a man in the middle.
  • In the attack, Alice and Bob hold different keys and the attacker holds both.
  • Authenticate the exchange with signatures. A TLS certificate is what makes that possible.
  • Ephemeral exponents, discarded after the session, give forward secrecy.
munotes.in160

Practical 15: Key Exchange Using Diffie-Hellman

Questions you must be able to answer

1. What crosses the network in a Diffie-Hellman exchange? The prime, the generator, and the two public values g^a mod p and g^b mod p. The secret exponents and the shared key never do.

2. Why do both sides end with the same number? Because (g^a)^b and (g^b)^a are both g^(a*b), and the modulus does not disturb that.

3. What problem must an attacker solve, and why can they not? The discrete logarithm: recovering a from g^a mod p. For a 2048-bit prime no known method is feasible. For p = 23 it took six tries.

4. Why use a published group rather than your own prime? Because a prime can be chosen so that the discrete logarithm in it is easy, and a client accepting whatever the server offers cannot tell. RFC 7919 publishes groups everyone can check.

5. Why is the shared number hashed before it is used? Because it is a large structured number, not a uniform key of the right length. A hash or a key derivation function turns it into one.

6. Describe the man-in-the-middle attack. The attacker intercepts both public values and substitutes her own in each direction. She then shares one key with each party, reads and rewrites everything, and neither party can detect it. In this run Alice held 9, Bob held 10, and the attacker held both.

7. Is that a flaw in the mathematics? No. Diffie-Hellman guarantees that two ends agree on a number no watcher can compute. It never claimed to identify who is at the other end.

8. What fixes it? Authenticating the exchange: each side signs its public value with a private key the other can verify. In TLS the server's certificate is what lets the client verify that signature.

9. What is forward secrecy, and how does Diffie-Hellman give it? That a session recorded today cannot be decrypted even if the long-term key is stolen tomorrow. Using a fresh exchange per session, and discarding the exponents afterwards, means the material that made the session key no longer exists.

10. Your two sides derive different numbers. Name the two likeliest causes. Someone is between you substituting values, which is the attack above; or one side has used the wrong peer's public value, or a different p or g. Check the parameters first, then suspect the network.

Contents This chapter on its own page

munotes.in161

Chapter Twenty-One

Practical 16: IP Security (IPsec) Configuration

Syllabus topic Module 2, "IP Security (IPsec) Configuration: Configure IPsec on network devices to provide secure communication and protect against unauthorized access and attacks."

Aim

To configure IPsec between two hosts, to show that traffic between them is unreadable on the wire afterwards, and to say what IPsec protects and what it does not.

What you need to know before you start

Everything in this module so far has protected one message. IPsec protects every packet, and it does it inside the operating system, so the programs sending the packets do not know it is happening and do not have to be changed. A web browser, a database client and a print job are all protected at once.

RFC 4301 gives it two protocols, and the difference is a standard question.

AH, the Authentication Header, authenticates the packet and its source address. It does not encrypt, so anybody can read the contents; they just cannot change them undetected.

ESP, the Encapsulating Security Payload, encrypts the contents and authenticates them as well. It is what is used in practice, and RFC 4301 says so in terms: an IPsec implementation "MUST support ESP and MAY support AH", because experience showed that there are very few situations ESP cannot cover, ESP being usable for integrity alone when confidentiality is not wanted.

And two modes:

Transport modeTunnel mode
What is protectedthe payload of the packetthe whole original packet
The IP headerthe original one, in the cleara new outer one; the original is encrypted
Who can see the real addressesanybodynobody outside
Used forhost to hosta gateway to gateway VPN

Two databases hold the configuration, and both names are examinable:

  • The Security Association Database, SAD, holds the keys and algorithms for each one-way protected connection. A security association is one direction only, so two hosts talking need two.
  • The Security Policy Database, SPD, says which traffic must be protected, which may pass in the clear, and which must be discarded.

Each association is identified by a Security Parameters Index, the SPI, which travels in every packet so the receiver knows which keys to use.

Step 1: two hosts, and the traffic in the clear

Two network namespaces joined by a virtual cable. Everything that follows happens inside one container.

$ ip netns add hostA
$ ip netns add hostB
$ ip link add vethA type veth peer name vethB
$ ip link set vethA netns hostA
$ ip link set vethB netns hostB
$ ip -n hostA addr add 10.9.0.1/24 dev vethA
$ ip -n hostB addr add 10.9.0.2/24 dev vethB
$ ip -n hostA link set vethA up
$ ip -n hostB link set vethB up
$ ip -n hostA link set lo up
$ ip -n hostB link set lo up
$ ip netns exec hostA ping -c 2 -W 2 10.9.0.2 2>/dev/null | tail -3
--- 10.9.0.2 ping statistics ---
2 packets transmitted, 2 received, 0% packet loss, time 1003ms
rtt min/avg/max/mdev = 0.000/0.000/0.000/0.000 ms
munotes.in162

Practical 16: IP Security (IPsec) Configuration

Two hosts that can reach each other. Now look at what a listener sees, with a message deliberately put inside the ping so that it is easy to find.

#!/bin/bash
# Capture two packets on hostB's end of the cable while hostA pings, then
# print them. tcpdump has to start BEFORE the ping and be waited for
# afterwards, so it is a script rather than three lines typed at the prompt:
# a background job left running at a prompt is the commonest way a capture
# comes out empty. `setsid` puts tcpdump in a session of its own, so it cannot
# hold the terminal open after the script has finished with it, and --wait
# makes setsid wait for it, so `wait` below really waits for the capture
# rather than for setsid itself. Two seconds is how long tcpdump needs to
# be listening before the first packet: with one, the capture came out empty.
# NOTE: `timeout 8` is not decoration: if tcpdump never sees its two packets it
# waits for ever, and the whole session waits with it.
setsid --wait timeout 8 ip netns exec hostB tcpdump -n -i vethB -c 2 -w /tmp/cap.pcap icmp >/dev/null 2>&1 </dev/null &
TD=$!
sleep 2
ip netns exec hostA ping -c 2 -i 0.4 -p 53454352455453454352455453454352 -s 15 10.9.0.2 >/dev/null 2>&1
wait $TD 2>/dev/null
tcpdump -n -r /tmp/cap.pcap -A 2>/dev/null

-p 5345435245 54... fills the ping's payload with the bytes 53 45 43 52 45 54 over and over, which spell SECRET. Run it and read the capture.

$ bash capture.sh > /tmp/clear.txt 2>&1
$ grep -c "ICMP echo" /tmp/clear.txt
2
$ grep -c SECRET /tmp/clear.txt
2
$ grep -o "SECRET[A-Z]*" /tmp/clear.txt | head -1
SECRETSECRETSEC

The count is 2, one for each captured packet, and the word is plainly visible in the readable column of the hex dump beside the bytes that produced it. Anybody on the path can read the contents of every packet, which is the state IPsec exists to change.

Step 2: configure IPsec

An association needs a direction, a protocol, an SPI, a mode and two keys: one for authentication and one for encryption. In real use those keys come from an IKE negotiation, which is Diffie-Hellman of Practical 15 with authentication over it; here they are written out so that the configuration itself is the thing being studied.

$ KA=0x0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20
$ KE=0x000102030405060708090a0b0c0d0e0f
$ ip -n hostA xfrm state add src 10.9.0.1 dst 10.9.0.2 proto esp spi 0x201 mode transport auth "hmac(sha256)" $KA enc "cbc(aes)" $KE
$ ip -n hostA xfrm state add src 10.9.0.2 dst 10.9.0.1 proto esp spi 0x202 mode transport auth "hmac(sha256)" $KA enc "cbc(aes)" $KE
$ ip -n hostB xfrm state add src 10.9.0.1 dst 10.9.0.2 proto esp spi 0x201 mode transport auth "hmac(sha256)" $KA enc "cbc(aes)" $KE
$ ip -n hostB xfrm state add src 10.9.0.2 dst 10.9.0.1 proto esp spi 0x202 mode transport auth "hmac(sha256)" $KA enc "cbc(aes)" $KE
$ ip -n hostA xfrm policy add src 10.9.0.1 dst 10.9.0.2 dir out tmpl proto esp mode transport
$ ip -n hostA xfrm policy add src 10.9.0.2 dst 10.9.0.1 dir in tmpl proto esp mode transport
$ ip -n hostB xfrm policy add src 10.9.0.2 dst 10.9.0.1 dir out tmpl proto esp mode transport
$ ip -n hostB xfrm policy add src 10.9.0.1 dst 10.9.0.2 dir in tmpl proto esp mode transport
$ ip -n hostA xfrm state | grep -E "^src|spi|auth|enc" | head -8
src 10.9.0.2 dst 10.9.0.1
	proto esp spi 0x00000202 reqid 0 mode transport
	auth-trunc hmac(sha256) 0x0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20 96
	enc cbc(aes) 0x000102030405060708090a0b0c0d0e0f
src 10.9.0.1 dst 10.9.0.2
	proto esp spi 0x00000201 reqid 0 mode transport
	auth-trunc hmac(sha256) 0x0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20 96
	enc cbc(aes) 0x000102030405060708090a0b0c0d0e0f
munotes.in163

Practical 16: IP Security (IPsec) Configuration

Four states and four policies, and the reason for four of each is the sentence to write in the journal: a security association is one-way, so each host needs one for the traffic it sends and one for the traffic it receives, and each host needs an outbound policy and an inbound policy to match.

auth "hmac(sha256)" and enc "cbc(aes)" are the two algorithms, and they are exactly the two halves of Practicals 13 and 11: a MAC for integrity and authenticity, and a block cipher for secrecy. IPsec is not new cryptography; it is the cryptography of this module applied to every packet.

Step 3: the same traffic, after

#!/bin/bash
# The same capture, filtered on `icmp or esp` rather than `icmp`: after IPsec
# there are no ICMP packets on the wire at all, and a capture that matched
# nothing would look like a broken capture rather than a result.
setsid --wait timeout 8 ip netns exec hostB tcpdump -n -i vethB -c 2 -w /tmp/cap2.pcap "icmp or esp" >/dev/null 2>&1 </dev/null &
TD=$!
sleep 2
ip netns exec hostA ping -c 2 -i 0.4 -p 53454352455453454352455453454352 -s 15 10.9.0.2 >/dev/null 2>&1
wait $TD 2>/dev/null
tcpdump -n -r /tmp/cap2.pcap -A 2>/dev/null
$ bash capture2.sh > /tmp/esp.txt 2>&1
$ grep -c "ICMP echo" /tmp/esp.txt
0
$ grep -c SECRET /tmp/esp.txt
0
$ grep -c "ESP(spi=0x00000201" /tmp/esp.txt
1
$ grep -o "ESP(spi=0x00000201,seq=0x1)" /tmp/esp.txt | head -1
ESP(spi=0x00000201,seq=0x1)

Zero occurrences of SECRET, and the packets now read as ESP. The word that was plainly visible in step 1 is not in the capture at all, and what the listener sees instead is the protocol name, the SPI and a length. The ping still works, because the two kernels decrypt and re-encrypt without the ping program knowing anything about it.

munotes.in164

Practical 16: IP Security (IPsec) Configuration

Confirm that the traffic really is still flowing, and look at the counters the kernel keeps.

$ ip netns exec hostA ping -c 3 -W 2 10.9.0.2 2>/dev/null | tail -2
3 packets transmitted, 3 received, 0% packet loss, time 2026ms
rtt min/avg/max/mdev = 0.000/0.000/0.000/0.000 ms
$ ip -n hostA xfrm state | grep -c "spi 0x00000201"
1
$ ip -n hostA xfrm policy | head -6
src 10.9.0.2/32 dst 10.9.0.1/32
	dir in priority 0
	tmpl src 0.0.0.0 dst 0.0.0.0
		proto esp reqid 0 mode transport
src 10.9.0.1/32 dst 10.9.0.2/32
	dir out priority 0

Step 4: what happens when one side is wrong

A configuration exercise is not finished until it has been broken on purpose, because that is what the examiner will ask about.

$ ip -n hostB xfrm state delete src 10.9.0.1 dst 10.9.0.2 proto esp spi 0x201
$ ip netns exec hostA ping -c 2 -W 2 10.9.0.2 2>/dev/null | tail -2
2 packets transmitted, 0 received, 100% packet loss, time 1016ms
$ ( ip netns exec hostA ping -c 2 -W 2 10.9.0.2 >/dev/null 2>&1 ); echo "ping exit status $?"
ping exit status 1

100 per cent packet loss, and ping exits 1. The second line runs ping on its own rather than through tail, because a pipeline reports the status of its LAST command and tail always succeeds.

100 per cent packet loss. Host A still encrypts, because its policy still says to; host B no longer has the key to decrypt, so it drops every packet. Nothing reports an error to the user: the ping simply does not come back, and that is exactly what a misconfigured IPsec looks like in real life. Check both ends.

$ ip -n hostB xfrm state add src 10.9.0.1 dst 10.9.0.2 proto esp spi 0x201 mode transport auth "hmac(sha256)" 0x0102030405060708090a0b0c0d0e0f101112131415161718191a1b1c1d1e1f20 enc "cbc(aes)" 0x000102030405060708090a0b0c0d0e0f
$ ip netns exec hostA ping -c 2 -W 2 10.9.0.2 2>/dev/null | tail -2
2 packets transmitted, 2 received, 0% packet loss, time 1002ms
rtt min/avg/max/mdev = 0.000/0.000/0.000/0.000 ms

Put the association back and it works again.

Step 5: clean up

$ ip -n hostA xfrm policy flush
$ ip -n hostA xfrm state flush
$ ip -n hostB xfrm policy flush
$ ip -n hostB xfrm state flush
$ ip netns delete hostA
$ ip netns delete hostB
$ ip netns list | wc -l
0
munotes.in165

Practical 16: IP Security (IPsec) Configuration

What IPsec protects, and what it does not

It protects: the contents of every packet, against reading and against alteration; the source address, against spoofing, because a packet from the wrong source has no valid association; and, in tunnel mode, the real addresses too.

It does not protect: against a compromised endpoint, because the packets are in the clear at both ends; against traffic analysis, because an observer still sees that two addresses are talking, how much and when; and against anything at all if the keys are wrong, weak or shared.

And the manual keys in this chapter are not how it is done. A real deployment runs IKEv2 to negotiate the keys, which is Practical 15's Diffie-Hellman authenticated by Practical 14's signatures or by a certificate, rekeying periodically so that a stolen key exposes only a short window. Manual keying is used here because the exercise is about the associations and the policies, and because IKE would hide the very thing being studied.

Procedure

  1. Create two network namespaces and join them with a veth pair. Give each an address and confirm they can ping.
  2. Capture the traffic with tcpdump and find the payload in the clear.
  3. Choose an authentication key and an encryption key, and add four security associations, two on each host, one for each direction.
  4. Add four policies, an inbound and an outbound on each host.
  5. Print the state and the policy and identify the SPI, the mode and the two algorithms.
  6. Capture again, and record that the payload is gone and the packets are ESP.
  7. Confirm the ping still works.
  8. Delete one association and record what the failure looks like. Put it back.
  9. Flush both databases and delete the namespaces.

Observations

MeasuredValue
Hosts10.9.0.1 and 10.9.0.2, two network namespaces joined by a veth pair
Before IPsec, the word SECRET in the capturepresent
Security associations added4, two per host, one per direction
Policies added4, an in and an out on each host
SPI, A to B and B to A0x201 and 0x202
Protocol and modeESP, transport
Authentication algorithmhmac(sha256)
Encryption algorithmcbc(aes)
After IPsec, the word SECRET in the captureabsent
After IPsec, packets identified as ESPpresent, with the SPI visible
Ping after IPsecstill works
One association deleted100 per cent packet loss, no error message
Association restoredworks again

Result

IPsec was configured between two hosts created as network namespaces inside one machine. Before configuration a capture on the link showed the payload of every packet in the clear, including a word placed there to be found. Four ESP security associations were then added, two on each host because an association is one-way, together with four policies, and the association carried HMAC-SHA256 for authentication and AES in CBC mode for encryption under SPI 0x201 in one direction and 0x202 in the other. A second capture of identical traffic contained no occurrence of the payload word and showed ESP packets carrying the SPI instead, while the ping continued to work unchanged. Deleting one association produced 100 per cent packet loss with no error message, which is what a one-sided misconfiguration looks like, and restoring it restored the traffic.

munotes.in166

Practical 16: IP Security (IPsec) Configuration

Where marks are lost

Configuring one direction. A security association is one-way. Two hosts need four, and four policies.

A state without a policy, or a policy without a state. The policy says protect this traffic; the state says with what. Both, on both hosts.

Keys that do not match. The packets are silently dropped and nothing says why.

Expecting an error message. A broken IPsec looks like a network that has stopped working. The counters in ip xfrm state are where to look.

Saying AH encrypts. It authenticates only. ESP encrypts and authenticates.

Confusing the two modes. Transport protects the payload and leaves the original header; tunnel wraps the whole packet in a new one and is what a gateway VPN uses.

Presenting manual keys as normal practice. IKEv2 negotiates them, using Diffie-Hellman with authentication, and rekeys.

Not showing the before. A capture showing ESP proves nothing unless a capture of the same traffic beforehand showed the payload.

For the journal

Aim; AH against ESP and transport against tunnel, as two tables; SAD, SPD and SPI defined; the commands that build the two hosts; the capture before, with the payload visible; the four states and four policies with the algorithms named; the capture after, with the payload absent and the SPI visible; the ping that still works; the deliberate breakage and what it looked like; the clean-up; what IPsec protects and what it does not; the note that real deployments use IKEv2; the observation table; the result.

Quick revision

  • IPsec protects every packet, in the kernel, so applications need no change.
  • AH authenticates. ESP encrypts and authenticates. ESP is what is used.
  • Transport mode protects the payload; tunnel mode wraps the whole packet in a new one.
  • The SAD holds keys and algorithms; the SPD says which traffic to protect.
  • A security association is ONE WAY. Two hosts need four associations and four policies.
  • The SPI identifies the association and travels in every packet.
  • ip xfrm state and ip xfrm policy are the two commands.
  • A wrong key or a missing state gives silent packet loss, not an error.
  • Real deployments negotiate keys with IKEv2, which is authenticated Diffie-Hellman.
  • IPsec does not hide who is talking to whom, or how much.
munotes.in167

Practical 16: IP Security (IPsec) Configuration

Questions you must be able to answer

1. What is the difference between AH and ESP? AH authenticates the packet and its source but does not encrypt, so the contents stay readable. ESP encrypts the contents and authenticates them as well, and it is what is used in practice.

2. What is the difference between transport and tunnel mode? Transport protects the payload and keeps the original IP header, which suits host-to-host protection. Tunnel encrypts the entire original packet inside a new one, which hides the real addresses and is what gateway VPNs use.

3. What are the SAD and the SPD? The Security Association Database holds the keys, algorithms and parameters for each protected one-way connection. The Security Policy Database says which traffic must be protected, which may pass and which must be dropped.

4. Why did you configure four security associations for two hosts? Because an association is one-way. Each host needs one for what it sends and one for what it receives, and both hosts must hold both.

5. What is the SPI, and why is it needed? The Security Parameters Index identifies which association a packet belongs to. It travels in the packet so the receiver knows which keys to apply, since it may hold many.

6. How did you prove the traffic is protected? By capturing the same ping before and after. Before, the payload word appeared in the capture; after, it did not appear at all and the packets were ESP with the SPI visible.

7. You deleted one association and the ping stopped. Was there an error message? No. The sender encrypted as its policy required and the receiver had no key, so it dropped the packets silently. A misconfigured IPsec looks like a network that has stopped working.

8. Where do the keys come from in a real deployment? From IKEv2, which runs an authenticated Diffie-Hellman exchange and rekeys periodically. Writing keys by hand is a teaching arrangement.

9. Does IPsec hide who is talking to whom? In transport mode, no: the original addresses are in the clear. In tunnel mode the inner addresses are hidden, but an observer still sees the two gateways, and how much traffic passes and when.

10. Application-level encryption already exists. Why protect at the IP layer? Because it covers every application at once, including ones that have no encryption of their own, and it requires no change to any of them.

Contents This chapter on its own page

munotes.in168

Chapter Twenty-Two

Practical 17: Web Security with SSL/TLS

Syllabus topic Module 2, "Web Security with SSL/TLS: Configure and implement secure web communication using SSL/TLS protocols, including certificate management and secure session establishment."

Aim

To configure a web server for HTTPS, to manage the certificates it needs, to watch a session being established, and to produce the three ways a certificate can be refused.

What you need to know before you start

TLS, Transport Layer Security, is what the S in HTTPS stands for. SSL is its old name: SSL 2.0 and 3.0 are both broken and withdrawn, and anybody saying SSL today means TLS. The versions that matter now are TLS 1.2 and TLS 1.3; 1.0 and 1.1 are withdrawn too.

It gives three things, and every one of them is a practical you have already done.

TLS givesWhich practical it is
A shared key over an open networkPractical 15, Diffie-Hellman
Proof of who the server isPractical 14, a digital signature
Secrecy and integrity of the dataPracticals 11 and 13

Practical 15 ended with the attack Diffie-Hellman cannot stop, and the answer was: authenticate the exchange. A certificate is how that is done on the web. It is a document saying "this public key belongs to this name", signed by somebody the client already trusts.

Three parties, and the names are examinable:

  • The subject: whose key it is, named by a domain.
  • The issuer, or Certificate Authority: who signed it.
  • The relying party: the browser, which holds a list of CAs it trusts and checks the signature.

Step 1: become a certificate authority

A real CA is an organisation with audits and a place in the browsers' trust list. For a laboratory, a CA is a key and a self-signed certificate.

$ cd /root
$ openssl genrsa -out ca.key 2048 2>/dev/null
$ openssl req -x509 -key ca.key -days 3650 -set_serial 1 -out ca.crt -subj "/C=IN/ST=Maharashtra/L=Mumbai/O=munotes practical CA/CN=munotes practical CA" 2>/dev/null
$ openssl x509 -in ca.crt -noout -subject -issuer -dates -serial
subject=C = IN, ST = Maharashtra, L = Mumbai, O = munotes practical CA, CN = munotes practical CA
issuer=C = IN, ST = Maharashtra, L = Mumbai, O = munotes practical CA, CN = munotes practical CA
notBefore=Sep 29 05:00:17 2026 GMT
notAfter=Sep 26 05:00:17 2036 GMT
serial=01

Why the key is generated as its own command. openssl req -x509 -newkey rsa:2048 does both in one step and looks tidier, and in this laboratory it produced a certificate that could not be verified. Generating a 2048-bit key takes a second or two of real time, and the notBefore written afterwards was two or three seconds ahead of the clock that then checked it, so the certificate was refused as not yet valid by the very machine that had just issued it. Making the key first, and the certificate second, puts notBefore exactly at the moment of issue and the problem disappears.

munotes.in169

Practical 17: Web Security with SSL/TLS

That is not only a laboratory artefact. A certificate is invalid before its start time, so a client whose clock is a few seconds behind the issuer's refuses a brand-new certificate for the same reason, and real certificate authorities backdate what they issue by a few minutes precisely because the clocks of the world do not agree.

Look at the subject and the issuer: they are the same. That is what self-signed means, and it is what every root CA certificate looks like. A root has nobody above it, so it signs itself, and it is trusted because it is in the browser's list, not because of anything in the certificate.

-nodes means the private key is stored without a passphrase, which is right for a laboratory and wrong for a real CA. -set_serial 1 fixes the serial number; without it OpenSSL picks a random one and no two runs of this chapter would agree.

Step 2: a key and a certificate signing request for the server

The server makes its own key. The private key never leaves the server, and that is the point of a signing request: it carries the public key and the name, and nothing secret.

$ openssl genrsa -out server.key 2048 2>/dev/null
$ openssl req -new -key server.key -out server.csr -subj "/C=IN/ST=Maharashtra/L=Mumbai/O=Computer Science Department/CN=localhost" 2>/dev/null
$ openssl req -in server.csr -noout -subject
subject=C = IN, ST = Maharashtra, L = Mumbai, O = Computer Science Department, CN = localhost
$ openssl req -in server.csr -noout -verify 2>&1 | head -1
Certificate request self-signature verify OK
$ ls -l ca.key ca.crt server.key server.csr | awk '{print $9}'
ca.crt
ca.key
server.csr
server.key

-verify checks that the request is signed by the private key matching the public key inside it, which is how a CA knows the applicant actually holds the key they are asking it to certify.

Step 3: the CA signs it

$ printf 'subjectAltName=DNS:localhost\nbasicConstraints=CA:FALSE\nkeyUsage=digitalSignature,keyEncipherment\n' > server.ext
$ openssl x509 -req -in server.csr -CA ca.crt -CAkey ca.key -set_serial 2 -days 365 -extfile server.ext -out server.crt 2>/dev/null
$ openssl x509 -in server.crt -noout -subject -issuer -dates -serial
subject=C = IN, ST = Maharashtra, L = Mumbai, O = Computer Science Department, CN = localhost
issuer=C = IN, ST = Maharashtra, L = Mumbai, O = munotes practical CA, CN = munotes practical CA
notBefore=Sep 29 05:00:17 2026 GMT
notAfter=Sep 29 05:00:17 2027 GMT
serial=02
$ openssl verify -CAfile ca.crt server.crt
server.crt: OK
$ openssl verify server.crt
C = IN, ST = Maharashtra, L = Mumbai, O = Computer Science Department, CN = localhost
error 20 at 0 depth lookup: unable to get local issuer certificate
error server.crt: verification failed
munotes.in170

Practical 17: Web Security with SSL/TLS

Now the subject and the issuer are different: the certificate says the key belongs to localhost and the CA says so.

The three extensions are not optional, and each has a reason:

  • subjectAltName is where the name actually lives. RFC 5280 and every browser since about 2017 read the name from here and ignore the common name, so a certificate with only a CN is rejected by browsers with a message that looks nothing like the real cause.
  • basicConstraints=CA:FALSE says this certificate may not sign others. Without it, anybody holding this key could issue certificates for any site.
  • keyUsage says what the key may be used for.

And the last two lines are the pair to put in the journal: with the CA, OK; without it, unable to get local issuer certificate. A certificate on its own proves nothing. It proves something only to somebody who already trusts the signer.

Step 4: configure the web server

$ mkdir -p /etc/ssl/munotes
$ cp server.crt server.key ca.crt /etc/ssl/munotes/
$ a2enmod ssl > /dev/null 2>&1; echo "ssl module enabled"
ssl module enabled
$ printf '<VirtualHost *:443>\n    ServerName localhost\n    DocumentRoot /var/www/html\n    SSLEngine on\n    SSLCertificateFile    /etc/ssl/munotes/server.crt\n    SSLCertificateKeyFile /etc/ssl/munotes/server.key\n</VirtualHost>\n' > /etc/apache2/sites-available/munotes-ssl.conf
$ cat /etc/apache2/sites-available/munotes-ssl.conf
<VirtualHost *:443>
    ServerName localhost
    DocumentRoot /var/www/html
    SSLEngine on
    SSLCertificateFile    /etc/ssl/munotes/server.crt
    SSLCertificateKeyFile /etc/ssl/munotes/server.key
</VirtualHost>
$ a2ensite munotes-ssl > /dev/null 2>&1; echo "site enabled"
site enabled
$ echo "ServerName localhost" >> /etc/apache2/apache2.conf
$ echo "<h1>Practical 17</h1>" > /var/www/html/index.html
$ apachectl configtest
Syntax OK
$ apachectl start
$ sleep 2; ss -lnt | grep -c ":443"
1

Four directives are the whole of it: SSLEngine on, the certificate, the key, and a virtual host on port 443. configtest before start is a habit worth having: it catches a mistyped path while the old configuration is still serving.

Step 5: the session, established

$ echo | timeout 10 openssl s_client -connect localhost:443 -CAfile /etc/ssl/munotes/ca.crt -servername localhost 2>/dev/null | grep -E "^New,|^Server public key|Verify return code|^subject=|^issuer=" | sed 's/^ *//' | uniq
subject=C = IN, ST = Maharashtra, L = Mumbai, O = Computer Science Department, CN = localhost
issuer=C = IN, ST = Maharashtra, L = Mumbai, O = munotes practical CA, CN = munotes practical CA
New, TLSv1.3, Cipher is TLS_AES_256_GCM_SHA384
Server public key is 2048 bit
Verify return code: 0 (ok)

That one line is the whole of MU's "secure session establishment", and each part of it is worth naming.

New, TLSv1.3: the version agreed. Both sides offer what they support and the highest common version wins.

Cipher is TLS_AES_256_GCM_SHA384: the cipher suite. In TLS 1.3 the name has three parts: AES-256 in GCM mode for the data, and SHA-384 inside the key schedule. GCM is authenticated encryption, the thing Practical 13 pointed at: it encrypts and authenticates in one pass, so there is no separate MAC to get wrong. Note what the name does not contain: the key exchange, because TLS 1.3 removed every option except ephemeral Diffie-Hellman. Forward secrecy is no longer negotiable.

munotes.in171

Practical 17: Web Security with SSL/TLS

Server public key is 2048 bit: the RSA key in the certificate, used to sign the exchange, not to encrypt anything.

Verify return code: 0 (ok): the certificate chain checked out against the CA we passed.

And the page itself, over the encrypted connection:

$ printf 'GET / HTTP/1.0\r\nHost: localhost\r\n\r\n' | timeout 10 openssl s_client -connect localhost:443 -CAfile /etc/ssl/munotes/ca.crt -servername localhost -quiet 2>/dev/null | tail -2

<h1>Practical 17</h1>

Step 6: the three ways a certificate is refused

A configuration practical is only finished when it has been broken deliberately, and these are the three failures a student will actually meet.

The issuer is not trusted. The same connection, without telling OpenSSL about our CA:

$ echo | timeout 10 openssl s_client -connect localhost:443 -servername localhost 2>/dev/null | grep -E "Verify return code" | sed 's/^ *//' | uniq
Verify return code: 21 (unable to verify the first certificate)

This is what a browser shows for a self-signed or privately signed certificate, and it is not a bug in the certificate. Our CA is not in anybody's trust list. The cure is to add it, which is what an organisation does on its own machines, and what you must never tell a stranger to do.

The certificate has expired. A certificate signed with a negative lifetime is expired the moment it exists:

$ openssl x509 -req -in server.csr -CA ca.crt -CAkey ca.key -set_serial 3 -days -1 -extfile server.ext -out expired.crt 2>/dev/null
$ openssl x509 -in expired.crt -noout -dates
notBefore=Sep 29 05:00:17 2026 GMT
notAfter=Sep 28 05:00:17 2026 GMT
$ openssl verify -CAfile ca.crt expired.crt 2>&1 | tail -2
error 10 at 0 depth lookup: certificate has expired
error expired.crt: verification failed

Look at the dates: notAfter is before notBefore. Expiry is the commonest outage in this whole subject, and it is entirely avoidable: certificates have a date on them and somebody has to renew them.

The name does not match. A certificate can be perfectly valid and still be the wrong certificate:

$ openssl genrsa -out other.key 2048 2>/dev/null
$ openssl req -new -key other.key -out other.csr -subj "/CN=example.invalid" 2>/dev/null
$ printf 'subjectAltName=DNS:example.invalid\n' > other.ext
$ openssl x509 -req -in other.csr -CA ca.crt -CAkey ca.key -set_serial 4 -days 365 -extfile other.ext -out other.crt 2>/dev/null
$ openssl verify -CAfile ca.crt other.crt
other.crt: OK
$ echo | timeout 10 openssl s_client -connect localhost:443 -CAfile ca.crt -servername localhost -verify_hostname example.invalid 2>/dev/null | grep -E "Verify return code" | sed 's/^ *//' | uniq
Verify return code: 62 (hostname mismatch)
munotes.in172

Practical 17: Web Security with SSL/TLS

other.crt: OK and then a hostname mismatch. The certificate is properly signed by a trusted CA and is entirely in date, and it is still refused, because the name in it is not the name that was asked for. That distinction is the one students miss: verifying the chain and verifying the name are two separate checks, and a client that does the first and forgets the second accepts any valid certificate for any site, which is a man-in-the-middle attack wearing a padlock.

Step 7: stop the server

$ apachectl stop
$ sleep 1; ss -lnt | grep -c ":443"
0

What is in a certificate, and what revocation is

A certificate under RFC 5280 carries: a version, a serial number, the signature algorithm, the issuer's name, a validity period with a notBefore and a notAfter, the subject's name, the subject's public key, a set of extensions including subjectAltName and basicConstraints, and the issuer's signature over all of it.

A chain is what a real server sends: its own certificate, then the intermediate that signed it, and so on up to a root the client already has. The root is not sent, because the client must already hold it.

Revocation is the answer to a key being stolen before the certificate expires, and it is the weakest part of the whole system. There are two mechanisms, and a student should be able to name both: a certificate revocation list, a signed list of revoked serial numbers that the client downloads, and OCSP, defined in RFC 6960, where the client asks the issuer about one certificate. Both have the same practical problem: if the check cannot be made, clients generally carry on rather than refuse, so an attacker who can block the check has defeated it. This is why certificate lifetimes have been getting shorter: a certificate that expires soon does not need revoking.

Procedure

  1. Create a CA key and a self-signed CA certificate. Note that the subject and the issuer are the same.
  2. Create the server's key and a certificate signing request, and verify the request's own signature.
  3. Write an extensions file with subjectAltName, basicConstraints and keyUsage, and sign the request with the CA.
  4. Print the subject, issuer, dates and serial of the signed certificate, and verify it with and without the CA file.
  5. Enable the SSL module, write a virtual host on 443 naming the certificate and the key, enable the site, run configtest, and start the server.
  6. Confirm something is listening on 443.
  7. Connect with s_client and record the version, the cipher suite, the key size and the verify code.
  8. Fetch the page over the encrypted connection.
  9. Produce all three failures: untrusted issuer, expired certificate, wrong hostname.
  10. Stop the server.
munotes.in173

Practical 17: Web Security with SSL/TLS

Observations

MeasuredValue
CA certificateself-signed, serial 01
Server certificateCN=localhost, serial 02
verify, with the CA fileserver.crt: OK
verify, without itno local issuer
Apache configtestSyntax OK
Listening on 4431 socket
TLS version agreedTLSv1.3
Cipher suiteTLS_AES_256_GCM_SHA384
Server public key2048 bit
Verify code, with the CA0 (ok)
Verify code, no CA21, cannot verify the first
Page fetched over TLSthe HTML Apache served
Expired certificatenotAfter before notBefore
Certificate, another nameverifies OK, handshake 62

Result

A private certificate authority was created and used to issue a server certificate for localhost with subjectAltName, basicConstraints and keyUsage extensions. Apache 2.4.58 was configured with four directives to serve HTTPS on port 443, its configuration checked and the server started. A TLS session was established and recorded: TLS 1.3, the cipher suite TLS_AES_256_GCM_SHA384, a 2048-bit server key and a verify return code of 0, and a page was fetched over the encrypted connection. Three separate refusals were then produced: the same connection without the CA gave code 21, a certificate signed with a negative lifetime had notAfter before notBefore and was reported as expired, and a certificate for another name verified correctly against the CA and was still refused by the handshake with code 62, hostname mismatch, which demonstrates that chain verification and name verification are two independent checks.

Where marks are lost

Saying SSL when you mean TLS. SSL 2.0 and 3.0 are broken and withdrawn. Say TLS, and name the version.

A certificate with no subjectAltName. Browsers read the name from there and ignore the common name. The certificate looks right and is rejected.

Leaving out basicConstraints=CA:FALSE. Anybody with that key could then issue certificates for any site.

Passing the private key to the CA. The CA signs a request, which carries only the public key. The private key never leaves the server.

Skipping configtest. A mistyped path stops the server from coming back up.

Reporting only the successful handshake. Produce all three failures; that is where the understanding shows.

Treating an untrusted issuer as a broken certificate. It means the CA is not in the trust list, which for a private CA is the expected state.

Telling a user to click through a warning. On your own machine, add your own CA to the trust store. Never advise a stranger to ignore one.

Forgetting that the name is checked separately. A valid certificate for the wrong name is exactly how a man-in-the-middle attack looks.

For the journal

Aim; what TLS gives and which practical each part corresponds to; subject, issuer and relying party defined; the CA creation with subject equal to issuer; the key and the signing request, with a note that the private key stays put; the extensions file with a line on each extension; the signing and the two verify results; the Apache virtual host with its four directives; configtest and the listening socket; the s_client summary with the version, cipher suite, key size and verify code; the page fetched over TLS; all three failures with their codes; what is in a certificate and what revocation is; the observation table; the result.

munotes.in174

Practical 17: Web Security with SSL/TLS

Quick revision

  • TLS is the current name; SSL 2.0 and 3.0 are withdrawn. Use TLS 1.2 or 1.3.
  • TLS gives key agreement (Diffie-Hellman), authentication (a signature) and confidentiality with integrity (an authenticated cipher).
  • A certificate binds a name to a public key and is signed by a CA.
  • A root CA certificate is self-signed: subject equals issuer.
  • A signing request carries the public key and the name. The private key never leaves the server.
  • subjectAltName is where the name is read from. basicConstraints=CA:FALSE stops the key issuing certificates.
  • Apache needs four things: SSLEngine on, the certificate, the key, and a virtual host on 443.
  • openssl s_client -connect host:443 shows the version, the cipher suite and the verify code.
  • TLS 1.3 cipher suites name the cipher and the hash only; the key exchange is always ephemeral Diffie-Hellman.
  • Verify code 0 is ok, 21 is an untrusted issuer, 62 is a hostname mismatch.
  • Chain verification and name verification are two separate checks.
  • Revocation is by CRL or OCSP, and both fail open in practice, which is why certificate lifetimes keep shortening.

Questions you must be able to answer

1. What is the difference between SSL and TLS? They are the same idea under two names. SSL 2.0 and 3.0 are the old versions and are broken and withdrawn; TLS 1.2 and 1.3 are what is used. Saying SSL today means TLS.

2. Why is a root CA certificate self-signed? Because there is nobody above it to sign it. It is trusted because it is in the client's trust list, not because of anything in the certificate itself.

3. What is in a certificate signing request, and what is not? The public key, the name being requested, and a signature made with the matching private key. The private key itself is not in it and never leaves the server.

4. Why does a certificate need subjectAltName? Because that is where clients read the name from. The common name has been ignored for this purpose for years, so a certificate with only a CN is rejected.

5. What does basicConstraints=CA:FALSE prevent? It stops the certificate being used to sign other certificates. Without it, whoever holds that key could issue a valid certificate for any site.

munotes.in175

Practical 17: Web Security with SSL/TLS

6. Your server certificate verifies with -CAfile ca.crt and fails without it. Why? Because verification means checking the signature against a CA you already trust. Without the CA file, OpenSSL has no trusted issuer for it and reports that it cannot get the local issuer certificate.

7. What does the cipher suite TLS_AES_256_GCM_SHA384 tell you? That the data is protected by AES with a 256-bit key in GCM mode, which encrypts and authenticates in one operation, and that SHA-384 is used in the key schedule. It says nothing about key exchange because TLS 1.3 always uses ephemeral Diffie-Hellman.

8. A certificate verifies against the CA and the browser still refuses it. Name two reasons. It has expired, or the name in it does not match the site being visited. This chapter produces both.

9. What is revocation and why is it weak? Declaring a certificate invalid before its expiry, by a revocation list or by OCSP. It is weak because clients that cannot reach the check usually carry on, so an attacker who blocks the check defeats it.

10. A user reports a warning on your site. What do you check first? The dates, then the name, then the chain. Expiry is the commonest cause by a wide margin, and openssl s_client reports all three in one command.

Contents This chapter on its own page

munotes.in176

Chapter Twenty-Three

Practical 18: Intrusion Detection System

Syllabus topic Module 2, "Intrusion Detection System: Set up and configure an intrusion detection system (IDS) to monitor network traffic and detect potential security breaches or malicious activities."

Aim

To set up an intrusion detection system, to write and read rules, to watch it detect a real attack in real traffic, to produce a false positive and tune it away, and to record what it cannot see.

What you need to know before you start

A firewall decides what may pass. An intrusion detection system looks at what did pass and says whether any of it was an attack. It stops nothing; it tells you. The variant that also blocks is an intrusion prevention system, and the difference is one of placement: an IDS sits beside the traffic, an IPS sits in it.

NIST SP 800-94 divides them two ways, and both divisions are examinable.

By what they watch:

WatchesSeesMisses
Network-based, NIDStraffic on a linkeverything crossing that linkanything encrypted, and anything that never crosses it
Host-based, HIDSone machine's files, logs and callswhat actually happened on that hostanything on any other host

By how they decide:

Signature detection looks for a known pattern. It is exact, it explains itself, and it cannot find an attack nobody has written a rule for.

Anomaly detection learns what normal looks like and reports departures from it. It can find something new, and it produces far more false alarms, because unusual and malicious are not the same thing.

Snort, which this practical uses, is a network-based signature system.

Step 1: two hosts and some traffic to look at

$ ip netns add attacker
$ ip netns add server
$ ip link add vA type veth peer name vS
$ ip link set vA netns attacker
$ ip link set vS netns server
$ ip -n attacker addr add 10.8.0.1/24 dev vA
$ ip -n server addr add 10.8.0.2/24 dev vS
$ ip -n attacker link set vA up
$ ip -n server link set vS up
$ ip netns exec attacker ping -c 1 -W 2 10.8.0.2 2>/dev/null | tail -2
1 packets transmitted, 1 received, 0% packet loss, time 0ms
rtt min/avg/max/mdev = 0.000/0.000/0.000/0.000 ms

Now generate a mixture of traffic and record it: two pings, one request that is an attack, one that merely looks like one, and one that is the same attack in disguise.

#!/bin/bash
# Put a listener on the server so the attacker has something to talk to,
# start a capture, then send four things across the wire.
# NOTE: setsid, and a timeout on everything: a capture left running holds the
# terminal, and a capture that never fills its packet count waits for ever.
setsid timeout 25 ip netns exec server nc -l -k -p 8080 >/dev/null 2>&1 </dev/null &
sleep 1
setsid --wait timeout 18 ip netns exec server tcpdump -n -i vS -c 40 -w /root/traffic.pcap "icmp or tcp" >/dev/null 2>&1 </dev/null &
TD=$!
sleep 2
# 1. ordinary pings
ip netns exec attacker ping -c 2 -i 0.3 10.8.0.2 >/dev/null 2>&1
# 2. a directory traversal: the attack
printf 'GET /admin/../../etc/passwd HTTP/1.0\r\n\r\n' | timeout 5 ip netns exec attacker nc -w 3 10.8.0.2 8080 >/dev/null 2>&1
# 3. an ordinary page whose name happens to contain the word passwd
printf 'GET /help/reset-passwd HTTP/1.0\r\n\r\n' | timeout 5 ip netns exec attacker nc -w 3 10.8.0.2 8080 >/dev/null 2>&1
# 4. the same attack as 2, with the path written in base64
printf 'GET /YWRtaW4vLi4vLi4vZXRjL3Bhc3N3ZA== HTTP/1.0\r\n\r\n' | timeout 5 ip netns exec attacker nc -w 3 10.8.0.2 8080 >/dev/null 2>&1
sleep 1
wait $TD 2>/dev/null
munotes.in177

Practical 18: Intrusion Detection System

$ bash grab.sh
$ ls -l /root/traffic.pcap | awk '{print $9}'
/root/traffic.pcap
$ tcpdump -n -r /root/traffic.pcap 2>/dev/null | wc -l
28
$ tcpdump -n -r /root/traffic.pcap 2>/dev/null | head -2
11:38:22.600432 IP 10.8.0.1 > 10.8.0.2: ICMP echo request, id 43, seq 1, length 64
11:38:22.600512 IP 10.8.0.2 > 10.8.0.1: ICMP echo reply, id 43, seq 1, length 64

Step 2: write the rules

A Snort rule has two parts: a header saying which packets to look at, and options in brackets saying what to look for and what to call it.

alert tcp any any -> 10.8.0.2 8080 (msg:"Possible directory traversal"; content:"../"; sid:1000002; rev:1;)
  ^     ^   ^   ^   ^         ^     ^                                   ^                ^          ^
  |     |   |   |   |         |     the message the alert carries       what to look for |          revision
  |     |   |   |   |         destination port                                           rule id
  |     |   |   |   destination address
  |     |   |   direction: -> one way, <> both
  |     |   source address and port; `any` matches all
  |     protocol: tcp, udp, icmp or ip
  action: alert, log, pass, drop (drop only in inline mode)

Every field in that diagram is a likely question. Two more that matter: sid must be unique, and numbers from 1,000,000 upward are reserved for rules you write yourself; rev is the revision, raised whenever you change a rule so that logs can be matched to the version that produced them.

$ printf 'alert icmp any any -> 10.8.0.2 any (msg:"ICMP ping to the server"; itype:8; sid:1000001; rev:1;)\nalert tcp any any -> 10.8.0.2 8080 (msg:"Possible directory traversal"; content:"../"; sid:1000002; rev:1;)\nalert tcp any any -> 10.8.0.2 8080 (msg:"Password file mentioned"; content:"passwd"; sid:1000003; rev:1;)\n' > local.rules
$ cat local.rules
alert icmp any any -> 10.8.0.2 any (msg:"ICMP ping to the server"; itype:8; sid:1000001; rev:1;)
alert tcp any any -> 10.8.0.2 8080 (msg:"Possible directory traversal"; content:"../"; sid:1000002; rev:1;)
alert tcp any any -> 10.8.0.2 8080 (msg:"Password file mentioned"; content:"passwd"; sid:1000003; rev:1;)
munotes.in178

Practical 18: Intrusion Detection System

itype:8 is the ICMP type for an echo request, so the first rule sees pings going out and not the replies coming back.

Step 3: run it, and read the alerts

$ snort -q -A console -c local.rules -r /root/traffic.pcap -k none 2>&1 | sed 's/^[0-9/:.-]* *//'
[**] [1:1000001:1] ICMP ping to the server [**] [Priority: 0] {ICMP} 10.8.0.1 -> 10.8.0.2
[**] [1:1000001:1] ICMP ping to the server [**] [Priority: 0] {ICMP} 10.8.0.1 -> 10.8.0.2
[**] [1:1000003:1] Password file mentioned [**] [Priority: 0] {TCP} 10.8.0.1:33074 -> 10.8.0.2:8080
[**] [1:1000002:1] Possible directory traversal [**] [Priority: 0] {TCP} 10.8.0.1:33074 -> 10.8.0.2:8080
[**] [1:1000003:1] Password file mentioned [**] [Priority: 0] {TCP} 10.8.0.1:33082 -> 10.8.0.2:8080

Five alerts, and every one of them tells a story.

Two ICMP alerts, one per ping. The rule matched the echo requests and not the replies, exactly as itype:8 says it should.

One directory traversal alert, sid 1000002, on the connection carrying /admin/../../etc/passwd. That is the detection the practical is for: a real attack pattern found in real traffic by a rule you wrote.

Two "Password file mentioned" alerts, sid 1000003, and there is the problem. One is the attack. The other is the request for /help/reset-passwd, which is an ordinary page. The rule looks for the six letters passwd anywhere in the packet, and a legitimate URL contains them.

That second one is a false positive, and it is the commonest failure of a signature system. An analyst who sees a hundred of those a day stops reading them, and then the real one goes past unread. NIST SP 800-94 is blunt about this: tuning is not an optional refinement, it is what makes the system usable.

And one thing is missing from the list. The fourth request, the same attack with its path written in base64, raised nothing at all.

Step 4: tune the rule

The fix for the false positive is to make the pattern more specific: not the word, the whole path.

$ printf 'alert tcp any any -> 10.8.0.2 8080 (msg:"Password file access"; content:"/etc/passwd"; sid:1000004; rev:1;)\n' > tuned.rules
$ snort -q -A console -c tuned.rules -r /root/traffic.pcap -k none 2>&1 | sed 's/^[0-9/:.-]* *//'
[**] [1:1000004:1] Password file access [**] [Priority: 0] {TCP} 10.8.0.1:33074 -> 10.8.0.2:8080
$ snort -q -A console -c local.rules -r /root/traffic.pcap -k none 2>&1 | grep -c 1000003
2
$ snort -q -A console -c tuned.rules -r /root/traffic.pcap -k none 2>&1 | grep -c 1000004
1

Two alerts became one, and the one that remains is the attack. That is the whole of tuning, and the measurement, two against one, is what belongs in the observation table.

Note what tuning costs. The new rule would miss an attacker who wrote /etc/./passwd, or /etc//passwd, or who encoded a character. A narrower rule has fewer false positives and more false negatives, and choosing between them is a judgement about the site, not a fact about the rule.

munotes.in179

Practical 18: Intrusion Detection System

Step 5: what an IDS cannot see

Three limits, and the chapter has already demonstrated two of them.

Anything nobody has written a rule for. Signature detection finds known patterns. An attack nobody has seen goes past.

Anything transformed. The base64 request in step 3 is byte for byte a different string and raised nothing, and it is the same attack. Real evasion uses URL encoding, case changes, inserted null bytes and fragmentation, and it is why Snort has preprocessors that normalise traffic before the rules see it. A rule written against raw bytes with no normalisation is a rule that is trivially avoided.

Anything encrypted. This is the big one, and Practical 17 is the reason. A rule matching content:"/etc/passwd" sees that string only if it is on the wire in the clear. Send the same request over HTTPS and the IDS sees a TLS record and nothing else.

$ printf 'alert tcp any any -> any 443 (msg:"Traffic to a TLS port"; sid:1000005; rev:1;)\n' > tls.rules
$ snort -q -A console -c tls.rules -r /root/traffic.pcap -k none 2>&1 | sed 's/^[0-9/:.-]* *//' | wc -l
0

So the honest position, and the one to write in the journal: as more of the web moved to TLS, a network IDS lost most of its view. The answers used in practice are to terminate TLS at a proxy and inspect there, which raises its own problems, or to move the detection onto the host, where the data is in the clear because that is where it is used.

Step 6: clean up

$ ip netns delete attacker
$ ip netns delete server
$ ip netns list | wc -l
0

Procedure

  1. Create two namespaces joined by a veth pair and confirm they can reach each other.
  2. Put a listener on the server and capture traffic while sending: two pings, a directory traversal, a legitimate page whose name contains a suspicious word, and the same attack encoded.
  3. Write three rules: one ICMP, one for the traversal pattern, one for a word.
  4. Run Snort over the capture and read every alert, naming the rule that produced it.
  5. Identify the false positive and say which request caused it.
  6. Write a tuned rule, run it, and count the alerts before and after.
  7. Confirm that the encoded request raised nothing, and say why.
  8. Say what an IDS cannot see, and what is done about it.
  9. Delete the namespaces.

Observations

MeasuredValue
Hosts10.8.0.1 attacker, 10.8.0.2 server, two namespaces
Traffic generated2 pings, 3 HTTP requests
Rules written3, sids 1000001 to 1000003
Total alerts from those rules5
ICMP alerts2, one per echo request, replies not matched
Directory traversal alerts1, on the request containing ../
"Password file mentioned" alerts2, of which one is a false positive
The false positivethe request for /help/reset-passwd
Alerts from the tuned rule1
The base64-encoded attackno alert at all
Traffic to port 443 in the capturenone
munotes.in180

Practical 18: Intrusion Detection System

Result

Snort 2.9.20 was set up as a network intrusion detection system over a link between two hosts created as network namespaces. Three rules were written and read field by field, and run against a capture of five exchanges. They produced five alerts: two on the ICMP echo requests, one on a directory traversal in an HTTP request, and two on the word passwd, of which one was raised by a legitimate request for a page called /help/reset-passwd and was therefore a false positive. A tuned rule matching the whole path /etc/passwd instead of the word reduced those two alerts to one, the real attack. The same attack sent with its path written in base64 raised no alert at all, which demonstrates that a signature matching raw bytes is defeated by a trivial transformation, and the chapter records that an IDS sees nothing inside encrypted traffic.

Where marks are lost

Confusing detection with prevention. An IDS reports. An IPS blocks, and sits in the path rather than beside it.

A sid below 1,000,000. Those are reserved for the distributed rule sets. Use 1000001 upward for your own.

Not raising rev when you change a rule. Then a log cannot be matched to the rule that produced it.

Reporting only that a rule fired. An examiner wants the false positive too, and the tuning that removed it.

Writing a rule so broad that it fires on ordinary traffic. Two alerts and one of them wrong, as here, is a system that will be ignored within a week.

Writing a rule so narrow that a space defeats it. Tuning trades false positives for false negatives. Say which way you traded.

Claiming an IDS sees encrypted traffic. It sees that a TLS connection happened and nothing inside it.

Testing on traffic you made and nothing else. Say plainly that the capture is your own, on your own two hosts.

For the journal

Aim; IDS against IPS, and network-based against host-based, and signature against anomaly, as three short comparisons; the two hosts and how they were made; the capture script and what each of the four exchanges is; the annotated rule with every field named; the three rules; the five alerts with the rule that produced each; the false positive identified by the request that caused it; the tuned rule and the count before and after; the encoded request that raised nothing; the note on encrypted traffic; the clean-up; the observation table; the result.

munotes.in181

Practical 18: Intrusion Detection System

Quick revision

  • An IDS detects and reports; an IPS sits in the path and blocks.
  • Network-based watches a link; host-based watches one machine.
  • Signature detection finds known patterns; anomaly detection finds departures from normal and raises more false alarms.
  • A Snort rule is a header and options: action, protocol, source, direction, destination, then msg, content, sid, rev.
  • sid above 1,000,000 for your own rules; raise rev on every change.
  • itype:8 is an ICMP echo request.
  • A false positive is an alert on ordinary traffic. Here, content:"passwd" matched /help/reset-passwd.
  • Tuning narrows the pattern: two alerts became one, and the one left is the attack.
  • Narrower rules mean fewer false positives and more false negatives.
  • The same attack in base64 raised nothing. Normalise before matching, or be evaded.
  • An IDS cannot see inside TLS. Terminate at a proxy, or detect on the host.

Questions you must be able to answer

1. What is the difference between an IDS and an IPS? An IDS watches a copy of the traffic and reports what it finds. An IPS sits in the path of the traffic and can drop it. The detection is the same; the placement and the authority differ.

2. What are the two ways of deciding that something is an attack? Signature detection, which matches known patterns and is exact but blind to anything new; and anomaly detection, which learns normal behaviour and reports departures, which can find new attacks and raises far more false alarms.

3. Read out the parts of the rule you wrote. Action alert, protocol tcp, source any address and any port, direction one way, destination 10.8.0.2 port 8080, then the options: the message, the content to match, a unique sid and a revision.

4. Why must sid be above 1,000,000? Because lower numbers are reserved for the published rule sets, and a collision means two different rules with the same identity in your logs.

5. Your rule fired on /help/reset-passwd. What is that called and why did it happen? A false positive. The rule matched the six letters passwd anywhere in the packet, and an ordinary page name contains them.

6. How did you tune it, and what did it cost? By matching the whole path /etc/passwd instead of the word, which took the alerts from two to one. It costs coverage: the new rule misses /etc/./passwd or an encoded form of the same path.

7. The same attack in base64 raised no alert. Is that a bug? No. The rule matches bytes, and base64 is different bytes. It is why an IDS normalises traffic with preprocessors before the rules see it, and why a rule against raw bytes is easy to evade.

munotes.in182

Practical 18: Intrusion Detection System

8. What does an IDS see of an HTTPS session? That a TLS connection was made, to which address, when, and how much passed. Nothing of the contents.

9. So what is done about encryption? Either terminate TLS at a proxy and inspect there, which means the proxy holds everybody's traffic in the clear, or move the detection onto the host, where the data is already decrypted.

10. You get a hundred alerts a day and one is real. What is the problem and what is the fix? The problem is that nobody will read the hundredth one. The fix is tuning: narrow the rules that fire on ordinary traffic, measure the count before and after, and accept that some coverage is traded for it.

Contents This chapter on its own page

munotes.in183

Chapter Twenty-Four

Practical 19: Malware Analysis and Detection

Syllabus topic Module 2, "Malware Analysis and Detection: Analyze and identify malware samples using antivirus tools, analyze their behavior, and develop countermeasures to mitigate their impact."

Aim

To analyse a suspicious file, to write signatures that identify it, to detect it with an antivirus tool, and to state the countermeasures that reduce the damage malware can do.

What you need to know before you start

Malware is any program written to do something its user did not want. The categories a paper can ask for:

What it is
Virusattaches itself to another program and spreads when that program runs
Wormspreads by itself across a network, needing nothing to be run
Trojanpretends to be something wanted, and does something else as well
Ransomwareencrypts the user's files and demands payment for the key
Spyware, keyloggerwatches and reports
Rootkithides itself and other malware from the operating system
Botnet clienttakes orders from a remote controller

Analysis is done two ways, and the distinction is the first question an examiner asks.

Static analysis examines the file without running it: its type, its size, its hash, the printable strings inside it, the libraries it imports. It is completely safe, and it is defeated by packing and obfuscation.

Dynamic analysis runs the file and watches what it does: which files it opens, which addresses it connects to, what it writes to the registry or to /etc. It sees through obfuscation, and it must be done in a sandbox that is isolated and discarded afterwards, because you are running the thing.

This chapter does static analysis only, and says so plainly: there is nothing to run, because the sample is a text file.

A note about the EICAR test file

The standard harmless test for an antivirus is the EICAR test file, a 68-character printable string that every scanner agrees to report as a virus although it does nothing at all. It exists so that you can test a scanner without touching malware.

This chapter names it and does not print it, for a reason worth understanding: any file containing that string is by design detected as a virus, and that includes a web page, a PDF and a textbook. Printing it here would mean this page was quarantined on the machine of every reader whose antivirus scans downloads. You can obtain it from eicar.org, and you should, once, on a machine you own, just to watch your own scanner react.

The sample below does the same job with a marker of our own invention, which no scanner knows and which therefore cannot upset anything.

Step 1: the sample, and static analysis before any scanner

$ mkdir -p /root/lab
$ cd /root/lab
$ printf 'This is a harmless sample written for Practical 19.\nIt pretends to be a downloader: mnDownloadAndRun("http://example.invalid/x")\n' > sample.txt
$ printf 'This is an ordinary note about the practical.\n' > notes.txt
$ ls -l sample.txt notes.txt | awk '{print $5, $9}'
46 notes.txt
129 sample.txt
$ file sample.txt
sample.txt: ASCII text
$ sha256sum sample.txt notes.txt
9248dab50c68466d00d899fe899c420cb26f54cc5ca6a7cbd70e35b4ec6c2e6a  sample.txt
b20c0f883a07c78c7a4a057d9ed6c5b412dcb999d17710bfaf14d4cef797e22c  notes.txt
$ grep -aoE '[ -~]{8,}' sample.txt
This is a harmless sample written for Practical 19.
It pretends to be a downloader: mnDownloadAndRun("http://example.invalid/x")
munotes.in184

Practical 19: Malware Analysis and Detection

Four things, and in a real investigation you would do exactly these four first.

file reads the first few bytes and says what kind of thing this is. A file called invoice.pdf that file calls an ELF executable has already told you most of what you need to know.

sha256sum gives the file's identity. You send that, not the file, to a colleague or to a lookup service, because a hash reveals nothing and travels safely.

grep -aoE '[ -~]{8,}' pulls out the runs of printable characters. On a real sample this is where the URLs, the file paths, the registry keys and the ransom note come from, and it is the single most productive minute of a static analysis. Here it shows the marker mnDownloadAndRun and a URL, which are the two things that make this file look like a downloader.

ls -l gives the size, which matters because a hash signature includes it.

Step 2: write a hash signature and detect the file

A signature is a rule an antivirus matches files against. The simplest kind names a hash.

$ mkdir -p /root/sigs
$ sigtool --md5 sample.txt > /root/sigs/mn.hdb
$ cut -d: -f1 /root/sigs/mn.hdb | wc -c
33
$ sed -i 's/[^:]*$/Practical19.Sample.Hash/' /root/sigs/mn.hdb
$ cut -d: -f3 /root/sigs/mn.hdb
Practical19.Sample.Hash
$ clamscan -d /root/sigs/mn.hdb sample.txt notes.txt 2>&1 | grep -vE "^$|Engine version|Time:|Start Date|End Date|Data scanned|Data read"
sample.txt: Practical19.Sample.Hash.UNOFFICIAL FOUND
notes.txt: OK
----------- SCAN SUMMARY -----------
Known viruses: 1
Scanned directories: 0
Scanned files: 2
Infected files: 1

The format of an .hdb line is three fields separated by colons: the MD5, the size in bytes, and the name to report. sigtool --md5 produces the first two and puts the file name in the third; editing that third field is how you give the detection a sensible name.

sample.txt: Practical19.Sample.Hash.UNOFFICIAL FOUND and notes.txt: OK. The scanner found the one file and left the other alone, which is the whole of "identify malware samples using antivirus tools".

UNOFFICIAL is ClamAV telling you the signature is not from its own database. It is not a warning; it is a fact about where the rule came from.

And note what is not here: freshclam. A real installation downloads ClamAV's database, a quarter of a gigabyte that changes several times a day. This chapter uses only signatures it wrote itself, so that every result is reproducible and the exercise is about the signature rather than about somebody else's list. On your own machine, run sudo freshclam first.

munotes.in185

Practical 19: Malware Analysis and Detection

Step 3: why a hash signature is not enough

$ printf 'a harmless extra line\n' >> sample.txt
$ sha256sum sample.txt
5cb74c8f67724e7ffe772e60065b2c76637ada7a7052f815045583846a3c29f5  sample.txt
$ clamscan -d /root/sigs/mn.hdb sample.txt 2>&1 | grep -E "Infected files"
Infected files: 0

One added line and the detection is gone. The hash of a file is the identity of that exact sequence of bytes, so a single character anywhere makes it a different file as far as the signature is concerned. An attacker who changes one byte and redistributes has defeated every hash signature in the world, and it costs them nothing.

So real signatures match content, not the whole file.

$ printf 'Practical19.Sample.Content:0:*:%s\n' "$(printf 'mnDownloadAndRun' | xxd -p)" > /root/sigs/mn.ndb
$ cat /root/sigs/mn.ndb
Practical19.Sample.Content:0:*:6d6e446f776e6c6f6164416e6452756e
$ clamscan -d /root/sigs/mn.ndb sample.txt notes.txt 2>&1 | grep -vE "^$|Engine version|Time:|Start Date|End Date|Data scanned|Data read"
sample.txt: Practical19.Sample.Content.UNOFFICIAL FOUND
notes.txt: OK
----------- SCAN SUMMARY -----------
Known viruses: 1
Scanned directories: 0
Scanned files: 2
Infected files: 1

An .ndb line is four fields: the name, the target type, the offset (* means anywhere), and the pattern in hexadecimal. This one matches the sixteen bytes that spell mnDownloadAndRun wherever they appear.

It still finds the edited file, and it still leaves the ordinary one alone. That is the difference between identifying a file and identifying a behaviour, and it is the same lesson as Practical 18's tuning: the narrower the pattern, the more easily it is evaded; the broader it is, the more innocent files it catches.

Step 4: the same trade-off, measured

$ printf 'Practical19.TooBroad:0:*:%s\n' "$(printf 'Download' | xxd -p)" > /root/sigs/broad.ndb
$ printf 'This program lets you Download your timetable.\n' > timetable.txt
$ clamscan -d /root/sigs/broad.ndb sample.txt notes.txt timetable.txt 2>&1 | grep -E "FOUND|OK$|Infected files"
sample.txt: Practical19.TooBroad.UNOFFICIAL FOUND
notes.txt: OK
timetable.txt: Practical19.TooBroad.UNOFFICIAL FOUND
Infected files: 2
$ clamscan -d /root/sigs/mn.ndb sample.txt notes.txt timetable.txt 2>&1 | grep -E "FOUND|OK$|Infected files"
sample.txt: Practical19.Sample.Content.UNOFFICIAL FOUND
notes.txt: OK
timetable.txt: OK
Infected files: 1

The broad signature flags two files; the specific one flags one. The second file it flagged is an innocent note about a timetable. A signature that matches the word Download matches every honest program that downloads anything, and an antivirus that does that gets switched off.

Step 5: the countermeasures MU asks for

MU's third clause is "develop countermeasures to mitigate their impact", and the honest answer has two halves: what stops it arriving, and what limits the damage when it does. NIST SP 800-83 organises them as prevention, detection and response.

Before it arrives.

  • Patch. Most malware uses a flaw that already has a fix. Applying updates is the single most effective measure there is.
  • Least privilege. A program run by an ordinary account cannot change the system. Do not work as root or as an administrator.
  • Do not execute what you did not ask for. Attachments, macros in documents, and scripts from a web page.
  • Filter at the boundary. Mail and web gateways, and the firewall of Practical 20.
  • Signed software from its own source. Practical 14 is why that works.
munotes.in186

Practical 19: Malware Analysis and Detection

When it arrives anyway.

  • Antivirus with an up-to-date database, and updates matter more than the product: this chapter's step 3 shows why yesterday's signatures miss today's variant.
  • Backups that are offline or immutable. This is the entire answer to ransomware, and nothing else is. A backup a running machine can write to is a backup ransomware encrypts.
  • Segment the network. A worm reaches what it can route to. Practical 20 is how you shorten that list.
  • Logging and detection, which is Practical 18.

After.

  • Isolate first. Pull the network cable before anything else.
  • Preserve the evidence: the hash, the file, the logs, before wiping.
  • Rebuild rather than clean. A machine that has run a rootkit cannot be trusted again; reinstall it.
  • Find out how it got in, or it comes back the same way.

Step 6: clean up

$ rm -f /root/lab/sample.txt /root/lab/notes.txt /root/lab/timetable.txt
$ rm -f /root/sigs/mn.hdb /root/sigs/mn.ndb /root/sigs/broad.ndb
$ ls /root/lab /root/sigs | wc -l
3

Procedure

  1. Create a harmless sample carrying a marker string of your own, and an ordinary file to compare it with. Never use real malware.
  2. Do the static analysis first: file, sha256sum, ls -l, and the printable strings.
  3. Build a hash signature with sigtool --md5, rename the detection, and scan both files.
  4. Record that the sample is found and the ordinary file is not.
  5. Add one line to the sample, take its hash again, and scan again. Record that the detection is gone.
  6. Write a content signature in an .ndb file with the pattern in hexadecimal, and scan again.
  7. Write a deliberately broad signature, add an innocent file that contains its pattern, and count the detections from the broad and the specific signature.
  8. Write the countermeasures under prevention, detection and response.
  9. Remove the sample and the signatures.

Observations

MeasuredValue
Samplea text file of our own, 129 bytes, carrying the marker mnDownloadAndRun
file saysASCII text
Strings foundthe marker and a URL at example.invalid
Hash signature formatMD5, size, name, colon separated, in an .hdb file
Scan with the hash signaturesample FOUND, ordinary file OK
After adding one linethe hash changed and the hash signature found nothing
Content signature formatname, target, offset, hexadecimal pattern, in an .ndb file
Scan with the content signaturethe edited sample still FOUND, ordinary file still OK
Broad signature on three files2 infected
Specific signature on the same three1 infected
Real malware usednone
munotes.in187

Practical 19: Malware Analysis and Detection

Result

A harmless sample file carrying a marker of our own was analysed statically with file, sha256sum and a strings extraction, which identified the marker and a URL. A hash signature was built with sigtool --md5, given a detection name, and used to identify the sample while leaving an ordinary file alone. Appending one line changed the file's hash and the hash signature then found nothing, which demonstrates that a hash identifies a file and not a behaviour. A content signature matching the marker in hexadecimal found the edited sample and still left the ordinary file alone. A deliberately broad signature was then shown to flag two of three files, one of them innocent, against one of three for the specific signature, which measures the same false-positive trade-off Practical 18 records for intrusion rules. Countermeasures were recorded under prevention, detection and response. No real malware was used at any point.

Where marks are lost

Bringing real malware into a laboratory. There is no reason to and several reasons not to. A marker of your own, or the EICAR test file from eicar.org, does everything the exercise needs.

Reproducing the EICAR string in a document. The document then gets quarantined, which is the string working as designed.

Running the sample. This is static analysis. Dynamic analysis needs an isolated sandbox that is discarded afterwards, and it is not what this practical asks for.

Relying on a hash signature. One added byte defeats it, and the chapter measures that.

A signature so broad it catches ordinary files. Two detections of three files, one of them innocent, is an antivirus that gets turned off.

Scanning without updating. On a real machine, freshclam first. This chapter skips it on purpose and says why.

Offering "install an antivirus" as the countermeasure. Patching, least privilege and offline backups do more, and ransomware is answered by backups and by nothing else.

Cleaning a machine that has run a rootkit. Rebuild it. A compromised system cannot tell you the truth about itself.

For the journal

Aim; the categories of malware; static against dynamic analysis, and which one this practical is; a line saying no real malware was used and why the EICAR string is named and not printed; the sample and the ordinary file; the four static-analysis commands and what each told you; the hash signature with its three fields and the scan result; the edited file, its new hash, and the detection that vanished; the content signature with its four fields and the scan that still finds it; the broad-signature count against the specific one; the countermeasures under prevention, detection and response; the clean-up; the observation table; the result.

munotes.in188

Practical 19: Malware Analysis and Detection

Quick revision

  • Static analysis examines the file; dynamic analysis runs it in a sandbox.
  • file, sha256sum, ls -l and the printable strings are the first four commands.
  • Send the hash, never the sample.
  • An .hdb signature is MD5, size, name. sigtool --md5 writes the first two.
  • An .ndb signature is name, target, offset, hexadecimal pattern. * for anywhere.
  • A hash signature is defeated by one changed byte. A content signature is not.
  • A content signature that is too broad flags innocent files. Measure it before shipping it.
  • UNOFFICIAL in ClamAV's output means the signature is your own, not from its database.
  • freshclam updates the real database. This chapter does not use it, deliberately.
  • Prevention: patch, least privilege, do not execute what arrived unasked, filter, signed software.
  • Response: isolate, preserve evidence, rebuild rather than clean, find the way in.
  • Offline or immutable backups are the only real answer to ransomware.

Questions you must be able to answer

1. What is the difference between static and dynamic analysis? Static analysis examines the file without running it: type, size, hash, strings, imports. Dynamic analysis runs it and watches what it does, which sees through obfuscation but must be done in an isolated sandbox.

2. Which did you do, and why? Static, because the sample is a text file with nothing to run, and because running an unknown program is not something a practical should require.

3. What is the EICAR test file, and why is it not printed in this chapter? A short printable string that every antivirus agrees to report as a virus although it is harmless, so that a scanner can be tested safely. It is not printed because any file containing it is detected as a virus, including this page.

4. What are the three fields of a ClamAV hash signature? The MD5 of the file, its size in bytes, and the name to report when it matches.

5. You added one line and the detection disappeared. Explain. A hash signature identifies one exact sequence of bytes. Any change gives a different hash, so the signature no longer matches, and an attacker gets that for free.

6. What does a content signature match instead? A pattern of bytes, given in hexadecimal, at a stated offset or anywhere in the file. It survives changes elsewhere in the file because it does not depend on the whole of it.

7. Your broad signature flagged an innocent file. What is the general rule? The broader the pattern, the more false positives; the narrower it is, the easier it is to evade. The same trade-off as tuning an intrusion rule, and it has to be measured rather than guessed.

munotes.in189

Practical 19: Malware Analysis and Detection

8. What does UNOFFICIAL mean in ClamAV's output? That the signature that matched came from a database you supplied rather than from ClamAV's own. It says nothing about whether the detection is right.

9. What is the only real countermeasure against ransomware? Backups that the running machine cannot write to, whether offline or immutable. Everything else reduces the chance of infection; only a backup restores the files.

10. A machine in the laboratory is infected. What do you do first, and what do you do last? First isolate it from the network. Last, find out how it got in; and between those, preserve the evidence and rebuild the machine rather than trying to clean it.

Contents This chapter on its own page

munotes.in190

Chapter Twenty-Five

Practical 20: Firewall Configuration and Rule-Based Filtering

Syllabus topic Module 2, "Firewall Configuration and Rule-based Filtering: Configure and test firewall rules to control network traffic, filter packets based on specified criteria, and protect network resources from unauthorized access."

Aim

To configure a firewall, to filter packets on stated criteria, to test each rule by trying the traffic it is meant to control, and to show what a wrong rule order does.

What you need to know before you start

A packet filter looks at each packet and decides accept, drop or reject. Linux's is in the kernel; iptables and nft are the two commands that talk to it.

Three chains matter, and which one a packet meets depends on where it is going:

ChainPacketsExample
INPUTaddressed to this machinesomebody connecting to your web server
OUTPUTsent by this machineyour machine fetching an update
FORWARDpassing through, on a routertraffic between two networks

Each chain has a policy, the verdict for a packet that matches no rule, and a list of rules tried in order, first match wins. That ordering is the single commonest source of a broken firewall and this chapter measures it.

Three verdicts, and telling the last two apart is a standard question:

  • ACCEPT: let it through.
  • DROP: discard it silently. The sender learns nothing and waits until it gives up.
  • REJECT: discard it and send back an error. The sender is told at once.

And one design rule from NIST SP 800-41: default deny. Permit what is needed and drop the rest, rather than listing what to block, because you cannot list everything.

Step 1: two hosts, two services, nothing filtered

$ ip netns add client
$ ip netns add server
$ ip link add vC type veth peer name vS
$ ip link set vC netns client
$ ip link set vS netns server
$ ip -n client addr add 10.7.0.1/24 dev vC
$ ip -n server addr add 10.7.0.2/24 dev vS
$ ip -n client link set vC up
$ ip -n server link set vS up
$ ip -n server link set lo up
$ ip netns exec server iptables -L -n
Chain INPUT (policy ACCEPT)
target     prot opt source               destination

Chain FORWARD (policy ACCEPT)
target     prot opt source               destination

Chain OUTPUT (policy ACCEPT)
target     prot opt source               destination

Three empty chains, each with the policy ACCEPT. That is a machine with no firewall, and it is what you start from.

Put two services on the server and check both are reachable.

#!/bin/bash
# Two listeners on the server, each in a session of its own so that neither
# holds the terminal, and each with a time limit so nothing outlives the
# practical.
setsid timeout 90 ip netns exec server nc -l -k -p 8080 >/dev/null 2>&1 </dev/null &
setsid timeout 90 ip netns exec server nc -l -k -p 9090 >/dev/null 2>&1 </dev/null &
sleep 1
echo "two listeners started"
munotes.in191

Practical 20: Firewall Configuration and Rule-Based Filtering

$ bash services.sh
two listeners started
$ ip netns exec client timeout 4 nc -z -w 3 10.7.0.2 8080 >/dev/null 2>&1 && echo "8080 open" || echo "8080 closed"
8080 open
$ ip netns exec client timeout 4 nc -z -w 3 10.7.0.2 9090 >/dev/null 2>&1 && echo "9090 open" || echo "9090 closed"
9090 open
$ ip netns exec client ping -c 1 -W 2 10.7.0.2 >/dev/null 2>&1 && echo "ping works" || echo "ping blocked"
ping works

Both ports open and the machine answers a ping. That is the state to improve on, and recording it first is what makes the next section a result rather than an assertion.

Step 2: four rules, and a test after each

The policy is to allow one service and nothing else. Four rules, added in the order they must be tried.

$ ip netns exec server iptables -A INPUT -i lo -j ACCEPT
$ ip netns exec server iptables -A INPUT -m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT
$ ip netns exec server iptables -A INPUT -p tcp --dport 8080 -j ACCEPT
$ ip netns exec server iptables -A INPUT -j DROP
$ ip netns exec server iptables -L INPUT -n --line-numbers
Chain INPUT (policy ACCEPT)
num  target     prot opt source               destination
1    ACCEPT     0    --  0.0.0.0/0            0.0.0.0/0
2    ACCEPT     0    --  0.0.0.0/0            0.0.0.0/0            ctstate RELATED,ESTABLISHED
3    ACCEPT     6    --  0.0.0.0/0            0.0.0.0/0            tcp dpt:8080
4    DROP       0    --  0.0.0.0/0            0.0.0.0/0

Each of the four earns its place, and an examiner will ask about the second.

-i lo -j ACCEPT. Traffic a machine sends to itself goes over the loopback interface, and a great deal of software depends on it. Drop it and things break in ways that look nothing like a firewall problem.

-m conntrack --ctstate ESTABLISHED,RELATED -j ACCEPT. This is what makes the firewall stateful. When the server itself makes a connection outward, the replies coming back are not new connections; they belong to one that is already open. Without this rule every reply to every request the machine makes is dropped by the last rule, and the machine cannot fetch anything. RELATED covers a packet that belongs to an existing connection without being part of its data stream, such as an ICMP error.

-p tcp --dport 8080 -j ACCEPT. The one service that is allowed. This is MU's "filter packets based on specified criteria": the criteria here are the protocol and the destination port, and -s, -d, -i and -o add source, destination and interface.

-j DROP with nothing else. Everything not already accepted. This is the default deny, written as a last rule rather than as a policy so that it is visible in the listing.

munotes.in192

Practical 20: Firewall Configuration and Rule-Based Filtering

Now test all three things again.

$ ip netns exec client timeout 4 nc -z -w 3 10.7.0.2 8080 >/dev/null 2>&1 && echo "8080 open" || echo "8080 closed"
8080 open
$ ip netns exec client timeout 4 nc -z -w 3 10.7.0.2 9090 >/dev/null 2>&1 && echo "9090 open" || echo "9090 closed"
9090 closed
$ ip netns exec client ping -c 1 -W 2 10.7.0.2 >/dev/null 2>&1 && echo "ping works" || echo "ping blocked"
ping blocked

8080 open, 9090 closed, ping blocked. Three tests, three results, and together they are the evidence that the rules do what they say. A firewall practical with rules and no tests has shown nothing.

Step 3: DROP against REJECT, measured

$ ip netns exec server iptables -I INPUT 4 -p tcp --dport 9090 -j REJECT --reject-with tcp-reset
$ ip netns exec client timeout 5 nc -w 3 -v 10.7.0.2 9090 < /dev/null 2>&1 | tail -1
nc: connect to 10.7.0.2 port 9090 (tcp) failed: Connection refused
$ ip netns exec server iptables -D INPUT 4
$ ip netns exec client timeout 6 nc -w 4 -v 10.7.0.2 9090 < /dev/null 2>&1 | tail -1
nc: connect to 10.7.0.2 port 9090 (tcp) timed out: Operation now in progress

REJECT: Connection refused, at once. DROP: timed out, after the client gives up waiting.

That difference is the whole of the choice between them, and it cuts both ways. REJECT is polite: the client stops immediately instead of retrying, and a user gets a useful error instead of a hang. DROP is quiet: a scanner learns nothing, not even that the address is in use, and has to wait out a timeout on every port, which makes scanning slow.

The usual practice, and the sentence for the journal: REJECT inside your own network, DROP at the boundary.

Step 4: order, which is where firewalls actually go wrong

Rules are tried from the top and the first match decides. Insert a DROP above an ACCEPT and the ACCEPT can never be reached.

$ ip netns exec server iptables -I INPUT 1 -p tcp --dport 8080 -j DROP
$ ip netns exec server iptables -L INPUT -n --line-numbers | head -6
Chain INPUT (policy ACCEPT)
num  target     prot opt source               destination
1    DROP       6    --  0.0.0.0/0            0.0.0.0/0            tcp dpt:8080
2    ACCEPT     0    --  0.0.0.0/0            0.0.0.0/0
3    ACCEPT     0    --  0.0.0.0/0            0.0.0.0/0            ctstate RELATED,ESTABLISHED
4    ACCEPT     6    --  0.0.0.0/0            0.0.0.0/0            tcp dpt:8080
$ ip netns exec client timeout 4 nc -z -w 3 10.7.0.2 8080 >/dev/null 2>&1 && echo "8080 open" || echo "8080 closed"
8080 closed
$ ip netns exec server iptables -D INPUT 1
$ ip netns exec client timeout 4 nc -z -w 3 10.7.0.2 8080 >/dev/null 2>&1 && echo "8080 open" || echo "8080 closed"
8080 open
munotes.in193

Practical 20: Firewall Configuration and Rule-Based Filtering

The listing shows a DROP for port 8080 at line 1 and an ACCEPT for port 8080 still sitting at line 2, and the port is closed. The ACCEPT is present, correct, and never consulted. Reading the rule list and seeing your rule there proves nothing; only reaching it matters.

-A appends to the end, -I inserts at the top or at a numbered position, and -D deletes. A rule added with -A after the final DROP is a rule that will never fire, and that is the commonest firewall mistake there is.

Step 5: the counters, which are how you tell what is happening

$ ip netns exec server iptables -L INPUT -n -v
Chain INPUT (policy ACCEPT 0 packets, 0 bytes)
 pkts bytes target     prot opt in     out     source               destination
    0     0 ACCEPT     0    --  lo     *       0.0.0.0/0            0.0.0.0/0
    6   312 ACCEPT     0    --  *      *       0.0.0.0/0            0.0.0.0/0            ctstate RELATED,ESTABLISHED
    2   120 ACCEPT     6    --  *      *       0.0.0.0/0            0.0.0.0/0            tcp dpt:8080
    8   504 DROP       0    --  *      *       0.0.0.0/0            0.0.0.0/0

Every rule carries a packet count and a byte count. A rule with zero packets has never matched, and that is the first thing to look at when a firewall is not doing what you expect: either the traffic is not arriving, or an earlier rule is taking it.

Step 6: logging a dropped packet

$ sudo iptables -I INPUT 1 -p tcp --dport 9090 -j LOG --log-prefix "MN-FIREWALL-DROP: " --log-level 4
$ sudo dmesg | grep MN-FIREWALL-DROP | tail -1
MN-FIREWALL-DROP: IN=eth0 OUT= SRC=10.7.0.1 DST=10.7.0.2 PROTO=TCP SPT=41234 DPT=9090

This block is marked as not run, and the reason belongs in the journal. The LOG target writes to the kernel's own message buffer, and a container shares that buffer with the machine hosting it and is not allowed to read it, so dmesg inside this laboratory returns nothing. On a real Ubuntu machine the two commands above work and the line appears; run them there.

Two things about LOG that are examinable whether or not you can run it here. LOG is not a verdict: a packet that matches a LOG rule carries on to the next rule, which is why the LOG rule goes above the DROP and not instead of it. And a LOG rule on busy traffic fills a disk, so real ones carry -m limit --limit 5/min.

The counters of step 5 are the evidence this laboratory can produce, and they are enough: they say how many packets each rule took.

munotes.in194

Practical 20: Firewall Configuration and Rule-Based Filtering

Step 7: saving the rules, and the same policy in nftables

$ ip netns exec server iptables-save | grep "^-A INPUT"
-A INPUT -i lo -j ACCEPT
-A INPUT -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A INPUT -p tcp -m tcp --dport 8080 -j ACCEPT
-A INPUT -j DROP

Rules live in the kernel and vanish on reboot. iptables-save prints them in a form iptables-restore reads back, and on Ubuntu the package iptables-persistent does that automatically at boot. A firewall that is not saved is a firewall that lasts until the next power cut.

And the same policy written for nftables, which is what Ubuntu 24.04 actually runs underneath:

$ ip netns exec server nft add table inet mnfilter
$ ip netns exec server nft add chain inet mnfilter input '{ type filter hook input priority 0 ; policy drop ; }'
$ ip netns exec server nft add rule inet mnfilter input ct state established,related accept
$ ip netns exec server nft add rule inet mnfilter input tcp dport 8080 accept
$ ip netns exec server nft list table inet mnfilter
table inet mnfilter {
	chain input {
		type filter hook input priority filter; policy drop;
		ct state established,related accept
		tcp dport 8080 accept
	}
}

Two things worth saying. The policy is declared on the chain here rather than written as a last rule, which is the tidier form. And inet means the table handles IPv4 and IPv6 together, where iptables needs a separate ip6tables and a second set of rules that people forget to write.

The iptables command on this system is iptables-nft: it speaks the old syntax and writes nftables rules underneath. Both are therefore the same firewall, and a rule added one way is visible the other way.

Step 8: clean up

$ ip netns exec server iptables -F
$ ip netns exec server nft delete table inet mnfilter
$ ip netns delete client
$ ip netns delete server
$ ip netns list | wc -l
0

Procedure

  1. Create two namespaces joined by a veth pair, and start two listeners on the server.
  2. Show the empty chains and their policies, and test both ports and a ping. Record all three as open or working.
  3. Add four rules in order: loopback, established and related, the one permitted port, then drop.
  4. List the rules with line numbers and explain each.
  5. Test all three things again and record the three results.
  6. Insert a REJECT for the blocked port, test, record the message; remove it, test again, record the different message.
  7. Insert a DROP above the ACCEPT for the permitted port, show the listing, and record that the port is now closed although the ACCEPT is still there. Remove it.
  8. Print the rules with counters and say which rule is taking the traffic.
  9. Save the rules with iptables-save, and write the same policy in nftables.
  10. Flush everything and delete the namespaces.
munotes.in195

Practical 20: Firewall Configuration and Rule-Based Filtering

Observations

TestBefore any ruleAfter the four rules
Port 8080openopen
Port 9090openclosed
Pingworksblocked
MeasuredValue
Chains and their initial policyINPUT, FORWARD, OUTPUT, all ACCEPT
Rules added4: loopback, conntrack, dport 8080, drop
Blocked port with REJECTConnection refused, immediately
Blocked port with DROPtimed out, after the client gave up
DROP inserted above the ACCEPTport 8080 closed, with the ACCEPT still listed at line 2
After deleting the inserted DROPport 8080 open again
Countersevery rule carries a packet and a byte count
iptables-save lines for INPUT4
nftables equivalentone table, one chain with policy drop, two rules

Result

A firewall was configured on a host created as a network namespace and tested at every step. With empty chains both services were reachable and the host answered a ping. Four rules were then added, accepting loopback traffic, accepting established and related connections, accepting TCP to port 8080 and dropping everything else, after which port 8080 was still open, port 9090 was closed and ping was blocked. REJECT on the blocked port produced Connection refused at once while DROP produced a timeout, which is the measured difference between them. Inserting a DROP above the existing ACCEPT for port 8080 closed the port while the ACCEPT remained visible at line 2 of the listing, demonstrating that rules are tried in order and the first match decides, and deleting it restored the service. The rules were printed with their packet counters, exported with iptables-save, and the same policy was written in nftables with the policy declared on the chain and a single inet table covering both IP versions.

Where marks are lost

Rules with no tests. Adding a rule is half the exercise. MU says configure and test.

No before. Record the state with no firewall, or the after proves nothing.

Forgetting the conntrack rule. The machine can then make no outward connection, because the replies are dropped, and the symptom looks nothing like a firewall.

Forgetting loopback. Software that talks to itself breaks in confusing ways.

Appending a rule after the final DROP. It will never be reached. -A appends; use -I with a position.

Saying DROP and REJECT are the same. One times out, one refuses immediately. Produce both.

Blocking a list instead of permitting one. Default deny. You cannot enumerate everything that is bad.

Not saving the rules. They are gone at the next reboot.

Locking yourself out. On a remote machine, permit your own management port before the DROP, and have a way back in.

munotes.in196

Practical 20: Firewall Configuration and Rule-Based Filtering

For the journal

Aim; the three chains and what each sees; the three verdicts; the default-deny rule from NIST; the two hosts and two services; the empty chains and the three tests before; the four rules with a sentence on each; the listing with line numbers; the three tests after; the REJECT and the DROP messages side by side; the reordering experiment with the listing showing the unreachable ACCEPT; the counters; the LOG rule with a note on why it is not run here and what the log line looks like; iptables-save; the nftables version; the clean-up; both observation tables; the result.

Quick revision

  • INPUT for packets to this machine, OUTPUT from it, FORWARD through it.
  • Each chain has a policy and an ordered list of rules. First match wins.
  • ACCEPT, DROP (silent, times out), REJECT (sends an error, refuses at once).
  • Default deny: permit what is needed, drop the rest.
  • -A appends, -I inserts, -D deletes, -F flushes, -L -n -v --line-numbers lists.
  • Always accept loopback and established/related, or the machine breaks.
  • Criteria: -p protocol, --dport and --sport ports, -s and -d addresses, -i and -o interfaces.
  • A rule after the final DROP is never reached.
  • Counters show which rule is taking the traffic; zero means never matched.
  • LOG is not a verdict: the packet carries on. Rate-limit it.
  • iptables-save and iptables-restore; iptables-persistent does it at boot.
  • On Ubuntu 24.04 iptables is iptables-nft: nftables underneath, and inet covers both IP versions.

Questions you must be able to answer

1. Which chain does a packet arriving for a web server on this machine meet? INPUT. OUTPUT is for packets this machine sends, FORWARD for packets routed through it.

2. What is the difference between DROP and REJECT, and when would you use each? DROP discards silently, so the sender waits and times out; REJECT sends an error and the sender is refused at once. REJECT inside your own network so users get a useful error; DROP at the boundary so a scanner learns nothing and is slowed down.

3. Why do you need the conntrack rule? Because replies to connections the machine itself opened are not new connections. Without accepting ESTABLISHED and RELATED, the final DROP discards every reply and the machine cannot fetch anything.

4. What does default deny mean? That anything not explicitly permitted is dropped. It is the only workable direction, because you can list what should be allowed and you cannot list everything that should not.

5. Your ACCEPT rule is in the list and the port is still closed. Why? Because an earlier rule matched first. Rules are tried in order and the first match decides; this chapter inserts a DROP above the ACCEPT and the port closes while the ACCEPT sits visible at line 2.

munotes.in197

Practical 20: Firewall Configuration and Rule-Based Filtering

6. How do you find out which rule is stopping traffic? iptables -L -n -v, and read the packet counters. The rule taking the traffic is the one whose count is rising; a rule at zero has never matched.

7. Is LOG a verdict? No. A packet matching a LOG rule is logged and then carries on to the next rule, which is why the LOG rule is placed above the DROP rather than in place of it.

8. Why did the LOG demonstration not run in this laboratory? Because LOG writes to the kernel's message buffer, which a container shares with its host and is not permitted to read, so dmesg returns nothing inside it. On a real machine the same two commands work.

9. What happens to your rules when the machine reboots? They are gone. They live in the kernel. Save them with iptables-save, or install iptables-persistent to restore them at boot.

10. What does inet mean in an nftables table, and why does it matter? That the table handles IPv4 and IPv6 in one place. With iptables you need a second, separate set of rules under ip6tables, and forgetting them leaves the machine open over IPv6.

Contents This chapter on its own page

munotes.in198

Chapter Twenty-Six

The Two-Hour Paper: Sitting the Examination

Syllabus topic Module 1 and Module 2, and MU's assessment of a 2-credit practical course: "A Semester End Practical Examination of 2 hours duration for 30 marks", "Q. 1 Module 1 15", "Q. 2 Module 2 15", "Certified Journal is compulsory for appearing at the time of Practical Exam", "Minimum 80% practical are required to be completed."

Aim

To sit the Semester End Practical Examination: two hours, two questions, fifteen marks each, one from each module.

What the paper looks like

Chapter 1 has the figures; here they are again in the form they matter on the day.

Duration2 hours
Total30 marks
Q.1a practical question on Module 1, 15 marks
Q.2a practical question on Module 2, 15 marks
Required to sita certified journal, and at least 16 of the 20 practicals complete

Neither question is announced in advance. Q.1 is one of the ten AI exercises and Q.2 is one of the ten security exercises, and either could be any of them, which is why the 80 per cent rule is a floor and not a target.

What a full-marks answer contains

Whatever the question, the answer has the same seven parts. Write the headings on the answer sheet first, before touching the keyboard; they cost a minute and they stop you forgetting the parts that are not code.

  1. Aim. One line, in the question's own words.
  2. Theory. Three or four sentences. What the algorithm or the mechanism is.
  3. Algorithm or procedure. Numbered steps, in words, before any code.
  4. Program or configuration. Complete enough to run.
  5. Output. Exactly what the machine printed.
  6. Observations. The table, the counts, the accuracy, the comparison.
  7. Conclusion. Two lines saying what the run showed.

Parts 5 and 6 are the ones students lose marks on, and they are the cheapest to get right. A program with no output is an unfinished practical, and a practical whose question says "evaluate" or "compare" and whose answer has no table has not answered the question.

The two hours

MinutesQ.1Q.2
0 to 5read both questions and decide which to start with
5 to 12aim, theory, algorithm, on paper
12 to 45type, run, fix, get output
45 to 55write the output and the observations
55 to 62aim, theory, procedure
62 to 95type, run, fix, get output
95 to 105output and observations
105 to 120check both: output present, observations filled, conclusion written

Start with the one you are surer of, and do not spend more than an hour on either. Half the paper is in the other module, and a complete answer worth 12 beats two half answers worth 6 each.

Q.1, worked: a Module 1 question

Q.1 Implement the K-Nearest Neighbours algorithm. Apply it to the given dataset, predict the class for the test data, and evaluate the accuracy of the predictions. (15 marks)

Aim. To implement K-NN, predict the class of each test row, and measure the accuracy.

Theory. K-NN stores the training data and does all its work when asked. To classify a row it finds the k training rows nearest to it by Euclidean distance and takes the commonest label among them. It has no training phase. The features must be scaled first, because distance adds up the difference in every column and a column with larger numbers would decide the answer by itself.

munotes.in199

The Two-Hour Paper: Sitting the Examination

Algorithm.

  1. Read the dataset.
  2. Scale every feature to the range 0 to 1.
  3. Shuffle with a fixed seed and split 70 to 30.
  4. For each test row, compute its distance to every training row.
  5. Take the k nearest and let them vote.
  6. Count how many predictions are right, and build the confusion matrix.
  7. Compare with the baseline of always answering the commoner class.

Program and output.

name,attendance,practice,result
Aarav,50,22,Fail
Isha,48,16,Fail
Rohan,88,38,Pass
Sanya,97,39,Pass
Vikram,90,30,Fail
Meera,35,41,Fail
Farhan,45,9,Fail
Nikita,71,8,Fail
Omkar,92,2,Fail
Pooja,97,45,Pass
Rahul,75,15,Fail
Sneha,85,18,Pass
Tejas,79,24,Pass
Urmila,83,34,Pass
Varun,44,23,Fail
Yash,46,37,Pass
Zoya,72,20,Fail
Amit,44,9,Pass
Bhavna,74,3,Fail
Chirag,82,25,Pass
Deepa,94,29,Pass
Eshan,46,27,Fail
Gauri,98,9,Fail
Harsh,89,34,Pass
Ira,97,27,Pass
Jatin,68,28,Pass
Kavya,96,34,Pass
Lalit,38,38,Fail
Manav,63,10,Fail
Neha,41,35,Fail
"""Q.1 (Module 1, 15 marks): implement K-NN, predict, and evaluate."""
import csv
import math
import random

# ---- 1. read the dataset
with open("students.csv", newline="") as fh:
    rows = list(csv.DictReader(fh))
raw = [[float(r["attendance"]), float(r["practice"])] for r in rows]
y = [r["result"] for r in rows]

# ---- 2. scale, because distance adds up every column
lo = [min(r[k] for r in raw) for k in range(2)]
hi = [max(r[k] for r in raw) for k in range(2)]
X = [[(r[k] - lo[k]) / (hi[k] - lo[k]) for k in range(2)] for r in raw]

# ---- 3. split, seeded so the run repeats
random.seed(1)
order = list(range(len(X)))
random.shuffle(order)
cut = int(0.7 * len(X))
train, test = order[:cut], order[cut:]
print("training rows %d, test rows %d" % (len(train), len(test)))

# ---- 4. the algorithm
def distance(a, b):
    return math.sqrt(sum((a[k] - b[k]) ** 2 for k in range(len(a))))

def classify(i, k):
    near = sorted(train, key=lambda j: (distance(X[i], X[j]), j))[:k]
    votes = {}
    for j in near:
        votes[y[j]] = votes.get(y[j], 0) + 1
    best = max(votes.values())
    return sorted(c for c, v in votes.items() if v == best)[0]

# ---- 5. predict every test row
K = 5
print()
print("%-10s %11s %9s %10s %8s %7s"
      % ("student", "attendance", "practice", "predicted", "actual", "right"))
right = 0
for i in test:
    p = classify(i, K)
    ok = (p == y[i])
    right += ok
    print("%-10s %11s %9s %10s %8s %7s"
          % (rows[i]["name"], rows[i]["attendance"], rows[i]["practice"], p, y[i],
             "yes" if ok else "NO"))

# ---- 6. evaluate
print()
tp = sum(1 for i in test if classify(i, K) == "Pass" and y[i] == "Pass")
fp = sum(1 for i in test if classify(i, K) == "Pass" and y[i] != "Pass")
fn = sum(1 for i in test if classify(i, K) != "Pass" and y[i] == "Pass")
tn = sum(1 for i in test if classify(i, K) != "Pass" and y[i] != "Pass")
acc = (tp + tn) / len(test)
prec = tp / (tp + fp) if tp + fp else 0.0
rec = tp / (tp + fn) if tp + fn else 0.0
print("confusion matrix, positive class Pass")
print("             predicted Pass  predicted Fail")
print("actual Pass  %14d  %14d" % (tp, fn))
print("actual Fail  %14d  %14d" % (fp, tn))
print()
print("k          : %d" % K)
print("accuracy   : %d of %d = %.3f" % (right, len(test), acc))
print("precision  : %.3f" % prec)
print("recall     : %.3f" % rec)

from collections import Counter
commonest = Counter(y[i] for i in train).most_common(1)[0][0]
base = sum(1 for i in test if y[i] == commonest) / len(test)
print("baseline   : always answer %s, %.3f" % (commonest, base))
munotes.in200

The Two-Hour Paper: Sitting the Examination

training rows 21, test rows 9

student     attendance  practice  predicted   actual   right
Yash                46        37       Fail     Pass      NO
Sanya               97        39       Pass     Pass     yes
Omkar               92         2       Fail     Fail     yes
Rohan               88        38       Pass     Pass     yes
Ira                 97        27       Pass     Pass     yes
Jatin               68        28       Pass     Pass     yes
Lalit               38        38       Fail     Fail     yes
Bhavna              74         3       Fail     Fail     yes
Vikram              90        30       Pass     Fail      NO

confusion matrix, positive class Pass
             predicted Pass  predicted Fail
actual Pass               4               1
actual Fail               1               3

k          : 5
accuracy   : 7 of 9 = 0.778
precision  : 0.800
recall     : 0.800
baseline   : always answer Fail, 0.444

Observations.

MeasuredValue
Training rows, test rows21, 9
k5
Correct predictions7 of 9
Accuracy0.778
Precision, positive class Pass0.800
Recall0.800
Confusion matrixTP 4, FN 1, FP 1, TN 3
Baseline, always Fail0.444
Rows got wrongYash, predicted Fail and passed; Vikram, predicted Pass and failed

Conclusion. K-NN with k = 5 classified 7 of the 9 unseen rows correctly, an accuracy of 0.778 against a baseline of 0.444, with precision and recall both 0.800. The two it got wrong are the two students whose results do not follow the pattern of the rest: one with low attendance who passed and one with high attendance who failed.

Why that answer would score well. It names the aim, gives the theory in four sentences, lists the algorithm before the code, runs, prints the prediction for every test row and not just a total, reports accuracy and the confusion matrix and the baseline, and ends by saying which rows were wrong and why. None of that took longer than the program did.

Q.2, worked: a Module 2 question

Q.2 Implement the RSA algorithm. Generate a key pair, encrypt and decrypt a message, and demonstrate a digital signature using the same key. (15 marks)

Aim. To generate an RSA key pair, to encrypt and decrypt a message with it, and to sign and verify a message with the same pair.

munotes.in201

The Two-Hour Paper: Sitting the Examination

Theory. RSA uses two keys. The modulus n is the product of two primes p and q; the public exponent e is chosen coprime with phi(n) = (p-1)(q-1), and the private exponent d is the inverse of e modulo phi. Encryption is c = m^e mod n with the public key and decryption is m = c^d mod n with the private one. A signature is the same arithmetic the other way round: the signer raises the hash of the message to d, and anybody raising the result to e and comparing with their own hash has verified it.

Algorithm.

  1. Choose two primes p and q, and check they are prime.
  2. Compute n and phi.
  3. Choose e with gcd(e, phi) = 1, and check.
  4. Compute d as the inverse of e modulo phi, by the extended Euclidean algorithm.
  5. Encrypt each letter, then decrypt it, and show both.
  6. Hash the message, raise the hash to d, and verify with e.
  7. Verify the same signature against a different message and show that it fails.

Program and output.

"""Q.2 (Module 2, 15 marks): RSA key generation, encryption, decryption,
   and a digital signature over the same message."""
import hashlib
import math

def egcd(a, b):
    if b == 0:
        return a, 1, 0
    g, x, y = egcd(b, a % b)
    return g, y, x - (a // b) * y

def inverse_mod(a, m):
    g, x, _ = egcd(a % m, m)
    if g != 1:
        raise ValueError("e and phi are not coprime")
    return x % m

def is_prime(n):
    if n < 2:
        return False
    for f in range(2, int(math.isqrt(n)) + 1):
        if n % f == 0:
            return False
    return True

# ---- 1. key generation
p, q = 137, 149
assert is_prime(p) and is_prime(q), "p and q must be prime"
n = p * q
phi = (p - 1) * (q - 1)
# NOTE: e = 17 does NOT work with these primes: 17 divides 20128 exactly, and the
# assertion below caught it. Choose e again rather than ignoring the failure.
e = 7
assert math.gcd(e, phi) == 1, "e must be coprime with phi"
d = inverse_mod(e, phi)
print("KEY GENERATION")
print("  p = %d, q = %d" % (p, q))
print("  n = p * q = %d" % n)
print("  phi = (p-1)(q-1) = %d" % phi)
print("  e = %d, gcd(e, phi) = %d" % (e, math.gcd(e, phi)))
print("  d = %d, and e*d mod phi = %d" % (d, (e * d) % phi))
print("  public (e, n) = (%d, %d), private (d, n) = (%d, %d)" % (e, n, d, n))

# ---- 2. encryption and decryption
MESSAGE = "EXAM"
print()
print("ENCRYPTION AND DECRYPTION, one letter at a time")
print("  %-8s %6s %10s %10s" % ("letter", "m", "c", "back"))
for ch in MESSAGE:
    m = ord(ch)
    c = pow(m, e, n)
    print("  %-8s %6d %10d %10d" % (ch, m, c, pow(c, d, n)))

# ---- 3. a signature over the whole message
print()
print("DIGITAL SIGNATURE")
digest = hashlib.sha256(MESSAGE.encode()).hexdigest()
print("  sha256(%s) = %s" % (MESSAGE, digest))
h = int(digest, 16) % n
print("  reduced mod n so it fits this small key: h = %d" % h)
sig = pow(h, d, n)
print("  signature = h^d mod n = %d" % sig)
print("  verify: sig^e mod n = %d, h = %d, accepted: %s"
      % (pow(sig, e, n), h, pow(sig, e, n) == h))
bad = int(hashlib.sha256(b"EXAN").hexdigest(), 16) % n
print("  the same signature against a different message: accepted: %s"
      % (pow(sig, e, n) == bad))
munotes.in202

The Two-Hour Paper: Sitting the Examination

KEY GENERATION
  p = 137, q = 149
  n = p * q = 20413
  phi = (p-1)(q-1) = 20128
  e = 7, gcd(e, phi) = 1
  d = 5751, and e*d mod phi = 1
  public (e, n) = (7, 20413), private (d, n) = (5751, 20413)

ENCRYPTION AND DECRYPTION, one letter at a time
  letter        m          c       back
  E            69       7474         69
  X            88      14185         88
  A            65      11375         65
  M            77      14997         77

DIGITAL SIGNATURE
  sha256(EXAM) = bb86f6ac87986ae1e5c420ee82a911c5473444bdd34a0c2c679eb0ae904bae43
  reduced mod n so it fits this small key: h = 19475
  signature = h^d mod n = 13752
  verify: sig^e mod n = 19475, h = 19475, accepted: True
  the same signature against a different message: accepted: False

Observations.

MeasuredValue
p, q137, 149, both checked prime
n20413
phi20128
e7, because 17 divides 20128 and the assertion caught it
d5751, and 7 times 5751 mod 20128 is 1
MessageEXAM, encrypted and decrypted letter by letter, all four recovered
Signature13752, over the SHA-256 digest reduced modulo n
Verification, correct messageaccepted
Verification, altered messagerejected

Conclusion. An RSA key pair was generated from two primes and the relation between e and d confirmed. Each letter of the message was encrypted with the public key and recovered with the private one. A signature over the message's SHA-256 digest verified with the public key and failed against a different message, which shows the signature binds one message to one key.

Note the assertion that failed. The first attempt used e = 17, and gcd(17, 20128) is 17, not 1, so there is no inverse and the key cannot exist. The assertion stopped the program with a message instead of producing a key that silently did not work. Put those two assertions in your exam program: one that p and q are prime, one that e is coprime with phi. They cost two lines and they turn a silent wrong answer into a clear one.

munotes.in203

The Two-Hour Paper: Sitting the Examination

If Q.2 is a configuration question instead

Five of the ten Module 2 practicals are configuration rather than programming: IPsec, TLS, the IDS, the antivirus and the firewall. The seven parts are the same, with two differences.

The "program" is the commands and the configuration file, written out in full. Paste them; do not describe them.

The "output" is the terminal transcript, and the observation is a before and an after. Every one of those five chapters is built the same way: record the state before you change anything, make the change, record the state again. A firewall answer that lists rules and never shows a connection being refused has not tested anything, and MU's wording for that practical is "configure and test".

And take the two minutes to break it on purpose at the end. Deleting one IPsec association, or connecting without the CA file, or inserting a DROP above an ACCEPT, produces a second observation that shows you understand the mechanism rather than the recipe.

The journal, which decides whether you sit the paper at all

Before the examination, not on the day:

  • [ ] All twenty practicals written up, or at the very least sixteen.
  • [ ] Every one signed by the subject teacher, with a date.
  • [ ] Each write-up has all seven parts, including the output and the observations.
  • [ ] The index at the front lists all twenty with their page numbers.
  • [ ] Your name, roll number, class and subject on the cover.

An unsigned journal is not a certified journal, and MU's rule is that a certified journal is compulsory for appearing. Get each practical signed in the week you do it. A teacher asked to sign twenty write-ups the day before the examination is entitled to decline.

The things that lose marks in the hall, in order of how often they happen

No output. The commonest of all. If the program will not run, write down the exact error, say in one sentence what you think it means, and move on.

No observations. Six of the ten Module 1 exercises say evaluate or compare, and every Module 2 configuration says test. That part is not optional and it is not long.

Ninety minutes on one question. Two questions, fifteen marks each.

A program with no aim, theory or algorithm above it. Those three are quick, and they are marks.

Numbers with no labels. 0.778 on its own means nothing; accuracy: 0.778 on 9 test rows is a result.

munotes.in204

The Two-Hour Paper: Sitting the Examination

An accuracy measured on the training rows. Chapter 3, and it applies to whichever model the question asks for.

A key made from numbers that are not prime. It fails silently. Assert.

Rules added and never tested. Configure and test.

No conclusion. Two lines. They are the easiest marks on the paper.

Procedure

  1. A week before: count your write-ups, make sure at least sixteen are signed, and complete the index.
  2. Re-read Chapter 1 and this chapter.
  3. Work one question from each module against the clock, with no notes, and see how long you actually take.
  4. In the hall: read both questions, decide the order, write the seven headings on the answer sheet.
  5. Do the question you are surer of first, and stop at the hour.
  6. Write the output into the answer as soon as you have it, before improving anything.
  7. Fill in the observations while the program is still on the screen.
  8. Leave fifteen minutes to check that both answers have output, observations and a conclusion.

Observations

Measured, on the two worked answersValue
Q.1, K-NN at k = 57 of 9 correct, accuracy 0.778
Q.1, precision and recall0.800 and 0.800
Q.1, baseline0.444
Q.2, keyp 137, q 149, n 20413, phi 20128, e 7, d 5751
Q.2, encryption and decryptionall four letters recovered
Q.2, signatureverified against the message, rejected against another
Q.2, first choice of e17, rejected by an assertion because it divides phi
Time each answer takes at a steady paceabout 50 minutes

Result

A complete Semester End Practical Examination was worked under its printed conditions: two hours, two questions, fifteen marks each. Q.1 implemented K-Nearest Neighbours, predicted every one of nine unseen rows, and reported an accuracy of 0.778 with a confusion matrix, precision and recall of 0.800 each and a baseline of 0.444. Q.2 generated an RSA key pair from two primes, encrypted and decrypted a message letter by letter, and produced and verified a digital signature, with an assertion rejecting a public exponent that shared a factor with phi. Both answers were written in the seven-part form, and the journal requirements for being admitted to the examination were listed.

Quick revision

  • Two hours, 30 marks, Q.1 on Module 1 for 15 and Q.2 on Module 2 for 15.
  • A certified journal is compulsory, and at least 16 of the 20 practicals must be complete.
  • Seven parts: aim, theory, algorithm, program, output, observations, conclusion.
  • Write the seven headings before touching the keyboard.
  • Start with the question you are surer of, and stop at the hour.
  • Never leave the output blank. An error message with a sentence of diagnosis beats an empty page.
  • If the question says evaluate or compare, the observation table is the answer to it.
  • If the question says configure, record the state before and after, and test.
  • Label every number and say how many rows a score was measured on.
  • Assert what must be true: that p and q are prime, that e is coprime with phi.
  • Leave fifteen minutes at the end to check output, observations and conclusion on both answers.
munotes.in205

The Two-Hour Paper: Sitting the Examination

Questions you must be able to answer

1. How long is the paper and how is it divided? Two hours and 30 marks: one practical question on Module 1 for 15 marks and one on Module 2 for 15.

2. What do you need in order to be allowed to sit it? A certified journal, and at least eighty per cent of the practicals completed, which is sixteen of the twenty.

3. Your program will not run and twenty minutes remain. What do you do? Write the aim, theory and algorithm, paste the program as far as it goes, copy the exact error message, add one sentence saying what you think it means, and start the other question.

4. The question says "evaluate the accuracy". What must the answer contain? A number, what it was measured on, how many rows that was, and the baseline it beats. A confusion matrix if the model classifies.

5. The question says "configure and test a firewall". What is the test? Trying the traffic. Record what reaches the machine before the rules, add the rules, and record what reaches it after. An answer with rules and no connection attempt has tested nothing.

6. Why write the seven headings before starting? Because output, observations and conclusion are the parts that get forgotten when time runs short, and they are the cheapest marks on the paper.

7. Your key generation raised an assertion. Is that a failed answer? No, it is a correct one. The exponent shared a factor with phi and no inverse exists. Choose another exponent, say so in the answer, and you have shown you know why the condition is there.

8. How much of your two hours should the first question take? About an hour. If it has taken more, write down what you have and move on: the other fifteen marks are in the other module.

Contents This chapter on its own page

munotes.in206

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Issue
Done!