munotes®

Major Practical 3 Notes | B.Sc. (Information Technology) Semester 3 | Mumbai University | munotes

Official Notes munotes.in

Major Practical 3

B.SC. (INFORMATION TECHNOLOGY) · SEMESTER 3

Strictly as per the University of Mumbai NEP syllabus in force for B.Sc. (Information Technology)

For B.Sc. (Information Technology) students of the University of Mumbai and all its affiliated colleges

Open the book ↓

munotes.in Second Year

Major Practical 3

Copyright © 2026 munotes.in. All rights reserved.

Written and first published by munotes.in, 2026.

This book is free for individual students to read at munotes.in. No part of it may be reproduced, distributed, stored, translated or used for institutional or classroom purposes in any form without a prior written licence from munotes.in.

Licensing and permissions: contact@munotes.in

The text of statutes and of judgments reproduced in this book is in the public domain under section 52(1)(q) of the Copyright Act 1957. The commentary, arrangement, examples and questions are the original work of munotes.in.

munotes.in is an independent study resource for MU students. It is not affiliated with, endorsed by, or officially connected to the University of Mumbai. Course names and university references describe the students and syllabus the material relates to.

munotes.in

Contents

Module I 1. Write programs for the following:

  1. How This Practical Is Examined: the Slip, the Journal and the Viva 1
  2. The Laboratory from Zero: Python, IDLE, and Your First Program 5
  3. Practical 1: Input, Output, and the Year You Turn 100 10
  4. Practical 1 continued: Even or Odd, and the SGPI Grade Ladder 15
  5. Practical 2: the Fibonacci Series, and the Sum of the Digits 21
  6. Practical 3: Arrays, Basic Operations, Indexing and Slicing 27
  7. Practical 3 continued: Mathematical Functions, Aliasing and Copying 33
  8. Practical 4: NumPy Slicing, Basic and Advanced Indexing 41
  9. Practical 4 continued: the Dimensions and Attributes of an Array 48
  10. Practical 5: Functions, Armstrong Numbers and Palindromes 55
  11. Practical 5 continued: Recursion, and the Lambda 62
  12. Practical 6: Counting the Characters and the Words in a String 68
  13. Practical 6 continued: the Geometry Module, and pointyShapeVolume 73
  14. Practical 7: a Common Member, and a Dictionary Sorted by Value 79
  15. Practical 8: the Tuple Return, Area and Circumference 86
  16. Practical 8 continued: Text Files, Binary Files, and the Last n Lines 92
  17. Practical 9: Counting a Word in a File with a Regular Expression 100
  18. Practical 9 continued: Extracting Every Hyperlink from an HTML File 106
  19. Practical 10: Comparing Two Dates in DD/MM/YYYY Form 113
  20. Practical 10 continued: Measuring Execution Time, and the Calendar Module 119

Module II 1. Array Operations: Write a program to implement basic array operations:

  1. Python for Data Structures, and the Cost of an Operation 126
  2. Practical 1: Array Operations, Insert, Delete and Linear Search 137
  3. Practical 2: Building a Singly Linked List 144
  4. Practical 2 continued: Deleting a Node from a Linked List 152
  5. Practical 3: a Stack over an Array 160
  6. Practical 3 continued: Infix to Postfix with a Stack 167
  7. Practical 4: a Queue over an Array 177
  8. Practical 4 continued: Simulating a Customer Service Queue 186
  9. Practical 5: the Binary Search Tree, Create, Insert and Search 193
  10. Practical 6: Tree Traversal, Pre-order, In-order and Post-order 203
  11. Practical 7: a Hash Table with Separate Chaining 212
  12. Practical 8: Bubble, Insertion and Selection Sort Compared 224
  13. Practical 9: Linear and Binary Search Compared 233
  14. Practical 10: the Combined Application 243
  15. If Your College Runs Module 2 in C 253

Module J The journal, the viva and the practical examination

  1. Keeping the Journal, and What Goes on the Page 268
  2. A Worked Practical Slip: Q1, Q2 and Q3 275
  3. The Viva: the Questions Asked at the Table 285
munotes.in

Module I

1. Write programs for the following:

munotes.in

Chapter One

How This Practical Is Examined: the Slip, the Journal and the Viva

Syllabus topic Items 12, 13 and 14 of MU's own particulars table for Major Practical 3

Aim

To know, before touching a keyboard, exactly how this paper is marked: what the slip looks like, what the journal is worth, and what the viva is.

Why this chapter is first

Every other chapter in this book teaches a program. This one teaches the paper, and it comes first because several marks on it are lost before any program is written. A student who arrives without a signed journal is not allowed to sit the examination at all. A student who has done eighteen of the twenty exercises has already given away marks that no amount of good code recovers.

What the paper is

PaperMajor Practical 3
CourseB.Sc. (Information Technology), Second Year, Semester 3
VerticalMajor
TypePractical
Credits2
Marks50
Internal40 per cent, which is 20 marks
Semester end60 per cent, which is 30 marks
Duration of the examination2 hours

The 50 marks are the whole paper. They arrive in two pieces, and each is passed separately. MU's scheme of examination for this course prints "Individual Passing in Internal and External Examination" with a standard of passing of 40 per cent. So 8 of the 20 internal marks and 12 of the 30 external marks are the two floors, and a brilliant examination does not rescue an empty journal.

The 20 internal marks: what MU asks of you week by week

MU's item 13 for this paper prints it plainly. Students are expected to attend each practical and to submit the written practical of the previous session. Performing the practical and submitting the writeup is the continuous internal evaluation, and 2.5 marks can be awarded for each practical performance and writeup submission, totalling to 50 marks, which is then converted to 20 marks.

Read that arithmetic, because it settles a question students argue about:

Marks per practical2.5
Practicals MU prints20
Marks available in all2.5 × 20 = 50
Converted to20

Two and a half marks multiplied by twenty practicals is exactly the 50 MU prints. There is no slack in that arithmetic. All twenty exercises are expected in the journal, ten from Module 1 and ten from Module 2, each performed in the laboratory and each written up and submitted at the following session.

The phrase "the written practical of the previous session" is the one most often missed. The writeup is not something to be assembled the week before the examination. It is due at the next practical, and a journal submitted in one heap at the end has usually already lost part of these marks.

The 30 external marks: the practical slip

MU's item 14 prints the format of the question paper for this course. It is a practical examination of two hours, and the slip you are handed carries three questions:

munotes.in1

How This Practical Is Examined: the Slip, the Journal and the Viva

QuestionSet onMarks
Q1Module 113
Q2Module 212
Q3Journal and Viva05
Total30

Three things in that table are worth saying out loud.

The two programming questions are not equal. Q1 carries 13 and Q2 carries 12. It is a single mark, but it tells you which module to start with if you are short of time: Module 1, the Python half, is worth marginally more, and it is also the half where a program either runs or does not, with less to argue about.

Q3 is not a formality. Five marks out of thirty is a sixth of the external paper, and it is awarded for your journal and for what you say at the table. That is more than the difference between Q1 and Q2 five times over. [The Viva: the Questions Asked at the Table] is a whole chapter on it, because no other chapter in the book owns it.

The journal is counted twice. It earns part of the 20 internal marks through item 13, and it is examined again inside Q3 of the external slip. One document, two sets of marks.

The rule that stops you at the door

MU's item 14 prints this in the same breath as the format:

Certified copy of Journal is compulsory to appear for the practical examination.

Certified means signed. A journal that is complete but not signed by your practical teacher is not a certified journal, and the requirement is to appear, not merely to score. This is the one line in the whole scheme that can cost a student an entire paper, and it is settled weeks before the examination, in the laboratory, by getting each entry signed as it is completed. [Keeping the Journal, and What Goes on the Page] sets out the form of an entry and what the signature is given for.

What the two modules are

MU sets ten exercises in each module, and they are two different subjects. This paper is the practical half of two theory papers you are sitting in the same semester.

Module 1Module 2
The subjectPython programmingData structures
Its theory paperPython ProgrammingData Structures
Exercises1010
The slip questionQ1, 13 marksQ2, 12 marks
Chapters herepracticals 1 to 10 of Module 1practicals 1 to 10 of Module 2

MU's own description of this paper is "a comprehensive exploration of advanced Python programming concepts", which is written about Module 1, and her Course Objectives then run on into arrays, linked lists, stacks, queues, trees and graphs, which is Module 2. One paper, two halves, one slip.

munotes.in2

How This Practical Is Examined: the Slip, the Journal and the Viva

The language question, answered

Module 1 is Python and there is nothing to decide: every row MU prints names a Python feature, from NumPy indexing to the calendar module.

Module 2 does not name a language, and MU leaves the choice to you in writing. The Data Structures theory paper that this module pairs with prints, as its third Course Outcome:

Translate algorithmic solutions into correctly functioning code using their chosen

programming language.

This book writes Module 2 in Python, because it is the language of the other half of the same slip and of the theory paper you are sitting alongside it. But three of MU's five text books for this paper are C books, and this degree taught you Programming with C in Semester 1, so some colleges run the Module 2 laboratory in C. If yours does, [If Your College Runs Module 2 in C] gives all ten exercises in C, compiled and run, so nothing in this book is out of reach.

Write your journal in the language your own laboratory uses, and answer Q2 in that language. The structures and the algorithms are identical either way, and an examiner marks the structure, not the syntax.

How the 30 marks of a practical question are earned

MU does not print a breakdown inside Q1 and Q2, and no honest book can invent one. What every practical examiner in this system looks for is the same five things, and they are what the journal format in this book is built around:

  1. The aim, in MU's own words where possible.
  2. The program, written out, correctly indented, with the variable names a reader can

follow.

  1. The run, the output as it actually appeared, not as it should have.
  2. The conclusion, one or two sentences saying what the exercise showed.
  3. Being able to answer for it, which is Q3.

A program that runs and prints the right answer, with no aim and no conclusion written underneath, leaves marks on the table in every one of those five.

Where marks are lost

  • No certified journal, so no examination. The worst and the most avoidable.
  • Eighteen practicals instead of twenty. Item 13's own arithmetic is 2.5 times 20.
  • Writeups submitted in a heap at the end instead of at the following session.
  • Treating Q3 as a formality. It is five marks, more than the gap between Q1 and Q2.
  • Starting with Module 2 out of habit when Module 1 carries the extra mark and is

faster to finish.

  • Assuming the 15 and 15 split from another course's practical paper. This one is

13, 12 and 5.

munotes.in3

How This Practical Is Examined: the Slip, the Journal and the Viva

  • Writing the program and not the output. The output is the evidence the program ran.

For the journal

This chapter is not a practical and has no journal entry. Copy the three-row table of Q1, Q2 and Q3 into the front of your journal anyway, on the page before the index, with the line about the certified copy under it. It takes two minutes and it is the only part of the scheme you need to remember.

Quick revision

  • Major Practical 3, Semester 3, 2 credits, 50 marks: 20 internal and 30 external.
  • Individual passing in internal and external, standard of passing 40 per cent.
  • Internal: 2.5 marks per practical performance and writeup, 20 practicals, total 50,

converted to 20. The writeup is due at the next session.

  • External: a two hour practical examination for 30 marks.
  • The slip: Q1 Module 1 13 marks, Q2 Module 2 12 marks, Q3 Journal and Viva 5 marks.
  • A certified copy of the journal is compulsory to appear. Signed, not just complete.
  • Module 1 is Python, ten exercises. Module 2 is data structures, ten exercises.
  • Module 2's language is yours to choose: MU's own Data Structures OC 3 says so.
  • Source: MU item 6.9 (N), AC 28 March 2025, in force from 2025-26.

Questions you should be able to answer

1. What are the three questions on the practical slip, and what is each worth? Q1 on Module 1 for 13 marks, Q2 on Module 2 for 12 marks, and Q3 on the journal and the viva for 5 marks. Thirty in all, in two hours.

2. How many marks is the journal worth? More than one figure. It earns part of the 20 internal marks under item 13, at 2.5 marks per practical and writeup, and it is examined again inside Q3 of the 30 mark external slip.

3. What happens if your journal is complete but unsigned? It is not a certified copy, and a certified copy is compulsory to appear for the practical examination.

4. How many practicals does MU expect in the journal, and how do you know? Twenty. Item 13 awards 2.5 marks for each practical performance and writeup to a total of 50 marks, and 2.5 multiplied by 20 is exactly 50.

5. Which module is worth more, and by how much? Module 1, by one mark: 13 against 12.

6. What language should Module 2 be written in? The one your laboratory uses. MU does not fix it: the Data Structures theory paper's Course Outcome 3 is expressly "using their chosen programming language". This book uses Python and gives the C versions in a chapter of their own.

7. What is the standard of passing, and why does it matter here? Forty per cent, with individual passing in internal and external. So the journal and the examination cannot cover for each other.

Contents This chapter on its own page

munotes.in4

Chapter Two

The Laboratory from Zero: Python, IDLE, and Your First Program

Syllabus topic Module 1's own subject, and MU's description of this paper as "a comprehensive exploration of advanced Python programming concepts"

Aim

To get Python, NumPy and an editor onto a machine, to run a program, and to read an error message, so that no time in the laboratory is spent on the tools.

Which Python

MU's exercises need Python 3. Every version from 3.10 onward runs everything in this book. The listings here were run on three versions, 3.14, 3.13 and 3.12, and any line whose behaviour differs between them is called out where it appears.

There is no Python 2 in this syllabus and there is no reason to install it. If a machine answers python with Python 2.7, that machine has an old system Python and the command you want is python3.

Getting it onto the machine

On Windows. Download the installer from the official Python site and run it. On the first screen, tick Add python.exe to PATH before clicking Install Now. If that box is missed, python will not be found at the Command Prompt and the usual fix is to run the installer again and choose Modify. The installer brings IDLE with it.

On Ubuntu or any Debian based Linux. Python 3 is already there. What is usually missing is the editor and the package installer:

sudo apt update
sudo apt install python3 python3-pip idle3

On macOS. The system Python is old and should be left alone. Install a current one from the official installer or with Homebrew, brew install python.

Check it, on any of the three:

python3 --version
python3 -c "print(2 + 2)"

The second line printing 4 is the whole test: the interpreter exists, it can run a statement, and it can print.

NumPy, because two of MU's exercises will not run without it

MU's practicals 3 and 4 of Module 1 are about arrays, and the arrays she means are NumPy arrays. NumPy is not part of Python and has to be installed once:

python3 -m pip install numpy

On Ubuntu, sudo apt install python3-numpy also works and is often already done in a college laboratory. Check it the same way:

python3 -c "import numpy; print(numpy.__version__)"

Anything from 1.20 onward runs every array listing in this book. The listings here were proved on two different major versions, 2.4.2 and 2.0.0, on purpose: a laboratory machine will not be on the version a book was written on, and an output that changes between versions is an output a student cannot rely on.

If that command answers ModuleNotFoundError: No module named 'numpy', NumPy went into a different Python than the one you are running. The python3 -m pip form above prevents that, because it installs into the interpreter you named.

munotes.in5

The Laboratory from Zero: Python, IDLE, and Your First Program

The two ways to run Python, and when each is right

The interactive prompt. Type python3 and you get >>>. Every line you type runs at once and its value is printed back. It is the right tool for checking one thing:

>>> 17 % 5
2
>>> "munotes"[0]
'm'
>>> exit()

Nothing typed at >>> is saved. It is a calculator, not a place to write a practical.

A saved file. A file called practical1.py holds the program, and

python3 practical1.py

runs it from top to bottom. This is what the examination expects, this is what goes in your journal, and this is what every listing in this book is.

IDLE, the editor MU's laboratories have

IDLE arrives with Python and needs nothing installed. Open it and you get the same >>> shell. To write a program, choose File, then New File, type the program, save it with a .py ending, and press F5 to run it. The output appears in the shell window.

Three things about IDLE that save time in a laboratory:

  • F5 saves before it runs. If it asks, say yes. Running an unsaved file is not

possible, which is a feature: what ran is what is on disk.

  • The Tab key indents, and IDLE indents for you after a line ending in a colon.

Python decides what is inside a block by indentation alone, so this matters more than in any language you have used before.

  • File, then Save As, into your own folder. A file saved into the Python

installation folder is a file you will not find again on a shared machine.

Any editor works. VS Code, Notepad++, nano and gedit are all fine. IDLE is described here because it is what is installed on a machine you did not set up.

Your first program, and what each line does

print("munotes")
print("Major Practical 3")
print(2 + 2)
munotes
Major Practical 3
4

Three statements, run in order, each printing one line. print puts a newline at the end by itself, which is why three calls give three lines.

Indentation is the syntax

This is the one rule that catches every student moving from C to Python. C marks a block with braces and the indentation is decoration. In Python the indentation IS the block, and there are no braces at all.

total = 0
for n in [1, 2, 3, 4]:
    total = total + n
    print("running total", total)
print("final", total)
running total 1
running total 3
running total 6
running total 10
final 10

The two indented lines are inside the loop and run four times. The last line is not indented, so it is outside the loop and runs once. Move that last line four spaces right and the program prints five lines instead of one. Nothing else changes.

munotes.in6

The Laboratory from Zero: Python, IDLE, and Your First Program

Use four spaces for one level. Do not mix spaces and tabs in one file: Python 3 refuses to run a file that does, with TabError, and the two look identical on screen.

Reading an error message

Python prints errors as a traceback, and the useful line is the last one. Read from the bottom up.

numbers = [10, 20, 30]
print(numbers[3])
IndexError: list index out of range

That is the last line of the traceback, and it is the one that matters: it names the kind of error, IndexError, and what was wrong, list index out of range. On your own screen several lines come before it, giving the file and the line number, and this book prints only the last line because the lines above it are laid out slightly differently in different versions of Python. Read from the bottom up and most errors are fixed in seconds.

Here are the four you will meet in week one, with what they actually mean.

The last line saysWhat it meansThe usual cause
SyntaxErrorPython could not even read the linea missing colon after if, for, def or while, or a missing bracket
IndentationErrora colon promised a block and none camethe line after if ...: is not indented
NameErrora name was used before it was given a valuea typo in a variable name, or using it above where it is set
TypeErrora string and a number were addedinput() returns a string and it was not converted with int()

The last of those four is the single commonest error in this whole paper, and [Practical 1: Input, Output, and the Year You Turn 100] is where it is met and killed.

Why the prompts in this book look run together

Every output block in this book is what the program actually printed, captured by feeding the input in from a file rather than typing it. That changes how a prompt looks, and it is worth understanding once so it never puzzles you again.

At your own keyboard, a program that asks two questions shows this, where the values in bold are what you typed:

What is your name? Aarti
How old are you? 19
Hello, Aarti

The same program, fed from a file, prints this, which is what a proved output block in this book holds:

What is your name? How old are you? Hello, Aarti

Nothing is wrong. A prompt has no newline at the end, and when nobody is typing there is nothing to push the next prompt onto its own line, and the typed value is never echoed because it was never typed. So the two prompts sit side by side. Your own run will put your answers in between.

munotes.in7

The Laboratory from Zero: Python, IDLE, and Your First Program

What to put in the journal when a program will not run

Not nothing, and not a blank page. An entry that records a failure honestly is worth marks; a missing entry is worth none.

Write the aim, write the program as you wrote it, write the error message in full, and then one or two sentences on what you think it means and what you tried. If you fix it later, add the corrected program underneath with a line saying what the fix was. A teacher signing your journal is looking for evidence that you sat and worked on it, and a traceback with your own diagnosis under it is exactly that.

Procedure

  1. Install Python 3 and confirm the version with python3 --version.
  2. Install NumPy with python3 -m pip install numpy and confirm it imports.
  3. Open IDLE, create a new file, save it as first.py in your own folder.
  4. Type the three line print program, press F5, and read the output in the shell.
  5. Add a for loop with an indented body and confirm the indentation decides what

repeats.

  1. Cause an IndexError on purpose and read the traceback from the bottom line up.

Result

Python, NumPy and an editor are installed and working, a program has been saved and run from a file, the role of indentation is understood, and a traceback can be read.

Where marks are lost

  • Working at the >>> prompt and then having nothing to show. A practical is a file.
  • Not adding Python to PATH on Windows, then spending laboratory time on it.
  • Installing NumPy into a different Python than the one being run. Use

python3 -m pip.

  • Mixing tabs and spaces, which produces TabError and looks like nothing at all on

screen.

  • Reading a traceback from the top. The bottom line is the error.
  • Leaving a failed practical out of the journal instead of recording the error.

For the journal

This is the setup and it has no numbered entry. On the page facing your index, write the version of Python and the version of NumPy on your laboratory machine, and the command that runs a file on it. Every entry afterwards can then say "run as before" instead of repeating it twenty times.

Quick revision

  • Python 3, any version from 3.10, and python3 --version proves it.
  • NumPy is separate: python3 -m pip install numpy. MU's practicals 3 and 4 need it.
  • >>> is for checking one thing. A practical is a saved .py file run as
munotes.in8

The Laboratory from Zero: Python, IDLE, and Your First Program

python3 file.py.

  • IDLE: File, New File, save with .py, F5 to run. F5 saves first.
  • Indentation is the block. Four spaces, never tabs mixed with spaces.
  • A traceback is read from the bottom line, which names the error and the reason.
  • The four week one errors: SyntaxError (missing colon or bracket),

IndentationError, NameError (typo or used too early), TypeError (a string added to a number, usually an unconverted input()).

  • A program that will not run still gets a journal entry, with the error message in full.

Questions you should be able to answer

1. What is the difference between the interactive prompt and a saved file? At >>> each line runs as it is typed and nothing is kept. A saved file is a program that runs top to bottom, can be re-run, and is what the journal and the examination require.

2. Why does python3 -m pip install numpy avoid a common problem? It installs into the interpreter you named. A bare pip may belong to a different Python, and then import numpy fails in the one you are running.

3. What decides which statements are inside a loop in Python? The indentation, and nothing else. There are no braces.

4. Which line of a traceback do you read first, and what does it tell you? The last one. It gives the kind of error and the reason, for example IndexError: list index out of range.

5. You wrote if marks > 40 and Python said SyntaxError: invalid syntax. Why? The colon is missing. if, elif, else, for, while and def all end their line with a colon.

6. age = input("Age: ") then age + 1 gives a TypeError. What is the fix? input() always returns a string. Convert it: age = int(input("Age: ")).

7. Why do the prompts in this book's output blocks run together on one line? Because the input was fed from a file rather than typed, so nothing was echoed and no newline was added after a prompt. At a keyboard your answers appear in between.

8. What goes in the journal for a practical that would not run? The aim, the program as written, the complete error message, and your own sentence or two on what it means and what you tried.

Contents This chapter on its own page

munotes.in9

Chapter Three

Practical 1: Input, Output, and the Year You Turn 100

Syllabus topic Module 1, practical 1(a), "Write a program that asks the user to enter their name and their age. Print out a message addressed to them that tells them the year that they will turn 100 years old."

Aim

To accept a name and an age from the user, and to print a message addressed to them giving the year in which they will turn 100.

What you need to know before you start

Three things, and the third is where almost everybody loses a mark.

input() reads one line from the keyboard. The text you pass it is printed first, as a prompt, and the line the user types is handed back.

input() always returns a string. Always. Even when the user types 19, what comes back is the two character string "19", not the number 19. Python will not guess.

A string and a number cannot be added. So the age has to be converted before any arithmetic, with int(). Forget it and the program dies with a TypeError, which is the error this exercise exists to teach.

age = input("How old are you? ")
print("Next year you will be", age + 1)
19
How old are you? TypeError: can only concatenate str (not "int") to str

Read the last line. A string can be joined to a string and a number can be added to a number, and age is a string. The fix is one call:

age = int(input("How old are you? "))
print("Next year you will be", age + 1)
19
How old are you? Next year you will be 20

The three ways to print, and which to use

The exercise asks for a message addressed to them, so the name has to appear inside the sentence. Python gives three ways of doing that, and it is worth seeing all three once, because an examiner may ask for a particular one.

name = "Aarti"
year = 2107

print("Hello " + name + ", you will turn 100 in " + str(year) + ".")
print("Hello", name + ", you will turn 100 in", year, end=".\n")
print(f"Hello {name}, you will turn 100 in {year}.")
Hello Aarti, you will turn 100 in 2107.
Hello Aarti, you will turn 100 in 2107.
Hello Aarti, you will turn 100 in 2107.
FormHow it joinsWatch out for
+ concatenationstrings onlyevery number needs str() around it, or TypeError
print with commasany typesit inserts a space between each item, whether you want one or not
f-stringany typesthe f before the quote is not optional

The f-string is the one to use. It reads like the sentence it produces, it needs no str(), and it puts spaces exactly where you type them. It has existed since Python 3.6 and every laboratory machine has it.

Getting the year without hard coding it

The year they turn 100 is the current year plus the years they have left to go. That needs today's year, and typing 2026 into the program is what makes it wrong in January. Ask the machine:

munotes.in10

Practical 1: Input, Output, and the Year You Turn 100

from datetime import date

this_year = date.today().year
print("the current year is", this_year)
print("a 19 year old turns 100 in", this_year + 100 - 19)
the current year is 2026
a 19 year old turns 100 in 2107

date.today() gives today's date and .year takes the year out of it, as an ordinary integer you can do arithmetic with.

The two numbers in that output block were true on the day this page was built and they will not be true when you read it, which is exactly the point of the section: the program is right in every year because it asks the machine. Every other output in this chapter is a number that does not move, and is printed in full.

The program

from datetime import date

name = input("What is your name? ")
age = int(input("How old are you? "))

this_year = date.today().year
years_to_go = 100 - age
century_year = this_year + years_to_go

print(f"Hello {name}, you are {age} years old.")
print(f"You have {years_to_go} years to go until you are 100.")
print(f"You will turn 100 in the year {century_year}.")
Aarti
19
What is your name? How old are you? Hello Aarti, you are 19 years old.
You have 81 years to go until you are 100.
You will turn 100 in the year 2107.

That is the answer, and it is the version that goes in your journal.

The year in the last line moved on the day this page was built and will move again, so here is the same arithmetic with the year held still, which is how you check your own result and how you explain it at the table:

name = "Aarti"
age = 19
this_year = 2026

years_to_go = 100 - age
century_year = this_year + years_to_go

print(f"{name} is {age}, so {100} - {age} = {years_to_go} years to go")
print(f"and {this_year} + {years_to_go} = {century_year}")
Aarti is 19, so 100 - 19 = 81 years to go
and 2026 + 81 = 2107

100 - 19 = 81, and 2026 + 81 = 2107. A person who is 19 in 2026 turns 100 in 2107.

The mark almost everybody misses: the birthday this year

The program above is the answer MU's row asks for, and it is what goes in the journal. But it is only exactly right for somebody whose birthday has already passed this year.

Think about it with a date. If Aarti is 19 today and her birthday was in March, she turns 20 next March, 21 the March after, and 100 in the year 2026 + 81 = 2107. Correct. But if her birthday is in December and she is still 19 in this calendar year, she turns 20 this December, and she will reach 100 one year earlier than the simple sum says.

munotes.in11

Practical 1: Input, Output, and the Year You Turn 100

The complete version asks for the date of birth instead of the age, and then there is nothing to guess:

from datetime import date

dob = date(2006, 12, 15)
today = date(2026, 3, 1)

age = today.year - dob.year - ((today.month, today.day) < (dob.month, dob.day))
century_year = dob.year + 100

print(f"born {dob}, so today the age is {age}")
print(f"the 100th birthday falls in {century_year}")
print(f"the simple sum would have said {today.year + (100 - age)}")
born 2006-12-15, so today the age is 19
the 100th birthday falls in 2106
the simple sum would have said 2107

The comparison (today.month, today.day) < (dob.month, dob.day) is worth understanding, because it is a neat Python idiom and a viva question. Python compares two tuples item by item: it looks at the months first and only looks at the days if the months are equal. So the whole expression is True when this year's birthday has not happened yet, and True counts as 1 when subtracted, which takes one year off the age.

The last two lines show the point. Born in December 2006, the hundredth birthday is in December 2106, but the naive sum from an age of 19 in March 2026 says 2107. One year out.

Put the simple version in your journal, because that is what MU's row asks for, and add two sentences underneath saying that the answer assumes this year's birthday has passed, and that using the date of birth removes the assumption. That sentence is worth a mark and takes fifteen seconds.

Procedure

  1. Open a new file and save it as practical1a.py.
  2. Read the name with input().
  3. Read the age with input() and convert it with int().
  4. Take the current year from date.today().year.
  5. Compute the years to go as 100 minus the age, and add them to the current year.
  6. Print the message with an f-string so the name is inside the sentence.
  7. Run it with your own name and age, and check the sum by hand.

Result

The program accepts a name and an age and prints a message addressed to the user giving the year in which they will turn 100. Checked: age 19 in 2026 gives 2107, and 2026 + 81 = 2107.

Where marks are lost

  • Forgetting int() around the age, so the program dies with a TypeError.
  • Hard coding the current year. The answer is then wrong from the next 1st of January.
  • Not addressing the user by name. MU's row says "addressed to them", so the name
munotes.in12

Practical 1: Input, Output, and the Year You Turn 100

goes inside the sentence, not on a line of its own.

  • Writing print("Hello " + name + ", you are " + age) with age a number, which is

the same TypeError from the other direction.

  • Using input without a prompt, so the run shows a blank screen and nobody knows

what to type.

  • No output written in the journal. The output is the evidence that it ran.

For the journal

Write the aim in MU's own words. Then the program, the one that reads the year from date.today(). Then the run, with your own name and age and the year it printed. Then two sentences: that input() returns a string so the age must be converted with int(), and that the answer assumes this year's birthday has passed, which reading the date of birth would remove. The conclusion: a program that reads the year from the system rather than having it typed in stays correct next year.

Quick revision

  • input(prompt) prints the prompt and returns the line typed, always as a string.
  • int() converts it. Without that, age + 1 is a TypeError.
  • Three ways to print a sentence with a value in it: + with str(), print with

commas which inserts spaces, and the f-string, which is the one to use.

  • from datetime import date, then date.today().year is the current year. Never type

the year into the program.

  • Years to go = 100 - age. Year of the 100th birthday = current year + years to go.
  • Worked: 19 years old in 2026 gives 100 - 19 = 81 and 2026 + 81 = 2107.
  • The simple sum assumes this year's birthday has passed. From a date of birth,

age = today.year - dob.year - ((today.month, today.day) < (dob.month, dob.day)).

  • Python compares tuples item by item, which is why that one line works.

Questions you should be able to answer

1. What type does input() return? A string, always, even when the user types digits.

2. Why does int(input(...)) appear so often in this paper? Because almost every exercise does arithmetic on what was typed, and arithmetic on a string either fails with a TypeError or does something unwanted: multiplying the string "2" by 3 gives "222", not 6.

3. Write the line that prints a name and a year in one sentence. print(f"Hello {name}, you will turn 100 in {year}.")

4. Why is year = 2026 in the program a defect? Because the program gives a wrong answer from the next 1st of January. date.today().year asks the machine instead.

5. A student is 19 in 2026. In which year do they turn 100, and show the sum. 100 - 19 = 81 years to go, and 2026 + 81 = 2107.

munotes.in13

Practical 1: Input, Output, and the Year You Turn 100

6. When is that answer one year out, and why? When this year's birthday has not happened yet. The person is still 19 but will turn 20 later this year, so they reach 100 one year sooner than the sum from the age says.

7. What does the tuple comparison in the date version do? (today.month, today.day) < (dob.month, dob.day) compares item by item, months first and days only if the months are equal, so it is True exactly when this year's birthday is still to come. True counts as 1, so subtracting it takes one year off the age.

8. What is the difference between print(a, b) and print(a + b) for two strings? The first prints them with a space between; the second joins them with no space. The first also works when b is a number and the second does not.

Contents This chapter on its own page

munotes.in14

Chapter Four

Practical 1 continued: Even or Odd, and the SGPI Grade Ladder

Syllabus topic Module 1, practical 1(b), "Write a program to accept a number from the user and depending on whether the number is even or odd, print out an appropriate message to the user", and 1(c) to 1(l), "Write a program to accept the SGPI from the user and print corresponding grade based on the following", with her eight band table from O to F

Aim

To accept a number and report whether it is even or odd, and to accept an SGPI and print the grade for it from MU's own eight band table.

Part one: even or odd

What decides it

A whole number is even when dividing it by 2 leaves nothing over, and odd when it leaves 1 over. Python's remainder operator is %, called modulo:

for n in [10, 11, 0, 7, 100]:
    print(n, "% 2 =", n % 2)
10 % 2 = 0
11 % 2 = 1
0 % 2 = 0
7 % 2 = 1
100 % 2 = 0

So the test is n % 2 == 0. Note the two equals signs: one equals sign assigns a value, two compare. That is a SyntaxError in Python if you get it wrong inside an if, which is a mercy, because in C it compiles and does the wrong thing silently.

The program

number = int(input("Enter a number: "))

if number % 2 == 0:
    print(f"{number} is an even number.")
else:
    print(f"{number} is an odd number.")
17
Enter a number: 17 is an odd number.

The two cases that catch a careless answer

Zero. Zero divided by 2 leaves nothing over, so zero is even. Many students expect a program to say something special about it, and it should not: the mathematics is settled.

Negative numbers. In C, -7 % 2 is -1, so a C programmer's habit of testing n % 2 == 1 fails for negative odd numbers. Python does not behave that way. Python's remainder always takes the sign of the divisor, so -7 % 2 is 1, and testing for 1 happens to work. Testing == 0 for even works in both languages and is the habit to keep.

for n in [0, -7, -8, 7, 8]:
    print(f"{n:>3} % 2 = {n % 2:>2}   even? {n % 2 == 0}")
  0 % 2 =  0   even? True
 -7 % 2 =  1   even? False
 -8 % 2 =  0   even? True
  7 % 2 =  1   even? False
  8 % 2 =  0   even? True

The {n:>3} inside the f-string means "right align this in 3 characters", which is how a table of numbers is lined up in output. It is worth knowing for every later practical.

What if the user types something that is not a number

int("hello") raises ValueError. The complete answer catches it:

raw = input("Enter a number: ")
try:
    number = int(raw)
except ValueError:
    print(f"'{raw}' is not a whole number.")
else:
    print(f"{number} is", "even." if number % 2 == 0 else "odd.")
twelve
Enter a number: 'twelve' is not a whole number.
munotes.in15

Practical 1 continued: Even or Odd, and the SGPI Grade Ladder

try and except are worth one sentence in the journal and one line of code. MU does not ask for them in this row, so the plain version is the answer, and this is the improvement to mention underneath.

Part two: MU's SGPI grade ladder

The table MU prints

This is her own table, reproduced exactly as the rows appear in the syllabus:

SGPIGrade
9.00 to 10.00O
8.00 to 8.99A+
7.00 to 7.99A
6.00 to 6.99B+
5.50 to 5.99B
5.00 to 5.49C
4.00 to 4.99P
Below 4F

SGPI is the Semester Grade Performance Index, the weighted average of the grade points you earned in one semester, on a ten point scale. O is Outstanding, A+ and A are the two first classes, B+ and B the two seconds, C is a pass class, P is a bare pass and F is a fail.

The idea that makes the program short

Eight bands sounds like eight tests. It is not, and seeing why is the whole exercise.

Work downwards from the top. If the SGPI is 9.00 or more, the grade is O and nothing else needs testing. If it was not, then it is below 9.00, so "8.00 or more" is enough to mean 8.00 to 8.99 without ever writing the upper limit. Each rung of the ladder only has to name its own floor, because every rung above it has already been ruled out.

That is what elif is for. It means "else, if", and only the first branch whose condition is true runs.

The program

sgpi = float(input("Enter your SGPI: "))

if sgpi > 10:
    grade = "invalid, the scale ends at 10.00"
elif sgpi >= 9.00:
    grade = "O"
elif sgpi >= 8.00:
    grade = "A+"
elif sgpi >= 7.00:
    grade = "A"
elif sgpi >= 6.00:
    grade = "B+"
elif sgpi >= 5.50:
    grade = "B"
elif sgpi >= 5.00:
    grade = "C"
elif sgpi >= 4.00:
    grade = "P"
else:
    grade = "F"

print(f"SGPI {sgpi:.2f} gives grade {grade}")
7.85
Enter your SGPI: SGPI 7.85 gives grade A

Three things to notice in that listing.

float, not int. An SGPI is 7.85, not 7. int("7.85") raises ValueError, and int(7.85) would throw the decimals away and turn an A into an A. Every mark in the middle of a band would land on the wrong rung.

The bands are written in descending order. Write them ascending and the first condition, sgpi >= 4.00, is true for almost everybody, so almost everybody gets a P. That is the single commonest wrong answer to this exercise.

The sgpi > 10 guard comes first. Without it, an SGPI of 47 quietly gets an O.

munotes.in16

Practical 1 continued: Even or Odd, and the SGPI Grade Ladder

Every band tested, which is what proves it

One run proves one band. Testing all eight, and both sides of every junction, is what turns a program into an answer:

def grade_for(sgpi):
    if sgpi > 10:
        return "invalid"
    if sgpi >= 9.00:
        return "O"
    if sgpi >= 8.00:
        return "A+"
    if sgpi >= 7.00:
        return "A"
    if sgpi >= 6.00:
        return "B+"
    if sgpi >= 5.50:
        return "B"
    if sgpi >= 5.00:
        return "C"
    if sgpi >= 4.00:
        return "P"
    return "F"


for value in [10.00, 9.00, 8.99, 8.00, 7.99, 7.00, 6.99, 6.00,
              5.99, 5.50, 5.49, 5.00, 4.99, 4.00, 3.99, 0.00]:
    print(f"{value:>6.2f} -> {grade_for(value)}")
 10.00 -> O
  9.00 -> O
  8.99 -> A+
  8.00 -> A+
  7.99 -> A
  7.00 -> A
  6.99 -> B+
  6.00 -> B+
  5.99 -> B
  5.50 -> B
  5.49 -> C
  5.00 -> C
  4.99 -> P
  4.00 -> P
  3.99 -> F
  0.00 -> F

Read that output against MU's table row by row. Every band's floor and every band's ceiling appears, and each lands where her table says it should.

Notice that a series of plain if statements with a return in each does the same job as the elif ladder. It works because return leaves the function at once, so nothing below it can run. Either form is correct; the elif ladder is the one to write when there is no function.

The hole in MU's table, and what to do about it

Look at the two middle bands again: C is 5.00 to 5.49 and B is 5.50 to 5.99. Now ask what grade an SGPI of exactly 5.495 gets. It is above 5.49 and below 5.50, so MU's table does not say. The same gap sits between every pair of bands: 8.99 to 9.00, 7.99 to 8.00, and so on.

This is not a trick question, it is how marks are actually recorded: an SGPI is published to two decimal places, so 5.495 never occurs. The program above handles it anyway, and it is worth knowing how, because it is a viva question.

def grade_for(sgpi):
    if sgpi >= 9.00:
        return "O"
    if sgpi >= 8.00:
        return "A+"
    if sgpi >= 5.50:
        return "B"
    if sgpi >= 5.00:
        return "C"
    return "below C"


for value in [5.495, 5.50, 8.995, 9.00]:
    print(f"{value:>7.3f} -> {grade_for(value)}")
  5.495 -> C
  5.500 -> B
  8.995 -> A+
  9.000 -> O

Because each rung tests "greater than or equal to its own floor" and the ladder runs downwards, a value in a gap falls through to the band below it: 5.495 gets a C and 8.995 gets an A+. That is a decision the program makes, so say it out loud in the journal: at a value between two bands this ladder awards the lower grade, and at an exact band floor it awards the higher one. An examiner who asks about the gap is checking whether you know your own program's behaviour, not whether you can recite MU's table.

munotes.in17

Practical 1 continued: Even or Odd, and the SGPI Grade Ladder

Procedure

  1. Save the even or odd program as practical1b.py. Read the number with int(input()).
  2. Test number % 2 == 0 and print with an f-string in each branch.
  3. Run it with an odd number, an even number, zero and a negative number.
  4. Save the grade program as practical1c.py. Read the SGPI with float(input()).
  5. Write the eight bands as an elif ladder in descending order, each testing only

its own floor, with a guard above 10 and else for F.

  1. Run it once for each of MU's eight bands and at both ends of each band.
  2. Write both programs, both runs, and the note about the band boundaries in the journal.

Result

Even or odd is decided by n % 2 == 0, proved on positive, negative and zero. The grade ladder was run at the floor and the ceiling of all eight of MU's bands and every value landed in the band her table gives, with the boundary rule stated.

Where marks are lost

  • Writing the bands in ascending order. Then sgpi >= 4.00 catches nearly everybody

and the program prints P for an SGPI of 9.

  • Using int() for the SGPI. It refuses 7.85 outright, and truncating would put

marks in the wrong band.

  • Writing each band as two tests, if 8.00 <= sgpi <= 8.99, which is not wrong but is

twice the code and leaves the gaps between bands unhandled, so an SGPI of 8.995 gets no grade at all.

  • A chain of separate if statements instead of elif with no return in them: then

several branches run and the last one wins, which is usually F.

  • No guard above 10, so an SGPI of 47 gets an O.
  • Testing n % 2 == 1 for odd out of C habit. It happens to work in Python, but the

reason it works is worth knowing and == 0 for even is the safer habit.

  • Only one run. Eight bands need eight runs to be an answer.

For the journal

Two entries under one practical number. For the even or odd program: the aim in MU's words, the program, runs with an odd number, an even number, zero and a negative number, and the conclusion that % gives the remainder and zero is even. For the grade program: the aim, MU's own eight band table copied out, the program, a run for each band, and two sentences saying that the ladder is written downwards so each band needs only its floor, and that a value between two bands takes the lower grade.

munotes.in18

Practical 1 continued: Even or Odd, and the SGPI Grade Ladder

Quick revision

  • % is the remainder. Even is n % 2 == 0, odd is anything else. Zero is even.
  • In Python the remainder takes the sign of the divisor, so -7 % 2 is 1, not -1 as in C.
  • == compares, = assigns.
  • int(input()) for a whole number, float(input()) for an SGPI.
  • elif means else-if, and only the first true branch runs.
  • Write the bands downwards, each testing only its own floor: >= 9.00 is O, and if

that failed then >= 8.00 already means 8.00 to 8.99.

  • MU's bands: 9.00 O, 8.00 A+, 7.00 A, 6.00 B+, 5.50 B, 5.00 C, 4.00 P, below 4 F.
  • Guard above 10.00 first, else for F last.
  • A value in a gap between bands falls to the lower band; a value on a floor takes the

higher. Say so in the journal.

  • {value:>6.2f} in an f-string right aligns a number in 6 characters with 2 decimals.

Questions you should be able to answer

1. How do you test whether a number is even? n % 2 == 0. The remainder on dividing by 2 is zero.

2. Is zero even or odd? Even. Zero divided by 2 leaves no remainder.

3. What is -7 % 2 in Python, and why does that differ from C? It is 1. Python's remainder takes the sign of the divisor, so it is never negative for a positive divisor. C's takes the sign of the dividend, giving -1.

4. Why must the SGPI be read with float and not int? Because an SGPI has decimals. int("7.85") raises ValueError, and truncating to 7 would move a mark into the wrong band.

5. Why are the grade bands written in descending order? Because each elif then only needs its own floor: if >= 9.00 has already failed, >= 8.00 can only mean 8.00 to 8.99. Written upwards, the first test catches almost every value and everyone gets a P.

6. What is the difference between eight if statements and an elif ladder? In a ladder only the first true branch runs. With eight separate if statements every true one runs, so several assignments happen and the last wins, which for a high SGPI is F. Separate if statements are safe only when each one ends in return.

7. MU's table jumps from 5.49 to 5.50. What grade does 5.495 get, and why? A C. Each rung tests "greater than or equal to its own floor" going downwards, so a value in a gap falls through to the band below. Real SGPIs are published to two decimals, so the case does not arise, but the program has an answer for it.

munotes.in19

Practical 1 continued: Even or Odd, and the SGPI Grade Ladder

8. What does the guard if sgpi > 10 protect against? An impossible SGPI silently getting the top grade. Without it, 47 is greater than 9.00 and prints O.

Contents This chapter on its own page

munotes.in20

Chapter Five

Practical 2: the Fibonacci Series, and the Sum of the Digits

Syllabus topic Module 1, practical 2, "Write a program to generate the Fibonacci series" and "Write a program to accept a number from the user display sum of its digits"

Aim

To generate the Fibonacci series, and to accept a number from the user and display the sum of its digits.

Part one: the Fibonacci series

What the series is

Each term is the sum of the two before it. The series starts 0 and 1, and then:

TermWorked fromValue
1stgiven0
2ndgiven1
3rd0 + 11
4th1 + 12
5th1 + 23
6th2 + 35
7th3 + 58
8th5 + 813

Some books start the series at 1 and 1 instead of 0 and 1. Both are used. Say in your journal which start you used, and be ready to change it if the examiner asks for the other, which is a one character edit.

The answer: two variables and a loop

count = int(input("How many terms? "))

a, b = 0, 1
for _ in range(count):
    print(a, end=" ")
    a, b = b, a + b
print()
10
How many terms? 0 1 1 2 3 5 8 13 21 34

Two lines in that program are worth all the marks.

a, b = 0, 1 assigns both at once. It is Python's tuple assignment: the right hand side is built first, then unpacked into the names on the left.

a, b = b, a + b is the whole series in one line, and it is right for exactly that reason. The right hand side is worked out completely before anything is assigned, so a + b uses the old a, not the new one. In C you would need a temporary variable:

temp = a + b;
a = b;
b = temp;

Try it without the tuple assignment and get it wrong, because seeing the wrong answer is what fixes the idea:

a, b = 0, 1
for _ in range(8):
    print(a, end=" ")
    a = b          # a is overwritten first
    b = a + b      # so this adds the NEW a, which is wrong
print()
0 1 2 4 8 16 32 64

That prints the powers of two, not the Fibonacci series, because by the time the second line runs the old value of a is gone. It is the commonest bug in this exercise.

The _ in for _ in range(count) is an ordinary variable name, used by convention when the loop counter is not needed. Writing for i in range(count) is equally correct.

The same series three other ways

An examiner may ask for any of these, and each teaches something.

Into a list, when the terms are needed afterwards rather than just printed:

def fibonacci_list(count):
    series = [0, 1]
    while len(series) < count:
        series.append(series[-1] + series[-2])
    return series[:count]


print(fibonacci_list(10))
print(fibonacci_list(1))
print(fibonacci_list(0))
munotes.in21

Practical 2: the Fibonacci Series, and the Sum of the Digits

[0, 1, 1, 2, 3, 5, 8, 13, 21, 34]
[0]
[]

series[-1] is the last item and series[-2] the one before it. The [:count] at the end handles the small cases: the list starts with two terms, so asking for one or for none would otherwise give too many.

Up to a limit rather than a count, which MU's wording also permits:

limit = 100
a, b = 0, 1
while a <= limit:
    print(a, end=" ")
    a, b = b, a + b
print()
0 1 1 2 3 5 8 13 21 34 55 89

By recursion, which is the version to know about and not to use:

def fib(n):
    if n <= 1:
        return n
    return fib(n - 1) + fib(n - 2)


print([fib(n) for n in range(10)])
[0, 1, 1, 2, 3, 5, 8, 13, 21, 34]

That is a direct translation of the definition and it is beautifully short. It is also wasteful in a way worth measuring, because the examiner's question is "why is this slow".

calls = 0


def fib_counted(n):
    global calls
    calls += 1
    if n <= 1:
        return n
    return fib_counted(n - 1) + fib_counted(n - 2)


for n in [5, 10, 15, 20, 25, 30]:
    calls = 0
    value = fib_counted(n)
    print(f"fib({n:>2}) = {value:>6}   and it took {calls:>7} calls to work out")
fib( 5) =      5   and it took      15 calls to work out
fib(10) =     55   and it took     177 calls to work out
fib(15) =    610   and it took    1973 calls to work out
fib(20) =   6765   and it took   21891 calls to work out
fib(25) =  75025   and it took  242785 calls to work out
fib(30) = 832040   and it took 2692537 calls to work out

The call count is the explanation, and it is the same on every machine. Recursion recomputes the same terms over and over. To get fib(25) it computes fib(23) twice, fib(22) three times and so on down the tree, which is why the calls run into the hundreds of thousands. Look at the column: the calls roughly multiply by 11 each time n goes up by 5, while the loop version does n additions and nothing more.

That is what makes recursion here the wrong tool for a series, and saying it with the call count is a much better answer than saying it is slow.

Part two: the sum of the digits

The arithmetic method

Two operators do the whole job. % 10 gives the last digit, and // 10 throws the last digit away. // is floor division, which divides and discards the remainder; plain / would give a float and break everything.

munotes.in22

Practical 2: the Fibonacci Series, and the Sum of the Digits

n = 4729
while n > 0:
    print(f"n = {n:>5}   last digit {n % 10}   rest {n // 10}")
    n = n // 10
n =  4729   last digit 9   rest 472
n =   472   last digit 2   rest 47
n =    47   last digit 7   rest 4
n =     4   last digit 4   rest 0

So the program is that loop with the digits added up instead of printed:

number = int(input("Enter a number: "))

n = abs(number)
total = 0
while n > 0:
    total = total + n % 10
    n = n // 10

print(f"The sum of the digits of {number} is {total}.")
4729
Enter a number: The sum of the digits of 4729 is 22.

Check it by hand: 4 + 7 + 2 + 9 = 22.

Three details in that program earn marks.

abs(number) first. Without it, a negative number never enters the loop, because -4729 > 0 is false, and the program prints a sum of 0. With abs, the digits of -4729 add to 22.

The original number is kept for the message, and a separate n is consumed by the loop. Print number after the loop without that and it prints 0, because the loop ate it.

Zero. 0 > 0 is false, so the loop never runs and the total stays 0, which is the right answer for zero. That is luck rather than design, so it is worth one sentence in the journal saying you checked it.

The string method

number = int(input("Enter a number: "))

total = sum(int(digit) for digit in str(abs(number)))
print(f"The sum of the digits of {number} is {total}.")
4729
Enter a number: The sum of the digits of 4729 is 22.

str(abs(number)) turns the number into its digits as characters, int(digit) turns each character back into a number, and sum adds them. One line, and perfectly correct.

Which to put in the journal? The arithmetic one, with the string one underneath as an alternative. The arithmetic version is what the exercise is teaching, because % 10 and // 10 are how digits are taken apart in every language, and an examiner may well ask for it "without converting to a string".

Every case tested

def digit_sum(number):
    n = abs(number)
    total = 0
    while n > 0:
        total = total + n % 10
        n = n // 10
    return total


for value in [4729, 0, 7, -4729, 1000, 999999999]:
    digits = " + ".join(str(d) for d in str(abs(value)))
    print(f"{value:>10}  ->  {digits} = {digit_sum(value)}")
      4729  ->  4 + 7 + 2 + 9 = 22
         0  ->  0 = 0
         7  ->  7 = 7
     -4729  ->  4 + 7 + 2 + 9 = 22
      1000  ->  1 + 0 + 0 + 0 = 1
 999999999  ->  9 + 9 + 9 + 9 + 9 + 9 + 9 + 9 + 9 = 81
munotes.in23

Practical 2: the Fibonacci Series, and the Sum of the Digits

The digit sum is also useful to know for its own sake: a number is divisible by 3 exactly when its digit sum is, and by 9 exactly when its digit sum is. 4 + 7 + 2 + 9 = 22, which is not divisible by 3, so 4729 is not either.

Procedure

  1. Save the series program as practical2a.py. Read the number of terms with

int(input()).

  1. Set a, b = 0, 1 and loop count times, printing a and then doing

a, b = b, a + b.

  1. Run it for 10 terms and check the first eight against the table by hand.
  2. Deliberately replace the tuple assignment with two separate assignments and see the

wrong answer, then put it back.

  1. Save the digit sum as practical2b.py. Read the number, take abs, and loop with

% 10 and // 10.

  1. Run it with a positive number, zero and a negative number, and check the sum by hand.

Result

The series program printed 0 1 1 2 3 5 8 13 21 34 for ten terms, which matches the table term by term. The digit sum of 4729 was 22, and 4 + 7 + 2 + 9 = 22. Zero gave 0 and -4729 gave 22.

Where marks are lost

  • Two separate assignments instead of a, b = b, a + b, which prints the powers of

two and looks almost right.

  • Printing one term too many or too few. Ten terms means ten numbers; count the

output.

  • Using / instead of // for the digit sum, which turns the number into a float and

loops for ever or crashes.

  • No abs(), so a negative number gives a sum of 0.
  • Consuming the number in the loop and then printing it, so the message says 0.
  • Offering the recursive version as the answer without being able to say why it is

slow. The reason is that it recomputes the same terms, and the call count proves it.

  • Only one run. Zero and a negative number are two more lines of output and two more

marks.

For the journal

Two entries under practical 2. For the series: the aim, the table of the first eight terms worked from the two before, the program, the run for ten terms, and one sentence saying that a, b = b, a + b works because the right hand side is evaluated in full before anything is assigned. For the digit sum: the aim, the program, runs for a positive number, zero and a negative number, the hand check 4 + 7 + 2 + 9 = 22, and one sentence on why abs is needed.

munotes.in24

Practical 2: the Fibonacci Series, and the Sum of the Digits

Quick revision

  • Fibonacci: each term is the sum of the two before. Starts 0, 1, then 1, 2, 3, 5, 8, 13.
  • a, b = 0, 1 and a, b = b, a + b. The right side is built before anything is

assigned, so no temporary variable is needed.

  • Two separate assignments give the powers of two. That is the bug to recognise.
  • series[-1] is the last item, series[-2] the one before.
  • Recursive Fibonacci is the definition written out, and it recomputes terms, so it is

exponential. The call count is the evidence.

  • Digit sum: % 10 takes the last digit, // 10 removes it. Loop while the number is

greater than 0.

  • // is floor division. Plain / gives a float and breaks the loop.
  • abs() first, or negative numbers give 0. Keep the original number for the message.
  • One line alternative: sum(int(d) for d in str(abs(n))).
  • 4 + 7 + 2 + 9 = 22, so the digit sum of 4729 is 22.

Questions you should be able to answer

1. Write the Fibonacci series in one loop. a, b = 0, 1, then repeat: print a, and a, b = b, a + b.

2. Why does a, b = b, a + b not need a temporary variable? Because Python builds the whole right hand side first, as a tuple, and only then unpacks it into a and b. So a + b uses the old a.

3. What does the program print if you write a = b and then b = a + b? The powers of two: 0 1 2 4 8 and so on. The old a has already been lost when the second line runs.

4. Why is the recursive Fibonacci slow? It recomputes the same terms many times. Computing fib(25) recomputes fib(23) twice, fib(22) three times and so on, which here came to hundreds of thousands of calls.

5. Which two operators take a number apart into digits? % 10 gives the last digit, // 10 removes it.

6. Why // and not /? / always gives a float in Python 3, so the number never becomes exactly 0 and the loop misbehaves. // divides and discards the remainder.

7. What does your digit sum program print for -4729, and why? 22, because abs() is applied first. Without it the loop condition is false straight away and the answer would be 0.

munotes.in25

Practical 2: the Fibonacci Series, and the Sum of the Digits

8. What is the digit sum of 4729, and one thing it tells you? 22. Since 22 is not divisible by 3, neither is 4729.

Contents This chapter on its own page

munotes.in26

Chapter Six

Practical 3: Arrays, Basic Operations, Indexing and Slicing

Syllabus topic Module 1, practical 3(a), "Write a program to perform basic operations, indexing and slicing on arrays"

Aim

To create arrays, to perform the basic operations on them, and to index and slice them.

First, three things called an array

This confuses more students on this paper than any other single point, because Python has three different things that a teacher may call an array.

What it isImportHolds
listPython's built in sequencenothinganything, mixed
array.arraya compact array of one numeric typefrom array import arrayone type, fixed by a type code
numpy.ndarraythe numerical array, any number of dimensionsimport numpy as npone type, with mathematics built in

A list is what you get from [1, 2, 3]. It can hold a number, a string and another list all at once, because each slot holds a reference to an object rather than the value itself. That flexibility costs memory and speed.

An array.array is a real array in the C sense: a contiguous block of one numeric type. It is part of the standard library and needs nothing installed.

A NumPy array is the one every later exercise means. It is also one contiguous block of one type, and it brings element by element arithmetic, multiple dimensions, and a large library of mathematical functions.

from array import array
import numpy as np

py_list = [3, 1, 4, 1, 5]
std_array = array('i', [3, 1, 4, 1, 5])
np_array = np.array([3, 1, 4, 1, 5])

print("list        ", py_list, type(py_list).__name__)
print("array.array ", std_array.tolist(), std_array.typecode, "itemsize", std_array.itemsize)
print("numpy array ", np_array, np_array.dtype)

print("a list can hold anything:", [1, "two", 3.0, [4]])
list         [3, 1, 4, 1, 5] list
array.array  [3, 1, 4, 1, 5] i itemsize 4
numpy array  [3 1 4 1 5] int64
a list can hold anything: [1, 'two', 3.0, [4]]

The 'i' in array('i', ...) is the type code for a signed integer. 'd' is a double, 'f' a float, 'b' a signed byte. An array.array refuses a value of the wrong type, which is the point of it:

from array import array

numbers = array('i', [1, 2, 3])
numbers.append(4.5)
TypeError: 'float' object cannot be interpreted as an integer

Creating an array

import numpy as np

print(np.array([3, 1, 4, 1, 5, 9, 2, 6]))
print(np.array([[1, 2, 3], [4, 5, 6]]))
print(np.zeros(5))
print(np.ones(4, dtype=int))
print(np.full(4, 7))
print(np.arange(0, 10, 2))
print(np.linspace(0, 1, 5))
print(np.eye(3, dtype=int))
[3 1 4 1 5 9 2 6]
[[1 2 3]
 [4 5 6]]
[0. 0. 0. 0. 0.]
[1 1 1 1]
[7 7 7 7]
[0 2 4 6 8]
[0.   0.25 0.5  0.75 1.  ]
[[1 0 0]
 [0 1 0]
 [0 0 1]]
CallWhat it makes
np.array(list)an array from a list, or a list of lists
np.zeros(n)n zeros, as floats
np.ones(n)n ones
np.full(n, v)n copies of v
np.arange(start, stop, step)like range, and the stop is excluded
np.linspace(start, stop, count)count values evenly spaced, and the stop IS included
np.eye(n)the identity matrix, ones down the diagonal
munotes.in27

Practical 3: Arrays, Basic Operations, Indexing and Slicing

The difference between arange and linspace is a favourite question: arange takes a step and excludes the stop, linspace takes a count and includes it.

Indexing

import numpy as np

a = np.array([10, 20, 30, 40, 50, 60])

print("the whole array   ", a)
print("a[0]  first       ", a[0])
print("a[3]               ", a[3])
print("a[-1] last        ", a[-1])
print("a[-2] second last ", a[-2])
print("len(a)            ", len(a))
the whole array    [10 20 30 40 50 60]
a[0]  first        10
a[3]                40
a[-1] last         60
a[-2] second last  50
len(a)             6

Indexing starts at 0, so the first item is a[0] and the last is a[len(a) - 1]. A negative index counts from the end, so a[-1] is the last item. This is worth learning properly because it removes a whole class of off by one errors: to get the last item you never need to know the length.

Going past the end raises an error rather than returning rubbish:

import numpy as np

a = np.array([10, 20, 30])
print(a[5])
IndexError: index 5 is out of bounds for axis 0 with size 3

For a two dimensional array, one pair of brackets holds both indexes, separated by a comma:

import numpy as np

m = np.array([[1, 2, 3, 4],
              [5, 6, 7, 8],
              [9, 10, 11, 12]])

print(m)
print("m[1, 2] row 1 column 2 :", m[1, 2])
print("m[0]    the whole row 0:", m[0])
print("m[:, 1] the whole col 1:", m[:, 1])
print("m[-1, -1] bottom right :", m[-1, -1])
[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]
m[1, 2] row 1 column 2 : 7
m[0]    the whole row 0: [1 2 3 4]
m[:, 1] the whole col 1: [ 2  6 10]
m[-1, -1] bottom right : 12

m[1][2] also works and gives the same value, but m[1, 2] is the NumPy way and is faster, because m[1][2] builds the whole row as an intermediate array first.

Slicing

A slice takes a piece of the array. It is written start:stop:step, and every part may be left out.

The stop is excluded. a[1:4] gives items 1, 2 and 3, not 4. This is the single most important fact about slicing, and the reason it is designed that way is worth knowing: a[:n] and a[n:] between them give the whole array exactly once, with nothing repeated and nothing missed.

import numpy as np

a = np.array([10, 20, 30, 40, 50, 60, 70])

print("a          ", a)
print("a[1:4]     ", a[1:4])
print("a[:3]      ", a[:3])
print("a[4:]      ", a[4:])
print("a[:]       ", a[:])
print("a[::2]     ", a[::2])
print("a[1::2]    ", a[1::2])
print("a[::-1]    ", a[::-1])
print("a[-3:]     ", a[-3:])
print("a[2:2]     ", a[2:2], "an empty slice, not an error")
print("a[1:100]   ", a[1:100], "an out of range stop is clipped, not an error")
munotes.in28

Practical 3: Arrays, Basic Operations, Indexing and Slicing

a           [10 20 30 40 50 60 70]
a[1:4]      [20 30 40]
a[:3]       [10 20 30]
a[4:]       [50 60 70]
a[:]        [10 20 30 40 50 60 70]
a[::2]      [10 30 50 70]
a[1::2]     [20 40 60]
a[::-1]     [70 60 50 40 30 20 10]
a[-3:]      [50 60 70]
a[2:2]      [] an empty slice, not an error
a[1:100]    [20 30 40 50 60 70] an out of range stop is clipped, not an error
SliceMeans
a[1:4]items 1, 2, 3. The stop is excluded
a[:3]from the start up to but not including 3
a[4:]from 4 to the end
a[:]the whole array
a[::2]every second item
a[::-1]the array reversed
a[-3:]the last three items

Two behaviours in that output are worth remembering because they are the opposite of indexing: an empty slice is not an error, and nor is a stop past the end. a[5] on a three item array raises IndexError; a[1:100] quietly gives what there is. That is deliberate, and it is why slicing is the safe way to take a piece of something whose length you are not sure of.

Slicing a two dimensional array takes a slice in each direction:

import numpy as np

m = np.arange(1, 13).reshape(3, 4)

print(m)
print("m[0:2, 1:3] first two rows, middle two columns")
print(m[0:2, 1:3])
print("m[:, ::2] every second column")
print(m[:, ::2])
print("m[::-1] the rows reversed")
print(m[::-1])
[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]
m[0:2, 1:3] first two rows, middle two columns
[[2 3]
 [6 7]]
m[:, ::2] every second column
[[ 1  3]
 [ 5  7]
 [ 9 11]]
m[::-1] the rows reversed
[[ 9 10 11 12]
 [ 5  6  7  8]
 [ 1  2  3  4]]

The one that will be asked: a slice is a view

Slicing a Python list gives a new list. Slicing a NumPy array gives a view, which is a window onto the same memory. Write through the window and the original changes.

import numpy as np

py = [10, 20, 30, 40, 50]
piece = py[1:4]
piece[0] = 999
print("list  :", py, "and the slice", piece)

arr = np.array([10, 20, 30, 40, 50])
window = arr[1:4]
window[0] = 999
print("array :", arr, "and the slice", window)
print("does the slice own its data?", window.base is None)
list  : [10, 20, 30, 40, 50] and the slice [999, 30, 40]
array : [ 10 999  30  40  50] and the slice [999  30  40]
does the slice own its data? False
munotes.in29

Practical 3: Arrays, Basic Operations, Indexing and Slicing

The list is untouched and the array is changed. This is the most important difference between a list and a NumPy array, and it is the whole subject of the next chapter, which is MU's own third bullet on aliasing and copying. It is also not a defect: a view costs no memory and no copying, which is why NumPy is fast on large data.

window.base is the array a view looks into, and it is None for an array that owns its own memory. That one line is the quickest way to answer "is this a view or a copy" at the table.

The basic operations MU asks for

import numpy as np

a = np.array([3, 1, 4, 1, 5, 9, 2, 6])

print("array        ", a)
print("length       ", len(a), "or a.size", a.size)
print("sum          ", a.sum())
print("smallest     ", a.min(), "at index", a.argmin())
print("largest      ", a.max(), "at index", a.argmax())
print("mean         ", a.mean())
print("sorted       ", np.sort(a))
print("the original is untouched:", a)
print("reversed     ", a[::-1])
print("is 5 present ", 5 in a)
print("where a > 3  ", np.where(a > 3)[0])
print("cumulative   ", a.cumsum())
print("unique       ", np.unique(a))
print("append 7     ", np.append(a, 7))
print("insert 0 at 3", np.insert(a, 3, 0))
print("delete idx 2 ", np.delete(a, 2))
print("and still    ", a)
array         [3 1 4 1 5 9 2 6]
length        8 or a.size 8
sum           31
smallest      1 at index 1
largest       9 at index 5
mean          3.875
sorted        [1 1 2 3 4 5 6 9]
the original is untouched: [3 1 4 1 5 9 2 6]
reversed      [6 2 9 5 1 4 1 3]
is 5 present  True
where a > 3   [2 4 5 7]
cumulative    [ 3  4  8  9 14 23 25 31]
unique        [1 2 3 4 5 6 9]
append 7      [3 1 4 1 5 9 2 6 7]
insert 0 at 3 [3 1 4 0 1 5 9 2 6]
delete idx 2  [3 1 1 5 9 2 6]
and still     [3 1 4 1 5 9 2 6]

Two things in that run are the marks.

np.sort(a) returns a sorted copy and leaves a alone. a.sort(), with no np. in front, sorts in place and returns None. Writing a = a.sort() therefore destroys the array, which is a classic mistake.

append, insert and delete all return a new array. They do not change the original, as the last line proves. A NumPy array has a fixed size: there is no such thing as growing one. If you need to add items repeatedly, build a Python list and convert it once at the end.

munotes.in30

Practical 3: Arrays, Basic Operations, Indexing and Slicing

Procedure

  1. Save the program as practical3a.py and import numpy as np.
  2. Create arrays with np.array, np.zeros, np.arange and np.linspace and print each.
  3. Index the first, the last and a middle item, using a negative index for the last.
  4. Take at least six slices, including a[::2] and a[::-1], and print each with a label.
  5. Slice a two dimensional array in both directions.
  6. Show that a list slice is a copy and an array slice is a view, by writing through both.
  7. Run the basic operations: sum, min, max, mean, sort, search, append, insert, delete.
  8. Record the dtype and itemsize your own machine printed.

Result

Arrays were created five ways, indexed from both ends, sliced eleven ways including reversal and step, and sliced in two dimensions. Writing through a list slice left the list unchanged; writing through an array slice changed the array, proving the slice is a view. Every basic operation ran and np.sort, append, insert and delete all left the original array untouched.

Where marks are lost

  • Not saying which kind of array you used. Name it: a NumPy array, or array.array.
  • Expecting a[1:4] to include item 4. The stop is always excluded.
  • a = a.sort(), which sets a to None because the in place sort returns nothing.
  • Expecting np.append to change the array. It returns a new one.
  • Trying to grow an array in a loop. Build a list, then convert once.
  • Not knowing a slice is a view. This is the question on this exercise.
  • Using m[1][2] and not knowing that m[1, 2] is the NumPy form.

For the journal

The aim in MU's words. A short table of the three things called an array in Python, with one line on each. The program, with every print labelled so the output can be read against the code. The full output. Then the list against array comparison, with both printed before and after writing through the slice, and one sentence: a list slice is a copy, a NumPy slice is a view onto the same memory. The conclusion: indexing a single item raises an error when out of range, while slicing clips silently, and a NumPy array has a fixed size so append, insert and delete all return new arrays.

Quick revision

  • Three arrays in Python: list (anything, flexible), array.array (one numeric type,

standard library, type code such as 'i'), numpy.ndarray (one type, n dimensions, the mathematics). MU's exercises mean NumPy.

  • Create: np.array, np.zeros, np.ones, np.full, np.arange, np.linspace,

np.eye.

  • arange takes a step and excludes the stop; linspace takes a count and

includes it.

  • Index from 0. a[-1] is the last. Out of range raises IndexError.
  • Two dimensions: m[row, col]. m[:, 1] is a whole column.
  • Slice start:stop:step, stop excluded. a[::-1] reverses, a[::2] steps.
  • An empty slice and a stop past the end are not errors.
  • A list slice is a copy. A NumPy slice is a VIEW. x.base is None tells them apart.
  • np.sort(a) returns a copy; a.sort() sorts in place and returns None.
  • np.append, np.insert, np.delete all return new arrays. Array size is fixed.
munotes.in31

Practical 3: Arrays, Basic Operations, Indexing and Slicing

Questions you should be able to answer

1. Name the three things in Python that get called an array, and what each holds. A list, which holds anything; an array.array, a compact array of one numeric type from the standard library; and a NumPy ndarray, one type with n dimensions and mathematics built in.

2. What does a[1:4] give? Items at indexes 1, 2 and 3. The stop is excluded.

3. How do you reverse an array in one expression? a[::-1].

4. a[5] on a three item array raises an error, but a[1:100] does not. Why? Indexing asks for one item that must exist. Slicing asks for whatever is in a range and clips to what there is, so that code can take a piece of something of unknown length.

5. What is the difference between slicing a list and slicing a NumPy array? A list slice is a new list, so writing to it leaves the original alone. A NumPy slice is a view onto the same memory, so writing to it changes the original.

6. How do you find out whether an array is a view or owns its data? x.base is the array it looks into, and it is None when the array owns its own memory.

7. What is wrong with a = a.sort()? a.sort() sorts in place and returns None, so a becomes None. Use b = np.sort(a) for a sorted copy, or a.sort() on its own line.

8. Why does np.append not grow the array? A NumPy array is one fixed block of memory. np.append makes a new, bigger array and copies into it, which is why appending in a loop is slow and building a list first is right.

9. What is the difference between arange(0, 1, 0.25) and linspace(0, 1, 5)? The first takes a step of 0.25 and stops before 1, giving four values. The second asks for five values evenly spaced and includes 1.

Contents This chapter on its own page

munotes.in32

Chapter Seven

Practical 3 continued: Mathematical Functions, Aliasing and Copying

Syllabus topic Module 1, practical 3(b), "Write a program to implement mathematical functions on arrays", and 3(c), "Write a program to perform array aliasing and copying"

Aim

To apply mathematical functions to arrays, and to show the difference between aliasing an array, taking a view of it, and copying it.

Part one: mathematical functions on arrays

The idea: one operation, the whole array

A Python list needs a loop to add 1 to every item. A NumPy array does not. An operation written once applies to every element, and this is called vectorising.

import numpy as np

py = [1, 2, 3, 4]
doubled_the_hard_way = [x * 2 for x in py]

a = np.array([1, 2, 3, 4])

print("a list needs a loop :", doubled_the_hard_way)
print("an array does not   :", a * 2)
print("and it works for all:", a + 10, a - 1, a ** 2, a / 2)
print("careful, a list repeats instead of multiplying:", py * 2)
a list needs a loop : [2, 4, 6, 8]
an array does not   : [2 4 6 8]
and it works for all: [11 12 13 14] [0 1 2 3] [ 1  4  9 16] [0.5 1.  1.5 2. ]
careful, a list repeats instead of multiplying: [1, 2, 3, 4, 1, 2, 3, 4]

The last line is the trap. [1, 2, 3, 4] * 2 on a list gives the list twice over, because multiplying a sequence by an integer repeats it. On an array it multiplies every element. Same symbol, two meanings, and it depends entirely on the type.

Arithmetic between two arrays

import numpy as np

a = np.array([10, 20, 30, 40])
b = np.array([1, 2, 3, 4])

print("a      ", a)
print("b      ", b)
print("a + b  ", a + b)
print("a - b  ", a - b)
print("a * b  ", a * b)
print("a / b  ", a / b)
print("a // b ", a // b)
print("a % b  ", a % b)
print("a ** 2 ", a ** 2)
print("a > 25 ", a > 25)
a       [10 20 30 40]
b       [1 2 3 4]
a + b   [11 22 33 44]
a - b   [ 9 18 27 36]
a * b   [ 10  40  90 160]
a / b   [10. 10. 10. 10.]
a // b  [10 10 10 10]
a % b   [0 0 0 0]
a ** 2  [ 100  400  900 1600]
a > 25  [False False  True  True]

Every one of those works element by element: the first with the first, the second with the second. Note that a * b is NOT matrix multiplication, it is element by element multiplication. Matrix multiplication is the @ operator or np.dot.

import numpy as np

m = np.array([[1, 2], [3, 4]])
n = np.array([[5, 6], [7, 8]])

print("element by element, m * n")
print(m * n)
print("matrix product, m @ n")
print(m @ n)
print("the top left of the product is 1*5 + 2*7 =", 1 * 5 + 2 * 7)
munotes.in33

Practical 3 continued: Mathematical Functions, Aliasing and Copying

element by element, m * n
[[ 5 12]
 [21 32]]
matrix product, m @ n
[[19 22]
 [43 50]]
the top left of the product is 1*5 + 2*7 = 19

The universal functions

A universal function, or ufunc, is a mathematical function that NumPy applies to every element. There is one for everything the mathematics paper needs.

import numpy as np

a = np.array([1.0, 4.0, 9.0, 16.0])
angles = np.array([0.0, np.pi / 6, np.pi / 4, np.pi / 2])

print("sqrt     ", np.sqrt(a))
print("square   ", np.square(a))
print("exp      ", np.round(np.exp(np.array([0.0, 1.0, 2.0])), 4))
print("log      ", np.round(np.log(a), 4))
print("log10    ", np.log10(np.array([1.0, 10.0, 100.0, 1000.0])))
print("sin      ", np.round(np.sin(angles), 4))
print("cos      ", np.round(np.cos(angles), 4))
print("degrees  ", np.degrees(angles))
print("abs      ", np.abs(np.array([-3, 4, -5])))
print("floor    ", np.floor(np.array([1.2, 1.8, -1.2])))
print("ceil     ", np.ceil(np.array([1.2, 1.8, -1.2])))
print("round    ", np.round(np.array([1.24, 1.25, 1.26]), 1))
print("power    ", np.power(np.array([2, 3, 4]), 3))
sqrt      [1. 2. 3. 4.]
square    [  1.  16.  81. 256.]
exp       [1.     2.7183 7.3891]
log       [0.     1.3863 2.1972 2.7726]
log10     [0. 1. 2. 3.]
sin       [0.     0.5    0.7071 1.    ]
cos       [1.     0.866  0.7071 0.    ]
degrees   [ 0. 30. 45. 90.]
abs       [3 4 5]
floor     [ 1.  1. -2.]
ceil      [ 2.  2. -1.]
round     [1.2 1.2 1.3]
power     [ 8 27 64]

np.round(x, 4) is used above so the output fits the page and does not depend on how many digits a version chooses to print. In your own journal print them unrounded once, so you have seen the full values.

Look at the round row: 1.24 gave 1.2, 1.26 gave 1.3, and 1.25 gave 1.2, not 1.3. That is not a bug. NumPy, like Python's own round, rounds a value sitting exactly halfway to the nearest even last digit, so 1.25 goes down to 1.2 and 2.25 would go up to 2.2 as well. It is called round half to even, and it exists because always rounding halves upward makes a long column of figures drift high. Expect it, and do not treat it as an error in your program.

sin 30 degrees is 0.5, and the second value in the sin row is exactly that, which is the check that the angles were given in radians. NumPy's trigonometric functions take radians, never degrees. np.radians() converts if your input is in degrees.

Aggregate functions, and the axis

An aggregate reduces many numbers to one. On a two dimensional array you can reduce down the columns, along the rows, or over everything.

import numpy as np

marks = np.array([[78, 65, 80],
                  [55, 72, 60],
                  [90, 88, 95],
                  [40, 51, 45]])

print(marks)
print("total of everything      ", marks.sum())
print("mean of everything       ", round(marks.mean(), 4))
print("axis=0, down the columns ", marks.sum(axis=0))
print("axis=1, along the rows   ", marks.sum(axis=1))
print("each student's mean      ", marks.mean(axis=1))
print("each subject's highest   ", marks.max(axis=0))
print("each subject's lowest    ", marks.min(axis=0))
print("standard deviation       ", round(marks.std(), 4))
print("row 2 is the best student, index", marks.sum(axis=1).argmax())
munotes.in34

Practical 3 continued: Mathematical Functions, Aliasing and Copying

[[78 65 80]
 [55 72 60]
 [90 88 95]
 [40 51 45]]
total of everything       819
mean of everything        68.25
axis=0, down the columns  [263 276 280]
axis=1, along the rows    [223 187 273 136]
each student's mean       [74.33333333 62.33333333 91.         45.33333333]
each subject's highest    [90 88 95]
each subject's lowest     [40 51 45]
standard deviation        17.5979
row 2 is the best student, index 2

axis=0 goes down, axis=1 goes across. That is the fact to memorise, and here is the way to keep it straight: the axis you name is the one that disappears. marks has shape (4, 3); marks.sum(axis=0) removes the 4 and leaves 3 numbers, one per column.

Check the first column by hand: 78 + 55 + 90 + 40 = 263, and that is the first number in the axis=0 row.

Broadcasting

Two arrays of different shapes can still be combined, if one of them can be stretched to fit. That stretching is called broadcasting, and it is what makes a * 2 work: the 2 is broadcast to every element.

The rule, in one line: compare the shapes from the right, and a dimension of 1, or a missing dimension, is stretched.

import numpy as np

marks = np.array([[78, 65, 80],
                  [55, 72, 60],
                  [90, 88, 95]])
weights = np.array([2, 1, 1])

print("marks shape  ", marks.shape)
print("weights shape", weights.shape)
print("weighted marks, each column scaled")
print(marks * weights)
print("add 5 to everybody")
print(marks + 5)
print("subtract each subject's mean from its column")
print(np.round(marks - marks.mean(axis=0), 4))
marks shape   (3, 3)
weights shape (3,)
weighted marks, each column scaled
[[156  65  80]
 [110  72  60]
 [180  88  95]]
add 5 to everybody
[[ 83  70  85]
 [ 60  77  65]
 [ 95  93 100]]
subtract each subject's mean from its column
[[  3.6667 -10.       1.6667]
 [-19.3333  -3.     -18.3333]
 [ 15.6667  13.      16.6667]]

And when the shapes cannot be made to fit, NumPy says so rather than guessing:

import numpy as np

a = np.array([[1, 2, 3], [4, 5, 6]])
b = np.array([1, 2])
print(a + b)
ValueError: operands could not be broadcast together with shapes (2,3) (2,)

Read that message: it names both shapes. (2, 3) and (2,) compared from the right gives 3 against 2, and neither is 1, so there is nothing to stretch. b of shape (3,) would have worked.

munotes.in35

Practical 3 continued: Mathematical Functions, Aliasing and Copying

Part two: aliasing, views and copying

This is MU's third bullet and it is the part of this practical that gets asked about. There are three cases, not two, and they have to be told apart by writing through each one and looking at the original.

What it isMade byOwn memoryWriting through it changes the original
Aliasa second name for the same arrayb = ano, there is one arrayyes, always
Viewa window onto the same memorya[1:4], a.reshape(), a.Tnoyes
Copya new array with the same valuesa.copy(), np.array(a)yesno

An alias is a second name

import numpy as np

a = np.array([1, 2, 3, 4, 5])
b = a

print("a is", a)
print("b is", b)
print("are they the same object?", b is a)
print("same id?", id(a) == id(b))

b[0] = 999
print("after b[0] = 999")
print("a is", a)
print("b is", b)
a is [1 2 3 4 5]
b is [1 2 3 4 5]
are they the same object? True
same id? True
after b[0] = 999
a is [999   2   3   4   5]
b is [999   2   3   4   5]

There is one array in that program and two names for it. b = a copies nothing at all; it writes the same reference into a second name. So of course changing one changes the other, and b is a proves they are the same object.

This is not special to NumPy. A Python list behaves identically, and it is the source of one of the nastiest bugs a beginner writes:

first = [1, 2, 3]
second = first
second.append(4)
print("first ", first)
print("second", second)
print("the same object?", second is first)
first  [1, 2, 3, 4]
second [1, 2, 3, 4]
the same object? True

A view is a window

import numpy as np

a = np.array([1, 2, 3, 4, 5, 6])
window = a[1:4]
reshaped = a.reshape(2, 3)

print("a       ", a)
print("window  ", window, " is it a view?", window.base is not None)
print("reshaped\n", reshaped, " is it a view?", reshaped.base is not None)

window[0] = 999
print("after window[0] = 999, a is", a)

reshaped[1, 2] = 777
print("after reshaped[1, 2] = 777, a is", a)
print("are they the same object? ", window is a, reshaped is a)
a        [1 2 3 4 5 6]
window   [2 3 4]  is it a view? True
reshaped
 [[1 2 3]
 [4 5 6]]  is it a view? True
after window[0] = 999, a is [  1 999   3   4   5   6]
after reshaped[1, 2] = 777, a is [  1 999   3   4   5 777]
are they the same object?  False False
munotes.in36

Practical 3 continued: Mathematical Functions, Aliasing and Copying

Note the last line: a view is not the same object as the array, so is says False, and yet writing through it still changes the array. That is exactly why is is not the test here. base is.

A copy is a new array

import numpy as np

a = np.array([1, 2, 3, 4, 5])

shallow_name = a
sliced_view = a[:]
real_copy = a.copy()
also_a_copy = np.array(a)

for label, other in [("b = a          ", shallow_name),
                     ("a[:]           ", sliced_view),
                     ("a.copy()       ", real_copy),
                     ("np.array(a)    ", also_a_copy)]:
    print(f"{label} same object {other is a!s:>5}   owns its data {other.base is None!s:>5}")

real_copy[0] = 999
also_a_copy[1] = 888
print("after writing 999 and 888 into the two copies, a is", a)

sliced_view[2] = 777
print("after writing 777 into a[:], a is         ", a)
b = a           same object  True   owns its data  True
a[:]            same object False   owns its data False
a.copy()        same object False   owns its data  True
np.array(a)     same object False   owns its data  True
after writing 999 and 888 into the two copies, a is [1 2 3 4 5]
after writing 777 into a[:], a is          [  1   2 777   4   5]

There is the whole exercise in one output. a.copy() and np.array(a) own their data and leave a alone. a[:] does not, and this is the line that catches people. On a Python list a[:] is the classic way to make a copy. On a NumPy array it is a view. A student who carries the list habit across corrupts their data silently.

Deep copy, and when you actually need it

a.copy() is enough for any array of numbers. It is not enough for an array or a list holding other containers, because copying the outer one still leaves the inner ones shared.

from copy import deepcopy

original = [[1, 2], [3, 4]]
shallow = list(original)
deep = deepcopy(original)

shallow[0][0] = 999
print("after writing through the shallow copy:", original)

deep[1][1] = 888
print("after writing through the deep copy   :", original)
print("the deep one changed only itself      :", deep)
after writing through the shallow copy: [[999, 2], [3, 4]]
after writing through the deep copy   : [[999, 2], [3, 4]]
the deep one changed only itself      : [[1, 2], [3, 888]]

list(original) made a new outer list, but its two slots still refer to the same two inner lists, so writing into shallow[0][0] reached the original. deepcopy copied all the way down. For NumPy numeric arrays this never arises, because the numbers are stored in the array itself rather than referred to.

The one line summary to say at the table

import numpy as np

a = np.arange(6)

for label, other in [("alias  b = a   ", a),
                     ("view   a[1:4]  ", a[1:4]),
                     ("copy   a.copy()", a.copy())]:
    print(f"{label}  is a: {other is a!s:>5}"
          f"   shares memory: {np.shares_memory(other, a)!s:>5}")
munotes.in37

Practical 3 continued: Mathematical Functions, Aliasing and Copying

alias  b = a     is a:  True   shares memory:  True
view   a[1:4]    is a: False   shares memory:  True
copy   a.copy()  is a: False   shares memory: False

is a true means an alias. False but sharing memory means a view, and false and sharing nothing means a copy. Two tests, three answers.

One caution about base, because [Practical 4: NumPy Slicing, Basic and Advanced Indexing] meets it head on. base says whether an array owns its own buffer, which for a plain slice or a plain copy is the same question as "is this a view of that array". For a subscript that mixes a slice with a list of positions it is not the same question, and base is then not None even though nothing is shared. Where the answer matters, ask np.shares_memory(x, a).

Procedure

  1. Save as practical3b.py. Apply +, -, *, /, //, % and ** to an array and

to two arrays, and print each result.

  1. Show that list 2 repeats and array 2 multiplies.
  2. Apply sqrt, exp, log, sin, cos, abs, floor, ceil and power, and check

that sin of pi over 6 is 0.5.

  1. Build a two dimensional array of marks. Take sum, mean, max and min over

everything, then with axis=0 and axis=1. Check one column total by hand.

  1. Show broadcasting with a row of weights, and show the shape error when it cannot work.
  2. Save as practical3c.py. Make an alias, a view and a copy of one array. For each, print

whether it is the same object and whether it shares memory with the original, then write through it and print the original.

  1. Show that a[:] copies a list but views an array.

Result

Every operator and nine mathematical functions were applied element by element; sin of pi over 6 came out 0.5, confirming radians. Column totals matched the hand check 78 + 55 + 90 + 40 = 263. Broadcasting worked for shapes (3, 3) and (3,) and raised ValueError for (2, 3) and (2,). Writing through the alias and through the view both changed the original; writing through a.copy() and np.array(a) did not; and a[:] behaved as a view on the array and as a copy on the list.

Where marks are lost

  • Using a loop to apply a function to an array. The whole point is that you do not.
  • Calling a * b matrix multiplication. It is element by element. @ is the product.
  • Giving degrees to np.sin. It takes radians.
  • Getting axis backwards. The axis you name is the one that disappears.
  • Calling a view a copy. a[1:4] is a view, and writing to it changes a.
  • Using a[:] to copy an array out of list habit. On an array it is a view.
  • Testing with is to tell a view from a copy. A view is a different object; use
munotes.in38

Practical 3 continued: Mathematical Functions, Aliasing and Copying

np.shares_memory.

  • Showing only the code. Aliasing is only proved by printing the original after the

write.

For the journal

Two entries under practical 3. For the mathematical functions: the aim, a table of the operators with one example each, the ufunc program and its output, and the axis table of marks with one column total checked by hand. One sentence on broadcasting: shapes are compared from the right and a dimension of 1 is stretched.

For aliasing and copying: the aim, the three row table of alias, view and copy, then the program that writes through each one, and the original printed after each write. Record np.shares_memory for each of the three, because that is the line that proves the table. The conclusion in one sentence: b = a makes a second name, a slice makes a window, and only copy() makes a new array, so the first two change the original and the third does not.

Quick revision

  • An operation on an array applies to every element. No loop.
  • list 2 repeats the list; array 2 multiplies each element.
  • Between two arrays every operator is element by element. a * b is NOT the matrix

product; a @ b is.

  • ufuncs: sqrt, square, exp, log, log10, sin, cos, abs, floor, ceil,

round, power. Trigonometry takes radians; np.radians() converts.

  • Aggregates: sum, mean, min, max, std, argmax.

axis=0 goes down the columns and axis=1 along the rows, and the axis named is the one that disappears.

  • Broadcasting: compare shapes from the right; a 1 or a missing dimension stretches.

Otherwise ValueError, and the message names both shapes.

  • Alias b = a: one array, two names. b is a is True.
  • View a[1:4], a.reshape(), a.T: different object, same memory.
  • Copy a.copy() or np.array(a): new memory.
  • The test is np.shares_memory(x, a). x.base is a quick indicator only: on a mixed

subscript such as m[1:, [0, 2]] the base is a temporary, so it is not None while nothing is shared.

  • Writing through an alias or a view changes the original. Through a copy it does not.
  • a[:] copies a list and views an array.
  • deepcopy is only needed for containers holding containers.
munotes.in39

Practical 3 continued: Mathematical Functions, Aliasing and Copying

Questions you should be able to answer

1. Why does an array not need a loop to add 1 to every element? Because arithmetic on a NumPy array is applied element by element by the library itself, in compiled code. That is called vectorising.

2. What does [1, 2] 3 give, and what does np.array([1, 2]) 3 give? [1, 2, 1, 2, 1, 2], because multiplying a list repeats it, and [3 6], because multiplying an array scales every element.

3. What is the difference between a b and a @ b? a b multiplies element by element. a @ b is the matrix product.

4. marks has shape (4, 3). What does marks.sum(axis=0) give? An array of shape (3,): one total per column, that is, per subject. The axis named is the one that disappears.

5. State the broadcasting rule. Compare the shapes from the right. A dimension that is 1, or absent, is stretched to match. If two dimensions differ and neither is 1, it is an error.

6. What are the three ways a second name can relate to an array? An alias, which is the same object; a view, a different object sharing the same memory; and a copy, a different object with its own memory.

7. How do you tell a view from a copy at the keyboard? np.shares_memory(x, a), which is True exactly when writing to one is seen in the other. x.base is a quick indicator of whether x owns its buffer, but on a mixed subscript the base is a temporary array, so it can be not None while nothing at all is shared.

8. Why is b = a never a copy? Because assignment binds a name to the object that is already there. Nothing is duplicated, so both names reach one array.

9. A student writes backup = a[:] and then changes backup. What happens? The array a changes too, because on a NumPy array a[:] is a view. On a Python list the same line would have made a real copy, which is why the mistake is so easy.

10. When do you need deepcopy rather than copy? Only when the container holds other containers. copy duplicates the outer one and leaves the inner ones shared. An array of numbers never needs it.

Contents This chapter on its own page

munotes.in40

Chapter Eight

Practical 4: NumPy Slicing, Basic and Advanced Indexing

Syllabus topic Module 1, practical 4(a) and 4(b), "Write a program to perform slicing, basic and advanced indexing on NumPy arrays"

Aim

To perform slicing, basic indexing and advanced indexing on NumPy arrays, and to show what each one returns.

The two kinds of indexing, and the real difference

NumPy divides every way of picking elements out of an array into two families, and they are its own published names for them.

Written withReturns
Basic indexingintegers, slices start:stop:step, ..., np.newaxisa view onto the same memory
Advanced indexingan array or list of integers, or a boolean arraya copy

That is the whole distinction, and it decides whether writing to the result changes the original array. Everything else in this chapter follows from it.

import numpy as np

a = np.arange(10) * 10
basic = a[2:6]
advanced = a[[2, 3, 4, 5]]

print("a       ", a)
print("basic   ", basic, "  shares memory with a?", np.shares_memory(basic, a))
print("advanced", advanced, "  shares memory with a?", np.shares_memory(advanced, a))

basic[0] = 999
advanced[1] = 888

print("after writing 999 into the basic result and 888 into the advanced one")
print("a       ", a)
a        [ 0 10 20 30 40 50 60 70 80 90]
basic    [20 30 40 50]   shares memory with a? True
advanced [20 30 40 50]   shares memory with a? False
after writing 999 into the basic result and 888 into the advanced one
a        [  0  10 999  30  40  50  60  70  80  90]

The two results hold the same four numbers and behave completely differently. Writing 999 through the basic slice changed a. Writing 888 through the advanced result did not, because that result is a copy.

Part one: slicing, which is basic indexing

Slicing was introduced in [Practical 3: Arrays, Basic Operations, Indexing and Slicing]. Here it is on more than one dimension, which is what MU's row is about.

import numpy as np

m = np.arange(1, 25).reshape(4, 6)
print(m)
[[ 1  2  3  4  5  6]
 [ 7  8  9 10 11 12]
 [13 14 15 16 17 18]
 [19 20 21 22 23 24]]

Now every kind of slice on it, each labelled:

import numpy as np

m = np.arange(1, 25).reshape(4, 6)

print("m[1]        one whole row   ", m[1])
print("m[:, 2]     one whole column", m[:, 2])
print("m[1, 3]     one element     ", m[1, 3])
print("m[0:2, 0:3] a block")
print(m[0:2, 0:3])
print("m[::2, ::3] every 2nd row, every 3rd column")
print(m[::2, ::3])
print("m[:, -1]    the last column ", m[:, -1])
print("m[::-1, :]  rows reversed")
print(m[::-1, :])
print("m[1:3, 2:5] rows 1 to 2, columns 2 to 4")
print(m[1:3, 2:5])
m[1]        one whole row    [ 7  8  9 10 11 12]
m[:, 2]     one whole column [ 3  9 15 21]
m[1, 3]     one element      10
m[0:2, 0:3] a block
[[1 2 3]
 [7 8 9]]
m[::2, ::3] every 2nd row, every 3rd column
[[ 1  4]
 [13 16]]
m[:, -1]    the last column  [ 6 12 18 24]
m[::-1, :]  rows reversed
[[19 20 21 22 23 24]
 [13 14 15 16 17 18]
 [ 7  8  9 10 11 12]
 [ 1  2  3  4  5  6]]
m[1:3, 2:5] rows 1 to 2, columns 2 to 4
[[ 9 10 11]
 [15 16 17]]
munotes.in41

Practical 4: NumPy Slicing, Basic and Advanced Indexing

The shape a slice returns, which is the part students get wrong

import numpy as np

m = np.arange(1, 25).reshape(4, 6)

print("m         shape", m.shape, "ndim", m.ndim)
print("m[1]      shape", m[1].shape, "ndim", m[1].ndim, "  one index removes a dimension")
print("m[1:2]    shape", m[1:2].shape, "ndim", m[1:2].ndim, "  a slice KEEPS it")
print("m[1, 3]   shape", np.shape(m[1, 3]), "a single element, no dimensions at all")
print("m[:, 2]   shape", m[:, 2].shape, "  a column comes back as a flat row")
print("m[:, 2:3] shape", m[:, 2:3].shape, "  as a column, because it is a slice")
m         shape (4, 6) ndim 2
m[1]      shape (6,) ndim 1   one index removes a dimension
m[1:2]    shape (1, 6) ndim 2   a slice KEEPS it
m[1, 3]   shape () a single element, no dimensions at all
m[:, 2]   shape (4,)   a column comes back as a flat row
m[:, 2:3] shape (4, 1)   as a column, because it is a slice

An integer index removes that dimension. A slice keeps it. m[1] is a one dimensional array of six numbers; m[1:2] is a two dimensional array with one row. They print differently and they behave differently in arithmetic, and knowing which you have is what prevents a broadcasting error later.

Note also that m[:, 2] gives a flat array, not a column. If you need it shaped as a column, slice instead of indexing: m[:, 2:3].

The ellipsis

... stands for "as many full slices as are needed here". On a three dimensional array it saves writing the colons out:

import numpy as np

cube = np.arange(24).reshape(2, 3, 4)

print("cube.shape", cube.shape)
print("cube[0, :, :] and cube[0, ...] are the same:")
print(cube[0, ...])
print("the last element along the last axis, cube[..., -1]")
print(cube[..., -1])
print("same thing written out, cube[:, :, -1]")
print(cube[:, :, -1])
print("equal?", np.array_equal(cube[..., -1], cube[:, :, -1]))
cube.shape (2, 3, 4)
cube[0, :, :] and cube[0, ...] are the same:
[[ 0  1  2  3]
 [ 4  5  6  7]
 [ 8  9 10 11]]
the last element along the last axis, cube[..., -1]
[[ 3  7 11]
 [15 19 23]]
same thing written out, cube[:, :, -1]
[[ 3  7 11]
 [15 19 23]]
equal? True

Part two: advanced indexing

With an array of integers

Give NumPy a list of positions and it gives back those elements, in the order you asked for, including repeats.

munotes.in42

Practical 4: NumPy Slicing, Basic and Advanced Indexing

import numpy as np

a = np.array([10, 20, 30, 40, 50, 60, 70])

print("a                 ", a)
print("a[[0, 2, 4]]      ", a[[0, 2, 4]])
print("a[[4, 2, 0]]      ", a[[4, 2, 0]], "the order you asked for")
print("a[[1, 1, 1]]      ", a[[1, 1, 1]], "repeats are allowed")
print("a[[-1, -2]]       ", a[[-1, -2]], "negatives work too")
print("shares with a?    ", np.shares_memory(a[[0, 2, 4]], a))
a                  [10 20 30 40 50 60 70]
a[[0, 2, 4]]       [10 30 50]
a[[4, 2, 0]]       [50 30 10] the order you asked for
a[[1, 1, 1]]       [20 20 20] repeats are allowed
a[[-1, -2]]        [70 60] negatives work too
shares with a?     False

Three things a slice cannot do are in that output. Out of order, repeated, and from any list of positions. That is what advanced indexing buys, and the price is the copy.

On two dimensions, two index arrays are paired element by element:

import numpy as np

m = np.arange(1, 13).reshape(3, 4)
print(m)

rows = [0, 1, 2]
cols = [3, 2, 0]
print("m[rows, cols] pairs them up:", m[rows, cols])
print("which is m[0,3], m[1,2], m[2,0] =", m[0, 3], m[1, 2], m[2, 0])

print("whole rows, in a chosen order")
print(m[[2, 0]])
[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]
m[rows, cols] pairs them up: [4 7 9]
which is m[0,3], m[1,2], m[2,0] = 4 7 9
whole rows, in a chosen order
[[ 9 10 11 12]
 [ 1  2  3  4]]

Two index arrays are zipped, not crossed. m[[0, 1, 2], [3, 2, 0]] gives three elements, not a three by three block. Expecting a block is the commonest wrong answer here. To get the block, index the rows and then the columns with a slice between:

import numpy as np

m = np.arange(1, 13).reshape(3, 4)

print("zipped, three elements :", m[[0, 2], [1, 3]])
print("crossed, a 2 by 2 block:")
print(m[np.ix_([0, 2], [1, 3])])
print("or with a reshaped row index:")
print(m[[[0], [2]], [1, 3]])
zipped, three elements : [ 2 12]
crossed, a 2 by 2 block:
[[ 2  4]
 [10 12]]
or with a reshaped row index:
[[ 2  4]
 [10 12]]

With a boolean mask

A boolean array of the same shape picks out the elements where it is True. This is the form that is actually used in real work, because the mask is usually a comparison.

import numpy as np

a = np.array([15, 42, 8, 23, 4, 16, 50])

mask = a > 20
print("a          ", a)
print("a > 20     ", mask, "a boolean array of the same length")
print("a[a > 20]  ", a[mask])
print("how many   ", mask.sum(), "because True counts as 1")
print("any? all?  ", mask.any(), mask.all())
print("even ones  ", a[a % 2 == 0])
print("between    ", a[(a > 10) & (a < 45)])
print("either or  ", a[(a < 10) | (a > 45)])
print("not        ", a[~mask])
print("shares?    ", np.shares_memory(a[mask], a))
munotes.in43

Practical 4: NumPy Slicing, Basic and Advanced Indexing

a           [15 42  8 23  4 16 50]
a > 20      [False  True False  True False False  True] a boolean array of the same length
a[a > 20]   [42 23 50]
how many    3 because True counts as 1
any? all?   True False
even ones   [42  8  4 16 50]
between     [15 42 23 16]
either or   [ 8  4 50]
not         [15  8  4 16]
shares?     False

Use &, | and ~, not and, or and not. The words work on one True or False at a time and cannot handle a whole array; the symbols work element by element. And the brackets around each comparison are not optional, because & binds more tightly than >.

import numpy as np

a = np.array([15, 42, 8])
print(a[a > 10 and a < 45])
ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all()

Read that message. It is one of NumPy's most recognisable errors and it always means the same thing: a word operator was used where a symbol was needed.

A mask can also be assigned through, and that is how a whole class of values is changed at once:

import numpy as np

marks = np.array([78, 32, 65, 18, 90, 39])
print("before      ", marks)

marks[marks < 40] = 0
print("fails zeroed", marks)

grades = np.where(marks >= 75, "distinction", np.where(marks >= 40, "pass", "fail"))
print("grades      ", grades)
before       [78 32 65 18 90 39]
fails zeroed [78  0 65  0 90  0]
grades       ['distinction' 'fail' 'pass' 'fail' 'distinction' 'fail']

np.where(condition, a, b) picks from a where the condition is True and from b where it is False, element by element. Nested, as above, it is a whole grade ladder in one line.

Mixing the two

When a slice and an index array appear in the same subscript, the result is a copy, because advanced indexing is present.

import numpy as np

m = np.arange(1, 13).reshape(3, 4)

mixed = m[1:, [0, 2]]
print("m[1:, [0, 2]]")
print(mixed)
print("shares memory with m?", np.shares_memory(mixed, m))

mixed[0, 0] = 999
print("after writing 999 into it, m is unchanged:")
print(m)
m[1:, [0, 2]]
[[ 5  7]
 [ 9 11]]
shares memory with m? False
after writing 999 into it, m is unchanged:
[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]

The test that lies, and the one that does not

Earlier chapters used x.base is None to ask whether an array owns its memory, and for a plain slice or a plain copy it answers correctly. On a mixed subscript it does not answer the question you think you asked, and it is worth seeing once:

munotes.in44

Practical 4: NumPy Slicing, Basic and Advanced Indexing

import numpy as np

m = np.arange(1, 13).reshape(3, 4)
mixed = m[1:, [0, 2]]

print("mixed.base is None?          ", mixed.base is None)
print("but np.shares_memory(mixed, m)", np.shares_memory(mixed, m))
print("and mixed.base is actually:")
print(mixed.base)
mixed.base is None?           False
but np.shares_memory(mixed, m) False
and mixed.base is actually:
[[ 5  9]
 [ 7 11]]

mixed.base is not m. It is a temporary array NumPy built while working the subscript out, and mixed is a view onto that, which is itself a copy. So base is not None is true and yet nothing is shared with m.

np.shares_memory(x, y) is the test that answers the real question, which is whether writing into one will be seen in the other. Use base as a quick look at whether an array owns its buffer, and np.shares_memory whenever the answer matters.

Summary of every form

WrittenFamilyReturns
a[3]basicone element
a[1:5]basica view
a[::2]basica view
m[1, 2]basicone element
m[:, 1]basica view, flat
m[..., -1]basica view
a[[1, 3, 5]]advanceda copy
a[[3, 3, 3]]advanceda copy, with repeats
m[rows, cols]advanceda copy, the two zipped
m[np.ix_(rows, cols)]advanceda copy, the two crossed
a[a > 20]advanced, booleana copy
m[1:, [0, 2]]mixeda copy

Procedure

  1. Save as practical4a.py. Build a 4 by 6 array with np.arange(1, 25).reshape(4, 6).
  2. Take one row, one column, one element and two blocks by slicing, printing each labelled.
  3. Print the shape of m[1] and of m[1:2] and say why they differ.
  4. Use ... on a three dimensional array and prove it equals the written out form.
  5. Index with a list of positions, out of order and with a repeat.
  6. Index a two dimensional array with two lists and show they are zipped, then get the

crossed block with np.ix_.

  1. Build a boolean mask from a comparison, combine masks with &, | and ~, and assign

through a mask.

  1. For each result print np.shares_memory(result, m), and write through it to show what

changes.

Result

Slicing and integer indexing returned views: writing 999 through a[2:6] changed the array. A list of positions, a boolean mask and a mixed subscript all returned copies: writing through them left the array unchanged. m[1] had shape (6,) and m[1:2] had shape (1, 6), confirming that an integer removes a dimension and a slice keeps it. m[[0, 1, 2], [3, 2, 0]] gave three elements, the pairs (0,3), (1,2) and (2,0), and np.ix_ gave the crossed block instead.

munotes.in45

Practical 4: NumPy Slicing, Basic and Advanced Indexing

Where marks are lost

  • Not knowing which family is which. Slices are basic and give views; index arrays and

masks are advanced and give copies. That is the exercise.

  • Expecting m[[0, 1], [2, 3]] to be a block. It is two elements, zipped.
  • Using and, or and not on arrays. It raises the ambiguous truth value error.
  • Forgetting the brackets in (a > 10) & (a < 45).
  • Expecting m[:, 2] to be a column. It comes back flat; m[:, 2:3] is the column.
  • Showing the code and not the memory check. The view against copy claim has to be

demonstrated, by writing through the result and printing the original.

  • Trusting base on a mixed subscript. Use np.shares_memory.

For the journal

The aim in MU's words. The two row table of basic against advanced with what each returns. The 4 by 6 array printed once, then every slice labelled with its output. The shape comparison of m[1] and m[1:2] with one sentence on why. Then advanced indexing: the list of positions out of order and repeated, the zipped pair, and the boolean mask with &, | and ~. For each one, the np.shares_memory line and the write through test. The conclusion in one sentence: basic indexing gives a view onto the same memory, advanced indexing gives a copy, so writing through the first changes the original and writing through the second does not.

Quick revision

  • Basic indexing: integers, slices, ..., np.newaxis. Returns a view.
  • Advanced indexing: a list or array of integers, or a boolean array. Returns a

copy.

  • np.shares_memory(x, a) is the honest test of whether writing to x changes a.

x.base is None is a quick look at whether x owns its buffer, and on a mixed subscript its base is a temporary, so it is not None even though nothing is shared.

  • An integer index removes a dimension; a slice keeps it. m[1] is (6,), m[1:2] is

(1, 6).

  • m[:, 2] is flat. m[:, 2:3] is a column.
  • ... fills in as many : as are needed.
  • An index array may be out of order, may repeat, and may use negatives.
  • Two index arrays are zipped, not crossed. np.ix_(rows, cols) crosses them.
  • A boolean mask must be the same shape. Combine with &, |, ~, each comparison in

brackets. and, or, not raise the ambiguous truth value error.

  • mask.sum() counts the Trues. a[mask] = 0 assigns through a mask.
  • np.where(cond, a, b) chooses element by element and nests.
  • A subscript mixing a slice and an index array gives a copy.
munotes.in46

Practical 4: NumPy Slicing, Basic and Advanced Indexing

Questions you should be able to answer

1. What is the difference between basic and advanced indexing? Basic indexing uses integers and slices and returns a view onto the same memory. Advanced indexing uses an integer array or a boolean array and returns a copy.

2. Why does that difference matter? Because writing to a view changes the original array and writing to a copy does not.

2a. Which test tells you which you have? np.shares_memory(result, original). base is not enough: for m[1:, [0, 2]] the base is a temporary array NumPy built, so it is not None even though the result shares nothing with m.

3. a is [10 20 30 40 50]. What does a[[3, 1, 1]] give? [40 20 20]. An index array gives the elements in the order asked for and may repeat them.

4. m is 3 by 4. What does m[[0, 1], [2, 3]] give? Two elements, m[0, 2] and m[1, 3]. The two lists are paired, not crossed.

5. How do you get the crossed 2 by 2 block instead? m[np.ix_([0, 1], [2, 3])].

6. Why can you not write a[a > 10 and a < 45]? and needs a single True or False and an array of several is ambiguous, so it raises ValueError. Use a[(a > 10) & (a < 45)].

7. What shape does m[1] have, and what shape does m[1:2] have, for a 4 by 6 array? (6,) and (1, 6). The integer removes the first dimension; the slice keeps it with a length of 1.

8. How do you set every mark below 40 to zero? marks[marks < 40] = 0. A mask can be assigned through.

9. How many elements are True in a mask, in one expression? mask.sum(), because True counts as 1.

10. Is m[1:, [0, 2]] a view or a copy, and why? A copy. The subscript contains an index array, so advanced indexing applies to the whole thing.

Contents This chapter on its own page

munotes.in47

Chapter Nine

Practical 4 continued: the Dimensions and Attributes of an Array

Syllabus topic Module 1, practical 4(c), "Write a program to analyze dimensions and attributes of arrays"

Aim

To analyse the dimensions and the attributes of arrays of one, two and three dimensions.

The vocabulary, in one picture

Three words get muddled, and untangling them is most of this exercise.

WordWhat it isExample
dimensionone direction the array extends in. Also called an axisa table has 2
shapehow long the array is in each direction, as a tuple(4, 6)
sizehow many elements there are in total24

A one dimensional array is a row of numbers. A two dimensional array is a table of rows and columns. A three dimensional array is a stack of tables. The number of dimensions is ndim, and it is exactly the length of shape.

import numpy as np

one = np.array([1, 2, 3, 4])
two = np.array([[1, 2, 3], [4, 5, 6]])
three = np.arange(24).reshape(2, 3, 4)
scalar = np.array(7)

for name, a in [("scalar", scalar), ("one", one), ("two", two), ("three", three)]:
    print(f"{name:>7}  ndim {a.ndim}   shape {str(a.shape):<12} size {a.size}"
          f"   len(shape) {len(a.shape)}")
 scalar  ndim 0   shape ()           size 1   len(shape) 0
    one  ndim 1   shape (4,)         size 4   len(shape) 1
    two  ndim 2   shape (2, 3)       size 6   len(shape) 2
  three  ndim 3   shape (2, 3, 4)    size 24   len(shape) 3

Read the last two columns together: ndim is always the length of shape. And a zero dimensional array is a legal thing: it holds one number and its shape is the empty tuple.

The size is the shape multiplied out. For the three dimensional array above, 2 × 3 × 4 = 24, which is what size printed.

Every attribute MU could mean

import numpy as np

m = np.arange(1, 13).reshape(3, 4)

print(m)
print("m.ndim     ", m.ndim, "    how many dimensions")
print("m.shape    ", m.shape, " the length along each one")
print("m.size     ", m.size, "   how many elements in total")
print("m.dtype    ", m.dtype, " the type of one element")
print("m.itemsize ", m.itemsize, "    bytes in one element")
print("m.nbytes   ", m.nbytes, "   bytes in the whole array")
print("m.T        ", "the transpose, rows and columns swapped")
print(m.T)
print("m.flags['C_CONTIGUOUS']", m.flags["C_CONTIGUOUS"])
print("check: size times itemsize =", m.size, "*", m.itemsize, "=", m.size * m.itemsize,
      "and nbytes is", m.nbytes)
[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]
m.ndim      2     how many dimensions
m.shape     (3, 4)  the length along each one
m.size      12    how many elements in total
m.dtype     int64  the type of one element
m.itemsize  8     bytes in one element
m.nbytes    96    bytes in the whole array
m.T         the transpose, rows and columns swapped
[[ 1  5  9]
 [ 2  6 10]
 [ 3  7 11]
 [ 4  8 12]]
m.flags['C_CONTIGUOUS'] True
check: size times itemsize = 12 * 8 = 96 and nbytes is 96
munotes.in48

Practical 4 continued: the Dimensions and Attributes of an Array

AttributeWhat it givesNote
ndimthe number of dimensionsequals len(shape)
shapea tuple of lengthsthe one MU means by "dimensions"
sizetotal elementsthe shape multiplied out
dtypethe element typeone type for the whole array
itemsizebytes per element8 for int64, 4 for int32
nbytesbytes in allalways size × itemsize
Tthe transposea view, not a copy
flagshow the memory is laid outC_CONTIGUOUS means row by row

None of these is a method, so none of them takes brackets. m.shape is the shape; m.shape() raises TypeError: 'tuple' object is not callable. This is the single most common error on this exercise.

import numpy as np

m = np.arange(6).reshape(2, 3)
print(m.shape())
TypeError: 'tuple' object is not callable

The two that ARE methods are m.reshape(...), m.ravel(), m.flatten(), m.astype(...), m.sum() and the other calculations. The rule is simple: an attribute describes the array as it is, and a method does something.

dtype, and what it costs

An array holds one type, and the type decides the memory.

import numpy as np

for dtype in [np.int8, np.int16, np.int32, np.int64,
              np.float32, np.float64, np.bool_]:
    a = np.ones(1000, dtype=dtype)
    print(f"{str(a.dtype):<9} itemsize {a.itemsize}   1000 elements use {a.nbytes:>5} bytes")
int8      itemsize 1   1000 elements use  1000 bytes
int16     itemsize 2   1000 elements use  2000 bytes
int32     itemsize 4   1000 elements use  4000 bytes
int64     itemsize 8   1000 elements use  8000 bytes
float32   itemsize 4   1000 elements use  4000 bytes
float64   itemsize 8   1000 elements use  8000 bytes
bool      itemsize 1   1000 elements use  1000 bytes

The same thousand numbers take 1000 bytes as int8 and 8000 as int64. That is the reason NumPy exists and the reason it asks you to choose.

Changing the type is done with astype, which makes a copy:

import numpy as np

a = np.array([1.7, 2.3, -1.7, 3.9])

print("a          ", a, a.dtype)
print("astype(int)", a.astype(int), a.astype(int).dtype, "  it TRUNCATES, it does not round")
print("np.round   ", np.round(a).astype(int), "  round first to round")
print("a is unchanged:", a)

b = np.array([1, 0, 3, 0])
print("astype(bool)", b.astype(bool), "  zero is False, anything else True")
a           [ 1.7  2.3 -1.7  3.9] float64
astype(int) [ 1  2 -1  3] int64   it TRUNCATES, it does not round
np.round    [ 2  2 -2  4]   round first to round
a is unchanged: [ 1.7  2.3 -1.7  3.9]
astype(bool) [ True False  True False]   zero is False, anything else True

astype(int) truncates towards zero. 1.7 becomes 1 and -1.7 becomes -1, not -2. If you want rounding, round first. Marks are lost here every year.

Reshaping

reshape gives the same data a different shape. The total size must not change.

import numpy as np

a = np.arange(12)

print("a            ", a)
print("reshape(3, 4)")
print(a.reshape(3, 4))
print("reshape(4, 3)")
print(a.reshape(4, 3))
print("reshape(2, 2, 3)")
print(a.reshape(2, 2, 3))
print("reshape(-1, 4), the -1 is worked out for you")
print(a.reshape(-1, 4))
print("reshape(2, -1)")
print(a.reshape(2, -1))
print("shares memory with a?", np.shares_memory(a.reshape(3, 4), a))
munotes.in49

Practical 4 continued: the Dimensions and Attributes of an Array

a             [ 0  1  2  3  4  5  6  7  8  9 10 11]
reshape(3, 4)
[[ 0  1  2  3]
 [ 4  5  6  7]
 [ 8  9 10 11]]
reshape(4, 3)
[[ 0  1  2]
 [ 3  4  5]
 [ 6  7  8]
 [ 9 10 11]]
reshape(2, 2, 3)
[[[ 0  1  2]
  [ 3  4  5]]

 [[ 6  7  8]
  [ 9 10 11]]]
reshape(-1, 4), the -1 is worked out for you
[[ 0  1  2  3]
 [ 4  5  6  7]
 [ 8  9 10 11]]
reshape(2, -1)
[[ 0  1  2  3  4  5]
 [ 6  7  8  9 10 11]]
shares memory with a? True

A single -1 means "work this one out from the others". It saves arithmetic and it saves a bug when the length changes. Two of them is an error, because then there is nothing to work out from.

import numpy as np

a = np.arange(12)
print(a.reshape(5, 3))
ValueError: cannot reshape array of size 12 into shape (5,3)

Read the message: it names the size and the shape asked for. 5 × 3 = 15 and there are only 12 elements.

reshape returns a view where it can, as the last line of the run above shows. So reshaping is free, and writing into the reshaped array writes into the original.

Flattening, and the one difference that matters

import numpy as np

m = np.arange(1, 7).reshape(2, 3)

print(m)
print("ravel   ", m.ravel(), "   shares memory?", np.shares_memory(m.ravel(), m))
print("flatten ", m.flatten(), "   shares memory?", np.shares_memory(m.flatten(), m))

r = m.ravel()
f = m.flatten()
r[0] = 999
f[1] = 888
print("after writing 999 through ravel and 888 through flatten")
print(m)
[[1 2 3]
 [4 5 6]]
ravel    [1 2 3 4 5 6]    shares memory? True
flatten  [1 2 3 4 5 6]    shares memory? False
after writing 999 through ravel and 888 through flatten
[[999   2   3]
 [  4   5   6]]

ravel gives a view where it can, and flatten always gives a copy. They print the same thing, so the only way to tell them apart is to write through them, which the run above does. Use ravel when you only want to read, and flatten when you want your own copy.

Transpose, and the axes of a three dimensional array

import numpy as np

m = np.arange(1, 7).reshape(2, 3)
cube = np.arange(24).reshape(2, 3, 4)

print("m shape", m.shape)
print(m)
print("m.T shape", m.T.shape)
print(m.T)
print("m.T shares memory with m?", np.shares_memory(m.T, m))
print()
print("cube shape       ", cube.shape)
print("cube.T shape     ", cube.T.shape, "  T reverses ALL the axes")
print("transpose(1, 0, 2)", cube.transpose(1, 0, 2).shape, " or name the order you want")
print("swapaxes(0, 1)   ", cube.swapaxes(0, 1).shape)
munotes.in50

Practical 4 continued: the Dimensions and Attributes of an Array

m shape (2, 3)
[[1 2 3]
 [4 5 6]]
m.T shape (3, 2)
[[1 4]
 [2 5]
 [3 6]]
m.T shares memory with m? True

cube shape        (2, 3, 4)
cube.T shape      (4, 3, 2)   T reverses ALL the axes
transpose(1, 0, 2) (3, 2, 4)  or name the order you want
swapaxes(0, 1)    (3, 2, 4)

For two dimensions the transpose is what you expect: rows become columns. For three, T reverses every axis, so (2, 3, 4) becomes (4, 3, 2). When you want something else, name the new order with transpose or swap one pair with swapaxes.

The transpose is a view, which surprises people: no data moves at all. NumPy simply records that the axes are to be read in a different order.

Adding and removing a dimension

import numpy as np

a = np.array([1, 2, 3])

print("a               shape", a.shape)
print("a[np.newaxis, :] shape", a[np.newaxis, :].shape, "a row")
print(a[np.newaxis, :])
print("a[:, np.newaxis] shape", a[:, np.newaxis].shape, "a column")
print(a[:, np.newaxis])
print("np.expand_dims   shape", np.expand_dims(a, 0).shape, "the same thing, named")

fat = np.array([[[1, 2, 3]]])
print("fat             shape", fat.shape)
print("squeeze         shape", fat.squeeze().shape, "every length 1 axis dropped")
a               shape (3,)
a[np.newaxis, :] shape (1, 3) a row
[[1 2 3]]
a[:, np.newaxis] shape (3, 1) a column
[[1]
 [2]
 [3]]
np.expand_dims   shape (1, 3) the same thing, named
fat             shape (1, 1, 3)
squeeze         shape (3,) every length 1 axis dropped

np.newaxis inserts an axis of length 1 wherever you put it, which is how a flat array is turned into a row or into a column. squeeze does the opposite and drops every axis of length 1. Both are views.

The whole analysis, as one program

This is the program for the journal: it takes any array and reports everything.

import numpy as np


def describe(name, a):
    print(f"--- {name}")
    print(a)
    print(f"  ndim      {a.ndim}")
    print(f"  shape     {a.shape}")
    print(f"  size      {a.size}   (the shape multiplied out)")
    print(f"  dtype     {a.dtype}")
    print(f"  itemsize  {a.itemsize} bytes")
    print(f"  nbytes    {a.nbytes} bytes  ({a.size} x {a.itemsize})")
    print(f"  T shape   {a.T.shape}")


describe("one dimension, 4 integers", np.array([10, 20, 30, 40]))
describe("two dimensions, 2 by 3 floats", np.array([[1.5, 2.5, 3.5], [4.5, 5.5, 6.5]]))
describe("three dimensions, 2 by 2 by 2", np.arange(8).reshape(2, 2, 2))
--- one dimension, 4 integers
[10 20 30 40]
  ndim      1
  shape     (4,)
  size      4   (the shape multiplied out)
  dtype     int64
  itemsize  8 bytes
  nbytes    32 bytes  (4 x 8)
  T shape   (4,)
--- two dimensions, 2 by 3 floats
[[1.5 2.5 3.5]
 [4.5 5.5 6.5]]
  ndim      2
  shape     (2, 3)
  size      6   (the shape multiplied out)
  dtype     float64
  itemsize  8 bytes
  nbytes    48 bytes  (6 x 8)
  T shape   (3, 2)
--- three dimensions, 2 by 2 by 2
[[[0 1]
  [2 3]]

 [[4 5]
  [6 7]]]
  ndim      3
  shape     (2, 2, 2)
  size      8   (the shape multiplied out)
  dtype     int64
  itemsize  8 bytes
  nbytes    64 bytes  (8 x 8)
  T shape   (2, 2, 2)
munotes.in51

Practical 4 continued: the Dimensions and Attributes of an Array

Procedure

  1. Save as practical4c.py. Build arrays of zero, one, two and three dimensions.
  2. For each, print ndim, shape and size, and confirm ndim equals len(shape) and

size is the shape multiplied out.

  1. Print dtype, itemsize and nbytes, and check that nbytes is size times

itemsize.

  1. Build the same thousand values as int8, int32, int64, float32 and float64 and

compare nbytes.

  1. Use astype and show that converting a float to an int truncates rather than rounds.
  2. Reshape one array three ways, use -1 once, and show the size mismatch error.
  3. Compare ravel with flatten by writing through each one.
  4. Print T for a two and a three dimensional array and show it is a view.
  5. Add an axis with np.newaxis to make a row and a column, and drop axes with squeeze.
  6. Write the describe function and run it on three arrays.

Result

ndim equalled len(shape) for every array, and size equalled the shape multiplied out, 2 × 3 × 4 = 24 for the three dimensional case. nbytes equalled size times itemsize in every run. int8 used 1000 bytes for a thousand elements against 8000 for int64. astype(int) turned 1.7 into 1 and -1.7 into -1, confirming truncation. reshape(5, 3) on 12 elements raised ValueError. ravel shared memory with the array and flatten did not: writing 999 through ravel changed the array and writing 888 through flatten did not. T was a view for both two and three dimensions, and cube.T turned (2, 3, 4) into (4, 3, 2).

Where marks are lost

  • Writing m.shape() with brackets. These are attributes, not methods.
  • Confusing shape with size. Shape is a tuple, size is one number.
  • Saying astype(int) rounds. It truncates towards zero.
  • Claiming ravel and flatten are the same. One is a view, one is a copy, and only a

write through each one shows it.

  • Expecting cube.T to swap only the first two axes. It reverses all of them.
  • Reshaping to a size that does not match and not being able to read the error.
  • Copying a book's int64 instead of recording what your own machine printed.

For the journal

The aim in MU's words. The three word table of dimension, shape and size. The describe function and its output for arrays of one, two and three dimensions, which covers every attribute in one place. Then the nbytes = size × itemsize check written out with your own machine's numbers, the dtype memory comparison, and the ravel against flatten write through test. The conclusion in one sentence: ndim is the length of shape, size is the shape multiplied out, nbytes is size × itemsize, and reshape, ravel and T are views while flatten and astype are copies.

munotes.in52

Practical 4 continued: the Dimensions and Attributes of an Array

Quick revision

  • dimension is a direction, shape is the length in each direction, size is the

total count.

  • ndim equals len(shape). size is the shape multiplied out.
  • nbytes is always size × itemsize.
  • dtype is one type for the whole array. itemsize is 8 for int64, 4 for int32,

1 for int8.

  • These are attributes, so no brackets: ndim, shape, size, dtype, itemsize,

nbytes, T, flags.

  • astype makes a copy and truncates when going from float to int. Round first if you

want rounding.

  • reshape keeps the size and returns a view. One -1 is worked out for you; two is an

error.

  • ravel is a view where it can be; flatten is always a copy.
  • T is a view. On three dimensions it reverses all the axes; use transpose or

swapaxes to choose.

  • np.newaxis adds an axis of length 1; squeeze removes every axis of length 1.

Questions you should be able to answer

1. What is the difference between shape and size? Shape is a tuple giving the length in each direction. Size is one number, the total count of elements, which is the shape multiplied out.

2. How is ndim related to shape? ndim is the length of the shape tuple.

3. An array has shape (3, 4) and dtype int64. What is nbytes? 3 × 4 = 12 elements, and 12 × 8 = 96 bytes.

4. Why does m.shape() fail? shape is an attribute holding a tuple, not a method. The brackets try to call a tuple.

5. What does np.array([1.7, -1.7]).astype(int) give? [1 -1]. It truncates towards zero rather than rounding.

6. What does the -1 in reshape(-1, 4) mean? Work that length out from the size and the other dimensions. Only one -1 is allowed.

7. What is the difference between ravel and flatten? Both give a flat array. ravel returns a view where possible, so writing to it changes the original; flatten always returns a copy.

8. Is the transpose a copy? No, it is a view. No data moves; NumPy records that the axes are read in a different order.

9. cube has shape (2, 3, 4). What shape is cube.T? (4, 3, 2). T reverses every axis.

munotes.in53

Practical 4 continued: the Dimensions and Attributes of an Array

10. How do you turn a flat array of three numbers into a column? a[:, np.newaxis], which gives shape (3, 1).

Contents This chapter on its own page

munotes.in54

Chapter Ten

Practical 5: Functions, Armstrong Numbers and Palindromes

Syllabus topic Module 1, practical 5(a), "Write a function to check the input value is Armstrong and also write the function for Palindrome"

Aim

To write a function that reports whether a value is an Armstrong number, and a function that reports whether a value is a palindrome.

What a function is, and the four parts of one

A function is a named piece of program that takes values in and hands a value back. Writing one is the whole point of this exercise, so here are the four parts with their proper names.

def is_even(number):
    """True when number divides by 2 exactly."""
    return number % 2 == 0


print(is_even(10))
print(is_even(7))
print(is_even.__doc__)
True
False
True when number divides by 2 exactly.
PartIn the exampleWhat it does
def and the namedef is_evendeclares the function and names it
the parameternumberthe name the value arrives under
the docstringthe line in triple quotessays what it does, and help() reads it
returnreturn number % 2 == 0hands a value back and ends the function

Three rules follow, and each one is a mark.

A function that does not return gives back None. So a function that prints its answer instead of returning it cannot be used in an if, cannot be tested in a loop, and cannot be built on.

def print_even(number):
    print(number % 2 == 0)


def return_even(number):
    return number % 2 == 0


result_of_printing = print_even(10)
result_of_returning = return_even(10)
print("what print_even gave back :", result_of_printing)
print("what return_even gave back:", result_of_returning)
print("so only one of them can be used in a test:", "even" if result_of_returning else "odd")
True
what print_even gave back : None
what return_even gave back: True
so only one of them can be used in a test: even

An argument is what you pass; a parameter is what it arrives as. Examiners ask for the difference, and that is the whole of it.

A return ends the function at once. Nothing after it in that branch runs, which is what lets a function of several ifs work without elif.

Part one: the Armstrong number

The definition, properly

A number is an Armstrong number when the sum of its digits, each raised to the power of how many digits there are, equals the number itself.

NumberDigitsWorkedArmstrong
15331^3 + 5^3 + 3^3 = 1 + 125 + 27 = 153yes
370327 + 343 + 0 = 370yes
371327 + 343 + 1 = 371yes
407364 + 0 + 343 = 407yes
12331 + 8 + 27 = 36no
947446561 + 256 + 2401 + 256 = 9474yes
820844096 + 16 + 0 + 4096 = 8208yes
163441 + 1296 + 81 + 256 = 1634yes
919^1 = 9yes
munotes.in55

Practical 5: Functions, Armstrong Numbers and Palindromes

Look at the last three rows. The power is the number of digits, not always 3. Every single digit number is an Armstrong number, because any number to the power 1 is itself. A function written for three digits only gets 9474 and 8208 wrong and 9 wrong too.

The function

def is_armstrong(number):
    """True when the digits raised to the count of digits sum to the number."""
    digits = str(abs(number))
    power = len(digits)
    total = sum(int(d) ** power for d in digits)
    return total == abs(number)


for value in [153, 370, 371, 407, 123, 9474, 8208, 9, 0, 1, 100]:
    print(f"{value:>5}  armstrong? {is_armstrong(value)}")
  153  armstrong? True
  370  armstrong? True
  371  armstrong? True
  407  armstrong? True
  123  armstrong? False
 9474  armstrong? True
 8208  armstrong? True
    9  armstrong? True
    0  armstrong? True
    1  armstrong? True
  100  armstrong? False

Three lines do the work. str(abs(number)) gives the digits as characters, power = len(digits) counts them, and the sum(...) adds each digit raised to that power.

Showing the working, which is what the journal needs

A function that answers yes or no is correct. A function that shows the sum is a better journal entry, because the examiner can see that you know what is being computed.

def armstrong_working(number):
    digits = str(abs(number))
    power = len(digits)
    parts = [f"{d}^{power}" for d in digits]
    values = [int(d) ** power for d in digits]
    total = sum(values)
    sums = " + ".join(str(v) for v in values)
    verdict = "IS an Armstrong number" if total == abs(number) else "is not"
    return f"{number}: {' + '.join(parts)} = {sums} = {total}, so {number} {verdict}"


for value in [153, 9474, 8208, 123, 9]:
    print(armstrong_working(value))
153: 1^3 + 5^3 + 3^3 = 1 + 125 + 27 = 153, so 153 IS an Armstrong number
9474: 9^4 + 4^4 + 7^4 + 4^4 = 6561 + 256 + 2401 + 256 = 9474, so 9474 IS an Armstrong number
8208: 8^4 + 2^4 + 0^4 + 8^4 = 4096 + 16 + 0 + 4096 = 8208, so 8208 IS an Armstrong number
123: 1^3 + 2^3 + 3^3 = 1 + 8 + 27 = 36, so 123 is not
9: 9^1 = 9 = 9, so 9 IS an Armstrong number

Check one against the table by hand: for 8208 the parts are 8^4, 2^4, 0^4, 8^4, which is 4096 + 16 + 0 + 4096, and 4096 + 16 + 0 + 4096 = 8208.

Without converting to a string

An examiner may ask for it with arithmetic only, which is the same digit peeling as [Practical 2: the Fibonacci Series, and the Sum of the Digits].

munotes.in56

Practical 5: Functions, Armstrong Numbers and Palindromes

def is_armstrong_arithmetic(number):
    n = abs(number)
    power = 0
    counting = n
    while counting > 0:
        power += 1
        counting //= 10
    if n == 0:
        power = 1

    total = 0
    left = n
    while left > 0:
        total += (left % 10) ** power
        left //= 10

    return total == n


for value in [153, 9474, 8208, 123, 9, 0]:
    print(value, is_armstrong_arithmetic(value))
153 True
9474 True
8208 True
123 False
9 True
0 True

Two loops: the first counts the digits, the second raises and adds them. The if n == 0 guard is needed because zero has no digits to count but is one digit long.

Every Armstrong number up to 10000

def is_armstrong(number):
    digits = str(number)
    power = len(digits)
    return sum(int(d) ** power for d in digits) == number


found = [n for n in range(10000) if is_armstrong(n)]
print("how many:", len(found))
print(found)
how many: 17
[0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 153, 370, 371, 407, 1634, 8208, 9474]

That list is worth putting in the journal, because it makes a point no single test does. There are only 17 Armstrong numbers below 10000: the ten single digit numbers 0 to 9, then nothing at all until 153, then four in the whole three digit range (153, 370, 371, 407), then three in the whole four digit range (1634, 8208, 9474).

Note 1634, which is in that list and is not in the table above: 1^4 + 6^4 + 3^4 + 4^4 = 1 + 1296 + 81 + 256 = 1634. A program written for three digits misses it, and so does a student who memorised four examples.

Part two: the palindrome

The definition

A palindrome reads the same forwards and backwards. For a number, 121 is one and 123 is not. For a string, "level" is one, and "Madam" is one once case is ignored, and "A man a plan a canal Panama" is one once case, spaces and punctuation are ignored.

MU's row does not say whether she means a number or a string, so a complete answer does both.

For a number

def is_palindrome_number(number):
    """True when the number reads the same in reverse."""
    text = str(abs(number))
    return text == text[::-1]


for value in [121, 123, 1221, 7, 10, 1001, -121]:
    print(f"{value:>6}  reversed {str(abs(value))[::-1]:>6}  palindrome? "
          f"{is_palindrome_number(value)}")
   121  reversed    121  palindrome? True
   123  reversed    321  palindrome? False
  1221  reversed   1221  palindrome? True
     7  reversed      7  palindrome? True
    10  reversed     01  palindrome? False
  1001  reversed   1001  palindrome? True
  -121  reversed    121  palindrome? True

text[::-1] is the reversed string, and comparing a string with its own reverse is the whole test. Note what abs does to -121: it makes the answer yes. Whether a negative number can be a palindrome is a matter of definition, so say which you chose. Written strictly, "-121" reversed is "121-", so it is not one, and dropping abs gives that answer instead.

munotes.in57

Practical 5: Functions, Armstrong Numbers and Palindromes

For a string

def is_palindrome_text(text):
    """True when the letters and digits read the same in reverse, ignoring case."""
    cleaned = "".join(ch.lower() for ch in text if ch.isalnum())
    return cleaned == cleaned[::-1]


tests = ["level", "Madam", "Munotes", "A man, a plan, a canal: Panama",
         "Was it a car or a cat I saw?", "hello world", "", "a"]
for text in tests:
    cleaned = "".join(ch.lower() for ch in text if ch.isalnum())
    print(f"{text!r:<35} cleaned {cleaned!r:<25} palindrome? {is_palindrome_text(text)}")
'level'                             cleaned 'level'                   palindrome? True
'Madam'                             cleaned 'madam'                   palindrome? True
'Munotes'                           cleaned 'munotes'                 palindrome? False
'A man, a plan, a canal: Panama'    cleaned 'amanaplanacanalpanama'   palindrome? True
'Was it a car or a cat I saw?'      cleaned 'wasitacaroracatisaw'     palindrome? True
'hello world'                       cleaned 'helloworld'              palindrome? False
''                                  cleaned ''                        palindrome? True
'a'                                 cleaned 'a'                       palindrome? True

ch.isalnum() is True for a letter or a digit and False for a space, a comma or a colon, so the join keeps only what should be compared. ch.lower() makes the comparison ignore case.

The empty string comes out as a palindrome, which is correct: there is nothing in it that differs from its reverse. Say so in the journal rather than treating it as a bug.

Without slicing, which is the version to know

text[::-1] is Python being generous. An examiner who wants to see the algorithm asks for two pointers, one from each end, walking inwards.

def is_palindrome_two_pointers(text):
    cleaned = [ch.lower() for ch in text if ch.isalnum()]
    left = 0
    right = len(cleaned) - 1
    steps = 0
    while left < right:
        steps += 1
        if cleaned[left] != cleaned[right]:
            return False, steps
        left += 1
        right -= 1
    return True, steps


for text in ["level", "Madam", "hello", "A man, a plan, a canal: Panama"]:
    answer, steps = is_palindrome_two_pointers(text)
    print(f"{text!r:<35} palindrome? {answer!s:<6} after {steps} comparison(s)")
'level'                             palindrome? True   after 2 comparison(s)
'Madam'                             palindrome? True   after 2 comparison(s)
'hello'                             palindrome? False  after 1 comparison(s)
'A man, a plan, a canal: Panama'    palindrome? True   after 10 comparison(s)

Two things to say about it at the table. The loop runs while left < right, so it stops in the middle and never compares a character with itself. And it returns as soon as a pair differs, so "hello" is decided in one comparison, while the slicing version always builds the whole reversed string first.

The reverse of a number by arithmetic

The same idea without strings, which is how it is done in C:

def reverse_number(number):
    n = abs(number)
    reversed_value = 0
    while n > 0:
        reversed_value = reversed_value * 10 + n % 10
        n //= 10
    return reversed_value


for value in [121, 123, 1200, 7]:
    print(f"{value:>5} reversed is {reverse_number(value):>5}  "
          f"palindrome? {reverse_number(value) == value}")
munotes.in58

Practical 5: Functions, Armstrong Numbers and Palindromes

  121 reversed is   121  palindrome? True
  123 reversed is   321  palindrome? False
 1200 reversed is    21  palindrome? False
    7 reversed is     7  palindrome? True

reversed_value * 10 + n % 10 pushes each digit on at the right hand end. Note 1200 becomes 21, because the leading zeros of a reversed number do not exist, and 1200 is correctly not a palindrome.

Procedure

  1. Save as practical5a.py.
  2. Write is_armstrong(number) with a docstring, using the count of digits as the power.
  3. Test it on 153, 370, 371, 407, 9474, 8208, 9 and 123, and print the verdict for each.
  4. Write a version that prints the working, and check 8208 against the table by hand.
  5. Write is_palindrome_number(number) and test it on 121, 123, 1221 and 7.
  6. Write is_palindrome_text(text) that ignores case, spaces and punctuation, and test it

on "level", "Madam" and a full sentence.

  1. Write the two pointer version and print the number of comparisons.
  2. Both functions must return, not print, and be called from a loop of test values.

Result

is_armstrong reported yes for 153, 370, 371, 407, 9474, 8208, 9, 0 and 1 and no for 123 and 100. The working for 8208 was 4096 + 16 + 0 + 4096 = 8208, matching the table. Counted by the program, there are 17 Armstrong numbers below 10000. is_palindrome_number was correct on 121, 1221, 7 and 1001 and rejected 123 and 10. is_palindrome_text accepted "level", "Madam" and "A man, a plan, a canal: Panama" after cleaning, and rejected "Munotes" and "hello world". The two pointer version rejected "hello" after one comparison.

Where marks are lost

  • Not writing a function. MU's row says function, twice.
  • print instead of return, so the answer cannot be used in a test.
  • Hard coding the power as 3. Then 9474, 8208 and every single digit number are wrong.
  • Forgetting that every single digit number is an Armstrong number.
  • Not cleaning the string for a palindrome, so "Madam" fails on the capital and a

sentence fails on the spaces.

  • Testing once. A function is proved by a table of values, not by one call.
  • No docstring. It costs one line and it is what a function is expected to have.

For the journal

The aim in MU's words. The Armstrong table with the working shown for at least 153, 9474 and 8208, and one sentence saying the power is the count of digits so every one digit number qualifies. Both functions in full, with docstrings. The loop of test values and its output for each. For the palindrome, the cleaned string printed beside the original so the cleaning is visible, and one sentence on the two pointer version returning early. The conclusion: a function returns a value so it can be tested and reused, and both of these are one line of real work around a definition that has to be got right first.

munotes.in59

Practical 5: Functions, Armstrong Numbers and Palindromes

Quick revision

  • def name(parameter): then a docstring then return. An argument is passed, a

parameter receives.

  • A function with no return gives back None, which cannot be used in an if.
  • return leaves the function at once.
  • Armstrong: each digit raised to the count of digits, summed, equals the number.
  • 153, 370, 371, 407 are the three digit ones. 1634, 8208 and 9474 are the four digit ones.

Every one digit number is one. Counted: 17 below 10000.

  • sum(int(d) ** len(s) for d in s) where s = str(n) is the whole function.
  • Palindrome: reads the same reversed. text == text[::-1].
  • For a string, clean first: "".join(ch.lower() for ch in text if ch.isalnum()).
  • The two pointer version walks in from both ends while left < right and returns False at

the first mismatch.

  • Reversing a number by arithmetic: rev = rev * 10 + n % 10, then n //= 10.

Questions you should be able to answer

1. What is the difference between an argument and a parameter? The parameter is the name in the def line; the argument is the value passed in when the function is called.

2. What does a function return if it has no return statement? None.

3. Why is print inside the function not good enough here? Because the caller gets None, so the answer cannot be used in an if, stored, or tested in a loop.

4. Define an Armstrong number. A number equal to the sum of its own digits, each raised to the power of the number of digits.

5. Is 8208 an Armstrong number? Show it. Yes. It has four digits, and 8^4 + 2^4 + 0^4 + 8^4 = 4096 + 16 + 0 + 4096 = 8208.

6. Is 7 an Armstrong number? Yes. It has one digit and 7^1 = 7. Every single digit number qualifies.

7. Write the palindrome test in one line. return text == text[::-1], after cleaning the text if it may contain case or punctuation.

8. Why must a string be cleaned before the test? Because a capital letter, a space or a comma is a character like any other, so "Madam" and any sentence would fail on characters that a reader does not count.

munotes.in60

Practical 5: Functions, Armstrong Numbers and Palindromes

9. What advantage does the two pointer version have? It stops at the first mismatch, so a long string that differs early is decided at once, and it builds no reversed copy.

10. What does reversing 1200 by arithmetic give, and is 1200 a palindrome? 21, because the trailing zeros become leading zeros and do not exist. 1200 is not a palindrome.

Contents This chapter on its own page

munotes.in61

Chapter Eleven

Practical 5 continued: Recursion, and the Lambda

Syllabus topic Module 1, practical 5(b), "Write a recursive function to print the factorial for a given number", and 5(c), "Write a lambda function that checks whether a given string starts with a specific character"

Aim

To write a recursive function for the factorial, and a lambda function that reports whether a string starts with a given character.

Part one: recursion

What recursion is

A recursive function is one that calls itself. It works because each call is given a smaller problem than the one before, and because there is a case small enough to answer without calling again.

Every recursive function has exactly two parts, and leaving out either one breaks it:

PartWhat it isFor the factorial
the base casethe smallest problem, answered directly0! is 1
the recursive stepthe same problem, made smallern! is n times (n-1)!

The factorial

The factorial of n, written n!, is every whole number from 1 to n multiplied together.

nWritten outValue
0by definition1
111
22 x 12
33 x 2 x 16
44 x 3 x 2 x 124
55 x 4 x 3 x 2 x 1120
66 x 5 x 4 x 3 x 2 x 1720

0! is 1, not 0. It is a definition, and it is the base case, and it is the mark that gets lost.

The function

def factorial(n):
    """n! by recursion. The base case is 0, which is 1."""
    if n == 0:
        return 1
    return n * factorial(n - 1)


for value in range(9):
    print(f"{value}! = {factorial(value)}")
0! = 1
1! = 1
2! = 2
3! = 6
4! = 24
5! = 120
6! = 720
7! = 5040
8! = 40320

Three lines. The if is the base case and the return n * factorial(n - 1) is the recursive step, and because return ends the function there is no need for an else.

Traced call by call, which is what earns the marks

An examiner asking about recursion wants to hear what the calls actually do. Make the program say it:

def factorial(n, depth=0):
    pad = "  " * depth
    print(f"{pad}factorial({n}) called")
    if n == 0:
        print(f"{pad}base case reached, returning 1")
        return 1
    result = n * factorial(n - 1, depth + 1)
    print(f"{pad}factorial({n}) returns {n} x factorial({n - 1}) = {result}")
    return result


print("answer:", factorial(4))
factorial(4) called
  factorial(3) called
    factorial(2) called
      factorial(1) called
        factorial(0) called
        base case reached, returning 1
      factorial(1) returns 1 x factorial(0) = 1
    factorial(2) returns 2 x factorial(1) = 2
  factorial(3) returns 3 x factorial(2) = 6
factorial(4) returns 4 x factorial(3) = 24
answer: 24

Read the indentation. The calls go down to the base case first, each one waiting, and then the answers come back up, each multiplication happening on the way out. Nothing is multiplied until the bottom is reached. That is the whole mechanism, and the picture of it is worth drawing in the journal.

munotes.in62

Practical 5 continued: Recursion, and the Lambda

Each waiting call is held on the call stack, which is a real, finite piece of memory. That is what the next section is about.

Where recursion stops

def factorial(n):
    if n == 0:
        return 1
    return n * factorial(n - 1)


print(factorial(5000))
RecursionError: maximum recursion depth exceeded

Python refuses at around a thousand nested calls, and it says so rather than crashing the machine. The limit is a guard: each call holds its own variables on the stack, and a runaway recursion would otherwise use all the memory there is.

import sys

print("the limit is", sys.getrecursionlimit())
the limit is 1000

The limit can be raised with sys.setrecursionlimit, and you should not, because the real answer is a loop.

The same factorial three other ways

from math import factorial as math_factorial


def factorial_loop(n):
    result = 1
    for i in range(2, n + 1):
        result *= i
    return result


def factorial_while(n):
    result = 1
    while n > 1:
        result *= n
        n -= 1
    return result


for value in [0, 1, 5, 10, 20]:
    print(f"{value:>3}!  for loop {factorial_loop(value):<20} "
          f"while loop {factorial_while(value):<20} math {math_factorial(value)}")
  0!  for loop 1                    while loop 1                    math 1
  1!  for loop 1                    while loop 1                    math 1
  5!  for loop 120                  while loop 120                  math 120
 10!  for loop 3628800              while loop 3628800              math 3628800
 20!  for loop 2432902008176640000  while loop 2432902008176640000  math 2432902008176640000

All four agree. The loop has no depth limit and uses no stack, so factorial_loop(5000) works where the recursive one does not. math.factorial is the one to use in real work, and it is the one to mention and not submit, because MU asked for a recursive function.

Notice 20! in that output. Python integers have no fixed width, so a factorial that would overflow a 64 bit integer in C simply keeps going here. That is worth one sentence in the journal, because in C this exercise breaks at 21!.

from math import factorial

digits = str(factorial(100))
print("100! has", len(digits), "digits")
print("and it is exact, ending in", digits[-30:])
100! has 158 digits
and it is exact, ending in 916864000000000000000000000000

Recursion against iteration, honestly

RecursionLoop
Reads like the definitionyesno
Extra memoryone stack frame per callnone
Depth limitabout 1000none
Right for the factorialfor the exercisefor real work
Right for a treeyes, see Module 2awkward, needs your own stack

Recursion is not a worse loop. It is the right tool when the data itself is nested, which is exactly what a tree is, and [Practical 6: Tree Traversal, Pre-order, In-order and Post-order] is where it earns its place.

munotes.in63

Practical 5 continued: Recursion, and the Lambda

Part two: the lambda

What a lambda is

A lambda is a function with no name, written in one expression. These two are the same function:

def starts_with_a(text):
    return text.startswith("a")


starts_with_a_lambda = lambda text: text.startswith("a")

print(starts_with_a("apple"), starts_with_a_lambda("apple"))
print(starts_with_a("banana"), starts_with_a_lambda("banana"))
print(type(starts_with_a), type(starts_with_a_lambda))
True True
False False
<class 'function'> <class 'function'>

The shape is lambda parameters: one expression. Three rules follow:

  • There is no return. The expression IS the value handed back.
  • It is one expression. No if statement, no loop, no assignment, no several lines.

A conditional expression, a if cond else b, is allowed because it is an expression.

  • It is a function like any other, as type shows above, and it can be stored in a

variable, passed to another function, or put in a list.

MU's exercise

Her row asks for a lambda that checks whether a given string starts with a specific character, so the lambda takes two things: the string and the character.

starts_with = lambda text, char: text.startswith(char)

tests = [("munotes", "m"), ("munotes", "n"), ("Practical", "p"), ("Practical", "P"),
         ("", "a"), ("python", "py")]
for text, char in tests:
    print(f"does {text!r:<12} start with {char!r:<4}? {starts_with(text, char)}")
does 'munotes'    start with 'm' ? True
does 'munotes'    start with 'n' ? False
does 'Practical'  start with 'p' ? False
does 'Practical'  start with 'P' ? True
does ''           start with 'a' ? False
does 'python'     start with 'py'? True

Two results in there are worth a sentence each. "Practical" does not start with "p" because startswith is case sensitive, and "python" does start with "py" because startswith accepts a whole prefix, not only one character.

A case insensitive version, still one expression:

starts_with_ci = lambda text, char: text.lower().startswith(char.lower())

for text, char in [("Practical", "p"), ("MUNOTES", "m"), ("munotes", "X")]:
    print(f"{text!r:<12} starts with {char!r} ignoring case? {starts_with_ci(text, char)}")
'Practical'  starts with 'p' ignoring case? True
'MUNOTES'    starts with 'm' ignoring case? True
'munotes'    starts with 'X' ignoring case? False

And without startswith at all, which an examiner may ask for:

starts_with_index = lambda text, char: len(text) > 0 and text[0] == char

for text, char in [("munotes", "m"), ("munotes", "n"), ("", "m")]:
    print(f"{text!r:<10} {char!r} -> {starts_with_index(text, char)}")
'munotes'  'm' -> True
'munotes'  'n' -> False
''         'm' -> False

The len(text) > 0 is not decoration. ""[0] raises IndexError, and and stops as soon as the left side is False, so the guard protects the indexing. That is called short circuit evaluation and it is a viva question.

Where a lambda is actually used

A lambda is at its best as an argument to something else, and that is where you will meet it in the rest of this paper.

munotes.in64

Practical 5 continued: Recursion, and the Lambda

words = ["munotes", "apple", "Practical", "banana", "array", "Python"]

print("sorted plainly        ", sorted(words))
print("sorted ignoring case  ", sorted(words, key=lambda w: w.lower()))
print("sorted by length      ", sorted(words, key=lambda w: len(w)))
print("sorted by last letter ", sorted(words, key=lambda w: w[-1]))
print("only the a words      ", list(filter(lambda w: w.lower().startswith("a"), words)))
print("their lengths         ", list(map(lambda w: len(w), words)))
print("any starting with P?  ", any(map(lambda w: w.startswith("P"), words)))
print("all longer than 3?    ", all(map(lambda w: len(w) > 3, words)))
sorted plainly         ['Practical', 'Python', 'apple', 'array', 'banana', 'munotes']
sorted ignoring case   ['apple', 'array', 'banana', 'munotes', 'Practical', 'Python']
sorted by length       ['apple', 'array', 'banana', 'Python', 'munotes', 'Practical']
sorted by last letter  ['banana', 'apple', 'Practical', 'Python', 'munotes', 'array']
only the a words       ['apple', 'array']
their lengths          [7, 5, 9, 6, 5, 6]
any starting with P?   True
all longer than 3?     True
Used withWhat the lambda is for
sorted(items, key=...)what to sort each item by
filter(f, items)keep the items where it is True
map(f, items)apply it to every item
any and all with mapis it true of one, or of all
min and max with key=what to compare by

filter and map return an iterator, not a list, which is why list(...) is wrapped around them above. Print one without the list and you get <filter object at ...>, which is the commonest surprise here.

When NOT to use a lambda

# clear
by_length = sorted(["bb", "a", "ccc"], key=lambda w: len(w))
print(by_length)


# not clear: this should have been a def with a name and a docstring
def grade(mark):
    """The grade band for a mark out of 100."""
    if mark >= 75:
        return "distinction"
    if mark >= 60:
        return "first class"
    if mark >= 40:
        return "pass"
    return "fail"


print([grade(m) for m in [80, 65, 45, 30]])
['a', 'bb', 'ccc']
['distinction', 'first class', 'pass', 'fail']

The rule to state at the table. A lambda is for a one expression function used once, in place. Anything that needs a name, a docstring or a second line is a def. Writing the grade ladder as nested conditional expressions inside a lambda is legal and unreadable.

Procedure

  1. Save as practical5b.py. Write factorial(n) with a base case of 0 returning 1 and a

recursive step of n * factorial(n - 1), with a docstring.

  1. Print the factorial of 0 to 8 and check 5! = 120 by hand.
  2. Add a depth parameter and print each call and each return, then draw the down and up

picture in the journal.

  1. Call it with 5000 and record the RecursionError, and print

sys.getrecursionlimit().

  1. Write the loop version and show that it agrees and that it has no depth limit.
  2. Save as practical5c.py. Write starts_with = lambda text, char: text.startswith(char)
munotes.in65

Practical 5 continued: Recursion, and the Lambda

and test it, including a case difference and a two character prefix.

  1. Write the case insensitive version and the version without startswith.
  2. Use a lambda with sorted(key=), filter and map, and wrap the last two in list.

Result

The recursive factorial gave 1, 1, 2, 6, 24, 120, 720, 5040 and 40320 for 0 to 8, matching the table, and 5! = 120 by hand. The traced run showed four calls descending to the base case and the multiplications happening on the way back up. factorial(5000) raised RecursionError at a limit of 1000. The loop, the while loop and math.factorial all agreed with the recursion. The lambda reported that "munotes" starts with "m" and not with "n", that "Practical" does not start with a lower case "p", and that "python" does start with "py".

Where marks are lost

  • No base case, so the function recurses until RecursionError.
  • Saying 0! is 0. It is 1, and it is the base case.
  • Writing return n * factorial(n) instead of n - 1, which never gets smaller.
  • Using a loop when the row says recursive.
  • Putting return inside a lambda. It is a syntax error; the expression is the value.
  • Trying to fit an if statement into a lambda. Use the conditional expression, or a

def.

  • Printing filter(...) without list(...) and reporting the object address.
  • Not knowing where recursion stops. The limit is real; print it.

For the journal

Two entries under practical 5. For the recursion: the aim, the factorial table to 6, the two part table of base case and recursive step, the function, the run for 0 to 8, and the traced run with its indentation, which is the evidence that you know how the calls nest. Add the RecursionError and the limit. One sentence: the calls descend to the base case and the multiplications happen on the way back up.

For the lambda: the aim, the def and the lambda written side by side to show they are the same function, MU's starts_with lambda, the test table including the case difference, and one use of a lambda with sorted, filter and map. The conclusion: a lambda is a one expression function with no name, used where a function is being passed to something else.

Quick revision

  • A recursive function calls itself. It needs a base case and a recursive step that

makes the problem smaller.

  • Factorial: if n == 0: return 1 then return n * factorial(n - 1).
  • 0! = 1. 5! = 120. 20! is exact in Python because integers have no fixed width.
  • The calls go down to the base case and the answers come back up. Each waiting call sits on
munotes.in66

Practical 5 continued: Recursion, and the Lambda

the call stack.

  • The recursion limit is about 1000 and RecursionError is what exceeding it raises.

sys.getrecursionlimit() prints it.

  • A loop has no depth limit and uses no stack. Recursion is the right tool for nested data,

such as a tree.

  • A lambda is lambda params: expression. No return, one expression.
  • a if cond else b is allowed inside one, because it is an expression.
  • startswith is case sensitive and accepts a whole prefix, not just one character.
  • filter and map return iterators; wrap them in list() to print.
  • Use a lambda for a one expression function passed in place. Anything needing a name or a

docstring is a def.

Questions you should be able to answer

1. What two parts does every recursive function have? A base case, the smallest problem answered without recursing, and a recursive step that calls itself on a smaller problem.

2. What is the base case of the factorial, and what does it return? n equal to 0, and it returns 1.

3. In the traced run, when does the first multiplication happen? Only after the base case is reached. The calls descend first, each waiting, and the multiplications happen on the way back up.

4. What happens if the base case is missing or never reached? The function recurses until Python's limit is exceeded and raises RecursionError.

5. What is the recursion limit, and can it be changed? Around 1000 by default, printed by sys.getrecursionlimit(). It can be raised with sys.setrecursionlimit, but a loop is the right answer.

6. Why does this program not overflow at 21! as the same program in C does? Python integers have no fixed width, so they grow as needed.

7. Write a lambda that says whether a string starts with a character. lambda text, char: text.startswith(char).

8. Why is there no return in a lambda? Because the single expression after the colon is what it hands back. A return there is a syntax error.

9. What can a lambda not contain? A statement of any kind: no if statement, no loop, no assignment, no second line. A conditional expression is allowed.

10. Why does printing a filter(...) show an object rather than the values? Because filter returns an iterator that produces values on demand rather than a list. Wrap it in list().

11. Does "Practical".startswith("p") return True? No. startswith is case sensitive. Use text.lower().startswith(char.lower()).

Contents This chapter on its own page

munotes.in67

Chapter Twelve

Practical 6: Counting the Characters and the Words in a String

Syllabus topic Module 1, practical 6(a), "Write a program to compute number of characters and words in a string"

Aim

To compute the number of characters and the number of words in a string.

The question hiding in the exercise

Take the string "To be or not to be". How many characters does it have?

  • 18, if a space is a character. It is; a space is " ", a perfectly ordinary character.
  • 14, if you mean only the letters.

Both are defensible and MU's row does not choose. So a complete answer prints both, labelled. The same applies to words: does "well-known" count as one word or two? Does a trailing space create an empty word? The program below answers each of those and says what it answered.

The program

text = input("Enter a sentence: ")

characters_with_spaces = len(text)
characters_without_spaces = len(text.replace(" ", ""))
letters_only = sum(1 for ch in text if ch.isalpha())
words = len(text.split())

print(f"the string is {text!r}")
print(f"characters, including spaces : {characters_with_spaces}")
print(f"characters, excluding spaces : {characters_without_spaces}")
print(f"letters only                 : {letters_only}")
print(f"words                        : {words}")
To be or not to be
Enter a sentence: the string is 'To be or not to be'
characters, including spaces : 18
characters, excluding spaces : 13
letters only                 : 13
words                        : 6

Check them by hand. "To be or not to be" has 18 characters counting the five spaces, 18 - 5 = 13 without them, 13 letters, and 6 words. Note that the second and third figures agree here because there is no punctuation; the next section shows where they part company.

len, split and the three things they do

text = "  Munotes  makes   notes.  "

print(f"text          {text!r}")
print(f"len(text)     {len(text)}")
print(f"split()       {text.split()}")
print(f"len(split())  {len(text.split())}")
print(f"split(' ')    {text.split(' ')}")
print(f"len(split(' ')) {len(text.split(' '))}")
print(f"strip()       {text.strip()!r}")
text          '  Munotes  makes   notes.  '
len(text)     27
split()       ['Munotes', 'makes', 'notes.']
len(split())  3
split(' ')    ['', '', 'Munotes', '', 'makes', '', '', 'notes.', '', '']
len(split(' ')) 10
strip()       'Munotes  makes   notes.'

This is the most important output in the chapter.

split() with no argument is the right call. It treats any run of whitespace as one separator, ignores leading and trailing whitespace, and never produces an empty word. Three words in, three words out.

split(' ') splits on every single space. Two spaces together produce an empty string between them, and a leading space produces one at the front. On this string of three words it gives ten pieces, seven of which are empty.

The rule for the journal: use split() for counting words, never split(' '). It also handles a tab and a newline, which split(' ') does not.

Punctuation, and why the counts differ

text = "Well-known students don't stop; they keep going!"

print(f"text                  {text!r}")
print(f"characters with spaces {len(text)}")
print(f"characters no spaces   {len(text.replace(' ', ''))}")
print(f"letters only           {sum(1 for ch in text if ch.isalpha())}")
print(f"words by split()       {len(text.split())}  ->  {text.split()}")
munotes.in68

Practical 6: Counting the Characters and the Words in a String

text                  "Well-known students don't stop; they keep going!"
characters with spaces 48
characters no spaces   42
letters only           38
words by split()       7  ->  ['Well-known', 'students', "don't", 'stop;', 'they', 'keep', 'going!']

Four different counts of the same sentence, and every one of them is correct for the question it answers. "Well-known" is one word to split() because there is no space in it, and "don't" is one word for the same reason. Say which definition you used and the marks are safe.

The letter frequency table, which is the natural extension

from collections import Counter

text = "Munotes makes notes for Mumbai University students"

letters = [ch.lower() for ch in text if ch.isalpha()]
counts = Counter(letters)

print("total letters:", len(letters))
print("distinct letters:", len(counts))
print("the five commonest:", counts.most_common(5))
print()
for letter, count in sorted(counts.items()):
    print(f"  {letter}  {'*' * count}  {count}")
total letters: 44
distinct letters: 16
the five commonest: [('s', 6), ('t', 5), ('e', 5), ('m', 4), ('u', 4)]

  a  **  2
  b  *  1
  d  *  1
  e  *****  5
  f  *  1
  i  ***  3
  k  *  1
  m  ****  4
  n  ****  4
  o  ***  3
  r  **  2
  s  ******  6
  t  *****  5
  u  ****  4
  v  *  1
  y  *  1

Counter counts anything countable and most_common(n) gives the top n in order. It is in the standard library, so nothing needs installing. Counting by hand with a dictionary is the version to know as well, because an examiner may forbid the import:

text = "Munotes makes notes"

counts = {}
for ch in text.lower():
    if ch.isalpha():
        counts[ch] = counts.get(ch, 0) + 1

print(sorted(counts.items()))
print("the same by hand, without Counter")
[('a', 1), ('e', 3), ('k', 1), ('m', 2), ('n', 2), ('o', 2), ('s', 3), ('t', 2), ('u', 1)]
the same by hand, without Counter

counts.get(ch, 0) returns the current count, or 0 if the letter has not been seen. Without the default, the first occurrence of each letter would raise KeyError.

Counting words, not just how many

from collections import Counter

text = ("the notes are the notes that the students read "
        "and the notes the students read are these notes")

words = text.lower().split()
counts = Counter(words)

print("words in all      :", len(words))
print("different words   :", len(counts))
print("commonest three   :", counts.most_common(3))
print("words used once   :", sorted(w for w, c in counts.items() if c == 1))
print("longest word      :", max(words, key=len))
print("average word length:", round(sum(len(w) for w in words) / len(words), 2))
words in all      : 18
different words   : 8
commonest three   : [('the', 5), ('notes', 4), ('are', 2)]
words used once   : ['and', 'that', 'these']
longest word      : students
average word length: 4.28
munotes.in69

Practical 6: Counting the Characters and the Words in a String

max(words, key=len) returns the longest word, and key=len is the lambda idea from [Practical 5 continued: Recursion, and the Lambda] without needing to write a lambda at all, because len is already a function.

Sentences, lines and the other counts an examiner may ask for

text = """Munotes makes notes. The notes are free to read.
They cover the whole syllabus! Do they cover yours?
"""

print("characters          :", len(text))
print("characters no spaces:", len(text.replace(" ", "").replace("\n", "")))
print("words               :", len(text.split()))
print("lines               :", len(text.splitlines()))
print("sentences           :", sum(1 for ch in text if ch in ".!?"))
print("vowels              :", sum(1 for ch in text.lower() if ch in "aeiou"))
print("consonants          :", sum(1 for ch in text.lower()
                                   if ch.isalpha() and ch not in "aeiou"))
print("digits              :", sum(1 for ch in text if ch.isdigit()))
print("upper case          :", sum(1 for ch in text if ch.isupper()))
print("punctuation         :", sum(1 for ch in text if not ch.isalnum()
                                   and not ch.isspace()))
characters          : 101
characters no spaces: 83
words               : 18
lines               : 2
sentences           : 4
vowels              : 31
consonants          : 48
digits              : 0
upper case          : 4
punctuation         : 4

splitlines() is the right way to count lines: it handles the last line whether or not the text ends with a newline, which split("\n") does not.

The string methods worth remembering, all of which return a value and none of which change the string:

MethodWhat it gives
len(s)how many characters, spaces included
s.split()a list of words, any whitespace as the separator
s.splitlines()a list of lines
s.strip()the string without leading and trailing whitespace
s.replace(a, b)a new string with every a replaced
s.lower(), s.upper()a new string in one case
s.count(sub)how many times sub occurs
s.isalpha(), s.isdigit(), s.isalnum(), s.isspace()tests on a character
s.title(), s.capitalize()a new string with capitals adjusted
" ".join(words)the words joined back into a string

A string cannot be changed. Every method above returns a new one, so text.replace(" ", "") on its own line does nothing at all: the result has to be assigned or used. That is worth a sentence in the journal.

text = "munotes"
text.upper()
print("after text.upper() on its own line:", text)
text = text.upper()
print("after assigning it back           :", text)
after text.upper() on its own line: munotes
after assigning it back           : MUNOTES

Procedure

  1. Save as practical6a.py. Read a sentence with input().
  2. Count characters with len, characters without spaces with len(text.replace(" ", "")),

letters with isalpha, and words with split().

  1. Print all four, each labelled, and check them by hand against your own sentence.
  2. Run split() and split(' ') on a string with double spaces and a leading space, and
munotes.in70

Practical 6: Counting the Characters and the Words in a String

record both results.

  1. Add a letter frequency count, with Counter and then again with a plain dictionary and

get.

  1. Add the line, sentence, vowel, consonant and punctuation counts.
  2. Show that a string method returns a new string and does not change the original.

Result

For "To be or not to be" the program counted 18 characters including spaces, 13 excluding them, 13 letters and 6 words, all confirmed by hand. On " Munotes makes notes. ", split() gave 3 words and split(' ') gave 10 pieces of which 7 were empty strings, which is why split() is the right call. The letter frequency table and the word frequency table both ran, and text.upper() on its own line left the string unchanged, confirming that strings cannot be modified in place.

Where marks are lost

  • One unlabelled character count. Say whether spaces are included. Both is better.
  • split(' ') instead of split(), which counts empty strings between double spaces as

words.

  • len(text.split(" ")) on a sentence with a trailing space, which counts one word too

many.

  • Calling text.replace(" ", "") and expecting text to change. Strings are

unchangeable; assign the result.

  • Counting punctuation as letters, or as words. Say which you counted.
  • Using text.count(" ") + 1 for the word count. It is wrong for double spaces, for a

leading space and for an empty string.

  • No hand check. One sentence counted by hand beside the output proves the program.

For the journal

The aim in MU's words. One sentence saying that "number of characters" has two answers and that both are given. The program, and the output with all four counts labelled. Your own sentence counted by hand beside it. Then the split() against split(' ') comparison with both outputs, and one sentence: split() with no argument treats any run of whitespace as one separator and never produces an empty word. The letter frequency table. The conclusion: len counts characters including spaces, split() counts words, and every string method returns a new string rather than changing the old one.

Quick revision

  • len(s) counts every character, spaces included.
  • Without spaces: len(s.replace(" ", "")). Letters only:

sum(1 for ch in s if ch.isalpha()).

  • s.split() is the word count. Any whitespace, any amount, no empty words.
  • s.split(' ') splits on each single space and produces empty strings. Do not use it to

count.

  • s.count(" ") + 1 is not a word count.
  • s.splitlines() counts lines correctly whether or not the text ends in a newline.
  • Counter(items) counts, and most_common(n) ranks. Or a dictionary with

counts.get(ch, 0) + 1.

  • max(words, key=len) is the longest word.
  • A string cannot be changed. Every method returns a new one; assign it.
  • Say which definition you counted by. That sentence is the mark.
munotes.in71

Practical 6: Counting the Characters and the Words in a String

Questions you should be able to answer

1. How many characters are in "To be or not to be"? 18 including the five spaces, 13 without them. Both are correct answers to different questions, so say which.

2. Why is split() better than split(' ') for counting words? split() treats any run of whitespace as one separator and ignores whitespace at the ends, so it never produces an empty word. split(' ') splits on each single space and returns empty strings between consecutive spaces.

3. What is wrong with text.count(" ") + 1 as a word count? It counts one word per space, so double spaces, a leading space, a trailing space and an empty string all give the wrong answer.

4. What does text.replace(" ", "") do to text? Nothing. It returns a new string. Strings cannot be changed, so the result must be assigned.

5. How do you count the letters only? sum(1 for ch in text if ch.isalpha()), which skips spaces, digits and punctuation.

6. How many words does split() find in "Well-known students don't stop"? Four. "Well-known" and "don't" each contain no space, so each is one word.

7. Which two calls give a letter frequency count? Counter(letters) from collections, or a plain dictionary built with counts[ch] = counts.get(ch, 0) + 1.

8. Why is the default needed in counts.get(ch, 0)? Because the first time a letter is seen it is not in the dictionary, and counts[ch] would raise KeyError.

9. How do you count lines, and why not split("\n")? splitlines(). Splitting on the newline gives an extra empty piece when the text ends with one.

10. Find the longest word in one expression. max(words, key=len).

Contents This chapter on its own page

munotes.in72

Chapter Thirteen

Practical 6 continued: the Geometry Module, and pointyShapeVolume

Syllabus topic Module 1, practical 6(b), MU's own wording in full: "Create a file geometry.py to calculate base areas for shapes square and circle. In another file, write a function pointyShapeVolume(x, y, squareBase) that calculates the volume of a square pyramid if squareBase is True and of a right circular cone if squareBase is False. x is the length of an edge on a square if squareBase is True and the radius of a circle when squareBase is False. y is the height of the object. First use squareBase to distinguish the cases. Use the circleArea and squareArea from the geometry module to calculate the base areas."

Aim

To write a module geometry.py holding squareArea and circleArea, and in a second file a function pointyShapeVolume(x, y, squareBase) that uses them to return the volume of a square pyramid or of a right circular cone.

What a module is

A module is a .py file whose functions another file can use. That is all. Any file you write is already a module; the only new thing is import.

Why bother? Because squareArea is useful in more than one program, and a function written once and imported is a function corrected once. MU's exercise is a small example of the habit, and it is the point of the row: not the geometry, but the separation.

The two shapes, and the one formula they share

A pyramid or a cone is a solid that rises from a flat base to a single point. Every such solid has the same volume rule, whatever the base is:

volume = one third × base area × height

So the only difference between MU's two cases is the base.

BaseBase areaVolume
Square pyramida square of edge xx squaredone third × x squared × y
Right circular conea circle of radius xpi × x squaredone third × pi × x squared × y

"Right circular" means the apex sits directly above the centre of the circle, so the height is measured straight up. MU's y is that height in both cases.

Read the two base areas against each other now, because it decides which solid is bigger. The square's edge is x, so it spans x. The circle's radius is x, so it spans 2x across. The circle is therefore the larger base, and for the same x and y the cone holds more than the pyramid.

This is why her row says to compute the base area in the module and the volume in the other file. The volume function does not need to know which shape it has beyond choosing which area function to call.

geometry.py

"""Base areas for the shapes Major Practical 3 needs.

Imported by pointy.py, which multiplies a base area by a height.
"""

from math import pi


def squareArea(edge):
    """The area of a square of the given edge length."""
    if edge < 0:
        raise ValueError("an edge cannot be negative")
    return edge * edge


def circleArea(radius):
    """The area of a circle of the given radius."""
    if radius < 0:
        raise ValueError("a radius cannot be negative")
    return pi * radius * radius


if __name__ == "__main__":
    print("geometry.py run directly, so here is a self test")
    print("squareArea(4) =", squareArea(4))
    print("circleArea(1) =", circleArea(1))

Four things in that file are worth a sentence each in the journal.

munotes.in73

Practical 6 continued: the Geometry Module, and pointyShapeVolume

The names are MU's, squareArea and circleArea, spelled exactly as she prints them. Python's own style would be square_area, and here the specification wins over the style guide. Say so; it shows you know both.

from math import pi, never pi = 3.14. The library value is exact to the last digit a computer can hold, and using 3.14 loses accuracy in the third decimal place of every answer.

Each function has a docstring, so help(geometry) is useful.

The if __name__ == "__main__": block at the end is the standard guard, and the next section is about what it does.

The second file, which is MU's pointyShapeVolume

from geometry import squareArea, circleArea


def pointyShapeVolume(x, y, squareBase):
    """The volume of a square pyramid when squareBase is True, else a cone.

    x is the edge of the square, or the radius of the circle.
    y is the height of the solid.
    """
    if squareBase:
        base = squareArea(x)
    else:
        base = circleArea(x)
    return base * y / 3


print("a square pyramid, edge 3, height 9")
print("  base area", squareArea(3))
print("  volume   ", pointyShapeVolume(3, 9, True))
print("a cone, radius 3, height 9")
print("  base area", round(circleArea(3), 6))
print("  volume   ", round(pointyShapeVolume(3, 9, False), 6))
a square pyramid, edge 3, height 9
  base area 9
  volume    27.0
a cone, radius 3, height 9
  base area 28.274334
  volume    84.823002

Check the pyramid by hand, because it comes out exact: the base is 3 × 3 = 9, the height is 9, and 9 × 9 / 3 = 27.

if squareBase: is MU's "first use squareBase to distinguish the cases", and it is written as a plain truth test rather than if squareBase == True. The two behave the same and the first is what Python programmers write.

The two shapes compared, which is the interesting part

from math import pi
from geometry import squareArea, circleArea


def pointyShapeVolume(x, y, squareBase):
    base = squareArea(x) if squareBase else circleArea(x)
    return base * y / 3


print(f"{'x':>3} {'y':>3}   {'pyramid':>12} {'cone':>12}   {'ratio':>8}")
for x, y in [(1, 1), (2, 6), (3, 9), (5, 12), (10, 10)]:
    pyramid = pointyShapeVolume(x, y, True)
    cone = pointyShapeVolume(x, y, False)
    print(f"{x:>3} {y:>3}   {pyramid:>12.4f} {cone:>12.4f}   {pyramid / cone:>8.6f}")

print()
print("the ratio never changes, and it is 1 / pi =", round(1 / pi, 6))
print("because the square of edge x has area x squared")
print("and the circle of RADIUS x has area pi times x squared, which is larger")
  x   y        pyramid         cone      ratio
  1   1         0.3333       1.0472   0.318310
  2   6         8.0000      25.1327   0.318310
  3   9        27.0000      84.8230   0.318310
  5  12       100.0000     314.1593   0.318310
 10  10       333.3333    1047.1976   0.318310

the ratio never changes, and it is 1 / pi = 0.31831
because the square of edge x has area x squared
and the circle of RADIUS x has area pi times x squared, which is larger
munotes.in74

Practical 6 continued: the Geometry Module, and pointyShapeVolume

Every row has the same ratio, and it is 1 divided by pi, about 0.3183. The reason is worth one line in the journal because it is the kind of thing a viva asks. A square of edge x has area x squared. A circle of radius x has area pi × x squared, which is larger, because the circle reaches x in every direction while the square reaches x only along its edge. The one third and the height cancel, so the cone is always pi × the pyramid, and the pyramid is 1 over pi of the cone.

Note the trap. If MU's x had been the square's edge against the circle's diameter, the square would have been the larger of the two and the ratio would have been 4 over pi. Her row says radius, so it is 1 over pi. Read the specification, not the picture in your head.

import, and the four forms of it

import geometry
from geometry import squareArea
from geometry import squareArea, circleArea
from geometry import circleArea as areaOfCircle

print("import geometry            ->", geometry.squareArea(4))
print("from geometry import name  ->", squareArea(4))
print("renamed with as            ->", round(areaOfCircle(1), 6))
print("what the module knows      ->", [n for n in dir(geometry) if not n.startswith("_")])
import geometry            -> 16
from geometry import name  -> 16
renamed with as            -> 3.141593
what the module knows      -> ['circleArea', 'pi', 'squareArea']
FormHow the name is then usedWhen
import geometrygeometry.squareArea(4)keeps it obvious where it came from
from geometry import squareAreasquareArea(4)when you use it a lot. MU's row implies this one
from geometry import a, bboth directlythe usual case
from geometry import x as yy(...)when the name clashes with one of yours

There is a fifth form, from geometry import *, which imports everything. Do not use it. It hides where a name came from and it will silently overwrite one of your own functions with the module's. An examiner may ask why; that is why.

dir(module) lists what a module holds, which is how you check an import worked.

if __name__ == "__main__": and what it is for

Every module has a variable called __name__. When the file is run, it holds "__main__". When the file is imported, it holds the module's own name.

import geometry

print("inside this program, __name__ is  ", __name__)
print("inside the imported module it is  ", geometry.__name__)
print("so geometry's self test did not run when we imported it")
inside this program, __name__ is   __main__
inside the imported module it is   geometry
so geometry's self test did not run when we imported it
munotes.in75

Practical 6 continued: the Geometry Module, and pointyShapeVolume

Look at what did not happen: geometry.py ends with a self test that prints two lines, and importing it printed nothing. That is the guard working. Without it, every import of geometry would run the test, so pointy.py would print geometry's output before its own.

That is the whole reason the idiom exists, and it is a standard viva question on any exercise that has two files.

The complete answer, with the checks MU's row implies

from geometry import squareArea, circleArea


def pointyShapeVolume(x, y, squareBase):
    """Volume of a square pyramid (squareBase True) or a right circular cone.

    x: the edge of the square, or the radius of the circle
    y: the height of the solid
    """
    if x < 0 or y < 0:
        raise ValueError("a length cannot be negative")
    if squareBase:
        base = squareArea(x)
        shape = "square pyramid"
    else:
        base = circleArea(x)
        shape = "right circular cone"
    volume = base * y / 3
    print(f"  {shape}: base area {base:.6f}, height {y}, "
          f"volume = {base:.6f} x {y} / 3 = {volume:.6f}")
    return volume


print("MU's two cases:")
pointyShapeVolume(3, 9, True)
pointyShapeVolume(3, 9, False)

print("a degenerate case, zero height:")
pointyShapeVolume(5, 0, True)

print("and a rejected one:")
try:
    pointyShapeVolume(-1, 5, True)
except ValueError as error:
    print("  refused:", error)
MU's two cases:
  square pyramid: base area 9.000000, height 9, volume = 9.000000 x 9 / 3 = 27.000000
  right circular cone: base area 28.274334, height 9, volume = 28.274334 x 9 / 3 = 84.823002
a degenerate case, zero height:
  square pyramid: base area 25.000000, height 0, volume = 25.000000 x 0 / 3 = 0.000000
and a rejected one:
  refused: a length cannot be negative

Procedure

  1. Save geometry.py with squareArea(edge) and circleArea(radius), both with docstrings,

from math import pi at the top, and an if __name__ == "__main__": self test.

  1. Run geometry.py on its own and confirm the self test prints.
  2. Save a second file, pointy.py, in the same folder.
  3. In it, write from geometry import squareArea, circleArea.
  4. Write pointyShapeVolume(x, y, squareBase) with MU's exact name and parameter order.
  5. Use if squareBase: to choose the area function, then return base times height over 3.
  6. Call it for a pyramid of edge 3 and height 9 and check by hand that 9 × 9 / 3 = 27.
  7. Call it for a cone of radius 3 and height 9 and note it is smaller by the factor 4 over pi.
  8. Import geometry and confirm its self test did NOT run, which proves the guard.

Result

geometry.py and pointy.py were saved in one folder. squareArea(3) gave 9 and pointyShapeVolume(3, 9, True) gave 27.0, which matches the hand check 9 × 9 / 3 = 27. The cone of the same radius and height was LARGER, 84.823002 against 27.0, and the ratio of pyramid to cone was the same for every pair of values tried, 0.318310, which is 1 divided by pi. Importing geometry printed nothing from its self test, confirming the __name__ guard works. A negative length was refused with ValueError.

munotes.in76

Practical 6 continued: the Geometry Module, and pointyShapeVolume

Where marks are lost

  • Renaming the function. MU prints pointyShapeVolume; write that, not

pointy_shape_volume.

  • Changing the parameter order. It is (x, y, squareBase).
  • Renaming squareArea or circleArea to Python's snake case style. Her row names them.
  • Writing everything in one file. Her row says "in another file", and the import is the

exercise.

  • pi = 3.14. Use from math import pi.
  • Forgetting the one third. Then you have computed a prism, not a pyramid.
  • Using the square's edge as a radius or the reverse. x means both, and which one

depends on squareBase.

  • from geometry import *, which hides where names came from.
  • No if __name__ == "__main__": guard, so the module's test output appears inside your

program's output.

  • Both files in different folders. The import then fails with ModuleNotFoundError.

For the journal

Two files, both written out in full, clearly labelled geometry.py and pointy.py. MU's own wording as the aim, because it is a specification. The volume rule stated once, one third times base area times height, and the two row table of what the base is in each case. The output of both calls, with the hand check 9 × 9 / 3 = 27 written beside the pyramid. One sentence on if __name__ == "__main__": and what it stopped from printing. One sentence on the ratio of the two volumes being 4 over pi whatever the numbers, and why. The conclusion: the base area lives in the module because it is useful on its own, and the volume function only has to choose which base to ask for.

Quick revision

  • A module is a .py file you can import from. Both files must be in the same folder.
  • from geometry import squareArea, circleArea is the form MU's row implies.
  • Never from geometry import *.
  • volume = one third × base area × height, for a pyramid and for a cone alike.
  • Square pyramid: base is x x. Cone: base is pi x * x. y is the height in both.
  • from math import pi, never 3.14.
  • if squareBase: is the plain truth test, and it is MU's "first use squareBase".
  • pointyShapeVolume(3, 9, True) is 27.0, because 9 × 9 / 3 = 27.
  • The cone is always pi × the pyramid for the same x and y, because the circle of
munotes.in77

Practical 6 continued: the Geometry Module, and pointyShapeVolume

radius x is the bigger base. The pyramid is 1 over pi of the cone, about 0.3183.

  • __name__ is "__main__" when a file is run and the module's name when it is imported, so

the guard keeps a module's self test out of an importer's output.

  • dir(module) lists what it holds.

Questions you should be able to answer

1. What is a module? A .py file whose names another file can import. Every file you write is one.

2. Why did MU's row split this into two files? Because the base areas are useful on their own, and separating them is the habit the exercise teaches. The volume function only has to decide which area to ask for.

3. State the volume rule for both shapes. One third × the base area × the height. The only difference is that the base is x x for the square and pi x * x for the circle.

4. In pointyShapeVolume(x, y, squareBase), what is x? The edge of the square when squareBase is True, and the radius of the circle when it is False.

5. Compute the volume of a square pyramid of edge 3 and height 9. The base is 3 × 3 = 9, so the volume is 9 × 9 / 3 = 27.

6. For the same x and y, which is bigger, and by how much? The cone. A square of edge x has area x squared; a circle of radius x has area pi × x squared, and pi is greater than 1. So the cone is exactly pi × the pyramid, and the ratio of pyramid to cone is 1 over pi, about 0.3183.

7. What does if __name__ == "__main__": do? The block runs only when the file is run directly, not when it is imported, because __name__ is "__main__" only in the file being run.

8. What happens without that guard? Whatever the module prints at the bottom is printed every time somebody imports it, so your program's output is preceded by the module's.

9. Why not pi = 3.14? Because math.pi is accurate to the full precision a float holds, and 3.14 is wrong in the third decimal place, which spoils every answer computed from it.

10. The import fails with ModuleNotFoundError: No module named 'geometry'. Why? The two files are not in the same folder, or the module file is not named exactly geometry.py.

Contents This chapter on its own page

munotes.in78

Chapter Fourteen

Practical 7: a Common Member, and a Dictionary Sorted by Value

Syllabus topic Module 1, practical 7(a), "Write a program that takes two lists and returns True if they have at least one common member", and 7(b), "Write a Python script to sort (ascending and descending) a dictionary by value"

Aim

To write a function that reports whether two lists have at least one member in common, and a script that sorts a dictionary by its values, ascending and descending.

Part one: at least one common member

The obvious way, and what it costs

def have_common_member_nested(first, second):
    """True if any item of first is also in second. Checked pair by pair."""
    for a in first:
        for b in second:
            if a == b:
                return True
    return False


print(have_common_member_nested([1, 2, 3, 4, 5], [5, 6, 7]))
print(have_common_member_nested([1, 2, 3], [4, 5, 6]))
print(have_common_member_nested([], [1, 2]))
True
False
False

Two loops, one inside the other, comparing every item of the first list with every item of the second. It is correct, and the return True in the middle matters: the function stops the moment it finds a match rather than carrying on to the end.

The Python way

def have_common_member_in(first, second):
    """True if any item of first is also in second."""
    for a in first:
        if a in second:
            return True
    return False


def have_common_member_any(first, second):
    """The same thing as one expression."""
    return any(a in second for a in first)


def have_common_member_set(first, second):
    """True if the two sets intersect."""
    return bool(set(first) & set(second))


for a, b in [([1, 2, 3, 4, 5], [5, 6, 7]), ([1, 2, 3], [4, 5, 6]), ([], [1]),
             (["p", "q"], ["q", "r"])]:
    print(f"{str(a):<18} {str(b):<12} "
          f"in-loop {have_common_member_in(a, b)!s:<6} "
          f"any {have_common_member_any(a, b)!s:<6} "
          f"set {have_common_member_set(a, b)}")
[1, 2, 3, 4, 5]    [5, 6, 7]    in-loop True   any True   set True
[1, 2, 3]          [4, 5, 6]    in-loop False  any False  set False
[]                 [1]          in-loop False  any False  set False
['p', 'q']         ['q', 'r']   in-loop True   any True   set True

All three agree on every pair. any(...) is the one to write: it reads as the specification does, and it also stops at the first True.

The set version deserves a word. set(first) & set(second) is the intersection, the items in both. An empty set is falsy, so bool(...) of it is the answer. set(...).isdisjoint(...) says the opposite and is the most direct of all:

def have_common_member_disjoint(first, second):
    return not set(first).isdisjoint(second)


print(have_common_member_disjoint([1, 2, 3], [3, 4]))
print(have_common_member_disjoint([1, 2, 3], [4, 5]))
print("and the common members themselves:", set([1, 2, 3, 4]) & set([3, 4, 5]))
True
False
and the common members themselves: {3, 4}

Which is faster, counted rather than guessed

The three approaches do different amounts of work, and the way to show it is to count the comparisons rather than to time them, because a count is the same on every machine.

class Counting:
    """A number that records how many times it was compared."""

    comparisons = 0

    def __init__(self, value):
        self.value = value

    def __eq__(self, other):
        Counting.comparisons += 1
        return self.value == other.value

    def __hash__(self):
        return hash(self.value)

    def __repr__(self):
        return str(self.value)


def nested(first, second):
    for a in first:
        for b in second:
            if a == b:
                return True
    return False


def with_in(first, second):
    for a in first:
        if a in second:
            return True
    return False


def with_sets(first, second):
    return bool(set(first) & set(second))


size = 200
for label, worst in [("no member in common", False), ("the LAST items match", True)]:
    left = [Counting(n) for n in range(size)]
    right = [Counting(n + size) for n in range(size)]
    if worst:
        right[-1] = Counting(size - 1)
    for name, fn in [("nested loops", nested), ("in operator", with_in), ("sets", with_sets)]:
        Counting.comparisons = 0
        answer = fn(left, right)
        print(f"{label:<22} {name:<14} answer {answer!s:<6} "
              f"comparisons {Counting.comparisons}")
munotes.in79

Practical 7: a Common Member, and a Dictionary Sorted by Value

no member in common    nested loops   answer False  comparisons 40000
no member in common    in operator    answer False  comparisons 40000
no member in common    sets           answer False  comparisons 0
the LAST items match   nested loops   answer True   comparisons 40000
the LAST items match   in operator    answer True   comparisons 40000
the LAST items match   sets           answer True   comparisons 1

Read those six lines carefully, because the figures are more extreme than most notes claim.

The nested loops and the in operator both did 40000 comparisons in both cases, which is 200 × 200, the whole product of the two lengths. That is the same whether the answer is no or whether the match is sitting at the very last pair.

The set version did 0 comparisons when nothing was in common and 1 when one pair matched. Not a few hundred: zero.

The reason is the general lesson of the whole Module 2 half of this paper. A set finds a member by hashing, not by comparing. It computes a hash of each item, uses it to go straight to a bucket, and only calls == when two items land in the same bucket. With no common member, no two items ever land together, so == is never called at all. With one common member, it is called exactly once, on the pair that matched.

ApproachEquality comparisons hereReads as
nested loops40000, the full productthe definition
a in second with a list40000, hidden inside inshorter
any(a in second for a in first)40000 againthe specification
set(first) & set(second)0 or 1: it hashes insteadthe mathematics

So the rule is short. Write the any version for readability and the set version when the lists are large, and say which and why. That sentence is the difference between a correct answer and a good one.

A set cannot hold an unhashable item, so set([[1], [2]]) raises TypeError. Lists inside lists rule the set version out, and then the any version is the answer.

munotes.in80

Practical 7: a Common Member, and a Dictionary Sorted by Value

Part two: sorting a dictionary by value

What a dictionary is, in one paragraph

A dictionary holds pairs: a key and the value it maps to. Keys are unique, and a value is found from its key in about one step, by hashing, exactly as for a set. Since Python 3.7 a dictionary keeps the order in which keys were first inserted, which is what makes "sorting a dictionary" mean something at all.

marks = {"Aarti": 78, "Divya": 55, "Chetan": 90, "Bhavesh": 55, "Eshan": 67}

print("the dictionary      ", marks)
print("keys                ", list(marks.keys()))
print("values              ", list(marks.values()))
print("pairs               ", list(marks.items()))
print("one value by key    ", marks["Chetan"])
print("a missing key safely", marks.get("Farhan", "not enrolled"))
print("how many            ", len(marks))
the dictionary       {'Aarti': 78, 'Divya': 55, 'Chetan': 90, 'Bhavesh': 55, 'Eshan': 67}
keys                 ['Aarti', 'Divya', 'Chetan', 'Bhavesh', 'Eshan']
values               [78, 55, 90, 55, 67]
pairs                [('Aarti', 78), ('Divya', 55), ('Chetan', 90), ('Bhavesh', 55), ('Eshan', 67)]
one value by key     90
a missing key safely not enrolled
how many             5

items() is the one this exercise needs: it gives each pair as a tuple (key, value).

The point students miss

A dictionary has no sort method. A list does, a dictionary does not. sorted(anything) returns a list, so sorting a dictionary gives back a list of pairs, and if you want a dictionary again you build one from that list.

marks = {"Aarti": 78, "Bhavesh": 55, "Chetan": 90}

print("sorted(marks) sorts the KEYS and returns a list:", sorted(marks))
print("type of that                                   :", type(sorted(marks)).__name__)
print("sorted(marks.items()) sorts pairs by key       :", sorted(marks.items()))
print("marks itself is untouched                      :", marks)
sorted(marks) sorts the KEYS and returns a list: ['Aarti', 'Bhavesh', 'Chetan']
type of that                                   : list
sorted(marks.items()) sorts pairs by key       : [('Aarti', 78), ('Bhavesh', 55), ('Chetan', 90)]
marks itself is untouched                      : {'Aarti': 78, 'Bhavesh': 55, 'Chetan': 90}

Ascending, which is MU's first half

marks = {"Aarti": 78, "Divya": 55, "Chetan": 90, "Bhavesh": 55, "Eshan": 67}

pairs_ascending = sorted(marks.items(), key=lambda pair: pair[1])
print("as a list of pairs :", pairs_ascending)

ascending = dict(pairs_ascending)
print("as a dictionary    :", ascending)

print("printed as a table:")
for name, mark in pairs_ascending:
    print(f"  {name:<10} {mark:>3}")
as a list of pairs : [('Divya', 55), ('Bhavesh', 55), ('Eshan', 67), ('Aarti', 78), ('Chetan', 90)]
as a dictionary    : {'Divya': 55, 'Bhavesh': 55, 'Eshan': 67, 'Aarti': 78, 'Chetan': 90}
printed as a table:
  Divya       55
  Bhavesh     55
  Eshan       67
  Aarti       78
  Chetan      90

key=lambda pair: pair[1] is the whole answer. sorted calls that function on each pair and sorts by what it returns, and pair[1] is the value. Using pair[0] would sort by the name, which is the default anyway.

munotes.in81

Practical 7: a Common Member, and a Dictionary Sorted by Value

Descending, which is MU's second half

marks = {"Aarti": 78, "Divya": 55, "Chetan": 90, "Bhavesh": 55, "Eshan": 67}

descending = dict(sorted(marks.items(), key=lambda pair: pair[1], reverse=True))
print("reverse=True   :", descending)

also_descending = dict(sorted(marks.items(), key=lambda pair: -pair[1]))
print("negate the key :", also_descending)

print("the top scorer  :", max(marks, key=marks.get), max(marks.values()))
print("the lowest      :", min(marks, key=marks.get), min(marks.values()))
print("the top two     :", sorted(marks.items(), key=lambda p: p[1], reverse=True)[:2])
reverse=True   : {'Chetan': 90, 'Aarti': 78, 'Eshan': 67, 'Divya': 55, 'Bhavesh': 55}
negate the key : {'Chetan': 90, 'Aarti': 78, 'Eshan': 67, 'Divya': 55, 'Bhavesh': 55}
the top scorer  : Chetan 90
the lowest      : Divya 55
the top two     : [('Chetan', 90), ('Aarti', 78)]

reverse=True is the right way. Negating the key works for numbers and fails for strings and for dates, so it is a habit worth not forming.

Three ways to write the same key function

from operator import itemgetter

marks = {"Aarti": 78, "Divya": 55, "Chetan": 90}

print("lambda on the pair :", sorted(marks.items(), key=lambda pair: pair[1]))
print("itemgetter(1)      :", sorted(marks.items(), key=itemgetter(1)))
print("the KEYS by value  :", sorted(marks, key=marks.get))
lambda on the pair : [('Divya', 55), ('Aarti', 78), ('Chetan', 90)]
itemgetter(1)      : [('Divya', 55), ('Aarti', 78), ('Chetan', 90)]
the KEYS by value  : ['Divya', 'Aarti', 'Chetan']

itemgetter(1) from the operator module does exactly what the lambda does and is slightly faster. The third line sorts the keys by their values, which is often all you wanted, and it is the only one of the three with no pair in it.

One trap, and it is better met on the page than in the hall. key=marks.get looks as though it should work on the items too, and it does not:

marks = {"Aarti": 78, "Divya": 55, "Chetan": 90}
print(sorted(marks.items(), key=marks.get))
TypeError: '<' not supported between instances of 'NoneType' and 'NoneType'

sorted hands the key function each pair, and ("Aarti", 78) is not a key of the dictionary, so get returns None every time and the sort then tries to compare None with None. Read the message: it names the two NoneTypes. marks.get is a key function for sorted(marks, ...) only, never for the items.

Ties, and making the order predictable

Two students in the example share a mark of 55. Which comes first?

marks = {"Aarti": 78, "Divya": 55, "Chetan": 90, "Bhavesh": 55, "Eshan": 67}

print("by value only        :", sorted(marks.items(), key=lambda p: p[1]))
print("value then name      :", sorted(marks.items(), key=lambda p: (p[1], p[0])))
print("value down, name up  :", sorted(marks.items(), key=lambda p: (-p[1], p[0])))
by value only        : [('Divya', 55), ('Bhavesh', 55), ('Eshan', 67), ('Aarti', 78), ('Chetan', 90)]
value then name      : [('Bhavesh', 55), ('Divya', 55), ('Eshan', 67), ('Aarti', 78), ('Chetan', 90)]
value down, name up  : [('Chetan', 90), ('Aarti', 78), ('Eshan', 67), ('Bhavesh', 55), ('Divya', 55)]

Python's sort is stable, which means items that compare equal keep the order they were in. So sorting by value alone leaves the two 55s in dictionary order. To make the tie break explicit, sort by a tuple: (p[1], p[0]) sorts by value and then by name, because tuples are compared item by item.

munotes.in82

Practical 7: a Common Member, and a Dictionary Sorted by Value

The last line is the one an examiner likes: marks descending with names ascending inside each tie. Note that reverse=True cannot do it, because it would reverse the names too. That is the one place negating the number is the right tool.

Procedure

  1. Save as practical7a.py. Write the nested loop version with a docstring, returning True

from inside the loops.

  1. Write the in version, the any version and the set version, and print all four for the

same pairs of lists, including two empty lists.

  1. Count the comparisons for 200 items with nothing in common and with only the last items

matching, and record both figures.

  1. Save as practical7b.py. Build a dictionary of five names and marks with a tie in it.
  2. Print keys(), values() and items().
  3. Sort ascending with sorted(marks.items(), key=lambda p: p[1]), print the list of pairs

and then dict() of it.

  1. Sort descending with reverse=True.
  2. Sort by a tuple key to break the tie, and print the result.
  3. Print the highest and lowest with max(marks, key=marks.get) and min.

Result

All four common member functions agreed on every pair of lists tested, including the empty list. The comparison counts were 40000 for the nested loops and for the in operator in both cases, the full 200 × 200 product, against 0 for the set version with nothing in common and 1 when a single pair matched. The dictionary sorted ascending and descending by value, dict() rebuilt a dictionary from the sorted list of pairs, and the two students tied on 55 kept their insertion order, Divya before Bhavesh, until a tuple key was used, after which the tie broke by name. sorted(marks) returned a list of keys, confirming that a dictionary has no sort of its own.

Where marks are lost

  • Writing a script instead of a function for 7(a). MU's row says "returns True".
  • print("True") instead of return True.
  • Not stopping at the first match. A return inside the loop is the whole point.
  • Calling marks.sort(). A dictionary has no sort method; that is an AttributeError.
  • Forgetting .items() and sorting the keys by accident.
  • Sorting by pair[0] and calling it sorted by value.
  • Losing the dictionary. sorted gives a list; wrap it in dict() if a dictionary is

wanted.

  • Giving only ascending. MU's row says ascending and descending.
  • No tie in the test data, so the stability question never comes up.
munotes.in83

Practical 7: a Common Member, and a Dictionary Sorted by Value

For the journal

Two entries under practical 7. For the common member: the aim in MU's words, the nested loop version and the any version, the output for at least three pairs of lists including an empty one, the comparison counts for both the worst cases, and one sentence saying a set finds a member by hashing while a list looks at every item, so the set version wins on long lists and the any version reads best on short ones.

For the dictionary: the aim, the dictionary with a deliberate tie, the ascending sort as a list of pairs and then as a dictionary, the descending sort with reverse=True, and the tuple key that breaks the tie. One sentence: a dictionary cannot be sorted in place, sorted returns a list of pairs, and dict() turns it back.

Quick revision

  • 7(a) is a function that returns True or False.
  • Four ways: nested loops, a in second, any(a in second for a in first),

bool(set(first) & set(second)). set(first).isdisjoint(second) is the direct negative.

  • A return inside the loop stops at the first match.
  • & is set intersection. An empty set is falsy.
  • A set finds a member by hashing; a list looks at every item. So use sets on long lists.

Counted on two lists of 200: 40000 comparisons for the list versions, 0 or 1 for the set.

  • A set cannot hold an unhashable item such as a list.
  • A dictionary maps a key to a value. items() gives (key, value) pairs.
  • A dictionary has no sort method, and sorted(...) always returns a list.
  • Ascending by value: sorted(marks.items(), key=lambda p: p[1]).
  • Descending: add reverse=True. Back to a dictionary: dict(...) around it.
  • itemgetter(1) is the other key function for pairs. marks.get is a key function for

sorted(marks, ...) only: handed a pair it returns None, and the sort then fails trying to compare None with None.

  • Python's sort is stable, so equal items keep their order. Break a tie with a tuple key,

(p[1], p[0]), and use (-p[1], p[0]) for value down and name up.

  • max(marks, key=marks.get) is the key with the largest value.

Questions you should be able to answer

1. Write the one line answer to MU's 7(a). return any(a in second for a in first).

2. Why does the nested loop version return from inside the loops? So that it stops at the first match instead of comparing every remaining pair.

3. What does set(first) & set(second) give? The intersection: the items that are in both. An empty result is falsy, so bool of it answers the question.

munotes.in84

Practical 7: a Common Member, and a Dictionary Sorted by Value

4. Why is the set version faster on long lists? Because a set finds a member by hashing rather than by comparing, and only calls == when two items land in the same bucket. Counted here on two lists of 200: the list versions did 40000 comparisons, the set version 0 with nothing in common and 1 with a single match.

5. When can the set version not be used? When the lists hold unhashable items, such as lists, because a set cannot contain them.

6. Does marks.sort() work? No. A dictionary has no sort method; that raises AttributeError. Use sorted(marks.items(), ...).

7. What does sorted(marks) return? A list of the keys, sorted. Not a dictionary, and not the values.

8. Sort a dictionary by value, ascending. dict(sorted(marks.items(), key=lambda p: p[1])).

9. Two students have the same mark. Which appears first, and why? The one that was inserted into the dictionary first, because Python's sort is stable and leaves equal items in the order they were already in.

10. Sort by mark descending with names alphabetical inside each tie. sorted(marks.items(), key=lambda p: (-p[1], p[0])). reverse=True cannot do it, because it would reverse the names as well.

Contents This chapter on its own page

munotes.in85

Chapter Fifteen

Practical 8: the Tuple Return, Area and Circumference

Syllabus topic Module 1, practical 8(a), "Write a program to accept and pass radius to a function that returns area and circumference (using tuple)"

Aim

To accept a radius and pass it to a function that returns the area and the circumference of the circle as a tuple.

The two formulas

FormulaFor radius 7
Areapi × r × rpi × 49
Circumference2 × pi × r14 × pi

Both come from math.pi. Never from 3.14, and never from 22 divided by 7. Both of those are approximations and the output below prints how far out each one is.

from math import pi

print("math.pi          ", pi)
print("3.14 is out by   ", round(pi - 3.14, 10))
print("22/7 is out by   ", round(pi - 22 / 7, 10))
math.pi           3.141592653589793
3.14 is out by    0.0015926536
22/7 is out by    -0.0012644893

What a tuple is

A tuple is an ordered collection, written with commas, that cannot be changed once made. That last property is the whole reason it exists.

pair = (12.5, 7.0)

print("the tuple     ", pair)
print("first item    ", pair[0])
print("second item   ", pair[1])
print("how many       ", len(pair))
print("the type       ", type(pair).__name__)
the tuple      (12.5, 7.0)
first item     12.5
second item    7.0
how many        2
the type        tuple

The brackets are optional. What makes a tuple is the comma:

with_brackets = (1, 2, 3)
without_brackets = 1, 2, 3
one_item = (7,)
not_a_tuple = (7)
empty = ()

print("with brackets   ", with_brackets, type(with_brackets).__name__)
print("without brackets", without_brackets, type(without_brackets).__name__)
print("one item (7,)   ", one_item, type(one_item).__name__)
print("just (7)        ", not_a_tuple, type(not_a_tuple).__name__)
print("empty ()        ", empty, type(empty).__name__)
with brackets    (1, 2, 3) tuple
without brackets (1, 2, 3) tuple
one item (7,)    (7,) tuple
just (7)         7 int
empty ()         () tuple

(7) is the number 7 in brackets. (7,) is a tuple of one item. The trailing comma is what tells them apart, and it is worth knowing because it is a real source of bugs and a standard viva question.

And a tuple cannot be changed:

pair = (12.5, 7.0)
pair[0] = 99
TypeError: 'tuple' object does not support item assignment

Returning two values, which in Python is one tuple

A Python function returns one value, always. When it looks as though it returns two, what it actually returns is one tuple holding two things.

from math import pi


def circle(radius):
    """The area and the circumference of a circle, as a tuple."""
    area = pi * radius * radius
    circumference = 2 * pi * radius
    return area, circumference


result = circle(7)
print("what the function returned :", result)
print("its type                   :", type(result).__name__)
print("its length                 :", len(result))
what the function returned : (153.93804002589985, 43.982297150257104)
its type                   : tuple
its length                 : 2

There is no bracket in return area, circumference, and there is still a tuple, because the comma made one.

munotes.in86

Practical 8: the Tuple Return, Area and Circumference

Unpacking, which is how you use it

from math import pi


def circle(radius):
    """The area and the circumference of a circle, as a tuple."""
    return pi * radius * radius, 2 * pi * radius


area, circumference = circle(7)
print(f"unpacked: area {area:.4f}  circumference {circumference:.4f}")

both = circle(7)
print(f"indexed : area {both[0]:.4f}  circumference {both[1]:.4f}")

area_only, _ = circle(7)
print(f"only one: area {area_only:.4f}, the other thrown away")
unpacked: area 153.9380  circumference 43.9823
indexed : area 153.9380  circumference 43.9823
only one: area 153.9380, the other thrown away

area, circumference = circle(7) is tuple unpacking: the tuple on the right is taken apart and its items assigned to the names on the left, in order. The number of names must match the number of items, or Python says so:

def circle(radius):
    return radius * radius, 2 * radius


a, b, c = circle(7)
ValueError: not enough values to unpack (expected 3, got 2)

The _ in the third example is an ordinary name used by convention for a value you are deliberately ignoring.

Unpacking is the same mechanism as a, b = b, a + b in [Practical 2: the Fibonacci Series, and the Sum of the Digits], and the same as swapping two variables in one line:

first, second = "left", "right"
print("before:", first, second)
first, second = second, first
print("after :", first, second)
before: left right
after : right left

The program MU asks for

from math import pi


def circle_properties(radius):
    """Return the area and the circumference of a circle as a tuple.

    radius: the radius, which may not be negative
    """
    if radius < 0:
        raise ValueError("a radius cannot be negative")
    area = pi * radius * radius
    circumference = 2 * pi * radius
    return area, circumference


radius = float(input("Enter the radius: "))
area, circumference = circle_properties(radius)

print(f"radius        : {radius}")
print(f"area          : {area:.4f}")
print(f"circumference : {circumference:.4f}")
7
Enter the radius: radius        : 7.0
area          : 153.9380
circumference : 43.9823

Check one of them by hand: the circumference of a circle of radius 7 is 2 × 7 = 14 lots of pi, and 14 × 3.14159265 is about 43.98. The output agrees.

Why a tuple and not a list

An examiner asks this, and the answer is short.

TupleList
Written(a, b) or a, b[a, b]
Can be changednoyes
Can be a dictionary keyyesno
Fora fixed number of different thingsany number of similar things
Herearea and circumference: two different things, always twothe wrong shape

A tuple says "this is exactly two things and they mean different things". A list says "this is a collection of similar things and there may be any number of them". Area and circumference are the first, so a tuple is right, and MU's bracketed "(using tuple)" is not an arbitrary requirement.

munotes.in87

Practical 8: the Tuple Return, Area and Circumference

The unchangeability is not just a rule. Because a tuple cannot change, it can be hashed, so it can be a dictionary key or go in a set, which a list cannot:

cache = {}
cache[(7, "cm")] = "a circle of radius 7 centimetres"
print("a tuple as a dictionary key works:", cache)

try:
    cache[[7, "cm"]] = "this will not work"
except TypeError:
    print("a list as a key raises TypeError, because a list is unhashable")

print("hash of the tuple is a number:", isinstance(hash((7, "cm")), int))
a tuple as a dictionary key works: {(7, 'cm'): 'a circle of radius 7 centimetres'}
a list as a key raises TypeError, because a list is unhashable
hash of the tuple is a number: True

The wording of that TypeError changed between Python 3.12 and 3.14, so this listing prints a sentence of its own rather than the message. Your own run will show your version's wording, and the reason is the same in both: a key must be hashable, and only something that cannot change can safely be hashed.

The readable form: a named tuple

For three or more values, result[0] and result[1] stop being readable. A named tuple fixes that and is still a tuple.

from collections import namedtuple
from math import pi

Circle = namedtuple("Circle", ["radius", "area", "circumference"])


def circle_named(radius):
    """The same answer, with the fields named."""
    return Circle(radius, pi * radius * radius, 2 * pi * radius)


result = circle_named(7)
print("the whole thing :", f"Circle(radius={result.radius}, "
                           f"area={result.area:.4f}, "
                           f"circumference={result.circumference:.4f})")
print("by name         :", round(result.area, 4))
print("by index still  :", round(result[1], 4))
print("still a tuple?  :", isinstance(result, tuple))
print("unpacks too     :", [round(x, 4) for x in result])
the whole thing : Circle(radius=7, area=153.9380, circumference=43.9823)
by name         : 153.938
by index still  : 153.938
still a tuple?  : True
unpacks too     : [7, 153.938, 43.9823]

It is worth one sentence in the journal: a named tuple gives the fields names without giving up anything, because it still is a tuple, still unpacks, and still indexes.

A table of radii, which makes a better journal entry than one run

from math import pi


def circle_properties(radius):
    return pi * radius * radius, 2 * pi * radius


print(f"{'radius':>7} {'area':>12} {'circumference':>14} {'area/circ':>10}")
for radius in [0.5, 1, 2, 7, 10, 100]:
    area, circumference = circle_properties(radius)
    print(f"{radius:>7} {area:>12.4f} {circumference:>14.4f} "
          f"{area / circumference:>10.4f}")
print()
print("the last column is always half the radius, because")
print("area / circumference = (pi r r) / (2 pi r) = r / 2")
 radius         area  circumference  area/circ
    0.5       0.7854         3.1416     0.2500
      1       3.1416         6.2832     0.5000
      2      12.5664        12.5664     1.0000
      7     153.9380        43.9823     3.5000
     10     314.1593        62.8319     5.0000
    100   31415.9265       628.3185    50.0000

the last column is always half the radius, because
area / circumference = (pi r r) / (2 pi r) = r / 2
munotes.in88

Practical 8: the Tuple Return, Area and Circumference

Read the last column against the radius. It is exactly half the radius every time, because the pi and one r cancel. That is the kind of check that shows an examiner you understand the formulas rather than having copied them.

Procedure

  1. Save as practical8a.py. from math import pi at the top.
  2. Write circle_properties(radius) with a docstring, computing the area and the

circumference, and return area, circumference with no brackets.

  1. Read the radius with float(input()).
  2. Unpack the result into two names on one line.
  3. Print both to four decimal places, labelled.
  4. Check the circumference by hand: for radius 7 it is 14 × pi, about 43.98.
  5. Show that (7) is a number and (7,) is a tuple, and that a tuple cannot be assigned to.
  6. Add the table of several radii and the area over circumference column.

Result

circle_properties(7) returned a tuple of length 2, which unpacked into the area and the circumference. The circumference matched the hand check, 2 × 7 = 14 lots of pi, about 43.98. (7) was an int and (7,) a tuple, assigning to pair[0] raised TypeError, and unpacking into the wrong number of names raised ValueError. Across six radii the ratio of area to circumference was exactly half the radius each time, as the algebra requires.

Where marks are lost

  • Printing inside the function instead of returning. MU's row says "returns".
  • Returning a list when the row says tuple.
  • pi = 3.14. Use math.pi. The run above shows 3.14 is out by about 0.0016 and 22 over 7

by about 0.0013, and both errors are multiplied by the radius squared in an area.

  • Two separate functions, one for area and one for circumference. The row asks for one that

returns both.

  • (7) for a one item tuple. It needs the trailing comma, (7,).
  • Unpacking into the wrong number of names, which raises ValueError.
  • int(input()) for the radius, which refuses 2.5.
  • No hand check. One formula checked by hand beside the output is a mark.

For the journal

The aim in MU's words, including "(using tuple)". The two formulas in a table. The function with its docstring, and return area, circumference with the comma visible. The unpacking line. The run for a radius you chose, with the circumference checked by hand. Then the three line demonstration that (7) is a number and (7,) is a tuple, and the TypeError from trying to change one. One sentence on why a tuple: it is exactly two different things, it cannot be changed, and so it can be a dictionary key. The conclusion: a Python function returns one value, and returning two means returning one tuple, which the caller unpacks.

munotes.in89

Practical 8: the Tuple Return, Area and Circumference

Quick revision

  • Area = pi × r × r. Circumference = 2 × pi × r. Both from math.pi.
  • A tuple is ordered and cannot be changed. The comma makes it, not the brackets.
  • (7) is a number. (7,) is a tuple of one. () is the empty tuple.
  • return a, b returns one tuple of two items.
  • x, y = f() is tuple unpacking. The count of names must match, or ValueError.
  • _ is the conventional name for a value being ignored.
  • a, b = b, a swaps in one line, and it is the same mechanism.
  • A tuple can be a dictionary key because it cannot change; a list cannot.
  • namedtuple gives the fields names and is still a tuple: it unpacks and indexes as before.
  • area / circumference = r / 2, always. A useful self check.

Questions you should be able to answer

1. How many values can a Python function return? One. Returning two means returning one tuple with two items in it.

2. What makes a tuple, the brackets or the comma? The comma. 1, 2, 3 is a tuple with no brackets at all.

3. What is the difference between (7) and (7,)? (7) is the integer 7 in brackets. (7,) is a tuple of one item.

4. What is tuple unpacking? Assigning the items of a tuple to several names at once, as in area, circumference = circle(7). The names are filled in order.

5. What happens if the number of names does not match? ValueError, saying how many values there were and how many were expected.

6. Why does MU's row insist on a tuple rather than a list? Because the answer is exactly two different things, always two, and a tuple says that. A list says "any number of similar things".

7. Give one thing a tuple can do that a list cannot. Be a dictionary key, or go in a set, because it cannot be changed and so can be hashed.

8. Why not use 3.14, or 22 over 7? Because math.pi is exact to the full precision of a float. Measured in the first listing, 3.14 is out by about 0.0016 and 22 over 7 by about 0.0013, and in an area that error is multiplied by the radius squared.

9. What is the circumference of a circle of radius 7, without a calculator? Twice the radius is 2 × 7 = 14, so it is 14 lots of pi, about 43.98.

munotes.in90

Practical 8: the Tuple Return, Area and Circumference

10. What is the area divided by the circumference, for any circle? Half the radius, because (pi r r) / (2 pi r) cancels to r / 2.

Contents This chapter on its own page

munotes.in91

Chapter Sixteen

Practical 8 continued: Text Files, Binary Files, and the Last n Lines

Syllabus topic Module 1, practical 8(b), "Write a program to perform basic file operations on text files and binary files", and 8(c), "Write a Python program to read last n lines of a file"

Aim

To perform the basic operations on a text file and on a binary file, and to read the last n lines of a file.

The file every listing below uses

Munotes makes notes for MU students.
The notes follow the printed syllabus.
Every chapter is free to read.
Data structures come in Module 2.
Python comes in Module 1.
The journal is compulsory.

That file is written into the folder before each program runs, so the outputs below are the real outputs for exactly that content. Make the same file yourself and your run will match.

open, and the modes

handle = open("notes.txt", "r")
print("the object      :", type(handle).__name__)
print("its name        :", handle.name)
print("its mode        :", handle.mode)
print("is it closed?   :", handle.closed)
handle.close()
print("after close     :", handle.closed)
the object      : TextIOWrapper
its name        : notes.txt
its mode        : r
is it closed?   : False
after close     : True
ModeMeansIf the file existsIf it does not
"r"read, the defaultread itFileNotFoundError
"w"writethe contents are destroyedcreated
"a"appendwritten at the endcreated
"x"createFileExistsErrorcreated
"r+"read and writeopened at the startFileNotFoundError
add "b"binary, as in "rb", "wb"bytes instead of text

"w" empties the file the moment it is opened, before a single byte is written. That is the one mode that destroys work, and a student meets it by opening a file to check something and losing it. Use "r" to look and "a" to add.

with, which is not optional

with open("notes.txt", "r") as handle:
    first = handle.readline()
    print("inside the with :", first.strip())
    print("closed inside?  :", handle.closed)

print("closed after?   :", handle.closed)
inside the with : Munotes makes notes for MU students.
closed inside?  : False
closed after?   : True

with closes the file when the block ends. It closes it even if the program raises an error inside the block. Without that, a program which fails halfway leaves the file open, and on a write that can mean the last part was never saved to disk.

The rule for every file in this paper: always with. It is one line and it removes a whole class of bug.

Reading a text file, four ways

with open("notes.txt") as handle:
    whole = handle.read()
print("read() gives one string of", len(whole), "characters")

with open("notes.txt") as handle:
    lines = handle.readlines()
print("readlines() gives a list of", len(lines), "strings")
print("the first is", repr(lines[0]))

with open("notes.txt") as handle:
    one = handle.readline()
print("readline() gives one line:", repr(one))

print("looping over the handle, line by line:")
with open("notes.txt") as handle:
    for number, line in enumerate(handle, 1):
        print(f"  {number}: {line.rstrip()}")
read() gives one string of 194 characters
readlines() gives a list of 6 strings
the first is 'Munotes makes notes for MU students.\n'
readline() gives one line: 'Munotes makes notes for MU students.\n'
looping over the handle, line by line:
  1: Munotes makes notes for MU students.
  2: The notes follow the printed syllabus.
  3: Every chapter is free to read.
  4: Data structures come in Module 2.
  5: Python comes in Module 1.
  6: The journal is compulsory.
munotes.in92

Practical 8 continued: Text Files, Binary Files, and the Last n Lines

CallGivesMemory
read()the whole file as one stringthe whole file
read(n)the next n charactersn characters
readlines()a list of every line, newlines keptthe whole file
readline()the next line onlyone line
for line in handleeach line in turnone line at a time

The loop is the right way to read a file you did not write. It never holds more than one line, so it works on a file bigger than the machine's memory, and it is also the shortest to write.

Every line from a text file ends with "\n" except possibly the last. strip() removes whitespace from both ends and rstrip() from the right only, which is what you usually want.

Writing and appending

with open("scratch.txt", "w") as handle:
    handle.write("first line\n")
    handle.write("second line\n")
    handle.writelines(["third\n", "fourth\n"])

with open("scratch.txt") as handle:
    print("after w:", repr(handle.read()))

with open("scratch.txt", "a") as handle:
    handle.write("appended\n")

with open("scratch.txt") as handle:
    print("after a:", repr(handle.read()))

with open("scratch.txt", "w") as handle:
    handle.write("everything before this is gone\n")

with open("scratch.txt") as handle:
    print("after w again:", repr(handle.read()))
after w: 'first line\nsecond line\nthird\nfourth\n'
after a: 'first line\nsecond line\nthird\nfourth\nappended\n'
after w again: 'everything before this is gone\n'

Three things to record in the journal.

write does not add a newline. print does; write does not. Forget the "\n" and the whole file is one line.

writelines does not add newlines either, despite the name. It writes the strings one after another, so each must carry its own.

The third block proves what "w" does. The file had five lines and after reopening in "w" it has one. That is the demonstration to put in the journal beside the mode table.

The other basic operations

import os

print("does notes.txt exist?  ", os.path.exists("notes.txt"))
print("how big is it?         ", os.path.getsize("notes.txt"), "bytes")
print("is it a file?          ", os.path.isfile("notes.txt"))

with open("scratch.txt", "w") as handle:
    handle.write("something to rename\n")

os.rename("scratch.txt", "renamed.txt")
print("after rename, old gone?", not os.path.exists("scratch.txt"))
print("and the new one exists?", os.path.exists("renamed.txt"))

os.remove("renamed.txt")
print("after remove, exists?  ", os.path.exists("renamed.txt"))

print("the .txt files here    :", sorted(f for f in os.listdir(".") if f.endswith(".txt")))

if os.path.exists("gone.txt"):
    os.remove("gone.txt")
else:
    print("checking before removing avoids FileNotFoundError")
does notes.txt exist?   True
how big is it?          194 bytes
is it a file?           True
after rename, old gone? True
and the new one exists? True
after remove, exists?   False
the .txt files here    : ['notes.txt']
checking before removing avoids FileNotFoundError
munotes.in93

Practical 8 continued: Text Files, Binary Files, and the Last n Lines

os.path.exists before os.remove is the habit to form: removing a file that is not there raises FileNotFoundError.

Binary files, and what actually differs

A text file holds characters, and Python encodes them to bytes on the way out and decodes them on the way in. A binary file holds bytes, and nothing is translated.

data = bytes([72, 101, 108, 108, 111, 10, 0, 1, 2, 255])

with open("raw.bin", "wb") as handle:
    handle.write(data)

with open("raw.bin", "rb") as handle:
    back = handle.read()

print("what we wrote  :", data)
print("what came back :", back)
print("identical?     :", back == data)
print("its type       :", type(back).__name__)
print("byte by byte   :", list(back))
print("one byte       :", back[0], "which is an int, not a string")
what we wrote  : b'Hello\n\x00\x01\x02\xff'
what came back : b'Hello\n\x00\x01\x02\xff'
identical?     : True
its type       : bytes
byte by byte   : [72, 101, 108, 108, 111, 10, 0, 1, 2, 255]
one byte       : 72 which is an int, not a string

In binary mode you read and write bytes, not str. That is the whole difference, and mixing them up is what the two errors below are:

with open("raw.bin", "wb") as handle:
    handle.write("this is a string, not bytes")
TypeError: a bytes-like object is required, not 'str'
with open("notes.txt", "w") as handle:
    handle.write(b"these are bytes, not a string")
TypeError: write() argument must be str, not bytes

To turn one into the other, encode and decode, and always name the encoding:

text = "Munotes"
encoded = text.encode("utf-8")

print("the string    :", repr(text), len(text), "characters")
print("encoded       :", encoded, len(encoded), "bytes")
print("decoded again :", repr(encoded.decode("utf-8")))

with open("written.bin", "wb") as handle:
    handle.write(text.encode("utf-8"))

with open("written.bin", "rb") as handle:
    print("read back     :", repr(handle.read().decode("utf-8")))
the string    : 'Munotes' 7 characters
encoded       : b'Munotes' 7 bytes
decoded again : 'Munotes'
read back     : 'Munotes'

seek and tell

tell() says where you are in the file, as a number of bytes from the start. seek() moves there.

with open("notes.txt", "rb") as handle:
    print("at the start, tell() =", handle.tell())
    print("first 7 bytes        =", handle.read(7))
    print("now tell() =", handle.tell())

    handle.seek(0)
    print("after seek(0), tell() =", handle.tell())

    handle.seek(8)
    print("after seek(8), read 5 =", handle.read(5))

    handle.seek(0, 2)
    print("seek(0, 2) is the END, tell() =", handle.tell())

    handle.seek(-20, 2)
    print("the last 20 bytes     =", handle.read())
at the start, tell() = 0
first 7 bytes        = b'Munotes'
now tell() = 7
after seek(0), tell() = 0
after seek(8), read 5 = b'makes'
seek(0, 2) is the END, tell() = 194
the last 20 bytes     = b'rnal is compulsory.\n'

The second argument to seek is where to measure from: 0 the start, 1 the current position, 2 the end. Seeking relative to the end needs binary mode, which is why this listing opens with "rb"; in text mode only seek(0) and a position tell() gave you are allowed.

munotes.in94

Practical 8 continued: Text Files, Binary Files, and the Last n Lines

pickle, for saving a Python object

pickle writes any Python object to a binary file and reads it back as the same object. It is the answer to "how do I save a dictionary".

import pickle

marks = {"Aarti": 78, "Divya": 55, "Chetan": 90}
series = [0, 1, 1, 2, 3, 5]

with open("saved.pkl", "wb") as handle:
    pickle.dump(marks, handle)
    pickle.dump(series, handle)

with open("saved.pkl", "rb") as handle:
    marks_back = pickle.load(handle)
    series_back = pickle.load(handle)

print("the dictionary came back:", marks_back, marks_back == marks)
print("the list came back      :", series_back, series_back == series)
print("and it is a dict again  :", type(marks_back).__name__)
print("so this works           :", marks_back["Chetan"])
the dictionary came back: {'Aarti': 78, 'Divya': 55, 'Chetan': 90} True
the list came back      : [0, 1, 1, 2, 3, 5] True
and it is a dict again  : dict
so this works           : 90

Note that the objects come back in the order they were written, one load per dump. The file is binary, so the mode is "wb" and "rb"; pickling to a text file is an error.

One warning that belongs in the journal: never unpickle a file you did not create. Loading a pickle can run code, so it is not a safe way to accept data from somebody else. For data that crosses between programs or machines, json is the right choice.

Part two: the last n lines

The simple way

def last_n_readlines(path, n):
    """The last n lines. Reads the WHOLE file into memory."""
    with open(path) as handle:
        return handle.readlines()[-n:]


for line in last_n_readlines("notes.txt", 3):
    print(line.rstrip())
Data structures come in Module 2.
Python comes in Module 1.
The journal is compulsory.

[-n:] is the last n items of a list, and it is safe when the file has fewer than n lines: a slice clips rather than raising, as [Practical 3: Arrays, Basic Operations, Indexing and Slicing] showed.

This is the answer to write in the journal, and it has one defect worth naming: it holds the whole file in memory to keep three lines of it.

The better way

from collections import deque


def last_n_deque(path, n):
    """The last n lines, holding only n lines in memory."""
    with open(path) as handle:
        return list(deque(handle, maxlen=n))


for line in last_n_deque("notes.txt", 3):
    print(line.rstrip())

print("and asking for more lines than the file has:")
print([line.rstrip() for line in last_n_deque("notes.txt", 100)])
Data structures come in Module 2.
Python comes in Module 1.
The journal is compulsory.
and asking for more lines than the file has:
['Munotes makes notes for MU students.', 'The notes follow the printed syllabus.', 'Every chapter is free to read.', 'Data structures come in Module 2.', 'Python comes in Module 1.', 'The journal is compulsory.']
munotes.in95

Practical 8 continued: Text Files, Binary Files, and the Last n Lines

A deque with maxlen=n is a queue that throws away from the front when it is full. Feed it every line and what is left at the end is the last n, and it never held more than n. It still reads every line, so the time is the size of the file, but the memory is n lines whatever the file's size.

The way that reads only the tail

def last_n_seek(path, n, block=1024):
    """The last n lines, reading only the end of the file."""
    with open(path, "rb") as handle:
        handle.seek(0, 2)
        size = handle.tell()
        data = b""
        read = 0
        while data.count(b"\n") <= n and read < size:
            read = min(read + block, size)
            handle.seek(size - read)
            data = handle.read(read)
    lines = data.decode("utf-8").splitlines()
    return lines[-n:]


for line in last_n_seek("notes.txt", 3):
    print(line)

print("one line :", last_n_seek("notes.txt", 1))
print("all of it:", len(last_n_seek("notes.txt", 100)), "lines")
Data structures come in Module 2.
Python comes in Module 1.
The journal is compulsory.
one line : ['The journal is compulsory.']
all of it: 6 lines

This one seeks to the end, then reads backwards a block at a time until it has counted more than n newlines. On a file of a gigabyte it touches a kilobyte. It is the answer to give when the examiner asks about a large file, and it is the version a tool like tail actually uses.

The three compared

WayTimeMemoryUse it when
readlines()[-n:]the whole filethe whole filethe file is small, and it is the clearest
deque(handle, maxlen=n)the whole filen linesthe file is large but you may read it all
seek from the endthe tail onlya blockthe file is very large

All three were run on the same file above and printed the same three lines, which is what makes the comparison a comparison rather than a claim.

Procedure

  1. Create notes.txt with six lines of your own.
  2. Save as practical8b.py. Open it with with open(...) as handle and read it four ways:

read(), readlines(), readline() and a for loop.

  1. Write a new file in "w" mode, append to it in "a" mode, then reopen it in "w" and show

that the contents are gone.

  1. Use os.path.exists, os.path.getsize, os.rename and os.remove.
  2. Write a bytes object to a .bin file in "wb" and read it back in "rb", and show that

it is identical and that one byte is an int.

  1. Cause both mixing errors: a str written in binary mode and bytes written in text mode.
  2. Use encode and decode with "utf-8" named explicitly.
  3. Use tell() and seek(), including seek(0, 2) for the end.
  4. Pickle a dictionary and a list into one file and load them back in order.
  5. Save as practical8c.py. Write all three last n line functions and run all three on the
munotes.in96

Practical 8 continued: Text Files, Binary Files, and the Last n Lines

same file with n of 3, 1 and 100.

Result

All four read methods returned the same content of notes.txt. Writing in "w", appending in "a" and reopening in "w" showed the file emptied on the third open, which is the behaviour to remember. The binary round trip returned bytes identical to those written, and indexing one byte gave an int. Writing a str in binary mode and bytes in text mode each raised TypeError. seek(0, 2) reported the file size and a negative seek from the end returned the tail. pickle returned a dictionary and a list equal to the originals, in the order written. All three last n functions printed the same last three lines, and asking for 100 lines from a six line file returned six rather than raising.

Where marks are lost

  • No with. Then a file stays open when something fails, and on a write the data may not

reach the disk.

  • Opening in "w" to look at a file. It is empty before you read a byte.
  • Forgetting "\n" in write. The whole file becomes one line.
  • Expecting writelines to add newlines. It does not.
  • Mixing str and bytes. Binary mode needs bytes; text mode needs str.
  • Not naming the encoding in encode and decode.
  • Seeking from the end in text mode, which is not allowed.
  • Pickling to a text file. pickle needs "wb" and "rb".
  • Only giving readlines()[-n:] with nothing to say about a large file.
  • Not testing n larger than the file. A slice clips; say so.

For the journal

Two entries under practical 8. For the file operations: the aim, notes.txt written out as you made it, the mode table, and the four read methods with their output. Then the write, append and overwrite sequence with the file printed after each step, because that is the proof of what "w" does. Then the binary section: the bytes written and read back, the two TypeErrors, and one sentence saying text mode moves str and binary mode moves bytes. Then pickle with the dictionary that came back as a dictionary.

For the last n lines: all three functions, all three run on the same file, and the three row comparison table. One sentence: readlines() is clearest, deque(maxlen=n) holds only n lines, and seeking from the end touches only the tail, which is what a large file needs.

Quick revision

  • open(path, mode), and always inside with, which closes even on an error.
  • Modes: "r" read, "w" writes and EMPTIES first, "a" append, "x" create only,
munotes.in97

Practical 8 continued: Text Files, Binary Files, and the Last n Lines

"r+" both. Add "b" for binary.

  • Reading: read() whole, read(n) n characters, readlines() a list, readline() one,

for line in handle one line at a time and the right way for a big file.

  • write and writelines do not add newlines. print does.
  • rstrip() removes the trailing newline.
  • os.path.exists, os.path.getsize, os.rename, os.remove, os.listdir.
  • Text mode moves str; binary mode moves bytes. s.encode("utf-8") and

b.decode("utf-8") convert.

  • Indexing a bytes gives an int.
  • tell() is the position, seek(n) moves, seek(0, 2) is the end. Seeking from the end needs

binary mode.

  • pickle.dump(obj, handle) and pickle.load(handle), in "wb" and "rb", one load per

dump, in order. Never unpickle a file you did not create.

  • Last n lines: readlines()[-n:] simplest, deque(handle, maxlen=n) for memory, seek from

the end for a huge file. A slice clips, so n bigger than the file is safe.

Questions you should be able to answer

1. Why is with open(...) better than open(...) and close()? Because with closes the file when the block ends even if an error is raised inside it, so a file is never left open and a write is never left unflushed.

2. What does opening a file in "w" do to what was in it? It empties it immediately, before anything is written.

3. Which read method should you use on a file too big for memory? for line in handle, which holds one line at a time.

4. Does write add a newline? No, and neither does writelines. Each string must carry its own "\n".

5. What is the difference between text mode and binary mode? Text mode reads and writes str, encoding and decoding on the way; binary mode reads and writes bytes with no translation.

6. handle.write("hello") on a file opened "wb" fails. Why? Binary mode needs bytes. Write b"hello" or "hello".encode("utf-8").

7. What does back[0] give for a bytes object? An int, the value of that byte, not a one character string.

8. What is seek(0, 2)? Move to the end: offset 0 measured from the end of the file. It needs binary mode.

9. Why pickle rather than write the text yourself? Because pickle returns the object as the same type, so a dictionary comes back a dictionary and can be indexed straight away.

10. What is the danger of pickle? Loading a pickle can execute code, so a pickle from somebody else is not safe. Use json for data that crosses between programs.

munotes.in98

Practical 8 continued: Text Files, Binary Files, and the Last n Lines

11. Give the three ways to read the last n lines and what each costs. readlines()[-n:] reads the whole file into memory; deque(handle, maxlen=n) reads it all but keeps only n lines; seeking from the end reads only the tail, which is the one to use on a very large file.

12. What happens if you ask for the last 100 lines of a 6 line file? You get 6. A slice clips rather than raising, and the deque and the seek versions were tested for it too.

Contents This chapter on its own page

munotes.in99

Chapter Seventeen

Practical 9: Counting a Word in a File with a Regular Expression

Syllabus topic Module 1, practical 9(a), "Write a program to count the occurrences of a specific word in a file using regular expressions"

Aim

To count the occurrences of a specific word in a file, using a regular expression.

The file every listing below uses

Notes help students. A note is short; notes are longer.
The notebook on the table holds no notes at all.
NOTE: notes are free. Note the word note appears often.
Denote and footnote contain the letters but are other words.

Count "note" in that file by eye before reading on, and decide what you are counting: the letters n, o, t, e anywhere, or the whole word "note". They give very different answers.

Why a regular expression, and not count

with open("report.txt") as handle:
    text = handle.read()

print("str.count('note')                   :", text.count("note"))
print("str.count('note') ignoring case     :", text.lower().count("note"))
print("split and count whole words         :",
      sum(1 for word in text.split() if word == "note"))
print("split, lowered, punctuation left on :",
      sum(1 for word in text.lower().split() if word == "note"))
str.count('note')                   : 8
str.count('note') ignoring case     : 11
split and count whole words         : 2
split, lowered, punctuation left on : 3

Four answers to the same question, and not one of them is right.

count finds the letters, so it counts them inside "notes", "notebook", "Denote" and "footnote". Splitting on whitespace and comparing gives too few, because the word appears as "note;" and "note." with punctuation stuck to it, and those do not equal "note".

That is the gap a regular expression fills, and it is why MU's row says to use one.

A regular expression, built up from nothing

A regular expression is a small pattern language for describing text. re is the module, and it is in the standard library.

Build the pattern one piece at a time, because that is how it is understood and how it is explained at the table.

import re

text = "note notes NOTE Note notebook denote"

print("'note'      ", re.findall(r"note", text))
print("'notes'     ", re.findall(r"notes", text))
print("'note.'     ", re.findall(r"note.", text), "  a dot is ANY character")
print("'note\\.'    ", re.findall(r"note\.", text), "  backslash dot is a real dot")
print("'[Nn]ote'   ", re.findall(r"[Nn]ote", text), "  a class: one of these")
print("'note[sd]'  ", re.findall(r"note[sd]", text))
print("'notes?'    ", re.findall(r"notes?", text), "  ? means the s is optional")
print("'note\\w*'   ", re.findall(r"note\w*", text), "  \\w is a word character, * is any number")
print("'\\bnote\\b'  ", re.findall(r"\bnote\b", text), "  \\b is a word BOUNDARY")
'note'       ['note', 'note', 'note', 'note']
'notes'      ['notes']
'note.'      ['note ', 'notes', 'noteb']   a dot is ANY character
'note\.'     []   backslash dot is a real dot
'[Nn]ote'    ['note', 'note', 'Note', 'note', 'note']   a class: one of these
'note[sd]'   ['notes']
'notes?'     ['note', 'notes', 'note', 'note']   ? means the s is optional
'note\w*'    ['note', 'notes', 'notebook', 'note']   \w is a word character, * is any number
'\bnote\b'   ['note']   \b is a word BOUNDARY
munotes.in100

Practical 9: Counting a Word in a File with a Regular Expression

PieceMeans
notethose four letters, in that order
.any one character
\.a real full stop
[Nn]one character out of the set
?the thing before it is optional
*any number of the thing before it, including none
+one or more
\wa word character: a letter, a digit or an underscore
\sa whitespace character
\da digit
\ba word boundary, which matches no character at all
^ and $the start and the end of a line

\b is the whole exercise. It matches the empty place between a word character and a non-word character, so \bnote\b matches "note" standing alone and refuses it inside "notes", "notebook" and "denote". Notice in the output above that \bnote\b found exactly the standalone ones.

The raw string, and why every pattern has an r in front

import re

print("a plain string   :", "\bnote\b".encode("unicode_escape"))
print("a raw string     :", r"\bnote\b".encode("unicode_escape"))
print("with the r       :", re.findall(r"\bnote\b", "one note here"))
print("without the r    :", re.findall("\bnote\b", "one note here"))
a plain string   : b'\\x08note\\x08'
a raw string     : b'\\\\bnote\\\\b'
with the r       : ['note']
without the r    : []

In an ordinary Python string \b is the backspace character, so "\bnote\b" is not the pattern you typed at all and it matches nothing. The r prefix makes a raw string, in which a backslash is a backslash. Every regular expression in this book has the r, and writing a pattern without it is the second commonest error on this exercise after forgetting \b.

Python warns about an unrecognised escape in a plain string, so a pattern with \d in it and no r may raise a SyntaxWarning today and stop working in a future version. Use the r.

The program MU asks for

import re

word = "note"
pattern = re.compile(r"\b" + re.escape(word) + r"\b", re.IGNORECASE)

with open("report.txt") as handle:
    text = handle.read()

matches = pattern.findall(text)
print(f"the word {word!r} appears {len(matches)} time(s)")
print("the matches, exactly as they appeared:", matches)
the word 'note' appears 4 time(s)
the matches, exactly as they appeared: ['note', 'NOTE', 'Note', 'note']

Four things in those three lines of pattern.

r"\b" + word + r"\b" puts a boundary on each side, which is what makes it a word count.

re.escape(word) protects against a word containing a character the pattern language uses. Searching for "C++" without it is a broken pattern, because + means "one or more"; with it, the plus signs are treated as plain characters.

re.IGNORECASE makes it find "Note", "NOTE" and "note" alike. Whether you want that is a decision, so say which you chose in the journal. The output above shows the matches as they appeared in the file, which is how you demonstrate that the flag worked.

munotes.in101

Practical 9: Counting a Word in a File with a Regular Expression

re.compile builds the pattern once. Inside a loop over thousands of lines, compiling each time is wasted work, and a compiled pattern also reads better: pattern.findall(text).

Case sensitive, for comparison

import re

with open("report.txt") as handle:
    text = handle.read()

for label, flags in [("case sensitive ", 0), ("ignoring case  ", re.IGNORECASE)]:
    found = re.findall(r"\bnote\b", text, flags)
    print(f"{label} {len(found)} match(es): {found}")
case sensitive  2 match(es): ['note', 'note']
ignoring case   4 match(es): ['note', 'NOTE', 'Note', 'note']

Two defensible answers, and the difference is one flag. An examiner who asks "how many times does the word appear" will accept either as long as you say which you counted.

findall, finditer, search and match

import re

with open("report.txt") as handle:
    text = handle.read()

pattern = re.compile(r"\bnote\b", re.IGNORECASE)

print("findall gives the matched strings:", pattern.findall(text))
print()
print("finditer gives match objects, with WHERE:")
for found in pattern.finditer(text):
    print(f"  {found.group()!r} at characters {found.start()} to {found.end()}")
print()
first = pattern.search(text)
print("search gives the first match only :", first.group(), "at", first.start())
print("match anchors at the start        :", pattern.match(text))
print("and that is None, because the file does not begin with 'note'")
findall gives the matched strings: ['note', 'NOTE', 'Note', 'note']

finditer gives match objects, with WHERE:
  'note' at characters 23 to 27
  'NOTE' at characters 105 to 109
  'Note' at characters 127 to 131
  'note' at characters 141 to 145

search gives the first match only : note at 23
match anchors at the start        : None
and that is None, because the file does not begin with 'note'
CallGivesUse it for
findalla list of the matched stringscounting
finditera match object per matchthe position, or a big file
searchthe first match, or Noneis it there at all
matcha match only at the startchecking a whole field
subthe text with matches replacedediting
splitthe text split on the patternparsing

match is not "does it match"; it is "does it match here, at position zero". Confusing match with search is a standard trap, and the output above shows match returning None on a file that plainly contains the word.

Counting line by line, which is what a large file needs

import re

pattern = re.compile(r"\bnote\b", re.IGNORECASE)

total = 0
lines_with_it = 0
with open("report.txt") as handle:
    for number, line in enumerate(handle, 1):
        found = pattern.findall(line)
        if found:
            lines_with_it += 1
            total += len(found)
            print(f"  line {number}: {len(found)} -> {found}")

print(f"{total} occurrence(s) on {lines_with_it} line(s)")
  line 1: 1 -> ['note']
  line 3: 3 -> ['NOTE', 'Note', 'note']
4 occurrence(s) on 2 line(s)

This never holds more than one line, so it works on a file of any size, and it reports which lines matched, which is more useful than a single number. It is the version to put in the journal if you have room for only one.

munotes.in102

Practical 9: Counting a Word in a File with a Regular Expression

Counting every word, with the same tool

import re
from collections import Counter

with open("report.txt") as handle:
    text = handle.read()

words = re.findall(r"\b\w+\b", text.lower())
counts = Counter(words)

print("words in all    :", len(words))
print("different words :", len(counts))
print("commonest five  :", counts.most_common(5))
print("'note' exactly  :", counts["note"])
print("'notes' exactly :", counts["notes"])
print("words containing note:", sorted(w for w in counts if "note" in w))
words in all    : 40
different words : 29
commonest five  : [('notes', 4), ('note', 4), ('the', 4), ('are', 3), ('help', 1)]
'note' exactly  : 4
'notes' exactly : 4
words containing note: ['denote', 'footnote', 'note', 'notebook', 'notes']

r"\b\w+\b" is "a run of word characters standing alone", which is a better definition of a word than splitting on spaces, because it drops the punctuation by itself. The last line is the proof of the whole chapter: several different words contain the letters, and only one of them is the word.

A few more patterns worth having

import re

sample = ("Contact 9820012345 or 022-24567890. Email help@munotes.in on 05/10/2026. "
          "Marks: 78, 92 and 45 out of 100.")

print("phone like 10 digits :", re.findall(r"\b\d{10}\b", sample))
print("any run of digits    :", re.findall(r"\d+", sample))
print("a date dd/mm/yyyy    :", re.findall(r"\b\d{2}/\d{2}/\d{4}\b", sample))
print("something@something  :", re.findall(r"\b[\w.]+@[\w.]+\b", sample))
print("capitalised words    :", re.findall(r"\b[A-Z]\w*", sample))
print("replace digits with #:", re.sub(r"\d", "#", "Marks: 78 and 92"))
print("split on punctuation :", re.split(r"[;,.]\s*", "one, two; three. four")[:4])
phone like 10 digits : ['9820012345']
any run of digits    : ['9820012345', '022', '24567890', '05', '10', '2026', '78', '92', '45', '100']
a date dd/mm/yyyy    : ['05/10/2026']
something@something  : ['help@munotes.in']
capitalised words    : ['Contact', 'Email', 'Marks']
replace digits with #: Marks: ## and ##
split on punctuation : ['one', 'two', 'three', 'four']

{10} means exactly ten of the thing before it, {2,4} means two to four, and {2,} means two or more. Those, with \b, \d, \w and the character class, cover almost everything this paper can ask for.

Procedure

  1. Create report.txt with several lines in which the word appears alone, inside longer words,

in capitals and with punctuation attached.

  1. Save as practical9a.py. First count with str.count and with split, and record both wrong

answers.

  1. import re. Build the pattern up: the plain word, then [Nn], then \w*, then \bword\b,

printing findall for each so the difference is visible.

  1. Show that "\bnote\b" without the r matches nothing, and that r"\bnote\b" works.
  2. Write the answer: re.compile(r"\b" + re.escape(word) + r"\b", re.IGNORECASE), then

pattern.findall(text) and len(...).

  1. Print the matches themselves, not only the count, so the flag can be seen working.
  2. Run it once case sensitive and once ignoring case and record both numbers.
  3. Rewrite it to count line by line and report which lines matched.
munotes.in103

Practical 9: Counting a Word in a File with a Regular Expression

Result

str.count("note") and the two split counts all gave different and wrong answers on report.txt, because count finds the letters inside longer words and split leaves punctuation attached. re.findall(r"\bnote\b", text) with re.IGNORECASE gave the correct count, and printing the matches showed which capitalisations were included. Without the r prefix the same pattern found nothing, because \b in a plain string is a backspace character. match returned None while search found the word, confirming that match anchors at position zero. The line by line version gave the same total and also named the lines.

Where marks are lost

  • No \b. Then "note" is found inside "notes", "notebook" and "denote", and the count is too

high. This is the exercise.

  • No r prefix. "\b" is a backspace, so the pattern matches nothing at all.
  • Using str.count and calling it a regular expression.
  • split() and ==, which misses "note." and "note;".
  • Not saying whether case was ignored. Two different answers; name yours.
  • No re.escape around a word the user typed, which breaks on a word containing . or +.
  • Confusing match with search. match only looks at the start.
  • Compiling inside the loop. Compile once, outside.
  • Printing only the number. Printing the matches is what proves the pattern.

For the journal

The aim in MU's words. report.txt written out as you made it, with the word appearing alone, inside another word, capitalised and next to punctuation, because the file is what makes the exercise real. Then the wrong counts from str.count and split, with one sentence on why each is wrong. Then the pattern built up piece by piece with its findall output at each step, which is the best single thing on this page for an examiner. Then the answer, the count and the matches themselves. Then the case sensitive and case insensitive totals, and the line by line version. The conclusion in one sentence: \b matches the boundary between a word character and a non-word character, so \bnote\b counts the word and not the letters, and the pattern must be a raw string or \b is a backspace.

Quick revision

  • import re. Every pattern is a raw string, r"...", or \b is a backspace.
  • \b is a word boundary and matches no character. r"\bword\b" is the word count.
  • . any character, \. a real dot, [abc] one of these, [^abc] none of these.
  • ? optional, * any number, + one or more, {n} exactly n, {n,m} n to m.
  • \w word character, \s whitespace, \d digit. Capitals negate them: \W, \S, \D.
  • ^ start, $ end.
  • re.escape(word) before putting a user's word in a pattern.
  • re.IGNORECASE as the third argument, or a flag to re.compile.
  • findall the strings, finditer the positions, search the first anywhere.
munotes.in104

Practical 9: Counting a Word in a File with a Regular Expression

match looks only at the start. sub replaces, split splits.

  • re.compile once, outside any loop.
  • r"\b\w+\b" with Counter is a word frequency table that drops punctuation by itself.

Questions you should be able to answer

1. Why is text.count("note") the wrong answer to MU's row? Because it counts the letters wherever they appear, so it also counts them inside "notes", "notebook" and "denote".

2. Why is splitting on whitespace and comparing also wrong? Because punctuation stays attached, so "note." and "note;" do not equal "note" and are missed.

3. What does \b match? A word boundary: the empty position between a word character and a non-word character. It consumes nothing.

4. Write the pattern that counts the word "note". r"\bnote\b".

5. Why must the pattern be a raw string? Because in a normal Python string \b is the backspace character, so the pattern is not what you typed and matches nothing.

6. What is re.escape for? To treat a word's characters as plain text. Without it, searching for "C++" is a broken pattern, since + means "one or more".

7. What is the difference between re.match and re.search? match only matches at the very start of the string; search looks anywhere. A file containing the word in the middle gives None from match.

8. Which call do you use to count, and which to find out where? findall to count, because it returns the list of matched strings. finditer for positions, through start() and end().

9. Why compile the pattern? So it is built once rather than on every line, and so the code reads as pattern.findall(line).

10. Give a pattern for a ten digit phone number and one for a date as dd/mm/yyyy. r"\b\d{10}\b" and r"\b\d{2}/\d{2}/\d{4}\b".

Contents This chapter on its own page

munotes.in105

Chapter Nineteen

Practical 10: Comparing Two Dates in DD/MM/YYYY Form

Syllabus topic Module 1, practical 10(a), "Write a program that compares two dates (in DD/MM/YYYY format) and prints which one is earlier"

Aim

To compare two dates given in DD/MM/YYYY form and print which one is earlier.

The wrong answer that looks right

first = "05/10/2026"
second = "12/03/2020"

print(f"{first} < {second} as strings? {first < second}")
print("so the string comparison says the earlier date is:",
      first if first < second else second)
print("but 2020 is plainly before 2026")
05/10/2026 < 12/03/2020 as strings? True
so the string comparison says the earlier date is: 05/10/2026
but 2020 is plainly before 2026

The string comparison is not broken; it is answering a different question. Python compares strings character by character, so it looks at "0" against "1" first, decides that "05..." sorts before "12..." and stops. It compared the days, because in DD/MM/YYYY the day comes first.

A DD/MM/YYYY string cannot be compared as a string. In YYYY-MM-DD it can, and that is exactly why that format exists, but MU's row specifies DD/MM/YYYY, so the string has to be turned into a date first.

strptime, which parses a string into a date

from datetime import datetime

parsed = datetime.strptime("05/10/2026", "%d/%m/%Y")

print("what came back    :", parsed)
print("its type          :", type(parsed).__name__)
print("just the date     :", parsed.date())
print("day, month, year  :", parsed.day, parsed.month, parsed.year)
what came back    : 2026-10-05 00:00:00
its type          : datetime
just the date     : 2026-10-05
day, month, year  : 5 10 2026

The second argument is the format string, and it has to describe the input exactly.

CodeMeansFor 05/10/2026
%dday, two digits05
%mmonth, two digits10
%Yyear, four digits2026
%yyear, two digitswould be 26
%Bmonth by nameOctober
%bmonth abbreviatedOct
%Aweekday by name
%H:%M:%Shours, minutes, seconds

%Y is four digits and %y is two. Getting them the wrong way round is a common error and the message it gives is not obvious, so it is worth remembering.

strptime returns a datetime, which carries a time as well. For a date with no time, .date() gives a plain date, which is cleaner to compare and to print.

The program MU asks for

from datetime import datetime


def parse(text):
    """A date from a DD/MM/YYYY string."""
    return datetime.strptime(text.strip(), "%d/%m/%Y").date()


first_text = input("Enter the first date  (DD/MM/YYYY): ")
second_text = input("Enter the second date (DD/MM/YYYY): ")

first = parse(first_text)
second = parse(second_text)

print(f"first  : {first_text} parsed as {first}")
print(f"second : {second_text} parsed as {second}")

if first < second:
    print(f"{first_text} is EARLIER than {second_text}")
elif second < first:
    print(f"{second_text} is EARLIER than {first_text}")
else:
    print("the two dates are the same")

print(f"the difference is {abs((first - second).days)} day(s)")
05/10/2026
12/03/2020
Enter the first date  (DD/MM/YYYY): Enter the second date (DD/MM/YYYY): first  : 05/10/2026 parsed as 2026-10-05
second : 12/03/2020 parsed as 2020-03-12
12/03/2020 is EARLIER than 05/10/2026
the difference is 2398 day(s)
munotes.in113

Practical 10: Comparing Two Dates in DD/MM/YYYY Form

Once the strings are dates, < means what you want it to mean, and so do >, ==, <= and >=. A date knows how to compare itself with another date.

Subtracting two dates gives a timedelta

from datetime import datetime, date

first = datetime.strptime("05/10/2026", "%d/%m/%Y").date()
second = datetime.strptime("12/03/2020", "%d/%m/%Y").date()

gap = first - second
print("first - second    :", gap)
print("its type          :", type(gap).__name__)
print("in days           :", gap.days)
print("in weeks          :", gap.days // 7, "weeks and", gap.days % 7, "days")
print("about years       :", round(gap.days / 365.25, 2))
print("the other way     :", (second - first).days, "which is negative")
print("always positive   :", abs((second - first).days))
first - second    : 2398 days, 0:00:00
its type          : timedelta
in days           : 2398
in weeks          : 342 weeks and 4 days
about years       : 6.57
the other way     : -2398 which is negative
always positive   : 2398

A timedelta is a length of time. Subtracting two dates gives one, and .days is the number of whole days in it. Note that subtracting the later from the earlier gives a negative number, so abs(...) is what you want when you only care how far apart they are.

Adding a timedelta to a date gives another date, which is how you answer "what is the date 90 days from now":

from datetime import date, timedelta

start = date(2026, 10, 5)

print("the day           :", start)
print("plus 1 day        :", start + timedelta(days=1))
print("plus 90 days      :", start + timedelta(days=90))
print("minus 30 days     :", start - timedelta(days=30))
print("plus 3 weeks      :", start + timedelta(weeks=3))
print("crossing a year   :", date(2026, 12, 25) + timedelta(days=10))
the day           : 2026-10-05
plus 1 day        : 2026-10-06
plus 90 days      : 2027-01-03
minus 30 days     : 2026-09-05
plus 3 weeks      : 2026-10-26
crossing a year   : 2027-01-04

timedelta takes days, weeks, hours, minutes and seconds. It does not take months or years, and the reason is worth a sentence in the journal: a month has no fixed length, so "one month after the 31st of January" has no single right answer.

What happens when the date is not valid

from datetime import datetime

try:
    print(datetime.strptime("31/02/2026", "%d/%m/%Y"))
except ValueError as error:
    print("refused, and the kind of error is", type(error).__name__)
    print("the wording of the message changed between Python 3.12 and 3.14,")
    print("so read your own; both say the day is out of range for February")
refused, and the kind of error is ValueError
the wording of the message changed between Python 3.12 and 3.14,
so read your own; both say the day is out of range for February

strptime checks the calendar, so it refuses the 31st of February. It also refuses a string that does not match the format at all:

munotes.in114

Practical 10: Comparing Two Dates in DD/MM/YYYY Form

from datetime import datetime

print(datetime.strptime("2026-10-05", "%d/%m/%Y"))
ValueError: time data '2026-10-05' does not match format '%d/%m/%Y'

The complete answer catches both:

from datetime import datetime


def parse(text):
    """A date from a DD/MM/YYYY string, or None when it cannot be read."""
    try:
        return datetime.strptime(text.strip(), "%d/%m/%Y").date()
    except ValueError:
        return None


for text in ["05/10/2026", "5/10/2026", "29/02/2024", "29/02/2023",
             "31/02/2026", "2026-10-05", "hello"]:
    answer = parse(text)
    if answer is None:
        print(f"  {text!r:<14} REFUSED")
    else:
        print(f"  {text!r:<14} -> {answer}")
  '05/10/2026'   -> 2026-10-05
  '5/10/2026'    -> 2026-10-05
  '29/02/2024'   -> 2024-02-29
  '29/02/2023'   REFUSED
  '31/02/2026'   REFUSED
  '2026-10-05'   REFUSED
  'hello'        REFUSED

In your own program print the message from the ValueError as well, because it says what was wrong. It is left out here only because its exact wording changed between Python 3.12 and 3.14 and this page has to be true on both.

Three of those lines deserve a note.

"5/10/2026" with one digit is accepted. %d is documented as zero padded, and strptime is lenient about the padding when it parses, though strftime always writes two digits.

"29/02/2024" is accepted and "29/02/2023" is not, because 2024 is a leap year and 2023 is not. That single pair is the best possible test of a date program and it belongs in the journal.

The refusals come from the library, not from our own checking. strptime applies the calendar itself, which is why the program has no leap year test in it at all.

Leap years, since the 29th of February decides this exercise

from calendar import isleap

for year in [2000, 1900, 2023, 2024, 2026, 2100]:
    divisible_by_4 = year % 4 == 0
    divisible_by_100 = year % 100 == 0
    divisible_by_400 = year % 400 == 0
    print(f"{year}  /4 {divisible_by_4!s:<5} /100 {divisible_by_100!s:<5} "
          f"/400 {divisible_by_400!s:<5} leap? {isleap(year)}")
2000  /4 True  /100 True  /400 True  leap? True
1900  /4 True  /100 True  /400 False leap? False
2023  /4 False /100 False /400 False leap? False
2024  /4 True  /100 False /400 False leap? True
2026  /4 False /100 False /400 False leap? False
2100  /4 True  /100 True  /400 False leap? False

The rule in one line. A year is a leap year when it divides by 4, except that a century year must divide by 400. So 2000 is a leap year and 1900 is not, which is the pair that catches a program written with only the divide-by-4 test. calendar.isleap applies the whole rule, so use it rather than writing the test out.

Printing a date the way you want it

from datetime import date

day = date(2026, 10, 5)

print("the default          :", day)
print("%d/%m/%Y             :", day.strftime("%d/%m/%Y"))
print("%d %B %Y             :", day.strftime("%d %B %Y"))
print("%A, %d %b %Y         :", day.strftime("%A, %d %b %Y"))
print("%Y-%m-%d, sortable   :", day.strftime("%Y-%m-%d"))
print("the ISO form         :", day.isoformat())
print("the weekday number   :", day.weekday(), "with Monday as 0")
print("or Monday as 1       :", day.isoweekday())
print("day of the year      :", day.timetuple().tm_yday)
munotes.in115

Practical 10: Comparing Two Dates in DD/MM/YYYY Form

the default          : 2026-10-05
%d/%m/%Y             : 05/10/2026
%d %B %Y             : 05 October 2026
%A, %d %b %Y         : Monday, 05 Oct 2026
%Y-%m-%d, sortable   : 2026-10-05
the ISO form         : 2026-10-05
the weekday number   : 0 with Monday as 0
or Monday as 1       : 1
day of the year      : 278

strptime reads a string into a date and strftime writes a date out as a string. The two names are almost identical and the p and the f are the only difference. P is for parse and f is for format. Getting them the wrong way round is the commonest error with dates in any language.

Sorting a list of DD/MM/YYYY dates

from datetime import datetime

dates = ["05/10/2026", "12/03/2020", "01/01/2026", "31/12/2019", "29/02/2024"]

print("sorted as STRINGS, which is wrong:")
print(" ", sorted(dates))

print("sorted as dates, which is right:")
print(" ", sorted(dates, key=lambda text: datetime.strptime(text, "%d/%m/%Y")))

parsed = [datetime.strptime(text, "%d/%m/%Y").date() for text in dates]
print("earliest :", min(parsed).strftime("%d/%m/%Y"))
print("latest   :", max(parsed).strftime("%d/%m/%Y"))
print("in order :", [d.strftime("%d/%m/%Y") for d in sorted(parsed)])
sorted as STRINGS, which is wrong:
  ['01/01/2026', '05/10/2026', '12/03/2020', '29/02/2024', '31/12/2019']
sorted as dates, which is right:
  ['31/12/2019', '12/03/2020', '29/02/2024', '01/01/2026', '05/10/2026']
earliest : 31/12/2019
latest   : 05/10/2026
in order : ['31/12/2019', '12/03/2020', '29/02/2024', '01/01/2026', '05/10/2026']

Read the two sorted lists against each other. The string sort puts the 1st of January 2026 before the 5th of October 2026 by luck and the 31st of December 2019 after both of them, which is wrong. A key that parses each string fixes it, and it is the same key= argument as in [Practical 7: a Common Member, and a Dictionary Sorted by Value].

Procedure

  1. Save as practical10a.py. First compare the two dates as plain strings and record the wrong

answer, with one sentence saying why.

  1. from datetime import datetime. Parse each string with

datetime.strptime(text, "%d/%m/%Y").date().

  1. Read both dates with input().
  2. Compare with < and print which is earlier, with an elif and an else for equal dates.
  3. Subtract them and print abs(...).days.
  4. Try an invalid date, 31/02/2026, and record the ValueError and its message on your machine.
  5. Try 29/02/2024 and 29/02/2023 and record that one is accepted and one is not.
  6. Print one of the dates back out with strftime("%d/%m/%Y").
  7. Sort a list of five such strings as strings and then with a parsing key, and compare.

Result

Comparing "05/10/2026" with "12/03/2020" as strings said the first was earlier, which is wrong, because the string comparison reaches the day before the year. Parsed with strptime and compared as dates, the program correctly reported 12/03/2020 as the earlier and printed the gap in days. strptime refused 31/02/2026 and 2026-10-05 with ValueError, accepted 29/02/2024 and refused 29/02/2023, confirming that it checks the calendar including the leap year rule. isleap gave True for 2000 and False for 1900, which is the century rule. Sorting the five strings as strings gave a wrong order and sorting with a parsing key gave the right one.

munotes.in116

Practical 10: Comparing Two Dates in DD/MM/YYYY Form

Where marks are lost

  • Comparing the strings. It reaches the day first and is wrong. This is the exercise.
  • Splitting on / and comparing the pieces one at a time without converting them to numbers,

which has the same fault.

  • %y instead of %Y. One is two digits, the other four.
  • Getting the order wrong in the format, "%m/%d/%Y", which quietly reads the 5th of October

as the 10th of May.

  • Confusing strptime with strftime. p parses, f formats.
  • No try around the parse, so one bad input ends the program with a traceback.
  • Not testing the 29th of February. It is the one test that proves the parse checks the

calendar.

  • Using timedelta(months=1), which does not exist.

For the journal

The aim in MU's words, with her DD/MM/YYYY. The string comparison first, with its wrong answer and one sentence saying it compares the day because the day comes first. Then the parse with strptime and the format string explained code by code. The program, the run for two dates you chose, and the number of days between them. Then the three refusals: 31/02/2026, the wrong format, and 29/02/2023 against 29/02/2024 accepted. One sentence on leap years: divisible by 4, except a century year must divide by 400, and calendar.isleap applies the whole rule. The conclusion: a DD/MM/YYYY string cannot be compared as a string, and once parsed into a date the ordinary comparison operators are correct.

Quick revision

  • A DD/MM/YYYY string compares wrongly as a string, because the day comes first. YYYY-MM-DD

compares correctly, which is why that format exists.

  • datetime.strptime(text, "%d/%m/%Y") parses. .date() drops the time.
  • %d day, %m month, %Y four digit year, %y two digit, %B month name, %A weekday

name.

  • strptime parses, strftime formats. p for parse, f for format.
  • Once parsed, <, >, == all work on dates.
  • first - second gives a timedelta; .days is the whole days; it can be negative, so use

abs.

  • timedelta(days=, weeks=, hours=). There is no months or years, because a month has no

fixed length.

  • strptime checks the calendar: 31/02 and 29/02 in a non leap year both raise ValueError.
  • Leap year: divisible by 4, except a century year must be divisible by 400. So 2000 yes, 1900 no.
munotes.in117

Practical 10: Comparing Two Dates in DD/MM/YYYY Form

calendar.isleap knows.

  • weekday() has Monday as 0; isoweekday() has Monday as 1.
  • Sort DD/MM/YYYY strings with key=lambda t: datetime.strptime(t, "%d/%m/%Y").

Questions you should be able to answer

1. Why can two DD/MM/YYYY strings not be compared directly? Because a string comparison goes character by character from the left, and the leftmost part is the day, so it compares the days first and never reaches the years unless the days are equal.

2. Which date format does compare correctly as a string, and why? YYYY-MM-DD, because its parts run from the largest unit to the smallest with fixed widths.

3. Write the line that turns "05/10/2026" into a date. datetime.strptime("05/10/2026", "%d/%m/%Y").date().

4. What is the difference between %Y and %y? %Y is a four digit year and %y a two digit one.

5. What is the difference between strptime and strftime? strptime parses a string into a date; strftime formats a date into a string. The p is for parse and the f for format.

6. What do you get by subtracting two dates? A timedelta, whose .days is the number of whole days. It is negative if you subtract the later from the earlier.

7. Why does timedelta have no months argument? Because a month has no fixed length, so adding one to the 31st of January has no single correct answer.

8. What does strptime("31/02/2026", "%d/%m/%Y") do? It raises ValueError, because it checks the calendar and February has no 31st.

9. State the leap year rule and give the pair that tests it. Divisible by 4, except that a century year must be divisible by 400. 2000 is a leap year and 1900 is not.

10. Sort a list of DD/MM/YYYY strings into date order. sorted(dates, key=lambda text: datetime.strptime(text, "%d/%m/%Y")).

Contents This chapter on its own page

munotes.in118

Chapter Twenty

Practical 10 continued: Measuring Execution Time, and the Calendar Module

Syllabus topic Module 1, practical 10(b), "Write a program to measure program execution time", and 10(c), "Write a program using the calendar module to print the weekday of the first day of a given month and year"

Aim

To measure the execution time of a program, and to print the weekday of the first day of a given month and year using the calendar module.

Part one: measuring execution time

The three clocks, and which one to use

import time

print("time.time()          ", type(time.time()).__name__,
      "seconds since 1 January 1970")
print("time.perf_counter()  ", type(time.perf_counter()).__name__,
      "a counter for measuring, with no fixed zero")
print("time.process_time()  ", type(time.process_time()).__name__,
      "CPU time used by this process only")
print()
print("perf_counter resolution:", f'{time.get_clock_info("perf_counter").resolution:.0e}')
print("time resolution        :", f'{time.get_clock_info("time").resolution:.0e}')
print("is perf_counter monotonic?", time.get_clock_info("perf_counter").monotonic)
print("is time monotonic?        ", time.get_clock_info("time").monotonic)
time.time()           float seconds since 1 January 1970
time.perf_counter()   float a counter for measuring, with no fixed zero
time.process_time()   float CPU time used by this process only

perf_counter resolution: 4e-08
time resolution        : 1e-06
is perf_counter monotonic? True
is time monotonic?         False
ClockMeasuresUse it for
time.time()the wall clock, seconds since 1970what time is it
time.perf_counter()elapsed time, at the best resolution availablemeasuring how long something took
time.process_time()CPU time used, sleeping not countedhow much work was done

The two resolutions above are this machine's and are printed in exponent form so they fit; yours may differ, and the smaller number is the finer clock.

perf_counter is the right clock for this exercise, and the last two lines of that output are why. It is monotonic, which means it never goes backwards. time.time() is not: it follows the system clock, so if the clock is corrected, or a time zone changes, or the machine synchronises with a time server in the middle of your measurement, the difference you compute can be wrong and can even be negative.

process_time answers a different question, and the difference shows when a program waits:

import time

start_perf = time.perf_counter()
start_cpu = time.process_time()

time.sleep(0.2)

print(f"perf_counter saw   {time.perf_counter() - start_perf:.2f} seconds pass")
print(f"process_time saw   {time.process_time() - start_cpu:.2f} seconds of CPU used")
print("because sleeping uses no CPU at all")
perf_counter saw   0.21 seconds pass
process_time saw   0.00 seconds of CPU used
because sleeping uses no CPU at all

Timing a whole program

import time

start = time.perf_counter()

total = 0
for n in range(1_000_000):
    total += n

elapsed = time.perf_counter() - start

print(f"the sum is {total}")
print(f"it took {elapsed:.4f} seconds on this machine")
print(f"which is {elapsed * 1000:.1f} milliseconds")
the sum is 499999500000
it took 0.0759 seconds on this machine
which is 75.9 milliseconds

The sum itself is exact and the same everywhere: the sum of 0 to 999999. The time is not, and the .4f is deliberate, because printing fifteen digits of a measurement implies a precision the measurement does not have.

1_000_000 with underscores is the same number as 1000000. Python ignores the underscores and they make a long number readable, which is worth knowing.

munotes.in119

Practical 10 continued: Measuring Execution Time, and the Calendar Module

The wrong way to compare two pieces of code

import time


def with_loop(n):
    total = 0
    for i in range(n):
        total += i
    return total


def with_sum(n):
    return sum(range(n))


for label, function in [("explicit loop", with_loop), ("built in sum", with_sum)]:
    start = time.perf_counter()
    answer = function(100_000)
    elapsed = time.perf_counter() - start
    print(f"{label:<14} answer {answer} in {elapsed * 1000:.3f} ms")
explicit loop  answer 4999950000 in 3.368 ms
built in sum   answer 4999950000 in 0.856 ms

Both answers are right and the two times are a single measurement each. That is not enough to conclude anything, and saying so is the difference between a measurement and a guess. The next section is the fix.

timeit, which is the right tool

import timeit


def with_loop(n):
    total = 0
    for i in range(n):
        total += i
    return total


def with_sum(n):
    return sum(range(n))


runs = 200
loop_total = timeit.timeit(lambda: with_loop(100_000), number=runs)
sum_total = timeit.timeit(lambda: with_sum(100_000), number=runs)

print(f"over {runs} runs")
print(f"  explicit loop {loop_total / runs * 1000:.3f} ms per run")
print(f"  built in sum  {sum_total / runs * 1000:.3f} ms per run")
print(f"  sum was about {loop_total / sum_total:.1f} times faster on this machine")
over 200 runs
  explicit loop 3.458 ms per run
  built in sum  0.973 ms per run
  sum was about 3.6 times faster on this machine

timeit runs the code many times and gives the total, so dividing by the count gives a per run figure that is far steadier than one reading. It also turns off the garbage collector during the run, which removes another source of noise.

The ratio is the finding, not the milliseconds. Your own machine will print different times and a similar ratio, and the reason the built in sum wins is that its loop runs in compiled code rather than in the interpreter.

The way to report a measurement honestly

import statistics
import timeit


def with_loop(n):
    total = 0
    for i in range(n):
        total += i
    return total


samples = [timeit.timeit(lambda: with_loop(50_000), number=20) / 20 * 1000
           for _ in range(7)]

print(f"seven samples, milliseconds each:")
print("  ", [round(s, 3) for s in samples])
print(f"  best    {min(samples):.3f} ms")
print(f"  median  {statistics.median(samples):.3f} ms")
print(f"  worst   {max(samples):.3f} ms")
print(f"  spread  {max(samples) - min(samples):.3f} ms")
seven samples, milliseconds each:
   [1.475, 1.554, 1.556, 1.492, 1.468, 1.439, 1.461]
  best    1.439 ms
  median  1.475 ms
  worst   1.556 ms
  spread  0.117 ms

Report the best or the median, never the mean, and never a single reading. The worst sample usually means something else on the machine took the processor for a moment, which says nothing about your program. Quoting the spread as well shows you know how much to trust the figure.

munotes.in120

Practical 10 continued: Measuring Execution Time, and the Calendar Module

A decorator, which is the tidy way to time a function

import time
from functools import wraps


def timed(function):
    """Print how long the function took, every time it is called."""

    @wraps(function)
    def wrapper(*args, **kwargs):
        start = time.perf_counter()
        result = function(*args, **kwargs)
        elapsed = time.perf_counter() - start
        print(f"  {function.__name__}{args} took {elapsed * 1000:.3f} ms")
        return result

    return wrapper


@timed
def factorial(n):
    result = 1
    for i in range(2, n + 1):
        result *= i
    return result


digits = len(str(factorial(1000)))
print("1000! has", digits, "digits")
  factorial(1000,) took 0.273 ms
1000! has 2568 digits

@timed above def factorial means "pass this function through timed and use what comes back". The wrapper times the call and passes the answer on, so nothing else in the program changes. @wraps copies the original name and docstring onto the wrapper, so function.__name__ still says factorial.

The digit count is exact and the same on every machine; the time is not.

Part two: the calendar module

MU's question

import calendar

year = 2026
month = 10

first_weekday, days_in_month = calendar.monthrange(year, month)

print(f"{calendar.month_name[month]} {year}")
print(f"  monthrange gives two values, unpacked above")
print(f"  the first day is weekday number {int(first_weekday)}")
print(f"  which is {calendar.day_name[first_weekday]}")
print(f"  and the month has {days_in_month} days")
October 2026
  monthrange gives two values, unpacked above
  the first day is weekday number 3
  which is Thursday
  and the month has 31 days

calendar.monthrange(year, month) returns a tuple of two things, which is [Practical 8: the Tuple Return, Area and Circumference] again: the weekday of the first day and the number of days in the month. That one call answers MU's row.

The int() around the weekday is not decoration. From Python 3.12 the weekday comes back as a member of an enumeration called calendar.Day, so printing the raw tuple shows (calendar.THURSDAY, 31) on a new Python and (3, 31) on an older one. It behaves as the number 3 in every way that matters, including indexing day_name, and int() makes the printed output the same on every version.

Monday is 0 and Sunday is 6 in the calendar module, and calendar.day_name is the list that turns the number into a name.

There is also a direct call:

import calendar

print("weekday(2026, 10, 5) =", calendar.weekday(2026, 10, 5),
      "which is", calendar.day_name[calendar.weekday(2026, 10, 5)])
print("weekday(2026, 10, 1) =", calendar.weekday(2026, 10, 1),
      "which is", calendar.day_name[calendar.weekday(2026, 10, 1)])
print()
print("the seven day names:", list(calendar.day_name))
print("abbreviated        :", list(calendar.day_abbr))
print("the month names    :", list(calendar.month_name)[1:])
weekday(2026, 10, 5) = 0 which is Monday
weekday(2026, 10, 1) = 3 which is Thursday

the seven day names: ['Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday', 'Sunday']
abbreviated        : ['Mon', 'Tue', 'Wed', 'Thu', 'Fri', 'Sat', 'Sun']
the month names    : ['January', 'February', 'March', 'April', 'May', 'June', 'July', 'August', 'September', 'October', 'November', 'December']
munotes.in121

Practical 10 continued: Measuring Execution Time, and the Calendar Module

calendar.month_name has an empty string at index 0, so that month_name[1] is January and the index matches the month number. That is why the last line above slices from 1. It is a small thing and it produces a blank first entry in anybody's output who forgets it.

The program

import calendar


def first_day_of(year, month):
    """The name of the weekday on which the given month starts."""
    if not 1 <= month <= 12:
        raise ValueError("a month is 1 to 12")
    weekday_number, _ = calendar.monthrange(year, month)
    return calendar.day_name[weekday_number]


year = int(input("Enter the year  (for example 2026): "))
month = int(input("Enter the month (1 to 12)        : "))

print(f"{calendar.month_name[month]} {year} begins on a {first_day_of(year, month)}")
print(f"and has {calendar.monthrange(year, month)[1]} days")
2026
10
Enter the year  (for example 2026): Enter the month (1 to 12)        : October 2026 begins on a Thursday
and has 31 days

Printing the month, which is what makes it a journal entry

import calendar

print(calendar.month(2026, 10))
    October 2026
Mo Tu We Th Fr Sa Su
          1  2  3  4
 5  6  7  8  9 10 11
12 13 14 15 16 17 18
19 20 21 22 23 24 25
26 27 28 29 30 31

calendar.month(year, month) returns the whole month as a string, laid out as a calendar. Read the first row against the program above: the 1st sits under the column the program named.

import calendar

print("as a list of weeks, with 0 for a day in another month:")
for week in calendar.monthcalendar(2026, 10):
    print("  ", week)

print()
print("only the real days of the first week:",
      [day for day in calendar.monthcalendar(2026, 10)[0] if day != 0])
as a list of weeks, with 0 for a day in another month:
   [0, 0, 0, 1, 2, 3, 4]
   [5, 6, 7, 8, 9, 10, 11]
   [12, 13, 14, 15, 16, 17, 18]
   [19, 20, 21, 22, 23, 24, 25]
   [26, 27, 28, 29, 30, 31, 0]

only the real days of the first week: [1, 2, 3, 4]

monthcalendar gives the month as a list of weeks, each a list of seven day numbers. A day belonging to another month is printed as 0. That is the shape to use when a program needs to work with the weeks rather than print them.

The rest of the module, briefly

import calendar

print("is 2024 a leap year?     ", calendar.isleap(2024))
print("is 1900 a leap year?     ", calendar.isleap(1900))
print("leap years 2000 to 2030  :", calendar.leapdays(2000, 2030))
print("days in February 2024    :", calendar.monthrange(2024, 2)[1])
print("days in February 2023    :", calendar.monthrange(2023, 2)[1])
print("the first day of the week:", calendar.firstweekday(), "which is Monday")
print()
print("every month of 2026 and the day it starts on:")
for month in range(1, 13):
    starts, days = calendar.monthrange(2026, month)
    print(f"  {calendar.month_name[month]:<10} starts {calendar.day_name[starts]:<10} "
          f"{days} days")
munotes.in122

Practical 10 continued: Measuring Execution Time, and the Calendar Module

is 2024 a leap year?      True
is 1900 a leap year?      False
leap years 2000 to 2030  : 8
days in February 2024    : 29
days in February 2023    : 28
the first day of the week: 0 which is Monday

every month of 2026 and the day it starts on:
  January    starts Thursday   31 days
  February   starts Sunday     28 days
  March      starts Sunday     31 days
  April      starts Wednesday  30 days
  May        starts Friday     31 days
  June       starts Monday     30 days
  July       starts Wednesday  31 days
  August     starts Saturday   31 days
  September  starts Tuesday    30 days
  October    starts Thursday   31 days
  November   starts Sunday     30 days
  December   starts Tuesday    31 days

calendar.leapdays(a, b) counts the leap years in a range, with the start included and the end excluded, which is worth noticing before you quote the number.

The February rows are the leap year rule showing up as a number of days, which is a neater demonstration than isleap on its own.

Procedure

  1. Save as practical10b.py. Print the resolution and the monotonic flag of perf_counter and of

time, and say which you will use and why.

  1. Time a loop that sums a million numbers with perf_counter, printing the elapsed time to four

decimal places.

  1. Time two ways of doing the same work once each, and say why one reading each proves nothing.
  2. Use timeit with number=200 and report the per run figure and the ratio.
  3. Take seven samples and report the best, the median, the worst and the spread.
  4. Save as practical10c.py. import calendar.
  5. Read a year and a month with int(input()).
  6. Call calendar.monthrange(year, month), unpack the two values, and print

calendar.day_name[first].

  1. Print the month with calendar.month(year, month) and check that the 1st is in the column you

named.

  1. Print every month of the year with the day it starts on, and the days in February for a leap

year and a common year.

Result

perf_counter reported itself monotonic and time did not, which is why the measurements use perf_counter. Sleeping for two tenths of a second was seen by perf_counter and not by process_time, confirming that sleeping uses no CPU. The built in sum was faster than the explicit loop over 200 runs of timeit, and the ratio rather than the milliseconds is the finding. Seven samples of the same work differed, which is why the median and the spread are reported and a single reading is not. On the calendar half, monthrange returned the weekday of the first day and the number of days, and the printed calendar put the 1st in the column the program named. monthrange(2024, 2)[1] was 29 and monthrange(2023, 2)[1] was 28.

munotes.in123

Practical 10 continued: Measuring Execution Time, and the Calendar Module

Where marks are lost

  • Using time.time() to measure. It is not monotonic; perf_counter is.
  • One reading. A single measurement proves nothing; use timeit or take several.
  • Reporting the mean of several samples. Use the best or the median: a slow sample is usually

another program on the machine.

  • Printing fifteen digits of a measurement, which claims a precision the clock does not have.
  • Quoting a time from a book as if it were yours. Say "on this machine".
  • Forgetting that Monday is 0 in the calendar module.
  • Forgetting the blank at index 0 of month_name, which prints an empty first entry.
  • Using monthrange(year, month)[0] as the number of days. It is the weekday; the days are

[1].

  • Reading the month as a name and not converting it. int(input()) gives a number.

For the journal

Two entries under practical 10. For the timing: the aim, the three clock table with one sentence on monotonic, the timed loop with the elapsed time to four decimal places, and then the timeit comparison with the ratio. The seven samples with the best, median, worst and spread. Write "on this machine" beside every time you record, and one sentence saying that the ratio is the result and the milliseconds are the machine's.

For the calendar: the aim in MU's words, the monthrange call with its tuple printed, the weekday name, and calendar.month(year, month) printed underneath so the answer can be checked by eye against the calendar. One sentence: Monday is 0 in the calendar module, and monthrange returns the weekday of the first day and the number of days as a tuple.

Quick revision

  • Three clocks: time.time() the wall clock, time.perf_counter() for measuring,

time.process_time() for CPU used.

  • perf_counter is monotonic; time.time() is not. A clock correction can make a time.time()

difference wrong or negative.

  • Sleeping is seen by perf_counter and not by process_time.
  • Pattern: start = perf_counter(), do the work, elapsed = perf_counter() - start.
  • 1_000_000 is 1000000. The underscores are ignored.
  • timeit.timeit(callable, number=n) gives the total for n runs; divide by n. It also disables

the garbage collector.

  • Report the best or the median of several samples, plus the spread. Never one reading, never

the mean.

  • The ratio is what transfers to another machine; the milliseconds do not.
  • A timing decorator wraps a function; @wraps keeps its name and docstring.
  • calendar.monthrange(year, month) returns (weekday of the 1st, number of days).
  • Monday is 0 in calendar. calendar.day_name[n] is the name.
  • calendar.month_name has a blank at index 0, so index 1 is January.
  • calendar.month(y, m) prints the month; calendar.monthcalendar(y, m) gives weeks as lists with

0 for a day in another month.

munotes.in124

Practical 10 continued: Measuring Execution Time, and the Calendar Module

  • calendar.isleap, calendar.leapdays(a, b) with the end excluded.

Questions you should be able to answer

1. Which clock should you measure with, and why? time.perf_counter(), because it has the best resolution available and it is monotonic, so it never goes backwards.

2. What can go wrong with time.time() for a measurement? It follows the system clock, so a correction or a time server synchronisation during your measurement makes the difference wrong, and it can even come out negative.

3. What is the difference between perf_counter and process_time? perf_counter measures time passing; process_time measures CPU used by this process. A program that sleeps advances the first and not the second.

4. Why is timeit better than one perf_counter reading? It runs the code many times and disables the garbage collector, so the per run figure is far steadier than a single measurement.

5. Of several samples, which one should you report? The best or the median, with the spread. A slow sample usually means something else on the machine took the processor.

6. Why report a ratio rather than milliseconds? Because the milliseconds are your machine's and will not reproduce, while the ratio between two ways of doing the same work usually will.

7. What does calendar.monthrange(2026, 10) return? A tuple: the weekday number of the 1st and the number of days in the month.

8. What number is Monday in the calendar module? 0. Sunday is 6.

9. Write the two lines that print the day a month starts on. first, days = calendar.monthrange(year, month) then print(calendar.day_name[first]).

10. Why does calendar.month_name have an empty string at index 0? So that the index matches the month number, with January at 1.

11. How does monthcalendar show a day that belongs to the previous or next month? As a 0 in the week's list.

Contents This chapter on its own page

munotes.in125

Module II

1. Array Operations: Write a program to implement basic array operations:

munotes.in

Chapter Twenty-One

Python for Data Structures, and the Cost of an Operation

Syllabus topic Module 2's own subject, and the five things MU's Course Objectives and Outcomes name that no practical row names: CO 6 "arrays, linked lists, stacks, queues, trees, and graphs", CO 8 "choose appropriate data structures for different applications and justify their choices", CO 9 "dynamic memory allocation and efficient data management techniques", CO 10 "debug and optimize code for data structure operations", OC 8 "analyze the time and space complexity of algorithms for various data structures"

Aim

To learn the Python a data structure is built from, and to be able to say what an operation costs and which structure to choose.

Why Module 2 cannot start at exercise one

Every exercise in Module 2 builds a structure out of nodes and references. That needs four things this book has not needed yet: a class, self, a reference, and None. Twenty minutes here saves the whole module.

A class, in one page

A class is a template. An object is one thing made from it. A node in a linked list is an object, and so is the list itself.

class Student:
    """One student: a name and a mark."""

    def __init__(self, name, mark):
        self.name = name
        self.mark = mark

    def passed(self):
        return self.mark >= 40

    def __repr__(self):
        return f"Student({self.name!r}, {self.mark})"


aarti = Student("Aarti", 78)
divya = Student("Divya", 32)

print("one object      :", aarti)
print("its fields      :", aarti.name, aarti.mark)
print("a method on it  :", aarti.passed(), divya.passed())
print("two objects     :", [aarti, divya])
print("the class       :", Student.__name__)
print("its docstring   :", Student.__doc__)
one object      : Student('Aarti', 78)
its fields      : Aarti 78
a method on it  : True False
two objects     : [Student('Aarti', 78), Student('Divya', 32)]
the class       : Student
its docstring   : One student: a name and a mark.
PartWhat it is
class Student:declares the template
__init__the constructor, run when an object is made
selfthe object the method was called on
self.name = namean attribute, stored in the object
def passed(self):a method, a function belonging to the class
__repr__what print and a list show for the object

self is the first parameter of every method and it is not passed by the caller. aarti.passed() calls passed(aarti). Forgetting self in the definition is the commonest error a beginner makes with a class:

class Broken:
    def greet():
        return "hello"


Broken().greet()
TypeError: Broken.greet() takes 0 positional arguments but 1 was given

Read the message: it says one argument was given and none expected, which is self arriving without a parameter to land in.

__repr__ is not optional in this module. Without it, printing a node gives <__main__.Node object at 0x102f3d9a0>, and a list of ten nodes is ten such lines. Every structure in Module 2 defines one, so that a program can print itself and the journal entry can show what happened.

A reference, and None

This is the idea the whole module rests on. A variable does not hold an object; it holds a reference to one. Two variables can therefore refer to the same object, which is what [Practical 3 continued: Mathematical Functions, Aliasing and Copying] proved for arrays, and a reference can point at nothing, which is None.

munotes.in126

Python for Data Structures, and the Cost of an Operation

class Node:
    def __init__(self, data):
        self.data = data
        self.nxt = None

    def __repr__(self):
        return f"Node({self.data!r})"


first = Node("a")
second = Node("b")
first.nxt = second

print("first          :", first)
print("first.nxt      :", first.nxt)
print("second.nxt     :", second.nxt, "which is None, so this is the end")
print("the same object?", first.nxt is second)
print("how many nodes exist? two. How many names? three:",
      [first, second, first.nxt])
print("is second.nxt None?", second.nxt is None)
first          : Node('a')
first.nxt      : Node('b')
second.nxt     : None which is None, so this is the end
the same object? True
how many nodes exist? two. How many names? three: [Node('a'), Node('b'), Node('b')]
is second.nxt None? True
first -> [ a | * ] -> [ b | None ]

That picture is the whole of a linked list, and it is worth drawing in the journal for every structure that has nodes in it. The * is a reference and the None is the end.

Test for None with is, not with ==. is asks "are these the same object", and there is only one None in a running Python program. == can be redefined by a class and then the test means something else.

class Odd:
    def __eq__(self, other):
        return True


strange = Odd()
print("strange == None :", strange == None, "  which is a lie the class told")
print("strange is None :", strange is None, "  which is the truth")
strange == None : True   which is a lie the class told
strange is None : False   which is the truth

Raising an error rather than returning a sentinel

Every structure in this module has an operation that can fail: popping an empty stack, dequeuing an empty queue, deleting a node that is not there. There are two ways to report it and only one of them is right.

class BadStack:
    """Returns None on an empty pop, which cannot be told from a stored None."""

    def __init__(self):
        self.items = []

    def push(self, item):
        self.items.append(item)

    def pop(self):
        if not self.items:
            return None
        return self.items.pop()


class GoodStack:
    """Raises on an empty pop, which cannot be mistaken for anything."""

    def __init__(self):
        self.items = []

    def push(self, item):
        self.items.append(item)

    def pop(self):
        if not self.items:
            raise IndexError("pop from an empty stack")
        return self.items.pop()


bad = BadStack()
bad.push(None)
print("BadStack: pushed None, popped", bad.pop(), "and popping empty gives", bad.pop())
print("so the caller cannot tell a stored None from an empty stack")

good = GoodStack()
try:
    good.pop()
except IndexError as error:
    print("GoodStack: popping empty raised IndexError:", error)
BadStack: pushed None, popped None and popping empty gives None
so the caller cannot tell a stored None from an empty stack
GoodStack: popping empty raised IndexError: pop from an empty stack

Raise. The BadStack cannot distinguish "the item was None" from "there was no item", and the output above shows both printing the same thing. This is the rule for every structure in the rest of the module, and an examiner who asks "what does your pop do when the stack is empty" is asking exactly this.

munotes.in127

Python for Data Structures, and the Cost of an Operation

MU's CO 9: dynamic memory allocation

Her Course Objective 9 is "dynamic memory allocation and efficient data management techniques". In C that is malloc and free, and [If Your College Runs Module 2 in C] shows them. In Python it is automatic, and you should still be able to say what happens.

Every object is allocated when it is created and freed when nothing refers to it any more. Python keeps a count of how many references each object has; when the count reaches zero the memory is released.

import sys


class Node:
    def __init__(self, data):
        self.data = data
        self.nxt = None


node = Node("a")
print("references to this node:", sys.getrefcount(node) - 1,
      "(getrefcount counts its own argument, so subtract one)")

alias = node
print("after alias = node    :", sys.getrefcount(node) - 1)

holder = [node, node]
print("after putting it in a list twice:", sys.getrefcount(node) - 1)

del alias
del holder
print("after deleting both   :", sys.getrefcount(node) - 1)
references to this node: 1 (getrefcount counts its own argument, so subtract one)
after alias = node    : 2
after putting it in a list twice: 4
after deleting both   : 1

That is why unlinking a node from a list is enough to free it: after before.nxt = here.nxt nothing refers to here, its count reaches zero and the memory goes back. There is no free to call and no free to forget.

import gc


class Watched:
    """Prints when it is collected, so allocation can be seen."""

    def __init__(self, name):
        self.name = name

    def __del__(self):
        print(f"  {self.name} was freed")


print("making three nodes")
a = Watched("a")
b = Watched("b")
c = Watched("c")

print("unlinking b by deleting the only name for it")
del b

print("unlinking the other two")
del a
del c
gc.collect()
print("done")
making three nodes
unlinking b by deleting the only name for it
  b was freed
unlinking the other two
  a was freed
  c was freed
done

Read the order. b was freed at the exact moment the last name for it went away, not at the end of the program. That is reference counting, and it is the answer to "when is the memory released".

MU's OC 8: time and space complexity

Her Outcome 8 is "analyze the time and space complexity of algorithms for various data structures". It is not a practical row and it is examinable, and it is the vocabulary the whole rest of this module uses.

munotes.in128

Python for Data Structures, and the Cost of an Operation

Big O, in plain words

Big O describes how the work grows as the data grows. It is not a time in seconds and it does not say which of two programs is faster on ten items. It says what happens when there are ten thousand.

WrittenCalledMeansDoubling n makes the work
O(1)constantthe same however big n isthe same
O(log n)logarithmicone more step each time n doublesone step more
O(n)linearproportional to ntwice as much
O(n log n)linearithmica good sorta little over twice
O(n squared)quadraticevery item against every itemfour times as much
O(2 to the n)exponentialunusable past about 40squared

Three rules for reading it:

  • Constants are dropped. 3n steps and n steps are both O(n), because doubling n doubles both.
  • Only the largest term counts. n squared plus n is O(n squared).
  • It is the worst case unless something else is said.

Counted, not asserted

def linear_search(items, target):
    """Returns (index, comparisons)."""
    comparisons = 0
    for index, item in enumerate(items):
        comparisons += 1
        if item == target:
            return index, comparisons
    return -1, comparisons


def binary_search(items, target):
    """Returns (index, comparisons). The list must be sorted."""
    low, high, comparisons = 0, len(items) - 1, 0
    while low <= high:
        middle = (low + high) // 2
        comparisons += 1
        if items[middle] == target:
            return middle, comparisons
        if items[middle] < target:
            low = middle + 1
        else:
            high = middle - 1
    return -1, comparisons


print(f"{'n':>8} {'linear worst':>13} {'binary worst':>13}")
for n in [10, 100, 1000, 10000, 100000]:
    items = list(range(n))
    _, linear = linear_search(items, -1)
    _, binary = binary_search(items, -1)
    print(f"{n:>8} {linear:>13} {binary:>13}")
       n  linear worst  binary worst
      10            10             3
     100           100             6
    1000          1000             9
   10000         10000            13
  100000        100000            16

Read the two columns. The linear column is the value of n, exactly. The binary column grows by about three every time n is multiplied by ten, which is what a logarithm does. Those are the two shapes the whole module is about, and the counts are the same on every machine.

And what O(n squared) looks like

def bubble_pass_count(items):
    """Comparisons made by a full bubble sort."""
    data = list(items)
    comparisons = 0
    for i in range(len(data)):
        for j in range(len(data) - 1 - i):
            comparisons += 1
            if data[j] > data[j + 1]:
                data[j], data[j + 1] = data[j + 1], data[j]
    return comparisons


print(f"{'n':>6} {'comparisons':>12} {'n(n-1)/2':>10} {'growth':>8}")
previous = None
for n in [10, 20, 40, 80, 160]:
    count = bubble_pass_count(list(range(n, 0, -1)))
    formula = n * (n - 1) // 2
    growth = "" if previous is None else f"{count / previous:.2f}x"
    print(f"{n:>6} {count:>12} {formula:>10} {growth:>8}")
    previous = count
munotes.in129

Python for Data Structures, and the Cost of an Operation

     n  comparisons   n(n-1)/2   growth
    10           45         45
    20          190        190    4.22x
    40          780        780    4.11x
    80         3160       3160    4.05x
   160        12720      12720    4.03x

Two things in that output are the whole of Big O made concrete. The comparison count equals n(n-1)/2 exactly, and doubling n multiplies the work by very nearly four. That is what quadratic means, and it is why a quadratic sort is unusable on a large list.

Space complexity

The same idea for memory. An algorithm that needs a fixed number of extra variables is O(1) extra space, and one that builds a second copy of the data is O(n).

import sys


def reverse_in_place(items):
    """O(1) extra space: two indexes and a swap."""
    left, right = 0, len(items) - 1
    while left < right:
        items[left], items[right] = items[right], items[left]
        left += 1
        right -= 1
    return items


def reverse_with_a_copy(items):
    """O(n) extra space: a whole new list."""
    return items[::-1]


first = [1, 2, 3, 4, 5]
second = [1, 2, 3, 4, 5]

print("in place  :", reverse_in_place(first), "and the original is now", first)
print("with copy :", reverse_with_a_copy(second), "and the original is still", second)
print()
print("the size of a list of 1000 ints, in bytes:",
      sys.getsizeof(list(range(1000))))
print("the size of a list of 2000 ints, in bytes:",
      sys.getsizeof(list(range(2000))))
print("roughly double, which is what O(n) space looks like")
in place  : [5, 4, 3, 2, 1] and the original is now [5, 4, 3, 2, 1]
with copy : [5, 4, 3, 2, 1] and the original is still [1, 2, 3, 4, 5]

the size of a list of 1000 ints, in bytes: 8056
the size of a list of 2000 ints, in bytes: 16056
roughly double, which is what O(n) space looks like

sys.getsizeof measures the container itself, not the objects it refers to, which is a caution worth one line in the journal: the list of a thousand integers also has a thousand integer objects behind it.

MU's CO 8: which structure to choose, and how to justify it

Her Objective 8 is "choose appropriate data structures for different applications and justify their choices". This table is the answer, and it is the one page of this book to learn by heart.

The table is in two halves so that it still reads on a phone. First the two structures that hold items in a row:

OperationList or arrayLinked list
Read item nO(1)O(n)
Find a valueO(n)O(n)
Insert at the frontO(n)O(1)
Insert at the backO(1) amortisedO(1) with a tail
Insert in the middleO(n)O(1) at a held node
Delete at the frontO(n)O(1)
In sorted orderO(n log n) to sortO(n log n)
Smallest or largestO(n)O(n)
Extra memorynonea link per item
munotes.in130

Python for Data Structures, and the Cost of an Operation

Then the four that restrict what you may ask of them, and are fast in return. A dash means the structure does not offer that operation at all.

OperationStackQueueBST (balanced)Hash table
Read item nnonoO(log n) by keyO(1) by key
Find a valuenonoO(log n)O(1)
Insert at the frontO(1) pushnoO(log n)O(1)
Insert at the backnoO(1)O(log n)O(1)
Insert in the middlenonoO(log n)O(1)
Delete at the frontO(1) popO(1)O(log n)O(1)
In sorted ordernonoO(n), freeO(n log n)
Smallest or largestnonoO(log n)O(n)
Extra memorynonenonetwo links per itemspare buckets

And the justification in words, which is what an examiner actually asks for:

ChooseWhenBecause
array or listyou index by position, and the size is stableone step to any item
linked listyou insert and delete a lot, at ends or at held positionsno shifting, and it grows freely
stackthe last thing in must come out firstundo, brackets, a recursion made explicit
queuethe first thing in must come out firstwaiting lines, scheduling, level order
BSTyou need order AND fast lookupin-order traversal is sorted, for free
hash tableyou look up by a key and do not care about orderone step, on average
heapyou repeatedly need the smallest or largestthe top is O(1), removing it O(log n)
graphthe relation is between any two thingsnothing else can hold it

Never answer "which structure" with a name alone. Answer with the operation that dominates: "a hash table, because the program looks a student up by roll number far more often than it lists them in order." That sentence pattern is worth a mark on Q2 and in the viva.

MU's CO 6: graphs, which no practical row names

Her Objective 6 lists "arrays, linked lists, stacks, queues, trees, and graphs". Her ten exercises never mention a graph, so here is one, complete, and [Practical 10: the Combined Application] uses it.

A graph is a set of vertices and a set of edges joining pairs of them. A tree is a special graph with no cycle; a general graph may have cycles, which is the whole difficulty.

class Graph:
    """An undirected graph as an adjacency list: each vertex to its neighbours."""

    def __init__(self):
        self.neighbours = {}

    def add_vertex(self, vertex):
        self.neighbours.setdefault(vertex, [])

    def add_edge(self, a, b):
        self.add_vertex(a)
        self.add_vertex(b)
        if b not in self.neighbours[a]:
            self.neighbours[a].append(b)
        if a not in self.neighbours[b]:
            self.neighbours[b].append(a)

    def breadth_first(self, start):
        """Vertices in order of distance from start."""
        seen = {start}
        order = []
        waiting = [start]
        while waiting:
            here = waiting.pop(0)
            order.append(here)
            for neighbour in self.neighbours[here]:
                if neighbour not in seen:
                    seen.add(neighbour)
                    waiting.append(neighbour)
        return order

    def depth_first(self, start, seen=None):
        """Vertices by going as deep as possible first."""
        if seen is None:
            seen = set()
        seen.add(start)
        order = [start]
        for neighbour in self.neighbours[start]:
            if neighbour not in seen:
                order.extend(self.depth_first(neighbour, seen))
        return order

    def __repr__(self):
        return "\n".join(f"  {v}: {sorted(n)}"
                         for v, n in sorted(self.neighbours.items()))


stations = Graph()
for a, b in [("Churchgate", "Marine Lines"), ("Marine Lines", "Charni Road"),
             ("Charni Road", "Grant Road"), ("Grant Road", "Mumbai Central"),
             ("Mumbai Central", "Mahalaxmi"), ("Grant Road", "Byculla"),
             ("Byculla", "Dadar"), ("Mahalaxmi", "Dadar")]:
    stations.add_edge(a, b)

print("the graph:")
print(stations)
print()
print("breadth first from Churchgate:")
for name in stations.breadth_first("Churchgate"):
    print("  ", name)
print()
print("depth first from Churchgate  :", len(stations.depth_first("Churchgate")),
      "vertices reached")
print("vertices:", len(stations.neighbours),
      " edges:", sum(len(n) for n in stations.neighbours.values()) // 2)
munotes.in131

Python for Data Structures, and the Cost of an Operation

the graph:
  Byculla: ['Dadar', 'Grant Road']
  Charni Road: ['Grant Road', 'Marine Lines']
  Churchgate: ['Marine Lines']
  Dadar: ['Byculla', 'Mahalaxmi']
  Grant Road: ['Byculla', 'Charni Road', 'Mumbai Central']
  Mahalaxmi: ['Dadar', 'Mumbai Central']
  Marine Lines: ['Charni Road', 'Churchgate']
  Mumbai Central: ['Grant Road', 'Mahalaxmi']

breadth first from Churchgate:
   Churchgate
   Marine Lines
   Charni Road
   Grant Road
   Mumbai Central
   Byculla
   Mahalaxmi
   Dadar

depth first from Churchgate  : 8 vertices reached
vertices: 8  edges: 8

seen is what makes it work. A graph may have a cycle, and this one does: Grant Road reaches Dadar both through Byculla and through Mumbai Central and Mahalaxmi. Without the seen set both traversals would go round for ever. A tree needs no such set, and that is the practical difference between the two.

Breadth first uses a queue and depth first uses a stack, which is why graphs belong in this module at all. waiting.pop(0) above is a queue, taking from the front; change it to waiting.pop() and it becomes a stack and the traversal becomes depth first. The two structures of practicals 3 and 4 are the two traversals of a graph.

waiting.pop(0) on a Python list is O(n), because every remaining item shifts down. On a real graph use collections.deque and popleft(), which is O(1). [Practical 4: a Queue over an Array] is about exactly that cost.

MU's CO 10: debugging a structure

Her Objective 10 is "debug and optimize code for data structure operations". The single most useful technique, and the one every chapter of this module uses, is to make the structure print itself at every step.

class TracedStack:
    """A stack that shows itself after every operation."""

    def __init__(self):
        self.items = []

    def push(self, item):
        self.items.append(item)
        print(f"  push {item!r:<6} -> {self.items}")

    def pop(self):
        if not self.items:
            raise IndexError("pop from an empty stack")
        item = self.items.pop()
        print(f"  pop  {item!r:<6} -> {self.items}")
        return item


stack = TracedStack()
for character in "a(b)":
    stack.push(character)
stack.pop()
stack.pop()
print("what is left:", stack.items)
munotes.in132

Python for Data Structures, and the Cost of an Operation

  push 'a'    -> ['a']
  push '('    -> ['a', '(']
  push 'b'    -> ['a', '(', 'b']
  push ')'    -> ['a', '(', 'b', ')']
  pop  ')'    -> ['a', '(', 'b']
  pop  'b'    -> ['a', '(']
what is left: ['a', '(']

That trace is worth more in a journal than a correct answer with no working, and it is how you find a bug in a structure: print the structure, not the variable. Every chapter in Module 2 has a listing that does this, and every one of them is there to be copied into your own code when something goes wrong.

The timing harness the later chapters use

import timeit


def with_list_pop_front(n):
    items = list(range(n))
    while items:
        items.pop(0)


def with_list_pop_back(n):
    items = list(range(n))
    while items:
        items.pop()


n = 4000
runs = 20
front = timeit.timeit(lambda: with_list_pop_front(n), number=runs) / runs
back = timeit.timeit(lambda: with_list_pop_back(n), number=runs) / runs

print(f"emptying a list of {n} items")
print(f"  from the front  {front * 1000:.2f} ms")
print(f"  from the back   {back * 1000:.2f} ms")
print(f"  the front was about {front / back:.0f} times slower on this machine")
emptying a list of 4000 items
  from the front  1.12 ms
  from the back   0.14 ms
  the front was about 8 times slower on this machine

That is the same timeit as [Practical 10 continued: Measuring Execution Time, and the Calendar Module], and it is the last time this book quotes a millisecond. From here on the measurements are counts, because a count is the same on your machine as on this one, and because an examiner can check a count.

The finding itself is real and matters for practical 4: taking from the front of a Python list is O(n) and taking from the back is O(1), so a queue built on pop(0) is quadratic.

Procedure

  1. Write a small class with __init__, one method and __repr__, and make two objects from it.
  2. Leave self out of a method definition on purpose and record the TypeError.
  3. Build two nodes and link one to the other. Print the chain and draw the picture with the arrow

and the None.

  1. Show that x is None and x == None can disagree, using a class that redefines __eq__.
  2. Write two stacks, one returning None on an empty pop and one raising, and show that the first

cannot tell a stored None from an empty stack.

  1. Print sys.getrefcount as names are added and deleted, and add a __del__ to see when an object

is freed.

  1. Count the comparisons of linear and binary search for n from 10 to 100000 and compare the two
munotes.in133

Python for Data Structures, and the Cost of an Operation

columns.

  1. Count the comparisons of a bubble sort for n doubling, and check them against n(n-1)/2 and

against the factor of four.

  1. Build the graph, run breadth first and depth first, and say what seen is for.
  2. Write a traced stack that prints itself after every operation.

Result

A class with __init__, a method and __repr__ behaved as expected, and omitting self raised TypeError naming one argument given and none expected. Two nodes linked with first.nxt = second printed as a chain ending in None, and first.nxt is second was True. A class redefining __eq__ made x == None True while x is None stayed False. The stack returning None could not distinguish a stored None from an empty stack; the raising one could. sys.getrefcount rose as names were added and fell as they were deleted, and __del__ fired the moment the last name went away. Linear search made exactly n comparisons, and binary search made 3, 6, 9, 13 and 16 for n of 10 to 100000, growing by about three per tenfold increase. Bubble sort's comparisons equalled n(n-1)/2 exactly and multiplied by nearly four as n doubled. The graph traversals both reached every vertex and terminated, which the seen set is what makes true, since the graph has a cycle.

Where marks are lost

  • No __repr__, so every printed node is an address and the journal shows nothing.
  • Leaving self out of a method definition.
  • Returning None on an empty pop. Raise instead; a stored None is indistinguishable.
  • == None instead of is None.
  • Saying O(n) is "slow". It is a growth rate, not a speed. Say what happens when n doubles.
  • Quoting a millisecond from a book as if it were your machine's.
  • Naming a structure with no reason. Name the operation that dominates.
  • Forgetting the seen set in a graph traversal, which loops for ever on a cycle.
  • pop(0) on a list in a queue. It is O(n); use deque.popleft().

For the journal

This chapter has no numbered practical, so give it the page before your first Module 2 entry. Put three things on it, because every later entry refers back to them.

The node picture, with the arrow and the None, because every structure in this module is that picture repeated.

The Big O table, six rows, with the "doubling n" column, and underneath it the two counted columns from linear and binary search and the bubble sort's n(n-1)/2. Those three counts are the evidence for the table.

The "choose this when" table. It is the answer to the commonest viva question on this module and it takes one page.

Quick revision

  • A class is a template, an object is one thing made from it. __init__ is the constructor,
munotes.in134

Python for Data Structures, and the Cost of an Operation

self is the object, an attribute is stored in it, a method belongs to the class.

  • self is the first parameter of every method and the caller does not pass it.
  • __repr__ is compulsory in this module, or a node prints as an address.
  • A variable holds a reference. None is the reference to nothing. Test it with is.
  • Raise on an empty pop; never return None, which cannot be told from a stored None.
  • Python frees an object when the reference count reaches zero. No malloc, no free.
  • Big O is how the work grows, not a time. O(1), O(log n), O(n), O(n log n), O(n squared).
  • Doubling n: O(n) doubles the work, O(n squared) multiplies it by four, O(log n) adds one step.
  • Constants are dropped and only the largest term counts. It is the worst case unless stated.
  • Bubble sort makes exactly n(n-1)/2 comparisons.
  • Space complexity: O(1) extra is a few variables, O(n) extra is a second copy.
  • Choosing a structure: name the dominant operation, not the structure.
  • A graph is vertices and edges and may have a cycle, so a traversal needs a seen set.

Breadth first uses a queue, depth first uses a stack.

  • list.pop(0) is O(n); deque.popleft() is O(1).

Questions you should be able to answer

1. What is self? The object the method was called on. It is the first parameter of every method and Python passes it automatically.

2. Why does every class in this module need __repr__? So that printing the object shows what it holds instead of its memory address, which is what makes a journal entry readable and a bug findable.

3. What does a variable actually hold? A reference to an object. Two variables can refer to the same object, and a reference may be None, meaning nothing.

4. Why is None and not == None? Because a class can redefine == and then the test means whatever that class decided. is asks whether it is the one None object, which is the question.

5. What should pop do on an empty stack, and why? Raise an exception. Returning None cannot be distinguished from having stored a None, as the two stacks in this chapter show.

6. When does Python free the memory for a node? When the last reference to it goes away, because Python counts references and releases the object when the count reaches zero.

7. What does O(n squared) mean in one sentence? Doubling the amount of data multiplies the work by four.

8. What is the difference between O(n) and O(log n) in numbers? Counted here: searching 100000 items took 100000 comparisons linearly and 16 by binary search, and every tenfold increase in n adds only about three comparisons to the binary figure.

munotes.in135

Python for Data Structures, and the Cost of an Operation

9. How many comparisons does a bubble sort make on n items? n(n-1)/2, which the program in this chapter checks exactly for n from 10 to 160.

10. How do you answer "which data structure would you use"? By naming the operation the program does most and then the structure that makes it cheap, for example "a hash table, because we look up by roll number far more often than we list in order".

11. Why does a graph traversal need a seen set when a tree traversal does not? Because a graph may contain a cycle, so without it the traversal would revisit vertices for ever. A tree has no cycle.

12. Which structure does breadth first search use, and which does depth first? Breadth first uses a queue; depth first uses a stack, whether explicit or the call stack of the recursion.

Contents This chapter on its own page

munotes.in136

Chapter Twenty-Three

Practical 2: Building a Singly Linked List

Syllabus topic Module 2, practical 2, "Linked List Manipulation: Write a program to: Create a singly linked list. Insert a node at the beginning, end, and at a given position in a linked list."

Aim

To create a singly linked list, and to insert a node at the beginning, at the end and at a given position.

The picture, which is the whole idea

head -> [ 10 | * ] -> [ 20 | * ] -> [ 30 | None ]
                                        ^
                                        tail

A node holds one value and a reference to the next node. The last node's reference is None, which is how the end is recognised. The list itself is nothing but a reference to the first node, called the head.

Nothing in that picture is contiguous. The three nodes may be anywhere in memory and the arrows are all that hold them together. Every difference from an array follows from that one fact:

ArraySingly linked list
The items areside by side in one blockanywhere, joined by references
Item n is found byarithmetic, one stepwalking n references, O(n)
Inserting at the frontevery item shifts, O(n)one new node, O(1)
Inserting at the backO(1) if there is roomO(1) with a tail, O(n) without
Growinga new block and a copynothing to do
Memory for n itemsn slots, plus sparen values plus n references
Going backwardsyes, subtract 1no

The last row is the defect that a doubly linked list exists to cure, and it is why delete in the next chapter needs a trailing reference.

The three things the class keeps

  • head, the first node, or None when the list is empty.
  • tail, the last node. Optional, and worth having: without it, appending has to walk the whole

list.

  • count, how many nodes there are. Optional, and worth having, or len() has to walk.

Keeping a tail and a count means every operation must maintain them, and forgetting one is the commonest bug in this exercise. Inserting into an empty list must set the tail as well as the head.

The class

"""A singly linked list, built by hand for Major Practical 3, Module 2."""


class Node:
    """One item, and the reference to the next one."""

    def __init__(self, data, nxt=None):
        self.data = data
        self.nxt = nxt

    def __repr__(self):
        return f"Node({self.data!r})"


class SinglyLinkedList:
    """The container. It owns the head, the tail and the count."""

    def __init__(self, items=()):
        self.head = None
        self.tail = None
        self.count = 0
        self.steps = 0
        for item in items:
            self.insert_at_end(item)

    # ---- asking about it -------------------------------------------------
    def is_empty(self):
        return self.head is None

    def __len__(self):
        return self.count

    def __iter__(self):
        here = self.head
        while here is not None:
            yield here.data
            here = here.nxt

    def __repr__(self):
        if self.is_empty():
            return "head -> None   (empty, 0 nodes)"
        chain = " -> ".join(f"[ {item} | * ]" for item in self)
        return f"head -> {chain[:-9]}[ {list(self)[-1]} | None ]   ({self.count} nodes)"

    def as_chain(self):
        """A plain readable form for the journal."""
        if self.is_empty():
            return "None   (empty)"
        return " -> ".join(str(item) for item in self) + " -> None"

    # ---- inserting -------------------------------------------------------
    def insert_at_beginning(self, data):
        """MU's first case. One step, whatever the length."""
        node = Node(data, self.head)
        self.head = node
        if self.tail is None:
            self.tail = node
        self.count += 1
        return 0

    def insert_at_end(self, data):
        """MU's second case. One step, BECAUSE a tail is kept."""
        node = Node(data)
        if self.is_empty():
            self.head = node
        else:
            self.tail.nxt = node
        self.tail = node
        self.count += 1
        return 0

    def insert_at_position(self, position, data):
        """MU's third case. Walks to the node before the position.

        position 0 is the beginning and position len(self) is the end.
        Returns how many nodes had to be walked past.
        """
        if not 0 <= position <= self.count:
            raise IndexError(
                f"cannot insert at {position}; 0 to {self.count} allowed")
        if position == 0:
            return self.insert_at_beginning(data)
        if position == self.count:
            return self.insert_at_end(data)
        before = self.head
        walked = 0
        for _ in range(position - 1):
            before = before.nxt
            walked += 1
        before.nxt = Node(data, before.nxt)
        self.count += 1
        self.steps += walked
        return walked

    # ---- looking ---------------------------------------------------------
    def traverse(self):
        """Every value, front to back, with the nodes visited counted."""
        values = []
        here = self.head
        visited = 0
        while here is not None:
            values.append(here.data)
            visited += 1
            here = here.nxt
        return values, visited

    def search(self, value):
        """The position of the first match and the nodes visited, or (-1, n)."""
        here = self.head
        position = 0
        while here is not None:
            if here.data == value:
                return position, position + 1
            here = here.nxt
            position += 1
        return -1, position

    def get(self, position):
        """The value at a position. O(n), which is the whole point."""
        if not 0 <= position < self.count:
            raise IndexError(f"no position {position} in a list of {self.count}")
        here = self.head
        for _ in range(position):
            here = here.nxt
        return here.data
munotes.in144

Practical 2: Building a Singly Linked List

Five things in that file are the marks.

insert_at_beginning builds the node first and only then moves the head. Two assignments, in that order. Reverse them and the rest of the list is lost, which the section below shows.

insert_at_end uses the tail, so it is two assignments rather than a walk. Without a tail it would be O(n), and building n items by appending would be O(n squared).

Both of them check for an empty list, because inserting the first node has to set the head and the tail.

insert_at_position walks to the node BEFORE the position, because a new node is linked in by changing the nxt of the node in front of it. position - 1 steps, not position.

munotes.in145

Practical 2: Building a Singly Linked List

position 0 and position == count are handed to the other two methods, which removes the special cases from the middle of the walk.

Creating the list and inserting all three ways

from linked import SinglyLinkedList

items = SinglyLinkedList()
print("empty                    ", items.as_chain(), f"  len {len(items)}")

items.insert_at_end(10)
print("insert_at_end(10)        ", items.as_chain(), f"  len {len(items)}")

items.insert_at_end(20)
items.insert_at_end(30)
print("two more at the end      ", items.as_chain(), f"  len {len(items)}")

items.insert_at_beginning(5)
print("insert_at_beginning(5)   ", items.as_chain(), f"  len {len(items)}")

walked = items.insert_at_position(2, 15)
print("insert_at_position(2, 15)", items.as_chain(),
      f"  len {len(items)}  walked {walked}")

walked = items.insert_at_position(0, 1)
print("insert_at_position(0, 1) ", items.as_chain(),
      f"  len {len(items)}  walked {walked}")

walked = items.insert_at_position(len(items), 99)
print("insert at the very end   ", items.as_chain(),
      f"  len {len(items)}  walked {walked}")

print()
print("head is", items.head, "and tail is", items.tail)
print("the tail's nxt is", items.tail.nxt, "which is how the end is known")
empty                     None   (empty)   len 0
insert_at_end(10)         10 -> None   len 1
two more at the end       10 -> 20 -> 30 -> None   len 3
insert_at_beginning(5)    5 -> 10 -> 20 -> 30 -> None   len 4
insert_at_position(2, 15) 5 -> 10 -> 15 -> 20 -> 30 -> None   len 5  walked 1
insert_at_position(0, 1)  1 -> 5 -> 10 -> 15 -> 20 -> 30 -> None   len 6  walked 0
insert at the very end    1 -> 5 -> 10 -> 15 -> 20 -> 30 -> 99 -> None   len 7  walked 0

head is Node(1) and tail is Node(99)
the tail's nxt is None which is how the end is known

Read the walk counts. Inserting at the beginning or at the end walked past nothing at all, because the head and the tail are held. Inserting at position 2 walked past one node, to reach the node before it. That is the linked list's bargain: cheap at the ends, and a walk in the middle.

The two cases that break a wrong answer

from linked import SinglyLinkedList

print("an EMPTY list:")
empty = SinglyLinkedList()
print("  is_empty      ", empty.is_empty())
print("  len           ", len(empty))
print("  chain         ", empty.as_chain())
print("  traverse      ", empty.traverse())
print("  search for 10 ", empty.search(10))
empty.insert_at_beginning(10)
print("  after inserting one at the beginning:")
print("    head", empty.head, " tail", empty.tail, " same node?", empty.head is empty.tail)

print()
print("a ONE NODE list:")
one = SinglyLinkedList([42])
print("  chain         ", one.as_chain())
print("  head is tail? ", one.head is one.tail)
one.insert_at_end(43)
print("  after an append:", one.as_chain(), " tail now", one.tail)
one.insert_at_beginning(41)
print("  after a prepend:", one.as_chain(), " head now", one.head,
      " tail still", one.tail)
an EMPTY list:
  is_empty       True
  len            0
  chain          None   (empty)
  traverse       ([], 0)
  search for 10  (-1, 0)
  after inserting one at the beginning:
    head Node(10)  tail Node(10)  same node? True

a ONE NODE list:
  chain          42 -> None
  head is tail?  True
  after an append: 42 -> 43 -> None  tail now Node(43)
  after a prepend: 41 -> 42 -> 43 -> None  head now Node(41)  tail still Node(43)
munotes.in146

Practical 2: Building a Singly Linked List

In an empty list the head and the tail are the same node after the first insertion, and in a one node list they already are. An answer that sets only the head when inserting into an empty list leaves the tail as None, and then the next insert_at_end raises AttributeError on self.tail.nxt. That is the bug this section exists to prevent.

The order of the two assignments

class Node:
    def __init__(self, data, nxt=None):
        self.data = data
        self.nxt = nxt


def chain(head):
    parts = []
    here = head
    while here is not None:
        parts.append(str(here.data))
        here = here.nxt
    return " -> ".join(parts) + " -> None"


# build 20 -> 30
head = Node(20, Node(30))
print("before          ", chain(head))

# the RIGHT order: point the new node at the old head, then move the head
node = Node(10, head)
head = node
print("right order     ", chain(head))

# the WRONG order, on a fresh list
head2 = Node(20, Node(30))
wrong = Node(10)
head2 = wrong          # the head moves first
wrong.nxt = head2      # and now it points at ITSELF, not at the old list
print("wrong order     ", chain(head2) if head2.nxt is not head2 else
      "head points at itself, so the list is one node and the rest is lost")
before           20 -> 30 -> None
right order      10 -> 20 -> 30 -> None
wrong order      head points at itself, so the list is one node and the rest is lost

The wrong order loses everything after the new node, because by the time the link is set the old head has already been forgotten. Build the new node with its link, then move the head. In one line: self.head = Node(data, self.head).

What each operation costs, counted

from linked import SinglyLinkedList

sizes = [10, 100, 1000]
print(f"{'n':>6} {'get(0)':>8} {'get(n-1)':>10} {'search first':>14} {'search last':>13}"
      f" {'search absent':>15}")
for n in sizes:
    items = SinglyLinkedList(range(n))
    _, first_steps = items.search(0)
    _, last_steps = items.search(n - 1)
    _, absent_steps = items.search(-1)
    print(f"{n:>6} {1:>8} {n:>10} {first_steps:>14} {last_steps:>13} {absent_steps:>15}")

print()
items = SinglyLinkedList(range(10))
print("insert_at_position walk counts on a list of 10:")
for position in [0, 1, 5, 9, 10]:
    test = SinglyLinkedList(range(10))
    print(f"  position {position:>2} walked {test.insert_at_position(position, 99):>2} node(s)")
     n   get(0)   get(n-1)   search first   search last   search absent
    10        1         10              1            10              10
   100        1        100              1           100             100
  1000        1       1000              1          1000            1000

insert_at_position walk counts on a list of 10:
  position  0 walked  0 node(s)
  position  1 walked  0 node(s)
  position  5 walked  4 node(s)
  position  9 walked  8 node(s)
  position 10 walked  0 node(s)
munotes.in147

Practical 2: Building a Singly Linked List

OperationCostWhy
insert_at_beginningO(1)two assignments, and the head is held
insert_at_endO(1)two assignments, because a tail is held
insert_at_position(p)O(p)walk to the node before p
get(p)O(p), so O(n)there is no arithmetic on an address
searchO(n)up to every node is visited
traverseO(n)every node once
lenO(1)because a count is kept

The array against the linked list, measured

from linked import SinglyLinkedList


class CountingArray:
    """A fixed array, counting the values it shifts."""

    def __init__(self, capacity):
        self.slots = [None] * capacity
        self.size = 0

    def insert(self, index, value):
        moved = 0
        for i in range(self.size, index, -1):
            self.slots[i] = self.slots[i - 1]
            moved += 1
        self.slots[index] = value
        self.size += 1
        return moved


n = 500

array = CountingArray(n + 1)
array_moves = 0
for value in range(n):
    array_moves += array.insert(0, value)

items = SinglyLinkedList()
list_moves = 0
for value in range(n):
    list_moves += items.insert_at_beginning(value)

print(f"building {n} items AT THE FRONT")
print(f"  array       : {array_moves} value(s) moved")
print(f"  linked list : {list_moves} value(s) moved")
print()

array2 = CountingArray(n)
for value in range(n):
    array2.insert(array2.size, value)
items2 = SinglyLinkedList(range(n))

middle = n // 2
print(f"reading the middle item of {n}")
print(f"  array       : 1 step, it is arithmetic")
_, steps = items2.search(middle)
print(f"  linked list : {steps} step(s), it has to walk")
building 500 items AT THE FRONT
  array       : 124750 value(s) moved
  linked list : 0 value(s) moved

reading the middle item of 500
  array       : 1 step, it is arithmetic
  linked list : 251 step(s), it has to walk

Read those two blocks against each other, because together they are the answer to "which is better", and the answer is neither.

Building at the front, the array moved a value for every item already there and the linked list moved nothing at all. Reading the middle item, the array did it in one step and the linked list walked halfway.

So the choice is decided by what the program does most, which is MU's Course Objective 8 exactly. A linked list trades position arithmetic for cheap insertion. Say that sentence at the table and the follow up question is answered before it is asked.

Procedure

  1. Save linked.py with Node holding data and nxt, and SinglyLinkedList holding head,

tail and count, both with __repr__.

  1. Write insert_at_beginning: build the node pointing at the old head, then move the head, and

set the tail too if the list was empty.

  1. Write insert_at_end using the tail, handling the empty list.
  2. Write insert_at_position, checking the bounds, walking position - 1 nodes, and handing

position 0 and position count to the other two methods.

munotes.in148

Practical 2: Building a Singly Linked List

  1. Write traverse, search and get, each counting the nodes it visits.
  2. In a second file, build a list, insert at the beginning, at the end and at a position, and print

the chain after every step with the walk count.

  1. Test the empty list and the one node list explicitly, and print whether the head and the tail

are the same node.

  1. Write the wrong order of the two assignments and record what it does.
  2. Count the search steps for the first, the last and an absent value, for n of 10, 100 and 1000.
  3. Compare with an array on building at the front and on reading the middle.

Result

The list was created empty and grown to seven nodes by all three insertions. Inserting at the beginning and at the end walked past no nodes; inserting at position 2 walked past one. On the empty list, the first insertion set the head and the tail to the same node. Reversing the two assignments in insert_at_beginning left the new node pointing at itself and lost the rest of the list. Search steps were 1 for the first value, n for the last and n for an absent value, at every size tried. Building 500 items at the front moved 124750 values in the array and none at all in the linked list; reading the middle item took 1 step in the array and 251 in the list.

Where marks are lost

  • Using a Python list and calling it a linked list. The exercise is the nodes.
  • No tail, and then claiming insert_at_end is O(1). Without a tail it is O(n).
  • Not setting the tail when inserting into an empty list, so the next append raises

AttributeError.

  • Moving the head before linking the new node, which loses the rest of the list.
  • Walking position nodes instead of position - 1, which inserts one place too far along.
  • No bounds check on the position.
  • No __repr__, so the output is a column of memory addresses.
  • Not testing the empty and one node cases. They are where a wrong answer fails.
  • Saying a linked list is faster than an array. It is faster at one thing and thousands of times

slower at another, and the marks are in saying which.

For the journal

The aim in MU's words, all three insertions. The picture first, three nodes with their references and the trailing None, because that diagram is worth a mark on its own. Then the seven row table of array against linked list. Then linked.py in full and the driver, with the chain printed after every insertion and the walk count beside it. Then the empty and one node tests with the head and tail lines. Then the wrong order of the two assignments and what it produced, with one sentence: build the node with its link, then move the head. Then the measured comparison, both figures. The conclusion: a linked list is a chain of nodes joined by references, so it inserts at either end in one step and has to walk to reach position n, which is the exact opposite of an array.

munotes.in149

Practical 2: Building a Singly Linked List

Quick revision

  • A node holds a value and a reference to the next. The last reference is None.
  • The list is a reference to the first node, the head. Keep a tail and a count too.
  • Every operation must maintain all three. Inserting into an empty list sets the head and the

tail.

  • At the beginning: self.head = Node(data, self.head). Build first, then move the head.
  • At the end: self.tail.nxt = node then self.tail = node. O(1) only because of the tail.
  • At a position: walk position - 1 nodes to reach the one BEFORE it, then

before.nxt = Node(data, before.nxt).

  • Position 0 is the beginning and position count is the end; hand both to the other methods.
  • Costs: both ends O(1), at position p O(p), get and search O(n), len O(1) with a

count.

  • Measured: building 500 items at the front moved 124750 values in an array and 0 in a list;

reading the middle took 1 step in the array and 251 in the list.

  • A linked list has no way back from a node to the one before it. That is why the next chapter needs

a trailing reference.

Questions you should be able to answer

1. What does a node hold, and what marks the end of the list? A value and a reference to the next node. The last node's reference is None.

2. What is the head? The reference to the first node, which is the only thing the list itself holds. It is None when the list is empty.

3. Write insert_at_beginning in one line. self.head = Node(data, self.head), then update the tail if the list was empty and increase the count.

4. What goes wrong if you move the head before linking the new node? The old head has already been forgotten, so the new node ends up pointing at itself or at nothing and the rest of the list is lost.

5. Why does insert_at_end need a tail? Without one it has to walk from the head to the last node, which is O(n), so building n items by appending would cost n squared steps. With a tail it is two assignments.

munotes.in150

Practical 2: Building a Singly Linked List

6. Inserting at position 5, how many nodes do you walk past, and to which one? Four, to reach position 4, the node before the insertion point, because linking a node in means changing the nxt of the node in front of it.

7. What must happen when you insert into an empty list? Both the head and the tail must be set to the new node. Setting only the head leaves the tail None and the next append fails.

8. What does get(n) cost, and why is it not O(1)? O(n). There is no address arithmetic, because the nodes are not side by side, so the only way to position n is to follow n references.

9. Which is better, an array or a linked list? Neither. Counted here: building 500 items at the front moved 124750 values in the array and none in the list, while reading the middle item took 1 step in the array and 251 in the list.

10. Why can a singly linked node not reach the node before it? Because it holds only a forward reference. That is the defect a doubly linked list cures, and it is why deletion needs a trailing reference.

Contents This chapter on its own page

munotes.in151

Chapter Twenty-Four

Practical 2 continued: Deleting a Node from a Linked List

Syllabus topic Module 2, practical 2(c), "Delete a node from a given position in a linked list"

Aim

To delete a node from a given position in a singly linked list.

Why deletion is harder than insertion

To remove a node, the nxt of the node in front of it has to be changed to skip it.

before deleting 20:

head -> [ 10 | * ] -> [ 20 | * ] -> [ 30 | None ]
           before        here          here.nxt

after:

head -> [ 10 | ------------------> [ 30 | None ]
                     [ 20 | * ]  <- nothing points at this any more

And a singly linked node has no way back to the node before it. So the walk has to carry a second reference, one node behind, and that trailing reference is the whole difficulty of this exercise.

The three cases

CaseWhat changesThe trap
the headself.head = head.nxtthere is no node in front, so the general code does not apply
the middlebefore.nxt = here.nxtneeds the trailing reference
the tailbefore.nxt = None and self.tail = beforeforgetting the tail update

And a fourth, which is the head and the tail at once: deleting the only node must set both to None.

The class

"""A singly linked list with deletion, for Major Practical 3, Module 2."""


class Node:
    def __init__(self, data, nxt=None):
        self.data = data
        self.nxt = nxt

    def __repr__(self):
        return f"Node({self.data!r})"


class SinglyLinkedList:
    def __init__(self, items=()):
        self.head = None
        self.tail = None
        self.count = 0
        for item in items:
            self.insert_at_end(item)

    # ---- the parts from the previous chapter -----------------------------
    def is_empty(self):
        return self.head is None

    def __len__(self):
        return self.count

    def __iter__(self):
        here = self.head
        while here is not None:
            yield here.data
            here = here.nxt

    def as_chain(self):
        if self.is_empty():
            return "None   (empty)"
        return " -> ".join(str(item) for item in self) + " -> None"

    def insert_at_beginning(self, data):
        self.head = Node(data, self.head)
        if self.tail is None:
            self.tail = self.head
        self.count += 1

    def insert_at_end(self, data):
        node = Node(data)
        if self.is_empty():
            self.head = node
        else:
            self.tail.nxt = node
        self.tail = node
        self.count += 1

    # ---- MU's third bullet ----------------------------------------------
    def delete_at_position(self, position):
        """Remove the node at a position and return (value, nodes walked).

        Three cases: the head, the middle, and the tail.
        """
        if self.is_empty():
            raise IndexError("cannot delete from an empty list")
        if not 0 <= position < self.count:
            raise IndexError(
                f"no position {position} in a list of {self.count}")

        # case 1: the head. There is no node in front of it.
        if position == 0:
            here = self.head
            self.head = here.nxt
            if self.head is None:          # it was the only node
                self.tail = None
            self.count -= 1
            here.nxt = None                # unlink it completely
            return here.data, 0

        # cases 2 and 3: walk, carrying the node behind
        before = self.head
        walked = 0
        for _ in range(position - 1):
            before = before.nxt
            walked += 1
        here = before.nxt
        before.nxt = here.nxt
        if here is self.tail:              # case 3: it WAS the tail
            self.tail = before
        self.count -= 1
        here.nxt = None
        return here.data, walked

    def delete_value(self, value):
        """Remove the first node holding this value. Returns its position, or -1."""
        before = None
        here = self.head
        position = 0
        while here is not None:
            if here.data == value:
                if before is None:
                    self.head = here.nxt
                else:
                    before.nxt = here.nxt
                if here is self.tail:
                    self.tail = before
                if self.head is None:
                    self.tail = None
                self.count -= 1
                here.nxt = None
                return position
            before = here
            here = here.nxt
            position += 1
        return -1

    def reverse(self):
        """Reverse the list in place, with three references."""
        before = None
        here = self.head
        self.tail = self.head
        while here is not None:
            after = here.nxt       # save it BEFORE overwriting
            here.nxt = before
            before = here
            here = after
        self.head = before
munotes.in152

Practical 2 continued: Deleting a Node from a Linked List

Five things in that file are the marks.

The head case is separate, because there is no node in front of it to change.

before walks position - 1 steps, exactly as for insertion, and here is before.nxt.

if here is self.tail: self.tail = before is the line most answers are missing.

Deleting the only node sets both the head and the tail to None, which the head case does with if self.head is None.

here.nxt = None at the end unlinks the removed node completely. Not doing it leaves the removed node still pointing into the list, which is harmless here and is a real bug the moment anybody keeps a reference to the removed node.

All three cases run

from linked2 import SinglyLinkedList

items = SinglyLinkedList([10, 20, 30, 40, 50])
print("start                   ", items.as_chain(), f" len {len(items)}")

value, walked = items.delete_at_position(2)
print(f"delete position 2       ", items.as_chain(),
      f" removed {value}, walked {walked}")

value, walked = items.delete_at_position(0)
print(f"delete the HEAD         ", items.as_chain(),
      f" removed {value}, walked {walked}")

value, walked = items.delete_at_position(len(items) - 1)
print(f"delete the TAIL         ", items.as_chain(),
      f" removed {value}, walked {walked}")
print("  and the tail is now    ", items.tail, "with nxt", items.tail.nxt)

items.insert_at_end(99)
print("append after that delete", items.as_chain(),
      "  the append WORKED, so the tail was updated")

value, walked = items.delete_at_position(0)
value, walked = items.delete_at_position(0)
print("down to the last node   ", items.as_chain())
value, walked = items.delete_at_position(0)
print("delete the only node    ", items.as_chain(),
      f" removed {value}")
print("  head", items.head, " tail", items.tail, " len", len(items))
start                    10 -> 20 -> 30 -> 40 -> 50 -> None  len 5
delete position 2        10 -> 20 -> 40 -> 50 -> None  removed 30, walked 1
delete the HEAD          20 -> 40 -> 50 -> None  removed 10, walked 0
delete the TAIL          20 -> 40 -> None  removed 50, walked 1
  and the tail is now     Node(40) with nxt None
append after that delete 20 -> 40 -> 99 -> None   the append WORKED, so the tail was updated
down to the last node    99 -> None
delete the only node     None   (empty)  removed 99
  head None  tail None  len 0
munotes.in153

Practical 2 continued: Deleting a Node from a Linked List

Look at the line after the tail deletion. The append worked and 99 appeared at the end, which is the proof that the tail was moved back. The next section shows what happens when it is not.

The bug: forgetting to move the tail

class Node:
    def __init__(self, data, nxt=None):
        self.data = data
        self.nxt = nxt

    def __repr__(self):
        return f"Node({self.data!r})"


class BrokenList:
    """Deletes correctly but never moves the tail back."""

    def __init__(self, items):
        self.head = None
        self.tail = None
        for item in items:
            node = Node(item)
            if self.head is None:
                self.head = node
            else:
                self.tail.nxt = node
            self.tail = node

    def chain(self):
        parts, here = [], self.head
        while here is not None:
            parts.append(str(here.data))
            here = here.nxt
        return " -> ".join(parts) + " -> None"

    def delete_last(self):
        before, here = None, self.head
        while here.nxt is not None:
            before, here = here, here.nxt
        before.nxt = None
        # the tail is NOT updated. That is the bug.

    def append(self, value):
        node = Node(value)
        self.tail.nxt = node
        self.tail = node


items = BrokenList([10, 20, 30])
print("start            ", items.chain(), " tail", items.tail)

items.delete_last()
print("after delete_last", items.chain(), " tail", items.tail,
      "  <- the tail is a node NOT in the list")

detached = items.tail
items.append(99)
print("after append(99) ", items.chain(),
      "  <- where did 99 go?")
print("  it was linked on to", detached, "which is not in the list,")
print("  so detached.nxt is", detached.nxt, "and nothing reaches it")
start             10 -> 20 -> 30 -> None  tail Node(30)
after delete_last 10 -> 20 -> None  tail Node(30)   <- the tail is a node NOT in the list
after append(99)  10 -> 20 -> None   <- where did 99 go?
  it was linked on to Node(30) which is not in the list,
  so detached.nxt is Node(99) and nothing reaches it

There is the whole bug. The chain still reads 10 -> 20 -> None and 99 is nowhere in it, because the append linked it on to the node that had already been removed. Nothing raised, nothing printed an error, and the data is simply lost.

A stale tail is the most dangerous bug in this exercise precisely because it is silent. One line fixes it: if here is self.tail: self.tail = before.

munotes.in154

Practical 2 continued: Deleting a Node from a Linked List

Deleting by value, and the trailing reference in its clearest form

from linked2 import SinglyLinkedList

items = SinglyLinkedList(["Physics", "Chemistry", "Maths", "Biology"])
print("start            ", items.as_chain())

print("delete 'Maths'   ", end=" ")
position = items.delete_value("Maths")
print(f"was at position {position}:", items.as_chain())

print("delete 'Physics' ", end=" ")
position = items.delete_value("Physics")
print(f"was at position {position}:", items.as_chain(),
      " head now", items.head)

print("delete 'Biology' ", end=" ")
position = items.delete_value("Biology")
print(f"was at position {position}:", items.as_chain(),
      " tail now", items.tail)

print("delete 'History' ", end=" ")
position = items.delete_value("History")
print(f"returned {position}, which means not found:", items.as_chain())
start             Physics -> Chemistry -> Maths -> Biology -> None
delete 'Maths'    was at position 2: Physics -> Chemistry -> Biology -> None
delete 'Physics'  was at position 0: Chemistry -> Biology -> None  head now Node('Chemistry')
delete 'Biology'  was at position 1: Chemistry -> None  tail now Node('Chemistry')
delete 'History'  returned -1, which means not found: Chemistry -> None

before is None is how this version recognises the head case, which is neater than a separate block: if nothing is behind here, then here is the head.

Why a singly linked list cannot delete a node it is holding

This is the question an examiner asks to separate a student who has understood the structure from one who has copied it.

class Node:
    def __init__(self, data, nxt=None):
        self.data = data
        self.nxt = nxt

    def __repr__(self):
        return f"Node({self.data!r})"


def chain(head):
    parts, here = [], head
    while here is not None:
        parts.append(str(here.data))
        here = here.nxt
    return " -> ".join(parts) + " -> None"


third = Node(30)
second = Node(20, third)
head = Node(10, second)

print("the list          ", chain(head))
print("I am holding      ", second)
print("can I reach the node before it? ",
      "no: a node has only a forward reference")
print("so to delete it I must walk from the head again:")

before, here = None, head
steps = 0
while here is not second:
    before, here = here, here.nxt
    steps += 1
before.nxt = here.nxt
print(f"  walked {steps} node(s) from the head, then one assignment")
print("the list          ", chain(head))
the list           10 -> 20 -> 30 -> None
I am holding       Node(20)
can I reach the node before it?  no: a node has only a forward reference
so to delete it I must walk from the head again:
  walked 1 node(s) from the head, then one assignment
the list           10 -> 30 -> None

So holding the node does not help: deletion still costs a walk from the head, which is O(n). That is the single strongest argument for a doubly linked list, in which each node also holds a reference backwards and a held node can be removed in constant time.

It is also why MU's practical 3 and 4 are a stack and a queue: both only ever remove from a position they already hold, which a singly linked list does perfectly well.

munotes.in155

Practical 2 continued: Deleting a Node from a Linked List

What happens to the deleted node

import gc
from linked2 import SinglyLinkedList


class Watched:
    def __init__(self, name):
        self.name = name

    def __del__(self):
        print(f"    the object {self.name} was freed")

    def __repr__(self):
        return self.name


items = SinglyLinkedList([Watched("a"), Watched("b"), Watched("c")])
print("list built:", items.as_chain())

print("  deleting position 1")
value, _ = items.delete_at_position(1)
print("  delete returned", value, "so a name still holds it")

print("  now dropping that name")
del value
gc.collect()

print("list now  :", items.as_chain())
list built: a -> b -> c -> None
  deleting position 1
  delete returned b so a name still holds it
  now dropping that name
    the object b was freed
list now  : a -> c -> None
    the object a was freed
    the object c was freed

Read the order carefully. The object was not freed when it was unlinked, because delete_at_position returned it and a name still held it. It was freed the moment that name went away. That is reference counting, and it is MU's Course Objective 9 in one demonstration. There is no free to call, and the memory goes back when the last reference does. [If Your College Runs Module 2 in C] shows the same deletion with free, where forgetting the call is a leak.

Reversing the list, which is the other question on this exercise

from linked2 import SinglyLinkedList

items = SinglyLinkedList([10, 20, 30, 40, 50])
print("before  ", items.as_chain(), " head", items.head, " tail", items.tail)
items.reverse()
print("after   ", items.as_chain(), " head", items.head, " tail", items.tail)

items.insert_at_end(5)
print("append 5", items.as_chain(), "  the tail was updated by reverse")

one = SinglyLinkedList([7])
one.reverse()
print("one node", one.as_chain())

empty = SinglyLinkedList()
empty.reverse()
print("empty   ", empty.as_chain())
before   10 -> 20 -> 30 -> 40 -> 50 -> None  head Node(10)  tail Node(50)
after    50 -> 40 -> 30 -> 20 -> 10 -> None  head Node(50)  tail Node(10)
append 5 50 -> 40 -> 30 -> 20 -> 10 -> 5 -> None   the tail was updated by reverse
one node 7 -> None
empty    None   (empty)

The reversal uses three references: before, here and after. The order of the four lines inside the loop is the whole answer:

after = here.nxt     # save it FIRST
here.nxt = before    # then turn the link round
before = here        # then advance both
here = after

Overwrite here.nxt before saving after and the rest of the list is unreachable. And note that reverse sets self.tail = self.head at the start, before the head moves, because the old head becomes the new tail.

Procedure

  1. Save linked2.py with the list from the previous chapter plus delete_at_position,

delete_value and reverse.

  1. In delete_at_position, handle the head as its own case, walk position - 1 nodes for the rest,
munotes.in156

Practical 2 continued: Deleting a Node from a Linked List

and move the tail back when the deleted node was the tail.

  1. Set both the head and the tail to None when the only node is deleted.
  2. Raise IndexError for an empty list and for a position out of range.
  3. In a driver, delete from the middle, the head and the tail of a five node list, printing the chain

and the walk count each time.

  1. Append immediately after deleting the tail and confirm the new node appears. That is the test

for the stale tail.

  1. Write a deliberately broken version that does not move the tail, and record that the append is

silently lost.

  1. Delete every node one at a time down to the empty list, and print the head, the tail and the

length at the end.

  1. Write delete_value using before is None to recognise the head.
  2. Reverse the list with three references and confirm the tail was updated by appending afterwards.

Result

Deleting from position 2 walked past one node; deleting the head walked past none; deleting the tail walked to the node before it and moved the tail back, which an append immediately afterwards confirmed by appearing at the end. Deleting the last remaining node set both the head and the tail to None. The deliberately broken version lost the appended value silently: the chain still read 10 -> 20 -> None after appending 99, because the append had been linked on to the detached node. delete_value removed a middle, a head and a tail value and returned -1 for a value not present. Deleting a node did not free it while the returned value was still named; it was freed when that name was deleted. Reversing a five node list, a one node list and an empty list all behaved correctly and the tail was updated.

Where marks are lost

  • No trailing reference. A singly linked node cannot reach the one before it, so the walk must

carry before.

  • Not moving the tail when the deleted node was the tail. This is the silent one.
  • Not setting both head and tail to None when the only node goes.
  • Treating the head like any other node, which crashes because there is nothing in front of it.
  • Walking position nodes instead of position - 1.
  • No bounds check, and no error for an empty list.
  • Not returning the removed value.
  • Overwriting here.nxt before saving after in reverse, which loses the rest of the list.
  • Forgetting that reverse must also swap the head and the tail.

For the journal

The aim in MU's words. The two pictures, before and after, with the removed node shown detached and nothing pointing at it. The three case table with the tail line highlighted. The class in full and the driver, with the chain and the walk count printed after every deletion, and then the append immediately after the tail deletion, because that line is the evidence that the tail was updated. Then the broken version and its silently lost append, with one sentence: a stale tail is dangerous because nothing raises. Then the deletion down to empty, with the head, tail and length printed. One sentence on why holding a node does not help: deletion changes the previous node's link and there is no way back, which is why a doubly linked list exists. The conclusion: deletion needs the node in front of the one being removed, so it has three cases and a trailing reference.

munotes.in157

Practical 2 continued: Deleting a Node from a Linked List

Quick revision

  • Deletion changes the nxt of the node in front, so the walk carries a trailing reference

before.

  • Three cases: the head (self.head = here.nxt), the middle (before.nxt = here.nxt), the tail

(also self.tail = before).

  • A fourth: the only node, which sets head and tail to None.
  • before walks position - 1 steps; here is before.nxt.
  • A stale tail is silent. The next append links onto a detached node and the value disappears.
  • here.nxt = None after unlinking, so the removed node does not still point into the list.
  • delete_value: before is None means here is the head.
  • Deletion is O(n) even when you are holding the node, because the walk is needed to find the one

in front. A doubly linked list fixes that.

  • Python frees the node when the last reference goes, not when it is unlinked.
  • reverse: three references, and save after before overwriting here.nxt. Set the tail to

the old head first.

Questions you should be able to answer

1. Why does deletion need a trailing reference? Because removing a node means changing the nxt of the node in front of it, and a singly linked node has no reference backwards.

2. What are the three cases of deletion? The head, where there is no node in front; the middle, where before.nxt = here.nxt; and the tail, where the tail must also be moved back to before.

3. What is the fourth case? Deleting the only node, which must set both the head and the tail to None.

4. What goes wrong if you do not move the tail? The tail still points at a node that has been removed, so the next append links the new node on to that detached node and the value never appears in the list. Nothing raises, which is why it is dangerous.

munotes.in158

Practical 2 continued: Deleting a Node from a Linked List

5. How do you test for that bug in one line? Delete the tail and then append. If the appended value appears, the tail was updated.

6. To delete position 5, how many nodes do you walk and which one do you stop at? Four, stopping at position 4, the node before the one to be removed.

7. Why does holding the node not make deletion cheaper? Because the deletion changes the previous node's link, and there is no way from a node to its predecessor, so the walk from the head is still needed. That is O(n).

8. Which structure fixes that, and how? A doubly linked list: each node also holds a reference to the previous node, so a held node can be unlinked in constant time.

9. When is the removed node's memory released in Python? When the last reference to it goes away. Unlinking is not enough if the delete method returned it and a name still holds it.

10. Write the four lines that reverse a list. after = here.nxt, here.nxt = before, before = here, here = after. Saving after must come first, and when the loop ends before is the new head.

Contents This chapter on its own page

munotes.in159

Chapter Twenty-Five

Practical 3: a Stack over an Array

Syllabus topic Module 2, practical 3(a), "Stack Application: Write a program to: Implement a stack using an array"

Aim

To implement a stack using an array, with push, pop and peek, and to show both overflow and underflow.

What a stack is

A stack is a collection in which the last thing put in is the first thing taken out. That rule has a name, LIFO, last in first out, and it is the whole definition.

push 10, push 20, push 30:

           top -> [ 30 ]   index 2
                  [ 20 ]   index 1
                  [ 10 ]   index 0

pop gives 30, because it went in last.

Everything happens at one end, called the top. Nothing is ever inserted or removed anywhere else, and that is why a stack needs no shifting and no searching: every operation is O(1).

OperationWhat it doesCost
push(x)put x on the topO(1)
pop()remove and return the topO(1)
peek()return the top without removing itO(1)
is_empty()is there nothing in itO(1)
is_full()is the array fullO(1)
size()how many itemsO(1)

top starts at minus one

The array holds the items and one integer, top, holds the index of the topmost item. When the stack is empty there is no topmost item, so top is -1.

empty:      top = -1    [ . | . | . | . ]
push 10:    top =  0    [ 10| . | . | . ]
push 20:    top =  1    [ 10| 20| . | . ]
pop -> 20:  top =  0    [ 10| 20| . | . ]   20 is still there but top no longer reaches it

Two things follow and both are exam answers. top == -1 means empty, and top == capacity - 1 means full. And notice the last line: popping does not erase anything, it only moves top. The value is unreachable, which is all that matters.

The class

"""A stack over a fixed array, for Major Practical 3, Module 2."""


class StackOverflow(Exception):
    """Pushing onto a full stack."""


class StackUnderflow(Exception):
    """Popping or peeking an empty stack."""


class ArrayStack:
    """LIFO, over an array of fixed capacity. top is -1 when empty."""

    def __init__(self, capacity=8):
        if capacity <= 0:
            raise ValueError("capacity must be at least 1")
        self.capacity = capacity
        self.slots = [None] * capacity
        self.top = -1

    # ---- asking -----------------------------------------------------------
    def is_empty(self):
        return self.top == -1

    def is_full(self):
        return self.top == self.capacity - 1

    def size(self):
        return self.top + 1

    def __len__(self):
        return self.size()

    def peek(self):
        """The top item, left where it is."""
        if self.is_empty():
            raise StackUnderflow("peek on an empty stack")
        return self.slots[self.top]

    def __repr__(self):
        if self.is_empty():
            return f"empty stack, top = -1, capacity {self.capacity}"
        items = ", ".join(repr(self.slots[i]) for i in range(self.top + 1))
        return f"bottom [{items}] top   top = {self.top}, {self.size()}/{self.capacity}"

    # ---- changing ---------------------------------------------------------
    def push(self, item):
        if self.is_full():
            raise StackOverflow(
                f"push onto a full stack of capacity {self.capacity}")
        self.top += 1
        self.slots[self.top] = item

    def pop(self):
        if self.is_empty():
            raise StackUnderflow("pop from an empty stack")
        item = self.slots[self.top]
        self.slots[self.top] = None     # drop the reference, so it can be freed
        self.top -= 1
        return item
munotes.in160

Practical 3: a Stack over an Array

Five things in that file are the marks.

top is -1 when empty, so is_empty is one comparison and size is top + 1.

push increases top and then writes. pop reads and then decreases. In that order, both times.

Overflow and underflow raise, and they raise their own exception types so a caller can tell them apart. Returning None on an empty pop is the mistake [Python for Data Structures, and the Cost of an Operation] shows failing.

peek does not remove. That is the entire difference from pop, and examiners ask it.

pop sets the slot to None. Not needed for correctness, and right anyway: it drops the reference so the object can be freed, which matters when the items are large.

Every operation run

from stack import ArrayStack, StackOverflow, StackUnderflow

stack = ArrayStack(capacity=4)
print("new           ", stack)
print("is_empty      ", stack.is_empty(), " is_full", stack.is_full())

for value in [10, 20, 30]:
    stack.push(value)
    print(f"push {value:<3}      ", stack)

print("peek          ", stack.peek(), " and the stack is unchanged:", stack)
print("size          ", stack.size(), "which is top + 1 =", stack.top, "+ 1")

print("pop           ", stack.pop(), "->", stack)
print("pop           ", stack.pop(), "->", stack)

stack.push(99)
print("push 99       ", stack)

for value in [40, 50]:
    stack.push(value)
print("filled up     ", stack, " is_full", stack.is_full())

try:
    stack.push(70)
except StackOverflow as error:
    print("push on full  ", "StackOverflow:", error)

while not stack.is_empty():
    stack.pop()
print("emptied       ", stack)

try:
    stack.pop()
except StackUnderflow as error:
    print("pop on empty  ", "StackUnderflow:", error)

try:
    stack.peek()
except StackUnderflow as error:
    print("peek on empty ", "StackUnderflow:", error)
new            empty stack, top = -1, capacity 4
is_empty       True  is_full False
push 10        bottom [10] top   top = 0, 1/4
push 20        bottom [10, 20] top   top = 1, 2/4
push 30        bottom [10, 20, 30] top   top = 2, 3/4
peek           30  and the stack is unchanged: bottom [10, 20, 30] top   top = 2, 3/4
size           3 which is top + 1 = 2 + 1
pop            30 -> bottom [10, 20] top   top = 1, 2/4
pop            20 -> bottom [10] top   top = 0, 1/4
push 99        bottom [10, 99] top   top = 1, 2/4
filled up      bottom [10, 99, 40, 50] top   top = 3, 4/4  is_full True
push on full   StackOverflow: push onto a full stack of capacity 4
emptied        empty stack, top = -1, capacity 4
pop on empty   StackUnderflow: pop from an empty stack
peek on empty  StackUnderflow: peek on an empty stack
munotes.in161

Practical 3: a Stack over an Array

Read the top values through that run. They go up on every push and down on every pop, and they reach -1 exactly when the stack is empty and capacity - 1 exactly when it is full.

Application one: reversing

The simplest use of a stack, and the one that makes LIFO obvious.

from stack import ArrayStack


def reverse_text(text):
    stack = ArrayStack(capacity=max(len(text), 1))
    for character in text:
        stack.push(character)
    out = []
    while not stack.is_empty():
        out.append(stack.pop())
    return "".join(out)


for text in ["munotes", "stack", "a", ""]:
    print(f"{text!r:<12} reversed is {reverse_text(text)!r}")
'munotes'    reversed is 'setonum'
'stack'      reversed is 'kcats'
'a'          reversed is 'a'
''           reversed is ''

Pushing every character and then popping them all gives them back in the opposite order, because that is what LIFO means. Note that the empty string works: nothing is pushed, so the while loop never runs and "".join([]) is the empty string.

The max(len(text), 1) is there because a capacity of 0 is refused by the class, and an empty input is exactly the case a careless answer crashes on.

Application two: matching brackets

This is the application MU's practical 3 is really pointing at, because the next chapter's infix to postfix conversion is the same idea with precedence added.

from stack import ArrayStack, StackUnderflow

PAIRS = {")": "(", "]": "[", "}": "{"}
OPENERS = set(PAIRS.values())


def brackets_match(text):
    """(True, message) if every bracket is closed by its own kind, in order."""
    stack = ArrayStack(capacity=max(len(text), 1))
    for position, character in enumerate(text):
        if character in OPENERS:
            stack.push((character, position))
        elif character in PAIRS:
            if stack.is_empty():
                return False, f"a {character!r} at {position} closes nothing"
            opener, opened_at = stack.pop()
            if opener != PAIRS[character]:
                return (False,
                        f"a {character!r} at {position} closes a "
                        f"{opener!r} opened at {opened_at}")
    if not stack.is_empty():
        opener, opened_at = stack.peek()
        return False, f"a {opener!r} at {opened_at} is never closed"
    return True, "balanced"


tests = ["(a + b) * [c - d]", "{[()]}", "(a + b", "a + b)", "(a + [b)]",
         "", "((()))", "]["]
for text in tests:
    ok, why = brackets_match(text)
    print(f"{text!r:<20} {'OK   ' if ok else 'WRONG'}  {why}")
'(a + b) * [c - d]'  OK     balanced
'{[()]}'             OK     balanced
'(a + b'             WRONG  a '(' at 0 is never closed
'a + b)'             WRONG  a ')' at 5 closes nothing
'(a + [b)]'          WRONG  a ')' at 7 closes a '[' opened at 5
''                   OK     balanced
'((()))'             OK     balanced
']['                 WRONG  a ']' at 0 closes nothing

Three separate faults have to be caught and the output shows all three:

FaultExampleDetected by
a closer with nothing opena + b)the stack is empty when a closer arrives
the wrong kind of closer(a + [b)]the popped opener is not the matching one
an opener never closed(a + bthe stack is not empty at the end
munotes.in162

Practical 3: a Stack over an Array

The third is the one most answers miss. A program that only checks as it goes reports (a + b as balanced, because nothing ever went wrong; the fault is what is left over.

The error messages carry the position, which is what makes this useful rather than decorative, and that is only possible because the stack holds a tuple of the bracket and where it was.

Application three: undo

from stack import ArrayStack, StackUnderflow


class Editor:
    """A text box with undo, which is a stack of previous states."""

    def __init__(self, capacity=10):
        self.text = ""
        self.history = ArrayStack(capacity)

    def type_in(self, more):
        self.history.push(self.text)
        self.text += more

    def delete_last(self, n=1):
        self.history.push(self.text)
        self.text = self.text[:-n]

    def undo(self):
        try:
            self.text = self.history.pop()
            return True
        except StackUnderflow:
            return False


editor = Editor()
for action in ["munotes", " makes", " notes"]:
    editor.type_in(action)
    print(f"typed {action!r:<10} -> {editor.text!r}")

editor.delete_last(6)
print(f"deleted 6 chars -> {editor.text!r}")

while editor.undo():
    print(f"undo            -> {editor.text!r}")
print("undo on an empty history returned False, so there is nothing left to undo")
typed 'munotes'  -> 'munotes'
typed ' makes'   -> 'munotes makes'
typed ' notes'   -> 'munotes makes notes'
deleted 6 chars -> 'munotes makes'
undo            -> 'munotes makes notes'
undo            -> 'munotes makes'
undo            -> 'munotes'
undo            -> ''
undo on an empty history returned False, so there is nothing left to undo

Undo is a stack and nothing else: each action pushes the state before it, and undo pops. The most recent action is undone first, which is LIFO, and it is why every editor in the world has a stack inside it.

The three ways to build a stack, compared

MU says an array. It is worth knowing the other two and being able to say what each costs.

from collections import deque
from stack import ArrayStack


class ListStack:
    """A stack over a Python list, which grows as needed."""

    def __init__(self):
        self.items = []

    def push(self, item):
        self.items.append(item)

    def pop(self):
        if not self.items:
            raise IndexError("pop from an empty stack")
        return self.items.pop()


class Node:
    def __init__(self, data, nxt=None):
        self.data = data
        self.nxt = nxt


class LinkedStack:
    """A stack over a linked list. The top is the head, so push is a prepend."""

    def __init__(self):
        self.head = None
        self.count = 0

    def push(self, item):
        self.head = Node(item, self.head)
        self.count += 1

    def pop(self):
        if self.head is None:
            raise IndexError("pop from an empty stack")
        item = self.head.data
        self.head = self.head.nxt
        self.count -= 1
        return item


for name, stack in [("ArrayStack ", ArrayStack(5)),
                    ("ListStack  ", ListStack()),
                    ("LinkedStack", LinkedStack()),
                    ("deque      ", deque())]:
    push = stack.append if isinstance(stack, deque) else stack.push
    pop = stack.pop
    for value in [1, 2, 3]:
        push(value)
    print(f"{name}: popped {pop()}, {pop()}, {pop()}")
munotes.in163

Practical 3: a Stack over an Array

ArrayStack : popped 3, 2, 1
ListStack  : popped 3, 2, 1
LinkedStack: popped 3, 2, 1
deque      : popped 3, 2, 1
Built onpushpopFixed sizeExtra memory
an array (MU's)O(1)O(1)yes, it can overflownone
a Python listO(1) amortisedO(1)nospare capacity
a linked listO(1)O(1)noa reference per item
collections.dequeO(1)O(1)no, unless maxlenits own blocks

All four are O(1) at both ends, which is the point: a stack is cheap however it is built. The array version is the one MU asks for and the only one that can overflow, and the linked version is the answer when the maximum size is not known.

"O(1) amortised" for a Python list means almost every append is one step, and occasionally one append reallocates the whole block. Averaged over many appends it is still constant, which is what amortised means, and it is worth knowing the word.

Procedure

  1. Save stack.py with StackOverflow and StackUnderflow exceptions and the ArrayStack class.
  2. Keep capacity, a slots list and top, with top starting at -1.
  3. is_empty is top == -1; is_full is top == capacity - 1; size is top + 1.
  4. push raises on full, increases top, then writes. pop raises on empty, reads, blanks the slot,

then decreases top.

  1. peek returns the top without changing anything and raises on empty.
  2. __repr__ prints the items from the bottom up with the current top.
  3. In a driver, push three, peek, pop two, fill the stack, and push once more to force overflow.
  4. Empty the stack and pop once more to force underflow, and peek on the empty stack too.
  5. Write the reversal, the bracket matcher and the undo editor.
  6. Test the bracket matcher on all three faults: a closer with nothing open, the wrong closer, and an

opener never closed.

Result

top was -1 on the new stack, rose with each push and fell with each pop, and reached capacity - 1 exactly when is_full became True. peek returned the top and left the stack unchanged. Pushing onto the full stack raised StackOverflow; popping and peeking the empty stack both raised StackUnderflow. Reversal by push then pop returned every test string backwards, including the empty string. The bracket matcher accepted the balanced strings and reported the position and the reason for all three kinds of fault, including (a + b, where the fault is what is left on the stack at the end. The undo editor restored each previous state in reverse order and returned False when the history was empty.

munotes.in164

Practical 3: a Stack over an Array

Where marks are lost

  • No capacity, so there is no overflow to demonstrate, and MU said "using an array".
  • top starting at 0 instead of -1, which makes is_empty wrong and loses slot 0.
  • Increasing top after writing, which writes over the current top.
  • Returning None on an empty pop instead of raising.
  • pop and peek doing the same thing. peek must not remove.
  • Not checking the stack is empty at the END of the bracket matcher, so (a + b passes.
  • One kind of bracket only, when three kinds are what makes the matching interesting.
  • Not saying what the operations cost. Every one of them is O(1), and that is the point of a

stack.

For the journal

The aim in MU's words. The stack picture with the three items and the arrow at the top, and the second picture showing top going from -1 upwards. The class in full. Then the run, with the stack and the top printed after every operation, and both errors forced, which is what the entry is really for. Then at least the bracket matcher, with all three faults in the test list and the reason printed for each, and one sentence: the third fault is found by checking that the stack is empty at the end. The conclusion: a stack touches one end only, so push, pop and peek are all O(1), and top == -1 means empty while top == capacity - 1 means full.

Quick revision

  • A stack is LIFO: last in, first out. Everything happens at the top.
  • push, pop, peek, is_empty, is_full, size. All O(1).
  • top starts at -1. is_empty is top == -1; is_full is top == capacity - 1;

size is top + 1.

  • push: check full, top += 1, then write. pop: check empty, read, blank, top -= 1.
  • peek does not remove. That is the only difference from pop.
  • Pushing onto a full stack is overflow; popping an empty one is underflow. Both raise.
  • Reversal: push everything, pop everything.
  • Bracket matching has three faults: a closer with nothing open, the wrong closer, and an opener

left on the stack at the end.

  • Undo is a stack of previous states.
  • A stack can be built on an array (fixed, can overflow), a Python list, a linked list, or a deque.

All are O(1) at both ends.

Questions you should be able to answer

1. What does LIFO mean, and where does it happen? Last in, first out. Every operation is at one end, the top.

munotes.in165

Practical 3: a Stack over an Array

2. Why is top initialised to -1? Because an empty stack has no topmost item, so the index before the first slot is the natural marker, and then size is simply top + 1.

3. How do you know the stack is full? top == capacity - 1.

4. What is the difference between pop and peek? pop removes the top item and returns it; peek returns it and leaves it in place.

5. What should pop do on an empty stack? Raise. Returning None cannot be distinguished from having pushed a None.

6. What is overflow and what is underflow? Overflow is pushing onto a full stack; underflow is popping or peeking an empty one.

7. What does every stack operation cost? O(1). Nothing is shifted and nothing is searched, because only one end is ever touched.

8. How do you reverse a string with a stack? Push every character, then pop them all; they come out in the opposite order.

9. Name the three faults a bracket matcher must catch. A closing bracket when nothing is open; a closing bracket of the wrong kind; and an opening bracket never closed, which is detected by the stack not being empty at the end.

10. Why is undo a stack? Because the most recent action must be undone first, which is exactly LIFO.

Contents This chapter on its own page

munotes.in166

Chapter Twenty-Six

Practical 3 continued: Infix to Postfix with a Stack

Syllabus topic Module 2, practical 3(b), "Convert an infix expression to postfix notation using a stack"

Aim

To convert an infix expression to postfix notation using a stack, and to evaluate the result.

The three notations

NameWhere the operator goesExample
infixbetween its two operandsa + b
prefix, or Polishbefore them+ a b
postfix, or reverse Polishafter thema b +

All three mean the same thing. Infix is what people write and it needs brackets and precedence rules to be read without ambiguity. Postfix needs neither.

infix     a + b * c        needs the rule that * binds tighter than +
postfix   a b c * +        no rule needed: the shape says it

infix     (a + b) * c      needs the brackets
postfix   a b + c *        no brackets needed: the shape says it

That is why postfix exists. A machine evaluating postfix reads left to right, keeps a stack, and never looks ahead or backtracks. Compilers and calculators convert to postfix for exactly that reason.

Precedence and associativity

Two rules decide what infix means, and both are needed by the conversion.

OperatorPrecedenceAssociativity
^ (power)3, highestright to left
* / %2left to right
+ -1, lowestleft to right

Precedence says which operator takes its operands first: in a + b c the does.

Associativity says what happens between two operators of equal precedence. a - b - c is (a - b) - c, left to right. But a ^ b ^ c is a ^ (b ^ c), right to left, and that single exception is the part of this exercise almost every answer gets wrong.

print("2 - 3 - 4 left to right :", (2 - 3) - 4, " and Python gives", 2 - 3 - 4)
print("2 - 3 - 4 right to left :", 2 - (3 - 4), " which is NOT what Python gives")
print()
print("2 ^ 3 ^ 2 right to left :", 2 ** (3 ** 2), " and Python gives", 2 ** 3 ** 2)
print("2 ^ 3 ^ 2 left to right :", (2 ** 3) ** 2, " which is NOT what Python gives")
2 - 3 - 4 left to right : -5  and Python gives -5
2 - 3 - 4 right to left : 3  which is NOT what Python gives

2 ^ 3 ^ 2 right to left : 512  and Python gives 512
2 ^ 3 ^ 2 left to right : 64  which is NOT what Python gives

The algorithm, in six rules

Read the infix expression left to right and keep a stack of operators.

  1. An operand goes straight to the output.
  2. An opening bracket is pushed.
  3. A closing bracket: pop operators to the output until the matching opening bracket is popped.
munotes.in167

Practical 3 continued: Infix to Postfix with a Stack

The brackets themselves are never output.

  1. An operator: while the stack has an operator on top that should come first, pop it to the

output. Then push this one.

  1. At the end, pop everything left to the output.
  2. An opening bracket still on the stack at the end, or a closing bracket with no opening one, is an

unbalanced expression.

Rule 4 is the whole algorithm, and "should come first" is where precedence and associativity live:

  • higher precedence on the stack, pop it;
  • equal precedence and the operator is left associative, pop it;
  • equal precedence and the operator is right associative, leave it.

The program

"""Infix to postfix, and the evaluation of postfix, for Module 2 practical 3."""

PRECEDENCE = {"+": 1, "-": 1, "*": 2, "/": 2, "%": 2, "^": 3}
RIGHT_ASSOCIATIVE = {"^"}


def tokenize(text):
    """Split into numbers, names, operators and brackets."""
    tokens = []
    i = 0
    while i < len(text):
        character = text[i]
        if character.isspace():
            i += 1
        elif character.isdigit() or character == ".":
            j = i
            while j < len(text) and (text[j].isdigit() or text[j] == "."):
                j += 1
            tokens.append(text[i:j])
            i = j
        elif character.isalpha():
            j = i
            while j < len(text) and text[j].isalnum():
                j += 1
            tokens.append(text[i:j])
            i = j
        elif character in PRECEDENCE or character in "()":
            tokens.append(character)
            i += 1
        else:
            raise ValueError(f"cannot read {character!r} at position {i}")
    return tokens


def is_operand(token):
    return token not in PRECEDENCE and token not in "()"


def pops_first(on_stack, arriving):
    """Should the operator already on the stack come out before this one?"""
    if on_stack == "(":
        return False
    if PRECEDENCE[on_stack] > PRECEDENCE[arriving]:
        return True
    if PRECEDENCE[on_stack] < PRECEDENCE[arriving]:
        return False
    return arriving not in RIGHT_ASSOCIATIVE      # equal: left associative pops


def to_postfix(text, trace=False):
    """The postfix form of an infix expression. Prints a trace if asked."""
    tokens = tokenize(text)
    output = []
    stack = []
    rows = []

    def note(token, action):
        rows.append((token, action, " ".join(stack), " ".join(output)))

    for token in tokens:
        if is_operand(token):
            output.append(token)
            note(token, "operand, straight to the output")
        elif token == "(":
            stack.append(token)
            note(token, "push the bracket")
        elif token == ")":
            popped = []
            while stack and stack[-1] != "(":
                popped.append(stack.pop())
                output.append(popped[-1])
            if not stack:
                raise ValueError("a ')' with no matching '('")
            stack.pop()
            note(token, f"pop to the '(': {' '.join(popped) or 'nothing'}")
        else:
            popped = []
            while stack and pops_first(stack[-1], token):
                popped.append(stack.pop())
                output.append(popped[-1])
            stack.append(token)
            if popped:
                note(token, f"pop {' '.join(popped)} first, then push {token}")
            else:
                note(token, f"push {token}")

    while stack:
        top = stack.pop()
        if top == "(":
            raise ValueError("a '(' that is never closed")
        output.append(top)
    note("end", "pop the rest of the stack")

    if trace:
        print(f"  {'token':<7} {'stack':<12} {'output':<24} what happened")
        for token, action, stack_now, output_now in rows:
            print(f"  {token:<7} {stack_now:<12} {output_now:<24} {action}")

    return " ".join(output)


def evaluate_postfix(text):
    """The value of a postfix expression, using a stack."""
    stack = []
    for token in tokenize(text):
        if is_operand(token):
            stack.append(float(token))
            continue
        if len(stack) < 2:
            raise ValueError(f"the operator {token!r} has too few operands")
        right = stack.pop()
        left = stack.pop()
        if token == "+":
            stack.append(left + right)
        elif token == "-":
            stack.append(left - right)
        elif token == "*":
            stack.append(left * right)
        elif token == "/":
            if right == 0:
                raise ZeroDivisionError("division by zero in the expression")
            stack.append(left / right)
        elif token == "%":
            stack.append(left % right)
        elif token == "^":
            stack.append(left ** right)
    if len(stack) != 1:
        raise ValueError("the expression left more than one value on the stack")
    return stack[0]
munotes.in168

Practical 3 continued: Infix to Postfix with a Stack

Four things in that file are the marks.

pops_first is the whole algorithm, and it has four cases: an opening bracket never pops, higher precedence pops, lower precedence does not, and equal precedence pops only when the arriving operator is left associative.

( is never popped by rule 4. It is only removed by its own closing bracket, which is what if on_stack == "(" protects.

In evaluate_postfix the SECOND value popped is the LEFT operand. right = stack.pop() comes first. Getting that round the wrong way gives the right answer for + and * and the wrong answer for -, /, % and ^, which is exactly the bug that survives testing on addition.

The unbalanced cases raise, both a closer with nothing open and an opener never closed.

A conversion traced, which is the journal entry

from postfix import to_postfix

expression = "a + b * c - d"
print(f"converting {expression!r}")
answer = to_postfix(expression, trace=True)
print(f"  postfix: {answer}")
converting 'a + b * c - d'
  token   stack        output                   what happened
  a                    a                        operand, straight to the output
  +       +            a                        push +
  b       +            a b                      operand, straight to the output
  *       + *          a b                      push *
  c       + *          a b c                    operand, straight to the output
  -       -            a b c * +                pop * + first, then push -
  d       -            a b c * + d              operand, straight to the output
  end                  a b c * + d -            pop the rest of the stack
  postfix: a b c * + d -

Read the stack column. The + was pushed and stayed there while arrived, because binds tighter; then - arrived, which does not bind tighter than +, so * and + both came out before - was pushed. That is rule 4 happening, and copying that table into the journal is what the entry is for.

munotes.in169

Practical 3 continued: Infix to Postfix with a Stack

Brackets, which is the other half

from postfix import to_postfix

expression = "( a + b ) * c"
print(f"converting {expression!r}")
print(f"  postfix: {to_postfix(expression, trace=True)}")
converting '( a + b ) * c'
  token   stack        output                   what happened
  (       (                                     push the bracket
  a       (            a                        operand, straight to the output
  +       ( +          a                        push +
  b       ( +          a b                      operand, straight to the output
  )                    a b +                    pop to the '(': +
  *       *            a b +                    push *
  c       *            a b + c                  operand, straight to the output
  end                  a b + c *                pop the rest of the stack
  postfix: a b + c *

The brackets appear in neither the output nor the final stack. They exist only to hold the + back until the ) arrives, and then they are discarded. Brackets are never part of postfix, which is the whole point of the notation.

The right associative power operator

from postfix import to_postfix

for expression in ["2 ^ 3 ^ 2", "2 - 3 - 4", "a ^ b ^ c", "a - b - c"]:
    print(f"{expression:<12} -> {to_postfix(expression)}")

print()
print("read them against each other:")
print("  2 ^ 3 ^ 2 -> 2 3 2 ^ ^   the RIGHT ^ is applied first, so 3^2 then 2^9")
print("  2 - 3 - 4 -> 2 3 - 4 -   the LEFT - is applied first, so (2-3) then -4")
2 ^ 3 ^ 2    -> 2 3 2 ^ ^
2 - 3 - 4    -> 2 3 - 4 -
a ^ b ^ c    -> a b c ^ ^
a - b - c    -> a b - c -

read them against each other:
  2 ^ 3 ^ 2 -> 2 3 2 ^ ^   the RIGHT ^ is applied first, so 3^2 then 2^9
  2 - 3 - 4 -> 2 3 - 4 -   the LEFT - is applied first, so (2-3) then -4

Those two lines are the whole difference between left and right associativity, and they are the pair to put in the journal. In 2 3 2 ^ ^ the first ^ reached is applied to 3 and 2; in 2 3 - 4 - the first - is applied to 2 and 3.

Every conversion evaluated, which is what proves it

A conversion that looks right can still be wrong. The only way to be sure is to evaluate the postfix and compare it with the value of the infix.

from postfix import to_postfix, evaluate_postfix

cases = [
    "2 + 3",
    "2 + 3 * 4",
    "2 * 3 + 4",
    "( 2 + 3 ) * 4",
    "2 * ( 3 + 4 )",
    "10 - 4 - 3",
    "10 - ( 4 - 3 )",
    "2 ^ 3 ^ 2",
    "( 2 ^ 3 ) ^ 2",
    "100 / 5 / 2",
    "100 / ( 5 / 2 )",
    "2 + 3 * 4 - 5 / 5",
    "( ( 2 + 3 ) * ( 4 - 1 ) ) ^ 2",
    "17 % 5 + 1",
]

print(f"{'infix':<32} {'postfix':<26} {'value':>12} {'Python':>12}  same?")
for infix in cases:
    postfix = to_postfix(infix)
    mine = evaluate_postfix(postfix)
    theirs = eval(infix.replace("^", "**"))
    agree = abs(mine - theirs) < 1e-9
    print(f"{infix:<32} {postfix:<26} {mine:>12.4f} {theirs:>12.4f}  {agree}")
munotes.in170

Practical 3 continued: Infix to Postfix with a Stack

infix                            postfix                           value       Python  same?
2 + 3                            2 3 +                            5.0000       5.0000  True
2 + 3 * 4                        2 3 4 * +                       14.0000      14.0000  True
2 * 3 + 4                        2 3 * 4 +                       10.0000      10.0000  True
( 2 + 3 ) * 4                    2 3 + 4 *                       20.0000      20.0000  True
2 * ( 3 + 4 )                    2 3 4 + *                       14.0000      14.0000  True
10 - 4 - 3                       10 4 - 3 -                       3.0000       3.0000  True
10 - ( 4 - 3 )                   10 4 3 - -                       9.0000       9.0000  True
2 ^ 3 ^ 2                        2 3 2 ^ ^                      512.0000     512.0000  True
( 2 ^ 3 ) ^ 2                    2 3 ^ 2 ^                       64.0000      64.0000  True
100 / 5 / 2                      100 5 / 2 /                     10.0000      10.0000  True
100 / ( 5 / 2 )                  100 5 2 / /                     40.0000      40.0000  True
2 + 3 * 4 - 5 / 5                2 3 4 * + 5 5 / -               13.0000      13.0000  True
( ( 2 + 3 ) * ( 4 - 1 ) ) ^ 2    2 3 + 4 1 - * 2 ^              225.0000     225.0000  True
17 % 5 + 1                       17 5 % 1 +                       3.0000       3.0000  True

Every row agrees, and the last column is the point: the conversion is not being proof-read, it is being checked against Python's own arithmetic on the original infix expression. Fourteen expressions, including both bracketings of the power operator and both of the division, and every value matches.

eval is used here only to produce a second opinion inside a test on a string this program wrote itself. Never use eval on text that came from a user, because it runs whatever it is given. That is worth a sentence in the journal, because it is the one place this chapter uses a dangerous tool and it should be seen to be used carefully.

Evaluation traced

from postfix import tokenize, is_operand

def evaluate_traced(text):
    stack = []
    print(f"  {'token':<7} {'action':<34} stack after")
    for token in tokenize(text):
        if is_operand(token):
            stack.append(float(token))
            print(f"  {token:<7} {'push the operand':<34} {stack}")
            continue
        right = stack.pop()
        left = stack.pop()
        value = {"+": left + right, "-": left - right,
                 "*": left * right, "/": left / right,
                 "^": left ** right}[token]
        stack.append(value)
        print(f"  {token:<7} {f'{left} {token} {right} = {value}':<34} {stack}")
    return stack[0]


print("evaluating 2 3 4 * + 5 -")
print("  the answer is", evaluate_traced("2 3 4 * + 5 -"))
munotes.in171

Practical 3 continued: Infix to Postfix with a Stack

evaluating 2 3 4 * + 5 -
  token   action                             stack after
  2       push the operand                   [2.0]
  3       push the operand                   [2.0, 3.0]
  4       push the operand                   [2.0, 3.0, 4.0]
  *       3.0 * 4.0 = 12.0                   [2.0, 12.0]
  +       2.0 + 12.0 = 14.0                  [14.0]
  5       push the operand                   [14.0, 5.0]
  -       14.0 - 5.0 = 9.0                   [9.0]
  the answer is 9.0

Read the token column. Operands are pushed and an operator pops two and pushes one, so the stack shrinks by one each time an operator is reached. The expression is valid exactly when the stack holds one value at the end, which is how evaluate_postfix detects a malformed expression.

The bug that survives testing on addition

from postfix import tokenize, is_operand


def evaluate_wrong(text):
    """Pops the operands in the WRONG order."""
    stack = []
    for token in tokenize(text):
        if is_operand(token):
            stack.append(float(token))
            continue
        left = stack.pop()       # WRONG: the first pop is the RIGHT operand
        right = stack.pop()
        stack.append({"+": left + right, "-": left - right,
                      "*": left * right, "/": left / right}[token])
    return stack[0]


from postfix import evaluate_postfix

print(f"{'postfix':<14} {'correct':>10} {'wrong order':>13}  same?")
for text in ["2 3 +", "2 3 *", "10 4 -", "100 5 /"]:
    right_way = evaluate_postfix(text)
    wrong_way = evaluate_wrong(text)
    print(f"{text:<14} {right_way:>10.4f} {wrong_way:>13.4f}  "
          f"{abs(right_way - wrong_way) < 1e-9}")
postfix           correct   wrong order  same?
2 3 +              5.0000        5.0000  True
2 3 *              6.0000        6.0000  True
10 4 -             6.0000       -6.0000  False
100 5 /           20.0000        0.0500  False

There it is. The wrong order is correct for + and * and wrong for - and /, because addition and multiplication do not care which way round their operands are and subtraction and division do. A student who tests only on 2 3 + ships the bug.

right = stack.pop() first, then left = stack.pop(). The second value out is the left operand, because it went in first.

Unbalanced expressions

from postfix import to_postfix

for expression in ["( a + b", "a + b )", "( ( a )", "a + b"]:
    try:
        print(f"{expression!r:<12} -> {to_postfix(expression)}")
    except ValueError as error:
        print(f"{expression!r:<12} -> refused: {error}")
'( a + b'    -> refused: a '(' that is never closed
'a + b )'    -> refused: a ')' with no matching '('
'( ( a )'    -> refused: a '(' that is never closed
'a + b'      -> a b +
munotes.in172

Practical 3 continued: Infix to Postfix with a Stack

Both faults are caught, and they are caught in different places: a closing bracket with nothing open is found when the stack runs out during rule 3, and an opening bracket never closed is found when rule 5 empties the stack and meets a (.

Prefix to postfix, if the examiner asks for it instead

MU's row says infix, and an examiner may ask for prefix, so here it is: read the prefix expression right to left, push operands, and when an operator is reached pop two and push them followed by the operator.

from postfix import tokenize, is_operand, evaluate_postfix


def prefix_to_postfix(text):
    """Read right to left, pushing partial postfix strings."""
    stack = []
    for token in reversed(tokenize(text)):
        if is_operand(token):
            stack.append(token)
        else:
            first = stack.pop()
            second = stack.pop()
            stack.append(f"{first} {second} {token}")
    return stack.pop()


for prefix, expected in [("+ a b", "a b +"),
                         ("+ a * b c", "a b c * +"),
                         ("* + a b c", "a b + c *"),
                         ("- + 2 3 4", "2 3 + 4 -")]:
    got = prefix_to_postfix(prefix)
    print(f"{prefix:<14} -> {got:<14} expected {expected:<14} {got == expected}")

print("and the last one evaluates to", evaluate_postfix(prefix_to_postfix("- + 2 3 4")))
+ a b          -> a b +          expected a b +          True
+ a * b c      -> a b c * +      expected a b c * +      True
* + a b c      -> a b + c *      expected a b + c *      True
- + 2 3 4      -> 2 3 + 4 -      expected 2 3 + 4 -      True
and the last one evaluates to 1.0

It is shorter than the infix conversion because prefix, like postfix, needs no precedence rules at all: the shape already says what applies to what.

Procedure

  1. Save postfix.py with PRECEDENCE, RIGHT_ASSOCIATIVE, tokenize, pops_first, to_postfix

and evaluate_postfix.

  1. Write pops_first with its four cases, including the opening bracket and the right associative

exception.

  1. Make to_postfix able to print a trace table of token, stack, output and what happened.
  2. Raise for a closing bracket with nothing open and for an opening bracket never closed.
  3. In evaluate_postfix, pop the right operand first and the left second.
  4. Convert a + b * c - d with the trace on and copy the table out.
  5. Convert ( a + b ) * c with the trace on and note that the brackets appear nowhere in the output.
  6. Convert 2 ^ 3 ^ 2 and 2 - 3 - 4 and say why the two differ.
  7. Convert and then evaluate at least ten expressions, comparing each value with the value of the
munotes.in173

Practical 3 continued: Infix to Postfix with a Stack

infix, and record that every one agrees.

  1. Write the wrong operand order and record that it is correct for + and * and wrong for - and

/.

Result

a + b c - d converted to the expected postfix, and the trace showed and + both leaving the stack before - was pushed. ( a + b ) c converted with no bracket in the output. 2 ^ 3 ^ 2 gave 2 3 2 ^ ^ and 2 - 3 - 4 gave 2 3 - 4 -, which is the right and left associative difference. Fourteen expressions were converted and evaluated, and every value agreed with Python's own value of the infix expression to within a billionth. The traced evaluation showed the stack shrinking by one at every operator and holding exactly one value at the end. Popping the operands in the wrong order gave the correct answer for + and and the wrong answer for - and /. Both unbalanced bracket cases were refused.

Where marks are lost

  • No trace table. The table is the answer to this question; the postfix string alone is a

fraction of it.

  • Treating ^ as left associative. 2 ^ 3 ^ 2 then converts to 2 3 ^ 2 ^, which is wrong.
  • Popping the ( in rule 4. It must only be removed by its own ).
  • Outputting the brackets. Postfix has none.
  • Forgetting rule 5, so the last operators never reach the output.
  • Popping the operands in the wrong order when evaluating, which passes every test on + and

*.

  • Not evaluating the result at all. Converting and evaluating is what proves the conversion.
  • Single character tokens only, so 12 + 345 is read as five one digit operands.
  • No check for unbalanced brackets.

For the journal

The aim in MU's words. The three notation table, and the two lines showing that postfix needs neither brackets nor precedence. The precedence and associativity table, with ^ marked right associative. The six rules. Then postfix.py, and then the trace table for at least two expressions, one with an operator precedence decision and one with brackets, because those two tables are the entry. Then the 2 ^ 3 ^ 2 against 2 - 3 - 4 pair with one sentence on associativity. Then the table of conversions with their evaluated values beside Python's own values, and one sentence: the conversion is proved by evaluating it, not by reading it. The conclusion: the stack holds operators until an operator of lower or equal precedence arrives, brackets are discarded, and the operand popped second is the left one.

munotes.in174

Practical 3 continued: Infix to Postfix with a Stack

Quick revision

  • Infix a + b, prefix + a b, postfix a b +.

Postfix needs no brackets and no precedence rules.

  • Precedence: ^ 3, * / % 2, + - 1.
  • Associativity: everything left except ^, which is right.
  • Rule 1 operand to the output. Rule 2 push (. Rule 3 on ) pop to the ( and discard both.

Rule 4 pop while the top should come first, then push. Rule 5 pop the rest at the end.

  • "Should come first": higher precedence pops. Equal precedence pops only if the arriving

operator is left associative, and ( never pops.

  • 2 ^ 3 ^ 2 gives 2 3 2 ^ ^. 2 - 3 - 4 gives 2 3 - 4 -.
  • Evaluating postfix: push operands; on an operator pop two and push one, so the stack

shrinks by one.

  • right = pop() first, then left = pop(). The wrong order is right for + and * and wrong

for -, /, %, ^.

  • Valid exactly when one value is left on the stack at the end.
  • Unbalanced: a ) when the stack has no (, or a ( still there after rule 5.
  • Never eval text from a user.

Questions you should be able to answer

1. Why does postfix need no brackets? Because the position of the operator already says which operands it applies to, so there is nothing left ambiguous for brackets or precedence rules to settle.

2. Convert a + b c and (a + b) c. a b c + and a b + c .

3. State rule 4. When an operator arrives, pop to the output every operator on the stack that should come out first, then push the new one.

4. What does "should come out first" mean exactly? Higher precedence on the stack pops; equal precedence pops only when the arriving operator is left associative; an opening bracket never pops.

5. Which operator is right associative, and what difference does it make? ^. 2 ^ 3 ^ 2 becomes 2 3 2 ^ ^, so the right hand power is applied first. Treating it as left associative gives 2 3 ^ 2 ^, which is a different value.

6. What happens to the brackets? They are never output. An opening bracket is pushed and removed by its matching closing bracket, and both are discarded.

7. How do you evaluate postfix? Read left to right. Push operands. On an operator, pop two, apply it, and push the result. One value should remain.

8. Which popped value is the left operand? The second one popped, because it was pushed first. So right = pop() then left = pop().

munotes.in175

Practical 3 continued: Infix to Postfix with a Stack

9. Why does the wrong operand order pass most tests? Because addition and multiplication give the same answer either way round. It fails on subtraction, division, remainder and power.

10. How do you know a conversion is correct rather than plausible? Evaluate the postfix and compare the value with the value of the original infix expression. Doing it for fourteen expressions is what this chapter does.

11. How are the two unbalanced cases detected? A closing bracket with nothing open empties the stack during rule 3; an opening bracket never closed is still on the stack at rule 5.

Contents This chapter on its own page

munotes.in176

Chapter Twenty-Seven

Practical 4: a Queue over an Array

Syllabus topic Module 2, practical 4(a), "Queue Application: Write a program to: Implement a queue using an array"

Aim

To implement a queue using an array, with enqueue, dequeue and peek, and to show why a circular queue is needed.

What a queue is

A queue is a collection in which the first thing put in is the first thing taken out: FIFO, first in first out. That is the opposite of a stack, and it is the only difference.

enqueue 10, 20, 30:

     front                  rear
       v                      v
     [ 10 | 20 | 30 |  . |  . ]

dequeue gives 10, because it went in first.

Two ends are used, not one. Items enter at the rear and leave from the front.

StackQueue
RuleLIFO, last in first outFIFO, first in first out
Ends usedone, the toptwo, the front and the rear
Addpush at the topenqueue at the rear
Removepop from the topdequeue from the front
Lookpeek at the toppeek at the front
Forundo, brackets, recursionwaiting lines, scheduling, breadth first search

The linear queue, and why it is not enough

Keep two indexes: front, the index of the first item, and rear, the index of the last.

"""Queues over a fixed array, for Major Practical 3, Module 2."""


class QueueFull(Exception):
    """Enqueueing onto a full queue."""


class QueueEmpty(Exception):
    """Dequeueing or peeking an empty queue."""


class LinearQueue:
    """The obvious queue: front and rear only ever move forward.

    WARNING: this is the version with the defect. It is here to be measured,
    not used.
    """

    def __init__(self, capacity=5):
        self.capacity = capacity
        self.slots = [None] * capacity
        self.front = 0
        self.rear = -1
        self.count = 0

    def is_empty(self):
        return self.count == 0

    def is_full(self):
        return self.rear == self.capacity - 1     # the defect is in this line

    def __len__(self):
        return self.count

    def __repr__(self):
        cells = " | ".join("  . " if v is None else f"{v:>4}" for v in self.slots)
        return (f"[{cells}]  front {self.front}, rear {self.rear}, "
                f"{self.count} item(s)")

    def enqueue(self, item):
        if self.is_full():
            raise QueueFull(
                f"the queue reports itself full: rear is at {self.rear}")
        self.rear += 1
        self.slots[self.rear] = item
        self.count += 1

    def dequeue(self):
        if self.is_empty():
            raise QueueEmpty("dequeue from an empty queue")
        item = self.slots[self.front]
        self.slots[self.front] = None
        self.front += 1
        self.count -= 1
        return item

    def peek(self):
        if self.is_empty():
            raise QueueEmpty("peek at an empty queue")
        return self.slots[self.front]


class CircularQueue:
    """The cure: front and rear wrap round with the modulo operator.

    A count is kept, so full and empty are told apart without wasting a slot.
    """

    def __init__(self, capacity=5):
        if capacity <= 0:
            raise ValueError("capacity must be at least 1")
        self.capacity = capacity
        self.slots = [None] * capacity
        self.front = 0
        self.rear = -1
        self.count = 0

    def is_empty(self):
        return self.count == 0

    def is_full(self):
        return self.count == self.capacity

    def __len__(self):
        return self.count

    def __iter__(self):
        for i in range(self.count):
            yield self.slots[(self.front + i) % self.capacity]

    def __repr__(self):
        cells = " | ".join("  . " if v is None else f"{v:>4}" for v in self.slots)
        order = ", ".join(str(item) for item in self) or "nothing"
        return (f"[{cells}]  front {self.front}, rear {self.rear}, "
                f"{self.count}/{self.capacity}  order: {order}")

    def enqueue(self, item):
        if self.is_full():
            raise QueueFull(f"the queue is genuinely full at {self.capacity}")
        self.rear = (self.rear + 1) % self.capacity
        self.slots[self.rear] = item
        self.count += 1

    def dequeue(self):
        if self.is_empty():
            raise QueueEmpty("dequeue from an empty queue")
        item = self.slots[self.front]
        self.slots[self.front] = None
        self.front = (self.front + 1) % self.capacity
        self.count -= 1
        return item

    def peek(self):
        if self.is_empty():
            raise QueueEmpty("peek at an empty queue")
        return self.slots[self.front]
munotes.in177

Practical 4: a Queue over an Array

The defect, shown happening

from queue_array import LinearQueue, QueueFull

queue = LinearQueue(capacity=5)
print("new              ", queue)

for value in [10, 20, 30, 40, 50]:
    queue.enqueue(value)
print("five enqueued    ", queue)

print("dequeue          ", queue.dequeue(), "->", queue)
print("dequeue          ", queue.dequeue(), "->", queue)
print("dequeue          ", queue.dequeue(), "->", queue)

print()
print("three slots at the front are now empty, and there are only two items.")
try:
    queue.enqueue(60)
except QueueFull as error:
    print("but enqueue(60) fails:", error)
print("is_full says     ", queue.is_full(), "with only", len(queue), "item(s) in it")
new               [  .  |   .  |   .  |   .  |   . ]  front 0, rear -1, 0 item(s)
five enqueued     [  10 |   20 |   30 |   40 |   50]  front 0, rear 4, 5 item(s)
dequeue           10 -> [  .  |   20 |   30 |   40 |   50]  front 1, rear 4, 4 item(s)
dequeue           20 -> [  .  |   .  |   30 |   40 |   50]  front 2, rear 4, 3 item(s)
dequeue           30 -> [  .  |   .  |   .  |   40 |   50]  front 3, rear 4, 2 item(s)

three slots at the front are now empty, and there are only two items.
but enqueue(60) fails: the queue reports itself full: rear is at 4
is_full says      True with only 2 item(s) in it

There is the defect, in one run. Three of five slots are free and the queue refuses to take anything. rear has reached the end of the array and the linear queue has no way to go back to the beginning.

The two bad cures are worth naming, because both get offered:

CureWhy it is bad
shift everything down after each dequeuedequeue becomes O(n) instead of O(1)
use a bigger arrayit only postpones the same failure

The real cure is to let the indexes wrap round. That is a circular queue.

The circular queue

capacity 5, after enqueueing 5 and dequeueing 3 and enqueueing 2 more:

index      0     1     2     3     4
        [ 60 |  70 |  .  |  40 |  50 ]
                            ^front      rear = 1

the order is 40, 50, 60, 70:  front, then (front+1) % 5, and so on
munotes.in178

Practical 4: a Queue over an Array

The array has no beginning and no end any more. Moving on from index 4 gives index 0, which is what (i + 1) % capacity does.

capacity = 5
print("moving forward from each index, with % 5:")
for i in range(capacity):
    print(f"  ({i} + 1) % {capacity} = {(i + 1) % capacity}")
print()
print("so index 4 is followed by index 0, and the array is a ring")
moving forward from each index, with % 5:
  (0 + 1) % 5 = 1
  (1 + 1) % 5 = 2
  (2 + 1) % 5 = 3
  (3 + 1) % 5 = 4
  (4 + 1) % 5 = 0

so index 4 is followed by index 0, and the array is a ring

The circular queue run

from queue_array import CircularQueue, QueueFull, QueueEmpty

queue = CircularQueue(capacity=5)
print("new              ", queue)

for value in [10, 20, 30, 40, 50]:
    queue.enqueue(value)
print("five enqueued    ", queue)

for _ in range(3):
    print(f"dequeue {queue.dequeue():<4}     ", queue)

print()
print("now the same enqueue that the linear queue refused:")
queue.enqueue(60)
print("enqueue 60       ", queue)
queue.enqueue(70)
print("enqueue 70       ", queue, "  <- rear has WRAPPED to index 1")
queue.enqueue(80)
print("enqueue 80       ", queue, "  full again, and honestly so")

try:
    queue.enqueue(90)
except QueueFull as error:
    print("enqueue on full  ", error)

print()
print("and the order is still correct, first in first out:")
while not queue.is_empty():
    print(f"  dequeue {queue.dequeue()}")
print("emptied          ", queue)

try:
    queue.dequeue()
except QueueEmpty as error:
    print("dequeue on empty ", error)
new               [  .  |   .  |   .  |   .  |   . ]  front 0, rear -1, 0/5  order: nothing
five enqueued     [  10 |   20 |   30 |   40 |   50]  front 0, rear 4, 5/5  order: 10, 20, 30, 40, 50
dequeue 10        [  .  |   20 |   30 |   40 |   50]  front 1, rear 4, 4/5  order: 20, 30, 40, 50
dequeue 20        [  .  |   .  |   30 |   40 |   50]  front 2, rear 4, 3/5  order: 30, 40, 50
dequeue 30        [  .  |   .  |   .  |   40 |   50]  front 3, rear 4, 2/5  order: 40, 50

now the same enqueue that the linear queue refused:
enqueue 60        [  60 |   .  |   .  |   40 |   50]  front 3, rear 0, 3/5  order: 40, 50, 60
enqueue 70        [  60 |   70 |   .  |   40 |   50]  front 3, rear 1, 4/5  order: 40, 50, 60, 70   <- rear has WRAPPED to index 1
enqueue 80        [  60 |   70 |   80 |   40 |   50]  front 3, rear 2, 5/5  order: 40, 50, 60, 70, 80   full again, and honestly so
enqueue on full   the queue is genuinely full at 5

and the order is still correct, first in first out:
  dequeue 40
  dequeue 50
  dequeue 60
  dequeue 70
  dequeue 80
emptied           [  .  |   .  |   .  |   .  |   . ]  front 3, rear 2, 0/5  order: nothing
dequeue on empty  dequeue from an empty queue
munotes.in179

Practical 4: a Queue over an Array

Read the rear value on the line marked. It went from 4 to 0 to 1, wrapping round the end of the array, and the order of the items coming out is still 40, 50, 60, 70, 80: first in, first out, even though they are not in that order in the array.

Full against empty: the ambiguity and the two cures

This is the part of the exercise an examiner is most likely to probe.

When front and rear are the only state, a circular queue cannot tell full from empty. In both cases front and rear sit next to each other in the same way.

EMPTY:  nothing in it, front == (rear + 1) % capacity
FULL:   every slot used, front == (rear + 1) % capacity

the same condition. Two different states.

There are exactly two standard cures.

Cure one: keep a count. That is what the class above does, and count == 0 means empty while count == capacity means full. It costs one integer and it wastes no slots.

Cure two: waste one slot. Never fill the last free slot, so full becomes (rear + 2) % capacity == front and the two states are distinguishable without a count.

class WastesOneSlot:
    """The count free version: full when one slot is deliberately left empty."""

    def __init__(self, capacity):
        self.capacity = capacity
        self.slots = [None] * capacity
        self.front = 0
        self.rear = 0          # rear is the NEXT free slot here

    def is_empty(self):
        return self.front == self.rear

    def is_full(self):
        return (self.rear + 1) % self.capacity == self.front

    def usable(self):
        return self.capacity - 1

    def enqueue(self, item):
        if self.is_full():
            raise OverflowError("full")
        self.slots[self.rear] = item
        self.rear = (self.rear + 1) % self.capacity

    def dequeue(self):
        if self.is_empty():
            raise IndexError("empty")
        item = self.slots[self.front]
        self.slots[self.front] = None
        self.front = (self.front + 1) % self.capacity
        return item


queue = WastesOneSlot(capacity=5)
print("capacity 5, usable", queue.usable(), "because one slot is always kept free")
for value in [10, 20, 30, 40]:
    queue.enqueue(value)
print("after 4 enqueues, is_full:", queue.is_full(), "and the array is", queue.slots)
try:
    queue.enqueue(50)
except OverflowError:
    print("the fifth is refused, even though slot 4 is", queue.slots[4])
print("empty and full are now different conditions, with no count kept")
capacity 5, usable 4 because one slot is always kept free
after 4 enqueues, is_full: True and the array is [10, 20, 30, 40, None]
the fifth is refused, even though slot 4 is None
empty and full are now different conditions, with no count kept
munotes.in180

Practical 4: a Queue over an Array

CureCostsUsable slots
keep a countone integerall of them
waste one slotone slotcapacity - 1

Keep a count. An integer is cheaper than a slot when the items are anything bigger than an integer, and the code reads better. But be able to describe the other one, because textbooks teach it and examiners ask for it.

What the operations cost

from queue_array import CircularQueue


def cost_of_n(n):
    """Every enqueue and dequeue is one step, whatever n is."""
    queue = CircularQueue(capacity=n)
    steps = 0
    for value in range(n):
        queue.enqueue(value)
        steps += 1
    while not queue.is_empty():
        queue.dequeue()
        steps += 1
    return steps


print(f"{'n':>7} {'operations':>12} {'steps':>8} {'steps per operation':>21}")
for n in [10, 100, 1000]:
    steps = cost_of_n(n)
    print(f"{n:>7} {2 * n:>12} {steps:>8} {steps / (2 * n):>21.1f}")
print()
print("one step per operation at every size, which is O(1)")
      n   operations    steps   steps per operation
     10           20       20                   1.0
    100          200      200                   1.0
   1000         2000     2000                   1.0

one step per operation at every size, which is O(1)
OperationCost
enqueueO(1)
dequeueO(1)
peekO(1)
is_empty, is_full, lenO(1)
searching the queueO(n), and a queue is not for searching

The wrong way, which is what a Python list tempts you into

import timeit
from collections import deque


def with_list_pop_zero(n):
    items = []
    for value in range(n):
        items.append(value)
    while items:
        items.pop(0)          # O(n): every remaining item shifts down


def with_deque(n):
    items = deque()
    for value in range(n):
        items.append(value)
    while items:
        items.popleft()       # O(1)


for n in [2000, 4000, 8000, 16000]:
    runs = max(3, 40000 // n)
    slow = timeit.timeit(lambda: with_list_pop_zero(n), number=runs) / runs
    fast = timeit.timeit(lambda: with_deque(n), number=runs) / runs
    print(f"n = {n}")
    print(f"  list with pop(0), ms: {slow * 1000:.2f}")
    print(f"  deque popleft, ms:    {fast * 1000:.2f}")
    print(f"  the list was slower by a factor of about {slow / fast:.0f}")

print()
print("the ratio GROWS with n, and that is the finding: pop(0) shifts every")
print("remaining item down by one, so emptying n items costs n(n-1)/2 moves")
print("in all, against n for the deque. Quadratic against linear.")
n = 2000
  list with pop(0), ms: 0.67
  deque popleft, ms:    0.18
  the list was slower by a factor of about 4
n = 4000
  list with pop(0), ms: 2.01
  deque popleft, ms:    0.55
  the list was slower by a factor of about 4
n = 8000
  list with pop(0), ms: 8.78
  deque popleft, ms:    0.91
  the list was slower by a factor of about 10
n = 16000
  list with pop(0), ms: 31.28
  deque popleft, ms:    1.53
  the list was slower by a factor of about 20

the ratio GROWS with n, and that is the finding: pop(0) shifts every
remaining item down by one, so emptying n items costs n(n-1)/2 moves
in all, against n for the deque. Quadratic against linear.
munotes.in181

Practical 4: a Queue over an Array

list.pop(0) is the classic wrong answer to this exercise. It works, it is one line, and it turns an O(1) operation into an O(n) one, so a queue built on it is quadratic overall. The shift counting in [Practical 1: Array Operations, Insert, Delete and Linear Search] is exactly this cost.

collections.deque is a double ended queue in the standard library, with append, appendleft, pop and popleft all O(1). It is what you use in real work. MU's row asks for the array version, so write the array version and mention the deque.

The deque, and the other two things it gives you

from collections import deque

queue = deque([10, 20, 30])
print("as a queue   :", queue)
print("  append 40  :", (queue.append(40), queue)[1])
print("  popleft    :", queue.popleft(), "->", queue)
print()

both_ends = deque([20, 30])
both_ends.appendleft(10)
both_ends.append(40)
print("both ends    :", both_ends)
print("  pop right  :", both_ends.pop(), "->", both_ends)
print("  pop left   :", both_ends.popleft(), "->", both_ends)
print()

bounded = deque(maxlen=3)
for value in [1, 2, 3, 4, 5]:
    bounded.append(value)
    print(f"  append {value} with maxlen=3 ->", bounded)
print("a bounded deque throws away from the FRONT, which is the last n lines idea")

ring = deque([1, 2, 3, 4, 5])
ring.rotate(2)
print("rotate(2)    :", ring)
ring.rotate(-2)
print("rotate(-2)   :", ring)
as a queue   : deque([10, 20, 30])
  append 40  : deque([10, 20, 30, 40])
  popleft    : 10 -> deque([20, 30, 40])

both ends    : deque([10, 20, 30, 40])
  pop right  : 40 -> deque([10, 20, 30])
  pop left   : 10 -> deque([20, 30])

  append 1 with maxlen=3 -> deque([1], maxlen=3)
  append 2 with maxlen=3 -> deque([1, 2], maxlen=3)
  append 3 with maxlen=3 -> deque([1, 2, 3], maxlen=3)
  append 4 with maxlen=3 -> deque([2, 3, 4], maxlen=3)
  append 5 with maxlen=3 -> deque([3, 4, 5], maxlen=3)
a bounded deque throws away from the FRONT, which is the last n lines idea
rotate(2)    : deque([4, 5, 1, 2, 3])
rotate(-2)   : deque([1, 2, 3, 4, 5])

maxlen is the same idea as the last n lines of a file in [Practical 8 continued: Text Files, Binary Files, and the Last n Lines], and rotate is the circular queue's wrap-around offered as a method.

Procedure

  1. Save queue_array.py with QueueFull and QueueEmpty, LinearQueue and CircularQueue.
  2. In both, keep capacity, slots, front, rear and count. Start front at 0 and rear at

-1.

  1. In LinearQueue, make is_full test rear == capacity - 1, which is the defect.
  2. Fill the linear queue, dequeue three, then enqueue again and record the refusal with the number
munotes.in182

Practical 4: a Queue over an Array

of free slots.

  1. In CircularQueue, advance both indexes with (i + 1) % capacity and test full with

count == capacity.

  1. Run the same sequence on the circular queue and record that the enqueue succeeds and that rear

wraps to a lower index.

  1. Print the queue after every operation with front, rear and the logical order of the items.
  2. Force both errors: enqueue on full and dequeue on empty.
  3. Write the version that wastes one slot instead of keeping a count, and show that its usable

capacity is one less.

  1. Compare list.pop(0) with deque.popleft() and say why the first is quadratic.

Result

The linear queue took five items into five slots, and after three dequeues it refused a sixth item while three slots stood empty, with is_full returning True for a queue holding two items. The circular queue accepted the same enqueue, and rear wrapped from index 4 to 0 and then to 1 while the items still came out in the order they went in. Enqueueing onto the genuinely full circular queue raised QueueFull and dequeueing the empty one raised QueueEmpty. The version that wastes one slot refused the fifth item of a capacity of five, leaving one slot permanently free, and needed no count. Enqueue and dequeue took one step each at every size tried, confirming O(1). Emptying a list with pop(0) was slower than deque.popleft() at every size tried, and the ratio grew as n grew, which is what distinguishes a quadratic algorithm from a linear one rather than a merely slower constant.

Where marks are lost

  • A linear queue offered as the answer, with no mention of the wasted slots. The defect is the

exercise.

  • Shifting everything after each dequeue, which makes dequeue O(n).
  • Forgetting the % capacity on one of the two indexes, so only one of them wraps.
  • No way to tell full from empty. Keep a count, or waste a slot, and say which.
  • rear starting at 0 instead of -1 without adjusting the first enqueue, which leaves slot 0

empty.

  • list.pop(0), which is O(n) and makes the whole queue quadratic.
  • Not printing front and rear. The wrap-around is invisible without them.
  • Not forcing both errors.
  • Saying a queue is like a stack. It is the opposite: two ends, FIFO.

For the journal

The aim in MU's words. The queue picture with front and rear marked, and the stack against queue table. Then the linear queue's failure, printed: five slots, five items, three dequeued, and the refusal with three slots free. That run is the reason the rest of the entry exists. Then the circular picture with the wrap, the % capacity line, and the circular queue's run with front, rear and the logical order printed after every operation, including the line where rear wraps. Then both errors forced. Then one sentence on full against empty: with only front and rear the two states look identical, so either keep a count or leave one slot unused. The conclusion: a queue is FIFO with two ends, both operations are O(1), and the modulo operator is what stops the array running out of room at one end.

munotes.in183

Practical 4: a Queue over an Array

Quick revision

  • A queue is FIFO. Items enqueue at the rear and dequeue from the front. Two ends.
  • A stack is LIFO with one end. That is the only difference.
  • Linear queue: front and rear only move forward, so it reports itself

full while slots at the front are free.

  • Do not cure it by shifting: that makes dequeue O(n).
  • Circular queue: self.rear = (self.rear + 1) % self.capacity, and the same for front.
  • front starts at 0 and rear at -1, so the first enqueue puts the item in slot 0.
  • Full and empty look identical from front and rear alone. Keep a count (count == 0 empty,

count == capacity full) or waste one slot ((rear + 1) % capacity == front is full).

  • Keeping a count uses every slot; wasting one loses a slot and needs no count.
  • enqueue, dequeue, peek are all O(1).
  • list.pop(0) is O(n) and makes a queue quadratic. collections.deque.popleft() is O(1).
  • A deque also gives appendleft, pop, maxlen and rotate.

Questions you should be able to answer

1. What does FIFO mean, and which ends does a queue use? First in, first out. Items enter at the rear and leave from the front, so both ends are used.

2. What is the defect of a linear queue over an array? rear only moves forward, so once it reaches the last slot the queue reports itself full even though dequeues have freed slots at the front.

3. Why not simply shift everything down after each dequeue? Because that makes dequeue O(n) instead of O(1), which loses the queue's only real advantage.

4. Write the line that makes a queue circular. self.rear = (self.rear + 1) % self.capacity, and the same for self.front.

5. Why can a circular queue not tell full from empty with front and rear alone? Because in both states front sits immediately after rear in the ring, so the condition is identical.

6. Give the two standard cures. Keep a count of the items, so count == 0 is empty and count == capacity is full; or deliberately never use the last free slot, so full is (rear + 1) % capacity == front.

munotes.in184

Practical 4: a Queue over an Array

7. Which cure would you choose and why? Keeping a count, because it uses every slot and one integer is cheaper than a wasted slot for anything but the smallest items. The other is worth knowing because textbooks teach it.

8. What do enqueue and dequeue cost? O(1) each, which the step count in this chapter confirms at every size.

9. Why is a queue built on list.pop(0) quadratic? Because pop(0) shifts every remaining item down one place, so emptying n items costs about n squared over two moves in all.

10. What should you use in real Python, and why did MU ask for the array? collections.deque, whose append and popleft are both O(1). MU asked for the array version because the array is where the wrap-around and the full-against-empty problem can be seen.

Contents This chapter on its own page

munotes.in185

Chapter Twenty-Eight

Practical 4 continued: Simulating a Customer Service Queue

Syllabus topic Module 2, practical 4(b), "Simulate a simple queuing system (e.g., customer service queue)"

Aim

To simulate a simple customer service queue and report the waiting times, the queue length and how many counters are needed.

What a simulation is, and the four things it needs

A simulation steps a clock forward and, at each step, does what would happen in the real system. This one needs four things and nothing else.

In this program
a clockan integer minute, from 0 upwards
arrivalsat each minute, does a customer arrive
a queuethe customers waiting, in the order they arrived
serversone or more counters, each either free or busy until some minute

The queue is FIFO, which is the whole reason this is practical 4. The customer who has waited longest is served next, and nothing else would be fair.

One counter, traced minute by minute

import random
from collections import deque

random.seed(7)                     # so this run is reproducible

MINUTES = 20
ARRIVAL_CHANCE = 0.45              # a customer arrives with this probability each minute
SERVICE_RANGE = (2, 4)             # service takes 2 to 4 minutes

queue = deque()
busy_until = 0
serving = None
next_number = 1
waits = []
served = []

print(f"  {'min':>3} {'arrives':>8} {'queue':<14} {'counter':<22} waiting")
for minute in range(MINUTES):
    arrived = ""
    if random.random() < ARRIVAL_CHANCE:
        queue.append((next_number, minute))
        arrived = f"C{next_number}"
        next_number += 1

    if serving is not None and minute >= busy_until:
        served.append(serving[0])
        serving = None

    counter = "free"
    if serving is None and queue:
        number, arrived_at = queue.popleft()
        service = random.randint(*SERVICE_RANGE)
        busy_until = minute + service
        serving = (number, arrived_at)
        waits.append(minute - arrived_at)
        counter = f"start C{number}, {service} min, waited {minute - arrived_at}"
    elif serving is not None:
        counter = f"busy with C{serving[0]} until {busy_until}"

    waiting = " ".join(f"C{n}" for n, _ in queue) or "-"
    print(f"  {minute:>3} {arrived:>8} {len(queue):<14} {counter:<22} {waiting}")

print()
print(f"customers who arrived  : {next_number - 1}")
print(f"customers served       : {len(served)}")
print(f"still waiting at the end: {len(queue)}")
print(f"waits recorded         : {waits}")
print(f"average wait           : {sum(waits) / len(waits):.2f} minutes")
print(f"longest wait           : {max(waits)} minutes")
  min  arrives queue          counter                waiting
    0       C1 0              start C1, 2 min, waited 0 -
    1       C2 1              busy with C1 until 2   C2
    2       C3 1              start C2, 4 min, waited 1 C3
    3       C4 2              busy with C2 until 6   C3 C4
    4          2              busy with C2 until 6   C3 C4
    5          2              busy with C2 until 6   C3 C4
    6       C5 2              start C3, 2 min, waited 4 C4 C5
    7       C6 3              busy with C3 until 8   C4 C5 C6
    8       C7 3              start C4, 2 min, waited 5 C5 C6 C7
    9          3              busy with C4 until 10  C5 C6 C7
   10       C8 3              start C5, 4 min, waited 4 C6 C7 C8
   11       C9 4              busy with C5 until 14  C6 C7 C8 C9
   12      C10 5              busy with C5 until 14  C6 C7 C8 C9 C10
   13          5              busy with C5 until 14  C6 C7 C8 C9 C10
   14          4              start C6, 4 min, waited 7 C7 C8 C9 C10
   15          4              busy with C6 until 18  C7 C8 C9 C10
   16      C11 5              busy with C6 until 18  C7 C8 C9 C10 C11
   17      C12 6              busy with C6 until 18  C7 C8 C9 C10 C11 C12
   18          5              start C7, 2 min, waited 10 C8 C9 C10 C11 C12
   19      C13 6              busy with C7 until 20  C8 C9 C10 C11 C12 C13

customers who arrived  : 13
customers served       : 6
still waiting at the end: 6
waits recorded         : [0, 1, 4, 5, 4, 7, 10]
average wait           : 4.43 minutes
longest wait           : 10 minutes
munotes.in186

Practical 4 continued: Simulating a Customer Service Queue

Read the last column. It grows, because with a 45 per cent chance of an arrival every minute and a service taking two to four minutes, one counter cannot keep up. That is the finding the simulation exists to produce, and the next section turns it into a number.

Why one counter is not enough, in arithmetic

Before running anything, the answer can be estimated, and doing so is what makes a simulation a check rather than a guess.

arrival_chance = 0.45
service_low, service_high = 2, 4

arrivals_per_minute = arrival_chance
mean_service = (service_low + service_high) / 2
work_per_minute = arrivals_per_minute * mean_service

print(f"customers arriving per minute : {arrivals_per_minute}")
print(f"average minutes to serve one  : {mean_service}")
print(f"counter-minutes of work created each minute: "
      f"{arrivals_per_minute} × {mean_service} = {work_per_minute}")
print()
print(f"one counter supplies 1 counter-minute per minute.")
print(f"needed: {work_per_minute}, supplied: 1, so one counter is short by "
      f"{work_per_minute - 1:.2f}")
print(f"the smallest number of counters that keeps up is "
      f"{int(work_per_minute) + 1 if work_per_minute % 1 else int(work_per_minute)}")
customers arriving per minute : 0.45
average minutes to serve one  : 3.0
counter-minutes of work created each minute: 0.45 × 3.0 = 1.35

one counter supplies 1 counter-minute per minute.
needed: 1.35, supplied: 1, so one counter is short by 0.35
the smallest number of counters that keeps up is 2

That is the whole of queueing theory in five lines. If more work arrives each minute than the counters can do, the queue grows without limit. One counter can do one counter-minute per minute, and this system creates more than that, so the queue must grow. A simulation cannot contradict that arithmetic; it shows how bad it gets and how quickly.

Several counters

import random
from collections import deque


def simulate(counters, minutes=200, arrival_chance=0.45,
             service_range=(2, 4), seed=7):
    """Run the queue and return (average wait, longest wait, left waiting, served)."""
    random.seed(seed)
    queue = deque()
    free_at = [0] * counters
    waits = []
    served = 0
    longest_queue = 0
    next_number = 1

    for minute in range(minutes):
        if random.random() < arrival_chance:
            queue.append(minute)
            next_number += 1
        for i in range(counters):
            if free_at[i] <= minute and queue:
                arrived_at = queue.popleft()
                waits.append(minute - arrived_at)
                free_at[i] = minute + random.randint(*service_range)
                served += 1
        longest_queue = max(longest_queue, len(queue))

    average = sum(waits) / len(waits) if waits else 0.0
    return average, max(waits) if waits else 0, len(queue), served, longest_queue


print(f"  {'counters':>9} {'avg wait':>10} {'worst wait':>12} {'left waiting':>14}"
      f" {'served':>8} {'longest queue':>15}")
for counters in [1, 2, 3, 4]:
    average, worst, left, served, longest = simulate(counters)
    print(f"  {counters:>9} {average:>10.2f} {worst:>12} {left:>14} {served:>8} "
          f"{longest:>15}")
munotes.in187

Practical 4 continued: Simulating a Customer Service Queue

   counters   avg wait   worst wait   left waiting   served   longest queue
          1      32.27           67             26       67              26
          2       0.44            3              0       88               2
          3       0.09            1              0       97               1
          4       0.00            0              0       96               0

Read the table down. One counter is hopeless and two are enough, and the third and fourth buy almost nothing. That is the answer a manager actually wants, and it is why a simulation is written: the arithmetic said "more than one", and the simulation says how much more and what it is worth.

This is MU's Course Objective 8 in its purest form. The structure chosen is a queue, and the justification is that fairness requires the longest waiting customer to be served next, which is exactly FIFO.

The queue length over time, drawn

import random
from collections import deque


def queue_lengths(counters, minutes=60, arrival_chance=0.45,
                  service_range=(2, 4), seed=7):
    random.seed(seed)
    queue = deque()
    free_at = [0] * counters
    lengths = []
    for minute in range(minutes):
        if random.random() < arrival_chance:
            queue.append(minute)
        for i in range(counters):
            if free_at[i] <= minute and queue:
                queue.popleft()
                free_at[i] = minute + random.randint(*service_range)
        lengths.append(len(queue))
    return lengths


for counters in [1, 2]:
    lengths = queue_lengths(counters)
    print(f"{counters} counter(s), the queue length each minute for 60 minutes:")
    for start in range(0, 60, 20):
        row = lengths[start:start + 20]
        bars = "".join(str(min(n, 9)) for n in row)
        print(f"  minute {start:>2} to {start + 19:>2}: {bars}")
    print(f"  ends at {lengths[-1]}, worst {max(lengths)}")
    print()
1 counter(s), the queue length each minute for 60 minutes:
  minute  0 to 19: 01122223333455445656
  minute 20 to 39: 66566666665677678776
  minute 40 to 59: 67666777888777678889
  ends at 9, worst 9

2 counter(s), the queue length each minute for 60 minutes:
  minute  0 to 19: 00001111211000000000
  minute 20 to 39: 00000000000000000000
  minute 40 to 59: 00100000000000000000
  ends at 0, worst 2

With one counter the digits climb and never come back down. With two they stay near zero. That picture, drawn with nothing but digits, is worth more in a journal than a paragraph claiming the same thing.

munotes.in188

Practical 4 continued: Simulating a Customer Service Queue

A priority queue, for the question that follows

The examiner's next question on this exercise is usually "what if some customers are more urgent". That is no longer a queue: it is a priority queue, and the standard library has one.

import heapq

# (priority, arrival order, name). Lower priority number is served first.
counter = []
for priority, name in [(2, "Aarti"), (1, "Bhavesh, urgent"), (3, "Chetan"),
                       (1, "Divya, urgent"), (2, "Eshan")]:
    heapq.heappush(counter, (priority, len(counter), name))

print("served in this order:")
while counter:
    priority, order, name = heapq.heappop(counter)
    print(f"  priority {priority}  {name}")
served in this order:
  priority 1  Bhavesh, urgent
  priority 1  Divya, urgent
  priority 2  Aarti
  priority 2  Eshan
  priority 3  Chetan

The middle item of the tuple, the arrival order, is what keeps it fair inside one priority level: two customers of priority 1 are served in the order they arrived, because tuples compare item by item. Without it, a heap gives no promise about ties.

A heap is a tree, not an array of this kind, and [Practical 5: the Binary Search Tree, Create, Insert and Search] is where trees begin.

QueuePriority queue
Next servedthe one who arrived firstthe most urgent
Built onan array or a dequea heap, which is a tree
Cost of add and removeO(1)O(log n)
Fair toeverybody equallythe urgent, and the rest may wait

The one line that makes this checkable

import random

print("with the seed set, the same numbers every time:")
for _ in range(3):
    random.seed(7)
    print("  ", [random.randint(1, 100) for _ in range(6)])

print("without it, a different sequence every run, which nobody can check:")
random.seed()
first = [random.randint(1, 100) for _ in range(6)]
random.seed()
second = [random.randint(1, 100) for _ in range(6)]
print("   two unseeded runs agreed?", first == second)
with the seed set, the same numbers every time:
   [42, 20, 51, 84, 7, 10]
   [42, 20, 51, 84, 7, 10]
   [42, 20, 51, 84, 7, 10]
without it, a different sequence every run, which nobody can check:
   two unseeded runs agreed? False

random.seed(7) at the top of a simulation is not optional in a journal entry. Without it, the teacher cannot reproduce your output, you cannot reproduce it yourself the next day, and a bug that appears once can never be found again. Put the seed in, and say in the writeup that changing it gives a different run of the same system.

Procedure

  1. Save as practical2-4b.py. import random and random.seed(7) at the top.
  2. Choose an arrival chance and a service time range, and write them as named constants at the top so

they can be changed in one place.

  1. Use a deque as the queue, storing the minute each customer arrived.
  2. Loop over the minutes. Each minute: maybe add an arrival; free any counter whose service has
munotes.in189

Practical 4 continued: Simulating a Customer Service Queue

finished; start serving from the front of the queue if a counter is free.

  1. Record each customer's wait as the minute served minus the minute they arrived.
  2. Print a table of minute, arrival, queue length, what the counter did, and who is waiting.
  3. Report the average wait, the longest wait, the number served and the number still waiting.
  4. Work out on paper how many counters the system needs, from the arrival rate times the mean service

time, before running the multi-counter version.

  1. Run the simulation for one, two, three and four counters and tabulate the results.
  2. Print the queue length each minute as a row of digits for one counter and for two.

Result

With one counter, an arrival chance of 0.45 and a service time of 2 to 4 minutes, the queue grew throughout the run and customers were still waiting at the end. The arithmetic predicted it: 0.45 arrivals a minute × a mean service of 3 minutes is 1.35 counter-minutes of work created per minute against 1 supplied, so the queue must grow. Over 200 minutes, one counter left an average wait of 32.27 minutes, a worst wait of 67 and 26 customers still waiting; two counters brought the average to 0.44 with nobody left waiting; and three and four counters improved it only to 0.09 and 0.00, so two is the answer. Printing the queue length each minute as digits showed the one counter case climbing and not recovering while the two counter case stayed near zero. All figures are reproducible because random.seed(7) is set.

Where marks are lost

  • No seed. The output cannot be reproduced by the teacher or by you.
  • Serving from the wrong end. pop() instead of popleft() serves the newest arrival first,

which is a stack, not a queue.

  • Not recording the arrival minute, so the wait cannot be computed at all.
  • Only one counter tried. The interesting result is how many are needed.
  • No arithmetic. One line of arrival rate times mean service predicts the answer and shows you

understand it.

  • Reporting only the average wait. The worst wait and the number left waiting are what matter to a

customer.

  • Using list.pop(0) for the queue, which is O(n) as the previous chapter measured.
  • Constants buried in the loop instead of named at the top.

For the journal

The aim in MU's words, including her own example of a customer service queue. The four row table of what a simulation needs. The seed line, called out, with one sentence saying why. Then the minute by minute table for one counter, which is the entry's centrepiece, and the summary figures under it. Then the arithmetic: arrivals per minute × mean service time against the counters available, and the conclusion that one is not enough. Then the table for one to four counters and the sentence naming the smallest number that works. Then the digit rows of queue length for one and two counters. The conclusion: a queue is the right structure because fairness means serving the longest waiting customer, and the simulation turns a guess about how many counters are needed into a number.

munotes.in190

Practical 4 continued: Simulating a Customer Service Queue

Quick revision

  • A simulation needs a clock, arrivals, a queue and one or more servers.
  • The queue is FIFO because fairness means the longest waiting customer is next. popleft(), never

pop().

  • Store the minute of arrival with each customer; the wait is the minute served minus that.
  • random.seed(n) at the top, or nothing in the output can be reproduced or checked.
  • The stability test: arrivals per minute × mean service time against the number of counters. If

the work created exceeds the work supplied, the queue grows without limit.

  • Here 0.45 × 3 = 1.35 counter-minutes per minute against 1 from one counter, so one counter cannot

keep up and two can.

  • Report the average wait, the worst wait, the number served and the number left waiting. The

average alone hides the worst case.

  • A queue length printed as a row of digits per minute shows the trend at a glance.
  • If some customers are urgent it is a priority queue, built on a heap, O(log n) per

operation. Include the arrival order in the tuple to keep ties fair.

Questions you should be able to answer

1. Why is a queue the right structure for a service counter? Because fairness requires the customer who has waited longest to be served next, which is exactly FIFO.

2. What would happen if you used pop() instead of popleft()? The most recent arrival would be served first. That is a stack, and it is the opposite of fair.

3. Why is random.seed(7) in the program? So the run is reproducible: the teacher can get the same output, you can get it again tomorrow, and a bug that appears once can be found again.

4. How do you compute each customer's wait? Store the minute they arrived, and subtract it from the minute they begin to be served.

5. Without running anything, how do you tell whether one counter is enough? Multiply the arrivals per minute by the mean service time. That is the counter-minutes of work created per minute. If it exceeds the number of counters, the queue grows without limit.

munotes.in191

Practical 4 continued: Simulating a Customer Service Queue

6. Do the arithmetic for an arrival chance of 0.45 and a service time of 2 to 4 minutes. The mean service time is 3 minutes, and 0.45 × 3 = 1.35 counter-minutes per minute. One counter supplies 1, so one is not enough and two are.

7. Why report the worst wait as well as the average? Because an average hides the customer who waited longest, and that is the one who complains.

8. What changes if some customers are urgent? It becomes a priority queue rather than a queue, built on a heap, with O(log n) per operation instead of O(1).

9. In heapq, why put the arrival order in the tuple? Because a heap makes no promise about items of equal priority. Including the arrival order makes tuples compare by priority and then by order, so ties are served fairly.

10. Why does adding a third and fourth counter change so little? Because two already supply more work per minute than arrives, so the queue is already near zero and there is almost nothing left for extra counters to do.

Contents This chapter on its own page

munotes.in192

Chapter Thirty

Practical 6: Tree Traversal, Pre-order, In-order and Post-order

Syllabus topic Module 2, practical 6, "Tree Traversal: Write a program to: Implement pre-order, in-order, Post-order traversal of a binary tree."

Aim

To implement pre-order, in-order and post-order traversal of a binary tree.

The tree every listing uses

                  A
                /   \
              B       C
             / \     / \
            D   E   F   G
               /
              H

Seven letters and one at depth 3, so the three orders come out genuinely different and can be checked by hand.

The three orders are one shape with the visit moved

Every one of the three does exactly three things at each node: visit the node, walk the left subtree, walk the right subtree. The only difference is where the visit goes.

TraversalOrder of the threeAlso called
pre-ordervisit, left, rightNLR, node first
in-orderleft, visit, rightLNR, node in the middle
post-orderleft, right, visitLRN, node last

The names say it: "pre" means the node is visited before its subtrees, "in" means between them, "post" means after them. There is nothing else to remember.

def walk(node):
    if node is None:
        return
    # pre-order:  visit(node) HERE
    walk(node.left)
    # in-order:   visit(node) HERE
    walk(node.right)
    # post-order: visit(node) HERE

The class

"""Binary tree traversals, for Major Practical 3, Module 2."""

from collections import deque


class Node:
    def __init__(self, key, left=None, right=None):
        self.key = key
        self.left = left
        self.right = right

    def __repr__(self):
        return f"Node({self.key!r})"


def build_sample():
    """The tree this chapter uses, built from the leaves upwards."""
    h = Node("H")
    e = Node("E", h, None)
    b = Node("B", Node("D"), e)
    c = Node("C", Node("F"), Node("G"))
    return Node("A", b, c)


# ---- recursive, which is the answer -------------------------------------
def pre_order(node, out=None):
    """Visit, then left, then right."""
    if out is None:
        out = []
    if node is not None:
        out.append(node.key)
        pre_order(node.left, out)
        pre_order(node.right, out)
    return out


def in_order(node, out=None):
    """Left, then visit, then right."""
    if out is None:
        out = []
    if node is not None:
        in_order(node.left, out)
        out.append(node.key)
        in_order(node.right, out)
    return out


def post_order(node, out=None):
    """Left, then right, then visit."""
    if out is None:
        out = []
    if node is not None:
        post_order(node.left, out)
        post_order(node.right, out)
        out.append(node.key)
    return out


# ---- one function, three answers ---------------------------------------
def walk(node, order, out=None):
    """All three traversals from one piece of code. order is 'pre', 'in' or 'post'."""
    if out is None:
        out = []
    if node is None:
        return out
    if order == "pre":
        out.append(node.key)
    walk(node.left, order, out)
    if order == "in":
        out.append(node.key)
    walk(node.right, order, out)
    if order == "post":
        out.append(node.key)
    return out


# ---- iterative, with an explicit stack ----------------------------------
def pre_order_iterative(root):
    """A stack, with the RIGHT child pushed first so the left comes out first."""
    if root is None:
        return []
    out = []
    stack = [root]
    while stack:
        node = stack.pop()
        out.append(node.key)
        if node.right is not None:
            stack.append(node.right)
        if node.left is not None:
            stack.append(node.left)
    return out


def in_order_iterative(root):
    """Go as far left as possible, visit, then turn right."""
    out = []
    stack = []
    node = root
    while stack or node is not None:
        while node is not None:
            stack.append(node)
            node = node.left
        node = stack.pop()
        out.append(node.key)
        node = node.right
    return out


def post_order_iterative(root):
    """Pre-order with the children reversed, then the whole thing reversed."""
    if root is None:
        return []
    out = []
    stack = [root]
    while stack:
        node = stack.pop()
        out.append(node.key)
        if node.left is not None:
            stack.append(node.left)
        if node.right is not None:
            stack.append(node.right)
    out.reverse()
    return out


# ---- the fourth one a viva asks for -------------------------------------
def level_order(root):
    """Breadth first, level by level. A QUEUE, not a stack."""
    if root is None:
        return []
    out = []
    waiting = deque([root])
    while waiting:
        node = waiting.popleft()
        out.append(node.key)
        if node.left is not None:
            waiting.append(node.left)
        if node.right is not None:
            waiting.append(node.right)
    return out


def level_order_by_level(root):
    """The same, but grouped into levels."""
    if root is None:
        return []
    levels = []
    waiting = deque([root])
    while waiting:
        level = []
        for _ in range(len(waiting)):
            node = waiting.popleft()
            level.append(node.key)
            if node.left is not None:
                waiting.append(node.left)
            if node.right is not None:
                waiting.append(node.right)
        levels.append(level)
    return levels
munotes.in203

Practical 6: Tree Traversal, Pre-order, In-order and Post-order

All three, from one function

from traverse import build_sample, walk, pre_order, in_order, post_order

root = build_sample()

print("the tree:")
print("             A")
print("           /   \\")
print("         B       C")
print("        / \\     / \\")
print("       D   E   F   G")
print("          /")
print("         H")
print()
print(f"  {'traversal':<12} {'order of the three steps':<28} result")
for order, description in [("pre", "visit, left, right"),
                           ("in", "left, visit, right"),
                           ("post", "left, right, visit")]:
    result = " ".join(walk(root, order))
    print(f"  {order + '-order':<12} {description:<28} {result}")

print()
print("and the three separate functions agree with the one function:")
print("  pre  ", pre_order(root) == walk(root, "pre"))
print("  in   ", in_order(root) == walk(root, "in"))
print("  post ", post_order(root) == walk(root, "post"))
the tree:
             A
           /   \
         B       C
        / \     / \
       D   E   F   G
          /
         H

  traversal    order of the three steps     result
  pre-order    visit, left, right           A B D E H C F G
  in-order     left, visit, right           D B H E A F C G
  post-order   left, right, visit           D H E B F G C A

and the three separate functions agree with the one function:
  pre   True
  in    True
  post  True

Check one of them by hand against the picture. Pre-order visits A first, then the whole left subtree, then the whole right: A, then B's subtree (B, D, E, H), then C's (C, F, G). Post-order visits A last, because a node is visited only after both its subtrees are finished.

Traced, which is the journal entry

from traverse import build_sample

root = build_sample()
depth_of = {}


def trace(node, order, out, depth=0):
    if node is None:
        return
    pad = "  " * depth
    print(f"  {pad}enter {node.key}")
    if order == "pre":
        out.append(node.key)
        print(f"  {pad}VISIT {node.key}  -> {' '.join(out)}")
    trace(node.left, order, out, depth + 1)
    if order == "in":
        out.append(node.key)
        print(f"  {pad}VISIT {node.key}  -> {' '.join(out)}")
    trace(node.right, order, out, depth + 1)
    if order == "post":
        out.append(node.key)
        print(f"  {pad}VISIT {node.key}  -> {' '.join(out)}")


for order in ["pre", "in", "post"]:
    print(f"{order}-order, traced:")
    trace(root, order, [])
    print()
munotes.in204

Practical 6: Tree Traversal, Pre-order, In-order and Post-order

pre-order, traced:
  enter A
  VISIT A  -> A
    enter B
    VISIT B  -> A B
      enter D
      VISIT D  -> A B D
      enter E
      VISIT E  -> A B D E
        enter H
        VISIT H  -> A B D E H
    enter C
    VISIT C  -> A B D E H C
      enter F
      VISIT F  -> A B D E H C F
      enter G
      VISIT G  -> A B D E H C F G

in-order, traced:
  enter A
    enter B
      enter D
      VISIT D  -> D
    VISIT B  -> D B
      enter E
        enter H
        VISIT H  -> D B H
      VISIT E  -> D B H E
  VISIT A  -> D B H E A
    enter C
      enter F
      VISIT F  -> D B H E A F
    VISIT C  -> D B H E A F C
      enter G
      VISIT G  -> D B H E A F C G

post-order, traced:
  enter A
    enter B
      enter D
      VISIT D  -> D
      enter E
        enter H
        VISIT H  -> D H
      VISIT E  -> D H E
    VISIT B  -> D H E B
    enter C
      enter F
      VISIT F  -> D H E B F
      enter G
      VISIT G  -> D H E B F G
    VISIT C  -> D H E B F G C
  VISIT A  -> D H E B F G C A

Read the indentation. Every node is entered in exactly the same sequence in all three traversals; only the line on which it is visited moves. That is the point of this chapter, and the three traces side by side are the best thing to put in the journal.

In-order on a binary SEARCH tree gives sorted order

This is the property to quote, and it is why in-order is the traversal that matters most.

from traverse import Node, in_order


def insert(node, key):
    if node is None:
        return Node(key)
    if key < node.key:
        node.left = insert(node.left, key)
    elif key > node.key:
        node.right = insert(node.right, key)
    return node


root = None
for key in [50, 30, 70, 20, 40, 60, 80, 35]:
    root = insert(root, key)

walked = in_order(root)
print("in-order walk of a BST:", walked)
print("is it sorted?          ", walked == sorted(walked))
print()
print("the same keys inserted in a different order:")
root2 = None
for key in [35, 80, 20, 60, 40, 70, 30, 50]:
    root2 = insert(root2, key)
print("in-order walk         :", in_order(root2))
print("the same sorted list, from a completely different tree shape")
munotes.in205

Practical 6: Tree Traversal, Pre-order, In-order and Post-order

in-order walk of a BST: [20, 30, 35, 40, 50, 60, 70, 80]
is it sorted?           True

the same keys inserted in a different order:
in-order walk         : [20, 30, 35, 40, 50, 60, 70, 80]
the same sorted list, from a completely different tree shape

The in-order walk of a binary search tree is the keys in sorted order, whatever shape the tree has. Two different insertion orders gave two different trees and the same sorted output. That is a free sort, and it is the single most useful fact about in-order traversal.

It also runs the other way: if an in-order walk is not sorted, the tree is not a binary search tree, which is the check used in [Practical 5: the Binary Search Tree, Create, Insert and Search].

Iterative, with an explicit stack

Recursion uses the call stack. Writing the traversal with a stack of your own makes that stack visible, and it is what an examiner asks for when recursion is not allowed.

from traverse import (build_sample, walk, pre_order_iterative,
                      in_order_iterative, post_order_iterative)

root = build_sample()

print(f"  {'traversal':<12} {'recursive':<18} {'iterative':<18} same?")
for order, iterative in [("pre", pre_order_iterative),
                         ("in", in_order_iterative),
                         ("post", post_order_iterative)]:
    recursive = walk(root, order)
    by_stack = iterative(root)
    print(f"  {order + '-order':<12} {' '.join(recursive):<18} "
          f"{' '.join(by_stack):<18} {recursive == by_stack}")
  traversal    recursive          iterative          same?
  pre-order    A B D E H C F G    A B D E H C F G    True
  in-order     D B H E A F C G    D B H E A F C G    True
  post-order   D H E B F G C A    D H E B F G C A    True

Each iterative version has one idea in it, and each is worth a sentence in the journal.

Pre-order pushes the right child first, because a stack gives back the last thing pushed, so pushing right then left makes left come out first.

In-order goes as far left as it can, pushing as it goes, then pops, visits, and turns right. The inner while is the "as far left as possible" part.

Post-order is the neat one: do a pre-order with the children pushed the other way round, which gives node, right, left, and then reverse the whole list, which gives left, right, node. Two lines instead of the awkward two-stack version most books print.

Level order, which is the fourth one

from traverse import build_sample, level_order, level_order_by_level

root = build_sample()
print("level-order (breadth first):", " ".join(level_order(root)))
print()
print("grouped by level:")
for depth, level in enumerate(level_order_by_level(root)):
    print(f"  depth {depth}: {' '.join(level)}")
print()
print("the height is", len(level_order_by_level(root)) - 1)
print("the widest level has", max(len(l) for l in level_order_by_level(root)), "nodes")
munotes.in206

Practical 6: Tree Traversal, Pre-order, In-order and Post-order

level-order (breadth first): A B C D E F G H

grouped by level:
  depth 0: A
  depth 1: B C
  depth 2: D E F G
  depth 3: H

the height is 3
the widest level has 4 nodes

Level order uses a queue; the other three use a stack. That is the whole difference, and it is the same contrast as breadth first against depth first in [Python for Data Structures, and the Cost of an Operation]. The for _ in range(len(waiting)) in level_order_by_level is how one level is taken off at a time: the queue's length at the start of a round is exactly the number of nodes on that level.

Level order also gives two things the other traversals do not: the height, as the number of levels minus one, and the widest level.

One traversal is not enough to rebuild the tree

An examiner's favourite. Given only a pre-order walk, the tree cannot be reconstructed, because different trees give the same walk.

from traverse import Node, walk

# two DIFFERENT trees
left_leaning = Node("A", Node("B", Node("C")), None)
right_leaning = Node("A", None, Node("B", None, Node("C")))
zig = Node("A", Node("B", None, Node("C")), None)

for name, tree in [("A with B, C all left  ", left_leaning),
                   ("A with B, C all right ", right_leaning),
                   ("A left to B, B right to C", zig)]:
    print(f"  {name:<26} pre {' '.join(walk(tree, 'pre')):<8} "
          f"in {' '.join(walk(tree, 'in')):<8} post {' '.join(walk(tree, 'post'))}")

print()
print("all three have the SAME pre-order, so pre-order alone is not enough.")
print("but their in-orders differ, and pre-order PLUS in-order is enough:")
  A with B, C all left       pre A B C    in C B A    post C B A
  A with B, C all right      pre A B C    in A B C    post C B A
  A left to B, B right to C  pre A B C    in B C A    post C B A

all three have the SAME pre-order, so pre-order alone is not enough.
but their in-orders differ, and pre-order PLUS in-order is enough:
from traverse import Node, walk


def rebuild(pre, ino):
    """Rebuild a binary tree from its pre-order and in-order walks."""
    if not pre:
        return None
    root_key = pre[0]
    cut = ino.index(root_key)
    node = Node(root_key)
    node.left = rebuild(pre[1:1 + cut], ino[:cut])
    node.right = rebuild(pre[1 + cut:], ino[cut + 1:])
    return node


def build_sample():
    h = Node("H")
    e = Node("E", h, None)
    b = Node("B", Node("D"), e)
    c = Node("C", Node("F"), Node("G"))
    return Node("A", b, c)


original = build_sample()
pre = walk(original, "pre")
ino = walk(original, "in")
post = walk(original, "post")

rebuilt = rebuild(pre, ino)

print("pre  of the original:", " ".join(pre))
print("in   of the original:", " ".join(ino))
print()
print("rebuilt from those two:")
print("  pre :", " ".join(walk(rebuilt, "pre")), " same?",
      walk(rebuilt, "pre") == pre)
print("  in  :", " ".join(walk(rebuilt, "in")), " same?",
      walk(rebuilt, "in") == ino)
print("  post:", " ".join(walk(rebuilt, "post")), " same?",
      walk(rebuilt, "post") == post)
print()
print("all three match, so the rebuilt tree is the original tree.")
munotes.in207

Practical 6: Tree Traversal, Pre-order, In-order and Post-order

pre  of the original: A B D E H C F G
in   of the original: D B H E A F C G

rebuilt from those two:
  pre : A B D E H C F G  same? True
  in  : D B H E A F C G  same? True
  post: D H E B F G C A  same? True

all three match, so the rebuilt tree is the original tree.

Pre-order plus in-order determines the tree, and so does post-order plus in-order. Pre-order plus post-order does not. The reason is that pre-order gives the root and in-order says how many nodes are on each side of it, and that pair of facts is enough to split the problem and recurse.

What each traversal is actually for

TraversalUsed for
pre-ordercopying a tree, writing it to a file, printing a directory listing with the folder before its contents
in-ordergetting a BST's keys in sorted order. Nothing else does this for free
post-orderfreeing or deleting a tree, evaluating an expression tree, computing a folder's total size
level orderthe shortest path in an unweighted graph, printing a tree by rows, finding the height

Post-order is the one to use when a node's work depends on its children. Deleting a tree is the clearest case: a node can only be freed after both its subtrees have been, so any other order leaves a dangling reference. Here it is on a folder size:

from traverse import Node, walk


class Folder:
    def __init__(self, name, size=0, children=()):
        self.name = name
        self.size = size
        self.children = list(children)


def total_size(folder, depth=0):
    """POST-order: a folder's total needs its children's totals first."""
    subtotal = sum(total_size(child, depth + 1) for child in folder.children)
    total = folder.size + subtotal
    print(f"  {'  ' * depth}{folder.name:<14} own {folder.size:>5} "
          f"children {subtotal:>5}  total {total:>5}")
    return total


tree = Folder("notes", 0, [
    Folder("module1", 0, [Folder("practical1.py", 400),
                          Folder("practical2.py", 350)]),
    Folder("module2", 0, [Folder("bst.py", 1200),
                          Folder("queue_array.py", 900)]),
    Folder("README.txt", 120),
])

print("computing the total size, post-order:")
grand = total_size(tree)
print(f"  the whole tree is {grand} bytes")
print(f"  check: 400 + 350 + 1200 + 900 + 120 = {400 + 350 + 1200 + 900 + 120}")
computing the total size, post-order:
      practical1.py  own   400 children     0  total   400
      practical2.py  own   350 children     0  total   350
    module1        own     0 children   750  total   750
      bst.py         own  1200 children     0  total  1200
      queue_array.py own   900 children     0  total   900
    module2        own     0 children  2100  total  2100
    README.txt     own   120 children     0  total   120
  notes          own     0 children  2970  total  2970
  the whole tree is 2970 bytes
  check: 400 + 350 + 1200 + 900 + 120 = 2970
munotes.in208

Practical 6: Tree Traversal, Pre-order, In-order and Post-order

Every folder's line is printed after its children's lines, which is post-order, and it has to be: the total cannot be known until the children's totals are.

Procedure

  1. Save traverse.py with Node, build_sample, the three recursive traversals, the one walk

function with an order argument, the three iterative versions and level_order.

  1. Build the seven node tree and draw it in the journal.
  2. Print all three traversals from the single walk function, and confirm the three separate

functions agree.

  1. Check one traversal by hand against the drawing.
  2. Print all three traced, showing every node entered and the line on which it is visited, and note

that the entering sequence is the same in all three.

  1. Build a binary search tree and show that its in-order walk is sorted, then build the same keys in a

different order and show the same sorted output from a different tree.

  1. Write the three iterative versions and confirm each equals its recursive twin.
  2. Add level order with a queue, print it grouped by level, and take the height from the number of

levels.

  1. Show two different trees with the same pre-order walk.
  2. Rebuild the tree from its pre-order and in-order walks and confirm all three walks of the rebuilt

tree match.

Result

All three traversals of the seven node tree came out as the picture requires, and the three separate functions agreed with the single walk function in every case. The traced runs showed every node entered in the same sequence in all three, with only the visit line moving. The in-order walk of a binary search tree was sorted, and the same keys inserted in a different order gave a different tree shape and the same sorted walk. Each iterative traversal equalled its recursive twin. Level order, using a queue, printed the tree by rows and gave the height as the number of levels minus one. Three different trees shared one pre-order walk, so pre-order alone does not determine a tree; rebuilding from pre-order plus in-order reproduced a tree whose pre-order, in-order and post-order all matched the original. The folder size computation printed every parent after its children, which post-order requires, and the total matched the hand check 400 + 350 + 1200 + 900 + 120 = 2970.

Where marks are lost

  • Memorising three functions instead of seeing one with the visit moved.
  • Getting the names backwards. Pre means the node before its subtrees, post after.
  • No base case for None, so the recursion fails on a leaf's missing child.
  • Pushing the left child first in the iterative pre-order, which gives the right subtree first.
  • Using a stack for level order. It needs a queue.
  • Not showing that in-order gives sorted output on a BST. That is the property worth quoting.
  • Claiming one traversal determines the tree. It does not; pre-order plus in-order does.
  • Using pre-order to delete a tree, which frees a parent before its children.
munotes.in209

Practical 6: Tree Traversal, Pre-order, In-order and Post-order

For the journal

The aim in MU's words, all three. The tree drawn, because every traversal is checked against the drawing. The three row table of where the visit goes, and the walk skeleton with the three comment positions in it, which is the one thing to memorise. Then all three outputs side by side, and one checked by hand against the drawing. Then at least one traced run with the indentation, and one sentence saying every node is entered in the same order in all three. Then the BST in-order walk shown to be sorted. Then the iterative versions with one sentence each on the idea. Then level order grouped by level with the height taken from it. The conclusion: pre, in and post-order are one recursion with the visit before, between or after the two subtrees, they all cost O(n) because every node is visited once, and in-order on a binary search tree gives the keys in sorted order.

Quick revision

  • Three steps at each node: visit, walk left, walk right. Only the position of the visit changes.
  • pre-order visit, left, right. in-order left, visit, right. post-order left, right,

visit.

  • pre means the node before its subtrees, in means between, post means after.
  • All three are O(n) in time, because each node is visited once, and O(height) in extra space for

the stack.

  • In-order on a BST gives the keys sorted, whatever the tree's shape. That is the property to

quote, and it is also the test that a tree is a BST.

  • Iterative pre-order: a stack, and push the right child first.
  • Iterative in-order: go as far left as possible pushing, then pop, visit, turn right.
  • Iterative post-order: pre-order with the children swapped, then reverse the result.
  • Level order uses a QUEUE. The queue's length at the start of a round is one whole level.
  • Level order gives the height (levels minus one) and the widest level.
  • One traversal does not determine the tree. Pre-order plus in-order does, and so does post-order

plus in-order; pre plus post does not.

munotes.in210

Practical 6: Tree Traversal, Pre-order, In-order and Post-order

  • Use pre-order to copy, in-order to sort, post-order when a node needs its children first, such

as deleting a tree or totalling a folder.

Questions you should be able to answer

1. Give the three traversals in terms of the visit. Pre-order visits the node, then the left subtree, then the right. In-order does left, node, right. Post-order does left, right, node.

2. How do you remember which is which? The prefix says where the node is: pre is before its subtrees, in is between them, post is after them.

3. What is the in-order walk of a binary search tree? The keys in sorted order, whatever shape the tree has.

4. How does that give you a test? If an in-order walk is not sorted, the tree is not a binary search tree.

5. What does each traversal cost? O(n) time, since every node is visited once, and O(height) extra space for the stack, whether that is the call stack or your own.

6. In the iterative pre-order, why is the right child pushed first? Because a stack returns the last thing pushed, so pushing right and then left makes the left subtree come out first.

7. How do you get post-order with one stack? Do a pre-order but push the left child first, giving node, right, left, and then reverse the whole list.

8. Which structure does level order use? A queue. All three depth first traversals use a stack.

9. Can you rebuild a tree from its pre-order walk? No: this chapter shows three different trees with the same pre-order walk. Pre-order plus in-order is enough, because pre-order gives the root and in-order says how many nodes lie on each side of it.

10. Which traversal would you use to delete a whole tree, and why? Post-order, because a node can only be freed after both its subtrees have been freed. Any other order leaves a reference to freed memory.

Contents This chapter on its own page

munotes.in211

Chapter Thirty-One

Practical 7: a Hash Table with Separate Chaining

Syllabus topic Module 2, practical 7, "Hash Table: Write a program to: Implement a hash table with separate chaining for collision handling. Store and retrieve data from the hash table."

Aim

To implement a hash table with separate chaining for collision handling, and to store data in it and retrieve it.

The idea

An array finds item n in one step, by arithmetic on the index. A hash table finds the item for a key in one step, by arithmetic on the key.

key "Aarti"  --hash-->  1234567  --% 7-->  bucket 4

buckets:
  0:
  1:  ("Divya", 55)
  2:
  3:
  4:  ("Aarti", 78) -> ("Eshan", 67)      <- two keys, one bucket: a CHAIN
  5:  ("Chetan", 90)
  6:

Three parts, and each has its own job:

PartWhat it does
the hash functionturns a key of any kind into a whole number
the bucket indexthat number modulo the number of buckets
collision handlingwhat to do when two keys land in the same bucket

MU names separate chaining as the collision handling, which means each bucket holds a list of the pairs that landed there.

A hash function, written out

def simple_hash(key):
    """Add the character codes. Simple, and bad, which the next section shows."""
    total = 0
    for character in str(key):
        total += ord(character)
    return total


def polynomial_hash(key, base=31):
    """Multiply by a base at each step, so POSITION matters."""
    total = 0
    for character in str(key):
        total = total * base + ord(character)
    return total


for name in ["Aarti", "Divya", "Chetan", "abc", "cba", "bca"]:
    print(f"  {name:<8} sum of codes {simple_hash(name):>8}   "
          f"polynomial {polynomial_hash(name):>14}")

print()
print("look at abc, cba and bca: the sum is the SAME for all three,")
print("because addition does not care about order. The polynomial differs.")
  Aarti    sum of codes      497   polynomial       63031847
  Divya    sum of codes      509   polynomial       66044729
  Chetan   sum of codes      595   polynomial     2017322785
  abc      sum of codes      294   polynomial          96354
  cba      sum of codes      294   polynomial          98274
  bca      sum of codes      294   polynomial          97344

look at abc, cba and bca: the sum is the SAME for all three,
because addition does not care about order. The polynomial differs.

There is the difference between a bad hash function and a good one. Adding the character codes gives the same number for any anagram, so abc, cba and bca all collide, and in real data anagrams and rearrangements are common. Multiplying by a base at each step makes the position of each character matter.

The 31 is a small odd prime, and the choice matters less than people think; what matters is that the multiplication mixes the bits.

The class

"""A hash table with separate chaining, for Major Practical 3, Module 2."""


def string_hash(key, base=31):
    """A polynomial hash, written out so the bucket numbers are reproducible.

    NOTE: Python's built in hash() of a str is salted per process, so it gives
    a different number in every run. A textbook example needs a stable one.
    """
    total = 0
    for character in str(key):
        total = total * base + ord(character)
    return total


class HashTable:
    """Separate chaining: each bucket is a list of (key, value) pairs."""

    def __init__(self, buckets=7, load_limit=0.75):
        if buckets <= 0:
            raise ValueError("there must be at least one bucket")
        self.buckets = [[] for _ in range(buckets)]
        self.count = 0
        self.load_limit = load_limit
        self.comparisons = 0
        self.rehashes = 0

    # ---- the three parts -------------------------------------------------
    def bucket_of(self, key):
        """The index this key belongs in."""
        return string_hash(key) % len(self.buckets)

    def load_factor(self):
        """Items divided by buckets. The single number that predicts the cost."""
        return self.count / len(self.buckets)

    # ---- storing ---------------------------------------------------------
    def put(self, key, value):
        """Store or update. Returns the bucket used and whether it collided."""
        index = self.bucket_of(key)
        chain = self.buckets[index]
        collided = len(chain) > 0
        for position, (existing, _) in enumerate(chain):
            if existing == key:
                chain[position] = (key, value)      # an update, not a second entry
                return index, False
        chain.append((key, value))
        self.count += 1
        if self.load_factor() > self.load_limit:
            self._rehash()
        return index, collided

    def __setitem__(self, key, value):
        self.put(key, value)

    # ---- retrieving ------------------------------------------------------
    def get(self, key, default=None):
        """The value for a key. Counts the comparisons made inside the chain."""
        self.comparisons = 0
        chain = self.buckets[self.bucket_of(key)]
        for existing, value in chain:
            self.comparisons += 1
            if existing == key:
                return value
        return default

    def __getitem__(self, key):
        marker = object()
        value = self.get(key, marker)
        if value is marker:
            raise KeyError(key)
        return value

    def __contains__(self, key):
        marker = object()
        return self.get(key, marker) is not marker

    def __len__(self):
        return self.count

    # ---- removing --------------------------------------------------------
    def remove(self, key):
        """Delete a key. Raises KeyError if it is not there."""
        chain = self.buckets[self.bucket_of(key)]
        for position, (existing, _) in enumerate(chain):
            if existing == key:
                del chain[position]
                self.count -= 1
                return True
        raise KeyError(key)

    # ---- growing ---------------------------------------------------------
    def _rehash(self):
        """Double the buckets and put every pair back. Every index changes."""
        old = self.buckets
        self.buckets = [[] for _ in range(len(old) * 2)]
        self.count = 0
        self.rehashes += 1
        for chain in old:
            for key, value in chain:
                index = self.bucket_of(key)
                self.buckets[index].append((key, value))
                self.count += 1

    # ---- looking at it ---------------------------------------------------
    def show(self):
        lines = []
        for index, chain in enumerate(self.buckets):
            if chain:
                items = " -> ".join(f"({k!r}, {v!r})" for k, v in chain)
                lines.append(f"  {index:>3}: {items}")
            else:
                lines.append(f"  {index:>3}:")
        return "\n".join(lines)

    def chain_lengths(self):
        return [len(chain) for chain in self.buckets]

    def statistics(self):
        lengths = self.chain_lengths()
        used = sum(1 for n in lengths if n)
        return {
            "items": self.count,
            "buckets": len(self.buckets),
            "load factor": round(self.load_factor(), 3),
            "buckets used": used,
            "buckets empty": len(lengths) - used,
            "longest chain": max(lengths),
            "average non-empty chain": round(self.count / used, 3) if used else 0,
        }
munotes.in212

Practical 7: a Hash Table with Separate Chaining

Storing and retrieving

from hashtable import HashTable

table = HashTable(buckets=7)

marks = [("Aarti", 78), ("Divya", 55), ("Chetan", 90),
         ("Eshan", 67), ("Farhan", 41), ("Gauri", 83)]

print("storing:")
for name, mark in marks:
    index, collided = table.put(name, mark)
    note = "  COLLISION, chained" if collided else ""
    print(f"  {name:<8} -> bucket {index}{note}")

print()
print("the table:")
print(table.show())

print()
print("retrieving:")
for name in ["Aarti", "Chetan", "Gauri", "Nobody"]:
    value = table.get(name)
    found = value if value is not None else "not found"
    print(f"  {name:<8} -> {str(found):<12} after {table.comparisons} "
          f"comparison(s) in its chain")

print()
print("the dictionary style operations work too:")
table["Hiral"] = 72
print("  table['Hiral'] =", table["Hiral"])
print("  'Aarti' in table:", "Aarti" in table)
print("  'Nobody' in table:", "Nobody" in table)
print("  len(table):", len(table))
munotes.in213

Practical 7: a Hash Table with Separate Chaining

storing:
  Aarti    -> bucket 4
  Divya    -> bucket 2
  Chetan   -> bucket 2  COLLISION, chained
  Eshan    -> bucket 0
  Farhan   -> bucket 1
  Gauri    -> bucket 0  COLLISION, chained

the table:
    0: ('Gauri', 83)
    1:
    2:
    3:
    4:
    5:
    6:
    7: ('Eshan', 67)
    8: ('Farhan', 41)
    9: ('Divya', 55) -> ('Chetan', 90)
   10:
   11: ('Aarti', 78)
   12:
   13:

retrieving:
  Aarti    -> 78           after 1 comparison(s) in its chain
  Chetan   -> 90           after 2 comparison(s) in its chain
  Gauri    -> 83           after 1 comparison(s) in its chain
  Nobody   -> not found    after 1 comparison(s) in its chain

the dictionary style operations work too:
  table['Hiral'] = 72
  'Aarti' in table: True
  'Nobody' in table: False
  len(table): 7

Read the comparison counts. A retrieval compares only inside its own bucket, so it took one comparison where the bucket held one pair and more where a chain had formed. That is the whole reason a hash table is fast: the bucket is found by arithmetic and only its own short chain is searched.

An update is not a second entry

from hashtable import HashTable

table = HashTable(buckets=7)
table.put("Aarti", 78)
print("after storing 78:", table.get("Aarti"), " len", len(table))

table.put("Aarti", 91)
print("after storing 91:", table.get("Aarti"), " len", len(table),
      " <- still one item, the value was REPLACED")
print()
print(table.show())
after storing 78: 78  len 1
after storing 91: 91  len 1  <- still one item, the value was REPLACED

    0:
    1:
    2:
    3:
    4: ('Aarti', 91)
    5:
    6:

put searches the chain first and replaces the pair if the key is already there. Without that search, the same key would appear twice in the chain and get would return whichever came first, which is a real bug and an easy one to write.

A collision, and how likely it is

from hashtable import HashTable, string_hash

table = HashTable(buckets=7)
names = ["Aarti", "Divya", "Chetan", "Eshan", "Farhan", "Gauri", "Hiral",
         "Imran", "Jyoti"]

print(f"  {'key':<8} {'hash':>16} {'% 7':>5}")
for name in names:
    print(f"  {name:<8} {string_hash(name):>16} {string_hash(name) % 7:>5}")

print()
seen = {}
for name in names:
    index = string_hash(name) % 7
    seen.setdefault(index, []).append(name)

for index in sorted(seen):
    if len(seen[index]) > 1:
        print(f"  bucket {index} is shared by {', '.join(seen[index])}"
              f"  <- a COLLISION")
munotes.in214

Practical 7: a Hash Table with Separate Chaining

  key                  hash   % 7
  Aarti            63031847     4
  Divya            66044729     2
  Chetan         2017322785     2
  Eshan            67251975     0
  Farhan         2097121342     1
  Gauri            68575794     0
  Hiral            69734236     5
  Imran            70776923     0
  Jyoti            72055637     3

  bucket 0 is shared by Eshan, Gauri, Imran  <- a COLLISION
  bucket 2 is shared by Divya, Chetan  <- a COLLISION

Collisions are not an accident. With more keys than buckets they are certain, by the pigeonhole principle, and long before that they are likely:

def collision_probability(keys, buckets):
    """The chance that at least two of `keys` keys share a bucket."""
    if keys > buckets:
        return 1.0
    no_collision = 1.0
    for i in range(keys):
        no_collision *= (buckets - i) / buckets
    return 1 - no_collision


print(f"  {'buckets':>8} {'keys':>6} {'chance of a collision':>23}")
for buckets, keys in [(7, 2), (7, 4), (23, 5), (23, 10), (100, 10),
                      (365, 23), (1000, 40)]:
    chance = collision_probability(keys, buckets)
    print(f"  {buckets:>8} {keys:>6} {chance * 100:>22.1f}%")

print()
print("the 365 and 23 row is the birthday problem: in a class of 23,")
print("it is more likely than not that two students share a birthday.")
print("a hash table meets the same arithmetic, which is why collision")
print("handling is part of the design and not an afterthought.")
   buckets   keys   chance of a collision
         7      2                   14.3%
         7      4                   65.0%
        23      5                   37.3%
        23     10                   90.0%
       100     10                   37.2%
       365     23                   50.7%
      1000     40                   54.6%

the 365 and 23 row is the birthday problem: in a class of 23,
it is more likely than not that two students share a birthday.
a hash table meets the same arithmetic, which is why collision
handling is part of the design and not an afterthought.

That is why MU's row says "for collision handling". A design without it is not a hash table.

A good hash function against a bad one, measured

from hashtable import HashTable


def bad_hash(key):
    """Only the first character. Deliberately terrible."""
    return ord(str(key)[0])


def sum_hash(key):
    """The sum of the character codes."""
    return sum(ord(c) for c in str(key))


def polynomial_hash(key, base=31):
    total = 0
    for character in str(key):
        total = total * base + ord(character)
    return total


names = ["Aarti", "Anita", "Amit", "Ajay", "Divya", "Deepak", "Chetan",
         "Eshan", "Farhan", "Gauri", "Hiral", "Imran", "Jyoti", "Kiran",
         "Lata", "Manish", "Nisha", "Omkar", "Pooja", "Rahul"]
buckets = 7

print(f"  {'hash function':<14} {'chain lengths':<22} {'longest':>8} {'empty':>7}")
for label, function in [("first letter", bad_hash), ("sum of codes", sum_hash),
                        ("polynomial", polynomial_hash)]:
    lengths = [0] * buckets
    for name in names:
        lengths[function(name) % buckets] += 1
    print(f"  {label:<14} {str(lengths):<22} {max(lengths):>8} "
          f"{lengths.count(0):>7}")

print()
print(f"{len(names)} keys into {buckets} buckets, so a perfect spread is about "
      f"{len(names) / buckets:.1f} per bucket")
print("the longest chain is what a lookup costs in the worst case")
munotes.in215

Practical 7: a Hash Table with Separate Chaining

  hash function  chain lengths           longest   empty
  first letter   [2, 2, 6, 2, 2, 4, 2]         6       0
  sum of codes   [3, 2, 3, 2, 4, 3, 3]         4       0
  polynomial     [5, 1, 3, 3, 1, 3, 4]         5       0

20 keys into 7 buckets, so a perfect spread is about 2.9 per bucket
the longest chain is what a lookup costs in the worst case

Read that honestly, because it does not say what a textbook would predict.

The first letter hash is the worst, at 6, and the reason is visible in the data: four of these names begin with A, so all four are forced into one bucket. That is a structural fault and it will show on any list of names.

The sum of codes came out best on this sample, at 4, and the polynomial second, at 5. Twenty keys in seven buckets is far too small a sample to rank two reasonable hash functions, and this run is a good reminder not to claim otherwise from one measurement.

So why is the sum of codes still the wrong choice? Because its fault is structural, not statistical, and the next listing makes it show.

def sum_hash(key):
    return sum(ord(c) for c in str(key))


def polynomial_hash(key, base=31):
    total = 0
    for character in str(key):
        total = total * base + ord(character)
    return total


groups = [["Anil", "Lina", "Nail", "Lani"],
          ["listen", "silent", "enlist", "tinsel"],
          ["MU101", "MU110", "MU011"]]

print("  keys that are rearrangements of each other:")
for group in groups:
    sums = {sum_hash(k) for k in group}
    polys = {polynomial_hash(k) for k in group}
    print(f"    {', '.join(group)}")
    print(f"      distinct sums       : {len(sums)} out of {len(group)}")
    print(f"      distinct polynomials: {len(polys)} out of {len(group)}")

print()
print("every rearrangement of the same characters gets the SAME sum, so they")
print("all collide, whatever the bucket count. The polynomial separates them.")
  keys that are rearrangements of each other:
    Anil, Lina, Nail, Lani
      distinct sums       : 1 out of 4
      distinct polynomials: 4 out of 4
    listen, silent, enlist, tinsel
      distinct sums       : 1 out of 4
      distinct polynomials: 4 out of 4
    MU101, MU110, MU011
      distinct sums       : 1 out of 3
      distinct polynomials: 3 out of 3

every rearrangement of the same characters gets the SAME sum, so they
all collide, whatever the bucket count. The polynomial separates them.

That is the disqualifying fault. Adding the codes gives every rearrangement of the same characters one number, so they collide in every table of every size, and no amount of rehashing helps. Real keys contain rearrangements all the time: names, product codes, roll numbers.

So the lesson for the journal is in two parts. Judge a hash function by the longest chain it produces on your own data, and, before that, reject one whose collisions are built into its arithmetic rather than left to chance.

munotes.in216

Practical 7: a Hash Table with Separate Chaining

The load factor, and rehashing

The load factor is the number of items divided by the number of buckets. It is the one number that predicts the cost of a lookup.

from hashtable import HashTable

table = HashTable(buckets=4, load_limit=0.75)
names = ["Aarti", "Divya", "Chetan", "Eshan", "Farhan", "Gauri", "Hiral",
         "Imran", "Jyoti", "Kiran", "Lata", "Manish"]

print(f"  {'after':<10} {'items':>6} {'buckets':>8} {'load':>6} "
      f"{'longest chain':>14} {'rehashes':>9}")
for name in names:
    table.put(name, len(name))
    stats = table.statistics()
    print(f"  {name:<10} {stats['items']:>6} {stats['buckets']:>8} "
          f"{stats['load factor']:>6} {stats['longest chain']:>14} "
          f"{table.rehashes:>9}")

print()
print("the buckets doubled whenever the load factor passed 0.75,")
print("which is what keeps the chains short and the lookup O(1)")
print()
for key, value in table.statistics().items():
    print(f"  {key:<24} {value}")
  after       items  buckets   load  longest chain  rehashes
  Aarti           1        4   0.25              1         0
  Divya           2        4    0.5              1         0
  Chetan          3        4   0.75              2         0
  Eshan           4        8    0.5              2         1
  Farhan          5        8  0.625              2         1
  Gauri           6        8   0.75              2         1
  Hiral           7       16  0.438              2         2
  Imran           8       16    0.5              2         2
  Jyoti           9       16  0.562              2         2
  Kiran          10       16  0.625              2         2
  Lata           11       16  0.688              2         2
  Manish         12       16   0.75              2         2

the buckets doubled whenever the load factor passed 0.75,
which is what keeps the chains short and the lookup O(1)

  items                    12
  buckets                  16
  load factor              0.75
  buckets used             9
  buckets empty            7
  longest chain            2
  average non-empty chain  1.333

Read the buckets column: it doubles whenever the load factor passes the limit. Rehashing is not optional in a hash table that grows, and the reason is in the cost table below.

Rehashing means every key has to be placed again, because the bucket index is the hash modulo the number of buckets, and that number has changed. A rehash is O(n), and it happens rarely enough that the average insertion is still O(1), which is what "amortised" means.

What it all costs

from hashtable import HashTable


def chain_costs(items, buckets):
    """The average and worst comparisons for a successful lookup."""
    # load_limit must be BIGGER than any load factor reached here, or the
    # table rehashes and the "1 bucket" row is not one bucket at all.
    table = HashTable(buckets=buckets, load_limit=float("inf"))
    for n in range(items):
        table.put(f"key{n:05d}", n)
    lengths = table.chain_lengths()
    total = sum(length * (length + 1) // 2 for length in lengths)
    return total / items, max(lengths)


print(f"  {'items':>7} {'buckets':>8} {'load':>6} {'avg comparisons':>17} "
      f"{'worst':>7}")
for items, buckets in [(100, 1000), (500, 1000), (1000, 1000),
                       (2000, 1000), (10000, 1000), (10000, 1)]:
    average, worst = chain_costs(items, buckets)
    print(f"  {items:>7} {buckets:>8} {items / buckets:>6.1f} "
          f"{average:>17.2f} {worst:>7}")

print()
print("the average comparisons track the LOAD FACTOR, not the number of items.")
print("the last row is one bucket, which is a plain linked list: O(n).")
munotes.in217

Practical 7: a Hash Table with Separate Chaining

    items  buckets   load   avg comparisons   worst
      100     1000    0.1              1.00       1
      500     1000    0.5              1.31       3
     1000     1000    1.0              1.59       4
     2000     1000    2.0              2.20       7
    10000     1000   10.0              5.74      16
    10000        1 10000.0           5000.50   10000

the average comparisons track the LOAD FACTOR, not the number of items.
the last row is one bucket, which is a plain linked list: O(n).
SituationLookup cost
load factor kept small, good hashO(1)
load factor allowed to growO(load factor), so O(n / buckets)
a bad hash putting everything in one bucketO(n), a linked list

So the honest statement is: a hash table is O(1) on average, given a good hash function and a load factor kept low by rehashing, and O(n) in the worst case. The last row of that output is the worst case, made to happen.

Separate chaining against the other way

MU names separate chaining. The alternative is open addressing, where a colliding key goes into another bucket rather than into a chain, and being able to compare them is worth a mark.

class LinearProbing:
    """Open addressing: on a collision, try the next slot, and the next."""

    EMPTY = object()
    DELETED = object()

    def __init__(self, size=11):
        self.slots = [self.EMPTY] * size
        self.count = 0

    def _hash(self, key):
        total = 0
        for character in str(key):
            total = total * 31 + ord(character)
        return total % len(self.slots)

    def put(self, key, value):
        index = self._hash(key)
        probes = 0
        while probes < len(self.slots):
            here = self.slots[index]
            if here is self.EMPTY or here is self.DELETED:
                self.slots[index] = (key, value)
                self.count += 1
                return probes + 1
            if here[0] == key:
                self.slots[index] = (key, value)
                return probes + 1
            index = (index + 1) % len(self.slots)
            probes += 1
        raise OverflowError("the table is full")

    def get(self, key):
        index = self._hash(key)
        probes = 0
        while probes < len(self.slots):
            here = self.slots[index]
            if here is self.EMPTY:
                return None, probes + 1
            if here is not self.DELETED and here[0] == key:
                return here[1], probes + 1
            index = (index + 1) % len(self.slots)
            probes += 1
        return None, probes

    def remove(self, key):
        """A removed slot becomes DELETED, not EMPTY. This is the TOMBSTONE."""
        index = self._hash(key)
        for _ in range(len(self.slots)):
            here = self.slots[index]
            if here is self.EMPTY:
                return False
            if here is not self.DELETED and here[0] == key:
                self.slots[index] = self.DELETED
                self.count -= 1
                return True
            index = (index + 1) % len(self.slots)
        return False

    def show(self):
        out = []
        for i, slot in enumerate(self.slots):
            if slot is self.EMPTY:
                out.append(f"  {i:>3}: .")
            elif slot is self.DELETED:
                out.append(f"  {i:>3}: TOMBSTONE")
            else:
                out.append(f"  {i:>3}: {slot[0]!r} = {slot[1]!r}")
        return "\n".join(out)


table = LinearProbing(size=11)
for name, mark in [("Aarti", 78), ("Divya", 55), ("Chetan", 90),
                   ("Eshan", 67), ("Farhan", 41)]:
    probes = table.put(name, mark)
    print(f"  put {name:<8} took {probes} probe(s)")

print()
print(table.show())
print()
for name in ["Aarti", "Farhan", "Nobody"]:
    value, probes = table.get(name)
    print(f"  get {name:<8} -> {str(value):<6} in {probes} probe(s)")

print()
print("now the reason a tombstone is needed:")
table.remove("Divya")
print("  removed Divya, and its slot is a TOMBSTONE, not empty")
for name in ["Aarti", "Chetan", "Farhan"]:
    value, probes = table.get(name)
    print(f"  get {name:<8} still works -> {value}")
print("  if the slot had been set to EMPTY, a search that probed past it")
print("  would stop there and report a key that is still present as missing.")
munotes.in218

Practical 7: a Hash Table with Separate Chaining

  put Aarti    took 1 probe(s)
  put Divya    took 1 probe(s)
  put Chetan   took 1 probe(s)
  put Eshan    took 2 probe(s)
  put Farhan   took 1 probe(s)

    0: 'Eshan' = 67
    1: .
    2: .
    3: 'Divya' = 55
    4: .
    5: 'Chetan' = 90
    6: .
    7: .
    8: 'Farhan' = 41
    9: .
   10: 'Aarti' = 78

  get Aarti    -> 78     in 1 probe(s)
  get Farhan   -> 41     in 1 probe(s)
  get Nobody   -> None   in 2 probe(s)

now the reason a tombstone is needed:
  removed Divya, and its slot is a TOMBSTONE, not empty
  get Aarti    still works -> 78
  get Chetan   still works -> 90
  get Farhan   still works -> 41
  if the slot had been set to EMPTY, a search that probed past it
  would stop there and report a key that is still present as missing.

The tombstone is the point of that listing. In open addressing a search stops at the first empty slot, so emptying a slot in the middle of a probe sequence cuts the sequence in half and hides everything after it. Marking it DELETED keeps the search going while still allowing the slot to be reused.

Separate chaining, MU'sOpen addressing
A collision goesinto a list in the same bucketinto another slot
Load factor may exceed 1yesno, never
Deletionremove from the listneeds a tombstone
Extra memorya list per bucketnone
Cache behaviourworse, the chain is scatteredbetter, the slots are contiguous
Simpler to writeyesno

Separate chaining is simpler, allows a load factor above 1, and deletes cleanly. That is why it is what MU asks for, and why it is what you should write.

Where a hash table is the right answer

from hashtable import HashTable

print("a hash table gives, on average:")
print("  lookup by key      O(1)")
print("  insert             O(1)")
print("  delete             O(1)")
print("  keys in order      O(n log n), it has to sort them")
print("  the smallest key   O(n), it has to look at all of them")
print()
print("a binary search tree gives:")
print("  lookup by key      O(log n)")
print("  insert             O(log n)")
print("  keys in order      O(n), the in-order walk, FREE")
print("  the smallest key   O(log n), the leftmost node")
print()
print("so: a hash table when you look up by key and do not care about order;")
print("    a BST when you need the order as well.")
munotes.in219

Practical 7: a Hash Table with Separate Chaining

a hash table gives, on average:
  lookup by key      O(1)
  insert             O(1)
  delete             O(1)
  keys in order      O(n log n), it has to sort them
  the smallest key   O(n), it has to look at all of them

a binary search tree gives:
  lookup by key      O(log n)
  insert             O(log n)
  keys in order      O(n), the in-order walk, FREE
  the smallest key   O(log n), the leftmost node

so: a hash table when you look up by key and do not care about order;
    a BST when you need the order as well.

That contrast is MU's Course Objective 8 in one page. And Python's own dict and set are hash tables, which is why x in some_dict is O(1) while x in some_list is O(n), as [Practical 7: a Common Member, and a Dictionary Sorted by Value] measured.

Procedure

  1. Save hashtable.py with a hash function written out by hand, not Python's hash, and the

HashTable class with a list of lists as buckets.

  1. Write bucket_of as the hash modulo the number of buckets.
  2. Write put so that it searches the chain first and replaces an existing key rather than adding

a second entry.

  1. Write get counting the comparisons it makes inside the chain, and remove raising KeyError.
  2. Write show to print every bucket with its chain, including the empty ones.
  3. Store at least six pairs, printing the bucket each one went to and marking the collisions.
  4. Print the whole table and then retrieve several keys, printing the comparison count each time.
  5. Store the same key twice with different values and show that the length does not change.
  6. Compare three hash functions on the same twenty keys, tabulate the chain lengths and the longest

chain, then show separately that the sum of codes gives one number for every rearrangement.

  1. Add rehashing at a load factor of 0.75 and print the load factor and the bucket count after each

insertion.

Result

Six pairs stored into seven buckets, with the bucket printed for each and the collisions marked. The whole table printed with its chains. Retrieval took one comparison in a bucket of one pair and more in a chain, confirming that only the key's own bucket is searched. Storing an existing key replaced its value and left the length unchanged. On twenty names into seven buckets the longest chains were 6 for the first-letter hash, 4 for the sum of codes and 5 for the polynomial, reported as measured rather than as predicted: twenty keys in seven buckets cannot rank two reasonable hash functions. The sum of codes was nevertheless shown to be disqualified, because every rearrangement of the same characters hashes to one number, so Anil, Lina, Nail and Lani collide in any table of any size. The collision probability table reproduced the birthday problem: 23 keys in 365 buckets collide more often than not. With a load limit of 0.75 the bucket count doubled whenever the load factor passed it, and the longest chain stayed short. Isolating the load factor showed the average comparisons tracking the load factor rather than the number of items: 1.00 at a load of 0.1, 1.59 at 1.0, 5.74 at 10, and 5000.50 with a single bucket, which is n over two and is exactly a linked list.

munotes.in220

Practical 7: a Hash Table with Separate Chaining

Where marks are lost

  • Using Python's hash() on strings in an example whose bucket numbers are printed. It is salted

per process, so the output differs every run.

  • Not searching the chain in put, so the same key appears twice and get returns the older

value.

  • No collision in the test data. The exercise is collision handling; the entry must show one.
  • Calling a collision an error. It is expected, and with more keys than buckets it is certain.
  • Judging a hash function by how clever it looks instead of by the longest chain it produces.
  • Adding the character codes as the hash, which gives every anagram the same bucket.
  • No load factor and no rehashing, so the chains grow and the lookup quietly becomes O(n).
  • Saying a hash table is O(1) without "on average, with a good hash and a low load factor".
  • Confusing separate chaining with open addressing. MU asked for chaining.

For the journal

The aim in MU's words, both bullets. The bucket picture with one chain visible in it. The three part table: hash function, bucket index, collision handling. Then the hash function written out, with the abc, cba, bca run showing why adding the codes is wrong. Then the class, and the storing run with the bucket number and the collisions marked, and the whole table printed. Then the retrieval run with the comparison counts. Then the collision probability table with the birthday problem row, and one sentence: a collision is expected, not an accident. Then the three hash functions compared by longest chain, and the load factor run showing the buckets doubling. The conclusion: the hash gives the bucket by arithmetic and only that bucket's chain is searched, so a lookup costs the length of one chain, which the load factor keeps near one.

munotes.in221

Practical 7: a Hash Table with Separate Chaining

Quick revision

  • Three parts: a hash function, the bucket index (hash modulo the bucket count), and

collision handling.

  • Separate chaining: each bucket holds a list of (key, value) pairs.
  • Write your own hash function for an example: Python's hash() of a string is salted per process.
  • Adding the character codes is a bad hash: every anagram collides. Multiply by a base each step so

position matters.

  • Judge a hash function by the longest chain on your own data, and reject outright any whose

collisions are built into the arithmetic: the sum of codes collides on every rearrangement, in every table of every size.

  • put must search the chain first and replace, or a key appears twice.
  • get searches only its own bucket's chain.
  • A collision is expected. With more keys than buckets it is certain (pigeonhole). 23 keys in 365

buckets collide more often than not, which is the birthday problem.

  • load factor = items / buckets. It is the number that predicts the lookup cost.
  • Rehashing: when the load factor passes a limit (0.75 is usual), double the buckets and place

every key again, because the index depends on the bucket count. O(n), rare, so insertion stays O(1) amortised.

  • Cost: O(1) average, O(n) worst, and the worst case is everything in one bucket.
  • Open addressing is the alternative: a colliding key goes to another slot, the load factor cannot

exceed 1, and deletion needs a tombstone or a probe sequence is cut short.

  • Python's dict and set are hash tables, which is why in on them is O(1).

Questions you should be able to answer

1. What are the three parts of a hash table? A hash function that turns a key into a number, the modulo that turns the number into a bucket index, and a way of handling two keys landing in the same bucket.

2. What is separate chaining? Each bucket holds a list of the pairs that hashed to it, and a lookup searches only that list.

3. Why is adding the character codes a bad hash function? Because addition ignores order, so every anagram gives the same number. abc, cba and bca all land in the same bucket.

4. How do you judge a hash function? By the longest chain it produces on your own data, because that is what a lookup costs in the worst case. But first reject any whose collisions are structural: measured here, the sum of codes gave the shortest chains on twenty names and is still wrong, because every rearrangement of the same characters hashes to one number.

munotes.in222

Practical 7: a Hash Table with Separate Chaining

5. What is the load factor and why does it matter? Items divided by buckets. The average number of comparisons in a lookup tracks it, so keeping it near 1 keeps the lookup O(1).

6. What is rehashing and why is it needed? Doubling the bucket count and placing every key again. It is needed because the index is the hash modulo the bucket count, so every index changes when the count does, and it keeps the chains short.

7. Is a collision an error? No. It is expected, and once there are more keys than buckets it is certain by the pigeonhole principle. Long before that it is likely: 23 keys in 365 buckets collide more often than not.

8. What does a lookup cost? O(1) on average, given a good hash function and a load factor kept low. O(n) in the worst case, when every key lands in one bucket and the table is a linked list.

9. Why must put search the chain before appending? Otherwise storing an existing key adds a second pair with the same key, and get returns whichever it finds first.

10. What is a tombstone, and which method needs one? A marker left where a pair was deleted, needed by open addressing, because a search stops at the first empty slot and emptying a slot mid-sequence would hide every key after it. Separate chaining does not need one.

11. When would you choose a BST over a hash table? When the order of the keys matters. A BST gives the sorted order free in an in-order walk and the smallest key in O(log n); a hash table has to sort, at O(n log n), and scan for the smallest.

Contents This chapter on its own page

munotes.in223

Chapter Thirty-Two

Practical 8: Bubble, Insertion and Selection Sort Compared

Syllabus topic Module 2, practical 8, "Sorting Algorithms: Write programs to implement and compare the following sorting algorithms: Bubble sort. Insertion sort. Selection sort."

Aim

To implement bubble sort, insertion sort and selection sort, and to compare them.

The three in one sentence each

SortWhat it doesHow to remember it
bubblerepeatedly swap adjacent items that are out of order, so the largest "bubbles" to the endit only ever compares neighbours
insertiontake each item and slide it back into its place among the items already sortedthe way you sort a hand of playing cards
selectionfind the smallest of what is left and swap it into placethe way you would sort by eye

All three split the list into a sorted part and an unsorted part and move the boundary one step per pass. The difference is what each one does on a pass.

bubble     [ unsorted ................ | sorted ]   grows from the RIGHT, by swapping neighbours
insertion  [ sorted ... | unsorted .............]   grows from the LEFT, by sliding one item back
selection  [ sorted ... | unsorted .............]   grows from the LEFT, by finding the minimum

The three programs, instrumented

"""The three sorts MU names, each counting its comparisons and movements."""


class Counter:
    """How much work a sort did. Movements are what actually costs time."""

    def __init__(self, name):
        self.name = name
        self.comparisons = 0
        self.movements = 0
        self.passes = 0

    def __repr__(self):
        return (f"{self.name}: {self.comparisons} comparison(s), "
                f"{self.movements} movement(s), {self.passes} pass(es)")


def bubble_sort(data, counter=None, trace=False):
    """Swap adjacent items out of order. The largest bubbles to the end each pass."""
    data = list(data)
    counter = counter or Counter("bubble")
    n = len(data)
    for i in range(n - 1):
        swapped = False
        counter.passes += 1
        for j in range(n - 1 - i):
            counter.comparisons += 1
            if data[j] > data[j + 1]:
                data[j], data[j + 1] = data[j + 1], data[j]
                counter.movements += 2          # a swap moves two items
                swapped = True
        if trace:
            print(f"    pass {counter.passes}: {data}   "
                  f"{'swapped' if swapped else 'nothing swapped, STOP'}")
        if not swapped:
            break                               # already sorted: this is the cheap exit
    return data, counter


def bubble_sort_naive(data, counter=None):
    """The version with no early exit. Always makes every pass."""
    data = list(data)
    counter = counter or Counter("bubble, no early exit")
    n = len(data)
    for i in range(n - 1):
        counter.passes += 1
        for j in range(n - 1 - i):
            counter.comparisons += 1
            if data[j] > data[j + 1]:
                data[j], data[j + 1] = data[j + 1], data[j]
                counter.movements += 2
    return data, counter


def insertion_sort(data, counter=None, trace=False):
    """Slide each item back into its place among those already sorted."""
    data = list(data)
    counter = counter or Counter("insertion")
    for i in range(1, len(data)):
        counter.passes += 1
        value = data[i]
        j = i - 1
        while j >= 0:
            counter.comparisons += 1
            if data[j] <= value:
                break
            data[j + 1] = data[j]
            counter.movements += 1
            j -= 1
        data[j + 1] = value
        counter.movements += 1
        if trace:
            sorted_part = " ".join(str(x) for x in data[:i + 1])
            rest = " ".join(str(x) for x in data[i + 1:])
            print(f"    pass {counter.passes}: [{sorted_part}] {rest}")
    return data, counter


def selection_sort(data, counter=None, trace=False):
    """Find the smallest of what is left and swap it into place."""
    data = list(data)
    counter = counter or Counter("selection")
    n = len(data)
    for i in range(n - 1):
        counter.passes += 1
        smallest = i
        for j in range(i + 1, n):
            counter.comparisons += 1
            if data[j] < data[smallest]:
                smallest = j
        if smallest != i:
            data[i], data[smallest] = data[smallest], data[i]
            counter.movements += 2
        if trace:
            sorted_part = " ".join(str(x) for x in data[:i + 1])
            rest = " ".join(str(x) for x in data[i + 1:])
            print(f"    pass {counter.passes}: [{sorted_part}] {rest}"
                  f"   smallest was at {smallest}")
    return data, counter


SORTS = {
    "bubble": bubble_sort,
    "insertion": insertion_sort,
    "selection": selection_sort,
}
munotes.in224

Practical 8: Bubble, Insertion and Selection Sort Compared

Three details in that file earn marks on their own.

Bubble's inner loop is range(n - 1 - i), not range(n - 1). After pass i the last i items are already in their final places, so comparing them again is wasted work. Almost every printed answer gets this wrong.

Bubble's swapped flag is the early exit. If a pass makes no swap the list is sorted, and the sort stops. That is what makes bubble sort O(n) on already sorted data, and without it bubble sort has no redeeming feature at all.

A swap counts as two movements and a slide as one. Insertion sort slides rather than swaps, which is why it moves fewer items than bubble sort for the same number of inversions.

Each one traced on the same small array

from sorts import bubble_sort, insertion_sort, selection_sort

data = [5, 2, 9, 1, 6]
print("starting from", data)
print()

for name, function in [("bubble", bubble_sort), ("insertion", insertion_sort),
                       ("selection", selection_sort)]:
    print(f"  {name} sort:")
    result, counter = function(data, trace=True)
    print(f"    result {result}")
    print(f"    {counter}")
    print()
starting from [5, 2, 9, 1, 6]

  bubble sort:
    pass 1: [2, 5, 1, 6, 9]   swapped
    pass 2: [2, 1, 5, 6, 9]   swapped
    pass 3: [1, 2, 5, 6, 9]   swapped
    pass 4: [1, 2, 5, 6, 9]   nothing swapped, STOP
    result [1, 2, 5, 6, 9]
    bubble: 10 comparison(s), 10 movement(s), 4 pass(es)

  insertion sort:
    pass 1: [2 5] 9 1 6
    pass 2: [2 5 9] 1 6
    pass 3: [1 2 5 9] 6
    pass 4: [1 2 5 6 9]
    result [1, 2, 5, 6, 9]
    insertion: 7 comparison(s), 9 movement(s), 4 pass(es)

  selection sort:
    pass 1: [1] 2 9 5 6   smallest was at 3
    pass 2: [1 2] 9 5 6   smallest was at 1
    pass 3: [1 2 5] 9 6   smallest was at 3
    pass 4: [1 2 5 6] 9   smallest was at 4
    result [1, 2, 5, 6, 9]
    selection: 10 comparison(s), 6 movement(s), 4 pass(es)
munotes.in225

Practical 8: Bubble, Insertion and Selection Sort Compared

Read the three traces against each other.

Bubble moves the largest item to the end on each pass, and you can watch 9 travel right. Insertion has a sorted block on the left that grows by one item a pass, and the item joins it by sliding back. Selection also grows a sorted block on the left, but it finds the smallest remaining item and swaps it in, so it makes at most one swap per pass.

The three on the same data, counted

import random
from sorts import bubble_sort, bubble_sort_naive, insertion_sort, selection_sort

random.seed(7)
n = 40
datasets = {
    "already sorted": list(range(n)),
    "reverse sorted": list(range(n, 0, -1)),
    "random": random.sample(range(n * 3), n),
    "nearly sorted": list(range(n)),
}
datasets["nearly sorted"][5], datasets["nearly sorted"][6] = (
    datasets["nearly sorted"][6], datasets["nearly sorted"][5])

print(f"n = {n}, and n(n-1)/2 = {n * (n - 1) // 2}")
print()
for label, data in datasets.items():
    print(f"  {label}:")
    print(f"    {'sort':<22} {'comparisons':>12} {'movements':>11} {'passes':>8}")
    for name, function in [("bubble, early exit", bubble_sort),
                          ("bubble, no early exit", bubble_sort_naive),
                          ("insertion", insertion_sort),
                          ("selection", selection_sort)]:
        result, counter = function(data)
        assert result == sorted(data), f"{name} did not sort correctly"
        print(f"    {name:<22} {counter.comparisons:>12} "
              f"{counter.movements:>11} {counter.passes:>8}")
    print()
n = 40, and n(n-1)/2 = 780

  already sorted:
    sort                    comparisons   movements   passes
    bubble, early exit               39           0        1
    bubble, no early exit           780           0       39
    insertion                        39          39       39
    selection                       780           0       39

  reverse sorted:
    sort                    comparisons   movements   passes
    bubble, early exit              780        1560       39
    bubble, no early exit           780        1560       39
    insertion                       780         819       39
    selection                       780          40       39

  random:
    sort                    comparisons   movements   passes
    bubble, early exit              774         592       36
    bubble, no early exit           780         592       39
    insertion                       332         335       39
    selection                       780          68       39

  nearly sorted:
    sort                    comparisons   movements   passes
    bubble, early exit               77           2        2
    bubble, no early exit           780           2       39
    insertion                        40          40       39
    selection                       780           2       39

That table is the answer to MU's "and compare", and every row of it is worth a sentence.

On already sorted data, bubble sort with the early exit made n-1 comparisons and stopped after one pass. Insertion sort made n-1 comparisons too. Selection sort made the full n(n-1)/2, because it has no way to notice that the data is sorted: it looks for the minimum of the rest whatever happens.

On reverse sorted data, every sort is at its worst and all three make n(n-1)/2 comparisons. Look at the movements column: selection sort moved far fewer items than the other two, because it makes at most one swap per pass.

munotes.in226

Practical 8: Bubble, Insertion and Selection Sort Compared

On random data, insertion sort made about half as many comparisons as bubble or selection, because it stops sliding as soon as it meets something smaller.

On nearly sorted data, insertion sort is dramatically the best, which is the reason it is still used in real libraries for small or nearly sorted pieces.

The assert in that listing is not decoration

Every run above is checked against Python's own sorted(). A sort that returns the wrong answer would stop the program rather than print a plausible comparison count. Put that assert in your own program, because a count from a sort that does not sort is worse than no count at all.

Stability, which is the one property that is not about speed

A sort is stable when two items that compare equal keep the order they were already in. It matters whenever the items carry more than the key being sorted on.

def bubble_sort_pairs(data):
    data = list(data)
    n = len(data)
    for i in range(n - 1):
        for j in range(n - 1 - i):
            if data[j][1] > data[j + 1][1]:
                data[j], data[j + 1] = data[j + 1], data[j]
    return data


def insertion_sort_pairs(data):
    data = list(data)
    for i in range(1, len(data)):
        value = data[i]
        j = i - 1
        while j >= 0 and data[j][1] > value[1]:
            data[j + 1] = data[j]
            j -= 1
        data[j + 1] = value
    return data


def selection_sort_pairs(data):
    data = list(data)
    n = len(data)
    for i in range(n - 1):
        smallest = i
        for j in range(i + 1, n):
            if data[j][1] < data[smallest][1]:
                smallest = j
        data[i], data[smallest] = data[smallest], data[i]
    return data


# Aarti and Bhavesh are both on 55, and Chetan's 41 is BELOW both of them.
# That last detail is what makes selection sort show its instability: its first
# swap brings Chetan to the front and throws Aarti to where Chetan was, which is
# behind Bhavesh.
students = [("Aarti", 55), ("Bhavesh", 55), ("Chetan", 41),
            ("Farhan", 90), ("Gauri", 78)]

print("original order:")
for name, mark in students:
    print(f"    {name:<9} {mark}")
print()

for label, function in [("bubble", bubble_sort_pairs),
                        ("insertion", insertion_sort_pairs),
                        ("selection", selection_sort_pairs)]:
    result = function(students)
    tied = [name for name, mark in result if mark == 55]
    stable = tied == ["Aarti", "Bhavesh"]
    print(f"  {label:<10} order of the two 55s: {', '.join(tied):<18} "
          f"stable? {stable}")
    print(f"             {[n for n, _ in result]}")
original order:
    Aarti     55
    Bhavesh   55
    Chetan    41
    Farhan    90
    Gauri     78

  bubble     order of the two 55s: Aarti, Bhavesh     stable? True
             ['Chetan', 'Aarti', 'Bhavesh', 'Gauri', 'Farhan']
  insertion  order of the two 55s: Aarti, Bhavesh     stable? True
             ['Chetan', 'Aarti', 'Bhavesh', 'Gauri', 'Farhan']
  selection  order of the two 55s: Bhavesh, Aarti     stable? False
             ['Chetan', 'Bhavesh', 'Aarti', 'Gauri', 'Farhan']
munotes.in227

Practical 8: Bubble, Insertion and Selection Sort Compared

Bubble and insertion sort are stable; this selection sort is not. The reason is in the code: bubble and insertion only move an item past another when it is strictly greater, so equal items never cross. Selection sort swaps an item from far away into position, and that swap can jump it over an equal item.

That matters in practice. Sort students by name, then by mark with a stable sort, and within each mark they stay in name order. With an unstable sort the second sort destroys the first.

The comparison must be > and not >= for stability. Changing one character to >= makes bubble and insertion sort unstable and makes them do more work at the same time, which is an unusually cheap mistake to avoid.

The formulas, checked against the counts

from sorts import bubble_sort_naive, insertion_sort, selection_sort

print(f"  {'n':>5} {'n(n-1)/2':>10} {'bubble':>9} {'selection':>11} "
      f"{'insertion worst':>17}")
for n in [5, 10, 20, 40, 80]:
    formula = n * (n - 1) // 2
    reversed_data = list(range(n, 0, -1))
    _, bubble = bubble_sort_naive(reversed_data)
    _, selection = selection_sort(reversed_data)
    _, insertion = insertion_sort(reversed_data)
    print(f"  {n:>5} {formula:>10} {bubble.comparisons:>9} "
          f"{selection.comparisons:>11} {insertion.comparisons:>17}")

print()
print("all three worst cases equal n(n-1)/2 exactly.")
print()
print(f"  {'n':>5} {'comparisons':>13} {'growth when n doubles':>23}")
previous = None
for n in [10, 20, 40, 80, 160]:
    _, counter = selection_sort(list(range(n, 0, -1)))
    growth = "" if previous is None else f"{counter.comparisons / previous:.2f}x"
    print(f"  {n:>5} {counter.comparisons:>13} {growth:>23}")
    previous = counter.comparisons
print()
print("about 4x each time n doubles, which is what O(n squared) means")
      n   n(n-1)/2    bubble   selection   insertion worst
      5         10        10          10                10
     10         45        45          45                45
     20        190       190         190               190
     40        780       780         780               780
     80       3160      3160        3160              3160

all three worst cases equal n(n-1)/2 exactly.

      n   comparisons   growth when n doubles
     10            45
     20           190                   4.22x
     40           780                   4.11x
     80          3160                   4.05x
    160         12720                   4.03x

about 4x each time n doubles, which is what O(n squared) means
Comparisons, bestComparisons, worstMovements, worst
bubble, with early exitn - 1n(n-1)/2about n squared
insertionn - 1n(n-1)/2about n squared
selectionn(n-1)/2n(n-1)/2at most 2(n-1)
StableNotices sorted data
bubble, with early exityesyes
insertionyesyes
selectionnono

All three are O(n squared) in the worst case, and the counted growth of about 4x per doubling of n is what that means. Bubble and insertion are O(n) in the best case; selection is O(n squared) even then.

So which is best, and the honest answer

import random
from sorts import bubble_sort, insertion_sort, selection_sort

random.seed(7)
n = 60
cases = {
    "sorted": list(range(n)),
    "reversed": list(range(n, 0, -1)),
    "random": random.sample(range(n * 4), n),
}

print(f"  {'case':<10} {'fewest comparisons':<34} {'fewest movements':<30}")
for label, data in cases.items():
    results = {}
    for name, function in [("bubble", bubble_sort), ("insertion", insertion_sort),
                           ("selection", selection_sort)]:
        _, counter = function(data)
        results[name] = counter
    fewest_comparisons = min(c.comparisons for c in results.values())
    fewest_movements = min(c.movements for c in results.values())
    comparison_winners = [n for n, c in results.items()
                          if c.comparisons == fewest_comparisons]
    movement_winners = [n for n, c in results.items()
                        if c.movements == fewest_movements]
    print(f"  {label:<10} "
          f"{'/'.join(comparison_winners) + f' ({fewest_comparisons})':<30} "
          f"{'/'.join(movement_winners) + f' ({fewest_movements})':<30}")

print()
print("note the TIES, which a single 'winner' would have hidden:")
print("  on sorted data bubble and insertion tie at n-1 comparisons")
print("  on reversed data all three tie at n(n-1)/2, the worst case for each")
print("  on random data insertion sort is clearly ahead on comparisons")
print("  selection sort is never beaten on MOVEMENTS")
munotes.in228

Practical 8: Bubble, Insertion and Selection Sort Compared

  case       fewest comparisons                 fewest movements
  sorted     bubble/insertion (59)          bubble/selection (0)
  reversed   bubble/insertion/selection (1770) selection (60)
  random     insertion (788)                selection (112)

note the TIES, which a single 'winner' would have hidden:
  on sorted data bubble and insertion tie at n-1 comparisons
  on reversed data all three tie at n(n-1)/2, the worst case for each
  on random data insertion sort is clearly ahead on comparisons
  selection sort is never beaten on MOVEMENTS

Selection sort wins on movements in every case, and insertion sort wins on comparisons except on reversed data. That is the shape of the real answer, and a student who says "insertion sort is best" without qualification has answered a different question.

The practical rule:

ChooseWhen
insertion sortthe data is small, or nearly sorted. It is what real libraries use for short pieces
selection sortmoving an item is expensive and comparing is cheap, because it makes at most n-1 swaps
bubble sortnever, in real work. It is a teaching algorithm
a library sortevery other time. Python's sorted is O(n log n) and written in C
import random
import timeit
from sorts import bubble_sort, insertion_sort, selection_sort

random.seed(7)
data = random.sample(range(4000), 800)

print("800 random items:")
for name, function in [("bubble", bubble_sort), ("insertion", insertion_sort),
                       ("selection", selection_sort)]:
    _, counter = function(data)
    print(f"  {name:<10} {counter.comparisons:>8} comparisons")

print(f"  {'python sorted()':<10} makes about n log n, which is roughly "
      f"{int(800 * 9.6)} comparisons")
print()
print("the counted ratio of the O(n squared) sorts to n log n at this size:")
_, counter = insertion_sort(data)
print(f"  insertion {counter.comparisons} against about {int(800 * 9.6)}, "
      f"a factor of {counter.comparisons / (800 * 9.6):.0f}")
print()
print("that factor grows with n, which is why nobody writes an O(n squared)")
print("sort for real data. It is here so the cost is understood.")
800 random items:
  bubble       319447 comparisons
  insertion    154690 comparisons
  selection    319600 comparisons
  python sorted() makes about n log n, which is roughly 7680 comparisons

the counted ratio of the O(n squared) sorts to n log n at this size:
  insertion 154690 against about 7680, a factor of 20

that factor grows with n, which is why nobody writes an O(n squared)
sort for real data. It is here so the cost is understood.
munotes.in229

Practical 8: Bubble, Insertion and Selection Sort Compared

Procedure

  1. Save sorts.py with a Counter class and the three sorts, each taking the counter and a trace

flag.

  1. In bubble sort, make the inner loop range(n - 1 - i) and add the swapped early exit.
  2. In insertion sort, slide rather than swap, and count one movement per slide.
  3. In selection sort, find the index of the smallest and swap once per pass.
  4. Count a swap as two movements and a slide as one.
  5. Trace all three on the same five element array, printing the array after every pass.
  6. Run all four versions on already sorted, reverse sorted, random and nearly sorted data of the same

size, and tabulate the comparisons, movements and passes.

  1. assert every result equals sorted(data), so a wrong sort cannot print a plausible count.
  2. Check the worst case counts against n(n-1)/2 for several n, and check the growth factor as n doubles.
  3. Test stability with two items of equal key and record which sorts preserve their order.

Result

All three sorts produced correctly sorted output on every dataset, checked by assert against Python's sorted(). Traced on [5, 2, 9, 1, 6], bubble sort carried 9 to the end on the first pass, insertion sort grew a sorted block by sliding each item back, and selection sort made at most one swap per pass. On 40 items: bubble sort with the early exit stopped after one pass on sorted data while selection sort still made the full n(n-1)/2 = 780 comparisons, since it cannot notice sorted input. On reversed data all three made 780 comparisons, and selection sort moved far fewer items. On random data insertion sort made the fewest comparisons. The worst case comparison count equalled n(n-1)/2 exactly for all three at every n tried, and the count grew by about a factor of four as n doubled. Bubble and insertion sort preserved the order of Aarti and Bhavesh, both on 55; selection sort reversed them, because its first swap threw Aarti to the position Chetan had vacated, behind Bhavesh.

Where marks are lost

  • No comparison. MU's row says "and compare", and three programs alone is a fraction of the answer.
  • Bubble's inner loop as range(n - 1), which re-compares items already in place.
  • No early exit in bubble sort, which is its only redeeming feature.
  • Swapping in insertion sort instead of sliding, which doubles the movements.
  • Not counting anything. The comparison has to be numbers, not adjectives.
  • No assert against sorted(), so a broken sort can print a convincing table.
  • One dataset. Sorted, reversed and random behave completely differently, which is the point.
  • Claiming selection sort improves on sorted data. It does not: it is O(n squared) always.
  • Saying "insertion sort is best" without saying on what and by which measure.
  • Not testing stability, or claiming selection sort is stable.
munotes.in230

Practical 8: Bubble, Insertion and Selection Sort Compared

For the journal

The aim in MU's words, all three sorts and the word compare. The three sentence table of what each one does, and the three line picture of where the sorted part grows. Then the code, with bubble's range(n - 1 - i) and its swapped flag pointed out. Then all three traced on the same five element array, which is the clearest page in the entry. Then the counted table across sorted, reversed, random and nearly sorted data, with a sentence per row. Then the worst case counts beside n(n-1)/2 and the growth factor as n doubles. Then the stability test with the two tied students. The conclusion: all three are O(n squared) in the worst case, bubble and insertion are O(n) on sorted data and selection is not, selection makes the fewest movements, insertion makes the fewest comparisons on nearly sorted data, and bubble sort is a teaching algorithm rather than a useful one.

Quick revision

  • bubble: swap adjacent items out of order; the largest bubbles to the end each pass.
  • insertion: slide each item back into the sorted part. The playing card method.
  • selection: find the smallest of the rest and swap it into place.
  • Bubble's inner loop is range(n - 1 - i), and the swapped flag stops it early.
  • Insertion slides (one movement each); bubble and selection swap (two each).
  • Worst case comparisons: n(n-1)/2 for all three. Verified exactly for n from 5 to 80.
  • Best case: n - 1 for bubble with the early exit and for insertion. n(n-1)/2 for selection, which

cannot notice sorted data.

  • Movements, worst case: about n squared for bubble and insertion, at most 2(n-1) for selection.
  • Stable: bubble and insertion. Not stable: selection. Stability needs >, never >=.
  • All three are O(n squared), and the count grows by about 4 when n doubles.
  • Choose insertion for small or nearly sorted data, selection when moving is expensive, a library sort

otherwise.

  • assert result == sorted(data) in the program, or a wrong sort prints a plausible table.
munotes.in231

Practical 8: Bubble, Insertion and Selection Sort Compared

Questions you should be able to answer

1. Describe each of the three in one sentence. Bubble sort swaps adjacent out-of-order items so the largest moves to the end each pass. Insertion sort slides each item back into its place among those already sorted. Selection sort finds the smallest remaining item and swaps it into position.

2. Why is bubble sort's inner loop range(n - 1 - i)? Because after pass i the last i items are already in their final positions, so comparing them again is wasted work.

3. What does the swapped flag do, and why does it matter? It stops the sort when a pass makes no swap, which means the list is already sorted. It is what makes bubble sort O(n) on sorted data, and without it bubble sort has no advantage over the other two at all.

4. What are the worst case comparison counts? n(n-1)/2 for all three, confirmed exactly here for n from 5 to 80.

5. Which sorts improve on already sorted data, and which does not? Bubble with the early exit and insertion both drop to n-1 comparisons. Selection sort still makes n(n-1)/2, because it always searches the rest of the list for the minimum.

6. Which sort makes the fewest movements, and why? Selection sort, at most one swap per pass, so at most 2(n-1) movements. It finds the right item first and moves it once.

7. What is a stable sort? One in which two items that compare equal keep the order they were already in.

8. Which of the three are stable? Bubble and insertion. Selection sort is not, because it swaps an item in from a distance and that swap can jump it over an equal item.

9. Why does >= instead of > break stability? Because the sort then moves an item past one that is equal to it, so equal items can cross. It also does more work for no benefit.

10. So which of the three is best? None outright. Insertion sort makes the fewest comparisons except on reversed data, selection sort makes the fewest movements in every case, and bubble sort wins nothing and is taught because it is the easiest to explain. In real work you use the library sort, which is O(n log n).

11. Why assert result == sorted(data)? Because a sort that does not sort still produces comparison counts, and a plausible table from a broken program is worse than no table.

Contents This chapter on its own page

munotes.in232

Chapter Thirty-Three

Practical 9: Linear and Binary Search Compared

Syllabus topic Module 2, practical 9, "Searching Algorithms: Write programs to implement and compare: Linear search. Binary search (on a sorted array)."

Aim

To implement linear search and binary search on a sorted array, and to compare them.

The two ideas

Linear search looks at every item in turn until it finds the target or runs out. It needs nothing of the data.

Binary search looks at the middle item, and one comparison tells it which half cannot contain the target, so that half is discarded. It requires the data to be sorted.

linear, looking for 23 in [4, 8, 15, 16, 23, 42]:
  4? 8? 15? 16? 23 -> found, in 5 comparisons

binary, the same:
  low=0 high=5 middle=2 -> 15 < 23, so discard 4, 8, 15
  low=3 high=5 middle=4 -> 23 = 23, found, in 2 comparisons

That is the same halving as a binary search tree in [Practical 5: the Binary Search Tree, Create, Insert and Search], done with arithmetic on indexes instead of with references.

The programs

"""Linear and binary search, each counting its comparisons."""


def linear_search(data, target):
    """Look at every item in turn. Returns (index, comparisons), -1 if absent."""
    comparisons = 0
    for index, item in enumerate(data):
        comparisons += 1
        if item == target:
            return index, comparisons
    return -1, comparisons


def linear_search_sorted(data, target):
    """On SORTED data, stop as soon as the items pass the target."""
    comparisons = 0
    for index, item in enumerate(data):
        comparisons += 1
        if item == target:
            return index, comparisons
        if item > target:
            return -1, comparisons          # everything after this is larger
    return -1, comparisons


def binary_search(data, target, trace=False):
    """Halve the range each step. The data MUST be sorted.

    Returns (index, comparisons), -1 if absent.
    """
    low = 0
    high = len(data) - 1
    comparisons = 0
    while low <= high:
        middle = low + (high - low) // 2        # never overflows, see the note
        comparisons += 1
        if trace:
            print(f"    low {low:>3} high {high:>3} middle {middle:>3} "
                  f"data[middle] {data[middle]:>4}  "
                  f"{'FOUND' if data[middle] == target else ('go right' if data[middle] < target else 'go left')}")
        if data[middle] == target:
            return middle, comparisons
        if data[middle] < target:
            low = middle + 1
        else:
            high = middle - 1
    return -1, comparisons


def binary_search_recursive(data, target, low=0, high=None, comparisons=0):
    """The same thing recursively. Returns (index, comparisons)."""
    if high is None:
        high = len(data) - 1
    if low > high:
        return -1, comparisons
    middle = low + (high - low) // 2
    comparisons += 1
    if data[middle] == target:
        return middle, comparisons
    if data[middle] < target:
        return binary_search_recursive(data, target, middle + 1, high, comparisons)
    return binary_search_recursive(data, target, low, middle - 1, comparisons)


def binary_search_first(data, target):
    """The FIRST position holding the target, when duplicates exist."""
    low, high = 0, len(data) - 1
    found = -1
    comparisons = 0
    while low <= high:
        middle = low + (high - low) // 2
        comparisons += 1
        if data[middle] == target:
            found = middle
            high = middle - 1              # keep looking to the LEFT
        elif data[middle] < target:
            low = middle + 1
        else:
            high = middle - 1
    return found, comparisons


def insertion_point(data, target):
    """Where the target would go if inserted, keeping the array sorted."""
    low, high = 0, len(data)
    comparisons = 0
    while low < high:
        middle = (low + high) // 2
        comparisons += 1
        if data[middle] < target:
            low = middle + 1
        else:
            high = middle
    return low, comparisons
munotes.in233

Practical 9: Linear and Binary Search Compared

Four details in that file are marks.

low <= high, not low < high. With < the loop misses a range of one item, so a target sitting alone at the end of a range is never found. It is the commonest bug in binary search.

low = middle + 1 and high = middle - 1, not low = middle and high = middle. Leaving the middle in the range means a range of two can never shrink, and the loop never ends.

middle = low + (high - low) // 2. In Python (low + high) // 2 is equally correct, because Python integers do not overflow. Write the safer form anyway and be able to say why: in C, Java or C++, low + high can overflow for a large array, and this exact bug sat in the Java standard library for nine years. That is a genuinely good viva answer.

The while loop is not the only version. The recursive form is here too, and so are the two variants an examiner asks for: the first occurrence when there are duplicates, and the insertion point.

Both, traced

from searches import linear_search, binary_search

data = [4, 8, 15, 16, 23, 42, 55, 67, 71, 89]
print("the sorted array:", data)
print("its indexes     :", list(range(len(data))))
print()

for target in [23, 4, 89, 30]:
    print(f"looking for {target}:")
    index, comparisons = linear_search(data, target)
    print(f"  linear: index {index}, {comparisons} comparison(s)")
    print(f"  binary:")
    index, comparisons = binary_search(data, target, trace=True)
    print(f"    index {index}, {comparisons} comparison(s)")
    print()
the sorted array: [4, 8, 15, 16, 23, 42, 55, 67, 71, 89]
its indexes     : [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]

looking for 23:
  linear: index 4, 5 comparison(s)
  binary:
    low   0 high   9 middle   4 data[middle]   23  FOUND
    index 4, 1 comparison(s)

looking for 4:
  linear: index 0, 1 comparison(s)
  binary:
    low   0 high   9 middle   4 data[middle]   23  go left
    low   0 high   3 middle   1 data[middle]    8  go left
    low   0 high   0 middle   0 data[middle]    4  FOUND
    index 0, 3 comparison(s)

looking for 89:
  linear: index 9, 10 comparison(s)
  binary:
    low   0 high   9 middle   4 data[middle]   23  go right
    low   5 high   9 middle   7 data[middle]   67  go right
    low   8 high   9 middle   8 data[middle]   71  go right
    low   9 high   9 middle   9 data[middle]   89  FOUND
    index 9, 4 comparison(s)

looking for 30:
  linear: index -1, 10 comparison(s)
  binary:
    low   0 high   9 middle   4 data[middle]   23  go right
    low   5 high   9 middle   7 data[middle]   67  go left
    low   5 high   6 middle   5 data[middle]   42  go left
    index -1, 3 comparison(s)
munotes.in234

Practical 9: Linear and Binary Search Compared

Read the binary traces. low and high close in on each other and each comparison halves the range, so ten items are searched in at most four comparisons. On a failed search the range shrinks to nothing and low passes high, which is what ends the loop.

Notice the row for the target 4: linear search found it in one comparison and binary search took three, because it had to halve its way down to index 0. That is not a defect; it is the one case where linear search wins.

The comparison counts, across sizes

from searches import linear_search, binary_search
from math import log2

print(f"  {'n':>8} {'linear worst':>13} {'binary worst':>13} {'log2(n)':>9} "
      f"{'binary, best':>13}")
for n in [10, 100, 1000, 10000, 100000, 1000000]:
    data = list(range(n))
    _, linear_worst = linear_search(data, -1)
    _, binary_worst = binary_search(data, -1)
    _, binary_best = binary_search(data, data[(n - 1) // 2])
    print(f"  {n:>8} {linear_worst:>13} {binary_worst:>13} {log2(n):>9.1f} "
          f"{binary_best:>13}")

print()
print("the binary column is about log2(n) + 1, and the linear column is n.")
print("at a million items: one comparison against a million.")
         n  linear worst  binary worst   log2(n)  binary, best
        10            10             3       3.3             1
       100           100             6       6.6             1
      1000          1000             9      10.0             1
     10000         10000            13      13.3             1
    100000        100000            16      16.6             1
   1000000       1000000            19      19.9             1

the binary column is about log2(n) + 1, and the linear column is n.
at a million items: one comparison against a million.

There is the whole comparison in one table. Multiplying n by ten adds about three comparisons to binary search and multiplies linear search by ten. At a million items binary search never needs more than about twenty comparisons, and linear search needs a million.

Linear searchBinary search
Needs the data sortednoyes
Best case1 comparison1 comparison
Worst casenabout log2(n) + 1
Average, successfulabout n/2about log2(n)
Works on a linked listyesno, it needs index arithmetic
Cost of an insertion into the structureO(1) at the endO(n), the array must stay sorted

The fifth row is the one students forget. Binary search needs random access, so it works on an array and not on a linked list: there is no way to reach the middle of a chain without walking to it, which would cost more than the search saves.

munotes.in235

Practical 9: Linear and Binary Search Compared

The honest comparison: counting the sort

MU's row says binary search is on a sorted array. So if the data is not already sorted, getting it sorted is part of the cost, and a fair comparison has to include it.

import random
from searches import linear_search, binary_search


def insertion_sort_counting(data):
    """The insertion sort of practical 8, returning its comparison count."""
    data = list(data)
    comparisons = 0
    for i in range(1, len(data)):
        value = data[i]
        j = i - 1
        while j >= 0:
            comparisons += 1
            if data[j] <= value:
                break
            data[j + 1] = data[j]
            j -= 1
        data[j + 1] = value
    return data, comparisons


def searching_unsorted(data, targets):
    """Linear search, on the data as it comes."""
    total = 0
    for target in targets:
        _, comparisons = linear_search(data, target)
        total += comparisons
    return total


def sorting_then_binary(data, targets):
    """Sort once, then binary search each target."""
    ordered, sort_cost = insertion_sort_counting(data)
    total = sort_cost
    for target in targets:
        _, comparisons = binary_search(ordered, target)
        total += comparisons
    return total, sort_cost


random.seed(7)
n = 500
data = random.sample(range(n * 4), n)

print(f"n = {n}, unsorted to begin with")
print(f"{'searches':>10} {'linear total':>14} {'sort + binary total':>21} "
      f"{'which wins':>12}")
for how_many in [1, 2, 5, 10, 50, 200, 1000]:
    targets = random.sample(data, min(how_many, n)) if how_many <= n \
        else [random.choice(data) for _ in range(how_many)]
    linear_total = searching_unsorted(data, targets)
    binary_total, sort_cost = sorting_then_binary(data, targets)
    winner = "linear" if linear_total < binary_total else "sort+binary"
    print(f"{how_many:>10} {linear_total:>14} {binary_total:>21} {winner:>12}")

_, sort_cost = sorting_then_binary(data, [])
print()
print(f"the sort alone cost {sort_cost} comparisons, which is the fixed price")
print("binary search has to pay before it can be used at all.")
n = 500, unsorted to begin with
  searches   linear total   sort + binary total   which wins
         1             38                 57748       linear
         2            452                 57763       linear
         5           1479                 57786       linear
        10           2514                 57829       linear
        50          11627                 58144       linear
       200          49660                 59335       linear
      1000         243862                 65728  sort+binary

the sort alone cost 57745 comparisons, which is the fixed price
binary search has to pay before it can be used at all.

That is the finding, and it is the one that separates a good answer from a recited one. For a small number of searches, linear search on the unsorted data wins outright, because the sort has to be paid for first. Binary search only pays for itself once there are enough searches to amortise the sort.

So the justification MU's Course Objective 8 asks for is:

ChooseWhen
linear searchthe data is unsorted and will be searched only a few times, or it is a linked list, or it is tiny
binary searchthe data is already sorted, or will be searched many times, and it is an array
a hash tableyou look up by an exact key many times and do not need order. O(1), see practical 7
a BSTyou need fast lookup and the sorted order. O(log n) for both
munotes.in236

Practical 9: Linear and Binary Search Compared

The bug low < high produces

def broken_binary(data, target):
    """WRONG: `low < high` misses a range of one item."""
    low, high = 0, len(data) - 1
    while low < high:                # the bug is this line
        middle = (low + high) // 2
        if data[middle] == target:
            return middle
        if data[middle] < target:
            low = middle + 1
        else:
            high = middle - 1
    return -1


from searches import binary_search

data = [10, 20, 30, 40, 50]
print(f"  {'target':>7} {'correct':>9} {'broken':>8}  agree?")
for target in data + [15, 99]:
    right, _ = binary_search(data, target)
    wrong = broken_binary(data, target)
    print(f"  {target:>7} {right:>9} {wrong:>8}  {right == wrong}")

print()
print("the broken version misses several values that ARE in the array,")
print("because when low and high meet on the answer the loop has already ended.")
   target   correct   broken  agree?
       10         0        0  True
       20         1       -1  False
       30         2        2  True
       40         3        3  True
       50         4       -1  False
       15        -1       -1  True
       99        -1       -1  True

the broken version misses several values that ARE in the array,
because when low and high meet on the answer the loop has already ended.

There is the bug, and there is why it survives testing: it finds some of the values. A student who tests one target that happens to work ships it.

Duplicates, and finding the first one

from searches import binary_search, binary_search_first, insertion_point

data = [10, 20, 20, 20, 20, 30, 40, 40, 50]
print("the array:", data)
print("indexes  :", list(range(len(data))))
print()

for target in [20, 40, 10, 50]:
    plain, c1 = binary_search(data, target)
    first, c2 = binary_search_first(data, target)
    print(f"  {target}: plain binary search found index {plain} in {c1} "
          f"comparison(s); the FIRST is index {first}, in {c2}")

print()
print("plain binary search finds SOME occurrence, not the first one.")
print()
print("and the insertion point, where a new value would go:")
for target in [5, 15, 20, 35, 60]:
    where, comparisons = insertion_point(data, target)
    print(f"  {target:>3} would be inserted at index {where}, "
          f"giving {data[:where] + [target] + data[where:]}")
the array: [10, 20, 20, 20, 20, 30, 40, 40, 50]
indexes  : [0, 1, 2, 3, 4, 5, 6, 7, 8]

  20: plain binary search found index 4 in 1 comparison(s); the FIRST is index 1, in 3
  40: plain binary search found index 6 in 2 comparison(s); the FIRST is index 6, in 3
  10: plain binary search found index 0 in 3 comparison(s); the FIRST is index 0, in 3
  50: plain binary search found index 8 in 4 comparison(s); the FIRST is index 8, in 4

plain binary search finds SOME occurrence, not the first one.

and the insertion point, where a new value would go:
    5 would be inserted at index 0, giving [5, 10, 20, 20, 20, 20, 30, 40, 40, 50]
   15 would be inserted at index 1, giving [10, 15, 20, 20, 20, 20, 30, 40, 40, 50]
   20 would be inserted at index 1, giving [10, 20, 20, 20, 20, 20, 30, 40, 40, 50]
   35 would be inserted at index 6, giving [10, 20, 20, 20, 20, 30, 35, 40, 40, 50]
   60 would be inserted at index 9, giving [10, 20, 20, 20, 20, 30, 40, 40, 50, 60]
munotes.in237

Practical 9: Linear and Binary Search Compared

Plain binary search finds an occurrence, not the first one. When duplicates exist and the position matters, keep searching left after a hit, which binary_search_first does with high = middle - 1 instead of returning.

The insertion point is the same halving used for a different question, and it is what Python's own bisect module provides:

import bisect

data = [10, 20, 20, 20, 20, 30, 40, 40, 50]

print("bisect_left (the first equal position) :", bisect.bisect_left(data, 20))
print("bisect_right (just past the last equal):", bisect.bisect_right(data, 20))
print("how many 20s                           :",
      bisect.bisect_right(data, 20) - bisect.bisect_left(data, 20))
print("where 35 would go                      :", bisect.bisect_left(data, 35))
print()
print("insort keeps a list sorted as items are added:")
running = []
for value in [30, 10, 50, 20, 40]:
    bisect.insort(running, value)
    print(f"  insert {value:>3} -> {running}")
bisect_left (the first equal position) : 1
bisect_right (just past the last equal): 5
how many 20s                           : 4
where 35 would go                      : 6

insort keeps a list sorted as items are added:
  insert  30 -> [30]
  insert  10 -> [10, 30]
  insert  50 -> [10, 30, 50]
  insert  20 -> [10, 20, 30, 50]
  insert  40 -> [10, 20, 30, 40, 50]

bisect_right(x) - bisect_left(x) counting the occurrences in O(log n) is a genuinely useful trick, and insort is how a sorted list is maintained. Both are the standard library, so nothing needs installing. Write the loop yourself for MU's row and mention bisect.

Recursive against iterative

from searches import binary_search, binary_search_recursive

data = list(range(1, 2001, 2))       # the odd numbers, so some targets are absent
print("the array holds the odd numbers from 1 to 1999, so", len(data), "items")
print(f"  {'target':>8} {'iterative':>22} {'recursive':>22}  agree?")
for target in [1, 999, 1999, 1000, 4000]:
    index_i, count_i = binary_search(data, target)
    index_r, count_r = binary_search_recursive(data, target)
    print(f"  {target:>8} {f'index {index_i}, {count_i} comps':>22} "
          f"{f'index {index_r}, {count_r} comps':>22}  "
          f"{(index_i, count_i) == (index_r, count_r)}")

print()
print("identical, as they must be: the recursion is the loop written another way.")
print("the iterative version uses no stack, so it has no depth limit.")
munotes.in238

Practical 9: Linear and Binary Search Compared

the array holds the odd numbers from 1 to 1999, so 1000 items
    target              iterative              recursive  agree?
         1       index 0, 9 comps       index 0, 9 comps  True
       999     index 499, 1 comps     index 499, 1 comps  True
      1999    index 999, 10 comps    index 999, 10 comps  True
      1000      index -1, 9 comps      index -1, 9 comps  True
      4000     index -1, 10 comps     index -1, 10 comps  True

identical, as they must be: the recursion is the loop written another way.
the iterative version uses no stack, so it has no depth limit.

The two agree exactly, index and comparison count. The iterative version uses no call stack, so there is no recursion limit to worry about, and it is the one to submit. The recursive one is shorter and reads more like the definition, which is why books print it.

Procedure

  1. Save searches.py with linear_search, binary_search (iterative, able to trace),

binary_search_recursive, binary_search_first and insertion_point, each counting its comparisons.

  1. Use while low <= high and middle = low + (high - low) // 2, and be able to say why the

second form is written that way.

  1. Move the bounds with middle + 1 and middle - 1, never middle.
  2. Search a sorted array of ten items for a value at the start, in the middle, at the end and absent,

printing the trace of low, high and middle at every step.

  1. Tabulate the worst case comparisons for both searches for n from 10 to a million, beside log2(n).
  2. Count the sort as well, and find how many searches it takes before sorting and binary searching

beats plain linear search.

  1. Write the low < high version and record which targets it misses.
  2. Search an array with duplicates and show that plain binary search finds an occurrence, not the first.
  3. Write the recursive version and confirm it gives the same index and count as the iterative one.
  4. Show bisect_left, bisect_right and insort.

Result

Both searches found every present value and reported -1 for absent ones. The traces showed low and high closing in, each comparison halving the range, and ten items needing at most four comparisons. For the first element, linear search took 1 comparison and binary search 3. Across n from 10 to a million, the linear worst case equalled n exactly and the binary worst case tracked log2(n) + 1: at a million items, 1,000,000 comparisons against 20. Counting the sort changed the answer for small numbers of searches: with 500 unsorted items the insertion sort cost 57,745 comparisons, so linear search won outright up to 200 searches and sorting then binary searching only won at 1000. The low < high version failed to find several values that were present. On an array with duplicates, plain binary search returned an occurrence and not the first, while binary_search_first returned the first. The recursive and iterative versions agreed on both the index and the comparison count for every target.

munotes.in239

Practical 9: Linear and Binary Search Compared

Where marks are lost

  • while low < high, which misses a range of one item and finds only some of the values.
  • low = middle or high = middle, which can stop the range shrinking so the loop never ends.
  • Binary search on unsorted data. It returns a wrong answer, not a slow one.
  • Not counting the sort when comparing. Binary search has a fixed price to pay first.
  • No comparison counts at all. MU's row says compare, and the comparison must be numbers.
  • Claiming binary search is always better. For the first element, or for a few searches of unsorted

data, it is not.

  • Saying binary search works on a linked list. It needs random access to the middle.
  • Assuming plain binary search finds the first duplicate. It finds one of them.
  • Returning 0 for absent. Zero is a valid index; return -1.

For the journal

The aim in MU's words, including "on a sorted array". The two pictures of the same search, linear and binary, with the comparison counts. Both programs, with while low <= high and middle = low + (high - low) // 2 pointed out and one sentence on why the second form is written that way. Then the traced binary searches for a value at the start, the middle, the end and absent, with low, high and middle printed at each step. Then the table of worst cases up to a million beside log2(n). Then the table that counts the sort, and the sentence naming how many searches it takes before binary search pays for itself, with the sort's own cost written beside it. Then the low < high bug and which targets it missed. The conclusion: linear search needs nothing of the data and costs up to n comparisons; binary search needs a sorted array and costs about log2(n), so the choice depends on whether the data is already sorted and how many times it will be searched.

Quick revision

  • Linear search: every item in turn. O(n). Needs nothing of the data. Works on a linked list.
  • Binary search: compare the middle and discard half. O(log n). Needs a SORTED array.
  • while low <= high. With < a range of one item is skipped.
  • low = middle + 1, high = middle - 1. Never middle, or a range of two never shrinks.
  • middle = low + (high - low) // 2: safe in every language. (low + high) // 2 can overflow in C or
munotes.in240

Practical 9: Linear and Binary Search Compared

Java, a bug that was in the Java library for nine years.

  • Worst cases measured: n for linear, about log2(n) + 1 for binary. A million items: 1,000,000 against 20.
  • Multiplying n by ten multiplies linear search by ten and adds about three to binary search.
  • Binary search needs random access, so an array and not a linked list.
  • Count the sort. For a few searches of unsorted data, plain linear search wins: measured here, an

O(n squared) sort of 500 items cost 57,745 comparisons and linear search stayed ahead to 200 searches. If the array is already sorted, binary search wins from the first search.

  • Plain binary search finds an occurrence, not the first. For the first, continue left after a hit.
  • bisect_left, bisect_right, and bisect_right - bisect_left counts occurrences in O(log n).

insort keeps a list sorted.

  • Recursive and iterative binary search give identical results; the iterative one uses no call stack.

Questions you should be able to answer

1. What does binary search require that linear search does not? The data must be sorted, and it must support random access to the middle, so an array rather than a linked list.

2. What does each one cost? Linear search is O(n), up to one comparison per item. Binary search is O(log n), about log2(n) + 1 comparisons.

3. Give the measured figures for a million items. Linear search made 1,000,000 comparisons in the worst case; binary search made 20.

4. Why while low <= high and not low < high? Because when the range narrows to a single item, low equals high, and low < high is then false, so that last item is never compared. The broken version in this chapter misses several values that are present.

5. Why low = middle + 1 rather than low = middle? Because leaving the middle in the range means a range of two items can stop shrinking, and the loop never ends.

6. Why is middle written as low + (high - low) // 2? Because (low + high) can overflow a fixed width integer in C, C++ or Java for a large array. It cannot in Python, and the safer form is worth writing anyway.

7. When does linear search beat binary search? When the target is at the front; when the data is unsorted and will be searched only a few times, because the sort has to be paid for first (measured here: 500 items, and linear search still ahead at 200 searches); when the structure is a linked list; and when the data is tiny.

munotes.in241

Practical 9: Linear and Binary Search Compared

8. What happens if you binary search unsorted data? It returns a wrong answer, not a slow one, because the comparison that discards half the range is only valid if the data is in order.

9. Does plain binary search find the first of several equal values? No, it finds one of them. To find the first, record the hit and keep searching to the left.

10. How do you count how many times a value occurs in a sorted array, in O(log n)? bisect_right(data, x) - bisect_left(data, x).

11. Which version would you submit, recursive or iterative? The iterative one: it gives identical answers and uses no call stack, so there is no recursion limit.

Contents This chapter on its own page

munotes.in242

Chapter Thirty-Four

Practical 10: the Combined Application

Syllabus topic Module 2, practical 10, "Combined Application: Design a simple program that uses multiple data structures ."

Aim

To design a simple program that uses several data structures together, choosing each one for a reason.

The application: a library issue desk

Small enough to write, large enough to need real structures. The desk has to do six things:

What the desk doesHow oftenThe operation it needs
look a book up by its accession numberconstantlyfind by exact key
list the catalogue in title orderat the end of the daythe keys in sorted order
hold reservations for a book that is outwhile a book is outfirst come, first served
undo the last few desk actionsrarely, when a mistake is mademost recent first
report the most borrowed titlesonce a weeksorted by a count
suggest what else a borrower might likeon requestwho else borrowed what

Read that table again, because it is the whole design. Six operations, and each one names a structure. Nothing is chosen because it is on the syllabus.

NeedStructureWhy that one
find by exact keyhash tableO(1) average. Order does not matter here
catalogue in title orderbinary search treethe in-order walk is sorted, free
reservationsqueueFIFO is the only fair rule for a waiting list
undostackLIFO: the most recent action is undone first
most borrowedsorted arraybuilt once a week and then read; a sort is cheaper than maintaining order
suggestionsgraphthe relation is between any two borrowers, which nothing linear can hold

That last one is MU's Course Objective 6, which names graphs although no practical row does.

The program

"""A library issue desk built from five data structures.

Each structure is here because of the operation named in its docstring, which is
what MU's Course Objective 8 asks for: choose, and justify the choice.
"""

from collections import deque


# --------------------------------------------------------------------------
# 1. the CATALOGUE: a hash table, because a book is looked up by its exact
#    accession number far more often than anything else happens.
# --------------------------------------------------------------------------
def accession_hash(key, base=31):
    """A polynomial hash, written out so bucket numbers are reproducible."""
    total = 0
    for character in str(key):
        total = total * base + ord(character)
    return total


class Catalogue:
    """Hash table with separate chaining. Lookup by accession number, O(1)."""

    def __init__(self, buckets=11):
        self.buckets = [[] for _ in range(buckets)]
        self.count = 0

    def _index(self, key):
        return accession_hash(key) % len(self.buckets)

    def add(self, accession, title, author):
        chain = self.buckets[self._index(accession)]
        for position, (existing, _) in enumerate(chain):
            if existing == accession:
                chain[position] = (accession, (title, author))
                return False
        chain.append((accession, (title, author)))
        self.count += 1
        return True

    def find(self, accession):
        """(title, author) or None. Searches one bucket's chain only."""
        for existing, value in self.buckets[self._index(accession)]:
            if existing == accession:
                return value
        return None

    def __len__(self):
        return self.count

    def all_items(self):
        for chain in self.buckets:
            for accession, value in chain:
                yield accession, value

    def chain_lengths(self):
        return [len(chain) for chain in self.buckets]


# --------------------------------------------------------------------------
# 2. the TITLE INDEX: a binary search tree, because the catalogue has to be
#    listed in title order and an in-order walk gives that for nothing.
# --------------------------------------------------------------------------
class TitleNode:
    def __init__(self, title, accession):
        self.title = title
        self.accessions = [accession]
        self.left = None
        self.right = None


class TitleIndex:
    """BST keyed on title. in_order() is the sorted catalogue, free."""

    def __init__(self):
        self.root = None
        self.count = 0

    def add(self, title, accession):
        if self.root is None:
            self.root = TitleNode(title, accession)
            self.count += 1
            return
        here = self.root
        while True:
            if title == here.title:
                here.accessions.append(accession)
                return
            if title < here.title:
                if here.left is None:
                    here.left = TitleNode(title, accession)
                    self.count += 1
                    return
                here = here.left
            else:
                if here.right is None:
                    here.right = TitleNode(title, accession)
                    self.count += 1
                    return
                here = here.right

    def find(self, title):
        """(accessions, comparisons). O(log n) when the tree is balanced."""
        here = self.root
        comparisons = 0
        while here is not None:
            comparisons += 1
            if title == here.title:
                return here.accessions, comparisons
            here = here.left if title < here.title else here.right
        return None, comparisons

    def in_order(self):
        """Every title in order, which is the whole point of using a tree."""
        out = []

        def walk(node):
            if node is not None:
                walk(node.left)
                out.append((node.title, node.accessions))
                walk(node.right)

        walk(self.root)
        return out

    def height(self):
        def deepest(node):
            return -1 if node is None else 1 + max(deepest(node.left),
                                                   deepest(node.right))

        return deepest(self.root)


# --------------------------------------------------------------------------
# 3. RESERVATIONS: a queue per book, because a waiting list is fair only if
#    it is first come, first served.
# --------------------------------------------------------------------------
class Reservations:
    """A FIFO queue per accession number. deque, so both ends are O(1)."""

    def __init__(self):
        self.queues = {}

    def reserve(self, accession, borrower):
        queue = self.queues.setdefault(accession, deque())
        if borrower in queue:
            return -1
        queue.append(borrower)
        return len(queue)

    def next_in_line(self, accession):
        queue = self.queues.get(accession)
        if not queue:
            return None
        return queue.popleft()

    def waiting(self, accession):
        return list(self.queues.get(accession, ()))


# --------------------------------------------------------------------------
# 4. UNDO: a stack, because the most recent action is the one to undo.
# --------------------------------------------------------------------------
class UndoStack:
    """LIFO, over a fixed array, so it can overflow like a real one."""

    def __init__(self, capacity=20):
        self.slots = [None] * capacity
        self.top = -1

    def is_empty(self):
        return self.top == -1

    def push(self, action):
        if self.top == len(self.slots) - 1:
            # the oldest action is dropped rather than refusing the newest
            self.slots.pop(0)
            self.slots.append(None)
            self.top -= 1
        self.top += 1
        self.slots[self.top] = action

    def pop(self):
        if self.is_empty():
            raise IndexError("nothing left to undo")
        action = self.slots[self.top]
        self.slots[self.top] = None
        self.top -= 1
        return action

    def peek(self):
        if self.is_empty():
            return None
        return self.slots[self.top]

    def __len__(self):
        return self.top + 1


# --------------------------------------------------------------------------
# 5. BORROWER GRAPH: two borrowers are joined when they have borrowed the same
#    book. Nothing linear can hold a relation between any two things.
# --------------------------------------------------------------------------
class BorrowerGraph:
    """Undirected graph, adjacency list. Used for 'who else read this'."""

    def __init__(self):
        self.neighbours = {}
        self.borrowed = {}

    def record(self, borrower, accession):
        self.neighbours.setdefault(borrower, set())
        others = self.borrowed.setdefault(accession, set())
        for other in others:
            if other != borrower:
                self.neighbours[borrower].add(other)
                self.neighbours.setdefault(other, set()).add(borrower)
        others.add(borrower)

    def also_borrowed_by(self, borrower):
        return sorted(self.neighbours.get(borrower, ()))

    def suggestions(self, borrower, catalogue):
        """Books read by people who read what this borrower read. BFS, depth 1."""
        mine = {a for a, who in self.borrowed.items() if borrower in who}
        out = {}
        for neighbour in self.neighbours.get(borrower, ()):
            for accession, who in self.borrowed.items():
                if neighbour in who and accession not in mine:
                    found = catalogue.find(accession)
                    if found:
                        out.setdefault(found[0], set()).add(neighbour)
        return {title: sorted(who) for title, who in sorted(out.items())}

    def connected_group(self, start):
        """Everyone reachable from a borrower. BFS with a seen set."""
        if start not in self.neighbours:
            return []
        seen = {start}
        order = []
        waiting = deque([start])
        while waiting:
            here = waiting.popleft()
            order.append(here)
            for neighbour in sorted(self.neighbours[here]):
                if neighbour not in seen:
                    seen.add(neighbour)
                    waiting.append(neighbour)
        return order


# --------------------------------------------------------------------------
# the DESK, which ties them together
# --------------------------------------------------------------------------
class IssueDesk:
    def __init__(self):
        self.catalogue = Catalogue()
        self.titles = TitleIndex()
        self.reservations = Reservations()
        self.undo = UndoStack()
        self.graph = BorrowerGraph()
        self.on_loan = {}
        self.borrow_count = {}

    def add_book(self, accession, title, author):
        if self.catalogue.add(accession, title, author):
            self.titles.add(title, accession)
            self.borrow_count.setdefault(accession, 0)
            return f"added {accession} {title!r}"
        return f"{accession} was already in the catalogue, details updated"

    def issue(self, accession, borrower):
        book = self.catalogue.find(accession)
        if book is None:
            return f"no such accession number {accession}"
        if accession in self.on_loan:
            place = self.reservations.reserve(accession, borrower)
            if place == -1:
                return f"{borrower} is already waiting for {book[0]!r}"
            return (f"{book[0]!r} is out with {self.on_loan[accession]}; "
                    f"{borrower} is number {place} in the queue")
        self.on_loan[accession] = borrower
        self.borrow_count[accession] += 1
        self.graph.record(borrower, accession)
        self.undo.push(("issue", accession, borrower))
        return f"issued {book[0]!r} to {borrower}"

    def give_back(self, accession):
        if accession not in self.on_loan:
            return f"{accession} is not on loan"
        was = self.on_loan.pop(accession)
        self.undo.push(("return", accession, was))
        nxt = self.reservations.next_in_line(accession)
        title = self.catalogue.find(accession)[0]
        if nxt is None:
            return f"{title!r} returned by {was}, back on the shelf"
        self.on_loan[accession] = nxt
        self.borrow_count[accession] += 1
        self.graph.record(nxt, accession)
        return (f"{title!r} returned by {was} and issued straight to {nxt}, "
                f"who was first in the queue")

    def undo_last(self):
        if self.undo.is_empty():
            return "nothing left to undo"
        what, accession, who = self.undo.pop()
        title = self.catalogue.find(accession)[0]
        if what == "issue":
            self.on_loan.pop(accession, None)
            self.borrow_count[accession] -= 1
            return f"undid the issue of {title!r} to {who}"
        self.on_loan[accession] = who
        return f"undid the return of {title!r}, it is with {who} again"

    def sorted_catalogue(self):
        return self.titles.in_order()

    def most_borrowed(self, how_many=3):
        """A SORTED ARRAY, built once and then read. Ties broken by title."""
        rows = []
        for accession, count in self.borrow_count.items():
            title = self.catalogue.find(accession)[0]
            rows.append((count, title, accession))
        rows.sort(key=lambda row: (-row[0], row[1]))
        return rows[:how_many]
munotes.in243

Practical 10: the Combined Application

The desk, run

from library import IssueDesk

desk = IssueDesk()

print("adding books:")
books = [("A104", "Data Structures", "Karumanchi"),
         ("A101", "Let Us Python", "Kanetkar"),
         ("A109", "Operating System Concepts", "Silberschatz"),
         ("A102", "Learning Python", "Lutz"),
         ("A115", "Introduction to Algorithms", "Cormen"),
         ("A107", "Python: The Complete Reference", "Brown")]
for accession, title, author in books:
    print("  ", desk.add_book(accession, title, author))

print()
print(f"the catalogue holds {len(desk.catalogue)} books")
print("the hash table's chain lengths:", desk.catalogue.chain_lengths())
print("the title tree's height:", desk.titles.height())
munotes.in244

Practical 10: the Combined Application

adding books:
   added A104 'Data Structures'
   added A101 'Let Us Python'
   added A109 'Operating System Concepts'
   added A102 'Learning Python'
   added A115 'Introduction to Algorithms'
   added A107 'Python: The Complete Reference'

the catalogue holds 6 books
the hash table's chain lengths: [1, 0, 0, 1, 1, 1, 1, 0, 0, 1, 0]
the title tree's height: 3
munotes.in245

Practical 10: the Combined Application

from library import IssueDesk

desk = IssueDesk()
for accession, title, author in [("A104", "Data Structures", "Karumanchi"),
                                 ("A101", "Let Us Python", "Kanetkar"),
                                 ("A109", "Operating System Concepts", "Silberschatz"),
                                 ("A102", "Learning Python", "Lutz"),
                                 ("A115", "Introduction to Algorithms", "Cormen"),
                                 ("A107", "Python: The Complete Reference", "Brown")]:
    desk.add_book(accession, title, author)

print("1. LOOKUP by accession number, which is the hash table:")
for accession in ["A104", "A115", "A999"]:
    found = desk.catalogue.find(accession)
    print(f"   {accession} -> {found if found else 'not in the catalogue'}")

print()
print("2. the CATALOGUE IN TITLE ORDER, which is the tree's in-order walk:")
for title, accessions in desk.sorted_catalogue():
    print(f"   {title:<34} {accessions}")

print()
print("3. ISSUING, and the RESERVATION QUEUE when a book is already out:")
for accession, borrower in [("A104", "Aarti"), ("A104", "Bhavesh"),
                            ("A104", "Chetan"), ("A101", "Aarti"),
                            ("A101", "Divya")]:
    print("  ", desk.issue(accession, borrower))
print("   waiting for A104:", desk.reservations.waiting("A104"))

print()
print("4. RETURNING, and the queue serving itself:")
print("  ", desk.give_back("A104"))
print("   waiting for A104 now:", desk.reservations.waiting("A104"))
print("  ", desk.give_back("A104"))
print("   waiting for A104 now:", desk.reservations.waiting("A104"))

print()
print("5. UNDO, which is the stack:")
print("   the stack holds", len(desk.undo), "action(s); the top is",
      desk.undo.peek())
print("  ", desk.undo_last())
print("  ", desk.undo_last())
print("   the stack holds", len(desk.undo), "action(s) now")

print()
print("6. MOST BORROWED, which is the sorted array:")
for count, title, accession in desk.most_borrowed(3):
    print(f"   {count} loan(s)  {title} ({accession})")
1. LOOKUP by accession number, which is the hash table:
   A104 -> ('Data Structures', 'Karumanchi')
   A115 -> ('Introduction to Algorithms', 'Cormen')
   A999 -> not in the catalogue

2. the CATALOGUE IN TITLE ORDER, which is the tree's in-order walk:
   Data Structures                    ['A104']
   Introduction to Algorithms         ['A115']
   Learning Python                    ['A102']
   Let Us Python                      ['A101']
   Operating System Concepts          ['A109']
   Python: The Complete Reference     ['A107']

3. ISSUING, and the RESERVATION QUEUE when a book is already out:
   issued 'Data Structures' to Aarti
   'Data Structures' is out with Aarti; Bhavesh is number 1 in the queue
   'Data Structures' is out with Aarti; Chetan is number 2 in the queue
   issued 'Let Us Python' to Aarti
   'Let Us Python' is out with Aarti; Divya is number 1 in the queue
   waiting for A104: ['Bhavesh', 'Chetan']

4. RETURNING, and the queue serving itself:
   'Data Structures' returned by Aarti and issued straight to Bhavesh, who was first in the queue
   waiting for A104 now: ['Chetan']
   'Data Structures' returned by Bhavesh and issued straight to Chetan, who was first in the queue
   waiting for A104 now: []

5. UNDO, which is the stack:
   the stack holds 4 action(s); the top is ('return', 'A104', 'Bhavesh')
   undid the return of 'Data Structures', it is with Bhavesh again
   undid the return of 'Data Structures', it is with Aarti again
   the stack holds 2 action(s) now

6. MOST BORROWED, which is the sorted array:
   3 loan(s)  Data Structures (A104)
   1 loan(s)  Let Us Python (A101)
   0 loan(s)  Introduction to Algorithms (A115)
munotes.in246

Practical 10: the Combined Application

Read the reservation lines. Aarti took A104, and Bhavesh and Chetan joined the queue behind her. When Aarti gave it back, the book went straight to Bhavesh, who was first in line, and the queue then held only Chetan. That is FIFO doing the only fair thing, and it is why the reservations are a queue and not anything else.

Read the undo lines too: the last action is undone first, which is LIFO, and that is why undo is a stack.

The graph, which is MU's Course Objective 6

from library import IssueDesk

desk = IssueDesk()
for accession, title, author in [("A104", "Data Structures", "Karumanchi"),
                                 ("A101", "Let Us Python", "Kanetkar"),
                                 ("A109", "Operating System Concepts", "Silberschatz"),
                                 ("A102", "Learning Python", "Lutz"),
                                 ("A115", "Introduction to Algorithms", "Cormen")]:
    desk.add_book(accession, title, author)

# a small borrowing history: several people, several books
history = [("A104", "Aarti"), ("A101", "Aarti"), ("A104", "Bhavesh"),
           ("A109", "Bhavesh"), ("A101", "Chetan"), ("A115", "Chetan"),
           ("A102", "Divya"), ("A115", "Divya"), ("A109", "Eshan")]
for accession, borrower in history:
    desk.graph.record(borrower, accession)

print("who has read a book in common with whom:")
for borrower in ["Aarti", "Bhavesh", "Chetan", "Divya", "Eshan"]:
    print(f"   {borrower:<9} {desk.graph.also_borrowed_by(borrower)}")

print()
print("suggestions for Aarti, from what her neighbours read that she has not:")
for title, who in desk.graph.suggestions("Aarti", desk.catalogue).items():
    print(f"   {title:<34} read by {', '.join(who)}")

print()
print("everyone reachable from Aarti, breadth first:")
print("  ", desk.graph.connected_group("Aarti"))
print("   note that Eshan is reached, although he shares no book with Aarti:")
print("   Aarti -> Bhavesh (A104) -> Eshan (A109). That is a PATH, and only a")
print("   graph can hold it.")
who has read a book in common with whom:
   Aarti     ['Bhavesh', 'Chetan']
   Bhavesh   ['Aarti', 'Eshan']
   Chetan    ['Aarti', 'Divya']
   Divya     ['Chetan']
   Eshan     ['Bhavesh']

suggestions for Aarti, from what her neighbours read that she has not:
   Introduction to Algorithms         read by Chetan
   Operating System Concepts          read by Bhavesh

everyone reachable from Aarti, breadth first:
   ['Aarti', 'Bhavesh', 'Chetan', 'Eshan', 'Divya']
   note that Eshan is reached, although he shares no book with Aarti:
   Aarti -> Bhavesh (A104) -> Eshan (A109). That is a PATH, and only a
   graph can hold it.
munotes.in247

Practical 10: the Combined Application

That last line is the justification. A hash table can say what Aarti borrowed. A tree can list the catalogue. Neither can answer "who is connected to Aarti through somebody else", because the relation is between any two borrowers and has no order and no single key. Only a graph holds it, and the answer is a breadth first search with a seen set, exactly as [Python for Data Structures, and the Cost of an Operation] set out.

What each structure costs, in this program

from library import IssueDesk

desk = IssueDesk()
for n in range(60):
    desk.add_book(f"A{n:04d}", f"Title {n:02d}", "Author")

print(f"60 books in the catalogue")
print(f"  hash table chain lengths : {desk.catalogue.chain_lengths()}")
print(f"  longest chain            : {max(desk.catalogue.chain_lengths())}")
print(f"  load factor              : {len(desk.catalogue) / 11:.2f}")
print(f"  title tree height        : {desk.titles.height()}")
print()
print("finding a title in the tree, with the comparisons:")
for title in ["Title 00", "Title 30", "Title 59", "Title 99"]:
    accessions, comparisons = desk.titles.find(title)
    found = accessions if accessions else "not found"
    print(f"  {title:<10} {str(found):<12} {comparisons} comparison(s)")
60 books in the catalogue
  hash table chain lengths : [6, 5, 6, 5, 6, 5, 6, 5, 5, 6, 5]
  longest chain            : 6
  load factor              : 5.45
  title tree height        : 59

finding a title in the tree, with the comparisons:
  Title 00   ['A0000']    1 comparison(s)
  Title 30   ['A0030']    31 comparison(s)
  Title 59   ['A0059']    60 comparison(s)
  Title 99   not found    60 comparison(s)

Note the tree's height. The titles were inserted in ascending order, which is exactly the degenerate case of [Practical 5: the Binary Search Tree, Create, Insert and Search], so the tree is a linked list and the searches cost a walk rather than a halving.

That is a real defect in this program, honestly reported, and the cure is the one that chapter gives:

from library import IssueDesk
import random

random.seed(7)

titles = [f"Title {n:02d}" for n in range(60)]

ordered = IssueDesk()
for n, title in enumerate(titles):
    ordered.add_book(f"A{n:04d}", title, "Author")

shuffled_titles = titles[:]
random.shuffle(shuffled_titles)
shuffled = IssueDesk()
for n, title in enumerate(shuffled_titles):
    shuffled.add_book(f"B{n:04d}", title, "Author")

print(f"  titles added in ascending order : height {ordered.titles.height()}")
print(f"  titles added in shuffled order  : height {shuffled.titles.height()}")
print()
for desk, label in [(ordered, "ascending"), (shuffled, "shuffled ")]:
    _, comparisons = desk.titles.find("Title 59")
    print(f"  finding the last title, {label}: {comparisons} comparison(s)")
print()
print("both list the catalogue in the same order, which is the point:")
print("  ", [t for t, _ in ordered.sorted_catalogue()][:5], "...")
print("  ", [t for t, _ in shuffled.sorted_catalogue()][:5], "...")
munotes.in248

Practical 10: the Combined Application

  titles added in ascending order : height 59
  titles added in shuffled order  : height 10

  finding the last title, ascending: 60 comparison(s)
  finding the last title, shuffled : 5 comparison(s)

both list the catalogue in the same order, which is the point:
   ['Title 00', 'Title 01', 'Title 02', 'Title 03', 'Title 04'] ...
   ['Title 00', 'Title 01', 'Title 02', 'Title 03', 'Title 04'] ...

Both trees give the same sorted catalogue and one of them takes far fewer comparisons to search. In a real desk the books arrive in accession order, not title order, so the tree comes out reasonably balanced by itself; adding a batch from a sorted file is the case to watch for.

A second design, so you can pick either

MU's row is open, so here is a different application with different justifications. Write whichever you prefer; what is marked is the reasoning.

NeedStructureWhy
an expression typed by the userstackinfix to postfix and then evaluation, practical 3
the variables it useshash tablelooked up by name, constantly, O(1)
the history of expressions enteredstackthe most recent is recalled first
the results, in numeric orderBSTthe in-order walk is the sorted list of answers
the pending calculations in a batchqueuethey are worked in the order they were submitted
from collections import deque

PRECEDENCE = {"+": 1, "-": 1, "*": 2, "/": 2}


def to_postfix(tokens, variables):
    """Stack of operators; names are looked up in the hash table on the way out."""
    output, stack = [], []
    for token in tokens:
        if token in PRECEDENCE:
            while stack and stack[-1] != "(" and \
                    PRECEDENCE[stack[-1]] >= PRECEDENCE[token]:
                output.append(stack.pop())
            stack.append(token)
        elif token == "(":
            stack.append(token)
        elif token == ")":
            while stack and stack[-1] != "(":
                output.append(stack.pop())
            stack.pop()
        else:
            output.append(variables.get(token, token))
    while stack:
        output.append(stack.pop())
    return output


def evaluate(postfix):
    stack = []
    for token in postfix:
        if token in PRECEDENCE:
            right, left = stack.pop(), stack.pop()
            stack.append({"+": left + right, "-": left - right,
                          "*": left * right, "/": left / right}[token])
        else:
            stack.append(float(token))
    return stack.pop()


variables = {"marks": 78, "total": 100, "credits": 4}     # a hash table
pending = deque()                                          # a queue
history = []                                               # a stack
answers = []                                               # kept sorted

for text in ["marks / total * 100", "marks * credits", "( marks + 10 ) / 2",
             "total - marks"]:
    pending.append(text)

print("working the queue in the order the expressions were submitted:")
while pending:
    text = pending.popleft()
    tokens = text.split()
    postfix = to_postfix(tokens, variables)
    value = evaluate(postfix)
    history.append((text, value))
    answers.append(value)
    print(f"  {text:<22} postfix {' '.join(str(t) for t in postfix):<20} "
          f"= {value}")

print()
print("the history, most recent first, which is the stack:")
while history:
    text, value = history.pop()
    print(f"  {text:<22} = {value}")

print()
print("the answers in numeric order, which a BST would give as its in-order walk:")
print("  ", sorted(answers))
munotes.in249

Practical 10: the Combined Application

working the queue in the order the expressions were submitted:
  marks / total * 100    postfix 78 100 / 100 *       = 78.0
  marks * credits        postfix 78 4 *               = 312.0
  ( marks + 10 ) / 2     postfix 78 10 + 2 /          = 44.0
  total - marks          postfix 100 78 -             = 22.0

the history, most recent first, which is the stack:
  total - marks          = 22.0
  ( marks + 10 ) / 2     = 44.0
  marks * credits        = 312.0
  marks / total * 100    = 78.0

the answers in numeric order, which a BST would give as its in-order walk:
   [22.0, 44.0, 78.0, 312.0]

What to write when the examiner says "design a program"

The answer has three parts and the third is where the marks are.

  1. The application, in two sentences. Small and concrete.
  2. What it has to do, as a list of operations, with how often each happens.
  3. A table: need, structure, why. One row per structure, and the "why" names the operation that

structure makes cheap, not the structure's features.

A good "why" reads: a hash table, because the desk looks a book up by accession number on every transaction and never needs the catalogue in accession order.

A bad "why" reads: a hash table, because hash tables are fast.

Procedure

  1. Pick an application small enough to write and big enough to need three or more structures.
  2. Write the operations list first, with how often each one happens. That list chooses the

structures.

  1. Write the need, structure, why table before writing any code.
  2. Implement each structure as its own class, with a docstring naming the operation it is there for.
  3. Tie them together in one class whose methods are the desk's actions.
  4. Run a realistic session: add records, look one up, list them in order, queue a reservation, return

the item and watch the queue serve itself, undo two actions, and report the top three.

  1. Print the internal state as you go: the chain lengths, the tree height, the queue contents, the stack

depth.

  1. Add the graph query that nothing else can answer, and show a path of length two.
  2. Report each structure's cost in this program, including any defect, such as the tree's height when the

keys arrive sorted.

  1. Write the justification table into the journal beside the output.

Result

Six books were added to a hash table catalogue and a binary search tree title index. Lookup by accession number searched one bucket's chain. The in-order walk of the tree listed the catalogue in title order. Issuing a book that was already out put the borrower in a queue, and returning it issued the book straight to the first in line, leaving the rest waiting, which is FIFO. Undo removed the two most recent actions in reverse order, which is LIFO. The most borrowed report was a sorted array built from the counts. The graph answered a question none of the others can: Eshan is reachable from Aarti through Bhavesh, although Aarti and Eshan share no book. With 60 titles added in ascending order the tree's height was 59, a degenerate tree, and the same titles shuffled gave a much smaller height and far fewer comparisons while producing the same sorted catalogue.

munotes.in250

Practical 10: the Combined Application

Where marks are lost

  • No justification. The structures are the easy half; the reasons are the exercise.
  • A "why" that describes the structure instead of naming the operation it makes cheap.
  • Using every structure on the syllabus whether the application needs it or not.
  • One structure doing everything, usually a Python dictionary, with the others mentioned and unused.
  • No output. The session run is the evidence that the structures work together.
  • Not printing the internal state. The chain lengths, the tree height and the queue contents are what

show the structures are real.

  • Not noticing the sorted insertion defect in the tree, or not reporting it.
  • A queue served from the wrong end, which makes the reservation list unfair.

For the journal

The aim in MU's words. The operations table first, with how often each happens, because that table is what chooses the structures. Then the need, structure, why table, with each "why" naming an operation. Then the code, each class with its docstring. Then the session run in full, and beside it the internal state: chain lengths, tree height, the queue before and after a return, the stack depth. Then the graph query with the two step path, and one sentence saying no other structure here can answer it. Then the tree height with the titles inserted in order and shuffled, reported as a defect and a cure. The conclusion: each structure was chosen by naming the operation it makes cheap, and the same data in the wrong structure would have made one of the desk's six jobs slow.

Quick revision

  • The operations list chooses the structures, not the other way round.
  • hash table for lookup by exact key, O(1) average, when order does not matter.
  • BST when you need fast lookup and the sorted order, which the in-order walk gives free.
  • queue for a waiting list, because FIFO is the only fair rule.
  • stack for undo, because the most recent action goes first.
  • sorted array for a report built once and then read; sorting once beats keeping order all the time.
  • graph when the relation is between any two things, which nothing linear can hold.
  • A good justification names the operation: "a hash table, because we look up by accession number on
munotes.in251

Practical 10: the Combined Application

every transaction". A bad one names the feature: "because hash tables are fast".

  • Print the internal state: chain lengths, tree height, queue contents, stack depth.
  • Watch for the tree going degenerate when keys arrive sorted, and say so.

Questions you should be able to answer

1. How do you decide which structures a program needs? Write the list of operations it performs and how often each happens. Each frequent operation names the structure that makes it cheap.

2. Why a hash table for the catalogue? Because the desk looks a book up by its exact accession number on every transaction, which a hash table does in O(1) on average, and it never needs the books in accession order.

3. Why a tree as well, when the hash table already holds everything? Because the catalogue must be listed in title order, and a hash table has no order at all. A binary search tree's in-order walk gives the sorted list for nothing.

4. Why a queue for reservations and not a stack? Because a waiting list is only fair if the first person to ask is the first served, which is FIFO. A stack would serve the most recent request first.

5. Why a stack for undo? Because the most recent action must be undone first, which is LIFO.

6. Why a sorted array for the most borrowed report and not a tree? Because the report is built once a week and then read. Sorting once is cheaper than keeping a structure in order through every loan.

7. What can the graph answer that nothing else here can? Whether two borrowers are connected through other people. In the run, Eshan is reachable from Aarti through Bhavesh although they share no book, and that is a path, which only a graph holds.

8. What was wrong with the title tree when 60 books were added in title order? It degenerated into a linked list with a height of 59, so searching it walked instead of halving. Shuffling the insertion order fixed it and gave the same sorted catalogue.

9. What makes a good justification? Naming the operation the structure makes cheap, and how often that operation happens. Naming a property of the structure instead is what loses the mark.

10. What would you print to show the structures are really there? The hash table's chain lengths and load factor, the tree's height, the queue's contents before and after a return, and the stack's depth.

Contents This chapter on its own page

munotes.in252

Chapter Thirty-Five

If Your College Runs Module 2 in C

Syllabus topic Module 2's own subject, in the language three of MU's five text books for this paper use, and CO 9, "dynamic memory allocation and efficient data management techniques"

Aim

To implement all ten of MU's Module 2 exercises in C, and to show dynamic memory allocation properly.

What changes, and what does not

The structures and the algorithms do not change at all. An array still shifts, a linked list still needs the node in front of the one being deleted, a stack is still LIFO, and a binary search tree still halves. Everything learned in the Python chapters transfers.

Four things about the language change:

In PythonIn C
a recorda class with __init__a struct, and a typedef to name it
a new nodeNode(data), freed automaticallymalloc, and you must free it
"no node"NoneNULL
a list that growsany lengtha fixed array, or malloc and your own links
a stringstr, any lengthchar[], and you count the terminating '\0'
an errorraisea return code, and you must check it

The last row matters more than it looks. C has no exceptions, so every function returns a status and every caller checks it. A program that ignores the return of malloc is one allocation failure away from a crash.

Dynamic memory allocation, which is Course Objective 9

#include <stdio.h>
#include <stdlib.h>

typedef struct Node {
    int data;
    struct Node *next;
} Node;

int main(void) {
    printf("sizeof(int)   = %zu bytes\n", sizeof(int));
    printf("sizeof(Node)  = %zu bytes\n", sizeof(Node));
    printf("sizeof(Node*) = %zu bytes\n", sizeof(Node *));

    /* one node on the heap */
    Node *node = malloc(sizeof(Node));
    if (node == NULL) {                      /* ALWAYS check */
        fprintf(stderr, "out of memory\n");
        return 1;
    }
    node->data = 42;
    node->next = NULL;
    printf("the node holds %d and its next is %s\n",
           node->data, node->next == NULL ? "NULL" : "something");

    free(node);                              /* and ALWAYS free */
    node = NULL;                             /* so a later use is a clear crash */
    printf("freed, and the pointer set to NULL\n");

    /* an array whose size is decided at run time */
    int wanted = 5;
    int *values = malloc((size_t)wanted * sizeof(int));
    if (values == NULL) {
        return 1;
    }
    for (int i = 0; i < wanted; i++) {
        values[i] = (i + 1) * 10;
    }
    printf("the dynamic array:");
    for (int i = 0; i < wanted; i++) {
        printf(" %d", values[i]);
    }
    printf("\n");

    /* grow it */
    int *bigger = realloc(values, 8 * sizeof(int));
    if (bigger == NULL) {
        free(values);                        /* realloc failing does NOT free the old block */
        return 1;
    }
    values = bigger;
    for (int i = wanted; i < 8; i++) {
        values[i] = (i + 1) * 10;
    }
    printf("after realloc to 8:");
    for (int i = 0; i < 8; i++) {
        printf(" %d", values[i]);
    }
    printf("\n");

    free(values);
    printf("freed\n");
    return 0;
}
munotes.in253

If Your College Runs Module 2 in C

sizeof(int)   = 4 bytes
sizeof(Node)  = 16 bytes
sizeof(Node*) = 8 bytes
the node holds 42 and its next is NULL
freed, and the pointer set to NULL
the dynamic array: 10 20 30 40 50
after realloc to 8: 10 20 30 40 50 60 70 80
freed

Five rules in that listing, and every one of them is a mark:

malloc(sizeof(Node)), never malloc(20). The sizeof is computed by the compiler and stays right when the struct changes.

Check the return. malloc returns NULL when it cannot allocate, and using a NULL pointer is a crash with no message.

free every malloc. One free per malloc, no more and no fewer. Freeing twice is as bad as never freeing.

Set the pointer to NULL after freeing. Then a later use crashes immediately instead of quietly reading freed memory, which is far harder to find.

realloc failing does not free the old block. Assign to a new pointer, check it, and only then replace the old one. Writing values = realloc(values, ...) loses the old pointer when it fails, which is a leak.

Practical 1: array operations

#include <stdio.h>

#define CAPACITY 10

static void show(const char *label, const int a[], int size) {
    printf("%-22s [", label);
    for (int i = 0; i < size; i++) {
        printf("%s%d", i ? ", " : "", a[i]);
    }
    printf("]  size %d/%d\n", size, CAPACITY);
}

static int insert_at(int a[], int *size, int index, int value) {
    if (*size >= CAPACITY) return -1;            /* full */
    if (index < 0 || index > *size) return -2;   /* out of range */
    for (int i = *size; i > index; i--) {        /* BACKWARDS */
        a[i] = a[i - 1];
    }
    a[index] = value;
    (*size)++;
    return *size - index - 1;                    /* how many moved */
}

static int delete_at(int a[], int *size, int index, int *removed) {
    if (*size == 0) return -1;
    if (index < 0 || index >= *size) return -2;
    *removed = a[index];
    for (int i = index; i < *size - 1; i++) {    /* FORWARDS */
        a[i] = a[i + 1];
    }
    (*size)--;
    return *size - index;
}

static int linear_search(const int a[], int size, int target, int *comparisons) {
    *comparisons = 0;
    for (int i = 0; i < size; i++) {
        (*comparisons)++;
        if (a[i] == target) return i;
    }
    return -1;
}

int main(void) {
    int a[CAPACITY];
    int size = 0;

    for (int v = 10; v <= 50; v += 10) {
        a[size++] = v;
    }
    show("five appended", a, size);

    int moved = insert_at(a, &size, 2, 99);
    printf("insert 99 at 2         moved %d value(s)\n", moved);
    show("after", a, size);

    moved = insert_at(a, &size, 0, 5);
    printf("insert 5 at 0          moved %d value(s)\n", moved);
    show("after", a, size);

    int removed = 0;
    moved = delete_at(a, &size, 3, &removed);
    printf("delete index 3         removed %d, moved %d\n", removed, moved);
    show("after", a, size);

    int comparisons = 0;
    for (int i = 0; i < 2; i++) {
        int target = i ? 77 : 30;
        int found = linear_search(a, size, target, &comparisons);
        printf("search %-3d             %s after %d comparison(s)\n", target,
               found >= 0 ? "found" : "NOT FOUND", comparisons);
    }
    return 0;
}
munotes.in254

If Your College Runs Module 2 in C

five appended          [10, 20, 30, 40, 50]  size 5/10
insert 99 at 2         moved 3 value(s)
after                  [10, 20, 99, 30, 40, 50]  size 6/10
insert 5 at 0          moved 6 value(s)
after                  [5, 10, 20, 99, 30, 40, 50]  size 7/10
delete index 3         removed 99, moved 3
after                  [5, 10, 20, 30, 40, 50]  size 6/10
search 30              found after 4 comparison(s)
search 77              NOT FOUND after 6 comparison(s)

int size and int removed are how a C function returns more than one thing: it takes a pointer and writes through it. That is what Python's tuple return does in one line, and it is the commonest place a student new to C forgets the & at the call site or the * inside the function.

Practical 2: the singly linked list, with malloc and free

#include <stdio.h>
#include <stdlib.h>

typedef struct Node {
    int data;
    struct Node *next;
} Node;

typedef struct {
    Node *head;
    Node *tail;
    int count;
} List;

static void init(List *list) {
    list->head = NULL;
    list->tail = NULL;
    list->count = 0;
}

static void show(const char *label, const List *list) {
    printf("%-26s ", label);
    for (Node *here = list->head; here != NULL; here = here->next) {
        printf("%d -> ", here->data);
    }
    printf("NULL   (%d node%s)\n", list->count, list->count == 1 ? "" : "s");
}

static int insert_at_beginning(List *list, int value) {
    Node *node = malloc(sizeof(Node));
    if (node == NULL) return -1;
    node->data = value;
    node->next = list->head;
    list->head = node;
    if (list->tail == NULL) list->tail = node;    /* the empty list case */
    list->count++;
    return 0;
}

static int insert_at_end(List *list, int value) {
    Node *node = malloc(sizeof(Node));
    if (node == NULL) return -1;
    node->data = value;
    node->next = NULL;
    if (list->head == NULL) {
        list->head = node;
    } else {
        list->tail->next = node;
    }
    list->tail = node;
    list->count++;
    return 0;
}

static int insert_at_position(List *list, int position, int value) {
    if (position < 0 || position > list->count) return -2;
    if (position == 0) return insert_at_beginning(list, value);
    if (position == list->count) return insert_at_end(list, value);
    Node *before = list->head;
    for (int i = 0; i < position - 1; i++) {
        before = before->next;
    }
    Node *node = malloc(sizeof(Node));
    if (node == NULL) return -1;
    node->data = value;
    node->next = before->next;
    before->next = node;
    list->count++;
    return 0;
}

static int delete_at_position(List *list, int position, int *removed) {
    if (list->head == NULL) return -1;
    if (position < 0 || position >= list->count) return -2;
    Node *gone = NULL;
    if (position == 0) {
        gone = list->head;
        list->head = gone->next;
        if (list->head == NULL) list->tail = NULL;
    } else {
        Node *before = list->head;
        for (int i = 0; i < position - 1; i++) {
            before = before->next;
        }
        gone = before->next;
        before->next = gone->next;
        if (gone == list->tail) list->tail = before;      /* the tail case */
    }
    *removed = gone->data;
    free(gone);                                           /* PYTHON DOES THIS FOR YOU */
    list->count--;
    return 0;
}

static void destroy(List *list) {
    Node *here = list->head;
    while (here != NULL) {
        Node *next = here->next;      /* save it BEFORE freeing */
        free(here);
        here = next;
    }
    init(list);
}

int main(void) {
    List list;
    init(&list);
    show("empty", &list);

    for (int v = 10; v <= 30; v += 10) {
        insert_at_end(&list, v);
    }
    show("three at the end", &list);

    insert_at_beginning(&list, 5);
    show("insert 5 at the beginning", &list);

    insert_at_position(&list, 2, 15);
    show("insert 15 at position 2", &list);

    int removed = 0;
    delete_at_position(&list, 2, &removed);
    printf("deleted %d from position 2\n", removed);
    show("after", &list);

    delete_at_position(&list, 0, &removed);
    printf("deleted %d from position 0\n", removed);
    show("after", &list);

    delete_at_position(&list, list.count - 1, &removed);
    printf("deleted %d, which was the tail\n", removed);
    show("after", &list);
    insert_at_end(&list, 99);
    show("append after that", &list);
    printf("the append worked, so the tail was updated\n");

    destroy(&list);
    show("after destroy", &list);
    return 0;
}
munotes.in255

If Your College Runs Module 2 in C

empty                      NULL   (0 nodes)
three at the end           10 -> 20 -> 30 -> NULL   (3 nodes)
insert 5 at the beginning  5 -> 10 -> 20 -> 30 -> NULL   (4 nodes)
insert 15 at position 2    5 -> 10 -> 15 -> 20 -> 30 -> NULL   (5 nodes)
deleted 15 from position 2
after                      5 -> 10 -> 20 -> 30 -> NULL   (4 nodes)
deleted 5 from position 0
after                      10 -> 20 -> 30 -> NULL   (3 nodes)
deleted 30, which was the tail
after                      10 -> 20 -> NULL   (2 nodes)
append after that          10 -> 20 -> 99 -> NULL   (3 nodes)
the append worked, so the tail was updated
after destroy              NULL   (0 nodes)

Three lines in that listing are the whole of Course Objective 9.

free(gone) in delete_at_position. In Python, unlinking the node was enough. In C the node is still allocated and still yours, and not freeing it is a leak that grows every time a node is deleted.

Node *next = here->next; before free(here) in destroy. Reading here->next after freeing here is use after free, and it may appear to work, which is what makes it dangerous.

munotes.in256

If Your College Runs Module 2 in C

destroy at the end. A program that exits without freeing is forgiven by the operating system; a function that does it in a loop is not.

Practical 3: the stack, and infix to postfix

#include <stdio.h>
#include <string.h>

#define STACK_MAX 64

typedef struct {
    char slots[STACK_MAX];
    int top;                    /* -1 when empty */
} CharStack;

static void init(CharStack *s) { s->top = -1; }
static int is_empty(const CharStack *s) { return s->top == -1; }
static int is_full(const CharStack *s) { return s->top == STACK_MAX - 1; }

static int push(CharStack *s, char c) {
    if (is_full(s)) return -1;
    s->slots[++s->top] = c;
    return 0;
}

static int pop(CharStack *s, char *out) {
    if (is_empty(s)) return -1;
    *out = s->slots[s->top--];
    return 0;
}

static int peek(const CharStack *s, char *out) {
    if (is_empty(s)) return -1;
    *out = s->slots[s->top];
    return 0;
}

static int precedence(char op) {
    switch (op) {
        case '^': return 3;
        case '*': case '/': case '%': return 2;
        case '+': case '-': return 1;
        default: return 0;
    }
}

static int is_operator(char c) { return precedence(c) > 0; }

static int pops_first(char on_stack, char arriving) {
    if (on_stack == '(') return 0;
    if (precedence(on_stack) > precedence(arriving)) return 1;
    if (precedence(on_stack) < precedence(arriving)) return 0;
    return arriving != '^';                 /* equal: left associative pops */
}

static int to_postfix(const char *infix, char *out, size_t out_size) {
    CharStack s;
    init(&s);
    size_t written = 0;
    for (const char *p = infix; *p; p++) {
        char c = *p;
        if (c == ' ') continue;
        if (c == '(') {
            if (push(&s, c) != 0) return -1;
        } else if (c == ')') {
            char top = 0;
            while (peek(&s, &top) == 0 && top != '(') {
                pop(&s, &top);
                if (written + 2 >= out_size) return -2;
                out[written++] = top;
                out[written++] = ' ';
            }
            if (pop(&s, &top) != 0) return -3;        /* unbalanced */
        } else if (is_operator(c)) {
            char top = 0;
            while (peek(&s, &top) == 0 && pops_first(top, c)) {
                pop(&s, &top);
                if (written + 2 >= out_size) return -2;
                out[written++] = top;
                out[written++] = ' ';
            }
            if (push(&s, c) != 0) return -1;
        } else {
            if (written + 2 >= out_size) return -2;
            out[written++] = c;
            out[written++] = ' ';
        }
    }
    char top = 0;
    while (pop(&s, &top) == 0) {
        if (top == '(') return -3;                    /* unbalanced */
        if (written + 2 >= out_size) return -2;
        out[written++] = top;
        out[written++] = ' ';
    }
    if (written > 0) written--;                       /* drop the trailing space */
    out[written] = '\0';
    return 0;
}

int main(void) {
    const char *tests[] = {"a + b * c - d", "( a + b ) * c", "a ^ b ^ c",
                           "a - b - c", "( a + b", "a + b )"};
    char out[128];
    for (size_t i = 0; i < sizeof(tests) / sizeof(tests[0]); i++) {
        int status = to_postfix(tests[i], out, sizeof(out));
        if (status == 0) {
            printf("%-16s -> %s\n", tests[i], out);
        } else {
            printf("%-16s -> refused, status %d\n", tests[i], status);
        }
    }
    return 0;
}
munotes.in257

If Your College Runs Module 2 in C

a + b * c - d    -> a b c * + d -
( a + b ) * c    -> a b + c *
a ^ b ^ c        -> a b c ^ ^
a - b - c        -> a b - c -
( a + b          -> refused, status -3
a + b )          -> refused, status -3

Compare those results with [Practical 3 continued: Infix to Postfix with a Stack]: they are identical, because the algorithm is the same. What C adds is the out_size check on every write, because a C string has no length and nothing stops you writing past the end of the buffer. Every written + 2 >= out_size in that listing is there for that reason, and leaving them out is how a real program is exploited.

Practical 4: the circular queue

#include <stdio.h>

#define QUEUE_MAX 5

typedef struct {
    int slots[QUEUE_MAX];
    int front;
    int rear;
    int count;
} Queue;

static void init(Queue *q) {
    q->front = 0;
    q->rear = -1;
    q->count = 0;
}

static int is_empty(const Queue *q) { return q->count == 0; }
static int is_full(const Queue *q) { return q->count == QUEUE_MAX; }

static int enqueue(Queue *q, int value) {
    if (is_full(q)) return -1;
    q->rear = (q->rear + 1) % QUEUE_MAX;     /* the wrap */
    q->slots[q->rear] = value;
    q->count++;
    return 0;
}

static int dequeue(Queue *q, int *out) {
    if (is_empty(q)) return -1;
    *out = q->slots[q->front];
    q->front = (q->front + 1) % QUEUE_MAX;
    q->count--;
    return 0;
}

static void show(const char *label, const Queue *q) {
    printf("%-18s front %d rear %2d count %d  order:", label,
           q->front, q->rear, q->count);
    for (int i = 0; i < q->count; i++) {
        printf(" %d", q->slots[(q->front + i) % QUEUE_MAX]);
    }
    printf("\n");
}

int main(void) {
    Queue q;
    init(&q);
    show("new", &q);

    for (int v = 10; v <= 50; v += 10) {
        enqueue(&q, v);
    }
    show("five enqueued", &q);

    int out = 0;
    for (int i = 0; i < 3; i++) {
        dequeue(&q, &out);
        printf("dequeued %-9d ", out);
        show("", &q);
    }

    printf("now the enqueue a LINEAR queue would refuse:\n");
    enqueue(&q, 60);
    show("enqueue 60", &q);
    enqueue(&q, 70);
    show("enqueue 70", &q);
    printf("rear wrapped round to a lower index, and the order is still FIFO\n");

    if (enqueue(&q, 80) != 0) {
        printf("enqueue on a full queue refused\n");
    }
    init(&q);
    if (dequeue(&q, &out) != 0) {
        printf("dequeue on an empty queue refused\n");
    }
    return 0;
}
munotes.in258

If Your College Runs Module 2 in C

new                front 0 rear -1 count 0  order:
five enqueued      front 0 rear  4 count 5  order: 10 20 30 40 50
dequeued 10                           front 1 rear  4 count 4  order: 20 30 40 50
dequeued 20                           front 2 rear  4 count 3  order: 30 40 50
dequeued 30                           front 3 rear  4 count 2  order: 40 50
now the enqueue a LINEAR queue would refuse:
enqueue 60         front 3 rear  0 count 3  order: 40 50 60
enqueue 70         front 3 rear  1 count 4  order: 40 50 60 70
rear wrapped round to a lower index, and the order is still FIFO
dequeue on an empty queue refused

Practical 5 and 6: the binary search tree and its traversals

#include <stdio.h>
#include <stdlib.h>

typedef struct TreeNode {
    int key;
    struct TreeNode *left;
    struct TreeNode *right;
} TreeNode;

static TreeNode *make_node(int key) {
    TreeNode *node = malloc(sizeof(TreeNode));
    if (node == NULL) return NULL;
    node->key = key;
    node->left = NULL;
    node->right = NULL;
    return node;
}

static TreeNode *insert(TreeNode *root, int key) {
    if (root == NULL) return make_node(key);
    if (key < root->key) {
        root->left = insert(root->left, key);
    } else if (key > root->key) {
        root->right = insert(root->right, key);
    }
    return root;                              /* a duplicate changes nothing */
}

static int search(const TreeNode *root, int key, int *comparisons) {
    *comparisons = 0;
    while (root != NULL) {
        (*comparisons)++;
        if (key == root->key) return 1;
        root = key < root->key ? root->left : root->right;
    }
    return 0;
}

static void pre_order(const TreeNode *node) {
    if (node == NULL) return;
    printf(" %d", node->key);
    pre_order(node->left);
    pre_order(node->right);
}

static void in_order(const TreeNode *node) {
    if (node == NULL) return;
    in_order(node->left);
    printf(" %d", node->key);
    in_order(node->right);
}

static void post_order(const TreeNode *node) {
    if (node == NULL) return;
    post_order(node->left);
    post_order(node->right);
    printf(" %d", node->key);
}

static int height(const TreeNode *node) {
    if (node == NULL) return -1;
    int left = height(node->left);
    int right = height(node->right);
    return 1 + (left > right ? left : right);
}

static int minimum(const TreeNode *node, int *found) {
    if (node == NULL) return -1;
    while (node->left != NULL) node = node->left;
    *found = node->key;
    return 0;
}

static void destroy(TreeNode *node) {
    if (node == NULL) return;
    destroy(node->left);          /* POST-ORDER: the children go first */
    destroy(node->right);
    free(node);
}

int main(void) {
    TreeNode *root = NULL;
    int keys[] = {50, 30, 70, 20, 40, 60, 80};
    for (size_t i = 0; i < sizeof(keys) / sizeof(keys[0]); i++) {
        root = insert(root, keys[i]);
    }

    printf("pre-order :");
    pre_order(root);
    printf("\nin-order  :");
    in_order(root);
    printf("\npost-order:");
    post_order(root);
    printf("\n");
    printf("the in-order walk is sorted, which proves it is a BST\n");
    printf("height    : %d\n", height(root));

    int smallest = 0;
    minimum(root, &smallest);
    printf("minimum   : %d, the leftmost node\n", smallest);

    int comparisons = 0;
    int targets[] = {50, 20, 80, 45};
    for (size_t i = 0; i < sizeof(targets) / sizeof(targets[0]); i++) {
        int found = search(root, targets[i], &comparisons);
        printf("search %-3d: %-9s in %d comparison(s)\n", targets[i],
               found ? "found" : "NOT FOUND", comparisons);
    }

    destroy(root);
    printf("the tree was freed POST-order, children before parents\n");
    return 0;
}
munotes.in259

If Your College Runs Module 2 in C

pre-order : 50 30 20 40 70 60 80
in-order  : 20 30 40 50 60 70 80
post-order: 20 40 30 60 80 70 50
the in-order walk is sorted, which proves it is a BST
height    : 2
minimum   : 20, the leftmost node
search 50 : found     in 1 comparison(s)
search 20 : found     in 3 comparison(s)
search 80 : found     in 3 comparison(s)
search 45 : NOT FOUND in 3 comparison(s)
the tree was freed POST-order, children before parents

destroy must be post-order. Freeing a node before its children means the pointers to those children are gone, and they can never be freed: a leak the size of the whole subtree. That is the clearest example in this book of why the three traversals are three different tools rather than three styles.

Practical 7: the hash table with separate chaining

#include <stdio.h>
#include <stdlib.h>
#include <string.h>

#define BUCKETS 7
#define KEY_MAX 32

typedef struct Entry {
    char key[KEY_MAX];
    int value;
    struct Entry *next;         /* the CHAIN */
} Entry;

typedef struct {
    Entry *buckets[BUCKETS];
    int count;
} HashTable;

static void init(HashTable *table) {
    for (int i = 0; i < BUCKETS; i++) table->buckets[i] = NULL;
    table->count = 0;
}

static unsigned long string_hash(const char *key) {
    unsigned long total = 0;
    for (const char *p = key; *p; p++) {
        total = total * 31u + (unsigned char)*p;
    }
    return total;
}

static int bucket_of(const char *key) {
    return (int)(string_hash(key) % BUCKETS);
}

static int put(HashTable *table, const char *key, int value) {
    int index = bucket_of(key);
    for (Entry *e = table->buckets[index]; e != NULL; e = e->next) {
        if (strcmp(e->key, key) == 0) {
            e->value = value;                 /* an UPDATE, not a second entry */
            return index;
        }
    }
    Entry *entry = malloc(sizeof(Entry));
    if (entry == NULL) return -1;
    snprintf(entry->key, sizeof(entry->key), "%s", key);
    entry->value = value;
    entry->next = table->buckets[index];      /* insert at the head of the chain */
    table->buckets[index] = entry;
    table->count++;
    return index;
}

static int get(const HashTable *table, const char *key, int *value,
               int *comparisons) {
    *comparisons = 0;
    for (Entry *e = table->buckets[bucket_of(key)]; e != NULL; e = e->next) {
        (*comparisons)++;
        if (strcmp(e->key, key) == 0) {
            *value = e->value;
            return 1;
        }
    }
    return 0;
}

static void show(const HashTable *table) {
    for (int i = 0; i < BUCKETS; i++) {
        printf("  %d:", i);
        for (Entry *e = table->buckets[i]; e != NULL; e = e->next) {
            printf(" (%s, %d)%s", e->key, e->value, e->next ? " ->" : "");
        }
        printf("\n");
    }
}

static void destroy(HashTable *table) {
    for (int i = 0; i < BUCKETS; i++) {
        Entry *here = table->buckets[i];
        while (here != NULL) {
            Entry *next = here->next;         /* save it before freeing */
            free(here);
            here = next;
        }
        table->buckets[i] = NULL;
    }
    table->count = 0;
}

int main(void) {
    HashTable table;
    init(&table);

    const char *names[] = {"Aarti", "Divya", "Chetan", "Eshan", "Farhan", "Gauri"};
    int marks[] = {78, 55, 90, 67, 41, 83};
    for (size_t i = 0; i < sizeof(marks) / sizeof(marks[0]); i++) {
        int index = put(&table, names[i], marks[i]);
        printf("put %-8s -> bucket %d\n", names[i], index);
    }

    printf("the table:\n");
    show(&table);

    int value = 0;
    int comparisons = 0;
    const char *wanted[] = {"Aarti", "Gauri", "Nobody"};
    for (size_t i = 0; i < sizeof(wanted) / sizeof(wanted[0]); i++) {
        if (get(&table, wanted[i], &value, &comparisons)) {
            printf("get %-8s -> %d in %d comparison(s)\n", wanted[i], value,
                   comparisons);
        } else {
            printf("get %-8s -> not found, %d comparison(s)\n", wanted[i],
                   comparisons);
        }
    }

    put(&table, "Aarti", 91);
    get(&table, "Aarti", &value, &comparisons);
    printf("stored Aarti again: value now %d, and the count is still %d\n",
           value, table.count);

    destroy(&table);
    printf("every chain freed\n");
    return 0;
}
munotes.in260

If Your College Runs Module 2 in C

put Aarti    -> bucket 4
put Divya    -> bucket 2
put Chetan   -> bucket 2
put Eshan    -> bucket 0
put Farhan   -> bucket 1
put Gauri    -> bucket 0
the table:
  0: (Gauri, 83) -> (Eshan, 67)
  1: (Farhan, 41)
  2: (Chetan, 90) -> (Divya, 55)
  3:
  4: (Aarti, 78)
  5:
  6:
get Aarti    -> 78 in 1 comparison(s)
get Gauri    -> 83 in 1 comparison(s)
get Nobody   -> not found, 2 comparison(s)
stored Aarti again: value now 91, and the count is still 6
every chain freed

Two C specific points. strcmp(a, b) == 0 compares strings, and a == b compares the addresses, which is the single most common C bug in a hash table. And snprintf is used rather than strcpy, because it cannot write past the end of the buffer.

Practicals 8 and 9: the three sorts and the two searches

#include <stdio.h>

typedef struct {
    long comparisons;
    long movements;
    int passes;
} Counter;

static void copy_array(const int from[], int to[], int n) {
    for (int i = 0; i < n; i++) to[i] = from[i];
}

static void bubble_sort(int a[], int n, Counter *c) {
    c->comparisons = 0; c->movements = 0; c->passes = 0;
    for (int i = 0; i < n - 1; i++) {
        int swapped = 0;
        c->passes++;
        for (int j = 0; j < n - 1 - i; j++) {        /* n - 1 - i, not n - 1 */
            c->comparisons++;
            if (a[j] > a[j + 1]) {
                int t = a[j]; a[j] = a[j + 1]; a[j + 1] = t;
                c->movements += 2;
                swapped = 1;
            }
        }
        if (!swapped) break;                         /* the early exit */
    }
}

static void insertion_sort(int a[], int n, Counter *c) {
    c->comparisons = 0; c->movements = 0; c->passes = 0;
    for (int i = 1; i < n; i++) {
        c->passes++;
        int value = a[i];
        int j = i - 1;
        while (j >= 0) {
            c->comparisons++;
            if (a[j] <= value) break;
            a[j + 1] = a[j];
            c->movements++;
            j--;
        }
        a[j + 1] = value;
        c->movements++;
    }
}

static void selection_sort(int a[], int n, Counter *c) {
    c->comparisons = 0; c->movements = 0; c->passes = 0;
    for (int i = 0; i < n - 1; i++) {
        c->passes++;
        int smallest = i;
        for (int j = i + 1; j < n; j++) {
            c->comparisons++;
            if (a[j] < a[smallest]) smallest = j;
        }
        if (smallest != i) {
            int t = a[i]; a[i] = a[smallest]; a[smallest] = t;
            c->movements += 2;
        }
    }
}

static int linear_search(const int a[], int n, int target, long *comparisons) {
    *comparisons = 0;
    for (int i = 0; i < n; i++) {
        (*comparisons)++;
        if (a[i] == target) return i;
    }
    return -1;
}

static int binary_search(const int a[], int n, int target, long *comparisons) {
    int low = 0, high = n - 1;
    *comparisons = 0;
    while (low <= high) {                            /* <=, not < */
        int middle = low + (high - low) / 2;         /* cannot overflow */
        (*comparisons)++;
        if (a[middle] == target) return middle;
        if (a[middle] < target) low = middle + 1;
        else high = middle - 1;
    }
    return -1;
}

static int is_sorted(const int a[], int n) {
    for (int i = 1; i < n; i++) if (a[i - 1] > a[i]) return 0;
    return 1;
}

#define N 40

int main(void) {
    int sorted_data[N], reversed[N], scratch[N];
    for (int i = 0; i < N; i++) {
        sorted_data[i] = i;
        reversed[i] = N - i;
    }

    Counter c;
    printf("n = %d, and n(n-1)/2 = %d\n\n", N, N * (N - 1) / 2);

    const char *labels[] = {"already sorted", "reverse sorted"};
    int *datasets[] = {sorted_data, reversed};
    for (int d = 0; d < 2; d++) {
        printf("%s:\n", labels[d]);
        printf("  %-12s %12s %11s %7s\n", "sort", "comparisons", "movements",
               "passes");

        copy_array(datasets[d], scratch, N);
        bubble_sort(scratch, N, &c);
        printf("  %-12s %12ld %11ld %7d  sorted? %d\n", "bubble",
               c.comparisons, c.movements, c.passes, is_sorted(scratch, N));

        copy_array(datasets[d], scratch, N);
        insertion_sort(scratch, N, &c);
        printf("  %-12s %12ld %11ld %7d  sorted? %d\n", "insertion",
               c.comparisons, c.movements, c.passes, is_sorted(scratch, N));

        copy_array(datasets[d], scratch, N);
        selection_sort(scratch, N, &c);
        printf("  %-12s %12ld %11ld %7d  sorted? %d\n", "selection",
               c.comparisons, c.movements, c.passes, is_sorted(scratch, N));
        printf("\n");
    }

    printf("searching the sorted array of %d items:\n", N);
    printf("  %-8s %14s %14s\n", "target", "linear", "binary");
    long linear_count = 0, binary_count = 0;
    int targets[] = {0, 20, 39, 99};
    for (size_t i = 0; i < sizeof(targets) / sizeof(targets[0]); i++) {
        linear_search(sorted_data, N, targets[i], &linear_count);
        binary_search(sorted_data, N, targets[i], &binary_count);
        printf("  %-8d %14ld %14ld\n", targets[i], linear_count, binary_count);
    }
    return 0;
}
munotes.in261

If Your College Runs Module 2 in C

n = 40, and n(n-1)/2 = 780

already sorted:
  sort          comparisons   movements  passes
  bubble                 39           0       1  sorted? 1
  insertion              39          39      39  sorted? 1
  selection             780           0      39  sorted? 1

reverse sorted:
  sort          comparisons   movements  passes
  bubble                780        1560      39  sorted? 1
  insertion             780         819      39  sorted? 1
  selection             780          40      39  sorted? 1

searching the sorted array of 40 items:
  target           linear         binary
  0                     1              5
  20                   21              5
  39                   40              6
  99                   40              6
munotes.in262

If Your College Runs Module 2 in C

Compare those counts with [Practical 8: Bubble, Insertion and Selection Sort Compared] and [Practical 9: Linear and Binary Search Compared]. They are the same numbers, because a comparison count is a property of the algorithm and not of the language. That is the strongest possible argument that the language is the least interesting thing about this module.

Practical 10: the combined application in C

#include <stdio.h>
#include <stdlib.h>
#include <string.h>

#define BUCKETS 7
#define TITLE_MAX 40
#define QUEUE_MAX 8
#define UNDO_MAX 8

/* the CATALOGUE: a hash table, for lookup by accession number */
typedef struct Book {
    char accession[8];
    char title[TITLE_MAX];
    struct Book *next;
} Book;

/* the TITLE INDEX: a binary search tree, for the sorted listing */
typedef struct TitleNode {
    char title[TITLE_MAX];
    char accession[8];
    struct TitleNode *left;
    struct TitleNode *right;
} TitleNode;

typedef struct {
    Book *buckets[BUCKETS];
    TitleNode *titles;
    char queue[QUEUE_MAX][16];      /* RESERVATIONS: a circular queue */
    int front, rear, waiting;
    char undo[UNDO_MAX][32];        /* UNDO: a stack */
    int top;
} Desk;

static unsigned long string_hash(const char *key) {
    unsigned long total = 0;
    for (const char *p = key; *p; p++) total = total * 31u + (unsigned char)*p;
    return total;
}

static void desk_init(Desk *d) {
    for (int i = 0; i < BUCKETS; i++) d->buckets[i] = NULL;
    d->titles = NULL;
    d->front = 0; d->rear = -1; d->waiting = 0;
    d->top = -1;
}

static TitleNode *title_insert(TitleNode *node, const char *title,
                               const char *accession) {
    if (node == NULL) {
        TitleNode *fresh = malloc(sizeof(TitleNode));
        if (fresh == NULL) return NULL;
        snprintf(fresh->title, TITLE_MAX, "%s", title);
        snprintf(fresh->accession, 8, "%s", accession);
        fresh->left = fresh->right = NULL;
        return fresh;
    }
    if (strcmp(title, node->title) < 0) {
        node->left = title_insert(node->left, title, accession);
    } else if (strcmp(title, node->title) > 0) {
        node->right = title_insert(node->right, title, accession);
    }
    return node;
}

static void title_in_order(const TitleNode *node) {
    if (node == NULL) return;
    title_in_order(node->left);
    printf("   %-34s %s\n", node->title, node->accession);
    title_in_order(node->right);
}

static void title_destroy(TitleNode *node) {
    if (node == NULL) return;
    title_destroy(node->left);
    title_destroy(node->right);
    free(node);
}

static int add_book(Desk *d, const char *accession, const char *title) {
    int index = (int)(string_hash(accession) % BUCKETS);
    for (Book *b = d->buckets[index]; b != NULL; b = b->next) {
        if (strcmp(b->accession, accession) == 0) return -1;
    }
    Book *book = malloc(sizeof(Book));
    if (book == NULL) return -2;
    snprintf(book->accession, 8, "%s", accession);
    snprintf(book->title, TITLE_MAX, "%s", title);
    book->next = d->buckets[index];
    d->buckets[index] = book;
    d->titles = title_insert(d->titles, title, accession);
    return index;
}

static const char *find_title(const Desk *d, const char *accession) {
    int index = (int)(string_hash(accession) % BUCKETS);
    for (Book *b = d->buckets[index]; b != NULL; b = b->next) {
        if (strcmp(b->accession, accession) == 0) return b->title;
    }
    return NULL;
}

static int reserve(Desk *d, const char *borrower) {
    if (d->waiting == QUEUE_MAX) return -1;
    d->rear = (d->rear + 1) % QUEUE_MAX;
    snprintf(d->queue[d->rear], 16, "%s", borrower);
    d->waiting++;
    return d->waiting;
}

static int next_in_line(Desk *d, char *out, size_t out_size) {
    if (d->waiting == 0) return -1;
    snprintf(out, out_size, "%s", d->queue[d->front]);
    d->front = (d->front + 1) % QUEUE_MAX;
    d->waiting--;
    return 0;
}

static int remember(Desk *d, const char *action) {
    if (d->top == UNDO_MAX - 1) return -1;
    snprintf(d->undo[++d->top], 32, "%s", action);
    return 0;
}

static int undo_last(Desk *d, char *out, size_t out_size) {
    if (d->top == -1) return -1;
    snprintf(out, out_size, "%s", d->undo[d->top--]);
    return 0;
}

static void desk_destroy(Desk *d) {
    for (int i = 0; i < BUCKETS; i++) {
        Book *here = d->buckets[i];
        while (here != NULL) {
            Book *next = here->next;
            free(here);
            here = next;
        }
        d->buckets[i] = NULL;
    }
    title_destroy(d->titles);
    d->titles = NULL;
}

int main(void) {
    Desk desk;
    desk_init(&desk);

    printf("1. the CATALOGUE, a hash table keyed on accession number:\n");
    add_book(&desk, "A104", "Data Structures");
    add_book(&desk, "A101", "Let Us Python");
    add_book(&desk, "A109", "Operating System Concepts");
    add_book(&desk, "A102", "Learning Python");
    printf("   A104 -> %s\n", find_title(&desk, "A104"));
    printf("   A999 -> %s\n",
           find_title(&desk, "A999") ? find_title(&desk, "A999") : "not found");

    printf("2. the TITLE INDEX, a BST, listed by its in-order walk:\n");
    title_in_order(desk.titles);

    printf("3. RESERVATIONS, a circular queue:\n");
    printf("   Bhavesh is number %d in the queue\n", reserve(&desk, "Bhavesh"));
    printf("   Chetan  is number %d in the queue\n", reserve(&desk, "Chetan"));
    char who[16];
    next_in_line(&desk, who, sizeof(who));
    printf("   the book is returned and goes to %s, who was first\n", who);
    printf("   still waiting: %d\n", desk.waiting);

    printf("4. UNDO, a stack:\n");
    remember(&desk, "issued A104 to Aarti");
    remember(&desk, "returned A104");
    char action[32];
    undo_last(&desk, action, sizeof(action));
    printf("   undoing: %s\n", action);
    undo_last(&desk, action, sizeof(action));
    printf("   undoing: %s\n", action);
    if (undo_last(&desk, action, sizeof(action)) != 0) {
        printf("   nothing left to undo\n");
    }

    desk_destroy(&desk);
    printf("everything freed\n");
    return 0;
}
munotes.in263

If Your College Runs Module 2 in C

1. the CATALOGUE, a hash table keyed on accession number:
   A104 -> Data Structures
   A999 -> not found
2. the TITLE INDEX, a BST, listed by its in-order walk:
   Data Structures                    A104
   Learning Python                    A102
   Let Us Python                      A101
   Operating System Concepts          A109
3. RESERVATIONS, a circular queue:
   Bhavesh is number 1 in the queue
   Chetan  is number 2 in the queue
   the book is returned and goes to Bhavesh, who was first
   still waiting: 1
4. UNDO, a stack:
   undoing: returned A104
   undoing: issued A104 to Aarti
   nothing left to undo
everything freed
munotes.in264

If Your College Runs Module 2 in C

Compiling and running, and the one flag to add

cc -std=c17 -Wall -Wextra -o practical practical.c
./practical

-Wall -Wextra is not optional. A warning is a bug the compiler found for you, and every listing in this chapter compiles with none.

For anything that uses malloc, add a leak check. On Linux:

valgrind --leak-check=full ./practical

and on macOS, leaks --atExit -- ./practical. A clean run reports no leaks, and a journal entry that says "no leaks, checked with valgrind" is worth a mark on a C data structures practical.

The translation table

In PythonIn C
class Node:a struct holding the data and a next pointer, named with a typedef
node = Node(5)malloc(sizeof(Node)), then check it against NULL
nothing, it is automaticfree(node); node = NULL;
NoneNULL
node.datanode->data through a pointer, node.data on a value
len(list)a count field you maintain, or a walk
a, b = f()f(&a, &b), writing through pointers
raise ValueErrorreturn -1; and the caller checks
a == b for stringsstrcmp(a, b) compared against 0
f"{x}"printf with the right conversion
a list of any lengtha fixed array, or malloc plus realloc
for item in items:for (int i = 0; i < n; i++)
__repr__a show() function you call yourself

Procedure

  1. Install a compiler: gcc or clang on Linux, Xcode command line tools on macOS, MinGW or WSL on

Windows.

  1. Compile every program with cc -std=c17 -Wall -Wextra and fix every warning before running it.
  2. Write each exercise as its own .c file with a main that exercises it and prints the state.
  3. For every malloc: check the return against NULL, and write the matching free before you write

anything else.

  1. Set a pointer to NULL after freeing it.
  2. In destroy for a linked list, save here->next before free(here).
  3. In destroy for a tree, free post-order: both children before the parent.
  4. Use strcmp to compare strings, never ==, and snprintf rather than strcpy.
  5. Return a status code from every function that can fail, and check it at the call site.
  6. Run it under valgrind or leaks and record "no leaks" in the journal.
munotes.in265

If Your College Runs Module 2 in C

Result

All ten exercises compiled with -std=c17 -Wall -Wextra and no warnings, and each ran and printed the same results as its Python counterpart. The array shifted backwards on insertion and forwards on deletion. The linked list handled the empty, one node and tail cases, and an append immediately after deleting the tail appeared correctly, which proves the tail was updated. The infix to postfix conversion gave the same strings as the Python version, and the sorting and searching comparison counts were identical to the Python ones, because a comparison count is a property of the algorithm. Every structure was freed: the list node by node, the tree post-order, and each hash bucket's chain.

Where marks are lost

  • No free. Every malloc needs one, and a delete that does not free leaks on every call.
  • Not checking malloc. A NULL used as a pointer crashes with no message.
  • free in the wrong order, either reading here->next after freeing here, or freeing a tree

parent before its children.

  • values = realloc(values, ...), which loses the old pointer and leaks when realloc fails.
  • a == b for strings. That compares addresses. Use strcmp.
  • strcpy into a fixed buffer with no length check. Use snprintf.
  • Ignoring a function's status. C has no exceptions, so an unchecked return is an ignored error.
  • Compiling without -Wall -Wextra, then shipping a warning that was a real bug.
  • while (low < high) in binary search, exactly as in Python.
  • Forgetting the & when passing an output parameter, or the * when writing through it.

For the journal

Write the same twenty entries you would write in Python, with the C program in each. Add three things that are specific to C and are what a C practical is marked on.

One free for every malloc, pointed out in the listing where it happens, with a sentence saying what Python did automatically.

The order of the frees: next saved before free in a list, and post-order in a tree, each with one sentence on what goes wrong otherwise.

A leak check: the valgrind or leaks command and the line of its output saying no leaks were found.

Quick revision

  • The structures and algorithms are identical; only the language changes. The comparison counts are

the same numbers.

  • struct plus typedef for a record. -> through a pointer, . on a value.
  • malloc(sizeof(Node)), check against NULL, and one free for every malloc.
  • NULL is Python's None. Set a pointer to NULL after freeing.
  • realloc failing does not free the old block: assign to a new pointer and check it first.
  • A function returns extra values through pointer parameters: f(&a, &b).
  • No exceptions: return a status code and check it at every call site.
  • strcmp(a, b) == 0 for strings, never ==. snprintf, never strcpy.
  • A C string has no length: check every write against the buffer size.
  • Free a list with next saved before the free. Free a tree post-order, children before parent.
  • Compile with -std=c17 -Wall -Wextra and treat every warning as a bug.
  • Check for leaks with valgrind --leak-check=full or leaks --atExit.
munotes.in266

If Your College Runs Module 2 in C

Questions you should be able to answer

1. What changes between the Python and the C version of a linked list? The syntax and the memory management. The structure, the three deletion cases, the trailing reference and the tail update are identical.

2. Why malloc(sizeof(Node)) and not malloc(20)? Because sizeof is computed by the compiler and stays correct when the struct changes.

3. What must you do with malloc's return before using it? Check it against NULL. malloc returns NULL when it cannot allocate, and using that as a pointer crashes.

4. How many free calls does one malloc need? Exactly one. None is a leak and two is undefined behaviour.

5. Why set a pointer to NULL after freeing it? So that a later use crashes immediately instead of quietly reading memory that has been handed back, which is far harder to diagnose.

6. What is wrong with values = realloc(values, bigger)? If realloc fails it returns NULL and does not free the old block, so assigning its result over values loses the only pointer to that block. Assign to a new pointer and check it.

7. Why must a list's destroy save here->next before calling free(here)? Because after the free, reading here->next is use after free. It often appears to work, which is why it is dangerous.

8. In what order must a tree be freed, and why? Post-order. A parent holds the only pointers to its children, so freeing it first makes them unreachable and they can never be freed.

9. How do you compare two strings in C? strcmp(a, b) == 0. Using == compares the addresses of the two arrays, not their contents.

10. How does a C function return two values? Through pointer parameters: the caller passes &a and &b and the function writes through them.

11. Why are the comparison counts the same as the Python ones? Because a comparison count is a property of the algorithm, not of the language it is written in.

Contents This chapter on its own page

munotes.in267

Module J

The journal, the viva and the practical examination

munotes.in

Chapter Thirty-Six

Keeping the Journal, and What Goes on the Page

Syllabus topic MU's item 13, "2.5 marks can be awarded for each practical performance and writeup submission totaling to 50 marks and can be converted to 20 marks", and item 14, "Certified copy of Journal is compulsory to appear for the practical examination"

Aim

To know exactly what the journal must contain, what one entry looks like, and what the teacher's signature is given for.

What MU says about the journal, in her own words

Two rows of the paper's particulars table mention it, and they say different things.

Item 13, the internal 20 marks:

Students are expected to attend each practical and submit the written practical of the previous session.

Performing Practical and writeup submission will be continuous internal evaluation. 2.5 marks can be

awarded for each practical performance and writeup submission totaling to 50 marks and can be converted

to 20 marks.

Item 14, the external 30 marks:

Format of Question Paper: Duration 2 hours. Certified copy of Journal is compulsory to appear for the

practical examination

Q1. From Module 1 13 marks

Q2. From Module 2 12marks

Q3. Journal and Viva 05 marks

Three facts follow, and all three are worth acting on.

Twenty entries are expected. 2.5 marks for each performance and writeup, totalling 50. There are exactly twenty exercises printed, and 2.5 multiplied by 20 is 50. There is no slack in that arithmetic.

The writeup is due at the NEXT session, not at the end of the semester. "Submit the written practical of the previous session" is MU's own wording, and a journal handed in as one heap in the last week has usually already lost part of these marks.

A certified copy is compulsory to appear. Certified means signed. This is the only line in the whole scheme that can cost a student an entire paper, and it is settled in the laboratory weeks before the examination.

The shape of the journal

PageWhat goes on it
frontyour name, roll number, class, division, the college, the subject, the academic year
certificatethe printed page your college provides, signed by the teacher and the head
indexone row per practical: number, aim, date performed, date submitted, page, signature
the entriestwenty of them, in MU's order
at the backanything extra: a correction, a second attempt, a note about a machine

The index is the page the examiner looks at first. It should show twenty rows, every one dated and signed. A gap in the signature column is a question you will be asked.

One entry, in full

An entry has eight parts. The first five are the marks; the last three are what separates a good entry from a complete one.

PartWhat it is
1. Practical number and dateMU's number, the date performed, the date submitted
2. AimMU's own wording, copied
3. Theory, brieflythe idea in three to six lines, in your own words
4. Programthe listing, indented as it is on screen
5. Outputwhat it actually printed, not what it should have
6. Conclusionone or two sentences on what the exercise showed
7. The cost, where there is onethe comparison count, the shift count, the Big O
8. Teacher's signature and datetheirs, not yours
munotes.in268

Keeping the Journal, and What Goes on the Page

Here is a complete entry, written out, so there is a model to copy.

Practical No. 12                    Performed on: 05/10/2026
                                    Submitted on: 12/10/2026

AIM
Linked List Manipulation. Write a program to: create a singly linked list;
insert a node at the beginning, end, and at a given position in a linked
list; delete a node from a given position in a linked list.

THEORY
A singly linked list is a chain of nodes. Each node holds one value and a
reference to the next node, and the last node's reference is None, which is
how the end is recognised. The list itself holds only the head, the first
node. A tail and a count are kept as well, so that appending and len() are
one step instead of a walk. Deletion needs the node IN FRONT of the one being
removed, because a node has no reference backwards, so the walk carries a
trailing reference.

PROGRAM
    class Node:
        def __init__(self, data, nxt=None):
            self.data = data
            self.nxt = nxt
    ...

OUTPUT
    empty                     None   (empty)   len 0
    insert_at_end(10)         10 -> None   len 1
    insert_at_beginning(5)    5 -> 10 -> None   len 2
    insert_at_position(1, 7)  5 -> 7 -> 10 -> None   len 3  walked 0
    delete position 1         5 -> 10 -> None   removed 7, walked 0
    delete the TAIL           5 -> None   removed 10, walked 0
    append after that delete  5 -> 99 -> None
    the append worked, so the tail was updated

COST
    insert at the beginning  O(1), two assignments
    insert at the end        O(1), because a tail is kept
    insert at position p     O(p), the walk to the node before it
    delete at position p     O(p), same walk, plus the tail update
    get(p) and search        O(n), there is no address arithmetic

CONCLUSION
A linked list inserts and deletes at either end in one step and has to walk
to reach position n, which is the exact opposite of an array. The tail must
be updated when the last node is deleted; appending straight afterwards is
the test, and it worked.

                                    Teacher's signature: ____________

Four things about that entry are worth naming.

The aim is MU's wording. Copy it. An aim in your own words invites the question of whether you read hers.

The output is the real output, including the walk counts. An output written from what the program should do is the one thing an examiner can catch without reading the code at all.

munotes.in269

Keeping the Journal, and What Goes on the Page

The COST block is short and it is what makes the entry good. Two or three lines of Big O with a reason each. Most entries do not have it.

The conclusion answers "so what". Not "thus the program was executed successfully", which says nothing, but what the exercise actually showed.

The twenty entries, and the aim of each

This is the table to copy into the front of your journal. The aims are MU's own, shortened only where her row runs to several lines.

No.ModuleAim, in MU's wordsThe chapter
11name and age, the year they turn 100; even or odd; the SGPI grade[Practical 1: Input, Output, and the Year You Turn 100]
21the Fibonacci series; the sum of the digits of a number[Practical 2: the Fibonacci Series, and the Sum of the Digits]
31basic operations, indexing and slicing on arrays; mathematical functions; aliasing and copying[Practical 3: Arrays, Basic Operations, Indexing and Slicing]
41slicing, basic and advanced indexing on NumPy arrays; dimensions and attributes[Practical 4: NumPy Slicing, Basic and Advanced Indexing]
51Armstrong and palindrome functions; a recursive factorial; a lambda for startswith[Practical 5: Functions, Armstrong Numbers and Palindromes]
61characters and words in a string; geometry.py and pointyShapeVolume[Practical 6: Counting the Characters and the Words in a String]
71two lists with a common member; a dictionary sorted by value[Practical 7: a Common Member, and a Dictionary Sorted by Value]
81area and circumference as a tuple; text and binary files; the last n lines[Practical 8: the Tuple Return, Area and Circumference]
91count a word in a file with a regular expression; extract every hyperlink[Practical 9: Counting a Word in a File with a Regular Expression]
101compare two dates; measure execution time; the calendar module[Practical 10: Comparing Two Dates in DD/MM/YYYY Form]
112array operations: insert at a position, delete from a position, linear search[Practical 1: Array Operations, Insert, Delete and Linear Search]
122a singly linked list: create, insert at beginning, end and position, delete[Practical 2: Building a Singly Linked List]
132a stack using an array; infix to postfix with a stack[Practical 3: a Stack over an Array]
142a queue using an array; simulate a customer service queue[Practical 4: a Queue over an Array]
152a binary search tree: create, insert, search[Practical 5: the Binary Search Tree, Create, Insert and Search]
162pre-order, in-order and post-order traversal of a binary tree[Practical 6: Tree Traversal, Pre-order, In-order and Post-order]
172a hash table with separate chaining; store and retrieve[Practical 7: a Hash Table with Separate Chaining]
182implement and compare bubble, insertion and selection sort[Practical 8: Bubble, Insertion and Selection Sort Compared]
192implement and compare linear search and binary search[Practical 9: Linear and Binary Search Compared]
202design a simple program that uses multiple data structures[Practical 10: the Combined Application]
munotes.in270

Keeping the Journal, and What Goes on the Page

Ten and ten, and 2.5 marks each is MU's 50.

Nine of the twenty are split across two chapters in this book, because MU's own bullets under them set two or three different programs. They are still one journal entry each. Put both programs under the one practical number.

What the signature is given for

The teacher signs when they have seen three things:

That you performed it, in the laboratory, in that session. Not copied from a friend at home.

That the output is yours. A teacher who has watched twenty students run the same program knows what its output looks like, and knows what it looks like when a machine behaved differently.

That the writeup is complete, with the aim, the program, the output and the conclusion.

So the way to get the signatures is simple and not negotiable: be at the practical, run it yourself, and bring the previous week's writeup with you.

What to do when a program will not run

Not nothing, and not a blank page. An entry that records a failure honestly is worth marks; a missing entry is worth none.

Practical No. 17                    Performed on: 05/10/2026

AIM
Hash Table. Write a program to implement a hash table with separate chaining
for collision handling, and to store and retrieve data from the hash table.

PROGRAM
    (as written, in full)

WHAT HAPPENED
    Traceback (most recent call last):
      File "practical17.py", line 34, in <module>
        table.put("Aarti", 78)
      File "practical17.py", line 21, in put
        chain = self.buckets[index]
    IndexError: list index out of range

WHAT I THINK IS WRONG
The bucket index came out larger than the number of buckets, so the hash was
not reduced by the modulo. Line 18 is `return string_hash(key)` and it should
be `return string_hash(key) % len(self.buckets)`.

CORRECTED PROGRAM AND OUTPUT
    (the fix, and the run)

                                    Teacher's signature: ____________

Write the error message in full. It is evidence that you ran the program, and a teacher reading a real traceback with a real diagnosis under it will sign that page.

Presenting a program on paper

Nine rules, and they cost nothing.

RuleWhy
indent exactly as the file does, four spaces a levelin Python the indentation IS the program
never break a line of code across two lines of the pagea broken line is a program that would not run
copy the variable names exactlya renamed variable is a different program
keep the blank lines between functionsthe structure is readable
write the output below the program, labelled OUTPUTan examiner should not have to hunt for it
copy the output including the spacinga table's alignment is part of the result
if the output is long, write the first and last few lines and say how many were left outhonest, and readable
number the pages and put the number in the indexso an entry can be found
write in ink and leave a marginit is a submitted document
munotes.in271

Keeping the Journal, and What Goes on the Page

The three things that cost marks and take five minutes to fix

No conclusion. Two sentences. What did the exercise show.

No cost line. Two or three lines of Big O with a reason. Most entries have none, so it stands out.

No output. An entry with a program and no output has not shown that the program runs, which is the one thing a practical examination is about.

A checklist to run down before submitting

[ ] the index has a row for every practical, with both dates
[ ] every entry has: number, dates, aim in MU's words, theory, program, output,
    conclusion
[ ] the aim is MU's wording, not mine
[ ] the output is what it printed, spacing included
[ ] the cost is stated where the exercise has one
[ ] the conclusion says what the exercise showed, not "executed successfully"
[ ] a failed practical has its error message and my diagnosis, not a blank page
[ ] every page is numbered and the number is in the index
[ ] every entry is signed
[ ] the certificate page at the front is signed
[ ] twenty entries, ten from each module

That last line is the one that matters most, and MU's own arithmetic is the reason: 2.5 marks per practical, twenty practicals, 50 marks, converted to 20.

Procedure

  1. Get the certificate page from your college and fill in the front matter on day one.
  2. Rule up the index with a row for all twenty practicals before you start, so a gap is visible.
  3. After each laboratory session, write the entry up that week, not later.
  4. Use the eight part shape for every entry: number and dates, aim, theory, program, output, conclusion,

cost, signature.

  1. Copy MU's aim word for word from the syllabus.
  2. Copy the output from the screen, spacing included.
  3. Add two or three lines of cost with a reason each.
  4. Write a conclusion that says what the exercise showed.
  5. Get it signed at the next session, and record the date in the index.
  6. Before the examination, run down the checklist and count the entries. Twenty.
munotes.in272

Keeping the Journal, and What Goes on the Page

Result

A journal with twenty signed entries, ten from each module, each carrying MU's own aim, the program, the real output, a cost line and a conclusion, and an index in which every row is dated and signed. That document earns part of the internal 20 marks under item 13, is examined again inside Q3 of the external slip, and is what permits you into the hall at all.

Where marks are lost

  • Fewer than twenty entries. Item 13's own arithmetic is 2.5 times 20.
  • No signatures, or signatures missing from some entries. A certified copy is compulsory to appear.
  • Everything written up in the last week. MU's wording is "the written practical of the previous

session".

  • No output. The program alone does not show that it runs.
  • Output written from memory rather than copied. It is the easiest thing for an examiner to catch.
  • "Executed successfully" as the conclusion. It says nothing.
  • No aim, or an aim in your own words instead of MU's.
  • Code broken across lines or indented differently from the file.
  • A blank page for a practical that failed, instead of the error and a diagnosis.
  • No index, or an index that does not match the pages.

For the journal

This chapter is the journal, so what it asks for is the journal itself. Two pages are worth adding before the first entry, and both take a few minutes.

The twenty row aim table above, so the index can be filled in ahead of time and any missing entry is obvious at a glance.

The eight part entry shape, copied onto the inside cover, so no entry is written without a conclusion and a cost line.

Quick revision

  • Twenty entries, ten per module. 2.5 marks each, total 50, converted to 20.
  • The writeup is due at the next session, not at the end of the semester.
  • A certified, meaning signed, copy is compulsory to appear for the practical examination.
  • The journal is marked twice: in the internal 20 under item 13, and inside Q3 of the external 30.
  • One entry: number and dates, aim in MU's words, theory, program, output, conclusion, cost,

signature.

  • The output is what it printed, with the spacing.
  • A cost line of two or three lines of Big O with a reason makes an entry stand out.
  • A conclusion says what the exercise showed, never "executed successfully".
  • A practical that failed gets the error message in full and your own diagnosis, not a blank page.
  • Nine of the twenty exercises are two or three programs each; they are still one entry each.
  • Number every page and match the index.
munotes.in273

Keeping the Journal, and What Goes on the Page

Questions you should be able to answer

1. How many entries should the journal have, and how do you know? Twenty. Item 13 awards 2.5 marks for each practical performance and writeup to a total of 50, and 2.5 multiplied by 20 is exactly 50.

2. When is a writeup due? At the next practical session. MU's wording is "submit the written practical of the previous session".

3. What does "certified copy" mean, and why does it matter? Signed by your teacher. Item 14 makes a certified copy compulsory to appear for the practical examination, so an unsigned journal keeps you out of the hall.

4. How many times is the journal marked? Twice: within the internal 20 marks under item 13, and inside Q3 of the external 30 mark slip.

5. What are the parts of one entry? Practical number and both dates, the aim in MU's words, brief theory, the program, the real output, a conclusion, a cost line where there is one, and the teacher's signature.

6. Why copy MU's aim word for word? Because it is her specification of the exercise, and an aim in your own words invites the question of whether you read hers.

7. What is wrong with "the program was executed successfully" as a conclusion? It says nothing about what the exercise showed. A conclusion should state the finding: what the structure cost, or what the comparison proved.

8. What do you write for a practical that would not run? The aim, the program as written, the complete error message, and your own diagnosis of what is wrong, followed by the corrected program and its output if you fixed it later.

9. MU sets ten exercises per module but some of them have three parts. How many entries is that? Ten per module still. Each numbered exercise is one entry, with all its programs under that number.

10. What is the single most avoidable way to lose marks on this paper? An unsigned journal, which prevents you from appearing at all, followed by having fewer than twenty entries.

Contents This chapter on its own page

munotes.in274

Chapter Thirty-Seven

A Worked Practical Slip: Q1, Q2 and Q3

Syllabus topic MU's item 14 for Major Practical 3, "Format of Question Paper: Duration 2 hours. Q1. From Module 1 13 marks, Q2. From Module 2 12marks, Q3. Journal and Viva 05 marks"

Aim

To work a complete practical slip in MU's own shape, and to know how to spend the two hours.

The slip you will be handed

MU prints the format and not the questions, so the wording below is the shape her item 14 sets, with questions of the kind her twenty exercises support.

                University of Mumbai
    B.Sc. (Information Technology), Semester III
              Major Practical 3
    Duration: 2 hours                Marks: 30

    Q1. Attempt any ONE from Module 1                      [13]
        (a) Write a program to accept an SGPI from the user and print the
            corresponding grade.
        (b) Write a program to perform slicing, basic and advanced indexing
            on NumPy arrays.
        (c) Write a program to count the occurrences of a specific word in a
            file using regular expressions.
        (d) Write a program that compares two dates in DD/MM/YYYY format and
            prints which one is earlier.

    Q2. Attempt any ONE from Module 2                      [12]
        (a) Write a program to convert an infix expression to postfix
            notation using a stack.
        (b) Write a program to create a binary search tree, insert nodes
            into it, and search it for a node.
        (c) Write a program to implement a hash table with separate chaining
            for collision handling.
        (d) Write programs to implement and compare bubble sort, insertion
            sort and selection sort.

    Q3. Journal and Viva                                    [5]

    Note: a certified copy of the journal is compulsory.

Every one of those eight alternatives is answered below, in full, because on the day you get one of them and you cannot choose which.

How to spend the two hours

TimeWhat you are doing
0 to 5 minutesread both questions, pick your alternatives, and write the aim for each
5 to 45Q1: write it, run it, get the output
45 to 85Q2: write it, run it, get the output
85 to 105write both up properly: aim, program, output, conclusion
105 to 120the viva, and whatever is not finished

Four things about that plan are worth following.

Write both aims in the first five minutes. They are marks you cannot lose later, and they force you to read the question rather than the first line of it.

Start with Q1. It carries the extra mark, and Module 1 programs are shorter.

Save 20 minutes for the writeup. A program that runs and is not written up loses the aim, the output and the conclusion.

If a program will not run with 20 minutes left, stop debugging. Write down what you have, the error message, and one sentence on what you think is wrong. That is worth several marks; a blank page is worth none.

Q1 (a): the SGPI grade

Aim: to accept an SGPI from the user and print the corresponding grade.

munotes.in275

A Worked Practical Slip: Q1, Q2 and Q3

sgpi = float(input("Enter your SGPI: "))

if sgpi > 10:
    grade = "invalid, the scale ends at 10.00"
elif sgpi >= 9.00:
    grade = "O"
elif sgpi >= 8.00:
    grade = "A+"
elif sgpi >= 7.00:
    grade = "A"
elif sgpi >= 6.00:
    grade = "B+"
elif sgpi >= 5.50:
    grade = "B"
elif sgpi >= 5.00:
    grade = "C"
elif sgpi >= 4.00:
    grade = "P"
else:
    grade = "F"

print(f"SGPI {sgpi:.2f} gives grade {grade}")
7.85
Enter your SGPI: SGPI 7.85 gives grade A

Conclusion: float is needed because an SGPI has decimals, and the bands are written downwards so each one needs only its own floor: once >= 9.00 has failed, >= 8.00 can only mean 8.00 to 8.99.

If there is time, add the loop that proves every band, because it turns one run into eight:

def grade_for(sgpi):
    if sgpi > 10:
        return "invalid"
    for floor, grade in [(9.00, "O"), (8.00, "A+"), (7.00, "A"), (6.00, "B+"),
                         (5.50, "B"), (5.00, "C"), (4.00, "P")]:
        if sgpi >= floor:
            return grade
    return "F"


for value in [10.00, 9.00, 8.99, 8.00, 7.99, 7.00, 6.99, 6.00,
              5.99, 5.50, 5.49, 5.00, 4.99, 4.00, 3.99]:
    print(f"{value:>6.2f} -> {grade_for(value)}")
 10.00 -> O
  9.00 -> O
  8.99 -> A+
  8.00 -> A+
  7.99 -> A
  7.00 -> A
  6.99 -> B+
  6.00 -> B+
  5.99 -> B
  5.50 -> B
  5.49 -> C
  5.00 -> C
  4.99 -> P
  4.00 -> P
  3.99 -> F

Q1 (b): NumPy slicing and indexing

Aim: to perform slicing, and basic and advanced indexing, on NumPy arrays.

import numpy as np

m = np.arange(1, 25).reshape(4, 6)
print("the array:")
print(m)

print()
print("BASIC indexing, which returns a VIEW:")
print("  m[1]        ", m[1])
print("  m[:, 2]     ", m[:, 2])
print("  m[1, 3]     ", m[1, 3])
print("  m[0:2, 0:3] ")
print(m[0:2, 0:3])
print("  m[::2, ::3] ")
print(m[::2, ::3])

print()
print("ADVANCED indexing, which returns a COPY:")
print("  m[[0, 2], :] whole rows, chosen")
print(m[[0, 2], :])
print("  m[[0, 1, 2], [3, 2, 0]] the two lists ZIPPED:", m[[0, 1, 2], [3, 2, 0]])
print("  m[m > 18] a boolean mask            :", m[m > 18])

print()
print("the difference, proved:")
view = m[1:3, 1:3]
copy = m[[1, 2], :][:, [1, 2]]
print("  m[1:3, 1:3]   shares memory with m?", np.shares_memory(view, m))
print("  the advanced one shares memory?    ", np.shares_memory(copy, m))
view[0, 0] = 999
print("  after writing 999 into the slice, m[1, 1] is", m[1, 1])
the array:
[[ 1  2  3  4  5  6]
 [ 7  8  9 10 11 12]
 [13 14 15 16 17 18]
 [19 20 21 22 23 24]]

BASIC indexing, which returns a VIEW:
  m[1]         [ 7  8  9 10 11 12]
  m[:, 2]      [ 3  9 15 21]
  m[1, 3]      10
  m[0:2, 0:3]
[[1 2 3]
 [7 8 9]]
  m[::2, ::3]
[[ 1  4]
 [13 16]]

ADVANCED indexing, which returns a COPY:
  m[[0, 2], :] whole rows, chosen
[[ 1  2  3  4  5  6]
 [13 14 15 16 17 18]]
  m[[0, 1, 2], [3, 2, 0]] the two lists ZIPPED: [ 4  9 13]
  m[m > 18] a boolean mask            : [19 20 21 22 23 24]

the difference, proved:
  m[1:3, 1:3]   shares memory with m? True
  the advanced one shares memory?     False
  after writing 999 into the slice, m[1, 1] is 999
munotes.in276

A Worked Practical Slip: Q1, Q2 and Q3

Conclusion: basic indexing uses integers and slices and returns a view onto the same memory, so writing to it changes the original. Advanced indexing uses an array of positions or a boolean mask and returns a copy, so writing to it does not. np.shares_memory is the test.

Q1 (c): count a word in a file with a regular expression

Aim: to count the occurrences of a specific word in a file using regular expressions.

Notes help students. A note is short; notes are longer.
The notebook on the table holds no notes at all.
NOTE: notes are free. Note the word note appears often.
import re

word = "note"
pattern = re.compile(r"\b" + re.escape(word) + r"\b", re.IGNORECASE)

total = 0
with open("report.txt") as handle:
    for number, line in enumerate(handle, 1):
        found = pattern.findall(line)
        if found:
            total += len(found)
            print(f"  line {number}: {len(found)} -> {found}")

print(f"the word {word!r} appears {total} time(s)")

print()
print("without the word boundary, for comparison:")
with open("report.txt") as handle:
    text = handle.read()
print("  str.count      :", text.lower().count(word))
print("  re without \\b  :", len(re.findall(word, text, re.IGNORECASE)))
print("  re with \\b     :", len(pattern.findall(text)))
  line 1: 1 -> ['note']
  line 3: 3 -> ['NOTE', 'Note', 'note']
the word 'note' appears 4 time(s)

without the word boundary, for comparison:
  str.count      : 9
  re without \b  : 9
  re with \b     : 4

Conclusion: \b matches a word boundary, so \bnote\b counts the word and not the letters. Without it the count includes "notes", "notebook" and "denote". The pattern must be a raw string, because in an ordinary Python string \b is the backspace character.

Q1 (d): compare two dates in DD/MM/YYYY form

Aim: to compare two dates given in DD/MM/YYYY form and print which one is earlier.

from datetime import datetime


def parse(text):
    """A date from a DD/MM/YYYY string."""
    return datetime.strptime(text.strip(), "%d/%m/%Y").date()


first_text = input("Enter the first date  (DD/MM/YYYY): ")
second_text = input("Enter the second date (DD/MM/YYYY): ")

first = parse(first_text)
second = parse(second_text)

if first < second:
    print(f"{first_text} is EARLIER than {second_text}")
elif second < first:
    print(f"{second_text} is EARLIER than {first_text}")
else:
    print("the two dates are the same")

print(f"they are {abs((first - second).days)} day(s) apart")

print()
print("and why the strings cannot simply be compared:")
print(f"  {first_text!r} < {second_text!r} as strings is {first_text < second_text}")
print("  because a string comparison reaches the DAY first")
munotes.in277

A Worked Practical Slip: Q1, Q2 and Q3

05/10/2026
12/03/2020
Enter the first date  (DD/MM/YYYY): Enter the second date (DD/MM/YYYY): 12/03/2020 is EARLIER than 05/10/2026
they are 2398 day(s) apart

and why the strings cannot simply be compared:
  '05/10/2026' < '12/03/2020' as strings is True
  because a string comparison reaches the DAY first

Conclusion: a DD/MM/YYYY string cannot be compared as a string, because the comparison reaches the day before the year. strptime with the format "%d/%m/%Y" turns each string into a date, after which the ordinary comparison operators are correct. Subtracting two dates gives a timedelta whose .days is the gap.

Q2 (a): infix to postfix with a stack

Aim: to convert an infix expression to postfix notation using a stack.

PRECEDENCE = {"+": 1, "-": 1, "*": 2, "/": 2, "%": 2, "^": 3}
RIGHT = {"^"}


def pops_first(on_stack, arriving):
    if on_stack == "(":
        return False
    if PRECEDENCE[on_stack] > PRECEDENCE[arriving]:
        return True
    if PRECEDENCE[on_stack] < PRECEDENCE[arriving]:
        return False
    return arriving not in RIGHT


def to_postfix(text, trace=False):
    output, stack = [], []
    if trace:
        print(f"  {'token':<7} {'stack':<10} output")
    for token in text.split():
        if token == "(":
            stack.append(token)
        elif token == ")":
            while stack and stack[-1] != "(":
                output.append(stack.pop())
            if not stack:
                raise ValueError("a ')' with no matching '('")
            stack.pop()
        elif token in PRECEDENCE:
            while stack and pops_first(stack[-1], token):
                output.append(stack.pop())
            stack.append(token)
        else:
            output.append(token)
        if trace:
            print(f"  {token:<7} {' '.join(stack):<10} {' '.join(output)}")
    while stack:
        top = stack.pop()
        if top == "(":
            raise ValueError("a '(' that is never closed")
        output.append(top)
    if trace:
        print(f"  {'end':<7} {'':<10} {' '.join(output)}")
    return " ".join(output)


def evaluate(postfix):
    stack = []
    for token in postfix.split():
        if token in PRECEDENCE:
            right = stack.pop()
            left = stack.pop()
            stack.append({"+": left + right, "-": left - right,
                          "*": left * right, "/": left / right,
                          "%": left % right, "^": left ** right}[token])
        else:
            stack.append(float(token))
    return stack.pop()


print("converting 'a + b * c - d', traced:")
print("  postfix:", to_postfix("a + b * c - d", trace=True))
print()
print("more conversions, each one EVALUATED as the check:")
for text in ["2 + 3 * 4", "( 2 + 3 ) * 4", "2 ^ 3 ^ 2", "10 - 4 - 3",
             "100 / ( 5 / 2 )"]:
    postfix = to_postfix(text)
    print(f"  {text:<18} -> {postfix:<16} = {evaluate(postfix)}")
converting 'a + b * c - d', traced:
  token   stack      output
  a                  a
  +       +          a
  b       +          a b
  *       + *        a b
  c       + *        a b c
  -       -          a b c * +
  d       -          a b c * + d
  end                a b c * + d -
  postfix: a b c * + d -

more conversions, each one EVALUATED as the check:
  2 + 3 * 4          -> 2 3 4 * +        = 14.0
  ( 2 + 3 ) * 4      -> 2 3 + 4 *        = 20.0
  2 ^ 3 ^ 2          -> 2 3 2 ^ ^        = 512.0
  10 - 4 - 3         -> 10 4 - 3 -       = 3.0
  100 / ( 5 / 2 )    -> 100 5 2 / /      = 40.0
munotes.in278

A Worked Practical Slip: Q1, Q2 and Q3

Conclusion: the stack holds operators. An operand goes straight to the output; an operator pops everything on the stack that should come out first and is then pushed; a closing bracket pops to its opening bracket and both are discarded. "Should come out first" means higher precedence, or equal precedence when the arriving operator is left associative, which is why ^ behaves differently. The conversion is proved by evaluating the postfix.

Q2 (b): the binary search tree

Aim: to create a binary search tree, insert nodes into it, and search it for a node.

class Node:
    def __init__(self, key):
        self.key = key
        self.left = None
        self.right = None


class BinarySearchTree:
    def __init__(self):
        self.root = None
        self.count = 0

    def insert(self, key):
        """Left subtree smaller, right subtree larger. A duplicate is refused."""
        if self.root is None:
            self.root = Node(key)
            self.count += 1
            return True
        here = self.root
        while True:
            if key == here.key:
                return False
            if key < here.key:
                if here.left is None:
                    here.left = Node(key)
                    self.count += 1
                    return True
                here = here.left
            else:
                if here.right is None:
                    here.right = Node(key)
                    self.count += 1
                    return True
                here = here.right

    def search(self, key):
        """(found, comparisons, path). One comparison discards a whole subtree."""
        here, comparisons, path = self.root, 0, []
        while here is not None:
            comparisons += 1
            path.append(here.key)
            if key == here.key:
                return True, comparisons, path
            here = here.left if key < here.key else here.right
        return False, comparisons, path

    def in_order(self):
        out = []

        def walk(node):
            if node is not None:
                walk(node.left)
                out.append(node.key)
                walk(node.right)

        walk(self.root)
        return out

    def height(self):
        def deepest(node):
            return -1 if node is None else 1 + max(deepest(node.left),
                                                   deepest(node.right))

        return deepest(self.root)

    def draw(self):
        lines = []

        def down(node, depth):
            if node is None:
                return
            down(node.right, depth + 1)
            lines.append("    " * depth + str(node.key))
            down(node.left, depth + 1)

        down(self.root, 0)
        return "\n".join(lines)


tree = BinarySearchTree()
for key in [50, 30, 70, 20, 40, 60, 80]:
    tree.insert(key)

print("the tree, root at the left:")
print(tree.draw())
print()
print("in-order walk (must be sorted):", tree.in_order())
print("sorted?", tree.in_order() == sorted(tree.in_order()))
print("nodes", tree.count, " height", tree.height())
print()
print("searching:")
for key in [50, 20, 80, 45]:
    found, comparisons, path = tree.search(key)
    print(f"  {key:>3}  {'found' if found else 'NOT FOUND':<10} "
          f"{comparisons} comparison(s), path {' -> '.join(str(k) for k in path)}")
print()
print("a duplicate:", tree.insert(30), "and the count is still", tree.count)
munotes.in279

A Worked Practical Slip: Q1, Q2 and Q3

the tree, root at the left:
        80
    70
        60
50
        40
    30
        20

in-order walk (must be sorted): [20, 30, 40, 50, 60, 70, 80]
sorted? True
nodes 7  height 2

searching:
   50  found      1 comparison(s), path 50
   20  found      3 comparison(s), path 50 -> 30 -> 20
   80  found      3 comparison(s), path 50 -> 70 -> 80
   45  NOT FOUND  3 comparison(s), path 50 -> 30 -> 40

a duplicate: False and the count is still 7

Conclusion: the binary search tree rule is that every key in a node's left subtree is smaller and every key in its right subtree is larger. One comparison at a node therefore discards a whole subtree, so a search costs at most the height plus one. The in-order walk comes out sorted, which is the proof that the tree really is a binary search tree.

Q2 (c): the hash table with separate chaining

Aim: to implement a hash table with separate chaining for collision handling.

def string_hash(key, base=31):
    """A polynomial hash, written out so the bucket numbers are reproducible."""
    total = 0
    for character in str(key):
        total = total * base + ord(character)
    return total


class HashTable:
    def __init__(self, buckets=7):
        self.buckets = [[] for _ in range(buckets)]
        self.count = 0

    def bucket_of(self, key):
        return string_hash(key) % len(self.buckets)

    def put(self, key, value):
        """Store or update. Returns (bucket, collided)."""
        index = self.bucket_of(key)
        chain = self.buckets[index]
        collided = len(chain) > 0
        for position, (existing, _) in enumerate(chain):
            if existing == key:
                chain[position] = (key, value)          # an UPDATE
                return index, False
        chain.append((key, value))
        self.count += 1
        return index, collided

    def get(self, key):
        """(value, comparisons). Searches one chain only."""
        comparisons = 0
        for existing, value in self.buckets[self.bucket_of(key)]:
            comparisons += 1
            if existing == key:
                return value, comparisons
        return None, comparisons

    def show(self):
        for index, chain in enumerate(self.buckets):
            items = " -> ".join(f"({k}, {v})" for k, v in chain)
            print(f"  {index}: {items}")

    def load_factor(self):
        return self.count / len(self.buckets)


table = HashTable(buckets=7)
for name, mark in [("Aarti", 78), ("Divya", 55), ("Chetan", 90),
                   ("Eshan", 67), ("Farhan", 41), ("Gauri", 83)]:
    index, collided = table.put(name, mark)
    print(f"put {name:<8} -> bucket {index}"
          f"{'   COLLISION, chained' if collided else ''}")

print()
print("the table:")
table.show()
print()
print(f"items {table.count}, buckets {len(table.buckets)}, "
      f"load factor {table.load_factor():.2f}")
print("chain lengths", [len(c) for c in table.buckets])
print()
print("retrieving:")
for name in ["Aarti", "Gauri", "Nobody"]:
    value, comparisons = table.get(name)
    print(f"  {name:<8} -> {str(value):<8} after {comparisons} comparison(s)")
print()
table.put("Aarti", 91)
print("stored Aarti again:", table.get("Aarti")[0],
      "and the count is still", table.count)
put Aarti    -> bucket 4
put Divya    -> bucket 2
put Chetan   -> bucket 2   COLLISION, chained
put Eshan    -> bucket 0
put Farhan   -> bucket 1
put Gauri    -> bucket 0   COLLISION, chained

the table:
  0: (Eshan, 67) -> (Gauri, 83)
  1: (Farhan, 41)
  2: (Divya, 55) -> (Chetan, 90)
  3:
  4: (Aarti, 78)
  5:
  6:

items 6, buckets 7, load factor 0.86
chain lengths [2, 1, 2, 0, 1, 0, 0]

retrieving:
  Aarti    -> 78       after 1 comparison(s)
  Gauri    -> 83       after 2 comparison(s)
  Nobody   -> None     after 2 comparison(s)

stored Aarti again: 91 and the count is still 6
munotes.in280

A Worked Practical Slip: Q1, Q2 and Q3

Conclusion: the hash function turns the key into a number, the modulo turns it into a bucket index, and separate chaining puts every key that lands in the same bucket into a list there. A lookup searches only its own bucket's chain, so it costs the length of that chain, which the load factor keeps near one. put must search the chain first, or the same key would appear twice.

Q2 (d): the three sorts compared

Aim: to implement and compare bubble sort, insertion sort and selection sort.

def bubble_sort(data):
    data = list(data)
    comparisons = movements = passes = 0
    n = len(data)
    for i in range(n - 1):
        swapped = False
        passes += 1
        for j in range(n - 1 - i):
            comparisons += 1
            if data[j] > data[j + 1]:
                data[j], data[j + 1] = data[j + 1], data[j]
                movements += 2
                swapped = True
        if not swapped:
            break
    return data, comparisons, movements, passes


def insertion_sort(data):
    data = list(data)
    comparisons = movements = passes = 0
    for i in range(1, len(data)):
        passes += 1
        value = data[i]
        j = i - 1
        while j >= 0:
            comparisons += 1
            if data[j] <= value:
                break
            data[j + 1] = data[j]
            movements += 1
            j -= 1
        data[j + 1] = value
        movements += 1
    return data, comparisons, movements, passes


def selection_sort(data):
    data = list(data)
    comparisons = movements = passes = 0
    n = len(data)
    for i in range(n - 1):
        passes += 1
        smallest = i
        for j in range(i + 1, n):
            comparisons += 1
            if data[j] < data[smallest]:
                smallest = j
        if smallest != i:
            data[i], data[smallest] = data[smallest], data[i]
            movements += 2
    return data, comparisons, movements, passes


datasets = {
    "already sorted": [10, 20, 30, 40, 50, 60, 70, 80],
    "reverse sorted": [80, 70, 60, 50, 40, 30, 20, 10],
    "random":         [30, 10, 80, 50, 20, 70, 40, 60],
}

n = 8
print(f"n = {n}, and n(n-1)/2 = {n * (n - 1) // 2}")
for label, data in datasets.items():
    print()
    print(f"{label}: {data}")
    print(f"  {'sort':<11} {'comparisons':>12} {'movements':>11} {'passes':>8}")
    for name, function in [("bubble", bubble_sort), ("insertion", insertion_sort),
                           ("selection", selection_sort)]:
        result, comparisons, movements, passes = function(data)
        assert result == sorted(data), f"{name} did not sort"
        print(f"  {name:<11} {comparisons:>12} {movements:>11} {passes:>8}")

print()
print("all results checked against sorted(), so the counts are from correct sorts")
munotes.in281

A Worked Practical Slip: Q1, Q2 and Q3

n = 8, and n(n-1)/2 = 28

already sorted: [10, 20, 30, 40, 50, 60, 70, 80]
  sort         comparisons   movements   passes
  bubble                 7           0        1
  insertion              7           7        7
  selection             28           0        7

reverse sorted: [80, 70, 60, 50, 40, 30, 20, 10]
  sort         comparisons   movements   passes
  bubble                28          56        7
  insertion             28          35        7
  selection             28           8        7

random: [30, 10, 80, 50, 20, 70, 40, 60]
  sort         comparisons   movements   passes
  bubble                22          22        4
  insertion             17          18        7
  selection             28          14        7

all results checked against sorted(), so the counts are from correct sorts

Conclusion: all three are O(n squared) in the worst case and all three made n(n-1)/2 comparisons on reversed data. Bubble sort with its swapped early exit and insertion sort both drop to n-1 comparisons on already sorted data; selection sort does not, because it searches for the minimum whatever the data looks like. Selection sort makes the fewest movements, at most one swap per pass. Bubble and insertion are stable; selection is not.

The writeup, and the twenty minutes it needs

For each question, on the answer sheet:

Q1 (c)

AIM
Write a program to count the occurrences of a specific word in a file using
regular expressions.

PROGRAM
    (the listing, indented as typed)

OUTPUT
    line 1: 1 -> ['note']
    line 3: 3 -> ['NOTE', 'Note', 'note']
    the word 'note' appears 4 time(s)

CONCLUSION
\b matches a word boundary, so \bnote\b counts the word and not the letters:
without it the count also includes "notes", "notebook" and "denote". The
pattern must be a raw string, because in an ordinary Python string \b is the
backspace character.

Four sections, and the fourth is the one most candidates leave out.

If it will not run

With twenty minutes left, stop debugging and write this instead:

Q2 (b)

AIM
(as printed on the slip)

PROGRAM AS WRITTEN
    (all of it)

WHAT HAPPENED
    AttributeError: 'NoneType' object has no attribute 'left'

WHAT I THINK IS WRONG
The insert method walks to here.left without first checking whether here.left
is None, so on reaching a leaf it dereferences None. The check belongs before
the walk: if here.left is None, attach the new node there.

That page is worth several marks. A blank page is worth none, and an examiner can tell the difference between a candidate who did not know and one who ran out of time with a correct diagnosis in hand.

Nine things to do on the day

1take the signed journal. Without it you do not sit the paper
2read both questions before writing anything
3write both aims in the first five minutes
4start with Q1, which carries the extra mark
5run the program and copy the output, spacing included
6keep 20 minutes for the writeup
7write a conclusion for each question, two sentences
8add the cost in Big O if the question has one
9if it will not run, write the error and your diagnosis
munotes.in282

A Worked Practical Slip: Q1, Q2 and Q3

Procedure

  1. Practise against the clock: one Module 1 question and one Module 2 question, in 90 minutes, with the

writeup.

  1. Work all eight alternatives above at least once, because you do not choose which one appears.
  2. For each, write the aim, the program, the output and a two sentence conclusion.
  3. Type the programs from memory rather than copying them, and note which lines you get wrong.
  4. For Q2, be able to write the stack, the tree, the hash table and the three sorts without looking.
  5. Include the assert against sorted() in the sorting answer, and np.shares_memory in the NumPy one.
  6. Practise the failure writeup: take a program, break one line, and write the diagnosis page in five

minutes.

Result

Eight complete answers, four from each module, each with its aim, its program, its real output and a conclusion. The SGPI ladder graded every one of MU's bands. The NumPy answer proved the view against copy distinction with np.shares_memory. The word count gave the right answer with \b and a larger wrong one without it. The date comparison showed the string comparison giving the wrong answer and the parsed comparison the right one. The infix conversion was checked by evaluating its own output. The binary search tree's in-order walk was sorted. The hash table's collisions were chained and a repeated key updated rather than duplicated. The three sorts were checked against sorted() and their counts tabulated across three datasets.

Where marks are lost

  • Arriving without a signed journal. You do not sit the paper.
  • No aim, which is the cheapest mark on the sheet.
  • No output. The program alone does not show that it runs.
  • No conclusion. Two sentences, and most candidates leave them out.
  • Starting with Q2 and running out of time on the question worth more.
  • Debugging to the last minute instead of writing up what works.
  • A blank page for a program that would not run, rather than the error and a diagnosis.
  • Output copied from the journal rather than from the screen in front of you.
  • Treating Q3 as nothing. Five marks out of thirty is a sixth of the paper.

For the journal

This chapter is the examination, so nothing here is a numbered entry. Two pages are worth keeping with the journal anyway.

munotes.in283

A Worked Practical Slip: Q1, Q2 and Q3

The time plan: the five row table of how the two hours are spent, on the inside back cover.

The four section answer shape: AIM, PROGRAM, OUTPUT, CONCLUSION, with one worked example under it, so that on the day the shape is automatic.

Quick revision

  • The slip: Q1 Module 1 for 13, Q2 Module 2 for 12, Q3 journal and viva for 5. Two hours.
  • A certified journal is compulsory to appear.
  • Read both questions first. Write both aims in the first five minutes.
  • Start with Q1: one more mark, shorter programs.
  • Keep 20 minutes for the writeup. Four sections: AIM, PROGRAM, OUTPUT, CONCLUSION.
  • Copy the output from the screen, spacing included.
  • Add a cost line in Big O where the question has one.
  • If a program will not run with 20 minutes left: stop, and write the error and your diagnosis.
  • Practise all eight kinds of question, because you do not choose which appears.

Questions you should be able to answer

1. What is on the slip, and what is each part worth? Q1 from Module 1 for 13 marks, Q2 from Module 2 for 12 marks, and Q3 on the journal and the viva for 5 marks, over two hours.

2. Which question do you attempt first, and why? Q1: it is worth one mark more and Module 1's programs are shorter, so it is the safer half to finish.

3. What do you write in the first five minutes? The aim for each question, copied from the slip. They are marks that cannot be lost later and they force you to read the whole question.

4. How much time should the writeup have? About twenty minutes. A program that runs and is not written up loses the aim, the output and the conclusion.

5. What are the four sections of an answer? AIM, PROGRAM, OUTPUT, CONCLUSION.

6. What do you do if a program will not run with twenty minutes left? Stop debugging. Write the program as you have it, the complete error message, and one or two sentences on what you think is wrong.

7. Why is copying the output from the screen important? Because output written from what the program should do is the easiest error for an examiner to catch, and the output is the evidence that it ran.

8. What single thing keeps you out of the examination hall? An uncertified journal. Item 14 makes a certified copy compulsory to appear.

Contents This chapter on its own page

munotes.in284

Chapter Thirty-Eight

The Viva: the Questions Asked at the Table

Syllabus topic MU's item 14 for Major Practical 3, "Q3. Journal and Viva 05 marks"

Aim

To be able to answer for your own programs at the table.

What the viva actually is

Q3 is journal and viva for five marks, and it happens with your journal open in front of you and the examiner turning its pages. It is not a theory examination. Almost every question is one of four kinds, and knowing the four is most of the preparation.

The kindWhat it sounds likeWhat it is testing
1. what does it do"explain this program"that you wrote it
2. why that way"why a queue and not a stack?"that you chose rather than copied
3. what if"what happens if the list is empty?"that you thought about the edges
4. what does it cost"what is the complexity?"that you know Big O and can justify it

The third and the fourth are where marks are won, because most candidates prepare only for the first.

The rule for answering

Name the thing, then give the reason, then give the number if there is one. Three clauses, twenty seconds.

WeakStrong
"a hash table is fast""a hash table, because we look up by roll number on every transaction, and that is O(1) on average"
"it is O(n squared)""O(n squared): counted here, 780 comparisons for 40 items, which is n(n-1)/2"
"bubble sort is slow""bubble sort is O(n squared) in the worst case but O(n) on sorted data, because of the swapped flag"

And two things not to do. Do not guess a number; say "I did not measure that" and then say what you did measure. Do not recite a definition you cannot apply; an examiner who hears "amortised constant time" will ask what amortised means.

Module 1, the ten exercises

Practical 1: input, output, even or odd, the SGPI ladder

What type does input() return? A string, always, even when digits are typed.

Why does age + 1 fail then? A string cannot be added to a number. int(input(...)) converts it.

Why not type the year into the program? It would be wrong from the next 1st of January. date.today().year asks the machine.

A student is 19 in 2026. Show the sum. 100 - 19 = 81 years to go, and 2026 + 81 = 2107.

When is that one year out? When this year's birthday has not happened yet. Reading the date of birth removes the assumption.

How do you test for even? n % 2 == 0. Zero is even.

-7 % 2 in Python? 1. Python's remainder takes the sign of the divisor, unlike C's -1.

Why are the grade bands written downwards? So each rung needs only its own floor: once >= 9.00 has failed, >= 8.00 can only mean 8.00 to 8.99.

munotes.in285

The Viva: the Questions Asked at the Table

What if you write them upwards? >= 4.00 is true for almost everybody, so almost everybody gets a P.

Why float and not int for the SGPI? An SGPI is 7.85. int("7.85") raises ValueError, and truncating moves a mark into the wrong band.

Practical 2: Fibonacci and the digit sum

Write the series in one loop. a, b = 0, 1, then repeat: print a, and a, b = b, a + b.

Why does that need no temporary variable? Python builds the whole right hand side first, as a tuple, so a + b uses the old a.

What if you write a = b then b = a + b? You get the powers of two. The old a has already been lost.

Why is the recursive version slow? It recomputes the same terms. Counted here: fib(25) took 242,785 calls, and fib(30) took 2,692,537.

Which two operators take a number apart? % 10 gives the last digit, // 10 removes it.

Why // and not /? / gives a float, so the number never becomes exactly 0.

What does your program give for -4729? 22, because abs() is applied first. Without it the answer would be 0.

Practical 3 and 4: arrays, NumPy, indexing and attributes

Name the three things called an array in Python. A list, an array.array, and a NumPy ndarray. MU's exercises mean NumPy.

What does a[1:4] give? Items 1, 2 and 3. The stop is excluded.

Reverse an array in one expression. a[::-1].

Slicing a list against slicing a NumPy array? A list slice is a copy; a NumPy slice is a view onto the same memory.

How do you tell a view from a copy? np.shares_memory(x, a). base is not enough: on a mixed subscript its base is a temporary array, so it is not None although nothing is shared.

What is basic and what is advanced indexing? Basic is integers and slices and returns a view. Advanced is an integer array or a boolean mask and returns a copy.

m[[0, 1], [2, 3]] gives what? Two elements, m[0,2] and m[1,3]. The lists are zipped, not crossed. np.ix_ crosses them.

Why not a[a > 10 and a < 45]? and needs one True or False and an array of several is ambiguous. Use (a > 10) & (a < 45).

axis=0 or axis=1? axis=0 down the columns, axis=1 along the rows. The axis you name is the one that disappears.

What is nbytes? size × itemsize, always.

ravel against flatten? ravel returns a view where it can; flatten always returns a copy.

munotes.in286

The Viva: the Questions Asked at the Table

Does astype(int) round? No, it truncates towards zero: 1.7 becomes 1 and -1.7 becomes -1.

Why is there no bracket on m.shape? It is an attribute, not a method. m.shape() tries to call a tuple.

Practical 5: functions, Armstrong, palindrome, recursion, lambda

Argument or parameter? The parameter is the name in the def; the argument is the value passed in.

What does a function with no return give back? None, which cannot be used in an if.

Define an Armstrong number. A number equal to the sum of its digits each raised to the power of the number of digits.

Is 8208 one? Yes: 8^4 + 2^4 + 0^4 + 8^4 = 4096 + 16 + 0 + 4096 = 8208.

Is 7 one? Yes. One digit, and 7^1 = 7. Every single digit number qualifies.

How many are there below 10000? 17, counted by the program: the ten single digits, then 153, 370, 371, 407, then 1634, 8208, 9474.

Test for a palindrome in one line. text == text[::-1], after cleaning if the text has case or punctuation.

Why clean it first? A capital, a space or a comma is a character like any other, so "Madam" and any sentence would fail.

What two parts does a recursive function have? A base case and a recursive step that makes the problem smaller.

What is the base case of the factorial? n equal to 0, returning 1.

What is the recursion limit here? About 1000, and exceeding it raises RecursionError.

Why does this factorial not overflow at 21! as C does? Python integers have no fixed width.

What can a lambda not contain? Any statement: no if statement, no loop, no assignment, no return. A conditional expression is allowed.

Is "Practical".startswith("p") True? No. startswith is case sensitive.

Practical 6: string counts, and the geometry module

How many characters in "To be or not to be"? 18 with the spaces, 13 without. Both are correct answers to different questions, so say which.

Why split() and not split(' ')? split() treats any run of whitespace as one separator and never produces an empty word. On one test string split() gave 3 words and split(' ') gave 10 pieces, 7 of them empty.

What is wrong with count(" ") + 1? Double spaces, a leading space and an empty string all give the wrong answer.

What does text.replace(" ", "") do to text? Nothing. Strings cannot be changed; it returns a new one.

What is a module? A .py file whose names another file can import.

munotes.in287

The Viva: the Questions Asked at the Table

Why did MU split practical 6(b) into two files? The base areas are useful on their own, and separating them is the habit the exercise teaches.

State the volume rule. One third × the base area × the height, for a pyramid and a cone alike.

Compute the pyramid of edge 3 and height 9. Base 3 × 3 = 9, so 9 × 9 / 3 = 27.

Which is bigger for the same x and y, pyramid or cone? The cone, by exactly pi, because a circle of radius x has area pi x squared against the square's x squared. The ratio of pyramid to cone is 1 over pi, 0.3183.

What does if __name__ == "__main__": do? Its block runs only when the file is run directly, not when it is imported, which keeps a module's self test out of an importer's output.

Why not pi = 3.14? Measured: 3.14 is out by about 0.0016, and in an area that error is multiplied by the radius squared.

Practical 7 and 8: lists, dictionaries, tuples and files

One line for "do these lists share a member"? return any(a in second for a in first).

Why is the set version faster on long lists? A set finds a member by hashing, not by comparing. Counted on two lists of 200: the list versions did 40,000 comparisons, the set 0 with nothing in common and 1 with one match.

Does marks.sort() work on a dictionary? No, that is an AttributeError. sorted(marks.items(), ...) returns a list.

Sort a dictionary by value descending. Sort marks.items() with key=lambda p: p[1] and reverse=True, then wrap it in dict().

Two students tie. Which comes first? The one inserted first, because Python's sort is stable. Break the tie with a tuple key, (p[1], p[0]).

How many values can a function return? One. Returning two means returning one tuple.

What makes a tuple, brackets or the comma? The comma. (7) is the number 7; (7,) is a tuple of one.

Why a tuple and not a list for area and circumference? It is exactly two different things, always two, and it cannot be changed, so it can be a dictionary key.

Why with open(...)? It closes the file even if an error is raised inside the block.

What does "w" do to an existing file? Empties it immediately, before anything is written.

Does write add a newline? No, and neither does writelines.

Text mode against binary mode? Text moves str and encodes on the way; binary moves bytes with no translation.

Three ways to read the last n lines, and the cost of each? readlines()[-n:] holds the whole file; deque(handle, maxlen=n) reads it all but holds only n lines; seeking from the end reads only the tail, which is what a huge file needs.

munotes.in288

The Viva: the Questions Asked at the Table

What is the danger of pickle? Loading one can execute code, so a pickle from somebody else is not safe.

Practical 9 and 10: regular expressions, dates, timing, the calendar

Why is text.count("note") wrong? It counts the letters wherever they appear, including inside "notes", "notebook" and "denote". On the test file it gave 9 against the correct 4.

What does \b match? A word boundary: the empty position between a word character and a non-word character.

Why must a pattern be a raw string? In an ordinary string \b is the backspace character, so the pattern matches nothing.

re.match against re.search? match only matches at position zero; search looks anywhere.

What is re.escape for? So a word's own characters are treated as text. Searching for "C++" without it is a broken pattern.

Why does the naive href pattern miss links? It requires double quotes, requires href first, is case sensitive and cannot cross a newline. It also mangled one link by running on to the "> at the end of the tag.

Greedy against lazy? . takes as much as it can; .? as little. Matching up to a closing delimiter needs the lazy form.

Can a regular expression parse HTML? No. HTML is not a regular language. The pattern lost a link whose title attribute contained a > and wrongly reported one inside a comment; html.parser got both right.

Why can two DD/MM/YYYY strings not be compared? The comparison goes left to right and the leftmost part is the day, so it compares days first.

strptime or strftime? P for parse, F for format.

Why does timedelta have no months? A month has no fixed length.

Which clock do you measure with? time.perf_counter(), because it is monotonic. time.time() follows the system clock and can go backwards.

Why timeit rather than one reading? It runs the code many times and disables the garbage collector, so the per run figure is steady.

What does calendar.monthrange(y, m) return? A tuple: the weekday of the 1st and the number of days in the month.

What number is Monday? 0 in the calendar module.

Module 2, the ten exercises

Practical 1 and 2: arrays and linked lists

Why is reading a[500] no slower than a[0]? The address is the start plus 500 times the slot size: arithmetic, not searching.

Capacity or size? Capacity is how many slots exist and never changes; size is how many hold data.

Which way must an insertion shift, and why? Backwards, from the last used slot towards the gap. Forwards overwrites the next value before it has been copied and smears one value across the rest.

munotes.in289

The Viva: the Questions Asked at the Table

How many items shift when inserting at position i of n? n - i. None at the end, all n at the front.

What should a linear search return when absent? -1, never 0, because 0 is a valid index.

What does a node hold? A value and a reference to the next node. The last reference is None.

Why does insert_at_end need a tail? Without one it walks from the head, which is O(n), so building n items by appending would be n squared.

What must happen when you insert into an empty list? Both the head and the tail must be set.

To insert at position 5, how many nodes do you walk? Four, to reach position 4, the node before it.

Why does deletion need a trailing reference? Removing a node means changing the nxt of the node in front, and a singly linked node has no way back.

What are the three cases of deletion? The head, the middle, and the tail, where the tail must also be moved back. A fourth: the only node, which sets head and tail to None.

What goes wrong if you do not move the tail? The tail points at a removed node, so the next append is linked on to it and the value never appears. Nothing raises, which is why it is dangerous.

How do you test for that bug in one line? Delete the tail, then append. If the value appears, the tail was updated.

Does holding the node make deletion cheaper? No. The deletion changes the previous node's link and there is no way back, so the walk is still O(n). A doubly linked list fixes it.

Which is better, array or linked list? Neither. Measured: building 500 items at the front moved 124,750 values in an array and 0 in a list; reading the middle took 1 step in the array and 251 in the list.

Practical 3 and 4: the stack and the queue

LIFO or FIFO, and which is which? A stack is LIFO, last in first out, one end. A queue is FIFO, first in first out, two ends.

Why is top initialised to -1? An empty stack has no topmost item, and then size is simply top + 1.

pop against peek? pop removes and returns the top; peek returns it and leaves it.

What should pop do on an empty stack? Raise. Returning None cannot be told from having pushed a None.

What do stack operations cost? All O(1). Only one end is ever touched.

munotes.in290

The Viva: the Questions Asked at the Table

Three faults a bracket matcher must catch? A closer with nothing open; the wrong kind of closer; and an opener never closed, found by the stack not being empty at the end.

What is the defect of a linear queue? rear only moves forward, so it reports itself full while slots at the front are free. Shown: five slots, three dequeued, and the sixth item refused with three slots empty.

Why not shift everything after each dequeue? It makes dequeue O(n) instead of O(1).

Write the line that makes a queue circular. Advance rear as (rear + 1) % capacity, and the same for front.

Why can full and empty not be told apart? From front and rear alone the two states look identical: front sits just after rear in both.

What are the two cures? Keep a count, so count == 0 is empty and count == capacity full; or never use the last free slot. The first uses every slot.

Why is a queue on list.pop(0) quadratic? pop(0) shifts every remaining item down, so emptying n items costs about n squared over two moves. Measured: at n of 2000 to 16,000 the ratio against deque.popleft() grew from 4 to 20.

Why a queue for a service counter? Fairness means the longest waiting customer goes next, which is exactly FIFO.

How do you tell whether one counter is enough, before running anything? Arrivals per minute × mean service time against the number of counters. Here 0.45 × 3 = 1.35 against 1, so one is not enough and two are.

Why put the seed in a simulation? So the teacher can reproduce your output and so a bug that appears once can be found again.

Practical 3(b): infix to postfix

Why does postfix need no brackets? The position of the operator already says which operands it applies to.

Convert a + b c and (a + b) c. a b c + and a b + c .

State rule 4. When an operator arrives, pop to the output every operator on the stack that should come out first, then push the new one.

What does "should come out first" mean? Higher precedence pops; equal precedence pops only when the arriving operator is left associative; an opening bracket never pops.

Which operator is right associative? ^. So 2 ^ 3 ^ 2 becomes 2 3 2 ^ ^, and evaluates to 512, not 64.

What happens to the brackets? They are never output. An opening bracket is removed by its matching closing bracket and both are discarded.

Which popped value is the left operand when evaluating? The second one popped. right = pop() then left = pop().

munotes.in291

The Viva: the Questions Asked at the Table

Why does the wrong order pass most tests? Addition and multiplication give the same answer either way round. It fails on subtraction, division, remainder and power.

How do you know a conversion is right rather than plausible? Evaluate the postfix and compare the value with the value of the original infix. Fourteen expressions were checked that way.

Practical 5 and 6: the tree

State the BST rule. Every key in a node's left subtree is smaller than the node, and every key in its right subtree is larger.

Why subtree and not child? A key can be a legal child and still be on the wrong side of an ancestor. Shown: 6 as the left child of 15 under a root of 10, where a search for 6 goes left at the root and never finds it.

How do you prove a tree is a BST? Walk it in order: the keys come out sorted. Or check it with bounds passed down.

How many comparisons does a search cost? At most the height plus one, because each comparison moves down one level.

Where is the smallest key? The leftmost node. Follow left until it is None.

What happens with sorted input? Every node becomes the right child of the one before, so the tree is a linked list. Measured at 127 keys: height 6 and 7 comparisons balanced, against height 126 and 127 comparisons sorted.

So what is a BST's cost, honestly? O(log n) if balanced, O(n) if not, and plain insertion does nothing to keep it balanced.

Two cures? Shuffle the keys, or insert the middle first. For any input, an AVL or red-black tree.

Give the three traversals. Pre-order: visit, left, right. In-order: left, visit, right. Post-order: left, right, visit.

How do you remember which is which? The prefix says where the node goes: before its subtrees, between them, or after them.

Which traversal gives sorted order on a BST? In-order, whatever the tree's shape.

Which structure does level order use? A queue. All three depth first traversals use a stack.

Can you rebuild a tree from one traversal? No: three different trees share one pre-order walk. Pre-order plus in-order is enough.

Which traversal deletes a tree, and why? Post-order. A parent holds the only pointers to its children, so freeing it first makes them unreachable.

What are the three deletion cases? A leaf is removed; one child takes its place; two children means copy the in-order successor, the smallest in the right subtree, and delete it.

Practical 7: hashing

What are the three parts of a hash table? A hash function, the modulo that gives the bucket index, and collision handling.

munotes.in292

The Viva: the Questions Asked at the Table

What is separate chaining? Each bucket holds a list of the pairs that hashed to it, and a lookup searches only that list.

Why is adding the character codes a bad hash? Addition ignores order, so every rearrangement collides. Measured: Anil, Lina, Nail, Lani gave one distinct hash out of four.

How do you judge a hash function? By the longest chain on your own data. And reject outright one whose collisions are structural.

What is the load factor? Items divided by buckets. The average comparisons in a lookup track it.

What is rehashing, and why? Doubling the buckets and placing every key again, because the index is the hash modulo the bucket count, which has changed.

Is a collision an error? No. With more keys than buckets it is certain by the pigeonhole principle, and 23 keys in 365 buckets collide more often than not.

What does a lookup cost? O(1) on average with a good hash and a low load factor; O(n) in the worst case. Measured: one bucket and 10,000 keys gave 5000.50 average comparisons, which is n over two.

Why must put search the chain first? Otherwise the same key appears twice and get returns whichever it finds first.

What is a tombstone, and who needs one? A marker left where a pair was deleted. Open addressing needs it, because a search stops at the first empty slot. Separate chaining does not.

When would you choose a BST over a hash table? When the order matters. A BST gives the sorted order free and the smallest key in O(log n).

Practical 8 and 9: sorting and searching

Describe the three sorts in one sentence each. Bubble swaps adjacent out-of-order items so the largest moves to the end. Insertion slides each item back into the sorted part. Selection finds the smallest remaining and swaps it into place.

Why is bubble's inner loop range(n - 1 - i)? After pass i the last i items are already in place.

What does the swapped flag do? Stops the sort when a pass makes no swap, which makes bubble sort O(n) on sorted data. Without it, bubble sort has no advantage at all.

Worst case comparisons? n(n-1)/2 for all three, verified exactly for n from 5 to 80.

Which improves on sorted data? Bubble with the early exit and insertion, both down to n-1. Selection does not: it still made 780 comparisons on 40 sorted items.

Which makes the fewest movements? Selection, at most one swap per pass.

What is a stable sort, and which of the three are? Equal items keep their order. Bubble and insertion are stable; selection is not, because a swap can jump an item over an equal one.

munotes.in293

The Viva: the Questions Asked at the Table

So which sort is best? None outright. Insertion wins on comparisons on random data, selection on movements in every case, bubble wins nothing and is taught because it explains easily.

What does binary search require? A sorted array, and random access, so not a linked list.

Why while low <= high? With <, a range of one item is never compared, so the search misses values that are present.

Why low = middle + 1 and not middle? Leaving the middle in the range means a range of two can stop shrinking and the loop never ends.

Why write low + (high - low) // 2? (low + high) can overflow a fixed width integer in C or Java. That bug was in the Java standard library for nine years.

Give the measured figures for a million items. Linear search 1,000,000 comparisons in the worst case; binary search 20.

When does linear search win? When the target is at the front; on a linked list; on tiny data; and when the data is unsorted and searched only a few times, because the sort must be paid for. Measured: 500 items, the sort cost 57,745 comparisons, and linear search was still ahead at 200 searches.

Does plain binary search find the first duplicate? No, it finds one of them. Record the hit and keep searching left.

Practical 10: the combined application

How did you decide which structures to use? I wrote the list of operations and how often each happens, and each frequent operation named a structure.

Why a hash table for the catalogue? Lookup by exact accession number happens on every transaction, and that is O(1) on average; the order is never needed.

Why a tree as well? The catalogue has to be listed in title order, and a hash table has no order. The tree's in-order walk gives the sorted list free.

Why a queue for reservations? A waiting list is fair only if the first to ask is first served.

Why a stack for undo? The most recent action must be undone first, which is LIFO.

What can the graph answer that nothing else can? Whether two borrowers are connected through other people. In the run, Eshan is reachable from Aarti through Bhavesh although they share no book.

Was there a defect in your program? Yes: 60 titles added in title order gave a tree of height 59, a linked list. Shuffling the insertion order gave height 10 and 5 comparisons instead of 60, with the same sorted catalogue.

munotes.in294

The Viva: the Questions Asked at the Table

The four answers that lose the mark

"It is O(n)." Without saying of what, and why. Say what happens when n doubles.

"The program was executed successfully." As a conclusion it says nothing. Say what the exercise showed.

A number you did not measure. If you did not count it, say so, and then say what you did count. An invented figure is worse than none, because the next question will test it.

"Because it is faster." Faster at which operation, and than what. Name the operation.

The seven things to be able to say without thinking

If there is only time to prepare seven answers, these are the seven, because they cover most of the viva.

  1. Big O in one sentence. How the work grows as the data grows. Doubling n doubles O(n), quadruples

O(n squared), and adds one step to O(log n).

  1. Array against linked list. An array reads in one step and shifts to insert; a linked list inserts at

either end in one step and has to walk to reach position n.

  1. Stack against queue. LIFO with one end against FIFO with two.
  2. Why \b. It is a word boundary, so \bnote\b counts the word and not the letters.
  3. The BST rule, with the word subtree. And that the in-order walk is sorted, which is the proof.
  4. Hash table cost. O(1) on average with a good hash and a low load factor, O(n) in the worst case.
  5. Which sort, and why none of them. All three are O(n squared); insertion is best on nearly sorted

data, selection makes the fewest movements, and real work uses the library sort at O(n log n).

Procedure

  1. Read your own journal cover to cover the night before, one entry at a time.
  2. For each entry, say out loud: what it does, why you did it that way, what happens at the edges, and what

it costs.

  1. Write the four kinds of question on the inside cover so the shape of every answer is familiar.
  2. Learn the seven answers above until they need no thought.
  3. Check every number you intend to quote against the entry it came from. Quote only what you counted.
  4. Practise saying "I did not measure that" followed by what you did measure.
  5. Have your journal open, signed, and with the index filled in.

Result

A set of answers for all twenty exercises, each in the three clause shape of name, reason and number, with every number taken from a run recorded earlier in this book. Together with a complete signed journal, that is what Q3's five marks are given for.

Where marks are lost

  • Not preparing for the viva at all. Five marks out of thirty, and no other chapter covers it.
  • Quoting a number you did not measure. The next question tests it.
  • "It is fast" or "it is slow" with no operation named.
  • Reciting a definition you cannot apply. "Amortised" will be followed by "what does that mean".
  • Not knowing your own program's edge cases. The empty list and the full array are the two commonest
munotes.in295

The Viva: the Questions Asked at the Table

questions.

  • Not knowing your own program's defects. Naming one yourself is a strong answer, not a weak one.
  • An unsigned or incomplete journal, which is half of Q3 before a word is said.

For the journal

Nothing here is a numbered entry. Two pages are worth keeping inside the back cover.

The four kinds of question, so that every answer starts in the right place.

The seven answers, written in your own words. Written in your own words is the point: an answer copied from a book sounds like one.

Quick revision

  • Q3 is journal and viva for 5 marks, a sixth of the external paper.
  • Four kinds of question: what does it do, why that way, what if, what does it cost. The last two are

where marks are won.

  • Answer in three clauses: name the thing, give the reason, give the number.
  • Never quote a number you did not measure. "I did not measure that" is a better answer.
  • The seven to know cold: Big O, array against linked list, stack against queue, why \b, the BST rule and

the sorted in-order walk, the hash table's average and worst case, and why no sort of the three is best.

  • Know your program's edge cases: the empty list, the full array, the duplicate key, the sorted input.
  • Know your program's defects, and name one before you are asked.

Questions you should be able to answer

1. What is Q3 worth and what is it for? Five marks out of the external thirty, for the journal and the viva together.

2. What are the four kinds of viva question? What does the program do; why did you do it that way; what happens if something changes; and what does it cost.

3. What shape should an answer have? Name the thing, give the reason, give the number if there is one. About twenty seconds.

4. What do you say when you are asked for a figure you never measured? That you did not measure it, and then what you did measure. An invented number will be tested by the next question.

5. Why is "a hash table is fast" a weak answer? It names no operation. The strong form is "a hash table, because we look up by key on every transaction, and that is O(1) on average".

munotes.in296

The Viva: the Questions Asked at the Table

6. Give the honest cost of a binary search tree. O(log n) if it is balanced and O(n) if it is not, and plain insertion does nothing to keep it balanced. Measured at 127 keys: 7 comparisons balanced against 127 from sorted input.

7. Should you volunteer a defect in your own program? Yes. Naming the sorted-input problem in your tree, with the measured heights and the cure, is one of the strongest answers available at the table.

Contents This chapter on its own page

munotes.in297

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Report or request
Done!