munotes®

Data Structures Notes | B.Sc. (Computer Science) Semester 3 | Mumbai University | munotes

Get access to whole semester resourcesSemester Pass

Official Notes munotes.in

Data Structures

B.SC. (COMPUTER SCIENCE) · SEMESTER 3

Strictly as per the University of Mumbai NEP syllabus in force for B.Sc. (Computer Science)

For B.Sc. (Computer Science) students of the University of Mumbai and all its affiliated colleges

Open the book ↓

munotes.in Second Year

Data Structures

Copyright © 2026 munotes.in. All rights reserved.

Written and first published by munotes.in, 2026.

This book is free for individual students to read at munotes.in. No part of it may be reproduced, distributed, stored, translated or used for institutional or classroom purposes in any form without a prior written licence from munotes.in.

Licensing and permissions: contact@munotes.in

The text of statutes and of judgments reproduced in this book is in the public domain under section 52(1)(q) of the Copyright Act 1957. The commentary, arrangement, examples and questions are the original work of munotes.in.

munotes.in is an independent study resource for MU students. It is not affiliated with, endorsed by, or officially connected to the University of Mumbai. Course names and university references describe the students and syllabus the material relates to.

munotes.in

Contents

Module I Abstract Data Type, Linked Structures, Stacks and Queues

  1. What a Data Structure Is, and Why the One You Pick Decides Whether the Program Works 1
  2. How This Paper Is Examined, and How to Read This Book 5
  3. Data and Its Types: What a Type Actually Settles 8
  4. The Primitive Types, and Where They Stop 11
  5. Linear and Non-Linear, Static and Dynamic: The Map of the Subject 14
  6. Choosing a Structure: The Three Questions to Ask of Any of Them 17
  7. The Abstract Data Type: A Promise, and a Hidden Representation 20
  8. Why Hiding the Representation Is the Whole Point 23
  9. Writing an ADT of Your Own 26
  10. Judging an ADT: Complete, Minimal, and Honest About Cost 29
  11. The Array: What It Really Is in Memory 32
  12. Where the Array Stops: Insertion, Deletion and Growth 35
  13. The Linked List: The Node, the Chain and the Head 38
  14. The Linked List ADT, and Building an Empty One 41
  15. Traversing a Singly Linked List 44
  16. Searching a Singly Linked List 47
  17. Prepending a Node 50
  18. Appending a Node, and Why It Costs More 53
  19. Removing a Node 56
  20. Inserting and Deleting at Any Position 59
  21. What a Singly Linked List Is Good and Bad At 62
  22. The Array Against the Linked List, Measured 65
  23. A Polynomial as a Linked List 68
  24. Adding Two Polynomials 71
  25. Multiplying Polynomials, and What the Representation Costs 74
  26. The Doubly Linked List: The Second Link 77
  27. Insertion and Deletion With Two Links 80
  28. Traversing Both Ways, and the Applications That Need It 84
  29. What the Second Link Costs and What It Buys 87
  30. The Stack: One End, and Why That Is Enough 90
  31. The Stack ADT: Push, Pop, and the Errors 93
  32. A Stack on an Array, With Peek 96
  33. A Stack on Links 99
  34. What a Stack Is Good and Bad At 102
  35. Balanced Delimiters, and Why a Counter Is Not Enough 105
  36. Infix, Prefix and Postfix 108
  37. Infix to Postfix With a Stack 111
  38. Evaluating a Postfix Expression 114
  39. Prefix: Converting to It, and Evaluating It 118
  40. The Queue: Two Ends 122
  41. The Queue ADT 125
  42. A Queue on an Array, and the Drift That Ruins It 128
  43. A Queue on Links 131
  44. The Circular Queue: Wrap-Around 134
  45. Full or Empty: Telling Them Apart 137
  46. What a Queue Is Good and Bad At 140
  47. The Deque: Open at Both Ends 143
  48. Job Scheduling With a Queue 146
  49. Module 1 in One Sitting 149

Module II Trees, Priority Queues and Heaps, Graphs and Hashing

  1. From a Line to a Tree: Why Linear Structures Run Out 153
  2. The Tree ADT: The Words, Said Exactly 156
  3. Height, Depth, Level and Size, and the Relations Between Them 159
  4. What a Tree Buys, and What It Costs 162
  5. The Binary Tree 165
  6. Binary Tree Properties, Proved 168
  7. Full, Complete and Perfect, and Why the Difference Matters 171
  8. Implementing a Binary Tree With Links 174
  9. The Array Representation of a Binary Tree 177
  10. Inorder Traversal 180
  11. Preorder and Postorder Traversal 183
  12. Level Order Traversal, and the Queue It Needs 186
  13. Iterative Traversal, and the Stack It Needs 189
  14. Rebuilding a Tree From Two Traversals 192
  15. The Binary Search Tree: The Invariant 196
  16. Searching a Binary Search Tree 199
  17. Inserting Into a Binary Search Tree 202
  18. Deleting From a Binary Search Tree: The Three Cases 205
  19. Why a Binary Search Tree Degenerates 209
  20. What Balance Means 212
  21. Threaded Binary Trees 216
  22. The AVL Tree and the Balance Factor 220
  23. The Four Rotations 223
  24. Insertion Into an AVL Tree 227
  25. Deletion From an AVL Tree 231
  26. Huffman Coding: The Problem 235
  27. Building the Huffman Tree 238
  28. Why the Huffman Code Is Prefix-Free, and What It Saves 241
  29. The Priority Queue: When First In Is the Wrong Rule 246
  30. The Priority Queue ADT 249
  31. Three Ways to Build One, and What Each Costs 252
  32. The Heap: Shape and Order 255
  33. The Array That Holds a Heap 258
  34. Min-Heap and Max-Heap 262
  35. Heapify: Sifting Up and Sifting Down 265
  36. Building a Heap, and Why It Is Linear 269
  37. Where Priority Queues Are Used 273
  38. What a Graph Is 276
  39. The Vocabulary of Graphs 279
  40. The Graph ADT 282
  41. The Adjacency Matrix 285
  42. The Adjacency List 288
  43. Which Representation: The Costs, Measured 291
  44. Inserting and Deleting Vertices and Edges 296
  45. Breadth First Search 302
  46. Depth First Search 307
  47. Connectivity and Connected Components 312
  48. The Shortest Path in an Unweighted Graph 317
  49. Dijkstra's Algorithm 321
  50. The Idea of Hashing: A Key Turned Into an Address 328
  51. The Hash Table ADT 334
  52. Hash Functions 339
  53. What Makes a Hash Function Good 348
  54. Collisions Are Certain, Not Unlucky 354
  55. Chaining 361
  56. Linear Probing 367
  57. Quadratic Probing and Double Hashing 376
  58. Load Factor and Rehashing 384
  59. What Hashing Buys and What It Gives Up 391
  60. Where Hashing Is Used, and Where It Must Not Be 396
  61. Choosing the Right Structure: The Whole Paper on One Page 402
  62. Module 2 in One Sitting 407
munotes.in

Module I

Abstract Data Type, Linked Structures, Stacks and Queues

munotes.in

Chapter One

What a Data Structure Is, and Why the One You Pick Decides Whether the Program Works

Syllabus topic Module 1, "Abstract Data Type: Different Data Types, different types of data structures & their classifications"

In one line

A data structure is a way of arranging data in memory together with the operations that arrangement makes cheap, and choosing it well is the difference between a program that answers and a program that hangs.

The experiment this whole paper is about

Here is a problem you already have. You hold a list of student roll numbers. Somebody gives you a roll number and asks: is it in the list?

The data is the same in both programs below. The question is the same. The only difference is how the roll numbers are arranged.

import random, time

random.seed(7)
N = 40000
rolls = random.sample(range(1, 10_000_000), N)
queries = random.sample(rolls, 2000)

as_list = list(rolls)
as_set = set(rolls)

start = time.perf_counter()
found = sum(1 for q in queries if q in as_list)
list_seconds = time.perf_counter() - start

start = time.perf_counter()
found_again = sum(1 for q in queries if q in as_set)
set_seconds = time.perf_counter() - start

print("roll numbers held      :", N)
print("lookups done           :", len(queries))
print("found, list            :", found)
print("found, set             :", found_again)
print("the answers agree      :", found == found_again)
print("set at least 50x faster:", set_seconds * 50 < list_seconds)
roll numbers held      : 40000
lookups done           : 2000
found, list            : 2000
found, set             : 2000
the answers agree      : True
set at least 50x faster: True

Read the last line again. Both programs are correct. Both give the same answer. One of them does the job at least fifty times faster than the other, and the only thing that changed was the arrangement of the data.

Notice what the program prints and what it refuses to print. It does not print "3,195 times faster", because that number is different on every machine and on every run, and a number a student cannot reproduce is worse than no number. It prints a claim that is true with room to spare, and you can raise the fifty and run it again yourself. Every measurement in this book is printed in that form: something the machine settled, not something the author remembered.

Why that happened

The list keeps the roll numbers in a row. To answer "is 5142337 in here" it has no choice but to walk along, comparing, until it finds the number or reaches the end. With forty thousand numbers, an unsuccessful search compares forty thousand times.

The set computes where the number would be if it were present, and looks only there. It compares once or twice, whatever the size of the collection.

You will build both of those in this paper. The walking search is chapter 16. The computing search is chapter 99, where it is called hashing. Everything between the two is other arrangements, each one cheap at something and expensive at something else.

munotes.in1

What a Data Structure Is, and Why the One You Pick Decides Whether the Program Works

The definition, said properly

A data structure is a way of organising data in a computer's memory so that a particular set of operations on it is efficient.

Two halves of that sentence do work, and students usually keep only the first.

A way of organising data in memory. Contiguous cells, or scattered cells joined by addresses, or a branching arrangement. This is the part that is easy to draw.

So that a particular set of operations is efficient. No structure is fast at everything. Every structure in this paper is a bargain: it makes some operations cheap by making others expensive. The array indexes instantly and inserts slowly. The linked list inserts instantly and searches slowly. The hash table finds instantly and cannot tell you what comes next in order.

So the honest question is never "which data structure is best". It is "which operations does my problem do most often, and which structure is cheap at exactly those". A student who leaves this paper with that question in their head has got the value out of it.

A second experiment, so the first is not a fluke

Membership is one operation. Here is a different one, and this time the list wins.

import time

N = 40000

start = time.perf_counter()
row = []
for i in range(N):
    row.append(i)
append_seconds = time.perf_counter() - start

start = time.perf_counter()
front = []
for i in range(N):
    front.insert(0, i)
insert_seconds = time.perf_counter() - start

print("items added              :", N)
print("both hold the same       :", len(row) == len(front))
print("front at least 20x slower:", append_seconds * 20 < insert_seconds)
items added              : 40000
both hold the same       : True
front at least 20x slower: True

Same structure, same number of items, same machine. Adding at the end is cheap; adding at the front is not, because every existing item has to shuffle up one place to make room.

That is the second lesson, and it is sharper than the first: the cost belongs to the pair, the structure and the operation, never to the structure alone. "Lists are fast" is not a statement that means anything. "Adding at the end of a Python list is cheap, and adding at the front is not" is.

What you are going to build

This paper sets eight structures. Here they are in one sentence each, so you know where the road goes.

StructureThe arrangementWhat it is cheap at
Arrayone row of cells, side by sidereaching the nth item
Linked listcells anywhere, each holding the address of the nextinserting and removing
Stackanything, used at one end onlyundoing, matching, nesting
Queueanything, added at one end and taken from the otherserving in order of arrival
Treea branching arrangementsearching, while still inserting
Heapa tree in an array, loosely sortedfinding the largest or smallest
Graphanything joined to anythingrelationships, routes, reachability
Hash tablea computed addressfinding a key, without order
munotes.in2

What a Data Structure Is, and Why the One You Pick Decides Whether the Program Works

Every one of them is built from scratch in this book, run, and measured.

How to read the timings in this book

A timing in seconds is a fact about a machine, not a fact about a structure. This book therefore never says "a linked list takes 0.4 seconds". It says what happens to the time when the data gets bigger, because that is the part that belongs to the structure and travels to your machine.

The three shapes you will meet, over and over:

  • Constant. Double the data, the work stays the same. Written O(1).
  • Linear. Double the data, double the work. Written O(n).
  • Logarithmic. Double the data, add one step. Written O(log n).

The last one is the one worth feeling rather than memorising. Going from a thousand items to a million items multiplies the data by a thousand and adds about ten steps. That is what a tree buys you, and it is why half this paper is trees.

Machine used for every measurement in this book: the timings were produced by the program printed beside them, and the numbers you get will differ. What must not differ is the shape.

Quick revision

  • A data structure is an arrangement of data in memory together with the operations that arrangement

makes efficient.

  • No structure is good at everything. Every one is a bargain: cheap at some operations, expensive at

others.

  • Cost belongs to the pair, structure and operation, never to the structure by itself.
  • The right question is which operations the problem does most, not which structure is best.
  • Timings in seconds belong to a machine. What belongs to the structure is how the time grows when the

data grows: constant, linear or logarithmic.

  • Searching a row of 40,000 numbers one by one against computing where the number lives is the

difference this whole paper is about.

Test yourself

1. Define a data structure in one sentence, including both halves of the definition. A way of organising data in memory together with the operations that arrangement makes efficient. The second half is the half students drop, and it is the half that matters.

2. Is a linked list faster than an array? The question has no answer as asked. A linked list is cheaper at inserting and removing; an array is cheaper at reaching the nth item. Cost belongs to the pair of structure and operation.

munotes.in3

What a Data Structure Is, and Why the One You Pick Decides Whether the Program Works

3. In the first experiment both programs printed the same answer. What, then, was wrong with the slow one? Nothing was wrong with its correctness. It was wrong in its choice of arrangement for the operation it did most: it searched by walking when it could have searched by computing.

4. Why does this book refuse to say "a linked list search takes 0.4 seconds"? Because that is a fact about one machine on one day. The fact that belongs to the structure is that doubling the data doubles the work.

5. Going from 1,000 items to 1,000,000 items, how much extra work does an O(log n) operation do? About ten more steps. The data multiplied by a thousand; a logarithmic cost adds about ten.

6. Give one operation that an array is cheap at and a linked list is expensive at, and one the other way round. Reaching the nth item: the array computes the address, the linked list walks to it. Inserting at the front: the linked list moves one link, the array shuffles every element up.

Contents This chapter on its own page

munotes.in4

Chapter Two

How This Paper Is Examined, and How to Read This Book

Syllabus topic Module 1 and Module 2, the paper's own particulars (MU item 6.14 (N), rows 3 to 12)

In one line

Data Structures is a theory paper of 50 marks: a one hour written examination for 30, and 20 from two class tests and two assignments in college.

What the University sets

VerticalMajor
TypeTheory
Credits2
Hours30
Marks50
ModulesModule 1 (15 hours), Module 2 (15 hours)

Type: Theory. This matters, because three of your other Semester 3 papers are not. Computer Science Practical 3 is examined at a machine with a journal; this paper is examined on paper, in writing, in one hour.

The written examination, exactly

The pattern is not printed in this paper's own block. It is in the shared evaluation scheme at the back of the same circular, and it is this:

QuestionSet onChoiceMarks
Q. 1Module 1any 2 out of 410
Q. 2Module 2any 2 out of 410
Q. 3Modules 1 and 2any 2 out of 410
Total30, in 1 hour

Three things follow, and they shaped every chapter of this book.

Every sub-question is worth five marks. Twelve are printed, you answer six, in sixty minutes. That is about ten minutes an answer. So every chapter here ends with a Quick revision written at exactly that length: what a five mark answer contains, and nothing padded.

Both modules are guaranteed to be asked. Q.1 is Module 1 and Q.2 is Module 2. There is no such thing as a module you can skip.

Q.3 is set on the two modules together. This is the question students prepare for least and it is a third of the paper. It asks things that live between the modules: which structure would you use here, what does a tree give you that a list does not, why is a hash table not a substitute for a search tree. Chapter 110 of this book exists for Q.3 and for nothing else.

The internal 20

ComponentMarks
Class Test 1 on Module 110
Class Test 2 on Module 210
The two averaged10
Assignment on Module 15
Assignment on Module 25
The two together10

An assignment in this subject is almost always an implementation: write this structure, trace this operation, compare these two. That is why every technique in this book carries a complete program that was run, and not a fragment.

Passing: 40% in each component, and you must pass the internal and the external separately.

The practical is a different paper

Computer Science Practical 3 sets ten exercises on Data Structures in its Module 2, and it has its own 50 marks, its own two hour examination at a machine, and its own certified journal which you cannot sit the practical without. It is not examined here and this book is not its journal.

munotes.in5

How This Paper Is Examined, and How to Read This Book

But it is the boundary of this paper, and the University drew it: four things the practical sets are taught in this book because a student is examined on them and no theory label names them. They are evaluating a postfix expression (chapter 38), the array-backed stack and peek (chapter 32), the comparison of the array against the linked list (chapter 22), and chaining and linear probing by name (chapters 104 and 105).

How this book is built

Every program in it was run. Not checked by reading, run. The output printed under a program is what the interpreter printed, on three versions of Python, and if the three disagreed the book says so. A gate refuses to build this book if any printed output is not what the program produced.

Every trace was produced by the machine. Where you see a tree after each of six insertions, or an array after each swap of a heapify, the program printed it. Nobody typed it out and hoped.

Every cost claim was measured, at growing sizes, and what is printed is either a counted operation or a claim with room to spare. You will not find "this takes 0.4 seconds" anywhere in this book, because that is a fact about somebody else's laptop.

Every chapter has the same shape: what it is in one line, the substance, a Quick revision at five mark length, and Test yourself with the answers written out.

How to use it in the last week

  • Read the Quick revision of every chapter in a module. That is the module in about an hour.
  • Work the Test yourself questions with the answers covered.
  • Then read chapter 110 twice. It is the Q.3 chapter, it is a third of the paper, and it is the part

most students walk in without.

Quick revision

  • Data Structures is a Major, Type Theory, 2 credits, 30 hours, 50 marks.
  • External: 1 hour, 30 marks. Q.1 Module 1, Q.2 Module 2, Q.3 both, each any 2 of 4 for 10.
  • Every sub-question is 5 marks and about ten minutes.
  • Internal 20: two class tests averaged to 10, two assignments totalling 10.
  • Pass 40% in each component, internal and external separately.
  • Q.3 spans both modules and is a third of the paper; chapter 110 is written for it.
  • Computer Science Practical 3 is a separate paper with its own marks and its own journal.

Test yourself

1. How long is the written paper and how many marks is it? One hour, 30 marks, out of a total of 50 for the subject.

munotes.in6

How This Paper Is Examined, and How to Read This Book

2. How many sub-questions are printed, how many do you answer, and what is each worth? Twelve printed, six answered, five marks each.

3. Which question is set on both modules together? Q.3. It is 10 of the 30 external marks.

4. Can you skip Module 2 and still pass the written paper? No. Q.2 is set on Module 2 alone and Q.3 draws on both, so 20 of the 30 marks need it.

5. What is the internal 20 made of? Two class tests, one per module, averaged to 10; and two assignments, one per module, totalling 10.

6. Is the practical examined in this paper? No. Computer Science Practical 3 is a separate paper with its own examination and its own compulsory journal, though four things it sets are taught here because they are examinable and the theory labels do not name them.

Contents This chapter on its own page

munotes.in7

Chapter Three

Data and Its Types: What a Type Actually Settles

Syllabus topic Module 1, "Abstract Data Type: Different Data Types"

In one line

A type is three promises at once: which values are allowed, which operations are allowed, and how the value is stored.

Data, before types

Memory is bytes. A byte is eight bits and holds a number from 0 to 255, and that is all a machine has. The same eight bits 01000001 are the number 65, the letter A, or part of something larger. Nothing in the bits says which.

A type is the answer to that question. It is the label that tells the machine, and the reader, how to understand the bits.

The three promises

1. The set of values. An integer type allows whole numbers; a boolean allows exactly two things. Asking for a value outside the set is an error the type can catch.

2. The operations. You may add two integers. You may add two strings in Python, and it means joining them. You may not subtract one string from another, and the type is what makes that a refusal rather than nonsense.

3. The representation. How many bytes, in what arrangement. This is the promise students forget, and it is the one that decides what things cost.

Here are all three, observed rather than asserted:

x = 65
c = "A"
b = True

print("value      :", x, "| type:", type(x).__name__)
print("value      :", c, "| type:", type(c).__name__)
print("value      :", b, "| type:", type(b).__name__)
print()
print("65 + 1     :", x + 1)
print("'A' + '1'  :", c + "1")
print("True + True:", b + b, "  (a bool is a kind of int in Python)")
print()
try:
    print(c - c)
except TypeError as e:
    print("'A' - 'A'  : refused, and the type is why:", e)
value      : 65 | type: int
value      : A | type: str
value      : True | type: bool

65 + 1     : 66
'A' + '1'  : A1
True + True: 2   (a bool is a kind of int in Python)

'A' - 'A'  : refused, and the type is why: unsupported operand type(s) for -: 'str' and 'str'

Two things in that run are worth stopping on.

'A' + '1' gives A1, not 66. The plus sign means different operations on different types, and the type is what selects which. This is why "what does + do" has no answer without a type.

True + True gives 2. In Python a boolean really is a kind of integer, which is convenient and is also the sort of detail that makes a wrong answer in an examination if you assume every language does it.

The representation, which is the part that costs

Ask Python how big a value actually is:

import sys

for value, name in [(0, "int 0"), (255, "int 255"), (10**30, "int 10^30"),
                    (3.5, "float"), ("", "empty str"), ("A", "str 'A'"),
                    (True, "bool"), ([], "empty list"), ([1, 2, 3], "list of 3")]:
    print("%-12s %4d bytes" % (name, sys.getsizeof(value)))
munotes.in8

Data and Its Types: What a Type Actually Settles

int 0          28 bytes
int 255        28 bytes
int 10^30      40 bytes
float          24 bytes
empty str      41 bytes
str 'A'        42 bytes
bool           28 bytes
empty list     56 bytes
list of 3      88 bytes

Those numbers are this Python's, on this machine. Another build gives different ones and the book is not asking you to memorise them. What they show is the point: a value has a size, the size depends on the type, and an empty list is not free.

Notice that int 0 and int 255 are the same size, but 10**30 is bigger. Python's integers grow to fit the number, which is unusual and generous. In C or Java an int is a fixed number of bytes and a number too big for it wraps around silently. That difference is exactly why the next chapter exists.

Static and dynamic typing, in one paragraph

In C and Java you declare the type and it is fixed: int n; and n holds integers for ever. That is static typing, checked before the program runs. In Python the value carries the type and a name can point at anything, which is dynamic typing, checked as it runs.

This paper's structures are built in Python, so the types are dynamic. But every structure here is a structure in any language, and where a language's typing changes the design, this book says so.

Why a type is not a data structure

A type says what one value is. A data structure says how many values are arranged.

Put the other way: int is a type. "Ten thousand integers, in a row, in order" is a data structure. The paper is about the second, and it starts with the first because a structure is made of typed values and you cannot reason about the cost of the arrangement without knowing the cost of one item.

Quick revision

  • Memory is bytes; the bits do not say what they mean. A type is what says.
  • A type settles three things: the allowed values, the allowed operations, and the representation.
  • The same operator means different things on different types: + adds integers and joins strings.
  • Every value has a size, the size depends on the type, and an empty container is not free.
  • Python integers grow to fit; C and Java integers are fixed width and can overflow.
  • Static typing fixes the type of a name before the run (C, Java); dynamic typing carries the type on

the value (Python).

  • A type describes one value; a data structure describes an arrangement of many.
munotes.in9

Data and Its Types: What a Type Actually Settles

Test yourself

1. What three things does a type settle? The set of allowed values, the operations allowed on them, and how the value is represented in memory.

2. Why does 'A' + '1' not give 66? Because + is selected by the type. On strings it joins; on integers it adds. The characters are not being read as their numeric codes.

3. Which of the three promises decides what a structure costs? The representation. The set of values and the operations decide what is legal; the representation decides what is cheap.

4. Give one difference between a Python integer and a C integer that matters. A Python integer grows to fit any whole number; a C int is fixed width and overflows silently when the value is too large.

5. State the difference between static and dynamic typing. Static: the type belongs to the name and is fixed and checked before the run. Dynamic: the type belongs to the value and is checked while the program runs.

6. Is int a data structure? Explain. No. It is a type, describing one value. A data structure describes how many values are arranged, such as an array of integers.

Contents This chapter on its own page

munotes.in10

Chapter Four

The Primitive Types, and Where They Stop

Syllabus topic Module 1, "Abstract Data Type: Different Data Types"

In one line

The primitive types are the ones the language gives you ready made, each holding a single value, and the moment a problem has many values they are not enough on their own.

The four families

Every language you will meet has these, whatever it calls them.

FamilyHoldsPythonC
Integerwhole numbersintint, long, short
Realnumbers with a fractional partfloatfloat, double
Characterone letter or symbol(a str of length 1)char
Booleantrue or falsebool_Bool

Python has no separate character type: a single letter is just a string of length one. C has no separate boolean until C99, and used integers instead, 0 for false. These differences are examinable, so the table above says both.

What a primitive can and cannot do

n = 42
r = 3.75
c = "Z"
flag = False

print("integer  :", n, type(n).__name__, "| n * 2 =", n * 2)
print("real     :", r, type(r).__name__, "| r * 2 =", r * 2)
print("character:", c, type(c).__name__, "| ord(c) =", ord(c))
print("boolean  :", flag, type(flag).__name__, "| not flag =", not flag)
print()
print("integer division 7 // 2 :", 7 // 2)
print("real division    7 / 2  :", 7 / 2)
print("0.1 + 0.2 == 0.3        :", 0.1 + 0.2 == 0.3)
print("0.1 + 0.2 is actually   :", 0.1 + 0.2)
integer  : 42 int | n * 2 = 84
real     : 3.75 float | r * 2 = 7.5
character: Z str | ord(c) = 90
boolean  : False bool | not flag = True

integer division 7 // 2 : 3
real division    7 / 2  : 3.5
0.1 + 0.2 == 0.3        : False
0.1 + 0.2 is actually   : 0.30000000000000004

0.1 + 0.2 is not 0.3, and that is not a bug. A real number is stored in a fixed number of bits, and most decimal fractions cannot be written exactly in binary any more than one third can be written exactly in decimal. So a float is an approximation, and two floats should almost never be compared with ==. Compare the difference against a small tolerance instead. This costs students marks and costs programmers money, and it is a property of the type, which is why it belongs here.

Where the limits are

Integers in most languages have a largest value. Python's do not, and it is worth seeing the difference rather than being told it:

import sys

print("Python int: no fixed maximum. 2 ** 200 =")
print(2 ** 200)
print()
print("float: largest and smallest, from the machine itself")
print("  max float :", sys.float_info.max)
print("  min normal:", sys.float_info.min)
print("  digits it can be trusted to:", sys.float_info.dig)
print()
print("what happens past the top of a float:")
big = sys.float_info.max
print("  max * 2 =", big * 2)
munotes.in11

The Primitive Types, and Where They Stop

Python int: no fixed maximum. 2 ** 200 =
1606938044258990275541962092341162602522202993782792835301376

float: largest and smallest, from the machine itself
  max float : 1.7976931348623157e+308
  min normal: 2.2250738585072014e-308
  digits it can be trusted to: 15

what happens past the top of a float:
  max * 2 = inf

A float that grows past its largest value becomes inf rather than wrapping. A C integer that grows past its largest value wraps to a negative number silently, which is the more dangerous behaviour and the reason overflow is a standard examination topic.

Where the primitives stop, and the paper begins

Here is the wall, and the whole of the rest of this book is the answer to it.

A primitive holds one value. Consider the smallest realistic problem: the marks of the students in a class.

mark1 = 78
mark2 = 65
mark3 = 91

print("average of three:", (mark1 + mark2 + mark3) / 3)
print()
print("now do it for sixty students, with sixty named variables.")
print("and then sort them.")
print("and then find the student ranked seventh.")
average of three: 78.0

now do it for sixty students, with sixty named variables.
and then sort them.
and then find the student ranked seventh.

Three problems appear at once, and none of them is about arithmetic.

You cannot write sixty variables and mean it. The program would have to be rewritten for a class of sixty one.

You cannot loop over separately named variables. A loop needs a way to say "the next one", and separate names have no next.

You cannot sort them. Sorting means moving values around by position, and a name is not a position.

What is needed is a way to hold many values under one name, reachable by position. That is the array, it is the next structure in this book, and every other structure here is a response to something the array cannot do.

Quick revision

  • The primitive types are integer, real, character and boolean; each holds a single value.
  • Python has no separate character type (a one-letter string) and its booleans are a kind of integer.
  • // is integer division, / is real division.
  • Floats are approximations: 0.1 + 0.2 is not exactly 0.3, so never compare floats with ==;

compare the difference against a tolerance.

  • Floats have a largest value and overflow to inf. Fixed width integers in C wrap silently instead.
  • Python integers have no fixed maximum.
  • A primitive holds one value. A problem with many values needs many values under one name, reachable

by position, and that is where data structures begin.

munotes.in12

The Primitive Types, and Where They Stop

Test yourself

1. Name the four primitive families and what each holds. Integer (whole numbers), real or float (numbers with a fractional part), character (one symbol), boolean (true or false).

2. Why should two floats not be compared with ==? Because a float is a binary approximation of a decimal value, so arithmetic leaves tiny errors: 0.1 + 0.2 gives 0.30000000000000004. Compare the absolute difference against a small tolerance.

3. What is the result of 7 // 2 and how does it differ from 7 / 2? 7 // 2 is 3, integer division discarding the fraction. 7 / 2 is 3.5.

4. What happens when a float exceeds its maximum, and how does a fixed width C integer differ? The float becomes inf. The C integer wraps around to a negative value with no warning, which is the more dangerous of the two.

5. Give three reasons sixty separately named variables cannot hold a class's marks. The program cannot be written for an unknown class size; a loop has no way to move to the next one; and sorting needs positions, which names do not have.

6. What property must the next structure have, that a primitive does not? It must hold many values under one name and let any of them be reached by position.

Contents This chapter on its own page

munotes.in13

Chapter Five

Linear and Non-Linear, Static and Dynamic: The Map of the Subject

Syllabus topic Module 1, "different types of data structures & their classifications"

In one line

Data structures divide two ways at once: linear or non-linear by how the items are arranged, and static or dynamic by whether the size is fixed when the structure is made.

The first division: linear or non-linear

Linear means the items sit in a sequence. Each item has one item before it and one after it, except the two at the ends. You can say "the next one" and mean something.

Non-linear means they do not. An item may have several items after it, or be reachable from several directions.

LinearNon-linear
ArrayTree
Linked listGraph
StackHeap (a tree, held in an array)
QueueHash table (no order at all)

The test that settles it: can you write the items down in one row so that every item's neighbours in the row are its neighbours in the structure? If yes, linear.

A tree fails that test. The root has two children and they are not next to each other in any row you could write. A graph fails it harder: a city can be joined to five other cities.

The second division: static or dynamic

Static means the size is decided when the structure is created and does not change. Memory is allocated once, in one block.

Dynamic means the structure grows and shrinks while the program runs, taking memory as it needs it and releasing it when it does not.

StaticDynamic
Array (in C, Java)Linked list
Tree, graph
Python's list, which grows for you

The honest note for a Python student: Python's list looks dynamic and is built on a static array underneath. When it runs out of room it quietly makes a bigger array and copies everything across. That is why appending is usually cheap and occasionally expensive, and chapter 12 measures exactly that.

The two divisions are independent

This is the part worth drawing. The two questions are separate, so every structure sits in one of four boxes.

LinearNon-linear
Staticarraya tree held in a fixed array (a heap)
Dynamiclinked list, stack, queuetree, graph

A student who has learnt the two lists separately often thinks "linear means static". It does not: a linked list is linear and fully dynamic.

Primitive against non-primitive

MU's label says "different types of data structures", and there is one more cut that examiners ask for.

Primitive data structures are the types the language gives you for a single value: integer, real, character, boolean. Chapter 4 was about these.

Non-primitive are the ones built out of those: arrays, lists, stacks, queues, trees, graphs, files. Everything else in this paper is here.

Non-primitive splits again into linear and non-linear, which is the division above. So the full picture, which is the one to reproduce in an answer:

munotes.in14

Linear and Non-Linear, Static and Dynamic: The Map of the Subject

data structures

= primitive (int, float, char, bool)

= non-primitive

= non-primitive: linear (array, linked list, stack, queue)

= non-primitive: non-linear (tree, graph)

Where every structure in this paper sits

Here is the whole syllabus placed on the map, so the rest of the book has somewhere to hang.

StructureLinear?Static or dynamicChapter
Arraylinearstatic11
Singly linked listlineardynamic13
Doubly linked listlineardynamic26
Stacklineareither30
Queuelineareither40
Circular queuelinearstatic array, reused44
Dequelineareither47
Binary treenon-lineardynamic54
Binary search treenon-lineardynamic64
AVL treenon-lineardynamic71
Heapnon-linear, held linearlyarray81
Graphnon-lineareither87
Hash tableneitherarray of buckets100

Two rows in that table are the interesting ones, and both are examined.

The heap is a tree that lives in an array. It is non-linear as an idea and linear in memory, and chapter 82 shows the index arithmetic that makes that work.

The hash table is neither linear nor non-linear. It has no order at all. That is what it gives up in exchange for finding things instantly, and chapter 108 is about that bargain.

Quick revision

  • Linear: items in a sequence, each with one before and one after. Array, linked list, stack, queue.
  • Non-linear: an item may have many successors. Tree, graph.
  • The test for linear: can the items be written in one row with neighbours preserved?
  • Static: size fixed at creation, one block of memory. Dynamic: grows and shrinks during the run.
  • The two divisions are independent: a linked list is linear and dynamic.
  • Primitive structures hold one value; non-primitive are built from them and split into linear and

non-linear.

  • Python's list is dynamic on the outside and a static array on the inside, which it replaces with a

bigger one when it fills.

  • The heap is non-linear as an idea and linear in memory. The hash table has no order at all.

Test yourself

1. State the test that decides whether a structure is linear. Whether the items can be written in a single row so that every item's neighbours in the row are its neighbours in the structure.

2. Is a linked list static or dynamic? Is it linear? Dynamic and linear. It is the standard example that the two divisions are independent.

3. Give the four boxes of the linear/non-linear against static/dynamic table, with an example in each. Static linear: array. Dynamic linear: linked list. Static non-linear: a heap in a fixed array. Dynamic non-linear: tree or graph.

4. Why is a hash table hard to place on the linear/non-linear division? Because it has no order among its items at all. It is not a sequence and it is not a branching arrangement; the position of an item is computed from its key.

munotes.in15

Linear and Non-Linear, Static and Dynamic: The Map of the Subject

5. Draw the primitive against non-primitive classification. Data structures split into primitive (int, float, char, bool) and non-primitive; non-primitive splits into linear (array, linked list, stack, queue) and non-linear (tree, graph).

6. Python's list grows when you append to it. Does that make it a dynamic structure, and what is it underneath? It behaves dynamically, and underneath it is a static array that is replaced by a larger one, with everything copied across, when it fills.

Contents This chapter on its own page

munotes.in16

Chapter Six

Choosing a Structure: The Three Questions to Ask of Any of Them

Syllabus topic Module 1, "different types of data structures & their classifications"

In one line

Before choosing a structure, ask which operations the problem performs most often, what each of them costs in that structure, and what the structure costs in memory.

The three questions

1. Which operations does this problem actually do, and how often? Not which operations are possible. Which ones run in the inner loop. A program that inserts once and searches a million times has a completely different answer from one that does the reverse.

2. What does each of those operations cost in this structure? Not in seconds. In how the work grows when the data grows: constant, logarithmic or linear.

3. What does the structure cost in memory, beyond the data itself? Every structure carries overhead. A linked list stores an address beside every value. A hash table keeps empty space on purpose. Sometimes that is the deciding factor, especially on a small machine.

Working the questions on a real problem

A college wants to check, at the gate, whether a scanned ID is on the list of enrolled students. The list changes twice a year. There are 40,000 students and the gate is used 5,000 times a day.

Question 1. Search, 5,000 times a day. Insert, twice a year. Search wins by a factor of about a million, so the answer is whatever searches fastest, and insertion cost is nearly irrelevant.

Question 2. Searching an unsorted array is linear, 40,000 comparisons in the worst case. Searching a sorted array is logarithmic, about 16 comparisons at worst. Searching a hash table is constant, about 1.

Question 3. The hash table wastes some space on purpose. At 40,000 students that is a few hundred kilobytes. Nobody cares.

So: hash table. And notice that the reasoning did not need any of the three structures to be built. The questions settle it.

The same problem, one detail changed

Now the gate must also answer "which student comes next in roll number order", for a printed register.

Question 1 has changed: there is now an ordered operation, and it runs once a day.

Question 2 has changed with it. A hash table cannot answer "the next key in order" at all, without looking at every key. A sorted structure answers it instantly.

The answer is now a tree, or a hash table plus a sorted list. One line of the requirement changed the structure, and that is the whole lesson of this chapter.

Counting, rather than guessing

Two programs, same data, same question, with the work counted rather than timed. Counting is exact and identical on every machine, which is why this book prefers it to a stopwatch.

def linear_search(row, target):
    """Walk from the front. Returns (found, comparisons)."""
    comparisons = 0
    for item in row:
        comparisons += 1
        if item == target:
            return True, comparisons
    return False, comparisons


def binary_search(sorted_row, target):
    """Halve the range each time. Returns (found, comparisons)."""
    lo, hi, comparisons = 0, len(sorted_row) - 1, 0
    while lo <= hi:
        mid = (lo + hi) // 2
        comparisons += 1
        if sorted_row[mid] == target:
            return True, comparisons
        if sorted_row[mid] < target:
            lo = mid + 1
        else:
            hi = mid - 1
    return False, comparisons


for n in (1000, 2000, 4000, 8000):
    data = list(range(n))
    missing = -1
    _, linear = linear_search(data, missing)
    _, binary = binary_search(data, missing)
    print("n = %5d   linear search: %5d comparisons   binary search: %2d"
          % (n, linear, binary))
munotes.in17

Choosing a Structure: The Three Questions to Ask of Any of Them

n =  1000   linear search:  1000 comparisons   binary search:  9
n =  2000   linear search:  2000 comparisons   binary search: 10
n =  4000   linear search:  4000 comparisons   binary search: 11
n =  8000   linear search:  8000 comparisons   binary search: 12

Read the two columns down. The data doubles four times. The linear search doubles with it. The binary search adds one.

That is the difference between O(n) and O(log n), and it is the single most important pattern in this paper. You do not need to memorise it; you need to have seen it happen.

What the counts do not say

Binary search needs the data sorted. Sorting costs something, and if the data changes constantly you pay that cost over and over. This is question 1 again: how often does each operation run?

That is why the answer to "which is better, linear or binary search" is, correctly, it depends on how often the data changes, and an examiner asking the question is asking for exactly that.

The three questions, as a checklist

Keep this beside you while reading the rest of the book. Every structure here is introduced with an answer to it.

Ask
1What operations, how often?
2What does each cost here: constant, logarithmic, linear?
3What does the structure cost in memory beyond the data?
4What does it refuse to do?

The fourth is not a question so much as a warning, and it is the one that catches people. A hash table refuses order. A stack refuses the middle. A static array refuses to grow. A structure's refusals are as much a part of choosing it as its strengths.

Quick revision

  • Ask three things: which operations run most, what each costs here, and what the structure costs in

memory.

  • Cost means how the work grows with the data, not seconds.
  • Counted operations are exact and machine independent; prefer them to timings.
  • Linear search doubles when the data doubles; binary search adds one step. O(n) against O(log n).
  • Binary search needs sorted data, so "which is better" depends on how often the data changes.
  • Add a fourth question in practice: what does this structure refuse to do?
  • One line of a requirement can change the answer, as adding "in order" rules out a hash table.
munotes.in18

Choosing a Structure: The Three Questions to Ask of Any of Them

Test yourself

1. State the three questions. Which operations does the problem do most often; what does each cost in this structure; what does the structure cost in memory beyond the data.

2. Data doubles from 4,000 to 8,000. What happens to a linear search and to a binary search? The linear search doubles, from about 4,000 comparisons to about 8,000. The binary search adds one, from 11 to 12.

3. Why is "which is better, linear or binary search" not answerable as asked? Because binary search needs sorted data. If the data changes often the sorting cost dominates, and the right answer depends on how often each operation runs.

4. A system searches constantly and inserts twice a year. Which question settles the choice? Question 1. Search frequency outweighs insertion by such a margin that insertion cost is nearly irrelevant.

5. The requirement gains one line: "and list the students in roll number order". What does that rule out? A hash table, which has no order and cannot produce the next key without inspecting every key.

6. Why does this book count comparisons instead of timing the search? Because a count is exact and the same on every machine, while a timing belongs to one machine on one day.

Contents This chapter on its own page

munotes.in19

Chapter Seven

The Abstract Data Type: A Promise, and a Hidden Representation

Syllabus topic Module 1, "Abstract Data Type: Introduction to ADT"

In one line

An Abstract Data Type is a set of values together with the operations allowed on them, described by what the operations do and deliberately not by how they are stored.

The definition, taken apart

An ADT has exactly two halves.

The values. What the thing holds. A bag of integers. A sequence of names. A collection of students.

The operations, described by their behaviour. What each one does, what it needs, what it gives back, and what happens when it is asked for something impossible.

And one deliberate absence, which is the whole idea:

No representation. The ADT does not say array or linked list. It does not say how many bytes. A reader of the ADT learns what the structure promises and learns nothing about how it keeps that promise.

That absence is not vagueness. It is a decision, and the next chapter is about why it is the right one.

An ADT written out properly

Here is the ADT for a Bag: a collection that holds items, allows duplicates, and has no order.

OperationNeedsGives backBehaviour
Bag()nothinga new bagcreates an empty bag
add(item)an itemnothingputs the item in; duplicates allowed
remove(item)an itemnothingremoves one copy; error if absent
contains(item)an itemtrue or falseis at least one copy present
size()nothinga whole numberhow many items, counting duplicates
isEmpty()nothingtrue or falseis the size zero

Notice what the table settles and what it refuses to settle. It settles that remove on an absent item is an error rather than silently doing nothing, which is a real design decision a reader needs. It refuses to say whether the bag is an array or a list, because a user of the bag does not need to know and should not depend on it.

The same ADT, built two completely different ways

This is the demonstration that makes the idea concrete rather than a sentence to memorise.

class BagAsList:
    """A bag kept as one row of items."""

    def __init__(self):
        self._items = []

    def add(self, item):
        self._items.append(item)

    def remove(self, item):
        if item not in self._items:
            raise KeyError("not in the bag: %r" % (item,))
        self._items.remove(item)

    def contains(self, item):
        return item in self._items

    def size(self):
        return len(self._items)

    def is_empty(self):
        return self.size() == 0


class BagAsCounts:
    """The same bag, kept as a table of item to how many copies."""

    def __init__(self):
        self._counts = {}

    def add(self, item):
        self._counts[item] = self._counts.get(item, 0) + 1

    def remove(self, item):
        if self._counts.get(item, 0) == 0:
            raise KeyError("not in the bag: %r" % (item,))
        self._counts[item] -= 1
        if self._counts[item] == 0:
            del self._counts[item]

    def contains(self, item):
        return self._counts.get(item, 0) > 0

    def size(self):
        return sum(self._counts.values())

    def is_empty(self):
        return self.size() == 0


def stock_report(bag, wanted):
    """A user of the ADT. It never asks how the bag is built."""
    bag.add("pen")
    bag.add("pen")
    bag.add("book")
    bag.remove("pen")
    return (bag.size(), bag.contains(wanted), bag.is_empty())


for cls in (BagAsList, BagAsCounts):
    print("%-12s ->" % cls.__name__, stock_report(cls(), "pen"))
munotes.in20

The Abstract Data Type: A Promise, and a Hidden Representation

BagAsList    -> (2, True, False)
BagAsCounts  -> (2, True, False)

One function, stock_report, ran against two structures that share not one line of storage code. One keeps a row of items; the other keeps a table of counts. The caller could not tell the difference, because the ADT is what it was written against.

That is an ADT doing its job.

What the two representations actually cost

They are interchangeable to the caller and they are not equivalent to the machine, which is the second half of the lesson.

class BagAsList:
    def __init__(self):
        self._items = []

    def add(self, item):
        self._items.append(item)

    def size(self):
        return len(self._items)


class BagAsCounts:
    def __init__(self):
        self._counts = {}

    def add(self, item):
        self._counts[item] = self._counts.get(item, 0) + 1

    def size(self):
        return sum(self._counts.values())


import sys

N = 50000
kinds = 5

as_list = BagAsList()
as_counts = BagAsCounts()
for i in range(N):
    item = "kind-%d" % (i % kinds)
    as_list.add(item)
    as_counts.add(item)

print("items added            :", N)
print("distinct kinds         :", kinds)
print("both report the size   :", as_list.size(), as_counts.size())
print("counts version smaller :",
      sys.getsizeof(as_counts._counts) < sys.getsizeof(as_list._items))
items added            : 50000
distinct kinds         : 5
both report the size   : 50000 50000
counts version smaller : True

Fifty thousand items of five kinds. The row keeps fifty thousand entries; the table keeps five. Same ADT, same answers, wildly different memory.

So the ADT tells you what you may ask for, and the choice of representation tells you what it will cost. Both are needed, and separating them is what lets you change the second without breaking every program that relies on the first.

Why the University prints the syllabus this way

Read MU's own labels again. "ADT for linked list". "Stack ADT for Stack". "Queue ADT". "ADT for Tree Structure". "Priority Queue ADT". "Graph ADT". "Hash Table ADT".

Seven of the eight structures in this paper are introduced by their ADT. That is deliberate, and it means the examinable content for each structure is in two parts: what it promises and how it is built. A question asking you to "write the ADT for a stack" is asking for the first, and an answer that starts with an array has answered a different question.

Quick revision

  • An ADT is a set of values plus the operations on them, described by behaviour and not by

representation.

  • The missing representation is the point, not an omission.
  • The operations must say what happens in the impossible cases, such as removing an absent item.
  • One ADT can have many representations; a caller written against the ADT works with all of them.
  • The representations are not equivalent in cost, which is why the choice still matters.
  • MU introduces seven of this paper's eight structures by their ADT, so "write the ADT for X" is a
munotes.in21

The Abstract Data Type: A Promise, and a Hidden Representation

standard question and it is not asking for code.

Test yourself

1. Define an Abstract Data Type. A set of values together with the operations allowed on them, specified by what the operations do and deliberately not by how the values are stored.

2. What is deliberately left out of an ADT, and why is that not a weakness? The representation. Leaving it out is what lets the representation be changed without breaking any program written against the ADT.

3. An ADT for a bag says remove on an absent item is an error. Is that part of the ADT or part of the implementation? Part of the ADT. What happens in an impossible case is behaviour, and a user needs to know it.

4. Two bags, one a row of items and one a table of counts, gave identical answers. What does that show, and what does it not show? It shows they implement the same ADT, so a caller cannot tell them apart. It does not show they cost the same: with 50,000 items of 5 kinds the table is far smaller.

5. Why is the ADT written before the implementation is chosen? Because the ADT states the problem and the implementation is one answer to it. Choosing storage first fixes the costs before anyone has said what operations matter.

6. "Write the ADT for a stack." What should your answer contain and what should it not? It should contain the values it holds and each operation with what it needs, returns and does, including overflow and underflow. It should not contain an array, a linked list or any code.

Contents This chapter on its own page

munotes.in22

Chapter Eight

Why Hiding the Representation Is the Whole Point

Syllabus topic Module 1, "Abstract Data Type: Introduction to ADT"

In one line

A program written against the ADT survives a change of representation; a program written against the representation breaks, and the break is silent.

The experiment

Two programs do the same job on a collection of marks. One asks the collection questions. The other reaches inside it.

Then the collection changes its storage, exactly as a real system does when it outgrows its first design. Watch which program survives.

class MarksV1:
    """Version 1: marks kept in a plain row."""

    def __init__(self):
        self.items = []          # deliberately public, which is the trap

    def add(self, mark):
        self.items.append(mark)

    def count(self):
        return len(self.items)

    def total(self):
        return sum(self.items)


class MarksV2:
    """Version 2, six months later. Same ADT, different storage: the marks are
    kept as a running total and a count, because nothing needed the individual
    marks and this uses a fixed amount of memory."""

    def __init__(self):
        self._total = 0
        self._count = 0

    def add(self, mark):
        self._total += mark
        self._count += 1

    def count(self):
        return self._count

    def total(self):
        return self._total


def average_politely(marks):
    """Written against the ADT: it asks."""
    return marks.total() / marks.count()


def average_rudely(marks):
    """Written against the representation: it reaches inside."""
    return sum(marks.items) / len(marks.items)


for cls in (MarksV1, MarksV2):
    m = cls()
    for mark in (78, 65, 91, 40):
        m.add(mark)
    print(cls.__name__)
    print("   polite:", average_politely(m))
    try:
        print("   rude  :", average_rudely(m))
    except AttributeError as e:
        print("   rude  : BROKEN.", e)
MarksV1
   polite: 68.5
   rude  : 68.5
MarksV2
   polite: 68.5
   rude  : BROKEN. 'MarksV2' object has no attribute 'items'

The polite function did not change and did not care. The rude function broke the moment the storage changed, and it broke in a program nobody had touched.

Why this is the argument for the whole idea

Notice three things about that break.

Nothing about the ADT changed. add, count and total behave identically in both versions. The promise was kept perfectly. Only the storage moved.

The rude function was not wrong when it was written. It worked. It was correct against version 1. It became wrong because it depended on something it was never promised.

The break is at a distance. The person who changed the storage did not edit the rude function and may never have seen it. In a real system there are fifty such functions and you find them one by one.

This is why the ADT hides the representation. Not for tidiness. Because every program that depends on your storage is a program you can no longer change your storage without breaking.

The word for it, and the three levels

Hiding the representation is called encapsulation, and the promise it enables is called information hiding.

LevelWhat the user seesWhat they may depend on
The ADToperations and their behaviourall of it
The data structurethe arrangement chosenits costs, not its internals
The implementationthe actual code and fieldsnothing
munotes.in23

Why Hiding the Representation Is the Whole Point

A student is examined on the first two. A programmer is disciplined by the third.

How it is signalled in code

Python has no way to truly forbid access, so it uses a convention that every Python programmer understands, and this book follows it throughout.

A name beginning with a single underscore is private. self._items means this is mine, do not touch it, and if you do, nothing is promised. A name without one is public and is part of the ADT.

Look back at the previous chapter: BagAsList stored self._items with the underscore, and MarksV1 above stored self.items without one. That single character is the difference between a class that can be changed later and a class that cannot, and it was the trap in the experiment.

In C++ and Java the compiler enforces it with private. In C you achieve it by putting the structure's fields in one file and only the function declarations in the header.

What you still need to know

Encapsulation hides the representation from the caller. It does not hide it from the designer, and it must not hide it from you in an examination.

You still have to know that a linked list stores an address beside every value, because that is what makes its insertion cheap and its indexing expensive. Hiding is about what a program may depend on, not about what a student may know.

That is the balance this paper keeps: the ADT for the promise, the implementation for the cost, and both examinable.

Quick revision

  • A program written against the ADT survives a change of representation; one written against the

representation breaks.

  • The break is silent, at a distance, and happens in code nobody edited.
  • The rude function was correct when written; it became wrong because it depended on something never

promised.

  • Hiding the representation is called encapsulation; the guarantee is information hiding.
  • Python signals private with a leading underscore; C++ and Java enforce it with private; C does it

by keeping fields out of the header.

  • A user may depend on the ADT and on the documented costs, never on the internals.
  • Encapsulation hides the representation from the caller, not from the designer or the student.

Test yourself

1. Two functions computed the same average. Why did only one break when the storage changed? One asked the collection through its operations, which did not change. The other reached into a field it was never promised, and that field no longer existed.

2. Was the broken function wrong when it was written? No. It was correct against the first version. It became wrong when the storage changed, because it depended on something outside the ADT.

munotes.in24

Why Hiding the Representation Is the Whole Point

3. Why is this kind of break described as "at a distance"? Because the person who changed the storage never touched the broken function and may not know it exists. The fault appears somewhere else entirely.

4. What is encapsulation, and what does it buy? Hiding a structure's representation behind its operations. It buys the freedom to change the representation later without breaking the programs that use the structure.

5. What does a leading underscore mean in this book's code? That the name is private: part of the implementation, not part of the ADT, and nothing about it is promised to a caller.

6. If the representation is hidden, why must you still learn how a linked list is stored? Because the storage decides the costs, and choosing a structure is a decision about costs. Encapsulation limits what a program may depend on, not what a designer needs to know.

Contents This chapter on its own page

munotes.in25

Chapter Nine

Writing an ADT of Your Own

Syllabus topic Module 1, "Abstract Data Type: Creating user-specific ADT"

In one line

Design an ADT by listing the operations the problem actually needs, deciding what each does in its impossible cases, and only then choosing how to store anything.

The method, in five steps

  1. Name the thing and say what it holds, in one sentence.
  2. List the operations the problem needs. From the problem, not from a textbook list.
  3. For each: what it needs, what it returns, what it does.
  4. For each: what happens when it cannot. Empty, absent, full, duplicate.
  5. Only now, choose a representation, and check each operation's cost against step 2.

Steps 4 and 5 are the ones students skip, and they are the ones that carry the marks.

Worked: a Playlist ADT

The problem, stated as a user would: I want to keep songs in the order I added them, play them one at a time from the front, add a song to the end, and sometimes move a song to the front so it plays next.

Step 1. A Playlist holds songs in a definite order.

Step 2. From that sentence, and nothing else: add to the end, take from the front, move to the front, how many, is it empty, and see what is next without removing it.

Steps 3 and 4 together, which is the ADT:

OperationNeedsReturnsDoesWhen it cannot
Playlist()nothinga playlistcreates it emptynever fails
add(song)a songnothingputs it at the endnever fails
play_next()nothinga songremoves and returns the fronterror if empty
peek()nothinga songreturns the front, leaves it thereerror if empty
bump(song)a songnothingmoves it to the fronterror if absent
size()nothinga numberhow many songsnever fails
is_empty()nothingtrue or falseis the size zeronever fails

Notice step 4 doing real work. play_next on an empty playlist is an error, not a silent None, because a caller that gets None back has no way to tell "empty" from "a song called None" and will carry the mistake somewhere else.

Step 5, the representation. The operations are all at the ends except bump. A row of songs makes add cheap and play_next expensive; the reverse is also possible. This chapter uses a row and measures it, and by chapter 43 you will be able to do better.

The ADT, built and run

class Playlist:
    """Songs in the order they were added. See the ADT table in this chapter."""

    def __init__(self):
        self._songs = []

    def add(self, song):
        self._songs.append(song)

    def play_next(self):
        if self.is_empty():
            raise IndexError("play_next on an empty playlist")
        return self._songs.pop(0)

    def peek(self):
        if self.is_empty():
            raise IndexError("peek on an empty playlist")
        return self._songs[0]

    def bump(self, song):
        if song not in self._songs:
            raise KeyError("not in the playlist: %r" % (song,))
        self._songs.remove(song)
        self._songs.insert(0, song)

    def size(self):
        return len(self._songs)

    def is_empty(self):
        return self.size() == 0


p = Playlist()
for song in ("Tum Hi Ho", "Kal Ho Naa Ho", "Zinda", "Ilahi"):
    p.add(song)

print("size            :", p.size())
print("next up         :", p.peek())
p.bump("Ilahi")
print("after bump      :", p.peek())
print("played          :", p.play_next())
print("now next up     :", p.peek())
print("size            :", p.size())

empty = Playlist()
try:
    empty.play_next()
except IndexError as e:
    print("empty playlist  : refused, as the ADT says:", e)
munotes.in26

Writing an ADT of Your Own

size            : 4
next up         : Tum Hi Ho
after bump      : Ilahi
played          : Ilahi
now next up     : Tum Hi Ho
size            : 3
empty playlist  : refused, as the ADT says: play_next on an empty playlist

Every line of that run is the ADT table being obeyed, including the last one: the impossible case behaves the way the design said it would, because it was decided at step 4 and not improvised in the code.

The same method on the practical's own examples

The practical asks for a Student, a Book or an Employee. The method does not change; only step 2 does, because it comes from the problem.

Student. Holds a roll number, a name and marks per subject. Operations: create, record a mark, total, average, has passed. Impossible cases: a mark outside 0 to 100 is an error; average of no marks is an error, not zero, because zero is a real average and "no marks" is not.

Book. Holds an ISBN, a title, an author, a number of copies. Operations: create, lend, return, copies available. Impossible cases: lending when none are available is an error; returning more copies than were lent is an error.

That last one is the habit this chapter is teaching. A good ADT refuses impossible things loudly, and the place to decide that is the design table, before any code exists.

Quick revision

  • Design an ADT in five steps: what it holds, which operations the problem needs, what each does, what

happens when it cannot, and only then the representation.

  • The operations come from the problem statement, not from a standard list.
  • Every operation needs its impossible case decided: empty, absent, full, duplicate.
  • Returning None for an error is a defect: the caller cannot tell it from a real value.
  • Choose the representation last, and check it against the operations that run most.
  • A good ADT refuses impossible things loudly, and that refusal is part of the design, not of the code.

Test yourself

1. List the five steps of designing an ADT. Say what it holds; list the operations the problem needs; say what each needs, returns and does; decide every impossible case; then choose the representation.

munotes.in27

Writing an ADT of Your Own

2. Why is the representation chosen last? Because it is an answer to the question the ADT states. Choosing storage first fixes the costs before anyone has said which operations matter.

3. play_next on an empty playlist raises an error. Why is returning None worse? Because the caller cannot distinguish "the playlist was empty" from "a song whose value is None", so the mistake is carried on silently instead of being reported where it happened.

4. Design the operations for a Book ADT in a library, with one impossible case each. Book(isbn, title, author, copies); lend(), error when no copies are available; give_back(), error when more are returned than were lent; available(), never fails.

5. Which two of the five steps do students usually skip, and why do they matter? Deciding the impossible cases and choosing the representation last. The first is where most marks and most bugs are; the second is what keeps the design honest about cost.

6. In the Playlist ADT, which single operation is not at an end of the list, and why does that matter? bump, which reaches into the middle. It is the one operation a simple end-based structure cannot do cheaply, so it is the one that decides the representation.

Contents This chapter on its own page

munotes.in28

Chapter Ten

Judging an ADT: Complete, Minimal, and Honest About Cost

Syllabus topic Module 1, "Abstract Data Type: Creating user-specific ADT"

In one line

Judge an ADT by three tests: can it do everything the problem needs, does it contain anything it does not need, and does it tell the truth about what its operations cost.

The three tests

Complete. Every task the problem states can be done through the operations. If a user has to reach inside the structure to do something the problem requires, the ADT is incomplete.

Minimal. No operation is there that could be built from the others without loss. Every extra operation is a promise that must be kept for ever, in every future representation.

Honest about cost. Each operation's cost is documented, so a user can choose. An ADT that hides a linear operation among constant ones invites a program that is accidentally slow.

Two of those pull against each other, and that tension is the interesting part.

A deliberately bad ADT, audited

A "StudentRecord" ADT, as a first attempt might write it:

OperationWhat it does
add_mark(subject, mark)records a mark
get_marks_list()returns the internal list of marks
total()sum of the marks
sum_of_marks()sum of the marks
average()mean of the marks
print_report()prints a formatted report to the screen

Now the audit.

Is it complete? No. There is no way to find the mark for one subject, and no way to correct a mark entered wrongly. Both are things any real user needs, and neither can be done through these operations.

Is it minimal? No, twice over. total and sum_of_marks are the same operation with two names. Every future version must keep both working for ever. average can be built from total and a count, so it is strictly speaking not minimal. But here the right answer is keep it: it is used constantly, and the alternative is every caller writing the same division, and one of them writing it wrongly. Minimality is a guide, not a law.

Is it honest about cost? It cannot be, because get_marks_list hands out the internal list. A caller can now modify the record behind its back, and no cost the ADT documents means anything. This is the worst defect on the list.

And one more, which is a category error rather than a test. print_report does output. An ADT describes what data is and what may be done to it; formatting for a screen is a different job. Fold it in and this structure can never be used in a program that writes to a file, a web page or a printer.

The repaired ADT

OperationNeedsReturnsCost
record(subject, mark)a subject, a mark 0 to 100nothingconstant
correct(subject, mark)a subject, a marknothing; error if not recordedconstant
mark_for(subject)a subjectthe mark; error if absentconstant
subjects()nothinga copy of the subject nameslinear
total()nothingthe sumlinear
average()nothingthe mean; error if no markslinear
munotes.in29

Judging an ADT: Complete, Minimal, and Honest About Cost

Three changes carry all the value: the duplicate is gone, the two missing operations are in, and subjects() returns a copy, so a caller cannot reach into the record. The cost column is now part of the promise.

Handing out the insides, demonstrated

The defect above is worth seeing happen, because it looks harmless.

class LeakyRecord:
    def __init__(self):
        self._marks = {}

    def record(self, subject, mark):
        self._marks[subject] = mark

    def get_marks(self):
        return self._marks            # the defect: the real thing


class SafeRecord:
    def __init__(self):
        self._marks = {}

    def record(self, subject, mark):
        if not 0 <= mark <= 100:
            raise ValueError("a mark must be between 0 and 100")
        self._marks[subject] = mark

    def get_marks(self):
        return dict(self._marks)      # a copy


for cls in (LeakyRecord, SafeRecord):
    r = cls()
    r.record("Data Structures", 78)
    outside = r.get_marks()
    outside["Data Structures"] = 100          # a caller being careless
    outside["Invented Subject"] = 99
    print("%-12s after an outsider edited what it was handed: %s"
          % (cls.__name__, r.get_marks()))
LeakyRecord  after an outsider edited what it was handed: {'Data Structures': 100, 'Invented Subject': 99}
SafeRecord   after an outsider edited what it was handed: {'Data Structures': 78}

The leaky record now holds a mark of 100 that nobody recorded and a subject that does not exist, and no operation of the ADT was used to do it. Every guarantee the structure made is gone, including the range check, which record applies and a direct edit walks straight past.

The tension, said plainly

Minimal and convenient are in tension, and an examiner may ask you to argue it.

Add average and you have broken minimality but saved every caller a division and prevented a class of mistakes. Refuse it and the ADT is pure and everybody writes the same three lines.

The rule worth stating: add an operation when it is used often enough that callers would otherwise repeat it, and when the structure can implement it better than a caller could. average passes both. sum_of_marks beside total passes neither.

Quick revision

  • Three tests: complete (does everything the problem needs), minimal (nothing redundant), honest about

cost (every operation's cost documented).

  • Incomplete shows up as a user reaching inside the structure to do a required job.
  • Two names for one operation is the clearest failure of minimality.
  • Returning the internal container destroys every guarantee, including validation, and is the worst of

the common defects; return a copy.

  • Printing and formatting are a different job and do not belong in the ADT.
  • Minimality is a guide: add an operation when callers would otherwise repeat it and the structure can
munotes.in30

Judging an ADT: Complete, Minimal, and Honest About Cost

do it better.

Test yourself

1. State the three tests for judging an ADT. Complete, minimal, and honest about cost.

2. An ADT hands back its internal list. Name two guarantees that breaks. Any validation it performs, since a direct edit bypasses it; and every documented cost, since the structure can be changed behind its back by code it does not control.

3. Why is print_report wrong in a StudentRecord ADT? Formatting output is a different responsibility. Including it ties the structure to one output medium and prevents its use with files, pages or printers.

4. average can be built from total and a count. Should it be in the ADT? Yes. Minimality is a guide, not a law: it is used constantly and the structure implements it once and correctly, instead of every caller repeating the division.

5. total and sum_of_marks both exist. Which test fails, and why does it matter later? Minimality. Both must be kept working in every future version of the structure, for no benefit.

6. How does incompleteness usually reveal itself in practice? A user has to reach inside the structure to do something the problem requires, because no operation offers it.

Contents This chapter on its own page

munotes.in31

Chapter Eleven

The Array: What It Really Is in Memory

Syllabus topic Computer Science Practical 3, Module 2, "Compare static (array) vs dynamic (linked) approaches"

In one line

An array is a run of equal-sized cells side by side in memory, which is why the machine can compute the address of the nth item instead of looking for it.

The arrangement

Ask for an array of ten integers and the machine gives you one unbroken block of memory, divided into ten equal cells. Nothing separates them; nothing links them. They are simply next to each other.

That one property, equal cells, side by side, is the whole of the array, and everything it is good and bad at follows from it.

Why indexing is instant

Suppose the block starts at address 1000 and each cell is 4 bytes.

address of item 0 = 1000 + (0 x 4) = 1000

address of item 1 = 1000 + (1 x 4) = 1004

address of item 2 = 1000 + (2 x 4) = 1008

address of item i = 1000 + (i x 4)

So, in general:

address of item i = base + (i x size of one cell)

That is one multiplication and one addition. It does not depend on i, and it does not depend on how many items the array holds. Reaching item 9,999 of ten thousand costs exactly what reaching item 0 costs.

This is what O(1) access means, and it is the array's one great gift. It is also why array indices start at 0 in most languages: item 0 is at base + 0, so the arithmetic has no correction in it.

Seen in a real run:

BASE = 1000
CELL = 4

print("index | address")
for i in (0, 1, 2, 9, 9999):
    print("%5d | %d" % (i, BASE + i * CELL))

print()
print("the cost of the arithmetic does not depend on i:")
print("  item     0 needs 1 multiply and 1 add")
print("  item 9,999 needs 1 multiply and 1 add")
index | address
    0 | 1000
    1 | 1004
    2 | 1008
    9 | 1036
 9999 | 40996

the cost of the arithmetic does not depend on i:
  item     0 needs 1 multiply and 1 add
  item 9,999 needs 1 multiply and 1 add

The three consequences

Everything about arrays follows from equal cells side by side.

1. Access is constant. Shown above.

2. The cells must be the same size. Otherwise the multiplication is meaningless. This is why a C array holds one type only. Python's list appears to hold anything, because what it really stores is a row of equal-sized addresses, each pointing at the actual value elsewhere. The row is still uniform; the values are not in it.

3. The block must be contiguous, so the size is fixed at creation. You cannot extend a block if something else is sitting immediately after it in memory. That single fact is the whole of the next chapter.

munotes.in32

The Array: What It Really Is in Memory

Row major order, and why a 2D array is really 1D

A two dimensional array is stored as one row of cells too, one row of the table after another. This is row major order, and it is examinable because the address formula follows from it.

For an array with C columns, starting at base, with cells of size w:

address of A[i][j] = base + ((i x C) + j) x w

BASE, COLS, CELL = 2000, 5, 4

print("A is 3 rows x 5 columns, base 2000, 4 bytes a cell")
print()
print("element | offset in cells | address")
for i, j in [(0, 0), (0, 4), (1, 0), (2, 3)]:
    cells = i * COLS + j
    print("A[%d][%d]  | %14d | %d" % (i, j, cells, BASE + cells * CELL))
A is 3 rows x 5 columns, base 2000, 4 bytes a cell

element | offset in cells | address
A[0][0]  |              0 | 2000
A[0][4]  |              4 | 2016
A[1][0]  |              5 | 2020
A[2][3]  |             13 | 2052

Notice A[0][4] and A[1][0] are next to each other in memory, four bytes apart. The row boundary exists only in your head; the machine sees one long run of cells.

Some languages, notably Fortran, store column major instead, down the columns first. The formula then becomes base + ((j x R) + i) x w. An examination question naming a language is asking which of the two you use.

What an array promises, as an ADT

OperationCost
get(i)constant
set(i, value)constant
length()constant
insert in the middlelinear
delete from the middlelinear
grownot possible; a new array must be made

The top three are why arrays are everywhere. The bottom three are why this paper has ten more structures in it.

Quick revision

  • An array is a contiguous block of equal-sized cells.
  • Address of item i is base + (i x cell size): one multiply, one add, independent of i and of the

array's length. That is O(1) access.

  • Indices start at 0 so the arithmetic needs no correction.
  • Cells must be equal in size, so a C array holds one type; a Python list stores equal-sized addresses

pointing elsewhere.

  • The block is contiguous, so the size is fixed when it is created.
  • 2D arrays are stored row major: base + ((i x C) + j) x w. Fortran uses column major.
  • A row boundary is not in memory; the cells of one row run straight into the next.
munotes.in33

The Array: What It Really Is in Memory

Test yourself

1. Write the address formula for the ith item of a one dimensional array and say why it is O(1). base + (i x w), where w is the cell size. It is one multiplication and one addition whatever i is and however long the array is, so the cost does not grow with the data.

2. Why must all cells be the same size? Because the address is computed by multiplying the index by the cell size. Unequal cells make that multiplication meaningless.

3. A Python list can hold an integer and a string at once. How is that consistent with equal cells? The list stores equal-sized addresses, not the values. The values live elsewhere; the row itself is uniform.

4. Give the row major address of A[2][3] where A has 5 columns, base 2000 and 4 byte cells. Offset in cells is (2 x 5) + 3 = 13, so the address is 2000 + 13 x 4 = 2052.

5. Why can an array not grow? Because its cells must be contiguous, and the memory immediately after the block may already be in use. Growing means allocating a new block and copying.

6. Which single property of the array causes both its greatest strength and its greatest weakness? Contiguity. It makes the address computable, which gives O(1) access, and it fixes the size at creation, which makes growth and middle insertion expensive.

Contents This chapter on its own page

munotes.in34

Chapter Twelve

Where the Array Stops: Insertion, Deletion and Growth

Syllabus topic Computer Science Practical 3, Module 2, "Compare static (array) vs dynamic (linked) approaches"

In one line

An array pays for its instant access with three expensive operations: inserting in the middle, deleting from the middle, and growing at all.

Cost one: inserting in the middle

To put a value at position k of an array that already holds n items, every item from k onward has to move up one place, and it has to be done from the back or items overwrite each other.

def insert_at(row, k, value):
    """Insert into a fixed array by shifting. Returns the number of moves."""
    row.append(None)                     # make room at the end
    moves = 0
    for i in range(len(row) - 1, k, -1):
        row[i] = row[i - 1]
        moves += 1
    row[k] = value
    return moves


for n in (1000, 2000, 4000, 8000):
    data = list(range(n))
    at_front = insert_at(list(data), 0, -1)
    at_middle = insert_at(list(data), n // 2, -1)
    at_end = insert_at(list(data), n, -1)
    print("n = %5d   insert at front: %5d moves   middle: %5d   end: %d"
          % (n, at_front, at_middle, at_end))
n =  1000   insert at front:  1000 moves   middle:   500   end: 0
n =  2000   insert at front:  2000 moves   middle:  1000   end: 0
n =  4000   insert at front:  4000 moves   middle:  2000   end: 0
n =  8000   insert at front:  8000 moves   middle:  4000   end: 0

Read the columns down. Inserting at the front moves every item, and doubling the data doubles the work: that is O(n). Inserting in the middle moves half of them, which is also O(n), because a half of something growing is still growing. Inserting at the end moves nothing at all.

So "inserting into an array is expensive" is too crude a statement to be useful. Inserting at the end is free. It is inserting anywhere else that costs, and it costs in proportion to how much of the array sits after the insertion point.

Cost two: deleting from the middle

Deletion is the same problem in reverse: the hole has to be closed, so everything after it moves down.

def delete_at(row, k):
    """Delete from a fixed array by shifting. Returns the number of moves."""
    moves = 0
    for i in range(k, len(row) - 1):
        row[i] = row[i + 1]
        moves += 1
    row.pop()
    return moves


for n in (1000, 2000, 4000, 8000):
    data = list(range(n))
    print("n = %5d   delete at front: %5d moves   middle: %5d   end: %d"
          % (n, delete_at(list(data), 0), delete_at(list(data), n // 2),
             delete_at(list(data), n - 1)))
n =  1000   delete at front:   999 moves   middle:   499   end: 0
n =  2000   delete at front:  1999 moves   middle:   999   end: 0
n =  4000   delete at front:  3999 moves   middle:  1999   end: 0
n =  8000   delete at front:  7999 moves   middle:  3999   end: 0

The same shape. A deletion at the front of an array of eight thousand items moves 7,999 of them to close a single hole.

munotes.in35

Where the Array Stops: Insertion, Deletion and Growth

Cost three: growth, and the trick that hides it

A true static array cannot grow: its block is fixed. To hold one more item you allocate a bigger block and copy everything across.

Done naively, one cell at a time, that is ruinous. Python does something cleverer, and it is worth seeing, because it explains a puzzling thing every Python programmer has noticed.

import sys

print("appending one at a time, watching the underlying block:")
print()
row = []
last = sys.getsizeof(row)
grows = 0
for i in range(1000):
    row.append(i)
    now = sys.getsizeof(row)
    if now != last:
        grows += 1
        if grows <= 8:
            print("  at %4d items the block was replaced" % len(row))
        last = now

print()
print("items appended      : 1000")
print("times the block grew:", grows)
print("copies per append, on average: about", round(1000 / grows / 100) / 10 if grows else 0)
appending one at a time, watching the underlying block:

  at    1 items the block was replaced
  at    5 items the block was replaced
  at    9 items the block was replaced
  at   17 items the block was replaced
  at   25 items the block was replaced
  at   33 items the block was replaced
  at   41 items the block was replaced
  at   53 items the block was replaced

items appended      : 1000
times the block grew: 28
copies per append, on average: about 0.0

A thousand appends caused the block to be replaced thirty times, not a thousand. The list does not grow by one; it grows by a proportion, so the gaps between growths get wider as it gets bigger.

That is why appending is described as amortised constant time: any one append might be expensive, but the expensive ones are rare enough that the average over many appends is constant. Chapter 107 uses the same idea again for hash tables.

The three costs, together

OperationArrayWhy
Read or write item iO(1)the address is computed
Insert or delete at the endO(1) amortisednothing after it to move
Insert or delete at the frontO(n)everything moves
Insert or delete in the middleO(n)half of it moves
Search, unsortedO(n)no shortcut
Search, sortedO(log n)binary search
GrowO(n) for that one appenda new block, everything copied

Look at that table as a shopping list for the next structure. What would be better?

Something where inserting and deleting cost nothing wherever you are, because nothing has to move. The price will be the top row: you will lose the computed address, and reaching item i will mean walking to it.

munotes.in36

Where the Array Stops: Insertion, Deletion and Growth

That trade is the linked list, and it is the next chapter.

Quick revision

  • Inserting at position k moves every item after k: O(n) at the front, O(n) in the middle, free at the

end.

  • Deleting closes the hole the same way, with the same costs.
  • A static array cannot grow; growth means a new block and a full copy.
  • Python's list grows by a proportion, not by one, so 1,000 appends replaced the block about 30 times.
  • That makes appending amortised constant: rare expensive operations averaged over many cheap ones.
  • "Insertion into an array is expensive" is too crude: at the end it is free, and the cost elsewhere is

proportional to how much of the array sits after the insertion point.

  • The structure that fixes this must give up the computed address, which is the linked list's bargain.

Test yourself

1. How many items move when you insert at the front of an array of 8,000? All 8,000, and the moves must be done from the back forward or items overwrite one another.

2. Why is inserting at the end of an array cheap when inserting at the front is not? Because the cost is the number of items sitting after the insertion point. At the end there are none.

3. What does amortised constant time mean, and why does appending to a Python list qualify? That occasional expensive operations are rare enough for the average over many to be constant. The list grows by a proportion, so 1,000 appends caused about 30 block replacements, not 1,000.

4. A deletion from the front of an array of 8,000 items reported 7,999 moves. Why not 8,000? Because the item being deleted does not move; the 7,999 items after it each shift down one place.

5. Which array operation is O(log n), and what does it require? Search, when the array is sorted, using binary search. It requires the data to be in order.

6. What must the next structure give up in exchange for cheap insertion anywhere? The computed address, and so constant-time access to item i: reaching an item will mean walking to it.

Contents This chapter on its own page

munotes.in37

Chapter Thirteen

The Linked List: The Node, the Chain and the Head

Syllabus topic Module 1, "Linked Structures: ADT for linked list"

In one line

A linked list stores each value in its own node together with the address of the next node, so the items can sit anywhere in memory and still be in order.

The idea

The array's problem was contiguity: the cells had to be side by side, so nothing could be inserted without moving things.

The linked list gives that up completely. Each item lives in its own small parcel called a node, and the node holds two things:

PartHolds
datathe value itself
nextthe address of the following node

The nodes can be scattered anywhere in memory. What puts them in order is not where they sit but what each one points at.

One more thing is needed, and it is the thing students forget: something has to say where the chain starts. That is the head, a single variable holding the address of the first node. Lose the head and the whole list is unreachable, however intact the nodes are.

The last node points at nothing. In Python that is None; in C it is NULL. This is how the machine knows the chain has ended, and a list whose last node points at something by mistake never ends.

The node, built

class Node:
    """One link of the chain: a value, and where the next one is."""

    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node

    def __repr__(self):
        return "Node(%r)" % (self.data,)


# Three nodes, built from the back so each knows the one after it.
third = Node("Zinda")
second = Node("Ilahi", third)
first = Node("Tum Hi Ho", second)

head = first          # the one variable that makes the chain reachable

print("head          :", head)
print("head.next     :", head.next)
print("head.next.next:", head.next.next)
print("the end       :", head.next.next.next)
head          : Node('Tum Hi Ho')
head.next     : Node('Ilahi')
head.next.next: Node('Zinda')
the end       : None

Three separate objects, created in no particular place in memory, and yet they have an order. The order lives in the next fields.

Drawing it

The standard picture, and the one to draw in an examination:

head --. [ Tum Hi Ho | -]--. [ Ilahi | -]--. [ Zinda | X ]

Each box is a node with two compartments, data on the left and next on the right. The arrow from the right compartment shows what it points at. The cross in the last one is the null that ends the chain.

Two mistakes to avoid when drawing it: the head is not a node, it is an arrow into the first node; and the last next must be shown as null, not left blank, because that is a real value.

Why insertion is now cheap

This is the payoff, and the whole reason for the structure. To put a new node between the first and the second, nothing moves. Two assignments do it.

munotes.in38

The Linked List: The Node, the Chain and the Head

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


def as_row(head):
    out, walk = [], head
    while walk is not None:
        out.append(walk.data)
        walk = walk.next
    return out


head = Node("A", Node("B", Node("C")))
print("before:", as_row(head))

# Insert "NEW" between A and B. Nothing is moved or copied.
new = Node("NEW")
new.next = head.next          # 1. the new node points at B
head.next = new               # 2. A points at the new node
print("after :", as_row(head))
print("nodes moved in memory:", 0)
before: ['A', 'B', 'C']
after : ['A', 'NEW', 'B', 'C']
nodes moved in memory: 0

Two assignments, and it would be two assignments if the list held ten million nodes. Compare the array, which moved 8,000 items to do the same job.

The order of those two assignments is not free. Reverse them and head.next is overwritten before the new node has recorded it, and the rest of the list is lost. Chapter 20 does that deliberately and shows the wreckage.

What it costs

Nothing is free. The linked list pays in three places.

Memory. Every node carries an address beside its value. On a 64 bit machine that is 8 extra bytes per item, so a list of a million small values carries 8 MB of pure overhead.

Access. There is no address arithmetic any more. To reach item 500 you must start at the head and follow 500 links. Reaching item i is O(n), where the array was O(1).

Locality. The nodes are scattered, so the machine's cache, which fetches neighbouring memory in blocks, helps far less than it does with an array. This is why an array often beats a linked list in practice even when the counting says otherwise, and chapter 22 measures exactly that.

Quick revision

  • A node holds a value and the address of the next node.
  • The nodes may sit anywhere; the order lives in the next fields, not in the positions.
  • The head is a variable holding the address of the first node. It is not a node. Lose it and the list

is unreachable.

  • The last node's next is null (None in Python, NULL in C), and that is what ends the chain.
  • Insertion needs two assignments and moves nothing, whatever the length of the list.
  • The order of those assignments matters: point the new node forward first.
  • It costs one address per item in memory, O(n) access instead of O(1), and poorer cache behaviour.

Test yourself

1. What two parts does a node have? The data, and next: the address of the following node.

munotes.in39

The Linked List: The Node, the Chain and the Head

2. Is the head a node? What happens if it is lost? No, it is a variable holding the address of the first node. If it is lost the entire list becomes unreachable even though every node is intact.

3. What marks the end of a singly linked list, and why can it not be left blank? The last node's next field holds null. It is a real value that the traversal tests against; a blank or uninitialised field would be followed into nonsense.

4. How many assignments insert a node after a known node, and how many items move? Two assignments, and nothing moves, whatever the list's length.

5. Give the three costs of a linked list against an array. One address of memory per item; O(n) access to item i instead of O(1); and scattered nodes, which use the cache poorly.

6. Draw a three node list with the head and the terminator marked. An arrow labelled head into the first box; each box split into data and next; arrows from each next to the following box; and the last next shown as a cross or null rather than left empty.

Contents This chapter on its own page

munotes.in40

Chapter Fourteen

The Linked List ADT, and Building an Empty One

Syllabus topic Module 1, "Linked Structures: ADT for linked list"

In one line

The linked list ADT promises the usual collection operations, and its whole implementation difficulty is concentrated in two cases: the empty list and the first node.

The ADT, as an answer to the examination question

A Linked List holds a sequence of values in a definite order, of any length, with no fixed maximum.

OperationNeedsReturnsDoesWhen it cannot
LinkedList()nothinga listcreates it emptynever fails
is_empty()nothingtrue or falseis there no first nodenever fails
length()nothinga numberhow many nodesnever fails
prepend(v)a valuenothingadds a node at the frontnever fails
append(v)a valuenothingadds a node at the endnever fails
search(v)a valuetrue or falseis v presentnever fails
insert_at(k, v)position, valuenothinginserts at position kerror if k is out of range
delete(v)a valuenothingremoves the first node holding verror if absent
traverse()nothingthe values in ordervisits every node oncenever fails

That table is the answer. Note what it does not contain: no Node, no next, no head. Those belong to the implementation, and an answer that starts with them has answered the wrong question.

The empty list, which is the whole problem

An empty list is head = None. That sounds trivial and is the source of most of the bugs in this chapter's subject, because every operation has to work when there is no first node.

Three questions to ask of any linked list operation you write:

  1. What does it do when the list is empty?
  2. What does it do when the list has exactly one node?
  3. Does it change the head, and if so, has it actually assigned the head?

The third is the subtle one. In C you need a pointer to a pointer, or to return the new head; in Python the head is a field of the list object and must be written back to self.head.

The empty list and the one node list, run

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


class LinkedList:
    """A singly linked list. head is None exactly when the list is empty."""

    def __init__(self):
        self.head = None

    def is_empty(self):
        return self.head is None

    def length(self):
        n, walk = 0, self.head
        while walk is not None:
            n += 1
            walk = walk.next
        return n

    def prepend(self, value):
        self.head = Node(value, self.head)     # works when head is None too

    def to_list(self):
        out, walk = [], self.head
        while walk is not None:
            out.append(walk.data)
            walk = walk.next
        return out


empty = LinkedList()
print("empty: is_empty =", empty.is_empty(), "| length =", empty.length(),
      "| contents =", empty.to_list())

one = LinkedList()
one.prepend("only")
print("one  : is_empty =", one.is_empty(), "| length =", one.length(),
      "| contents =", one.to_list())

three = LinkedList()
for value in ("C", "B", "A"):
    three.prepend(value)
print("three: is_empty =", three.is_empty(), "| length =", three.length(),
      "| contents =", three.to_list())
munotes.in41

The Linked List ADT, and Building an Empty One

empty: is_empty = True | length = 0 | contents = []
one  : is_empty = False | length = 1 | contents = ['only']
three: is_empty = False | length = 3 | contents = ['A', 'B', 'C']

Look at prepend and see why it needs no special case at all:

self.head = Node(value, self.head)

When the list is empty, self.head is None, so the new node's next becomes None, which is exactly right for a single node list. When the list is not empty, the new node points at the old first node, which is also exactly right. One line, no if.

That is the mark of a well written linked list operation: the empty case falls out of the general case rather than being bolted on beside it. Where it cannot, the special case is written first and deliberately.

Why length is O(n) here, and what to do about it

length above walks the entire list counting nodes. For an array, length is stored and is O(1).

There are two honest answers, and an examiner may ask for both.

Walk and count. Simple, nothing to keep in step, O(n).

Keep a counter in the list object, incremented on every insertion and decremented on every deletion. O(1) to read, but now every operation that changes the list must remember to update it, and the day one forgets, the structure lies about its own size.

This book keeps the counter from chapter 20 onwards and says so where it does, because by then the list is used inside stacks and queues where length is asked for constantly.

Quick revision

  • The ADT is the operations and their behaviour: create, is_empty, length, prepend, append, search,

insert_at, delete, traverse. It names no node, no next and no head.

  • The empty list is head = None, and every operation must work on it.
  • Ask three questions of every operation: what if empty, what if one node, does it change the head.
  • self.head = Node(value, self.head) prepends correctly whether or not the list is empty, with no

special case.

  • A good linked list operation makes the empty case fall out of the general case.
  • length by walking is O(n); a stored counter makes it O(1) but must be maintained by every operation

that changes the list.

Test yourself

1. Write the linked list ADT. What must your answer not mention? The operations with what each needs, returns and does. It must not mention nodes, next pointers or the head, which are implementation.

munotes.in42

The Linked List ADT, and Building an Empty One

2. How is an empty linked list represented, and why is it the hard case? head is null. It is hard because every operation must behave correctly when there is no first node, and operations that change the head must write the new head back.

3. Why does self.head = Node(value, self.head) need no test for emptiness? Because when head is None the new node's next becomes None, which is correct for a one node list; and when head is a node the new node points at it, which is correct too.

4. State the three questions to ask of any linked list operation. What does it do when the list is empty; what when it has exactly one node; and does it change the head, and if so has it assigned it.

5. Why is length O(n) on this list but O(1) on an array? The array stores its length. The list must walk from the head to the end, counting, because nothing records how many nodes there are.

6. Give the cost of keeping a length counter in the list object. Reading becomes O(1), but every insertion and deletion must update it, and any operation that forgets leaves the structure reporting a size it does not have.

Contents This chapter on its own page

munotes.in43

Chapter Fifteen

Traversing a Singly Linked List

Syllabus topic Module 1, "Linked Structures: Singly Linked List-Traversing"

In one line

To traverse is to visit every node once, in order, using a walking pointer that starts at the head and stops when it becomes null.

The pattern

Three lines, and they never change:

walk = head

while walk is not null: do something with walk.data; walk = walk.next

Every linked list algorithm in this paper is that shape with something inserted. Searching adds a test. Counting adds a counter. Printing adds a print. Learn the shape once.

Traversal, run

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


def build(values):
    """Build a list from a row of values, front to back."""
    head = None
    for value in reversed(values):
        head = Node(value, head)
    return head


def traverse(head):
    """Visit every node once, in order. Returns the values and the steps taken."""
    values, steps, walk = [], 0, head
    while walk is not None:
        values.append(walk.data)
        steps += 1
        walk = walk.next
    return values, steps


for values in ([], ["only"], ["A", "B", "C", "D"]):
    got, steps = traverse(build(values))
    print("%-18s visited %s in %d step(s)" % (str(values) + ":", got, steps))
[]:                visited [] in 0 step(s)
['only']:          visited ['only'] in 1 step(s)
['A', 'B', 'C', 'D']: visited ['A', 'B', 'C', 'D'] in 4 step(s)

The empty list takes zero steps and needs no special case: the while simply never runs. That is the test of a correctly written traversal.

The three ways it is written wrongly

Each of these is a real mistake students make, and each fails differently. They are run here rather than described, with the failure caught.

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


head = Node("A", Node("B", Node("C")))


def wrong_one(head):
    """Forgetting to advance the pointer. Runs for ever, so it is capped here."""
    walk, seen = head, 0
    while walk is not None and seen < 5:
        seen += 1
        # walk = walk.next     <- the missing line
    return "visited the first node %d times and never moved" % seen


def wrong_two(head):
    """Testing the data instead of the node. Stops at the first falsy value."""
    walk, out = head, []
    while walk is not None and walk.data:
        out.append(walk.data)
        walk = walk.next
    return out


def wrong_three(head):
    """Advancing before using. Skips the first node."""
    walk, out = head, []
    while walk is not None:
        walk = walk.next
        if walk is not None:
            out.append(walk.data)
    return out


print("1. no advance   :", wrong_one(head))
print("2. testing data :", wrong_two(Node("A", Node("", Node("C")))))
print("3. advance first:", wrong_three(head))
1. no advance   : visited the first node 5 times and never moved
2. testing data : ['A']
3. advance first: ['B', 'C']

Number 1 never ends. Forgetting walk = walk.next is the commonest mistake in this subject, and it does not crash: it hangs. The program above caps it at five visits only so that the point can be printed at all. Written as a student writes it, with no cap, the loop tests the same node for ever and the program has to be killed. There is no error message and no line number, which is what makes it harder to find than a crash.

munotes.in44

Traversing a Singly Linked List

Number 2 stops early. Testing while walk.data instead of while walk is not None works on most data and then silently stops at the first empty string, zero or None in the list. Test the node, never the value inside it.

Number 3 skips the first node. Advancing before using visits nodes 2 to n. It is easy to write when converting a for loop out of habit.

What traversal costs

One visit per node, so n steps for n nodes: O(n), and there is no faster way. A linked list has no shortcut to the middle, so anything that needs to see the whole list costs a full walk.

That is worth stating plainly because it is the source of several later results. length is O(n). append without a tail pointer is O(n). Reaching item i is O(n). All three are the same walk.

Traversing with an index, and why it is a trap

Students who learned arrays first sometimes write this:

for i in range(length(list)): do something with item_at(i)

If item_at(i) walks from the head, that loop is not O(n). It is O(n squared): the first item costs 1 step, the second 2, the last n, and the total is about n squared over 2. On a list of 10,000 that is about 50 million steps to do what a single traversal does in 10,000.

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


def build(n):
    head = None
    for value in reversed(range(n)):
        head = Node(value, head)
    return head


def item_at(head, i):
    """Walk to item i. Returns (value, steps)."""
    walk, steps = head, 0
    while i > 0:
        walk = walk.next
        steps += 1
        i -= 1
    return walk.data, steps


def by_traversal(head):
    steps, walk = 0, head
    while walk is not None:
        steps += 1
        walk = walk.next
    return steps


def by_index(head, n):
    total = 0
    for i in range(n):
        _, steps = item_at(head, i)
        total += steps + 1
    return total


for n in (250, 500, 1000, 2000):
    head = build(n)
    print("n = %4d   one traversal: %5d steps   indexed loop: %8d steps"
          % (n, by_traversal(head), by_index(head, n)))
n =  250   one traversal:   250 steps   indexed loop:    31375 steps
n =  500   one traversal:   500 steps   indexed loop:   125250 steps
n = 1000   one traversal:  1000 steps   indexed loop:   500500 steps
n = 2000   one traversal:  2000 steps   indexed loop:  2001000 steps
munotes.in45

Traversing a Singly Linked List

Read the right column down: the data doubles and the work goes up four times. That is the signature of O(n squared), and it comes from writing an array's loop over a linked structure.

Quick revision

  • Traversal: walk = head, then while walk is not null, use walk.data and walk = walk.next.
  • Every linked list algorithm is that shape with something added.
  • The empty list needs no special case: the loop never runs.
  • Three classic errors: forgetting to advance (hangs for ever), testing walk.data instead of the node

(stops at the first falsy value), advancing before using (skips the first node).

  • Traversal is O(n) and there is no shortcut, which is why length, append without a tail, and access by

index are all O(n).

  • Looping by index over a linked list is O(n squared): doubling the data quadrupled the work in the run

above.

Test yourself

1. Write the traversal pattern in three lines. walk = head; while walk is not null: use walk.data; walk = walk.next.

2. Why does a correct traversal need no special case for the empty list? Because head is null, so the loop condition fails immediately and the body never runs.

3. What happens if walk = walk.next is omitted, and why is that worse than a crash? The loop never ends. It hangs rather than failing, so there is no error message and no line number to look at.

4. Why test walk is not None rather than walk.data? Because a node may legitimately hold an empty string, a zero or None, and testing the data would stop the traversal at that node instead of at the end of the list.

5. A loop calls item_at(i) for i from 0 to n-1. What is its cost, and why? O(n squared), because each item_at walks from the head: 1 step, then 2, up to n, which totals about n squared over 2.

6. In the run above, n doubled from 1000 to 2000. What happened to the indexed loop's steps? They went from about 500,000 to about 2,000,000, four times as many, which is the signature of O(n squared).

Contents This chapter on its own page

munotes.in46

Chapter Sixteen

Searching a Singly Linked List

Syllabus topic Module 1, "Linked Structures: Singly Linked List-Traversing, Searching"

In one line

Searching a linked list means traversing until the value is found or the chain ends, which is O(n), and sorting the list does not improve that.

The search

A traversal with one test added.

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


def build(values):
    head = None
    for value in reversed(values):
        head = Node(value, head)
    return head


def search(head, target):
    """Is target present? Returns (found, comparisons)."""
    walk, comparisons = head, 0
    while walk is not None:
        comparisons += 1
        if walk.data == target:
            return True, comparisons
        walk = walk.next
    return False, comparisons


head = build(["A", "B", "C", "D", "E"])
for target in ("A", "C", "E", "Z"):
    found, comparisons = search(head, target)
    print("search %-3s -> found=%-5s after %d comparison(s)"
          % (target, found, comparisons))
search A   -> found=True  after 1 comparison(s)
search C   -> found=True  after 3 comparison(s)
search E   -> found=True  after 5 comparison(s)
search Z   -> found=False after 5 comparison(s)

Three cases, and an examiner will ask for all three.

Best case: the value is first. One comparison. O(1). Worst case: the value is last, or absent. n comparisons. O(n). Average case, when present: about n/2 comparisons, which is still O(n).

An unsuccessful search always costs the full n, because the only way to be sure something is absent is to look at everything.

The negative result: sorting does not help

Sort an array and binary search becomes available: look at the middle, discard half, repeat. Sixteen comparisons for 40,000 items instead of 40,000.

Sort a linked list and you get almost none of that, because binary search needs one thing the linked list cannot provide: the ability to jump to the middle.

To look at the middle element of a linked list you must walk to it, which costs n/2 steps. Then to look at the middle of the remaining half, you must walk again. The jumping is the expensive part, and it is exactly what the array gave you for free.

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


def build(values):
    head = None
    for value in reversed(values):
        head = Node(value, head)
    return head


def linear_search_sorted(head, target):
    """On a SORTED list you may stop early when you pass the target."""
    walk, comparisons = head, 0
    while walk is not None:
        comparisons += 1
        if walk.data == target:
            return True, comparisons
        if walk.data > target:
            return False, comparisons      # passed it; it cannot be later
        walk = walk.next
    return False, comparisons


def binary_search_on_links(head, target):
    """Binary search over a linked list, counting the WALKING as well."""
    def length_and_steps(node):
        n, steps = 0, 0
        while node is not None:
            n += 1
            steps += 1
            node = node.next
        return n, steps

    n, steps = length_and_steps(head)
    lo, hi = 0, n - 1
    comparisons = 0
    while lo <= hi:
        mid = (lo + hi) // 2
        walk = head                        # no way in but from the front
        for _ in range(mid):
            walk = walk.next
            steps += 1
        comparisons += 1
        if walk.data == target:
            return True, comparisons, steps
        if walk.data < target:
            lo = mid + 1
        else:
            hi = mid - 1
    return False, comparisons, steps


for n in (1000, 2000, 4000):
    data = list(range(0, n * 2, 2))        # sorted, even numbers
    head = build(data)
    missing = n * 2 + 1
    _, linear_comparisons = linear_search_sorted(head, missing)
    _, bin_comparisons, bin_steps = binary_search_on_links(head, missing)
    print("n = %4d | sorted linear: %5d comparisons | binary: %2d comparisons "
          "but %6d pointer steps" % (n, linear_comparisons, bin_comparisons, bin_steps))
munotes.in47

Searching a Singly Linked List

n = 1000 | sorted linear:  1000 comparisons | binary: 10 comparisons but   9996 pointer steps
n = 2000 | sorted linear:  2000 comparisons | binary: 11 comparisons but  21995 pointer steps
n = 4000 | sorted linear:  4000 comparisons | binary: 12 comparisons but  47994 pointer steps

The comparisons did drop to about 10, exactly as binary search promises. But the pointer steps went up, to ten times the length of the list, because every one of those ten comparisons had to walk from the head to find its element.

So binary search on a linked list is not merely no better; it is considerably worse. The structure decides which algorithms are available to it, and that is the real lesson.

What this tells you about choosing

You needUse
frequent search, rare changesorted array, binary search, O(log n)
frequent change, rare searchlinked list, O(n) search but O(1) insert
frequent search and frequent changea tree, chapter 64, which does both in O(log n)
frequent search, no order neededa hash table, chapter 100, O(1)

The third row is why Module 2 exists. Module 1 gives you two structures, each good at one thing and bad at the other, and the binary search tree is the structure that refuses to choose.

Quick revision

  • Search is a traversal with a comparison: O(n).
  • Best case 1 comparison, worst case n, average about n/2 when present; an unsuccessful search always

costs n.

  • On a sorted list you may stop early once you pass the target, which halves the average unsuccessful

search but does not change the order of growth.

  • Binary search needs to jump to the middle; a linked list has no way in except from the front.
  • Measured, binary search over links used about 10 comparisons but ten times the list length in pointer

steps, so it is worse than the plain walk.

munotes.in48

Searching a Singly Linked List

  • The structure decides which algorithms are available to it.

Test yourself

1. Give the best, worst and average comparisons for searching a linked list of n nodes. Best 1, worst n, average about n/2 when the value is present. An unsuccessful search always costs n.

2. What early exit does a sorted linked list allow, and does it change the complexity? You may stop as soon as a value greater than the target is reached. It improves the average unsuccessful search but the complexity stays O(n).

3. Why can binary search not be used efficiently on a linked list? Binary search must jump to the middle of a range. A linked list can only be entered at the head, so each jump costs a walk, and the walking dominates.

4. In the measured run, what happened to comparisons and to pointer steps under binary search? Comparisons fell to about 10, as promised, but pointer steps rose to roughly ten times the length of the list, making it worse overall than a single traversal.

5. Which structure would you choose for frequent searching with frequent insertion, and why not a list or a sorted array? A binary search tree, which does both in O(log n). A sorted array searches well but inserts in O(n); a linked list inserts well but searches in O(n).

6. State the general lesson this chapter illustrates. The structure determines which algorithms are available. An algorithm's stated complexity assumes the operations it needs are cheap, and on a different structure they may not be.

Contents This chapter on its own page

munotes.in49

Chapter Seventeen

Prepending a Node

Syllabus topic Module 1, "Linked Structures: Prepending and Removing Nodes"

In one line

Prepending puts a new node at the front by pointing it at the old first node and then moving the head, which is two assignments and is O(1) however long the list is.

The operation

new = Node(value)

new.next = head

head = new

Three lines, or two if the node is built with its next already set. Nothing else in the list is touched.

The order matters and is the whole of the difficulty. new.next = head must happen before head = new, because once the head has moved, the address of the old first node is gone and with it the rest of the list.

Built and run, including the empty case

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


class LinkedList:
    def __init__(self):
        self.head = None
        self._count = 0

    def prepend(self, value):
        """Two assignments, in this order. Works on an empty list unchanged."""
        new = Node(value)
        new.next = self.head          # 1. the new node records the old front
        self.head = new               # 2. the front moves to the new node
        self._count += 1

    def to_list(self):
        out, walk = [], self.head
        while walk is not None:
            out.append(walk.data)
            walk = walk.next
        return out

    def length(self):
        return self._count


lst = LinkedList()
print("empty         :", lst.to_list())
for value in ("C", "B", "A"):
    lst.prepend(value)
    print("prepend %-3s ->" % value, lst.to_list())
print("length        :", lst.length())
empty         : []
prepend C   -> ['C']
prepend B   -> ['B', 'C']
prepend A   -> ['A', 'B', 'C']
length        : 3

Notice the first prepend on the empty list. self.head was None, so the new node's next became None, which is exactly a one node list. No special case was written and none is needed.

Notice also that the values come out in the reverse of the order they went in. Prepending reverses, and that is worth remembering: it is how chapter 33 builds a stack out of a list in one line.

Why it is O(1)

Nothing in those two assignments depends on the length of the list. No walk, no count, no shift. A list of ten nodes and a list of ten million cost exactly the same.

Counted, so it is not merely asserted:

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


class Counted:
    """A list that counts the pointer assignments each operation performs."""

    def __init__(self):
        self.head = None
        self.assignments = 0

    def prepend(self, value):
        new = Node(value)
        new.next = self.head
        self.assignments += 1
        self.head = new
        self.assignments += 1


for n in (100, 1000, 10000, 100000):
    lst = Counted()
    for i in range(n):
        lst.prepend(i)
    print("prepended %6d nodes: %7d pointer assignments, %d per node"
          % (n, lst.assignments, lst.assignments // n))
prepended    100 nodes:     200 pointer assignments, 2 per node
prepended   1000 nodes:    2000 pointer assignments, 2 per node
prepended  10000 nodes:   20000 pointer assignments, 2 per node
prepended 100000 nodes:  200000 pointer assignments, 2 per node
munotes.in50

Prepending a Node

Two per node, at every size. That is what O(1) per operation looks like when you count it rather than time it.

Against the array

The same job on an array, from chapter 12: inserting at the front moves every existing item, so inserting n items at the front of an array costs about n squared over 2 moves in total.

nArray, total movesLinked list, total assignments
100about 4,950200
1,000about 499,5002,000
10,000about 49,995,00020,000
100,000about 4,999,950,000200,000

The array column grows by a factor of a hundred each time the list column grows by ten. This one operation is the entire argument for the linked list.

Where prepending is the natural thing to do

Three places in this paper, all of them later chapters, use prepending because it is free:

  • A stack (chapter 33) pushes and pops at the front, so both operations are O(1).
  • Reversing a list is a traversal that prepends each node onto a new list.
  • Chaining in a hash table (chapter 104) prepends each colliding key onto the front of its bucket,

because the bucket's order does not matter and the front is cheapest.

Quick revision

  • Prepend: build the node, point it at the old head, then move the head. Two assignments.
  • The order is load bearing: pointing the new node forward must come first, or the rest of the list is

lost.

  • It works on an empty list with no special case, because the new node's next becomes null.
  • It is O(1): measured at 2 pointer assignments per node at every size from 100 to 100,000.
  • Prepending reverses the order of insertion.
  • Inserting n items at the front of an array costs about n squared over 2 moves; the linked list costs

2n assignments.

  • Stacks, list reversal and hash table chaining all prepend because it is free.

Test yourself

1. Write the prepend operation in three lines. new = Node(value); new.next = head; head = new.

2. What happens if the two assignments are done in the wrong order? head = new first would overwrite the only reference to the old first node, so new.next = head would then point the new node at itself and the rest of the list would be unreachable.

3. Why does prepend need no special case for the empty list? Because when head is null the new node's next becomes null, which is exactly the correct shape for a one node list.

4. What is the cost of prepending, and what evidence is given for it here? O(1). Counted, it is exactly 2 pointer assignments per node at every size from 100 to 100,000 nodes.

munotes.in51

Prepending a Node

5. Values C, B, A are prepended in that order. What does the list contain? A, B, C. Prepending reverses the order of insertion.

6. Give two places later in this paper where prepending is chosen because it is free. The stack, which pushes and pops at the front; and chaining in a hash table, where a colliding key is put at the front of its bucket because the bucket's order does not matter.

Contents This chapter on its own page

munotes.in52

Chapter Eighteen

Appending a Node, and Why It Costs More

Syllabus topic Module 1, "Linked Structures: Prepending and Removing Nodes"

In one line

Appending means adding at the end, which costs a full walk to find the end, unless the list keeps a pointer to its last node, in which case it costs the same as prepending.

Why the end is far away

A singly linked list knows where it starts. Nothing in it knows where it ends: the only way to find the last node is to follow the chain until a node's next is null.

So the naive append is a traversal plus two assignments, and the traversal is the whole cost.

if head is null: head = new; stop

walk = head

while walk.next is not null: walk = walk.next

walk.next = new

Note while walk.next is not null and not while walk is not null. The loop must stop on the last node, not past it, because the last node is the one that has to be modified. Writing the traversal condition out of habit here is a standard mistake.

Measured

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


class SlowList:
    """Append by walking to the end every time."""

    def __init__(self):
        self.head = None
        self.steps = 0

    def append(self, value):
        new = Node(value)
        if self.head is None:
            self.head = new
            return
        walk = self.head
        while walk.next is not None:
            walk = walk.next
            self.steps += 1
        walk.next = new


for n in (100, 200, 400, 800):
    lst = SlowList()
    for i in range(n):
        lst.append(i)
    print("appended %4d nodes by walking: %7d pointer steps" % (n, lst.steps))
appended  100 nodes by walking:    4851 pointer steps
appended  200 nodes by walking:   19701 pointer steps
appended  400 nodes by walking:   79401 pointer steps
appended  800 nodes by walking:  318801 pointer steps

Double the nodes, quadruple the steps. Building a list by appending this way is O(n squared), which is exactly as bad as the array was at prepending. The structure that was supposed to make insertion cheap has become expensive, and only at one end.

The tail pointer

The fix is to remember the answer instead of recomputing it. Keep a second variable, tail, holding the address of the last node.

Append then becomes: point the old last node at the new one, and move the tail. Two assignments, O(1).

The price is that every operation that changes the end of the list must now maintain the tail, and forgetting is a real defect: the list keeps working for a while and then appends into a node that is no longer last.

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


class FastList:
    """Append in O(1) by keeping a pointer to the last node."""

    def __init__(self):
        self.head = None
        self.tail = None
        self.assignments = 0

    def append(self, value):
        new = Node(value)
        if self.head is None:
            self.head = self.tail = new        # both, on the first node
            self.assignments += 2
            return
        self.tail.next = new
        self.assignments += 1
        self.tail = new
        self.assignments += 1

    def to_list(self):
        out, walk = [], self.head
        while walk is not None:
            out.append(walk.data)
            walk = walk.next
        return out


lst = FastList()
for value in ("A", "B", "C"):
    lst.append(value)
    print("append %s ->" % value, lst.to_list(), "| tail holds", lst.tail.data)

print()
for n in (100, 200, 400, 800):
    big = FastList()
    for i in range(n):
        big.append(i)
    print("appended %4d nodes with a tail: %5d assignments, %d per node"
          % (n, big.assignments, big.assignments // n))
munotes.in53

Appending a Node, and Why It Costs More

append A -> ['A'] | tail holds A
append B -> ['A', 'B'] | tail holds B
append C -> ['A', 'B', 'C'] | tail holds C

appended  100 nodes with a tail:   200 assignments, 2 per node
appended  200 nodes with a tail:   400 assignments, 2 per node
appended  400 nodes with a tail:   800 assignments, 2 per node
appended  800 nodes with a tail:  1600 assignments, 2 per node

Two per node at every size, the same as prepend. The 318,801 steps for 800 nodes became 1,600.

The general habit

This is worth naming, because the paper does it four more times.

When an operation recomputes something that could have been remembered, remember it. The cost is a little memory and the discipline of keeping the remembered thing correct.

WhereWhat is rememberedWhat it saves
this chapterthe last nodea full walk per append
chapter 14the lengtha full walk per length
chapter 82a heap's shape, in an arrayall the child pointers
chapter 107the count of items in a hash tablerecomputing the load factor

And the matching warning, which is the same every time: anything remembered must be updated by every operation that could invalidate it. A tail pointer that is not updated on deletion of the last node is worse than no tail pointer at all, because the list still looks correct.

What the tail cannot fix

A tail pointer makes appending cheap. It does not make deleting from the end cheap, and that catches people.

To delete the last node you must set the second to last node's next to null, and a singly linked list gives you no way to find the second to last node except by walking from the head. The tail pointer tells you where the end is, not what came before it.

Deleting from the end of a singly linked list is O(n) even with a tail pointer. Fixing that needs a backward link, which is the doubly linked list of chapter 26.

munotes.in54

Appending a Node, and Why It Costs More

Quick revision

  • A singly linked list has no way to find its end except by walking; the naive append is O(n).
  • Building a list by naive appending is O(n squared): 800 nodes cost 318,801 pointer steps.
  • The append traversal stops on the last node, testing walk.next is not null, not walk is not null.
  • A tail pointer makes append O(1): measured at 2 assignments per node at every size.
  • On the first node, head and tail are both set to it.
  • Anything remembered must be maintained by every operation that could invalidate it.
  • A tail pointer does not make deletion from the end cheap: that needs the second to last node, which

only a backward link gives you.

Test yourself

1. Why is appending to a singly linked list O(n) without a tail pointer? Because nothing records where the list ends, so the last node must be found by following the chain from the head.

2. What is the cost of building an n node list by naive appending, and what did the run show? O(n squared). Doubling the nodes quadrupled the steps: 100 nodes cost 4,851 steps and 800 cost 318,801.

3. Why is the loop condition walk.next is not null rather than walk is not null? Because the operation must stop on the last node in order to modify it. The ordinary traversal condition would walk past the end.

4. What must be set when the first node is appended to an empty list? Both head and tail must be set to the new node.

5. State the general habit this chapter introduces, and its matching danger. Remember what would otherwise be recomputed. The danger is that anything remembered must be updated by every operation that could invalidate it, or the structure lies while still appearing to work.

6. Does a tail pointer make deleting the last node O(1)? Explain. No. Deleting the last node requires setting the second to last node's next to null, and a singly linked list can only find that node by walking from the head. It stays O(n) until a backward link exists.

Contents This chapter on its own page

munotes.in55

Chapter Nineteen

Removing a Node

Syllabus topic Module 1, "Linked Structures: Prepending and Removing Nodes"

In one line

To remove a node you must change the node before it, and a singly linked list gives you no way to reach backwards, so the predecessor has to be tracked on the way in.

The predecessor problem

Removing node X means making the chain skip it: whatever pointed at X must now point at X's next.

The trouble is that X does not know what points at it. Its next field looks forward; nothing looks back. So finding X is not enough. You have to arrive at X remembering where you came from.

That is the shape of every deletion in a singly linked list, and it is why deletion code carries two walking pointers where search carried one.

previous = null

walk = head

while walk is not null and walk.data is not target: previous = walk; walk = walk.next

When the loop ends, walk is the node to remove and previous is the node before it, or previous is null meaning the node to remove is the first one.

The three cases

CaseWhat to do
The list is emptynothing to remove: an error
The target is the first nodemove the head: head = head.next
The target is anywhere elseprevious.next = walk.next
The target is absentan error

The second case is the one that needs care, because it changes the head, and an operation that changes the head must actually write the new head back.

Built and run, all cases

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


class LinkedList:
    def __init__(self, values=()):
        self.head = None
        for value in reversed(list(values)):
            self.head = Node(value, self.head)

    def remove(self, target):
        """Remove the first node holding target. Returns the nodes examined."""
        previous, walk, examined = None, self.head, 0
        while walk is not None:
            examined += 1
            if walk.data == target:
                if previous is None:
                    self.head = walk.next        # removing the first node
                else:
                    previous.next = walk.next    # removing any other node
                return examined
            previous = walk
            walk = walk.next
        raise KeyError("not in the list: %r" % (target,))

    def to_list(self):
        out, walk = [], self.head
        while walk is not None:
            out.append(walk.data)
            walk = walk.next
        return out


for target in ("A", "C", "E"):
    lst = LinkedList(["A", "B", "C", "D", "E"])
    examined = lst.remove(target)
    print("remove %-3s -> %-22s (examined %d node(s))"
          % (target, lst.to_list(), examined))

one = LinkedList(["only"])
one.remove("only")
print("remove from a one node list ->", one.to_list(), "| empty again:",
      one.head is None)

try:
    LinkedList(["A"]).remove("Z")
except KeyError as e:
    print("remove an absent value -> refused:", e)

try:
    LinkedList([]).remove("Z")
except KeyError as e:
    print("remove from an empty list -> refused:", e)
remove A   -> ['B', 'C', 'D', 'E']   (examined 1 node(s))
remove C   -> ['A', 'B', 'D', 'E']   (examined 3 node(s))
remove E   -> ['A', 'B', 'C', 'D']   (examined 5 node(s))
remove from a one node list -> [] | empty again: True
remove an absent value -> refused: "not in the list: 'Z'"
remove from an empty list -> refused: "not in the list: 'Z'"
munotes.in56

Removing a Node

Every case behaves: the first node, a middle node, the last node, the only node, an absent value and an empty list. The empty list needed no special case at all, because the loop simply never runs and the error is raised at the end.

What it costs, and the honest statement of it

Removing a known node, when you already hold its predecessor, is two assignments: O(1).

Removing a node by value requires finding it first, which is a search: O(n).

Students often say "deletion from a linked list is O(1)", and examiners mark it wrong, because the sentence is missing its condition. The correct statement is:

Deletion is O(1) given the predecessor, and O(n) if the node must be found.

That distinction returns in chapter 29, where a doubly linked list can delete a known node in O(1) without its predecessor being supplied, and in chapter 104, where a hash table's chains are kept short precisely so that this search stays cheap.

The node that is removed but not gone

In C, unlinking a node leaks its memory unless you free it, and freeing it before reading its next reads memory you have given back. The order is: remember walk.next, unlink, then free.

In Python the garbage collector handles it once nothing refers to the node. It is worth knowing that this is a service the language is performing, not a property of linked lists.

next_one = walk.next

previous.next = next_one

free(walk)

Quick revision

  • To remove a node you must modify the node before it, and a singly linked list cannot look backwards.
  • So deletion walks with two pointers, previous and walk.
  • Four cases: empty list (error), first node (move the head), any other node (`previous.next =

walk.next`), absent (error).

  • Removing the first node changes the head, and the new head must be written back.
  • Deletion is O(1) given the predecessor, O(n) if the node must be found by value. The unqualified

claim "deletion is O(1)" is wrong.

  • In C, unlink first and free afterwards, having already saved walk.next.

Test yourself

1. Why is finding the node to delete not enough? Because deletion changes the node before it, and a node in a singly linked list holds no reference back to whatever points at it. The predecessor must be tracked on the way in.

2. Write the two pointer loop for deletion. previous = null; walk = head; while walk is not null and walk.data is not the target: previous = walk; walk = walk.next.

munotes.in57

Removing a Node

3. What tells you that the node to be removed is the first one, and what must you do then? previous is still null. Set head = walk.next, and make sure the new head is written back to the list.

4. Give the correct complexity statement for deletion, with its condition. O(1) if the predecessor is already known, O(n) if the node has to be found by value.

5. In C, in what order must unlinking and freeing be done, and why? Save walk.next first, then unlink, then free. Freeing before reading next would read memory that has been given back.

6. Removing "E" from a five node list examined five nodes. Removing "A" examined one. What does that show? That the cost of removal by value is the cost of the search, so it depends on where the value sits, not on the unlinking, which is always two assignments.

Contents This chapter on its own page

munotes.in58

Chapter Twenty

Inserting and Deleting at Any Position

Syllabus topic Module 1, "Linked Structures: Insertion and deletion of nodes at various positions"

In one line

Insertion and deletion at position k are both a walk to position k-1 followed by two pointer assignments, so the position decides the cost and the assignments are always the same.

One procedure, every position

Rather than three separate routines for front, middle and end, write one that walks to the node before the position and then does the same two assignments. The front is the special case, because there is no node before it.

insert value at position k:

if k is 0: prepend; stop

walk to the node at position k-1

new.next = walk.next

walk.next = new

The same shape deletes:

delete at position k:

if k is 0: head = head.next; stop

walk to the node at position k-1

walk.next = walk.next.next

That is the whole of it. Everything else is checking that k is in range.

Built and run at every position

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


class LinkedList:
    def __init__(self, values=()):
        self.head = None
        self._count = 0
        for value in reversed(list(values)):
            self.head = Node(value, self.head)
            self._count += 1

    def length(self):
        return self._count

    def insert_at(self, k, value):
        """Insert so that the new node ends up at position k. Returns steps walked."""
        if not 0 <= k <= self._count:
            raise IndexError("position %d is outside 0 to %d" % (k, self._count))
        self._count += 1
        if k == 0:
            self.head = Node(value, self.head)
            return 0
        walk, steps = self.head, 0
        for _ in range(k - 1):
            walk = walk.next
            steps += 1
        new = Node(value)
        new.next = walk.next            # 1. the new node looks forward FIRST
        walk.next = new                 # 2. then the predecessor looks at it
        return steps

    def delete_at(self, k):
        if not 0 <= k < self._count:
            raise IndexError("position %d is outside 0 to %d" % (k, self._count - 1))
        self._count -= 1
        if k == 0:
            self.head = self.head.next
            return 0
        walk, steps = self.head, 0
        for _ in range(k - 1):
            walk = walk.next
            steps += 1
        walk.next = walk.next.next
        return steps

    def to_list(self):
        out, walk = [], self.head
        while walk is not None:
            out.append(walk.data)
            walk = walk.next
        return out


for k in (0, 2, 4):
    lst = LinkedList(["A", "B", "C", "D"])
    steps = lst.insert_at(k, "NEW")
    print("insert at %d -> %-30s (walked %d)" % (k, str(lst.to_list()), steps))

print()
for k in (0, 2, 3):
    lst = LinkedList(["A", "B", "C", "D"])
    steps = lst.delete_at(k)
    print("delete at %d -> %-30s (walked %d)" % (k, str(lst.to_list()), steps))

print()
lst = LinkedList(["A", "B"])
try:
    lst.insert_at(5, "X")
except IndexError as e:
    print("insert out of range -> refused:", e)
try:
    lst.delete_at(2)
except IndexError as e:
    print("delete out of range -> refused:", e)
insert at 0 -> ['NEW', 'A', 'B', 'C', 'D']    (walked 0)
insert at 2 -> ['A', 'B', 'NEW', 'C', 'D']    (walked 1)
insert at 4 -> ['A', 'B', 'C', 'D', 'NEW']    (walked 3)

delete at 0 -> ['B', 'C', 'D']                (walked 0)
delete at 2 -> ['A', 'B', 'D']                (walked 1)
delete at 3 -> ['A', 'B', 'C']                (walked 2)

insert out of range -> refused: position 5 is outside 0 to 2
delete out of range -> refused: position 2 is outside 0 to 1
munotes.in59

Inserting and Deleting at Any Position

Two details in that run are examinable.

Insertion allows k equal to the length (appending), while deletion does not (there is no item there). That is why the two range checks differ, and getting them the same way round is a common error.

The walk is k-1 steps, not k. Inserting at position 4 walked 3.

The wrong order, run

Chapter 13 warned that new.next = walk.next must come before walk.next = new. Here is what the other order actually does.

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


def as_row(head, cap=8):
    """Walk, but stop after cap nodes so a broken list cannot hang this."""
    out, walk, n = [], head, 0
    while walk is not None and n < cap:
        out.append(walk.data)
        walk = walk.next
        n += 1
    if walk is not None:
        out.append("... still going")
    return out


head = Node("A", Node("B", Node("C")))
walk = head                       # insert after A

new = Node("NEW")
walk.next = new                   # WRONG: the predecessor points at the new node first
new.next = walk.next              # and now the new node points at ITSELF

print("after the wrong order:", as_row(head))
print("does NEW point at itself?", new.next is new)
after the wrong order: ['A', 'NEW', 'NEW', 'NEW', 'NEW', 'NEW', 'NEW', 'NEW', '... still going']
does NEW point at itself? True

The list now contains a cycle: NEW points at itself, so any traversal runs for ever. B and C are still in memory, perfectly intact, and completely unreachable.

This is why the order is taught as a rule rather than left to be worked out each time: the failure is not an error message, it is a program that never returns.

The cost

PositionWalkAssignmentsTotal
Front (k = 0)none2O(1)
Middleabout n/22O(n)
End (k = n)n-12O(n)

The assignments are constant everywhere. It is the walking that costs, so the position decides the price. That is the reverse of the array, where the assignments were the cost and the position decided how many items had to move.

And a tail pointer helps only the last of those rows, and only for insertion, for the reason chapter 18 gave: the end can be reached, but the node before the end cannot.

munotes.in60

Inserting and Deleting at Any Position

Quick revision

  • One procedure covers every position: walk to k-1, then two assignments. Position 0 is the special

case, since nothing precedes it.

  • Insert allows k equal to the length; delete does not. The range checks are deliberately different.
  • The walk is k-1 steps, not k.
  • The assignment order is load bearing: point the new node forward first. Reversed, the new node points

at itself, the list gains a cycle, and every traversal runs for ever while the rest of the list stays intact and unreachable.

  • Assignments are O(1) at every position; the walk is what costs, so the cost is O(1) at the front and

O(n) elsewhere.

  • This is the reverse of the array, where the shifting was the cost.

Test yourself

1. Write the general insert-at-k procedure. If k is 0, prepend. Otherwise walk to the node at position k-1, set new.next = walk.next, then walk.next = new.

2. Why does insertion permit k equal to the length while deletion does not? Because inserting at the length means appending after the last item, which is a real position. Deleting at the length would delete an item that is not there.

3. How many steps does the walk take to insert at position 4? Three: the walk stops at position 3, the node before the insertion point.

4. Exactly what happens if the two assignments are reversed? The predecessor is pointed at the new node first, so walk.next is already the new node when new.next = walk.next runs, and the new node points at itself. The list has a cycle, traversals never end, and the remaining nodes become unreachable.

5. Why is this failure worse than an error message? Because nothing reports it. The program hangs, and the data that was lost is still in memory and perfectly intact, so nothing looks corrupted.

6. What is the cost of insertion at the front, the middle and the end of a singly linked list? O(1) at the front, O(n) in the middle and O(n) at the end. The two assignments are constant everywhere; the walk to position k-1 is what costs.

Contents This chapter on its own page

munotes.in61

Chapter Twenty-One

What a Singly Linked List Is Good and Bad At

Syllabus topic Module 1, "Linked Structures: Advantages & Disadvantages, Singly Linked List"

In one line

A singly linked list buys cheap insertion and deletion anywhere, and unlimited growth, by giving up computed access, extra memory per item, and any ability to look backwards.

The advantages

1. Insertion and deletion cost nothing, given the position. Two assignments, whatever the length. Chapter 17 counted exactly 2 per node from 100 to 100,000. The array moved every item.

2. The size is not fixed. No maximum is declared and no block is reserved. The list grows while memory lasts and shrinks node by node. An array must either guess its size in advance or be rebuilt.

3. No memory is wasted on unused capacity. An array sized for 10,000 that holds 12 items keeps 9,988 empty cells. A linked list of 12 items holds 12 nodes.

4. No large contiguous block is needed. A machine whose memory is fragmented may be unable to provide one block of 10,000 cells while easily providing 10,000 scattered nodes.

5. Several structures fall out of it almost free. The stack, the queue and the deque of this module are all a linked list with the operations restricted.

The disadvantages

1. No computed access. Reaching item i costs i steps. The array computed the address. This is the big one, and it is what rules the linked list out whenever the work is indexing.

2. Extra memory per item. Every node carries an address beside its value. On a 64 bit machine that is 8 bytes per item, plus whatever the language's object header costs, which in Python is a great deal more.

3. No way to look backwards. Deletion needs the predecessor, so it is tracked on the way in; there is no reaching back. Chapter 26 pays for a backward link and chapter 29 measures what it costs.

4. Poor cache behaviour. Memory is fetched in blocks, so an array's next item is usually already there. A linked list's next node may be anywhere, so each step may be a fresh fetch. This is invisible to the counting and very visible on a real machine, which is the whole of chapter 22.

5. No binary search, even when sorted. Chapter 16 measured this: the comparisons drop to about ten and the pointer steps rise to ten times the length of the list.

The table

OperationArraySingly linked list
Read or write item iO(1)O(n)
Insert at the frontO(n)O(1)
Insert at the endO(1) amortisedO(n), or O(1) with a tail
Insert in the middle, position knownO(n)O(1)
Delete at the frontO(n)O(1)
Delete at the endO(1)O(n), even with a tail
Search, unsortedO(n)O(n)
Search, sortedO(log n)O(n)
Memory per itemthe valuethe value plus an address
Growthnew block, full copyone node
munotes.in62

What a Singly Linked List Is Good and Bad At

Read the first two rows together, because they are the whole bargain: the array wins the first, the list wins the second, and no structure in Module 1 wins both.

The sentence that is almost always wrong

"Linked lists are better than arrays for insertion."

It needs a condition, and without it examiners mark it down. Inserting into a linked list is O(1) once you are standing at the right place. Getting there is O(n). So inserting into a linked list by position is O(n) overall, exactly like the array.

The linked list wins when the position is already known: while traversing, at the front, at a node you already hold. That is why chapter 104 uses a linked list for hash table chains, where insertion is always at the front, and why the stack and the queue are built on one.

When to choose which, as a rule

If the work is mostlyChoose
reading item i, or binary searcharray
inserting and deleting at the frontlinked list
inserting and deleting while traversinglinked list
appending and readingarray
unpredictable size, no indexinglinked list
small data, in a tight looparray, for the cache

Quick revision

  • Advantages: O(1) insertion and deletion given the position; no fixed size; no wasted capacity; no

large contiguous block needed; stacks, queues and deques build on it.

  • Disadvantages: O(n) access to item i; an address of memory per item; no backward link; poor cache

behaviour; no binary search even when sorted.

  • The array wins access; the list wins insertion; neither wins both.
  • "Better for insertion" is only true when the position is already known, because reaching a position

is O(n).

  • Deleting the last node is O(n) even with a tail pointer.

Test yourself

1. Give three advantages of a singly linked list over an array. Insertion and deletion at a known position are O(1); the size is not fixed and no capacity is wasted; and it needs no single large contiguous block of memory.

2. Give three disadvantages. Access to item i is O(n); every node costs an extra address of memory; and there is no backward link, so deletion needs the predecessor tracked on the way in. Cache behaviour and the loss of binary search are two more.

3. Why is "linked lists are better for insertion" incomplete? Because insertion is O(1) only once you are at the position. Reaching position k is O(n), so insertion by position costs the same as the array's.

4. Name two places in this paper where the linked list's insertion advantage is genuinely realised. Hash table chaining, where insertion is always at the front of a bucket; and the stack, where every operation is at one end. The queue with a tail pointer is a third.

munotes.in63

What a Singly Linked List Is Good and Bad At

5. Which operation is O(n) on a linked list even with a tail pointer, and why? Deleting the last node, because the node before it must be modified and only a forward walk can find it.

6. A program stores a few thousand integers and mostly reads them by index in a tight loop. Which structure and why? The array. Indexing is O(1) against O(n), and the contiguous layout uses the cache well, which matters more than the counting suggests.

Contents This chapter on its own page

munotes.in64

Chapter Twenty-Two

The Array Against the Linked List, Measured

Syllabus topic Computer Science Practical 3, Module 2, "Compare static (array) vs dynamic (linked) approaches"

In one line

Counted head to head, the array wins access by a factor of the list's length, the linked list wins front insertion by the same factor, and each is useless at what the other is good at.

Both structures, the same operations

class Node:
    def __init__(self, data, next_node=None):
        self.data = data
        self.next = next_node


class LinkedSeq:
    """A singly linked list, counting every pointer step and assignment."""

    def __init__(self):
        self.head = None
        self.n = 0

    def prepend(self, value):
        self.head = Node(value, self.head)
        self.n += 1
        return 2                                  # two assignments

    def get(self, i):
        walk, steps = self.head, 0
        while i > 0:
            walk = walk.next
            steps += 1
            i -= 1
        return walk.data, steps + 1


class ArraySeq:
    """A fixed array, counting every element moved and every index computed."""

    def __init__(self, capacity):
        self.cells = [None] * capacity
        self.n = 0

    def prepend(self, value):
        moves = 0
        for i in range(self.n, 0, -1):            # shift everything up one
            self.cells[i] = self.cells[i - 1]
            moves += 1
        self.cells[0] = value
        self.n += 1
        return moves + 1

    def get(self, i):
        return self.cells[i], 1                   # one address computation


for n in (500, 1000, 2000, 4000):
    linked, array = LinkedSeq(), ArraySeq(n)
    link_build = sum(linked.prepend(i) for i in range(n))
    array_build = sum(array.prepend(i) for i in range(n))
    _, link_get = linked.get(n - 1)
    _, array_get = array.get(n - 1)
    print("n = %4d | build by prepending: array %8d, linked %5d "
          "| read the last item: array %d, linked %4d"
          % (n, array_build, link_build, array_get, link_get))
n =  500 | build by prepending: array   125250, linked  1000 | read the last item: array 1, linked  500
n = 1000 | build by prepending: array   500500, linked  2000 | read the last item: array 1, linked 1000
n = 2000 | build by prepending: array  2001000, linked  4000 | read the last item: array 1, linked 2000
n = 4000 | build by prepending: array  8002000, linked  8000 | read the last item: array 1, linked 4000

Both columns of that table are the same fact seen from two sides.

Building by prepending: the array's work quadruples when n doubles, because each of n insertions moves about n items. The list's doubles, because each insertion costs 2 whatever happens. At n = 4000 the array did a thousand times the work.

Reading the last item: the array does 1 unit of work at every size. The list does n. At n = 4000 the list did four thousand times the work.

Neither structure is better. They are opposites.

The thing counting cannot see

There is a real effect that no count above captures, and the practical's comparison is incomplete without it.

Memory is fetched from main memory into cache in blocks, typically 64 bytes at a time. An array's items sit together, so fetching item 0 usually brings items 1 to 15 along with it for free. A linked list's nodes are wherever the allocator put them, so each step may be a fresh fetch from main memory, which is far slower than a cache hit.

munotes.in65

The Array Against the Linked List, Measured

The consequence, and it surprises people: an array often beats a linked list at operations the counting says the list should win, once the data is large enough for cache to matter, because the array's 1,000 cheap moves can beat the list's 2 expensive pointer chases.

This book does not print a timing to prove that, for the reason chapter 1 gave. What it does is state it and name it, because an examiner asking "is a linked list always better for insertion" is asking about exactly this.

The comparison, as the practical wants it written

Static arrayDynamic linked list
Sizefixed at creationgrows and shrinks at run time
Memoryone contiguous blockscattered nodes
Extra memoryunused capacity is wastedone address per item
Access item iO(1), computedO(n), walked
Insert at frontO(n)O(1)
Insert at endO(1)O(n), or O(1) with a tail
Insert at a known positionO(n)O(1)
Delete at frontO(n)O(1)
Delete at endO(1)O(n)
Search unsortedO(n)O(n)
Search sortedO(log n)O(n)
Cache behaviourgoodpoor
Needs contiguous memoryyesno

Choosing between them, as one sentence

If the work is dominated by reaching items by position, use the array. If it is dominated by inserting and removing at positions you are already standing at, use the linked list. If it is both, you need a structure from Module 2.

That last clause is the bridge. Every structure in Module 2 exists because Module 1's two structures each fail at half the job.

Quick revision

  • Counted head to head: building 4,000 items by prepending cost the array 8,002,000 moves and the list

8,000 assignments; reading the last item cost the array 1 and the list 4,000.

  • The array's prepending work quadruples when n doubles; the list's doubles.
  • They are opposites, not better and worse.
  • Counting cannot see cache: an array's items are fetched together in blocks, a list's nodes are

scattered, so an array often wins in practice even where the count says otherwise.

  • Array: fixed size, contiguous, O(1) access, expensive insertion except at the end.
  • Linked list: grows at run time, scattered, O(n) access, O(1) insertion at a known position.
  • If the work needs both cheap access and cheap insertion, neither is enough, and that is why Module 2

exists.

Test yourself

1. At n = 4000, what did building by prepending cost each structure, and why do they differ so much? The array did 8,002,000 moves and the list 8,000 assignments. Each array insertion shifts about n items while each list insertion costs two assignments whatever the length.

munotes.in66

The Array Against the Linked List, Measured

2. At the same size, what did reading the last item cost each? The array 1 unit, the list 4,000. The array computes the address; the list must walk.

3. What effect does counting fail to capture, and which structure does it favour? Cache behaviour. Memory is fetched in blocks, so an array's neighbouring items come along for free while a linked list's scattered nodes may each need a fresh fetch. It favours the array.

4. Give the four rows of the comparison that concern memory. The array is a fixed contiguous block and wastes unused capacity; the list is scattered nodes and costs one address per item.

5. State in one sentence when to choose each. The array when the work is reaching items by position; the linked list when the work is inserting and removing at positions you are already at.

6. Why does this comparison point forward to Module 2? Because each structure fails at half the job, and a problem needing both cheap access and cheap insertion cannot be solved by either. Trees and hash tables are the answers.

Contents This chapter on its own page

munotes.in67

Chapter Twenty-Three

A Polynomial as a Linked List

Syllabus topic Module 1, "Linked Structures: applications of linked list like polynomial equation"

In one line

A polynomial is stored as a linked list of its non-zero terms, each node holding a coefficient and an exponent, which costs nothing for the terms that are not there.

The two representations

Take the polynomial:

P(x) = 5x^4 + 3x^2 + 7

As an array, indexed by exponent. Cell i holds the coefficient of x to the i.

index01234
coefficient70305

Simple, and the exponent is the index so nothing needs to be stored twice. Adding two polynomials is adding the cells.

As a linked list, one node per non-zero term, kept in decreasing order of exponent:

head --. [5, 4 | -]--. [3, 2 | -]--. [7, 0 | X]

Each node holds a coefficient and an exponent, and the order is a promise the operations rely on.

Why the list, and when

For that small polynomial the array is obviously fine. Now take this one:

Q(x) = 3x^1000 + 5

The array needs 1,001 cells to hold two numbers. 999 of them are zero. The linked list needs two nodes.

That is the case the linked representation is for, and it has a name: a polynomial with few non-zero terms relative to its degree is called sparse. The array wastes space in proportion to the degree; the list uses space in proportion to the number of terms.

The structure, built

class Term:
    """One term of a polynomial: a coefficient and an exponent."""

    def __init__(self, coefficient, exponent, next_term=None):
        self.coefficient = coefficient
        self.exponent = exponent
        self.next = next_term


class Polynomial:
    """Terms in strictly decreasing order of exponent, no zero coefficients."""

    def __init__(self, terms=()):
        self.head = None
        for coefficient, exponent in sorted(terms, key=lambda t: t[1]):
            if coefficient != 0:
                self.head = Term(coefficient, exponent, self.head)

    def terms(self):
        out, walk = [], self.head
        while walk is not None:
            out.append((walk.coefficient, walk.exponent))
            walk = walk.next
        return out

    def degree(self):
        return self.head.exponent if self.head else None

    def count(self):
        n, walk = 0, self.head
        while walk is not None:
            n += 1
            walk = walk.next
        return n

    def evaluate(self, x):
        total, walk = 0, self.head
        while walk is not None:
            total += walk.coefficient * (x ** walk.exponent)
            walk = walk.next
        return total

    def __str__(self):
        if self.head is None:
            return "0"
        parts, walk = [], self.head
        while walk is not None:
            c, e = walk.coefficient, walk.exponent
            if e == 0:
                parts.append("%d" % c)
            elif e == 1:
                parts.append("%dx" % c)
            else:
                parts.append("%dx^%d" % (c, e))
            walk = walk.next
        return " + ".join(parts).replace("+ -", "- ")


p = Polynomial([(5, 4), (3, 2), (7, 0)])
print("P(x)        =", p)
print("degree      =", p.degree())
print("terms held  =", p.count())
print("P(2)        =", p.evaluate(2))

q = Polynomial([(3, 1000), (5, 0)])
print()
print("Q(x)        =", q)
print("degree      =", q.degree())
print("terms held  =", q.count())
print("an array for Q would need", q.degree() + 1, "cells to hold", q.count(), "numbers")
munotes.in68

A Polynomial as a Linked List

P(x)        = 5x^4 + 3x^2 + 7
degree      = 4
terms held  = 3
P(2)        = 99

Q(x)        = 3x^1000 + 5
degree      = 1000
terms held  = 2
an array for Q would need 1001 cells to hold 2 numbers

Check P(2) by hand: 5 times 16 is 80, plus 3 times 4 is 12, plus 7, which is 99. The program agrees.

The two invariants

The representation makes two promises, and every operation in the next two chapters depends on both.

1. Exponents are strictly decreasing. No two nodes share an exponent, and they are in order. This is what lets addition walk the two lists together in one pass.

2. No coefficient is zero. A term that cancels is removed, not kept with a coefficient of 0. This is what keeps a sparse polynomial sparse.

The second is easy to forget when writing addition, and chapter 24 shows what happens when it is.

Measured: where each representation wins

def array_cells(terms):
    """An array indexed by exponent needs degree + 1 cells."""
    return max(e for _, e in terms) + 1


def list_nodes(terms):
    """The linked list needs one node per non-zero term."""
    return sum(1 for c, _ in terms if c != 0)


cases = [
    ("dense, degree 4",      [(5, 4), (2, 3), (3, 2), (1, 1), (7, 0)]),
    ("sparse, degree 1000",  [(3, 1000), (5, 0)]),
    ("one term, degree 50",  [(1, 50)]),
    ("dense, degree 10",     [(i + 1, i) for i in range(11)]),
]

print("%-22s %8s %8s   %s" % ("polynomial", "array", "list", "better"))
for name, terms in cases:
    cells, nodes = array_cells(terms), list_nodes(terms)
    # a list node holds two numbers and an address, so count it as 3 units
    better = "list" if nodes * 3 < cells else "array"
    print("%-22s %8d %8d   %s" % (name, cells, nodes * 3, better))
polynomial                array     list   better
dense, degree 4               5       15   array
sparse, degree 1000        1001        6   list
one term, degree 50          51        3   list
dense, degree 10             11       33   array

The honest answer is it depends on sparseness, and the table says where the line is. A node costs about three units against an array cell's one, so the list wins once fewer than about a third of the terms are present.

That is the right way to answer "why use a linked list for a polynomial" in an examination: not "because it is dynamic", but because a sparse polynomial wastes most of an array, and the list charges only for the terms that exist.

Quick revision

  • A polynomial is stored as a linked list of terms, each node holding a coefficient and an exponent.
  • Two invariants: exponents strictly decreasing, and no zero coefficients.
  • The array representation indexes by exponent and needs degree + 1 cells whatever the number of terms.
  • A sparse polynomial, with few terms relative to its degree, wastes almost all of an array: 3x^1000 + 5
munotes.in69

A Polynomial as a Linked List

needs 1,001 cells to hold 2 numbers.

  • A node costs about three units to an array cell's one, so the list wins when fewer than about a third

of the terms are present.

  • For a dense polynomial the array is better, and saying so is part of the answer.

Test yourself

1. What does each node of a polynomial list hold? A coefficient, an exponent, and the address of the next term.

2. State the two invariants of this representation. Exponents are strictly decreasing, with no repeats; and no coefficient is zero.

3. How many cells does an array need for 3x^1000 + 5, and how many nodes does the list need? 1,001 cells against 2 nodes.

4. Define a sparse polynomial. One with few non-zero terms relative to its degree.

5. Give the examination answer to "why use a linked list for a polynomial". Because a sparse polynomial leaves almost all of an array empty, while the list charges only for the terms that exist. For a dense polynomial the array is the better choice.

6. Evaluate 5x^4 + 3x^2 + 7 at x = 2 and show the working. 5 times 16 is 80, 3 times 4 is 12, plus 7, giving 99.

Contents This chapter on its own page

munotes.in70

Chapter Twenty-Four

Adding Two Polynomials

Syllabus topic Module 1, "Linked Structures: applications of linked list like polynomial equation"

In one line

Adding two polynomials is a single walk down both lists at once, taking the larger exponent each time and adding the coefficients when the exponents are equal.

Why one pass is enough

Both lists are in strictly decreasing order of exponent. That means at any moment the two nodes you are looking at hold the largest remaining exponent in each polynomial, so the largest remaining exponent of the answer is one of those two. You never have to look further ahead.

This is the merge pattern, and it works on any two ordered sequences.

while both lists have nodes:

if exponent of A is greater: take A's term; advance A

if exponent of B is greater: take B's term; advance B

if they are equal: add the coefficients; if the sum is not zero, take it; advance both

then append whatever remains of the list that is not exhausted

Three cases, plus the tail. Missing the third is the usual error; missing the tail is the other one.

The cancellation trap

When the exponents are equal the coefficients are added, and the sum may be zero. Adding 5x^2 and -5x^2 gives 0x^2, which must not appear in the answer at all, because chapter 23's second invariant says no coefficient is zero.

A version that keeps it still computes the right value when evaluated, so it passes a casual test. It fails the invariant, so the polynomial is no longer sparse, and every later operation carries the useless term. The program below refuses it.

Built and run

class Term:
    def __init__(self, coefficient, exponent, next_term=None):
        self.coefficient = coefficient
        self.exponent = exponent
        self.next = next_term


class Polynomial:
    def __init__(self, terms=()):
        self.head = None
        for coefficient, exponent in sorted(terms, key=lambda t: t[1]):
            if coefficient != 0:
                self.head = Term(coefficient, exponent, self.head)

    def terms(self):
        out, walk = [], self.head
        while walk is not None:
            out.append((walk.coefficient, walk.exponent))
            walk = walk.next
        return out

    def evaluate(self, x):
        total, walk = 0, self.head
        while walk is not None:
            total += walk.coefficient * (x ** walk.exponent)
            walk = walk.next
        return total

    def __str__(self):
        if self.head is None:
            return "0"
        parts, walk = [], self.head
        while walk is not None:
            c, e = walk.coefficient, walk.exponent
            parts.append("%d" % c if e == 0 else
                         ("%dx" % c if e == 1 else "%dx^%d" % (c, e)))
            walk = walk.next
        return " + ".join(parts).replace("+ -", "- ")


def add(p, q):
    """Merge two ordered term lists in one pass. Returns (sum, comparisons)."""
    result = Polynomial()
    tail = None
    a, b, comparisons = p.head, q.head, 0

    def attach(coefficient, exponent):
        nonlocal tail
        node = Term(coefficient, exponent)
        if tail is None:
            result.head = node
        else:
            tail.next = node
        tail = node

    while a is not None and b is not None:
        comparisons += 1
        if a.exponent > b.exponent:
            attach(a.coefficient, a.exponent)
            a = a.next
        elif b.exponent > a.exponent:
            attach(b.coefficient, b.exponent)
            b = b.next
        else:
            total = a.coefficient + b.coefficient
            if total != 0:                       # the cancellation trap
                attach(total, a.exponent)
            a, b = a.next, b.next

    for walk in (a, b):                          # whatever is left over
        while walk is not None:
            attach(walk.coefficient, walk.exponent)
            walk = walk.next

    return result, comparisons


p = Polynomial([(5, 4), (3, 2), (7, 0)])
q = Polynomial([(2, 4), (4, 3), (1, 1)])
total, comparisons = add(p, q)
print("P(x)      =", p)
print("Q(x)      =", q)
print("P + Q     =", total)
print("comparisons:", comparisons)
print("check at x = 2:", p.evaluate(2), "+", q.evaluate(2), "=", p.evaluate(2) + q.evaluate(2),
      "| sum polynomial gives", total.evaluate(2))

print()
a = Polynomial([(5, 2), (3, 1)])
b = Polynomial([(-5, 2), (2, 0)])
cancelled, _ = add(a, b)
print("cancellation: (%s) + (%s) = %s" % (a, b, cancelled))
print("terms kept  :", cancelled.terms())

print()
zero, _ = add(Polynomial([(4, 3)]), Polynomial([(-4, 3)]))
print("everything cancels ->", zero, "| terms:", zero.terms())
munotes.in71

Adding Two Polynomials

P(x)      = 5x^4 + 3x^2 + 7
Q(x)      = 2x^4 + 4x^3 + 1x
P + Q     = 7x^4 + 4x^3 + 3x^2 + 1x + 7
comparisons: 4
check at x = 2: 99 + 66 = 165 | sum polynomial gives 165

cancellation: (5x^2 + 3x) + (-5x^2 + 2) = 3x + 2
terms kept  : [(3, 1), (2, 0)]

everything cancels -> 0 | terms: []

Three things in that run are the answer to the examination question.

The sum is correct, and it is checked independently: evaluating P and Q at x = 2 and adding gives 165, and evaluating the sum polynomial gives 165 too. That is a real check, not a restatement.

The cancelled term is gone. 5x^2 and -5x^2 produced no node at all, and the answer is 3x + 2, with two terms, not three.

Everything cancelling gives the zero polynomial, an empty list, printed as 0. An empty result is a legitimate answer and the code must produce it rather than failing.

The cost

One pass, and each comparison advances at least one of the two lists. So for polynomials with m and n terms, addition is O(m + n), which is the best possible: you cannot add them without looking at every term.

Compare the array representation: adding two arrays indexed by exponent costs O(larger degree), whether or not the terms exist. For the sparse case that is enormously worse. 3x^1000 + 5 added to itself is 2 comparisons on lists and 1,001 additions on arrays.

Subtraction, in one line

Subtraction needs no new algorithm: negate every coefficient of the second polynomial and add. An examiner asking for subtraction is checking that you noticed.

munotes.in72

Adding Two Polynomials

Quick revision

  • Both lists are in decreasing exponent order, so the largest remaining exponent of the answer is at one

of the two current nodes: one pass suffices.

  • Three cases: A's exponent larger, B's larger, or equal. When equal, add the coefficients.
  • If the coefficients cancel to zero, attach nothing: a zero term breaks the representation's invariant.
  • After the loop, attach whatever remains of the unexhausted list. Forgetting the tail is a standard

error.

  • Cost is O(m + n), which is optimal; the array version costs O(larger degree) whether the terms exist

or not.

  • Everything cancelling gives the empty list, which prints as 0 and is a valid answer.
  • Subtraction is negation of the second polynomial followed by addition.

Test yourself

1. Why does one pass suffice to add two polynomials? Because both lists are in decreasing exponent order, so the next term of the answer is always one of the two nodes currently being looked at.

2. Give the three cases inside the merge loop. A's exponent greater: take A's term. B's greater: take B's. Equal: add the coefficients and take the result if it is not zero.

3. What must happen when two coefficients cancel, and why? No node is attached. Keeping a zero term breaks the invariant that no coefficient is zero, and it makes a sparse polynomial less sparse for no benefit.

4. What is the commonest omission after the merge loop? Attaching the remaining terms of whichever list is not yet exhausted.

5. Give the complexity of list addition and of array addition, and say when the difference matters. O(m + n) for lists, where m and n are the term counts; O(larger degree) for arrays. The difference matters for sparse polynomials: adding 3x^1000 + 5 to itself is 2 comparisons against 1,001 additions.

6. How is subtraction implemented? Negate every coefficient of the second polynomial and then add. No separate algorithm is needed.

Contents This chapter on its own page

munotes.in73

Chapter Twenty-Five

Multiplying Polynomials, and What the Representation Costs

Syllabus topic Module 1, "Linked Structures: applications of linked list like polynomial equation"

In one line

Multiplying two polynomials means multiplying every term of one by every term of the other and collecting the terms that share an exponent, which the ordered list makes awkward in a way addition did not.

The arithmetic

To multiply, take each term of P against each term of Q. Coefficients multiply; exponents add.

(5x^4) x (2x^3) = 10x^7

(5x^4) x (1x^0) = 5x^4

So P with m terms and Q with n terms produces m times n partial products. Those partial products then have to be collected, because many of them will share an exponent.

Why this is harder than addition

Addition walked both lists once and the answer came out already in order. Multiplication does not have that luxury:

The partial products arrive out of order. Taking P's terms outer and Q's inner gives exponents 7, 4, then 6, 3, and so on. Nothing is sorted.

Many share an exponent and must be combined. 5x^4 times 2x^0 and 5x^2 times 2x^2 both land on x^4.

So multiplication is m times n multiplications followed by a collection step, and the collection is where the representation makes you work.

Built and run

The approach here is the one to write in an examination: multiply one term of P into a running answer using the addition of chapter 24, which keeps the answer ordered and collected at every step.

class Term:
    def __init__(self, coefficient, exponent, next_term=None):
        self.coefficient = coefficient
        self.exponent = exponent
        self.next = next_term


class Polynomial:
    def __init__(self, terms=()):
        self.head = None
        for coefficient, exponent in sorted(terms, key=lambda t: t[1]):
            if coefficient != 0:
                self.head = Term(coefficient, exponent, self.head)

    def terms(self):
        out, walk = [], self.head
        while walk is not None:
            out.append((walk.coefficient, walk.exponent))
            walk = walk.next
        return out

    def evaluate(self, x):
        total, walk = 0, self.head
        while walk is not None:
            total += walk.coefficient * (x ** walk.exponent)
            walk = walk.next
        return total

    def __str__(self):
        if self.head is None:
            return "0"
        parts, walk = [], self.head
        while walk is not None:
            c, e = walk.coefficient, walk.exponent
            parts.append("%d" % c if e == 0 else
                         ("%dx" % c if e == 1 else "%dx^%d" % (c, e)))
            walk = walk.next
        return " + ".join(parts).replace("+ -", "- ")


def add(p, q):
    result, tail = Polynomial(), None
    a, b = p.head, q.head

    def attach(coefficient, exponent):
        nonlocal tail
        node = Term(coefficient, exponent)
        if tail is None:
            result.head = node
        else:
            tail.next = node
        tail = node

    while a is not None and b is not None:
        if a.exponent > b.exponent:
            attach(a.coefficient, a.exponent)
            a = a.next
        elif b.exponent > a.exponent:
            attach(b.coefficient, b.exponent)
            b = b.next
        else:
            total = a.coefficient + b.coefficient
            if total != 0:
                attach(total, a.exponent)
            a, b = a.next, b.next
    for walk in (a, b):
        while walk is not None:
            attach(walk.coefficient, walk.exponent)
            walk = walk.next
    return result


def multiply(p, q):
    """Multiply, counting the term multiplications performed."""
    answer, products = Polynomial(), 0
    a = p.head
    while a is not None:
        row, b = [], q.head
        while b is not None:
            row.append((a.coefficient * b.coefficient, a.exponent + b.exponent))
            products += 1
            b = b.next
        answer = add(answer, Polynomial(row))
        a = a.next
    return answer, products


p = Polynomial([(5, 2), (3, 0)])
q = Polynomial([(2, 1), (4, 0)])
product, products = multiply(p, q)
print("P(x)   =", p)
print("Q(x)   =", q)
print("P x Q  =", product)
print("term multiplications:", products)
print()
print("independent check at x = 3:")
print("  P(3) x Q(3) =", p.evaluate(3), "x", q.evaluate(3), "=", p.evaluate(3) * q.evaluate(3))
print("  the product polynomial at 3 gives", product.evaluate(3))
print("  they agree:", p.evaluate(3) * q.evaluate(3) == product.evaluate(3))

print()
big_p = Polynomial([(1, 4), (1, 3), (1, 2), (1, 1), (1, 0)])
big_q = Polynomial([(1, 2), (1, 1), (1, 0)])
big, big_products = multiply(big_p, big_q)
print("(%s) x (%s)" % (big_p, big_q))
print("  =", big)
print("  term multiplications:", big_products, "= 5 x 3")
print("  terms in the answer :", len(big.terms()))
munotes.in74

Multiplying Polynomials, and What the Representation Costs

P(x)   = 5x^2 + 3
Q(x)   = 2x + 4
P x Q  = 10x^3 + 20x^2 + 6x + 12
term multiplications: 4

independent check at x = 3:
  P(3) x Q(3) = 48 x 10 = 480
  the product polynomial at 3 gives 480
  they agree: True

(1x^4 + 1x^3 + 1x^2 + 1x + 1) x (1x^2 + 1x + 1)
  = 1x^6 + 2x^5 + 3x^4 + 3x^3 + 3x^2 + 2x + 1
  term multiplications: 15 = 5 x 3
  terms in the answer : 7

The small product can be checked by hand: (5x^2 + 3)(2x + 4) is 10x^3 + 20x^2 + 6x + 12. The program agrees, and the evaluation check at x = 3 agrees independently.

The larger one shows the shape: 5 terms times 3 terms is 15 multiplications, and the answer has 7 terms, because the 15 partial products collapsed onto 7 distinct exponents.

The cost, and the honest verdict

Multiplication is O(m n) multiplications, and that is unavoidable: every pair of terms contributes.

But the collection costs more than it should here. Each row of partial products is merged into the answer, and the answer grows, so the merging is roughly O(m n) again in the best arrangement and worse in a careless one. Written naively, appending every partial product and then sorting and collecting, it is O(m n log(m n)).

Now compare the array representation, indexed by exponent. Multiplication there is:

for each i, for each j: result[i + j] = result[i + j] + p[i] x q[j]

Two loops, no merging, no ordering to maintain, no cancellation to check. It is simpler and faster for a dense polynomial, because the array's index does the collecting for free.

munotes.in75

Multiplying Polynomials, and What the Representation Costs

So the honest verdict, and the one worth writing in an answer:

OperationLinked listArray
Storage, sparseexcellentterrible
Storage, densewasteful (an address per term)excellent
AdditionO(m + n), naturalO(larger degree)
MultiplicationO(m n) plus awkward collectionO(m n), trivially simple
EvaluationO(terms)O(degree)

The linked representation is chosen for sparse polynomials and for addition. It is not chosen because it is better at everything, and multiplication is where it shows.

That is the pattern this whole paper keeps repeating. A structure is a bargain, and part of knowing it is knowing what you paid.

Quick revision

  • Multiplying: coefficients multiply, exponents add. m terms times n terms gives m n partial products.
  • The partial products arrive out of order and many share an exponent, so they must be collected.
  • The clean method is to multiply one term of P into a running answer using polynomial addition, which

keeps the answer ordered and collected throughout.

  • Cost is O(m n) multiplications, plus collection; a naive approach adds a sort and becomes

O(m n log(m n)).

  • The array representation multiplies with two loops and result[i + j] += p[i] * q[j], with the index doing the collecting for free, so it is simpler and faster for dense polynomials. - The linked representation is chosen for sparseness and for addition, not because it is better at everything.

Test yourself

1. What happens to coefficients and exponents when two terms are multiplied? The coefficients multiply and the exponents add.

2. How many partial products does multiplying an m term by an n term polynomial give? m times n.

3. Why is multiplication harder on a linked list than addition was? The partial products come out unordered and many share an exponent, so they must be collected, whereas addition's single merge produced an already ordered, already collected answer.

4. Describe the clean method used in this chapter. Multiply each term of P by the whole of Q to make a row, and add that row into a running answer using polynomial addition, which keeps the answer ordered and collected at every step.

5. Write the array version of polynomial multiplication in one line, and say why it is simpler. result[i + j] = result[i + j] + p[i] * q[j] over all i and j. The index collects terms of equal exponent automatically, so there is no merging, no ordering and no cancellation check.

6. State the honest verdict on the two representations. The linked list is for sparse polynomials and for addition; the array is better for dense polynomials and for multiplication. The linked representation is a bargain, not an improvement in every direction.

Contents This chapter on its own page

munotes.in76

Chapter Twenty-Eight

Traversing Both Ways, and the Applications That Need It

Syllabus topic Module 1, "Linked Structures: ADT of doubly linked list"

In one line

Backward traversal is the same loop from the tail following prev, and it is what makes a cursor possible: a position in the list that can move either way, which is exactly what history and undo need.

The two traversals

forward : walk = head; while walk is not null: use walk.data; walk = walk.next

backward: walk = tail; while walk is not null: use walk.data; walk = walk.prev

Identical shape, opposite ends, opposite field. A singly linked list can do only the first, and to produce the second it must either reverse itself or push every node onto a stack, both of which cost an extra O(n) pass and extra memory.

The cursor

The real gift is not printing backwards. It is that a position in the list can be held and moved in either direction, cheaply.

Hold a node. Move to node.next or node.prev. Both O(1). That is a cursor, and applications that need "the thing before this one" are built on it.

Browser history, built

The practical names this. A browser's back button needs the previous page; forward needs the next; and visiting a new page from the middle of the history throws away everything ahead.

class DNode:
    def __init__(self, data):
        self.prev = None
        self.data = data
        self.next = None


class History:
    """Browser history as a doubly linked list with a cursor."""

    def __init__(self):
        self.current = None

    def visit(self, page):
        """Go to a new page. Everything ahead of the cursor is discarded."""
        node = DNode(page)
        node.prev = self.current
        if self.current is not None:
            self.current.next = node        # drops whatever was ahead
        self.current = node

    def back(self):
        if self.current is None or self.current.prev is None:
            return None                      # nothing behind
        self.current = self.current.prev
        return self.current.data

    def forward(self):
        if self.current is None or self.current.next is None:
            return None                      # nothing ahead
        self.current = self.current.next
        return self.current.data

    def here(self):
        return self.current.data if self.current else None

    def trail(self):
        """The whole history, oldest first, with the cursor marked."""
        walk = self.current
        while walk is not None and walk.prev is not None:
            walk = walk.prev
        out = []
        while walk is not None:
            out.append("[%s]" % walk.data if walk is self.current else walk.data)
            walk = walk.next
        return " ".join(out)


h = History()
for page in ("home", "notes", "data-structures"):
    h.visit(page)
print("after visiting  :", h.trail())
print("back            :", h.back(), "->", h.trail())
print("back            :", h.back(), "->", h.trail())
print("forward         :", h.forward(), "->", h.trail())
print("back at the start:", History().back())
print()
h.visit("syllabus")
print("visit from the middle:", h.trail())
print("forward now gives    :", h.forward())
after visiting  : home notes [data-structures]
back            : notes -> home [notes] data-structures
back            : home -> [home] notes data-structures
forward         : notes -> home [notes] data-structures
back at the start: None

visit from the middle: home notes [syllabus]
forward now gives    : None
munotes.in84

Traversing Both Ways, and the Applications That Need It

Every behaviour of a real back button is there, including the one people forget: visiting a new page from the middle discards the forward history. "data-structures" is gone, and forward returns nothing.

That behaviour falls out of the structure. self.current.next = node overwrites the old forward link, and nothing else refers to what came after, so it is dropped.

Undo and redo, built

The same structure, one behaviour different: undo does not discard, it moves.

class DNode:
    def __init__(self, data):
        self.prev = None
        self.data = data
        self.next = None


class Editor:
    """Text with undo and redo, as a doubly linked list of states."""

    def __init__(self):
        self.current = DNode("")            # the empty document is a state

    def type(self, text):
        """A new state. Anything that had been undone is discarded."""
        node = DNode(self.current.data + text)
        node.prev = self.current
        self.current.next = node
        self.current = node

    def undo(self):
        if self.current.prev is None:
            return False
        self.current = self.current.prev
        return True

    def redo(self):
        if self.current.next is None:
            return False
        self.current = self.current.next
        return True

    def text(self):
        return repr(self.current.data)


e = Editor()
for word in ("Data ", "Structures ", "notes"):
    e.type(word)
print("typed        :", e.text())
print("undo         :", e.undo(), e.text())
print("undo         :", e.undo(), e.text())
print("redo         :", e.redo(), e.text())
print("undo to start:", e.undo(), e.undo(), e.text())
print("undo again   :", e.undo(), "(nothing left to undo)")
print()
e.redo()
e.type("REPLACED")
print("typing after an undo:", e.text())
print("redo now            :", e.redo(), "(the old future was discarded)")
typed        : 'Data Structures notes'
undo         : True 'Data Structures '
undo         : True 'Data '
redo         : True 'Data Structures '
undo to start: True True ''
undo again   : False (nothing left to undo)

typing after an undo: 'Data REPLACED'
redo now            : False (the old future was discarded)

Undo and redo are just prev and next. The whole feature is a cursor on a doubly linked list, and the rule that typing after an undo discards the redo history is the same one the browser had.

Why not an array for these?

Both could be arrays with an index. It would work, and for a browser's history it would be fine.

The doubly linked list wins when entries are removed from the middle. A cache that holds the most recently used items must, on every access, move an item to the front, which means deleting it from wherever it is. On an array that is O(n); with a doubly linked node in hand it is O(1). That structure, a hash table pointing at nodes of a doubly linked list, is how a least-recently-used cache is built, and it is the most common real use of this structure in working software.

Quick revision

  • Backward traversal is the forward loop from the tail, following prev.
  • A singly linked list can produce it only by reversing itself or using a stack, costing an extra pass
munotes.in85

Traversing Both Ways, and the Applications That Need It

and extra memory.

  • The real gift is a cursor: a held position that can move either way in O(1).
  • Browser history: visit pushes a node after the cursor and discards what was ahead; back and forward

move the cursor.

  • Undo and redo are the same structure; typing after an undo discards the redo history for the same

reason.

  • An array with an index would do for these; the doubly linked list wins when items are removed from the

middle, as in a least-recently-used cache.

Test yourself

1. Write the backward traversal. walk = tail; while walk is not null: use walk.data; walk = walk.prev.

2. How would a singly linked list produce a backward traversal, and at what cost? By reversing the list, or by pushing every node onto a stack and popping it. Either costs an extra O(n) pass, and the stack costs O(n) extra memory.

3. What is a cursor and why does the second link make one possible? A held position in the list that can move either way. The backward link makes moving to the previous node O(1), which a singly linked list cannot do at all.

4. In the browser history, what happens to the forward pages when you visit a new page from the middle? They are discarded. The cursor's next is overwritten by the new node and nothing else refers to the old continuation.

5. How are undo and redo implemented here? As moving the cursor to prev and to next over a list of states. Typing creates a new state after the cursor and discards anything that had been undone.

6. Give the case where a doubly linked list clearly beats an array with an index. When items must be removed from the middle. Moving an item to the front of a least-recently-used cache is O(1) with a node in hand and O(n) on an array.

Contents This chapter on its own page

munotes.in86

Chapter Thirty

The Stack: One End, and Why That Is Enough

Syllabus topic Module 1, "Stacks: Stack ADT for Stack"

In one line

A stack allows adding and removing at one end only, so the last thing put in is the first thing taken out, and that single restriction is what makes it the right tool for anything nested.

Last in, first out

Take a pile of plates. You add to the top and you take from the top. The plate you put down most recently is the one you pick up first. Nobody takes a plate from the middle.

That is a stack, and the rule has a name: LIFO, last in, first out.

Real stackThe operation
put a plate on the pilepush
take the top platepop
look at the top plate without taking itpeek
is the pile emptyis_empty

The vocabulary matters in an examination. The end you work at is the top, whichever way you draw it.

Why restricting a list is useful

A stack can do less than a linked list. It is not a weaker linked list, it is a different promise, and three things follow from making that promise.

1. Every operation is O(1). Push, pop and peek all happen at one known place, so nothing is searched and nothing is shifted. Chapters 32 and 33 build it two ways and both give constant time.

2. The structure matches the problem. Anything nested is a stack problem: brackets inside brackets, function calls inside function calls, folders inside folders. Nesting means the most recently opened thing must be closed first, which is exactly LIFO.

3. It cannot be misused. A structure with no way to reach the middle cannot be reached into by accident. That is the same argument as chapter 8's, applied to operations rather than to fields.

The three questions, on a stack

Chapter 6 said to ask three questions of any structure.

Which operations? Push, pop, peek, is_empty. That is all there is.

What do they cost? All O(1), on both representations.

What memory? Exactly the items, plus a little. An array-backed stack reserves capacity; a linked stack pays an address per item.

And the fourth, the refusals, which for a stack are the whole design: no access to the middle, no access to the bottom, no searching, no ordering. If the problem needs any of those, it is not a stack problem.

Seen before it is built

Python's own list already does LIFO if you use only append and pop. This is not the stack this paper builds, but it shows the behaviour in three lines before any structure is written.

pile = []
for plate in ("blue", "green", "red"):
    pile.append(plate)
    print("push %-6s -> pile is now %s" % (plate, pile))

print()
while pile:
    print("pop        -> %-6s, leaving %s" % (pile.pop(), pile))

print()
print("empty now  :", pile == [])
print()
print("what went in : blue, green, red")
print("what came out: red, green, blue")
print("reversed     :", ["blue", "green", "red"] == list(reversed(["red", "green", "blue"])))
munotes.in90

The Stack: One End, and Why That Is Enough

push blue   -> pile is now ['blue']
push green  -> pile is now ['blue', 'green']
push red    -> pile is now ['blue', 'green', 'red']

pop        -> red   , leaving ['blue', 'green']
pop        -> green , leaving ['blue']
pop        -> blue  , leaving []

empty now  : True

what went in : blue, green, red
what came out: red, green, blue
reversed     : True

A stack reverses. Items come out in the opposite order to the way they went in, and that is not a side effect: it is the most useful thing about it. Reversing anything is pushing it all and popping it all.

Where stacks are already working in your programs

Three of these come back later in this paper.

Function calls. Every call pushes a frame holding the local variables and where to return; every return pops it. This is literally called the call stack, and it is why infinite recursion gives a "stack overflow".

Undo. Chapter 28 used a doubly linked list because redo was wanted too. Undo alone is a stack.

Expression evaluation. Chapters 37 and 38 convert and evaluate expressions with a stack, because brackets nest.

Depth first search. Chapter 95 explores a graph with a stack, and chapter 62 walks a tree with one. The reason is the same in both: going deeper is nesting.

Quick revision

  • A stack allows insertion and removal at one end, the top, only. Last in, first out.
  • Operations: push, pop, peek, is_empty.
  • The restriction is the feature: every operation is O(1), the structure fits anything nested, and the

middle cannot be reached into by mistake.

  • A stack refuses the middle, the bottom, searching and ordering. A problem needing any of those is not

a stack problem.

  • A stack reverses: pushing everything and popping everything gives the reverse order.
  • Already at work in the call stack, undo, expression evaluation and depth first search.

Test yourself

1. What does LIFO stand for and what does it mean? Last in, first out: the most recently pushed item is the next one popped.

2. Name the four stack operations and what each does. Push adds at the top, pop removes and returns the top, peek returns the top without removing it, and is_empty reports whether there is anything there.

3. Give three consequences of restricting a list to one end. Every operation is O(1); the structure matches any nested problem; and the middle cannot be reached into by accident.

4. What does a stack refuse to do? Reach the middle, reach the bottom, search, and impose any order other than the order of insertion.

munotes.in91

The Stack: One End, and Why That Is Enough

5. Why does a stack reverse a sequence? Because the first item pushed is at the bottom and is therefore popped last, so the output order is the exact reverse of the input order.

6. Give three places a stack is already at work in an ordinary program. The call stack holding function frames and return addresses; the undo feature of an editor; and expression evaluation, where brackets nest.

Contents This chapter on its own page

munotes.in92

Chapter Thirty-One

The Stack ADT: Push, Pop, and the Errors

Syllabus topic Module 1, "Stacks: Stack ADT for Stack"

In one line

The stack ADT is push, pop, peek, is_empty and size, with two error conditions that are part of the specification and not an implementation detail.

The ADT

A Stack holds a sequence of items and permits access at one end only, called the top.

OperationNeedsReturnsDoesWhen it cannot
Stack()nothingan empty stackcreates itnever fails
push(item)an itemnothingputs it on the topoverflow if the stack is full
pop()nothingan itemremoves and returns the topunderflow if the stack is empty
peek()nothingan itemreturns the top, leaves iterror if the stack is empty
is_empty()nothingtrue or falseis there nothing in itnever fails
size()nothinga numberhow many itemsnever fails

That table is the answer to "write the ADT for a stack". Nothing in it mentions an array or a linked list, which is the point of chapter 7.

Underflow and overflow

These two words are examined, so they are worth being exact about.

Underflow is popping or peeking an empty stack. There is no top, so there is nothing to return. It can happen on any implementation, because emptiness is a property of the stack and not of its storage.

Overflow is pushing onto a full stack. It can happen only when the stack has a fixed capacity, which is the array implementation. A linked stack has no maximum and cannot overflow until the machine itself runs out of memory.

So the honest statement, which is what an examiner wants: underflow applies to every stack; overflow applies to a stack of fixed size.

Why they are errors and not quiet answers

The lazy implementation returns a special value, usually None or -1, instead of raising an error. It is worth seeing why that is a defect rather than a style choice.

class QuietStack:
    """Returns None on underflow. Looks harmless."""

    def __init__(self):
        self._items = []

    def push(self, item):
        self._items.append(item)

    def pop(self):
        if not self._items:
            return None
        return self._items.pop()


class LoudStack:
    """Raises on underflow, as the ADT says."""

    def __init__(self):
        self._items = []

    def push(self, item):
        self._items.append(item)

    def pop(self):
        if not self._items:
            raise IndexError("pop from an empty stack: underflow")
        return self._items.pop()


quiet = QuietStack()
quiet.push("a")
quiet.push(None)             # a legitimate item that happens to be None

print("quiet stack, popping three times from a stack of two:")
for _ in range(3):
    print("   got", repr(quiet.pop()))
print("   which of those was underflow? There is no way to tell.")

print()
loud = LoudStack()
loud.push("a")
print("loud stack:")
print("   got", repr(loud.pop()))
try:
    loud.pop()
except IndexError as e:
    print("   underflow reported:", e)
quiet stack, popping three times from a stack of two:
   got None
   got 'a'
   got None
   which of those was underflow? There is no way to tell.

loud stack:
   got 'a'
   underflow reported: pop from an empty stack: underflow
munotes.in93

The Stack ADT: Push, Pop, and the Errors

Two of those three None values mean different things: one is a real item that was pushed, one is the stack saying it is empty. The caller cannot distinguish them, so the error is carried forward into whatever uses the value, and it surfaces somewhere else entirely.

That is the same argument as chapter 9's, and it is why the ADT specifies the error rather than leaving it to the implementer.

A note on what pop returns

There are two conventions and both appear in textbooks:

Pop removes and returns the top. One operation. This is what this book uses and what most languages do.

Pop removes; top or peek returns. Two operations, so removing without looking is possible. C++'s std::stack works this way.

An examination answer may use either, as long as it is consistent and says which. What is wrong is a pop that returns the item in the specification and does not in the code.

The ADT in the form the examination wants

If the question says "write the stack ADT", the expected answer is roughly:

Stack: a collection of items with access at one end, the top.

push(item) : add item to the top. Overflow if full.

pop() : remove and return the top item. Underflow if empty.

peek() : return the top item without removing it. Error if empty.

is_empty() : true when the stack holds nothing.

size() : the number of items held.

All operations are O(1).

Adding the cost line is worth a mark and almost nobody writes it.

Quick revision

  • The stack ADT: push, pop, peek, is_empty, size. All O(1).
  • It names no array and no linked list.
  • Underflow: popping or peeking an empty stack. Applies to every implementation.
  • Overflow: pushing onto a full stack. Applies only to a fixed capacity stack, so to the array version

and not the linked one.

  • Errors are part of the specification. Returning None instead makes a real None item and an empty stack

indistinguishable to the caller.

  • Two conventions for pop exist; either is acceptable if stated and applied consistently.

Test yourself

1. Write the stack ADT. A collection with access at one end. push(item) adds at the top, overflow if full; pop() removes and returns the top, underflow if empty; peek() returns the top without removing it, error if empty; is_empty(); size(). All operations are O(1).

2. Define underflow and overflow, and say which implementations each applies to. Underflow is popping or peeking an empty stack and applies to every implementation. Overflow is pushing onto a full stack and applies only to a fixed capacity stack, which is the array version.

munotes.in94

The Stack ADT: Push, Pop, and the Errors

3. Why is returning None on underflow a defect? Because None may be a legitimate item. The caller cannot tell a real value from the stack reporting emptiness, so the mistake travels to somewhere unrelated before it causes a visible failure.

4. Can a linked stack overflow? Not in the sense the ADT means. It has no fixed capacity, so it fails only when the machine runs out of memory altogether.

5. Give the two conventions for pop, and what makes an answer wrong. Pop removes and returns the top; or pop removes while a separate top or peek returns it. An answer is wrong when the specification and the code disagree, not when it picks one convention.

6. What line does almost nobody write in this answer, and why is it worth a mark? That every operation is O(1). It is part of what the ADT promises and it is what makes the stack worth choosing.

Contents This chapter on its own page

munotes.in95

Chapter Thirty-Two

A Stack on an Array, With Peek

Syllabus topic Computer Science Practical 3, Module 2, "Implement push, pop, peek using arrays or linked lists"

In one line

An array stack keeps the items in a fixed block with an integer marking the top, so push writes at top + 1 and pop reads at top, both O(1).

The design

Two fields: the block of cells, and an integer top.

The usual convention, and the one to state in an answer: top is the index of the topmost item, and -1 means empty. An empty stack has no topmost item, so -1 is the value just below the first cell.

empty : top = -1

after push : top = top + 1, then cells[top] = item

after pop : item = cells[top], then top = top - 1

full : top = capacity - 1

The order inside push and pop is the mirror of each other, and getting it wrong by one is the classic error: incrementing after writing overwrites the item below.

Built and run

class ArrayStack:
    """A stack in a fixed block, with top as the index of the topmost item."""

    def __init__(self, capacity):
        self.capacity = capacity
        self.cells = [None] * capacity
        self.top = -1                       # -1 means empty

    def is_empty(self):
        return self.top == -1

    def is_full(self):
        return self.top == self.capacity - 1

    def size(self):
        return self.top + 1

    def push(self, item):
        if self.is_full():
            raise OverflowError("push onto a full stack: overflow (capacity %d)"
                                % self.capacity)
        self.top += 1
        self.cells[self.top] = item

    def pop(self):
        if self.is_empty():
            raise IndexError("pop from an empty stack: underflow")
        item = self.cells[self.top]
        self.cells[self.top] = None          # so a popped item is not kept alive
        self.top -= 1
        return item

    def peek(self):
        if self.is_empty():
            raise IndexError("peek at an empty stack: underflow")
        return self.cells[self.top]

    def snapshot(self):
        return self.cells[:self.top + 1]


s = ArrayStack(4)
print("new stack: empty =", s.is_empty(), "| size =", s.size(), "| top =", s.top)

for item in ("A", "B", "C"):
    s.push(item)
    print("push %s -> %-16s top = %d, size = %d"
          % (item, str(s.snapshot()), s.top, s.size()))

print()
print("peek   ->", s.peek(), "| size is still", s.size())
print("pop    ->", s.pop(), "| now", s.snapshot())
print("pop    ->", s.pop(), "| now", s.snapshot())

print()
s.push("X")
s.push("Y")
s.push("Z")
print("filled to capacity:", s.snapshot(), "| full =", s.is_full())
try:
    s.push("one too many")
except OverflowError as e:
    print("overflow reported:", e)

print()
empty = ArrayStack(2)
for operation in ("pop", "peek"):
    try:
        getattr(empty, operation)()
    except IndexError as e:
        print("%-5s on an empty stack ->" % operation, e)
new stack: empty = True | size = 0 | top = -1
push A -> ['A']            top = 0, size = 1
push B -> ['A', 'B']       top = 1, size = 2
push C -> ['A', 'B', 'C']  top = 2, size = 3

peek   -> C | size is still 3
pop    -> C | now ['A', 'B']
pop    -> B | now ['A']

filled to capacity: ['A', 'X', 'Y', 'Z'] | full = True
overflow reported: push onto a full stack: overflow (capacity 4)

pop   on an empty stack -> pop from an empty stack: underflow
peek  on an empty stack -> peek at an empty stack: underflow
munotes.in96

A Stack on an Array, With Peek

Both error conditions of chapter 31 are real here: underflow on the empty stack, and overflow, which the linked version of the next chapter cannot produce.

Two details worth marks

peek does not change top. It is the one operation that reads without moving, and writing self.top -= 1 into it by habit is a common slip. The run above checks it: size is still 3 after peek.

The popped cell is cleared. self.cells[self.top] = None is not required for correctness, because top already says that cell is not in use. It is there so the popped object is not kept alive by a reference nobody can reach. In C the equivalent is that the cell still holds a stale pointer, and reading it is a real bug.

The capacity problem

The array stack's one weakness is the number you have to choose at the start.

Too small and it overflows on valid input. Too large and memory sits reserved and unused. And there is often no good way to know: the depth of a stack used for bracket matching depends on the input.

Two answers exist.

Grow it, as Python's list does: when full, allocate a bigger block, copy, continue. This makes push amortised O(1) rather than O(1), by the argument of chapter 12, and it means overflow effectively disappears.

Use links, which is the next chapter, and have no capacity at all.

Quick revision

  • Fields: a fixed block of cells, and top, the index of the topmost item, with -1 meaning empty.
  • push: increment top, then write. pop: read, then decrement. The order is the mirror; reversing either

is an off-by-one that overwrites or re-reads.

  • full is top == capacity - 1; size is top + 1.
  • peek reads without changing top.
  • Both error conditions are real here: underflow and overflow.
  • Clear the popped cell so a dead item is not kept alive, and in C so a stale pointer is not left.
  • The capacity must be chosen in advance, which is the weakness; growing the block or using links are

the two answers.

Test yourself

1. What does top hold, and what value means the stack is empty? The index of the topmost item. -1 means empty, because an empty stack has no topmost item.

2. Write push and pop in terms of top. push: check full, top = top + 1, then cells[top] = item. pop: check empty, item = cells[top], then top = top - 1.

munotes.in97

A Stack on an Array, With Peek

3. What goes wrong if push writes before incrementing? It overwrites the item already at the top instead of adding above it, so the stack silently loses an item on every push.

4. Give the conditions for full and for size. Full is top == capacity - 1. Size is top + 1.

5. Which error can this implementation raise that a linked stack cannot, and why? Overflow, because the capacity is fixed at creation. A linked stack has no maximum until the machine runs out of memory.

6. Why clear the cell when popping, given that top already marks it unused? So the popped object is not kept alive by an unreachable reference; and in C, so a stale pointer is not left in the cell to be read by mistake.

Contents This chapter on its own page

munotes.in98

Chapter Thirty-Four

What a Stack Is Good and Bad At

Syllabus topic Module 1, "Stacks: Stack ADT for Stack, Advantages & Disadvantages"

In one line

A stack is good at everything that nests and bad at everything that does not, and that is a design decision rather than a defect.

The advantages

1. Every operation is O(1), on both implementations. No searching, no shifting, no walking. Chapters 32 and 33 built both and neither has a loop in push, pop or peek.

2. It matches nested problems exactly. Brackets, function calls, folders, expression evaluation, depth first search. In all of these the most recently opened thing must be closed first, which is the definition of LIFO. The structure is not merely usable; it is the shape of the problem.

3. It is simple enough to be correct. Five operations, no positions, no ordering. There is very little to get wrong, which matters for something the call stack of every program depends on.

4. It cannot be misused. No operation reaches the middle, so no caller can.

5. It reverses for free. Push everything, pop everything. Chapter 33 used this in three lines.

The disadvantages

1. No access except the top. The second item is unreachable without removing the first. If you need to look at the middle, you need a different structure.

2. No searching. Finding whether an item is present means popping until you find it, which destroys the stack. A stack has no contains.

3. No ordering beyond arrival. A stack does not know which item is largest or smallest. That is what chapter 78's priority queue is for.

4. Fixed capacity, on the array version. Overflow is real, and the maximum depth must be guessed. The linked version trades this for an address per item.

5. Traversal destroys it. Reading all the items means popping them all, and then you no longer have them unless you pushed them onto a second stack, which reverses them again.

The last one is worth demonstrating, because it is the disadvantage students most often miss.

class LinkedStack:
    def __init__(self):
        self._items = []

    def push(self, item):
        self._items.append(item)

    def pop(self):
        return self._items.pop()

    def is_empty(self):
        return not self._items

    def size(self):
        return len(self._items)


s = LinkedStack()
for item in ("A", "B", "C", "D"):
    s.push(item)
print("size before reading everything:", s.size())

seen = []
while not s.is_empty():
    seen.append(s.pop())
print("what we read                  :", seen)
print("size after reading            :", s.size())
print("the stack is gone             :", s.is_empty())

restored = LinkedStack()
for item in seen:
    restored.push(item)
restored_order = []
while not restored.is_empty():
    restored_order.append(restored.pop())
print()
print("pushing what we read into a second stack and popping gives:", restored_order)
print("which is the original order   :", restored_order == ["A", "B", "C", "D"])
size before reading everything: 4
what we read                  : ['D', 'C', 'B', 'A']
size after reading            : 0
the stack is gone             : True

pushing what we read into a second stack and popping gives: ['A', 'B', 'C', 'D']
which is the original order   : True
munotes.in102

What a Stack Is Good and Bad At

Reading a stack empties it. Two stacks restore the order, because reversing twice is the identity, and that trick is the standard way to copy a stack without a third structure.

The table

Stack
push, pop, peekO(1)
access item inot possible
searchnot possible without destroying it
traversedestroys it, unless a second stack is used
find the largestnot possible
memory, array versionfixed block, capacity reserved
memory, linked versionone address per item

Look at how many rows say "not possible". That is the point of the structure, and an examination answer that lists them as weaknesses without saying so has missed the argument.

When a stack is the wrong choice

Three signs, and each points at a specific later chapter.

You need the item that arrived first. That is a queue, chapter 40. You need the largest or most urgent item. That is a priority queue, chapter 78. You need to search, or to look at the middle. That is a list, a tree or a hash table.

Quick revision

  • Advantages: every operation O(1); an exact match for nested problems; simple enough to be correct;

cannot be reached into; reverses for free.

  • Disadvantages: no access except the top; no search; no ordering by value; fixed capacity on the array

version; traversal destroys it.

  • The refusals are deliberate, and saying so is part of the answer.
  • Reading a stack empties it; pushing what you read onto a second stack restores the original order,

because reversing twice is the identity.

  • Wrong choice when you need the oldest item (queue), the most urgent (priority queue), or the middle (a

list, tree or hash table).

Test yourself

1. Give three advantages of a stack. Every operation is O(1); it matches nested problems exactly, since the most recently opened thing must close first; and it is simple enough to be correct, with no positions or ordering to get wrong.

2. Give three disadvantages. Only the top is reachable; there is no search; and traversal destroys the stack. Fixed capacity on the array version is a fourth.

3. Why is "no access to the middle" not simply a defect? Because it is the restriction that makes every operation O(1) and makes the structure impossible to misuse. It is the design, not an omission.

4. How do you read every item of a stack and still have the stack afterwards? Pop into a second stack and then pop that back into the first. Reversing twice restores the original order.

5. A problem needs the item that has been waiting longest. Is a stack right? No. That is first in, first out, which is a queue.

munotes.in103

What a Stack Is Good and Bad At

6. Can a stack tell you whether an item is present? Not without popping until it is found, which destroys the stack. A stack has no contains operation.

Contents This chapter on its own page

munotes.in104

Chapter Thirty-Five

Balanced Delimiters, and Why a Counter Is Not Enough

Syllabus topic Module 1, "Stacks: Applications of stack like balanced delimiter"

In one line

Checking that brackets are balanced needs a stack, because a counter can tell you how many are open but not which kind, and mismatched kinds are the whole problem.

The problem

A string of brackets is balanced when every opening bracket has a matching closing bracket of the same kind, and the pairs do not cross.

StringBalanced?Why
( )yes
( [ ] )yesnested correctly
( [ ) ]nothe pairs cross
( (noone never closes
) (nocloses before it opens
( ] )nowrong kind

The counter that fails

The obvious first attempt is a counter: add one for every opening bracket, subtract one for every closing bracket, and check that it never goes negative and ends at zero.

def counter_check(text):
    """Count openings and closings. Looks right; is not."""
    count = 0
    for character in text:
        if character in "([{":
            count += 1
        elif character in ")]}":
            count -= 1
            if count < 0:
                return False
    return count == 0


cases = ["()", "([])", "([)]", "((", ")(", "(]", "{[()]}", "{[(])}"]
for case in cases:
    print("counter says %-6s for %r" % (counter_check(case), case))
counter says True   for '()'
counter says True   for '([])'
counter says True   for '([)]'
counter says False  for '(('
counter says False  for ')('
counter says True   for '(]'
counter says True   for '{[()]}'
counter says True   for '{[(])}'

Three of those are wrong, and they are the interesting three.

([)] is accepted and should not be. The pairs cross: the square bracket opened inside the round one and closed outside it.

(] is accepted and should not be. A round bracket was closed by a square one.

The counter got {[(])} right by luck, not by reasoning: it went negative at a point that happened to catch it.

The reason is simple and worth stating exactly: a counter remembers how many brackets are open and forgets which kinds they were. The kind is exactly what the crossing cases turn on.

The stack that works

A stack remembers the kinds, in order, and the ordering is what nesting means.

for each character:

if it opens : push it

if it closes: if the stack is empty, fail

pop; if the popped bracket does not match this one, fail

at the end: balanced exactly when the stack is empty

Three failures, and an examination answer needs all three: a closing bracket with nothing open, a closing bracket of the wrong kind, and something still open at the end.

def stack_check(text):
    """Balanced delimiters with a stack. Returns (balanced, reason)."""
    pairs = {")": "(", "]": "[", "}": "{"}
    stack = []
    for position, character in enumerate(text):
        if character in "([{":
            stack.append(character)
        elif character in ")]}":
            if not stack:
                return False, "'%s' at position %d closes nothing" % (character, position)
            opened = stack.pop()
            if opened != pairs[character]:
                return False, ("'%s' at position %d closes a '%s'"
                               % (character, position, opened))
    if stack:
        return False, "'%s' is never closed" % stack[-1]
    return True, "balanced"


cases = ["()", "([])", "([)]", "((", ")(", "(]", "{[()]}", "{[(])}",
         "", "a(b[c]d)e", "(()(()))"]
for case in cases:
    balanced, reason = stack_check(case)
    print("%-10r %-6s %s" % (case, balanced, reason))
munotes.in105

Balanced Delimiters, and Why a Counter Is Not Enough

'()'       True   balanced
'([])'     True   balanced
'([)]'     False  ')' at position 2 closes a '['
'(('       False  '(' is never closed
')('       False  ')' at position 0 closes nothing
'(]'       False  ']' at position 1 closes a '('
'{[()]}'   True   balanced
'{[(])}'   False  ']' at position 3 closes a '('
''         True   balanced
'a(b[c]d)e' True   balanced
'(()(()))' True   balanced

Every case is now right, including the three the counter got wrong, and the answer says why rather than just no. Two more worth noting: the empty string is balanced, which is correct and is the kind of edge case an examiner adds; and other characters are simply ignored, which is what makes it usable on real text.

Why the stack is the right structure, in one sentence

Because the bracket that must close next is always the one that opened most recently, and that is the definition of last in, first out.

That sentence is worth memorising. It is the answer to "why a stack" for this problem and for every other nested problem in the paper.

The cost

One pass over the text, with one push or one pop per bracket. O(n) time.

The stack holds at most the depth of nesting, so O(d) memory, where d is the deepest nesting. For (((((...))))) that is O(n); for normal text it is small.

Quick revision

  • Balanced means every opening bracket is matched by a closing bracket of the same kind, without

crossing.

  • A counter fails: it remembers how many are open and forgets which kinds, so it accepts ([)] and

(].

  • The stack algorithm: push openings; on a closing bracket, fail if empty, else pop and compare kinds;

at the end, balanced exactly when the stack is empty.

  • Three failure cases: closing with nothing open, closing the wrong kind, and something left open.
  • The empty string is balanced; other characters are ignored.
  • Why a stack: the bracket that must close next is always the one that opened most recently.
  • O(n) time, O(d) memory where d is the deepest nesting.

Test yourself

1. Why does a counter accept ([)]? Because it only counts how many brackets are open. It never records that the open bracket was round, so it cannot see that a square bracket closed it.

munotes.in106

Balanced Delimiters, and Why a Counter Is Not Enough

2. Give the three failure cases the stack algorithm detects. A closing bracket when the stack is empty; a closing bracket whose kind does not match the popped opening bracket; and a non-empty stack at the end.

3. Write the algorithm in four lines. For each character: push it if it opens. If it closes, fail when the stack is empty, else pop and fail when the kinds do not match. At the end, balanced exactly when the stack is empty.

4. In one sentence, why is a stack the right structure here? Because the bracket that must close next is always the one that opened most recently, which is last in, first out.

5. Is the empty string balanced? Is a(b[c]d)e? Both are. The empty string has nothing unmatched, and characters that are not brackets are ignored.

6. Give the time and memory costs, and say what determines the memory. O(n) time, one pass with one push or pop per bracket. O(d) memory, where d is the deepest nesting, because that is the most the stack ever holds at once.

Contents This chapter on its own page

munotes.in107

Chapter Thirty-Six

Infix, Prefix and Postfix

Syllabus topic Module 1, "Stacks: Applications of stack like prefix to postfix notation"

In one line

The three notations differ only in where the operator sits, and postfix and prefix need no brackets and no precedence rules at all, which is why machines use them.

The three notations

NotationOperator sitsExample
Infixbetween its operandsA + B
Prefix (Polish)before its operands+ A B
Postfix (Reverse Polish)after its operandsA B +

Infix is what people write. The other two are what machines use, and the reason is worth understanding rather than accepting.

Why infix needs rules and the others do not

Write A + B * C. What does it mean?

It is ambiguous as written. It could be (A + B) C or A + (B C). To read it you need two extra pieces of knowledge that are nowhere in the string:

Precedence. Multiplication binds tighter than addition, so it is A + (B * C). Associativity. For equal precedence, which side groups first. A - B - C is (A - B) - C, so minus is left associative. Exponent is right associative: A ^ B ^ C is A ^ (B ^ C).

And when the rules give the wrong answer you need brackets to override them.

Now write the same expression in postfix:

A + (B x C) = A B C x +

(A + B) x C = A B + C x

Two different strings. No brackets, no precedence rules, no ambiguity. The order of the symbols alone determines the meaning, which is exactly what a machine wants: it can be evaluated in one pass with a stack and no lookahead.

That is the answer to "why convert to postfix at all", and it is the examinable point of this chapter.

Converting by hand

The reliable hand method is fully bracket, then move the operators.

Take A + B * C - D.

Step 1, bracket every operation in precedence order, innermost first:

A + B x C - D

= A + (B x C) - D

= (A + (B x C)) - D

= ((A + (B x C)) - D)

Step 2 for postfix: move each operator to just after its closing bracket, then drop the brackets.

((A + (B x C)) - D)

= ((A (B C x) +) D -)

= A B C x + D -

Step 2 for prefix: move each operator to just before its opening bracket, then drop the brackets.

((A + (B x C)) - D)

= (- (+ A (x B C)) D)

= - + A x B C D

That method is slow but it is reliable under examination conditions, and it is the one to use when you are checking your own stack-based answer.

munotes.in108

Infix, Prefix and Postfix

The conversions, run

The program below converts by building the expression tree implicitly through full bracketing, so the three forms can be printed together and checked against each other.

def to_forms(expression):
    """Given a fully bracketed infix expression as nested tuples, print all three."""

    def infix(node):
        if not isinstance(node, tuple):
            return node
        left, operator, right = node
        return "(%s %s %s)" % (infix(left), operator, infix(right))

    def prefix(node):
        if not isinstance(node, tuple):
            return node
        left, operator, right = node
        return "%s %s %s" % (operator, prefix(left), prefix(right))

    def postfix(node):
        if not isinstance(node, tuple):
            return node
        left, operator, right = node
        return "%s %s %s" % (postfix(left), postfix(right), operator)

    return infix(node=expression), prefix(expression), postfix(expression)


cases = [
    ("A + B",              ("A", "+", "B")),
    ("A + B x C",          ("A", "+", ("B", "x", "C"))),
    ("(A + B) x C",        (("A", "+", "B"), "x", "C")),
    ("A + B x C - D",      (("A", "+", ("B", "x", "C")), "-", "D")),
    ("(A + B) x (C - D)",  (("A", "+", "B"), "x", ("C", "-", "D"))),
]

print("%-20s %-24s %-18s %s" % ("meaning", "infix, bracketed", "prefix", "postfix"))
for name, tree in cases:
    infix_form, prefix_form, postfix_form = to_forms(tree)
    print("%-20s %-24s %-18s %s" % (name, infix_form, prefix_form, postfix_form))
meaning              infix, bracketed         prefix             postfix
A + B                (A + B)                  + A B              A B +
A + B x C            (A + (B x C))            + A x B C          A B C x +
(A + B) x C          ((A + B) x C)            x + A B C          A B + C x
A + B x C - D        ((A + (B x C)) - D)      - + A x B C D      A B C x + D -
(A + B) x (C - D)    ((A + B) x (C - D))      x + A B - C D      A B + C D - x

Read rows two and three together. The infix strings differ only by brackets; the postfix strings differ by the order of the symbols, with no brackets at all. That is the whole idea.

The precedence table to memorise

For this paper, the operators and their precedence, highest first:

PrecedenceOperatorsAssociativity
highest^ (exponent)right
x / %left
lowest+ -left

Brackets are not an operator; they override the table.

The one to be careful about is ^. It is right associative, so 2 ^ 3 ^ 2 is 2 ^ (3 ^ 2), which is 2 to the 9th, 512, and not (2 ^ 3) ^ 2, which is 64. Examiners use this.

munotes.in109

Infix, Prefix and Postfix

Quick revision

  • Infix: operator between operands. Prefix (Polish): operator before. Postfix (Reverse Polish): operator

after.

  • Infix is ambiguous without precedence, associativity and brackets; the string alone does not say what

it means.

  • Prefix and postfix need none of those: the order of symbols fixes the meaning.
  • Hand method: fully bracket in precedence order, then move each operator just after its closing bracket

for postfix, or just before its opening bracket for prefix, and drop the brackets.

  • Precedence: ^ highest and right associative, then x / %, then + -, both left associative.
  • 2 ^ 3 ^ 2 is 512, not 64.

Test yourself

1. Write A + B * C in prefix and postfix. Prefix + A x B C; postfix A B C x +.

2. Write (A + B) * C in postfix, and say how it differs from the answer above. A B + C x. The symbols are in a different order; no brackets are needed to tell the two apart.

3. Why do prefix and postfix need no brackets? Because the position of each operator relative to its operands fixes which operands it applies to, so there is no ambiguity for precedence or brackets to resolve.

4. Give the hand method for converting infix to postfix. Fully bracket the expression in precedence order, move each operator to just after its closing bracket, then remove the brackets.

5. State the precedence and associativity of the operators in this paper. ^ is highest and right associative; x, / and % are next and left associative; + and - are lowest and left associative.

6. Evaluate 2 ^ 3 ^ 2 and say why the answer is not 64.

  1. Exponent is right associative, so it is 2 ^ (3 ^ 2), which is 2 to the 9th, not (2 ^ 3) ^ 2.

Contents This chapter on its own page

munotes.in110

Chapter Thirty-Seven

Infix to Postfix With a Stack

Syllabus topic Module 1, "Stacks: Applications of stack like prefix to postfix notation"

In one line

Operands go straight to the output and operators wait on a stack until an operator of the same or higher precedence arrives, which is the algorithm known as shunting yard.

The algorithm

for each token:

operand : send it to the output

'(' : push it

')' : pop to the output until '(' is popped, and discard the '('

operator o : while the stack top is an operator of higher precedence,

or equal precedence and o is left associative: pop it to the output

then push o

at the end : pop everything remaining to the output

The one line that carries all the difficulty is the operator rule, and it says: an operator waiting on the stack goes to the output when something arrives that does not need to wait for it.

Associativity is why the rule says "equal precedence and left associative". For A - B - C, the first minus must come out before the second is pushed, because minus groups leftwards. For A ^ B ^ C it must not, because exponent groups rightwards.

Built, with the trace printed

PRECEDENCE = {"+": 1, "-": 1, "x": 2, "/": 2, "%": 2, "^": 3}
RIGHT_ASSOCIATIVE = {"^"}


def to_postfix(tokens, trace=False):
    """Shunting yard. Returns the postfix tokens."""
    output, stack = [], []
    if trace:
        print("%-6s %-24s %s" % ("token", "output", "stack"))
    for token in tokens:
        if token not in PRECEDENCE and token not in "()":
            output.append(token)
        elif token == "(":
            stack.append(token)
        elif token == ")":
            while stack and stack[-1] != "(":
                output.append(stack.pop())
            if not stack:
                raise ValueError("a ')' with no matching '('")
            stack.pop()                                  # discard the '('
        else:
            while (stack and stack[-1] != "("
                   and (PRECEDENCE[stack[-1]] > PRECEDENCE[token]
                        or (PRECEDENCE[stack[-1]] == PRECEDENCE[token]
                            and token not in RIGHT_ASSOCIATIVE))):
                output.append(stack.pop())
            stack.append(token)
        if trace:
            print("%-6s %-24s %s" % (token, " ".join(output), " ".join(stack)))
    while stack:
        if stack[-1] == "(":
            raise ValueError("a '(' that is never closed")
        output.append(stack.pop())
    if trace:
        print("%-6s %-24s %s" % ("end", " ".join(output), ""))
    return output


print("A + B x C - D")
to_postfix("A + B x C - D".split(), trace=True)

print()
# The tokens are split on spaces, so a bracket MUST be its own token.
# "(A + B)".split() gives "(A" and "B)", and the brackets are then read as
# part of the operand names: the first draft of this listing printed
# "(A B) C x +" and the gate caught it.
for infix in ["A + B x C", "( A + B ) x C", "A + B x C - D",
              "( A + B ) x ( C - D )", "A ^ B ^ C", "A - B - C",
              "A x ( B + C ) / D"]:
    print("%-22s -> %s" % (infix, " ".join(to_postfix(infix.split()))))

print()
for bad in ["A + B )", "( A + B"]:
    try:
        to_postfix(bad.split())
    except ValueError as e:
        print("%-12s refused: %s" % (bad, e))
munotes.in111

Infix to Postfix With a Stack

A + B x C - D
token  output                   stack
A      A
+      A                        +
B      A B                      +
x      A B                      + x
C      A B C                    + x
-      A B C x +                -
D      A B C x + D              -
end    A B C x + D -

A + B x C              -> A B C x +
( A + B ) x C          -> A B + C x
A + B x C - D          -> A B C x + D -
( A + B ) x ( C - D )  -> A B + C D - x
A ^ B ^ C              -> A B C ^ ^
A - B - C              -> A B - C -
A x ( B + C ) / D      -> A B C + x D /

A + B )      refused: a ')' with no matching '('
( A + B      refused: a '(' that is never closed

Follow the trace at the -. The stack held + x. Minus has lower precedence than both, so both were popped to the output before minus was pushed. That single step is the algorithm.

Now compare the last two rows of the conversions, which is where associativity shows:

A ^ B ^ C gives A B C ^ ^. The first ^ stayed on the stack when the second arrived, because exponent is right associative, so the rightmost exponent is applied first.

A - B - C gives A B - C -. The first - came out when the second arrived, because minus is left associative.

Those two lines are the standard examination trap, and the program settles them.

Checking the conversions

The conversions above agree with chapter 36's, which were produced by a completely different method (walking a fully bracketed tree). Two independent methods agreeing is worth more than either one checked by eye.

Row by row: A + B x C gave A B C x + in both; ( A + B ) x C gave A B + C x in both; A + B x C - D gave A B C x + D - in both; ( A + B ) x ( C - D ) gave A B + C D - x in both.

The cost

Each token is pushed at most once and popped at most once, so the work is proportional to the number of tokens: O(n) time. The stack holds at most the operators currently waiting, so O(n) memory in the worst case, which is an expression like ((((A)))).

munotes.in112

Infix to Postfix With a Stack

Quick revision

  • Operands go straight to the output; operators wait on a stack.
  • On an operator, pop while the top has higher precedence, or equal precedence and the new operator is

left associative; then push.

  • ( is pushed; ) pops to the output until ( is found, and the ( is discarded, not output.
  • At the end, pop everything left; a ( still there means it was never closed.
  • A ^ B ^ C gives A B C ^ ^ (right associative) and A - B - C gives A B - C - (left

associative).

  • O(n) time, O(n) memory in the worst case.

Test yourself

1. What happens to an operand, and what happens to an operator? An operand goes straight to the output. An operator waits on the stack until an operator arrives that does not need to wait for it.

2. State the popping rule for an incoming operator o. Pop while the stack top is an operator with higher precedence, or with equal precedence when o is left associative. Then push o.

3. Convert (A + B) x C and show why brackets disappear. A B + C x. The ( is pushed and discarded when the ) arrives; the ) forces the + out before the x is considered, so the grouping is carried by the order of the output symbols.

4. Convert A ^ B ^ C and A - B - C, and explain the difference. A B C ^ ^ and A B - C -. Exponent is right associative so the waiting ^ is not popped when a second arrives; minus is left associative so the waiting - is.

5. What are the two bracket errors, and when is each detected? A ) with nothing matching, found when the stack empties while searching for a (; and a ( never closed, found when it is still on the stack at the end.

6. Give the time and memory costs and say why. O(n) time, because each token is pushed and popped at most once. O(n) memory in the worst case, since the stack may hold every operator, as in a deeply bracketed expression.

Contents This chapter on its own page

munotes.in113

Chapter Thirty-Eight

Evaluating a Postfix Expression

Syllabus topic Computer Science Practical 3, Module 2, "Convert expressions from prefix to postfix and evaluate them"

In one line

To evaluate postfix, push every operand and on every operator pop two, apply, and push the result; the answer is the single value left at the end.

The algorithm

for each token:

operand : push its value

operator: pop the right operand, pop the left operand,

apply, push the result

at the end: exactly one value should remain; it is the answer

Shorter than the conversion, with no precedence and no brackets, because the postfix form has already resolved all of that. That is what the conversion bought.

The order trap

The two pops come off in the reverse of the order the operands appeared.

5 3 -

The first pop gives 3, the second gives 5, and the answer is 5 - 3 = 2, not 3 - 5 = -2.

For + and x it makes no difference and the bug hides. For -, /, % and ^ it does, so the rule is: the first value popped is the right operand.

Built, with the trace printed

def evaluate_postfix(tokens, trace=False):
    """Evaluate a postfix expression. Returns the value."""
    stack = []
    if trace:
        print("%-6s %-28s %s" % ("token", "action", "stack"))
    for token in tokens:
        if token in ("+", "-", "x", "/", "%", "^"):
            if len(stack) < 2:
                raise ValueError("operator '%s' needs two operands" % token)
            right = stack.pop()                 # FIRST pop is the RIGHT operand
            left = stack.pop()
            if token == "+":
                value = left + right
            elif token == "-":
                value = left - right
            elif token == "x":
                value = left * right
            elif token == "/":
                if right == 0:
                    raise ZeroDivisionError("division by zero in the expression")
                value = left / right
            elif token == "%":
                value = left % right
            else:
                value = left ** right
            stack.append(value)
            action = "%s %s %s = %s" % (left, token, right, value)
        else:
            value = int(token)
            stack.append(value)
            action = "push %s" % value
        if trace:
            print("%-6s %-28s %s" % (token, action, stack))
    if len(stack) != 1:
        raise ValueError("the expression left %d values, not 1" % len(stack))
    return stack[0]


print("5 3 - 2 x")
print("answer:", evaluate_postfix("5 3 - 2 x".split(), trace=True))

print()
cases = [
    ("2 3 +",            "2 + 3"),
    ("5 3 -",            "5 - 3"),
    ("2 3 4 x +",        "2 + 3 x 4"),
    ("2 3 + 4 x",        "(2 + 3) x 4"),
    ("2 3 2 ^ ^",        "2 ^ (3 ^ 2)"),
    ("2 3 ^ 2 ^",        "(2 ^ 3) ^ 2"),
    ("10 2 / 3 -",       "10 / 2 - 3"),
]
for postfix, meaning in cases:
    print("%-14s = %-16s -> %s" % (postfix, meaning, evaluate_postfix(postfix.split())))

print()
for bad, why in [("2 +", "not enough operands"), ("2 3", "two values left"),
                 ("4 0 /", "division by zero")]:
    try:
        evaluate_postfix(bad.split())
    except (ValueError, ZeroDivisionError) as e:
        print("%-8s refused (%s): %s" % (bad, why, e))
munotes.in114

Evaluating a Postfix Expression

5 3 - 2 x
token  action                       stack
5      push 5                       [5]
3      push 3                       [5, 3]
-      5 - 3 = 2                    [2]
2      push 2                       [2, 2]
x      2 x 2 = 4                    [4]
answer: 4

2 3 +          = 2 + 3            -> 5
5 3 -          = 5 - 3            -> 2
2 3 4 x +      = 2 + 3 x 4        -> 14
2 3 + 4 x      = (2 + 3) x 4      -> 20
2 3 2 ^ ^      = 2 ^ (3 ^ 2)      -> 512
2 3 ^ 2 ^      = (2 ^ 3) ^ 2      -> 64
10 2 / 3 -     = 10 / 2 - 3       -> 2.0

2 +      refused (not enough operands): operator '+' needs two operands
2 3      refused (two values left): the expression left 2 values, not 1
4 0 /    refused (division by zero): division by zero in the expression

Two rows settle chapter 36's warning with arithmetic: 2 3 2 ^ ^ is 512 and 2 3 ^ 2 ^ is 64. Those are the postfix forms of 2 ^ (3 ^ 2) and (2 ^ 3) ^ 2, and the difference is entirely in the order of the symbols.

Rows three and four are the same check for precedence: 2 3 4 x + is 14 and 2 3 + 4 x is 20.

Conversion and evaluation together

The two chapters compose: convert an infix expression, then evaluate the result, and compare against the value the expression should have.

PRECEDENCE = {"+": 1, "-": 1, "x": 2, "/": 2, "^": 3}
RIGHT = {"^"}


def to_postfix(tokens):
    output, stack = [], []
    for token in tokens:
        if token not in PRECEDENCE and token not in "()":
            output.append(token)
        elif token == "(":
            stack.append(token)
        elif token == ")":
            while stack and stack[-1] != "(":
                output.append(stack.pop())
            stack.pop()
        else:
            while (stack and stack[-1] != "("
                   and (PRECEDENCE[stack[-1]] > PRECEDENCE[token]
                        or (PRECEDENCE[stack[-1]] == PRECEDENCE[token]
                            and token not in RIGHT))):
                output.append(stack.pop())
            stack.append(token)
    while stack:
        output.append(stack.pop())
    return output


def evaluate(tokens):
    stack = []
    for token in tokens:
        if token in PRECEDENCE:
            right, left = stack.pop(), stack.pop()
            stack.append({"+": left + right, "-": left - right,
                          "x": left * right, "/": left / right,
                          "^": left ** right}[token])
        else:
            stack.append(int(token))
    return stack[0]


checks = [
    ("2 + 3 x 4",            2 + 3 * 4),
    ("( 2 + 3 ) x 4",        (2 + 3) * 4),
    ("10 / 2 - 3",           10 / 2 - 3),
    ("2 ^ 3 ^ 2",            2 ** 3 ** 2),
    ("7 - 2 - 1",            7 - 2 - 1),
    ("2 x ( 3 + 4 ) / 7",    2 * (3 + 4) / 7),
]
print("%-22s %-20s %-8s %-8s %s" % ("infix", "postfix", "ours", "Python", "agree"))
for infix, expected in checks:
    postfix = to_postfix(infix.split())
    ours = evaluate(postfix)
    print("%-22s %-20s %-8s %-8s %s"
          % (infix, " ".join(postfix), ours, expected, ours == expected))
munotes.in115

Evaluating a Postfix Expression

infix                  postfix              ours     Python   agree
2 + 3 x 4              2 3 4 x +            14       14       True
( 2 + 3 ) x 4          2 3 + 4 x            20       20       True
10 / 2 - 3             10 2 / 3 -           2.0      2.0      True
2 ^ 3 ^ 2              2 3 2 ^ ^            512      512      True
7 - 2 - 1              7 2 - 1 -            4        4        True
2 x ( 3 + 4 ) / 7      2 3 4 + x 7 /        2.0      2.0      True

Every row is checked against Python's own evaluation of the same expression, which knows nothing about our stack. Six independent agreements, including both associativity cases.

The cost

One pass, one push or one pop pair per token: O(n) time. The stack holds at most the operands not yet consumed, so O(n) memory, worst case for an expression like 1 2 3 4 5 + + + +.

Quick revision

  • Push operands; on an operator pop two, apply, push the result; one value should remain at the end.
  • The first value popped is the RIGHT operand. This matters for -, /, % and ^ and hides for +

and x.

  • No precedence and no brackets are needed: the conversion already resolved them.
  • 2 3 2 ^ ^ is 512 and 2 3 ^ 2 ^ is 64.
  • Errors: too few operands for an operator; more than one value left at the end; division by zero.
  • O(n) time and O(n) memory.

Test yourself

1. Give the evaluation algorithm in two lines. Push each operand. On each operator, pop two values, apply the operator and push the result. At the end exactly one value remains and it is the answer.

2. Which popped value is the right operand, and which operators does it matter for? The first one popped. It matters for -, /, % and ^; for + and x the mistake is invisible.

3. Evaluate 5 3 - 2 x step by step. Push 5, push 3; the - pops 3 then 5 and pushes 5 - 3 = 2; push 2; the x pops 2 and 2 and pushes 4. The answer is 4.

4. Evaluate 2 3 2 ^ ^ and 2 3 ^ 2 ^, and say which infix expressions they are. 512 and 64. They are 2 ^ (3 ^ 2) and (2 ^ 3) ^ 2.

munotes.in116

Evaluating a Postfix Expression

5. Name three errors the evaluator must detect. An operator with fewer than two operands available; more than one value remaining at the end; and division by zero.

6. Why does evaluation need no precedence rules? Because the postfix form already encodes the grouping in the order of its symbols, which is exactly what the conversion did.

Contents This chapter on its own page

munotes.in117

Chapter Thirty-Nine

Prefix: Converting to It, and Evaluating It

Syllabus topic Module 1, "Stacks: Applications of stack like prefix to postfix notation"

In one line

Prefix is postfix backwards: convert to it by reversing the infix, swapping the brackets, running the postfix algorithm and reversing the result, and evaluate it by scanning right to left.

Infix to prefix, without a new algorithm

Writing a second shunting yard for prefix is unnecessary. The standard method reuses chapter 37's:

1. reverse the infix expression

2. swap every '(' with ')' and every ')' with '('

3. run the infix to postfix algorithm with EVERY operator's associativity flipped

4. reverse the result

Step 3 is the part that is quietly dropped in wrong answers, and it is more than the exponent. Reversing the expression reverses the direction associativity works in, so every left associative operator must be treated as right associative for the duration of the conversion, and the exponent as left associative. The first draft of this chapter flipped nothing, and the independent tree method disagreed on A - B - C immediately.

Built, with both methods agreeing

PRECEDENCE = {"+": 1, "-": 1, "x": 2, "/": 2, "^": 3}
RIGHT = {"^"}


def to_postfix(tokens, right_associative=RIGHT):
    output, stack = [], []
    for token in tokens:
        if token not in PRECEDENCE and token not in "()":
            output.append(token)
        elif token == "(":
            stack.append(token)
        elif token == ")":
            while stack and stack[-1] != "(":
                output.append(stack.pop())
            stack.pop()
        else:
            while (stack and stack[-1] != "("
                   and (PRECEDENCE[stack[-1]] > PRECEDENCE[token]
                        or (PRECEDENCE[stack[-1]] == PRECEDENCE[token]
                            and token not in right_associative))):
                output.append(stack.pop())
            stack.append(token)
    while stack:
        output.append(stack.pop())
    return output


def to_prefix(tokens):
    """Reverse, swap brackets, convert with EVERY associativity flipped, reverse.

    Reversing the expression reverses the direction associativity works in, so
    the flag passed to to_postfix must be the set of operators that are LEFT
    associative in the original: they are the ones that must now NOT pop on
    equal precedence. The first draft passed set(), flipping nothing, and the
    tree method disagreed on 'A - B - C' at once."""
    swapped = []
    for token in reversed(tokens):
        swapped.append(")" if token == "(" else "(" if token == ")" else token)
    flipped = set(PRECEDENCE) - RIGHT          # every left associative operator
    converted = to_postfix(swapped, right_associative=flipped)
    return list(reversed(converted))


def tree_prefix(node):
    """An independent method: walk a fully bracketed tree, operator first."""
    if not isinstance(node, tuple):
        return [node]
    left, operator, right = node
    return [operator] + tree_prefix(left) + tree_prefix(right)


cases = [
    ("A + B",               ("A", "+", "B")),
    ("A + B x C",           ("A", "+", ("B", "x", "C"))),
    ("( A + B ) x C",       (("A", "+", "B"), "x", "C")),
    ("A + B x C - D",       (("A", "+", ("B", "x", "C")), "-", "D")),
    ("( A + B ) x ( C - D )", (("A", "+", "B"), "x", ("C", "-", "D"))),
    ("A ^ B ^ C",           ("A", "^", ("B", "^", "C"))),
    ("A - B - C",           (("A", "-", "B"), "-", "C")),
]

print("%-24s %-20s %-20s %s" % ("infix", "prefix, by stack", "prefix, by tree", "agree"))
for infix, tree in cases:
    by_stack = " ".join(to_prefix(infix.split()))
    by_tree = " ".join(tree_prefix(tree))
    print("%-24s %-20s %-20s %s" % (infix, by_stack, by_tree, by_stack == by_tree))
munotes.in118

Prefix: Converting to It, and Evaluating It

infix                    prefix, by stack     prefix, by tree      agree
A + B                    + A B                + A B                True
A + B x C                + A x B C            + A x B C            True
( A + B ) x C            x + A B C            x + A B C            True
A + B x C - D            - + A x B C D        - + A x B C D        True
( A + B ) x ( C - D )    x + A B - C D        x + A B - C D        True
A ^ B ^ C                ^ A ^ B C            ^ A ^ B C            True
A - B - C                - - A B C            - - A B C            True

Every row was produced twice, by the reversal trick and by walking a bracketed tree, and the two agree. The last two rows are the associativity cases, and they are exactly where this chapter's first draft failed: with nothing flipped, A - B - C converted to - A - B C, which means A - (B - C) and is wrong.

Evaluating prefix

Postfix was evaluated left to right. Prefix is evaluated right to left, and the two pops come off in the opposite order.

for each token, scanning RIGHT to LEFT:

operand : push it

operator: pop the LEFT operand, pop the RIGHT operand, apply, push

The reason is symmetry: in prefix the operator comes before its operands, so scanning backwards means the operands are already on the stack when the operator is reached, and the nearer one is the left operand.

OPERATORS = {"+", "-", "x", "/", "^"}


def apply_operator(operator, left, right):
    if operator == "+":
        return left + right
    if operator == "-":
        return left - right
    if operator == "x":
        return left * right
    if operator == "/":
        return left / right
    return left ** right


def evaluate_prefix(tokens):
    """Scan right to left. The FIRST pop is the LEFT operand."""
    stack = []
    for token in reversed(tokens):
        if token in OPERATORS:
            left = stack.pop()             # note: the reverse of postfix
            right = stack.pop()
            stack.append(apply_operator(token, left, right))
        else:
            stack.append(int(token))
    if len(stack) != 1:
        raise ValueError("the expression left %d values, not 1" % len(stack))
    return stack[0]


def evaluate_postfix(tokens):
    stack = []
    for token in tokens:
        if token in OPERATORS:
            right = stack.pop()
            left = stack.pop()
            stack.append(apply_operator(token, left, right))
        else:
            stack.append(int(token))
    return stack[0]


checks = [
    ("+ 2 3",            "2 3 +",            2 + 3),
    ("- 5 3",            "5 3 -",            5 - 3),
    ("+ 2 x 3 4",        "2 3 4 x +",        2 + 3 * 4),
    ("x + 2 3 4",        "2 3 + 4 x",        (2 + 3) * 4),
    ("^ 2 ^ 3 2",        "2 3 2 ^ ^",        2 ** (3 ** 2)),
    ("^ ^ 2 3 2",        "2 3 ^ 2 ^",        (2 ** 3) ** 2),
    ("- - 7 2 1",        "7 2 - 1 -",        7 - 2 - 1),
]
print("%-14s %-14s %-8s %-8s %-8s %s" % ("prefix", "postfix", "prefix", "postfix",
                                         "Python", "all agree"))
for prefix, postfix, expected in checks:
    a = evaluate_prefix(prefix.split())
    b = evaluate_postfix(postfix.split())
    print("%-14s %-14s %-8s %-8s %-8s %s"
          % (prefix, postfix, a, b, expected, a == b == expected))
munotes.in119

Prefix: Converting to It, and Evaluating It

prefix         postfix        prefix   postfix  Python   all agree
+ 2 3          2 3 +          5        5        5        True
- 5 3          5 3 -          2        2        2        True
+ 2 x 3 4      2 3 4 x +      14       14       14       True
x + 2 3 4      2 3 + 4 x      20       20       20       True
^ 2 ^ 3 2      2 3 2 ^ ^      512      512      512      True
^ ^ 2 3 2      2 3 ^ 2 ^      64       64       64       True
- - 7 2 1      7 2 - 1 -      4        4        4        True

Three independent evaluations agree on every row: the prefix evaluator, the postfix evaluator, and Python's own arithmetic. The two exponent rows differ from each other by 448, which is the associativity difference surviving both notations correctly.

The four conversions, summarised

FromToMethod
infixpostfixshunting yard, chapter 37
infixprefixreverse, swap brackets, shunting yard with ^ left associative, reverse
postfixinfixevaluate with strings: pop two, build (left op right), push
prefixpostfixevaluate right to left with strings: pop two, build left right op, push

The bottom two rows use the evaluators of this chapter and chapter 38 with the arithmetic replaced by string building, which is a neat thing to notice and occasionally asked for.

Quick revision

  • Prefix puts the operator before its operands; it needs no brackets and no precedence, like postfix.
  • Infix to prefix: reverse the tokens, swap the brackets, run the infix to postfix algorithm with every

operator's associativity flipped, then reverse the output.

  • Forgetting the associativity flip in step 3 is the standard error, and it shows on A - B - C, which

comes out as - A - B C instead of - - A B C.

  • Prefix is evaluated scanning right to left, and the first value popped is the LEFT operand, the
munotes.in120

Prefix: Converting to It, and Evaluating It

reverse of postfix.

  • Every conversion here was produced twice, by the stack and by a bracketed tree, and every evaluation

three times, including by Python itself.

  • Converting back to infix is the same evaluator with the arithmetic replaced by string building.

Test yourself

1. Give the four steps of infix to prefix conversion. Reverse the expression; swap every opening bracket with a closing one; run the infix to postfix algorithm with every operator's associativity flipped; reverse the result.

2. Why must associativity be flipped in step 3? Because reversing the expression reverses the direction in which associativity groups. Left associative operators must be handled as right associative, and the exponent as left associative, for the duration of the conversion.

3. Convert A + B x C to prefix. + A x B C.

4. How is prefix evaluated, and how do the pops differ from postfix? Scanning right to left. The first value popped is the left operand, which is the reverse of postfix, where the first popped is the right operand.

5. Evaluate ^ 2 ^ 3 2 and ^ ^ 2 3 2. 512 and 64. They are 2 ^ (3 ^ 2) and (2 ^ 3) ^ 2.

6. How would you convert postfix back to infix? Run the postfix evaluator with the arithmetic replaced by string building: pop two operands and push the string (left operator right).

Contents This chapter on its own page

munotes.in121

Chapter Forty

The Queue: Two Ends

Syllabus topic Module 1, "Queues: Queue ADT"

In one line

A queue adds at one end and removes from the other, so the first thing in is the first thing out, which is what fairness means when things are served in the order they arrived.

First in, first out

Stand in a queue at a bank. You join at the back. You are served from the front. The person who has been waiting longest goes next.

That is a queue, and the rule is FIFO, first in, first out.

The queueThe operation
join at the backenqueue
serve the person at the frontdequeue
see who is next without serving themfront, or peek
is anyone waitingis_empty

The vocabulary differs between textbooks: rear, back and tail all mean the end you add to, and front and head mean the end you take from. MU prints "Queue ADT" without fixing the words, so an answer may use any consistent pair, and this book uses front and rear.

Stack against queue

This is the comparison to have ready.

StackQueue
Rulelast in, first outfirst in, first out
Add atthe topthe rear
Remove fromthe topthe front
Ends usedonetwo
Natural fornestingwaiting
Order outreversedpreserved
Everyday examplea pile of platesa queue at a counter

The deepest difference is the last two rows. A stack reverses; a queue preserves. A stack is right when the most recent thing must be dealt with first; a queue is right when the oldest must.

Seen side by side:

from collections import deque

items = ["first", "second", "third", "fourth"]

stack = []
for item in items:
    stack.append(item)
stack_out = []
while stack:
    stack_out.append(stack.pop())

queue = deque()
for item in items:
    queue.append(item)
queue_out = []
while queue:
    queue_out.append(queue.popleft())

print("put in       :", items)
print("stack gives  :", stack_out)
print("queue gives  :", queue_out)
print()
print("the stack reversed the order  :", stack_out == list(reversed(items)))
print("the queue preserved the order :", queue_out == items)
put in       : ['first', 'second', 'third', 'fourth']
stack gives  : ['fourth', 'third', 'second', 'first']
queue gives  : ['first', 'second', 'third', 'fourth']

the stack reversed the order  : True
the queue preserved the order : True

Same four items, same order in, opposite orders out. Nothing else about the two structures matters as much as that.

Why two ends is harder than one

The stack was easy because everything happened at one place. A queue works at both ends, and that causes the one real difficulty of the next three chapters.

On a linked list, adding at the head and removing at the head are both cheap, but a queue needs one operation at each end, and one of those ends is expensive unless a pointer is kept there. Chapter 43 solves it with head and tail.

munotes.in122

The Queue: Two Ends

On an array, removing from the front means shifting everything down, which is O(n), or leaving a gap and moving the front marker, which wastes the space in front and eventually reports the queue as full when it is nearly empty. Chapter 42 demonstrates that failure and chapter 44 fixes it.

Neither problem arose for the stack. The second end is what costs.

Where queues are already working

Job scheduling. MU names this application by name, and chapter 48 builds it: processes waiting for the processor, served in arrival order.

Printing. Documents sent to a printer are served in the order they were sent.

Breadth first search. Chapter 94 explores a graph level by level with a queue, exactly as chapter 62 walks a tree with one. The stack goes deep; the queue goes wide.

Buffers. Data arriving faster than it can be handled waits in a queue, and the circular queue of chapter 44 is the standard implementation.

Quick revision

  • A queue adds at the rear and removes from the front: first in, first out.
  • Operations: enqueue, dequeue, front, is_empty.
  • Rear, back and tail mean the adding end; front and head mean the removing end.
  • A stack reverses the order; a queue preserves it. That is the deepest difference.
  • A stack works at one end, a queue at two, and the second end is what makes the implementations harder.
  • On an array, removing from the front shifts everything or wastes the space in front; on links, one end

needs a pointer kept to it.

  • Used for job scheduling, printing, breadth first search and buffers.

Test yourself

1. What does FIFO mean and how does it differ from LIFO? First in, first out: the item that has waited longest is removed next. LIFO removes the most recently added item.

2. Name the queue operations and the two ends. Enqueue at the rear, dequeue from the front, front or peek to look without removing, and is_empty. Rear, back and tail name the adding end; front and head the removing end.

3. Four items go into a stack and a queue in the same order. How do the outputs differ? The stack gives them back reversed; the queue gives them back in the original order.

4. Why are queue implementations harder than stack implementations? Because a queue works at two ends. On a linked list one end needs a pointer kept to it; on an array, removing from the front either shifts everything or leaves unusable space at the front.

5. Give the two problems the array queue has, and which chapters solve them. Shifting everything down on every dequeue, which is O(n); or moving a front marker, which wastes the space in front and reports full when nearly empty. Chapter 42 demonstrates it and chapter 44 fixes it with wrap-around.

munotes.in123

The Queue: Two Ends

6. Name three places a queue is already at work. Job scheduling for the processor, documents waiting at a printer, and breadth first search of a graph or tree. Buffers for data arriving faster than it is handled are a fourth.

Contents This chapter on its own page

munotes.in124

Chapter Forty-One

The Queue ADT

Syllabus topic Module 1, "Queues: Queue ADT"

In one line

The queue ADT is enqueue, dequeue, front, is_empty and size, with underflow on an empty queue and overflow on a full one, and every operation should be O(1).

The ADT

A Queue holds a sequence of items, added at one end called the rear and removed from the other end called the front.

OperationNeedsReturnsDoesWhen it cannot
Queue()nothingan empty queuecreates itnever fails
enqueue(item)an itemnothingadds at the rearoverflow if full
dequeue()nothingan itemremoves and returns the frontunderflow if empty
front()nothingan itemreturns the front, leaves iterror if empty
is_empty()nothingtrue or falseis there nothing in itnever fails
size()nothinga numberhow many itemsnever fails

As with the stack, that table names no array and no linked list, and adding all operations are O(1) is worth a mark.

The names an examiner may use

Textbooks differ, and MU does not fix the words, so know the pairs:

This bookAlso written
enqueueinsert, add, push
dequeuedelete, remove, pop, serve
fronthead, peek, first
rearback, tail, last

push and pop for a queue are used by some textbooks and are a genuine source of confusion with stacks. This book does not use them for queues.

Underflow and overflow, again

Exactly as for the stack, and for the same reasons:

Underflow is dequeueing or reading the front of an empty queue. It applies to every implementation, because emptiness belongs to the queue and not to its storage.

Overflow is enqueueing onto a full queue. It applies only where the capacity is fixed, which is the array implementations of chapters 42 and 44. The linked queue of chapter 43 has no maximum.

The ADT obeyed, on Python's own queue

Before building anything, here is the behaviour the next three chapters must reproduce, using a structure the standard library already provides.

from collections import deque


class Queue:
    """The ADT, on top of a deque, so the behaviour can be seen before it is built."""

    def __init__(self):
        self._items = deque()

    def enqueue(self, item):
        self._items.append(item)

    def dequeue(self):
        if not self._items:
            raise IndexError("dequeue from an empty queue: underflow")
        return self._items.popleft()

    def front(self):
        if not self._items:
            raise IndexError("front of an empty queue: underflow")
        return self._items[0]

    def is_empty(self):
        return len(self._items) == 0

    def size(self):
        return len(self._items)

    def snapshot(self):
        return list(self._items)


q = Queue()
print("new queue: empty =", q.is_empty(), "| size =", q.size())

for person in ("Asha", "Rohit", "Meera"):
    q.enqueue(person)
    print("enqueue %-6s -> %-28s front is %s"
          % (person, str(q.snapshot()), q.front()))

print()
print("dequeue ->", q.dequeue(), "| now", q.snapshot())
print("dequeue ->", q.dequeue(), "| now", q.snapshot())
print("front   ->", q.front(), "| size is still", q.size())

print()
empty = Queue()
for operation in ("dequeue", "front"):
    try:
        getattr(empty, operation)()
    except IndexError as e:
        print("%-8s on an empty queue ->" % operation, e)
munotes.in125

The Queue ADT

new queue: empty = True | size = 0
enqueue Asha   -> ['Asha']                     front is Asha
enqueue Rohit  -> ['Asha', 'Rohit']            front is Asha
enqueue Meera  -> ['Asha', 'Rohit', 'Meera']   front is Asha

dequeue -> Asha | now ['Rohit', 'Meera']
dequeue -> Rohit | now ['Meera']
front   -> Meera | size is still 1

dequeue  on an empty queue -> dequeue from an empty queue: underflow
front    on an empty queue -> front of an empty queue: underflow

Notice the third enqueue line: the front is still Asha, not Meera. Adding at the rear never changes the front, which sounds obvious and is exactly what the array implementation of chapter 42 gets wrong in a subtle way.

The ADT in examination form

Queue: a collection with insertion at the rear and removal at the front.

enqueue(item): add item at the rear. Overflow if full.

dequeue() : remove and return the front item. Underflow if empty.

front() : return the front item without removing it. Error if empty.

is_empty() : true when the queue holds nothing.

size() : the number of items held.

All operations should be O(1).

The word should in the last line is deliberate, and it is what chapters 42 to 44 are about: a naive array implementation does not achieve it, and getting to O(1) is the work.

Quick revision

  • The queue ADT: enqueue at the rear, dequeue from the front, front, is_empty, size.
  • It names no implementation, and every operation should be O(1).
  • Underflow applies to every implementation; overflow only where capacity is fixed.
  • Know the alternative names: insert and add for enqueue, delete and serve for dequeue, head for front,

back and tail for rear.

  • Adding at the rear never changes the front.
  • The naive array implementation does not achieve O(1), which is the point of the next three chapters.

Test yourself

1. Write the queue ADT. A collection with insertion at the rear and removal at the front. enqueue(item) adds at the rear, overflow if full; dequeue() removes and returns the front, underflow if empty; front() returns the front without removing it; is_empty(); size(). All operations should be O(1).

2. Which implementations can overflow, and which can underflow? Any fixed capacity implementation can overflow, which means the array versions. Every implementation can underflow.

3. Give three alternative names for dequeue. Delete, remove, serve. Some textbooks also write pop, which is best avoided because of the stack.

4. Does enqueueing change the front of the queue? No, unless the queue was empty. Adding happens at the rear.

5. Why does the ADT say operations "should" be O(1) rather than "are"? Because it is a target the implementation must achieve. The naive array queue does not: its dequeue shifts every remaining item and is O(n).

munotes.in126

The Queue ADT

6. What is the difference between front() and dequeue()? front() returns the front item and leaves the queue unchanged; dequeue() removes it and returns it, reducing the size by one.

Contents This chapter on its own page

munotes.in127

Chapter Forty-Two

A Queue on an Array, and the Drift That Ruins It

Syllabus topic Module 1, "Queues: linked representations"

In one line

An array queue with a moving front marker drifts up the array and eventually reports itself full while almost empty, and the obvious fix, shifting everything down, makes every dequeue O(n).

The first attempt: two markers

Keep the items in a block with two integers: front, the index of the first item, and rear, the index after the last.

enqueue: cells[rear] = item; rear = rear + 1

dequeue: item = cells[front]; front = front + 1

empty : front == rear

full : rear == capacity

Both operations are O(1), which looks like success. It is not, and the reason is the full condition.

The drift, run

class DriftingQueue:
    """front and rear both move up and never come back."""

    def __init__(self, capacity):
        self.capacity = capacity
        self.cells = [None] * capacity
        self.front = 0
        self.rear = 0

    def is_empty(self):
        return self.front == self.rear

    def size(self):
        return self.rear - self.front

    def is_full(self):
        return self.rear == self.capacity      # the defect lives here

    def enqueue(self, item):
        if self.is_full():
            raise OverflowError("the queue reports itself full")
        self.cells[self.rear] = item
        self.rear += 1

    def dequeue(self):
        if self.is_empty():
            raise IndexError("dequeue from an empty queue")
        item = self.cells[self.front]
        self.cells[self.front] = None
        self.front += 1
        return item

    def picture(self):
        """The whole block, with the two markers shown."""
        cells = ["%-4s" % ("." if c is None else c) for c in self.cells]
        return "[%s]  front=%d rear=%d size=%d" % (
            " ".join(cells), self.front, self.rear, self.size())


q = DriftingQueue(5)
print("new       ", q.picture())

for person in ("A", "B", "C"):
    q.enqueue(person)
print("3 enqueued", q.picture())

q.dequeue()
q.dequeue()
print("2 dequeued", q.picture())

q.enqueue("D")
q.enqueue("E")
print("2 enqueued", q.picture())

print()
print("size is %d, capacity is %d, so %d cells are unused."
      % (q.size(), q.capacity, q.capacity - q.size()))
print("is_full says:", q.is_full())
try:
    q.enqueue("F")
except OverflowError as e:
    print("enqueue F ->", e)
print()
print("and the unused cells are all at the FRONT, where nothing can reach them.")
new        [.    .    .    .    .   ]  front=0 rear=0 size=0
3 enqueued [A    B    C    .    .   ]  front=0 rear=3 size=3
2 dequeued [.    .    C    .    .   ]  front=2 rear=3 size=1
2 enqueued [.    .    C    D    E   ]  front=2 rear=5 size=3

size is 3, capacity is 5, so 2 cells are unused.
is_full says: True
enqueue F -> the queue reports itself full

and the unused cells are all at the FRONT, where nothing can reach them.

A queue of capacity 5, holding 3 items, refusing a fourth. Two cells are free and both are at the front of the array, behind the front marker, where rear can never reach them.

Both markers only ever move up. Every dequeue abandons a cell for ever. Run this long enough and the queue reports itself full having done nothing but work correctly.

munotes.in128

A Queue on an Array, and the Drift That Ruins It

That is the drift, and it is not an edge case: it happens to every array queue built this way, always.

The obvious fix, and what it costs

The natural response is to keep the front at index 0: on every dequeue, shift everything down one place.

class ShiftingQueue:
    """Keep the front at 0 by shifting on every dequeue. No drift, but a cost."""

    def __init__(self, capacity):
        self.capacity = capacity
        self.cells = [None] * capacity
        self.rear = 0
        self.moves = 0

    def is_empty(self):
        return self.rear == 0

    def size(self):
        return self.rear

    def enqueue(self, item):
        if self.rear == self.capacity:
            raise OverflowError("genuinely full")
        self.cells[self.rear] = item
        self.rear += 1

    def dequeue(self):
        if self.is_empty():
            raise IndexError("dequeue from an empty queue")
        item = self.cells[0]
        for i in range(1, self.rear):          # everything moves down one
            self.cells[i - 1] = self.cells[i]
            self.moves += 1
        self.rear -= 1
        self.cells[self.rear] = None
        return item


for n in (500, 1000, 2000, 4000):
    q = ShiftingQueue(n)
    for i in range(n):
        q.enqueue(i)
    for _ in range(n):
        q.dequeue()
    print("n = %4d | %d items through the queue cost %8d element moves"
          % (n, n, q.moves))
n =  500 | 500 items through the queue cost   124750 element moves
n = 1000 | 1000 items through the queue cost   499500 element moves
n = 2000 | 2000 items through the queue cost  1999000 element moves
n = 4000 | 4000 items through the queue cost  7998000 element moves

The drift is gone and the space is used properly. But look at the column: doubling the traffic quadruples the work. Every dequeue is now O(n), so passing n items through the queue is O(n squared).

That is not an acceptable queue. The ADT of chapter 41 said every operation should be O(1), and this one has traded a space defect for a time defect.

The two failures, side by side

Drifting queueShifting queue
enqueueO(1)O(1)
dequeueO(1)O(n)
Space usedabandons a cell per dequeueall of it
Reports full whenrear reaches the end, whatever the sizegenuinely full
Usable?nono

Neither is acceptable, and each fails where the other succeeds. What is needed is the drifting queue's O(1) dequeue with the shifting queue's use of space, and the answer is to let rear wrap around to the beginning of the array when it reaches the end.

That is the circular queue, and it is chapter 44.

Quick revision

  • The simple array queue keeps front and rear as indices that only move up.
  • Both operations are O(1), but every dequeue abandons the cell it leaves behind.
  • The queue therefore reports itself full when rear reaches the end, however few items it holds: run

here at capacity 5, size 3, refusing a fourth item with two cells free.

munotes.in129

A Queue on an Array, and the Drift That Ruins It

  • The free cells are all in front of front, where rear can never reach them.
  • Shifting everything down on dequeue fixes the space and makes dequeue O(n), so passing n items through

costs O(n squared): 4,000 items cost 7,998,000 element moves.

  • The fix that keeps both is to let the indices wrap around, which is the circular queue.

Test yourself

1. Why does the simple array queue report itself full while nearly empty? Because full is tested as rear reaching the capacity, and both markers only move up. Every dequeue abandons a cell at the front that nothing can reuse.

2. In the run, what were the capacity, the size and the number of free cells when it refused an item? Capacity 5, size 3, two free cells, both in front of the front marker.

3. What does shifting on every dequeue fix, and what does it break? It fixes the wasted space, because the front stays at index 0. It breaks the cost: dequeue becomes O(n) and passing n items through the queue becomes O(n squared).

4. At n = 4,000, how many element moves did the shifting queue perform? 7,998,000, which is about n squared over 2.

5. State what a correct array queue must achieve that neither of these does. O(1) enqueue and O(1) dequeue, while using every cell of the array.

6. What single change achieves it? Letting the front and rear indices wrap around to the beginning of the array when they reach the end, which is the circular queue.

Contents This chapter on its own page

munotes.in130

Chapter Forty-Four

The Circular Queue: Wrap-Around

Syllabus topic Module 1, "Queues: Circular Queue operations"

In one line

A circular queue treats the array as a ring by advancing every index with modulo arithmetic, so the cells freed at the front are reused and both operations stay O(1).

The one idea

Chapter 42's queue failed because rear ran off the end of the array while cells at the front sat empty. The fix is to let it carry on from the beginning.

rear = (rear + 1) mod capacity

front = (front + 1) mod capacity

That is the whole change. The modulo turns the array into a ring: index 4 of a 5 cell array is followed by index 0.

Drawn, the array is a circle with the two markers on it:

cell 0 - cell 1 - cell 2 - cell 3 - cell 4 - back to cell 0

The queue is the stretch from front round to rear, which may wrap past the end.

The operations

enqueue: cells[rear] = item; rear = (rear + 1) mod capacity; count = count + 1

dequeue: item = cells[front]; front = (front + 1) mod capacity; count = count - 1

empty : count == 0

full : count == capacity

A count is kept, and chapter 45 explains why that choice is made rather than deducing full and empty from the markers alone.

Built and run, with the wrap visible

class CircularQueue:
    """The array as a ring. count distinguishes full from empty."""

    def __init__(self, capacity):
        self.capacity = capacity
        self.cells = [None] * capacity
        self.front = 0
        self.rear = 0
        self.count = 0

    def is_empty(self):
        return self.count == 0

    def is_full(self):
        return self.count == self.capacity

    def size(self):
        return self.count

    def enqueue(self, item):
        if self.is_full():
            raise OverflowError("enqueue onto a full queue: overflow")
        self.cells[self.rear] = item
        self.rear = (self.rear + 1) % self.capacity
        self.count += 1

    def dequeue(self):
        if self.is_empty():
            raise IndexError("dequeue from an empty queue: underflow")
        item = self.cells[self.front]
        self.cells[self.front] = None
        self.front = (self.front + 1) % self.capacity
        self.count -= 1
        return item

    def front_item(self):
        if self.is_empty():
            raise IndexError("front of an empty queue: underflow")
        return self.cells[self.front]

    def order(self):
        """The items from front to rear, following the ring."""
        out = []
        for step in range(self.count):
            out.append(self.cells[(self.front + step) % self.capacity])
        return out

    def picture(self):
        cells = ["%-4s" % ("." if c is None else c) for c in self.cells]
        return "[%s] f=%d r=%d n=%d" % (" ".join(cells), self.front,
                                        self.rear, self.count)


q = CircularQueue(5)
print("new        ", q.picture())

for person in ("A", "B", "C"):
    q.enqueue(person)
print("A B C in   ", q.picture(), "order:", q.order())

q.dequeue()
q.dequeue()
print("2 out      ", q.picture(), "order:", q.order())

q.enqueue("D")
q.enqueue("E")
print("D E in     ", q.picture(), "order:", q.order())

print()
print("now the wrap: rear is at the end of the array and there are free cells at the front")
q.enqueue("F")
print("F in       ", q.picture(), "order:", q.order())
q.enqueue("G")
print("G in       ", q.picture(), "order:", q.order())

print()
print("the queue is full now:", q.is_full(), "with size", q.size(),
      "of capacity", q.capacity)
try:
    q.enqueue("H")
except OverflowError as e:
    print("enqueue H ->", e)

print()
print("draining it in order:")
while not q.is_empty():
    print("   dequeue ->", q.dequeue(), " ", q.picture())
munotes.in134

The Circular Queue: Wrap-Around

new         [.    .    .    .    .   ] f=0 r=0 n=0
A B C in    [A    B    C    .    .   ] f=0 r=3 n=3 order: ['A', 'B', 'C']
2 out       [.    .    C    .    .   ] f=2 r=3 n=1 order: ['C']
D E in      [.    .    C    D    E   ] f=2 r=0 n=3 order: ['C', 'D', 'E']

now the wrap: rear is at the end of the array and there are free cells at the front
F in        [F    .    C    D    E   ] f=2 r=1 n=4 order: ['C', 'D', 'E', 'F']
G in        [F    G    C    D    E   ] f=2 r=2 n=5 order: ['C', 'D', 'E', 'F', 'G']

the queue is full now: True with size 5 of capacity 5
enqueue H -> enqueue onto a full queue: overflow

draining it in order:
   dequeue -> C   [F    G    .    D    E   ] f=3 r=2 n=4
   dequeue -> D   [F    G    .    .    E   ] f=4 r=2 n=3
   dequeue -> E   [F    G    .    .    .   ] f=0 r=2 n=2
   dequeue -> F   [.    G    .    .    .   ] f=1 r=2 n=1
   dequeue -> G   [.    .    .    .    .   ] f=2 r=2 n=0

Follow the line marked "D E in": rear has become 0. It reached 5, the capacity, and wrapped to the beginning. The next enqueue puts F in cell 0, which chapter 42 had abandoned for ever.

The queue now uses every cell, it holds 5 items in a 5 cell array, and it correctly refuses the sixth. The drift is gone and both operations are still two assignments.

Note the order line while the queue is wrapped: the items are C, D, E, F, G but they sit in the array as F, G, C, D, E. The order lives in the markers, not in the positions, which is why order walks with the modulo rather than reading the array left to right.

The cost

Circular queue
enqueueO(1)
dequeueO(1)
Space usedall of it
Capacityfixed at creation
Overflowgenuine, only when actually full

Compare chapter 42's two attempts: the drifting queue had O(1) operations and wasted space, the shifting queue had full space use and O(n) dequeue. The circular queue has both, and the only thing it costs is a little index arithmetic.

The trap in reading the array

Printing the cells left to right does not print the queue. In the run above the array read F G C D E while the queue was C, D, E, F, G. An examination answer that reads the block in order has misread the structure.

munotes.in135

The Circular Queue: Wrap-Around

Always report a circular queue from front, stepping with the modulo, for count items.

Quick revision

  • The array is treated as a ring: every index advance is (index + 1) mod capacity.
  • This reuses the cells freed at the front, which the drifting queue abandoned.
  • enqueue writes at rear then advances it; dequeue reads at front then advances it. Both O(1).
  • A separate count is kept for size, empty and full.
  • The queue occupies the stretch from front round to rear and may wrap past the end of the array.
  • The order lives in the markers, not the positions: reading the array left to right does not give the

queue's order.

  • It achieves what neither chapter 42 attempt did: O(1) both ways with every cell usable.

Test yourself

1. What single change turns the drifting queue into a circular one? Advancing the indices with modulo the capacity, so that after the last cell they continue at the first.

2. Write the enqueue and dequeue operations. enqueue: check full, cells[rear] = item, rear = (rear + 1) mod capacity, increment the count. dequeue: check empty, item = cells[front], front = (front + 1) mod capacity, decrement the count.

3. In the run, the array held F G C D E and the queue was C, D, E, F, G. Explain. The queue starts at front, which was 2, and wraps round the end of the array. The order is given by stepping from front with the modulo, not by reading the cells left to right.

4. Compare the circular queue with the two attempts of chapter 42. The drifting queue was O(1) but abandoned a cell per dequeue. The shifting queue used all the space but made dequeue O(n). The circular queue is O(1) both ways and uses every cell.

5. When does a circular queue genuinely overflow? Only when it actually holds capacity items, unlike the drifting queue, which reported full whenever rear reached the end.

6. How should a circular queue be printed? From front, stepping with (front + step) mod capacity for count items. Reading the array in index order misreports it whenever the queue is wrapped.

Contents This chapter on its own page

munotes.in136

Chapter Forty-Five

Full or Empty: Telling Them Apart

Syllabus topic Module 1, "Queues: Circular Queue operations"

In one line

In a circular queue with only front and rear, both an empty queue and a full one satisfy front == rear, and the three ways out are to keep a count, to waste one cell, or to keep a flag.

The ambiguity

Take a circular queue holding only front and rear, with rear being the next cell to write.

Empty. Nothing has been added, or everything has been removed. front and rear are at the same place, so front == rear.

Full. Every cell holds an item. rear has advanced all the way round and caught up with front, so front == rear.

The two states produce identical values of both markers. No test on front and rear alone can distinguish them, because the information is genuinely not there.

Run, so it is not merely asserted

class AmbiguousQueue:
    """front and rear only. No count, no flag, one cell per capacity."""

    def __init__(self, capacity):
        self.capacity = capacity
        self.cells = [None] * capacity
        self.front = 0
        self.rear = 0

    def enqueue(self, item):
        self.cells[self.rear] = item
        self.rear = (self.rear + 1) % self.capacity

    def dequeue(self):
        item = self.cells[self.front]
        self.cells[self.front] = None
        self.front = (self.front + 1) % self.capacity
        return item

    def markers(self):
        return "front=%d rear=%d  front==rear is %s" % (
            self.front, self.rear, self.front == self.rear)


empty = AmbiguousQueue(4)
print("a queue with nothing in it :", empty.markers())
print("   its cells               :", empty.cells)

full = AmbiguousQueue(4)
for item in ("A", "B", "C", "D"):
    full.enqueue(item)
print()
print("a queue holding four items :", full.markers())
print("   its cells               :", full.cells)

print()
print("the markers are identical  :",
      (empty.front, empty.rear) == (full.front, full.rear))
print("the queues are not         :", empty.cells != full.cells)
print()
print("so no test on front and rear alone can tell these two apart.")
a queue with nothing in it : front=0 rear=0  front==rear is True
   its cells               : [None, None, None, None]

a queue holding four items : front=0 rear=0  front==rear is True
   its cells               : ['A', 'B', 'C', 'D']

the markers are identical  : True
the queues are not         : True

so no test on front and rear alone can tell these two apart.

One queue holds nothing, the other holds four items, and their markers are indistinguishable.

The three standard answers

1. Keep a count

Hold an integer of how many items are in the queue.

empty: count == 0

full : count == capacity

For: all cells are usable, both tests are obvious, and size is free. Against: one more field, which every operation must update.

This is what chapter 44 uses, and it is what this book recommends, because size is wanted anyway and a count that is maintained in exactly two places is hard to get wrong.

munotes.in137

Full or Empty: Telling Them Apart

2. Sacrifice one cell

Never let the queue hold more than capacity - 1 items. Then a full queue always leaves one gap, so the markers never coincide when full.

empty: front == rear

full : (rear + 1) mod capacity == front

For: no extra field at all. Against: one cell of the array is never used, and the tests are less obvious. A queue declared with capacity 5 holds 4.

This is the version many textbooks print, so it must be recognised, and the full test is the line to memorise: (rear + 1) mod capacity == front.

3. Keep a flag

Hold a boolean saying whether the last operation was an enqueue. If the markers coincide and the last operation was an enqueue, it is full; if it was a dequeue, it is empty.

For: all cells usable, only one bit. Against: the flag has to be correct after every operation, and reasoning about it is harder than either of the others. Rarely used.

All three, run side by side

class CountQueue:
    def __init__(self, capacity):
        self.capacity, self.cells = capacity, [None] * capacity
        self.front = self.rear = self.count = 0

    def is_empty(self):
        return self.count == 0

    def is_full(self):
        return self.count == self.capacity

    def enqueue(self, item):
        if self.is_full():
            return False
        self.cells[self.rear] = item
        self.rear = (self.rear + 1) % self.capacity
        self.count += 1
        return True


class SacrificeQueue:
    def __init__(self, capacity):
        self.capacity, self.cells = capacity, [None] * capacity
        self.front = self.rear = 0

    def is_empty(self):
        return self.front == self.rear

    def is_full(self):
        return (self.rear + 1) % self.capacity == self.front

    def enqueue(self, item):
        if self.is_full():
            return False
        self.cells[self.rear] = item
        self.rear = (self.rear + 1) % self.capacity
        return True


for cls in (CountQueue, SacrificeQueue):
    q = cls(5)
    accepted = 0
    while q.enqueue(accepted):
        accepted += 1
    print("%-16s capacity 5 accepted %d items before reporting full"
          % (cls.__name__, accepted))
CountQueue       capacity 5 accepted 5 items before reporting full
SacrificeQueue   capacity 5 accepted 4 items before reporting full

There is the cost of the second method, measured: a five cell array holding four items.

Which to write in an examination

Either of the first two, stated clearly. The marks are for:

  1. Explaining why the ambiguity exists: both states give front == rear.
  2. Naming the method chosen.
  3. Giving the correct full and empty tests for that method, which is where answers go wrong by mixing

the count method's empty with the sacrifice method's full.

Mixing them is the error to avoid. If you keep a count, use count == 0 and count == capacity. If you sacrifice a cell, use front == rear and (rear + 1) mod capacity == front.

Quick revision

  • In a circular queue with only front and rear, empty and full both give front == rear.
  • Three answers: keep a count; sacrifice one cell; keep a flag.
  • Count: empty is count == 0, full is count == capacity. All cells usable, one field to maintain.
  • Sacrifice: empty is front == rear, full is (rear + 1) mod capacity == front. No extra field,
munotes.in138

Full or Empty: Telling Them Apart

one cell lost, so capacity 5 holds 4.

  • Flag: a boolean recording whether the last operation was an enqueue. Rarely used.
  • In an answer: say why the ambiguity exists, name the method, and give that method's own two tests

without mixing them.

Test yourself

1. Why can front and rear alone not distinguish full from empty? Because in both states the two markers hold the same value: an empty queue has never advanced rear past front, and a full one has advanced it all the way round to meet front again.

2. Give the full and empty tests for the count method. Empty is count == 0; full is count == capacity.

3. Give the full and empty tests for the sacrifice method. Empty is front == rear; full is (rear + 1) mod capacity == front.

4. What does the sacrifice method cost, exactly? One cell of the array. A queue declared with capacity 5 can hold only 4 items, as the run showed.

5. Describe the flag method and say why it is rare. Keep a boolean recording whether the last operation was an enqueue; when the markers coincide it distinguishes full from empty. It is rare because the flag must be maintained correctly by every operation and is harder to reason about than a count.

6. What is the standard error in answering this question? Mixing the methods: using the count method's empty test with the sacrifice method's full test, or the reverse. Each method has its own pair and they must be used together.

Contents This chapter on its own page

munotes.in139

Chapter Forty-Six

What a Queue Is Good and Bad At

Syllabus topic Module 1, "Queues: Queue ADT, Advantages & Disadvantages"

In one line

A queue is good at serving things in the order they arrived and bad at everything that requires knowing what is inside it, including which item is most urgent.

The advantages

1. Every operation is O(1), on the circular array and on the linked list alike. No searching, no shifting.

2. It is fair, in the precise sense. Every item is served after those that arrived before it and before those that arrived after. Nothing can overtake, and nothing can starve: an item's wait is bounded by the number of items ahead of it.

3. It matches anything that waits. Jobs for a processor, documents for a printer, packets for a link, requests for a server, people for a counter.

4. It decouples producer from consumer. One part of a program can add faster than another removes, for a while, and the queue absorbs the difference. That is what a buffer is.

5. It is the structure for breadth first work. Chapter 94 searches a graph with a queue and chapter 61 walks a tree level by level with one. Where a stack goes deep, a queue goes wide.

The disadvantages

1. Only the front is reachable. The second item cannot be read without removing the first.

2. No searching. A queue has no contains, for the same reason a stack has none.

3. Traversal destroys it, unless the items are put back at the rear as they are read, which needs the size to be known first or the walk never ends.

4. The array version needs a capacity in advance, and the circular implementation needs the care of chapters 44 and 45.

5. Fairness is not always what is wanted. This is the important one. A queue cannot let an urgent item go first, because the only thing it knows about an item is when it arrived.

That last disadvantage is worth demonstrating, because it is the bridge into Module 2.

Where fairness is the wrong rule

from collections import deque

# Jobs as (name, urgency), where 1 is most urgent.
arrivals = [("bulk print", 5), ("routine backup", 5), ("routine report", 5),
            ("ALARM", 1), ("another backup", 5)]

queue = deque()
for job in arrivals:
    queue.append(job)

served, waited_before_alarm = [], 0
while queue:
    name, urgency = queue.popleft()
    served.append(name)
    if name == "ALARM":
        break
    waited_before_alarm += 1

print("arrived in this order:")
for name, urgency in arrivals:
    print("   %-16s urgency %d" % (name, urgency))
print()
print("a queue serves them in arrival order, so the ALARM was served")
print("after %d less urgent jobs had gone first." % waited_before_alarm)
print("order served:", served)
print()
print("the queue cannot do better, because the only thing it knows")
print("about a job is WHEN IT ARRIVED. Urgency is invisible to it.")
munotes.in140

What a Queue Is Good and Bad At

arrived in this order:
   bulk print       urgency 5
   routine backup   urgency 5
   routine report   urgency 5
   ALARM            urgency 1
   another backup   urgency 5

a queue serves them in arrival order, so the ALARM was served
after 3 less urgent jobs had gone first.
order served: ['bulk print', 'routine backup', 'routine report', 'ALARM']

the queue cannot do better, because the only thing it knows
about a job is WHEN IT ARRIVED. Urgency is invisible to it.

The queue behaved perfectly correctly and produced the wrong result, because arrival order was the wrong rule for this problem. No implementation detail fixes that: it is the ADT.

What is needed is a structure that serves by priority rather than by arrival, and that is the priority queue of chapter 78, built on the heap of chapter 81. Module 1 ends here on purpose, and Module 2 opens by picking it up.

The table

Queue
enqueue, dequeue, frontO(1)
access item inot possible
searchnot possible without destroying it
find the most urgentnot possible
traversedestroys it, unless items are re-enqueued
memory, circular arrayfixed block, capacity reserved
memory, linkedone address per item

Stack, queue and deque, in one line each

The three linear restricted structures of Module 1, as an examiner may ask them compared:

Add atRemove fromRule
Stackone endthe same endlast in, first out
Queueone endthe other endfirst in, first out
Dequeeither endeither endboth, and neither

Quick revision

  • Advantages: all operations O(1); fair, with bounded waiting and no starvation; matches anything that

waits; decouples producer from consumer; the structure for breadth first work.

  • Disadvantages: only the front is reachable; no search; traversal destroys it; the array version needs

a capacity; and it cannot serve by urgency.

  • The last one is an ADT limitation, not an implementation one: a queue knows only when an item arrived.
  • Demonstrated: an alarm arriving fourth is served fourth, behind three routine jobs.
  • The fix is the priority queue of chapter 78, which is where Module 2 begins.

Test yourself

1. Give three advantages of a queue. Every operation is O(1); it is fair, with each item's wait bounded by the number ahead of it and no starvation; and it decouples a producer from a consumer, which is what a buffer does.

2. Give three disadvantages. Only the front is reachable; there is no search; and traversal destroys the queue unless items are re-enqueued as they are read.

3. Why can a queue not serve an urgent item first? Because the only thing it knows about an item is when it arrived. Urgency is not part of the ADT, so no implementation can recover it.

munotes.in141

What a Queue Is Good and Bad At

4. In the run, how many jobs went before the ALARM, and was the queue faulty? Three. The queue was not faulty; it applied arrival order correctly, and arrival order was the wrong rule for the problem.

5. What structure fixes that, and where is it built? The priority queue, chapter 78, implemented with the heap of chapter 81.

6. Compare the stack, queue and deque in one line each. A stack adds and removes at one end (last in, first out); a queue adds at one end and removes at the other (first in, first out); a deque adds and removes at either end.

Contents This chapter on its own page

munotes.in142

Chapter Forty-Seven

The Deque: Open at Both Ends

Syllabus topic Module 1, "Queues: Dequeues"

In one line

A deque allows adding and removing at both ends, so it contains both the stack and the queue as special cases, and the two restricted structures are the ones worth using when they suffice.

The name

Deque is short for double ended queue, and it is pronounced "deck". MU's printed label is "Dequeues", which looks like the plural of the dequeue operation and is not.

Keep the two apart in an answer:

WordMeans
dequeuethe operation that removes from the front of a queue
dequethe structure that is open at both ends

The ADT

Four operations instead of two, because everything can happen at either end.

OperationDoesWhen it cannot
add_front(item)insert at the frontoverflow if full
add_rear(item)insert at the rearoverflow if full
remove_front()remove and return the frontunderflow if empty
remove_rear()remove and return the rearunderflow if empty
front(), rear()read either end without removingerror if empty
is_empty(), size()as usualnever fail

It contains the other two

This is the point of the structure and the likeliest examination question about it.

Use onlyAnd you have
add_rear and remove_reara stack
add_front and remove_fronta stack, at the other end
add_rear and remove_fronta queue
add_front and remove_reara queue, the other way round
from collections import deque


class Deque:
    """Double ended queue. Built on Python's own deque so this chapter is about
    the ADT rather than about re-implementing chapter 26's links."""

    def __init__(self):
        self._items = deque()

    def add_front(self, item):
        self._items.appendleft(item)

    def add_rear(self, item):
        self._items.append(item)

    def remove_front(self):
        if not self._items:
            raise IndexError("remove_front from an empty deque: underflow")
        return self._items.popleft()

    def remove_rear(self):
        if not self._items:
            raise IndexError("remove_rear from an empty deque: underflow")
        return self._items.pop()

    def size(self):
        return len(self._items)

    def snapshot(self):
        return list(self._items)


items = ["first", "second", "third"]

as_stack = Deque()
for item in items:
    as_stack.add_rear(item)
stack_out = [as_stack.remove_rear() for _ in range(len(items))]

as_queue = Deque()
for item in items:
    as_queue.add_rear(item)
queue_out = [as_queue.remove_front() for _ in range(len(items))]

print("put in                      :", items)
print("used as a stack, gives      :", stack_out)
print("used as a queue, gives      :", queue_out)
print()
print("a deque used at one end IS a stack :", stack_out == list(reversed(items)))
print("a deque used at both ends IS a queue:", queue_out == items)

print()
d = Deque()
d.add_rear("B")
d.add_front("A")
d.add_rear("C")
print("add_rear B, add_front A, add_rear C ->", d.snapshot())
print("remove_front ->", d.remove_front(), "| now", d.snapshot())
print("remove_rear  ->", d.remove_rear(), "| now", d.snapshot())

print()
empty = Deque()
for operation in ("remove_front", "remove_rear"):
    try:
        getattr(empty, operation)()
    except IndexError as e:
        print("%-13s on an empty deque ->" % operation, e)
put in                      : ['first', 'second', 'third']
used as a stack, gives      : ['third', 'second', 'first']
used as a queue, gives      : ['first', 'second', 'third']

a deque used at one end IS a stack : True
a deque used at both ends IS a queue: True

add_rear B, add_front A, add_rear C -> ['A', 'B', 'C']
remove_front -> A | now ['B', 'C']
remove_rear  -> C | now ['B']

remove_front  on an empty deque -> remove_front from an empty deque: underflow
remove_rear   on an empty deque -> remove_rear from an empty deque: underflow
munotes.in143

The Deque: Open at Both Ends

Why not just use a deque for everything?

If a deque does the work of both, why does this paper teach three structures?

Chapter 34 answered it for the stack and the answer is the same here: the restriction is the feature.

A structure that permits only what the problem needs cannot be misused. Code that receives a stack cannot accidentally take from the bottom; code that receives a queue cannot accidentally serve the most recent arrival. Hand those same functions a deque and both mistakes become possible, and both would be silent.

There is a secondary reason: a deque needs a doubly linked list or a circular array with two moving ends, so it is more complex to implement correctly than either of the structures it contains.

The two restricted deques

Textbooks name two variants and examiners occasionally ask for them:

Input restricted deque. Insertion at one end only; deletion at both. Output restricted deque. Insertion at both ends; deletion at one end only.

Both exist to recover some of the safety that the full deque gives up.

Implementation

Two choices, both already built in this module.

A doubly linked list (chapters 26 to 29). Both ends are reachable and both directions are available, so all four operations are O(1) directly.

A circular array (chapter 44) with both front and rear able to move in either direction, using modulo arithmetic in both directions. front = (front - 1 + capacity) mod capacity is the backward step, and the + capacity is there because a negative index would otherwise result.

A singly linked list is not suitable: removing from the rear is O(n) by chapter 18, so one of the four operations cannot be made constant.

Quick revision

  • Deque is a double ended queue, pronounced "deck"; MU prints it "Dequeues", which is not the dequeue

operation.

  • Four operations: add_front, add_rear, remove_front, remove_rear.
  • Used at one end it is a stack; used at both it is a queue. It contains them both.
  • The restricted structures are still preferred, because a structure that permits only what is needed

cannot be misused.

  • Input restricted allows insertion at one end only; output restricted allows deletion at one end only.
  • Implement with a doubly linked list, or a circular array stepping both ways, with
munotes.in144

The Deque: Open at Both Ends

(front - 1 + capacity) mod capacity for the backward step.

  • A singly linked list cannot do it in O(1), because removing from the rear is O(n).

Test yourself

1. What does deque stand for and how is it different from dequeue? Double ended queue. Dequeue is the operation that removes from the front of a queue; a deque is a structure open at both ends.

2. Name the four operations. add_front, add_rear, remove_front, remove_rear.

3. Which pairs of operations make a deque behave as a stack and as a queue? Adding and removing at the same end gives a stack; adding at one end and removing at the other gives a queue.

4. If a deque can do both jobs, why keep the stack and the queue? Because the restriction prevents misuse: code given a stack cannot take from the bottom and code given a queue cannot serve the newest arrival. A deque allows both mistakes, silently.

5. Define the input restricted and output restricted deques. Input restricted permits insertion at one end only and deletion at both; output restricted permits insertion at both ends and deletion at one end only.

6. Why can a singly linked list not implement a deque efficiently, and what is the backward step in a circular array? Because removing from the rear requires the node before it, which is O(n). In a circular array the backward step is (index - 1 + capacity) mod capacity, with the + capacity preventing a negative index.

Contents This chapter on its own page

munotes.in145

Chapter Forty-Eight

Job Scheduling With a Queue

Syllabus topic Module 1, "Queues: applications of queue like job scheduling queues"

In one line

Jobs waiting for a processor are held in a queue and served in arrival order, which is the scheduling policy called first come first served, and its weakness has a name and a fix.

The model

A scheduler holds jobs that are ready to run. The processor takes one at a time. In the simplest policy the jobs are served in the order they arrived, which is exactly a queue.

Three quantities are measured for each job, and an examination question will ask for them:

Meaning
Arrival timewhen the job joined the queue
Burst timehow long the job needs the processor
Completion timewhen it finished
Turnaround timecompletion minus arrival: total time in the system
Waiting timeturnaround minus burst: time spent waiting, not running

The scheduler is judged on the average waiting time, and that is what the calculations below produce.

First come first served, computed

from collections import deque


def fcfs(jobs):
    """jobs: (name, arrival, burst). Serve in arrival order, one at a time."""
    ready = deque(sorted(jobs, key=lambda j: j[1]))
    clock, rows = 0, []
    while ready:
        name, arrival, burst = ready.popleft()
        start = max(clock, arrival)          # the processor may have waited
        completion = start + burst
        turnaround = completion - arrival
        waiting = turnaround - burst
        rows.append((name, arrival, burst, start, completion, turnaround, waiting))
        clock = completion
    return rows


def report(title, rows):
    print(title)
    print("  %-6s %7s %6s %6s %10s %11s %8s"
          % ("job", "arrive", "burst", "start", "completion", "turnaround", "waiting"))
    for name, arrival, burst, start, completion, turnaround, waiting in rows:
        print("  %-6s %7d %6d %6d %10d %11d %8d"
              % (name, arrival, burst, start, completion, turnaround, waiting))
    n = len(rows)
    print("  average turnaround: %.2f" % (sum(r[5] for r in rows) / n))
    print("  average waiting   : %.2f" % (sum(r[6] for r in rows) / n))


jobs = [("P1", 0, 5), ("P2", 1, 3), ("P3", 2, 8), ("P4", 3, 2)]
report("first come first served", fcfs(jobs))
first come first served
  job     arrive  burst  start completion  turnaround  waiting
  P1           0      5      0          5           5        0
  P2           1      3      5          8           7        4
  P3           2      8      8         16          14        6
  P4           3      2     16         18          15       13
  average turnaround: 10.25
  average waiting   : 5.75

Work one row by hand to check the machine. P4 arrived at time 3 and did not start until 16, so it waited 13 and its turnaround was 15, of which only 2 was its own work. The table agrees.

The convoy effect

Look at P4. It needs the processor for 2 units and it waited 13, because a job needing 8 units happened to arrive before it.

This is the convoy effect: one long job at the front delays every short job behind it, however small they are. It is the standard criticism of first come first served, and examiners ask for it by name.

munotes.in146

Job Scheduling With a Queue

Measured, by running the same four jobs in a different order:

from collections import deque


def fcfs_average_wait(jobs):
    ready = deque(sorted(jobs, key=lambda j: j[1]))
    clock, waits = 0, []
    while ready:
        name, arrival, burst = ready.popleft()
        start = max(clock, arrival)
        waits.append(start - arrival)
        clock = start + burst
    return sum(waits) / len(waits)


long_first = [("P1", 0, 8), ("P2", 1, 2), ("P3", 2, 3), ("P4", 3, 1)]
short_first = [("P1", 0, 1), ("P2", 1, 2), ("P3", 2, 3), ("P4", 3, 8)]

print("the SAME four burst times, arriving in two different orders")
print()
print("longest first : average waiting time %.2f" % fcfs_average_wait(long_first))
print("shortest first: average waiting time %.2f" % fcfs_average_wait(short_first))
print()
print("identical work, identical policy, and the average wait differs")
print("entirely because of the order the jobs happened to arrive in.")
the SAME four burst times, arriving in two different orders

longest first : average waiting time 6.25
shortest first: average waiting time 1.00

identical work, identical policy, and the average wait differs
entirely because of the order the jobs happened to arrive in.

The same four jobs, the same total work, the same policy. The average wait is more than six times worse in one case, decided by nothing but arrival order.

What the queue cannot do about it

The obvious improvement is to serve the shortest job first. That would give the best possible average waiting time, and it is a standard result.

A queue cannot do it. Chapter 46 gave the reason: a queue knows only the order in which items arrived. To serve the shortest job first, the structure would have to be able to find the smallest burst time among everything waiting, and a queue has no way to look at anything but its front.

So the structure has to change, and what is needed is:

  • add a job with a priority, in this case its burst time
  • remove the job with the best priority, not the oldest

That is the priority queue, its ADT is chapter 79, and the structure that makes both operations cheap is the heap of chapter 81.

Where Module 1 ends

Module 1 has built five structures and each one answered a limitation of the one before it.

StructureAnswered
Arraymany values under one name, reachable by position
Linked listthe array's expensive insertion and fixed size
Stackanything nested, in O(1), safely
Queueanything waiting, in arrival order, fairly
Dequeboth ends, when the restriction is not wanted

And Module 1 ends with a limitation of its own that none of them can answer: none of these structures can find anything. Not the most urgent job, not a particular value, not the largest item, without looking at everything.

munotes.in147

Job Scheduling With a Queue

That is what Module 2 is about, from its first chapter to its last.

Quick revision

  • Jobs waiting for the processor sit in a queue; serving them in arrival order is first come first

served.

  • Turnaround time is completion minus arrival; waiting time is turnaround minus burst.
  • A scheduler is judged on the average waiting time.
  • The convoy effect: one long job at the front delays every short job behind it. Measured here, the same

four jobs gave average waits of 6.25 and 1.00 depending only on arrival order.

  • Serving the shortest job first would be better, and a queue cannot do it, because it knows only

arrival order.

  • That needs a structure that removes by priority rather than by age: the priority queue of chapter 78,

built on the heap of chapter 81.

Test yourself

1. Define turnaround time and waiting time. Turnaround is completion time minus arrival time, the total time in the system. Waiting is turnaround minus burst time, the time spent not running.

2. For a job arriving at 3 with a burst of 2 that starts at 16, give all three figures. Completion 18, turnaround 15, waiting 13.

3. What is the convoy effect? One long job at the front of the queue delaying every short job behind it, however small those jobs are.

4. The same four jobs were run in two arrival orders. What were the average waits, and what does that show? 6.25 with the longest first and 1.00 with the shortest first. The policy and the total work were identical, so the difference came entirely from arrival order, which is what first come first served is at the mercy of.

5. Why can a queue not serve the shortest job first? Because a queue knows only the order of arrival and can see only its front. Finding the smallest burst time requires looking at everything waiting, which is not a queue operation.

6. What limitation does Module 1 end on, and where is it answered? None of its structures can find anything without looking at everything: not the most urgent item, not a particular value, not the largest. Module 2 answers it, beginning with the tree.

Contents This chapter on its own page

munotes.in148

Chapter Forty-Nine

Module 1 in One Sitting

Syllabus topic Module 1, the whole module

In one line

Module 1 is five linear structures, each answering a limitation of the one before it, with the array's computed access traded away for the linked list's cheap insertion, and then restricted into the stack, the queue and the deque.

The one table to know

StructureAccess item iInsert frontInsert endDelete frontDelete endSearch
ArrayO(1)O(n)O(1)O(n)O(1)O(n), O(log n) sorted
Singly linked listO(n)O(1)O(n), O(1) with tailO(1)O(n)O(n)
Doubly linked listO(n)O(1)O(1)O(1)O(1)O(n)
Stacktop onlyO(1)O(1)O(1)O(1)not possible
Queuefront onlyO(1)O(1)O(1)O(1)not possible

Every figure in it was measured somewhere in chapters 11 to 47.

Abstract Data Types

An ADT is a set of values plus the operations on them, specified by behaviour and deliberately not by representation. The missing representation is the point: a program written against the ADT survives a change of storage, and one written against the storage does not.

Judge an ADT by three tests: complete (it does everything the problem needs), minimal (nothing redundant), honest about cost (every operation's cost documented). Returning the internal container destroys every guarantee the structure makes, including its validation.

Design one in five steps: what it holds; which operations the problem needs; what each needs, returns and does; what happens when it cannot; and only then the representation.

Arrays

A contiguous block of equal cells. Address of item i is base + (i x w), which is one multiply and one add whatever i is: O(1) access, the array's one great gift.

Contiguity also fixes the size at creation and makes insertion and deletion anywhere but the end O(n), because everything after the point must shift. Python's list hides this by growing in proportion, which makes appending amortised O(1).

2D arrays are stored row major: base + ((i x C) + j) x w.

Linked lists

A node holds a value and the address of the next node. The head is a variable holding the address of the first node; it is not a node. The last node's next is null.

Insertion and deletion are two assignments and move nothing, given the position. Reaching a position is O(n), which is why "linked lists are better for insertion" is only true when you are already there.

The order of the assignments is load bearing: point the new node forward before redirecting its predecessor, or the new node points at itself, the list gains a cycle and the rest becomes unreachable while remaining intact.

Deletion needs the predecessor, which a singly linked list can only track on the way in. A tail pointer makes appending O(1) but does not make deleting the last node O(1).

munotes.in149

Module 1 in One Sitting

A doubly linked list adds a backward link: one address per node, insertion goes from two assignments to four, and it buys O(1) deletion of a held node, O(1) deletion of the last node, and backward traversal. Its invariant is that x.next.prev is x, and both chains must always agree.

Polynomials are the named application: one node per non-zero term, exponents strictly decreasing, no zero coefficients. Addition is a one pass merge, O(m + n). The array representation is better for dense polynomials and for multiplication.

Stacks

Access at one end only, the top. Last in, first out. push, pop, peek, is_empty, size, all O(1).

Underflow is popping an empty stack and applies to every implementation. Overflow is pushing onto a full one and applies only to a fixed capacity, so to the array version.

The array stack keeps top as the index of the topmost item, with -1 meaning empty. push increments then writes; pop reads then decrements. The linked stack makes the top the head of the list, because that is the only end where adding and removing are both free.

Applications. Balanced delimiters: a counter fails because it forgets which kind of bracket was opened, and accepts ([)]. Expression conversion: infix needs precedence, associativity and brackets, while postfix and prefix need none, which is why machines use them. Shunting yard converts; postfix is evaluated by pushing operands and applying operators to the top two, with the first value popped being the right operand.

2 ^ 3 ^ 2 is 512, not 64: the exponent is right associative.

Queues

Add at the rear, remove from the front. First in, first out. All operations should be O(1).

The naive array queue drifts: both markers only move up, so every dequeue abandons a cell, and the queue reports itself full while nearly empty. Shifting on dequeue fixes the space and makes dequeue O(n). The circular queue fixes both by advancing indices with (index + 1) mod capacity.

In a circular queue, empty and full both give front == rear. The three answers are: keep a count (count == 0, count == capacity); sacrifice one cell (front == rear, (rear + 1) mod capacity == front); or keep a flag. Do not mix one method's test with another's.

The linked queue takes the head as front and the tail as rear. Its two traps are enqueueing onto an empty queue, which must set both pointers, and the dequeue that empties it, which must clear the rear.

Job scheduling is the named application. Turnaround is completion minus arrival; waiting is turnaround minus burst. First come first served suffers the convoy effect: one long job delays every short one behind it.

munotes.in150

Module 1 in One Sitting

A deque is open at both ends and contains the stack and the queue as special cases; the restricted structures are still preferred, because a structure that permits only what is needed cannot be misused.

The six "advantages and disadvantages", in one line each

MU asks this per structure, and they are different answers.

Singly linked list. For: O(1) insertion and deletion at a known position, no fixed size, no wasted capacity, no contiguous block needed. Against: O(n) access, an address per item, no backward link, poor cache behaviour, no binary search even when sorted.

Doubly linked list. For: O(1) deletion of a held node and of the last node, O(1) insertion before a node, backward traversal. Against: a second address per item, four assignments per insertion, two chains to keep in step.

Stack. For: all operations O(1), an exact match for nested problems, impossible to misuse. Against: only the top is reachable, no search, traversal destroys it, fixed capacity on the array version.

Queue. For: all operations O(1), fair with bounded waiting, decouples producer from consumer. Against: only the front is reachable, no search, and it cannot serve by urgency.

What Module 1 cannot do

None of these structures can find anything without looking at everything: not the most urgent job, not a particular value, not the largest item. Searching is O(n) on every one of them except a sorted array, which cannot then be cheaply changed.

Module 2 is the answer to exactly that.

The five mark answers most likely to be asked

  1. Write the ADT for a stack, or a queue, or a linked list. (Operations, behaviour, errors, costs. No

code.)

  1. Advantages and disadvantages of one named structure. (Its own, not another's.)
  2. Convert an infix expression to postfix or prefix, with the stack traced.
  3. Evaluate a postfix or prefix expression, with the stack traced.
  4. Why does a circular queue need a count or a sacrificed cell?
  5. Insert or delete a node at a given position in a linked list, with the pointer assignments in order.
  6. Add two polynomials represented as linked lists.
  7. First come first served scheduling: compute turnaround and waiting times, and name the convoy effect.

Test yourself

1. Give the access, front insertion and search costs of an array and of a singly linked list. Array: O(1) access, O(n) front insertion, O(n) search or O(log n) if sorted. Singly linked list: O(n) access, O(1) front insertion, O(n) search.

2. State the ordering rule for linked list pointer assignments and what happens when it is broken. Point the new node forward before redirecting its predecessor. Reversed, the new node points at itself, the list gains a cycle and the following nodes become unreachable while remaining intact.

munotes.in151

Module 1 in One Sitting

3. Why does a counter fail to check balanced brackets? It records how many brackets are open and not which kinds, so it accepts crossing pairs such as ([)] and mismatched pairs such as (].

4. Give the empty and full tests for both standard circular queue methods. With a count: empty is count == 0, full is count == capacity. Sacrificing a cell: empty is front == rear, full is (rear + 1) mod capacity == front.

5. What does a doubly linked list buy, and what does it cost? It buys O(1) deletion of a held node and of the last node, O(1) insertion before a node, and backward traversal. It costs one address per node, four assignments per insertion, and a second chain to keep in step.

6. What single thing can none of Module 1's structures do? Find anything without examining everything: the most urgent item, a particular value or the largest. Only a sorted array searches quickly, and it cannot then be changed cheaply.

Contents This chapter on its own page

munotes.in152

Module II

Trees, Priority Queues and Heaps, Graphs and Hashing

munotes.in

Chapter Fifty

From a Line to a Tree: Why Linear Structures Run Out

Syllabus topic Module 2, opening the module

In one line

Every linear structure is fast at searching or fast at changing, never both, and the tree is the structure that refuses that choice.

The wall

Module 1 built five structures. Put them against the two operations almost every real problem needs: find something, and change the collection.

StructureSearchInsert
Unsorted arrayO(n)O(1) at the end
Sorted arrayO(log n)O(n)
Singly linked listO(n)O(1) at a known position
Doubly linked listO(n)O(1) at a known position
Stack, queuenot possibleO(1)

Read the first two rows together. Sorting the array is what makes searching fast, and it is the same thing that makes insertion slow, because an insertion has to preserve the order, which means moving things.

That is not an accident of implementation. It is the shape of a line: a sorted line can be searched by halving, and keeping a line sorted means shifting.

Measured

import bisect


def sorted_array_search(data, target):
    """Binary search. Returns comparisons."""
    lo, hi, comparisons = 0, len(data) - 1, 0
    while lo <= hi:
        mid = (lo + hi) // 2
        comparisons += 1
        if data[mid] == target:
            return comparisons
        if data[mid] < target:
            lo = mid + 1
        else:
            hi = mid - 1
    return comparisons


def sorted_array_insert(data, value):
    """Insert keeping order. Returns elements moved."""
    position = bisect.bisect_left(data, value)
    moved = len(data) - position
    data.insert(position, value)
    return moved


def linked_search(values, target):
    """A walk. Returns comparisons."""
    comparisons = 0
    for item in values:
        comparisons += 1
        if item == target:
            break
    return comparisons


print("%6s | %-28s | %s" % ("n", "sorted array", "linked list"))
print("%6s | %12s %15s | %12s %12s"
      % ("", "search", "insert moves", "search", "insert moves"))
for n in (1000, 2000, 4000, 8000):
    data = list(range(0, n * 2, 2))
    values = list(data)
    search_cost = sorted_array_search(data, n)          # a value in the middle
    insert_cost = sorted_array_insert(list(data), 1)    # a small value, near the front
    link_search = linked_search(values, n)
    print("%6d | %12d %15d | %12d %12d"
          % (n, search_cost, insert_cost, link_search, 1))
     n | sorted array                 | linked list
       |       search    insert moves |       search insert moves
  1000 |            9             999 |          501            1
  2000 |           10            1999 |         1001            1
  4000 |           11            3999 |         2001            1
  8000 |           12            7999 |         4001            1

Both columns of that run are already the wall. The sorted array finds a value in 9 to 12 comparisons and pays up to 7,999 element moves to insert a small one. The linked list inserts in a single assignment and takes thousands of comparisons to find the same value.

It is still a single sample, so the next run takes the array's worst case search and its average insertion over 200 random values, which is the honest form of the comparison.

munotes.in153

From a Line to a Tree: Why Linear Structures Run Out

import bisect, random


def insert_moves(data, value):
    position = bisect.bisect_left(data, value)
    return len(data) - position


random.seed(11)
print("%6s | %-26s | %s" % ("n", "sorted array", "linked list"))
print("%6s | %10s %14s | %10s %12s"
      % ("", "search", "avg moves", "search", "insert"))
for n in (1000, 2000, 4000, 8000):
    data = list(range(0, n * 2, 2))

    worst_search = 0
    lo, hi = 0, len(data) - 1
    while lo <= hi:
        worst_search += 1
        lo = (lo + hi) // 2 + 1

    sample = [random.randrange(0, n * 2) for _ in range(200)]
    average_moves = sum(insert_moves(data, v) for v in sample) // len(sample)

    print("%6d | %10d %14d | %10d %12d"
          % (n, worst_search, average_moves, n, 2))
     n | sorted array               | linked list
       |     search      avg moves |     search       insert
  1000 |         10            542 |       1000            2
  2000 |         11           1040 |       2000            2
  4000 |         12           1942 |       4000            2
  8000 |         13           3846 |       8000            2

Now the wall is visible. At n = 8,000:

The sorted array searches in 13 comparisons and pays 3,846 element moves for an average insertion. The linked list inserts in 2 assignments and pays 8,000 comparisons for a search.

Each is excellent at one column and hopeless at the other, and the gap widens as the data grows.

What would fix it

Look at what makes binary search fast: at every step it discards half the remaining candidates. That is the whole of it. Thirteen steps for eight thousand items, because thirteen halvings of 8,000 reach 1.

Now ask what the structure would have to look like to allow that without being a line.

It would need, at each point, a value and a way to go two directions: everything smaller this way, everything larger that way. Then searching is: compare, go one way, and half the data is gone.

That is a tree. The ability to discard half at every step is what it buys, and because the halves are joined by links rather than by position, inserting does not shift anything.

[ 50 ]

/ .

[ 25 ] [ 75 ]

/ . / .

[ 10 ] [ 30 ] [ 60 ] [ 90 ]

Searching for 30: compare with 50, go left, half the tree is gone. Compare with 25, go right. Compare with 30, found. Three comparisons for seven values.

The promise of Module 2

StructureSearchInsertDelete
Binary search tree, balancedO(log n)O(log n)O(log n)
HeapO(n)O(log n)O(log n) for the extreme
Hash tableO(1) averageO(1) averageO(1) average
Graphdepends on the questionO(1)O(1)
munotes.in154

From a Line to a Tree: Why Linear Structures Run Out

The first row is the answer to this chapter. The third row is better still and gives up order entirely, which chapter 108 is about.

And each one has its own bargain, which is the habit Module 1 taught and Module 2 keeps.

Quick revision

  • Every linear structure is fast at searching or fast at changing, never both.
  • A sorted array searches in O(log n) and inserts in O(n); a linked list inserts in O(1) at a known

position and searches in O(n).

  • Measured at n = 8,000: 13 comparisons for the array's worst case search against an average of 3,846

element moves to insert; 2 assignments to insert into the list against 8,000 comparisons to search it.

  • Binary search is fast because it discards half the candidates at every step.
  • A tree allows that without being a line: a value with a smaller side and a larger side, joined by

links, so nothing shifts when something is inserted.

  • A balanced binary search tree gives O(log n) for search, insert and delete; a hash table gives O(1)

average and gives up order.

Test yourself

1. Why can a sorted array not also be fast at insertion? Because its order is carried by position, so inserting in the right place requires moving everything after it.

2. Give the two measured costs at n = 8,000 from this chapter. The sorted array searched in 13 comparisons at worst and moved an average of 3,846 elements per insertion. The linked list inserted in 2 assignments and took 8,000 comparisons to search.

3. What exactly makes binary search fast? It discards half the remaining candidates at every comparison, so the number of steps is the number of halvings needed to reach one item, which is about log n.

4. What must a structure provide to allow halving without being a line? At each point, a value and two directions: everything smaller one way and everything larger the other, joined by links rather than by position.

5. Why does inserting into a tree not shift anything? Because the order is carried by the links, not by positions in a block, so a new node is attached by changing pointers.

6. Give the costs a balanced binary search tree promises. O(log n) for search, insertion and deletion.

Contents This chapter on its own page

munotes.in155

Chapter Fifty-One

The Tree ADT: The Words, Said Exactly

Syllabus topic Module 2, "Trees: ADT for Tree Structure"

In one line

A tree is a collection of nodes with one root and every other node having exactly one parent, so there are no cycles and exactly one path from the root to each node.

The definition

A tree is either empty, or it consists of a node called the root together with zero or more subtrees, each of which is itself a tree, and the root of each subtree is connected to the tree's root by an edge.

Three things are worth noticing in that sentence.

It is recursive. A tree is defined in terms of trees. That is not a trick: it is why almost every tree algorithm in this module is three lines of recursion.

The empty tree is a tree. Leaving it out breaks every recursion at the bottom.

Order is not required. The definition says nothing about which subtree comes first, or about values being sorted. That comes later, with the binary search tree.

The vocabulary

These are the words examinations ask for. Each is defined against the same small tree.

A level 0

/ .

B C level 1

/ . .

D E F level 2

|

G level 3

TermMeaningIn the tree above
Rootthe one node with no parentA
Parentthe node directly aboveB is the parent of D and E
Childa node directly belowD and E are children of B
Siblingsnodes with the same parentD and E; B and C
Leafa node with no childrenD, F, G
Internal nodea node with at least one childA, B, C, E
Ancestorany node on the path up to the rootA and B are ancestors of D
Descendantany node belowE and G are descendants of B
Subtreea node and all its descendantsB, D, E, G is the subtree at B
Degree of a nodehow many children it hasB has degree 2, C has degree 1
Degree of the treethe largest degree of any node2
Edgea link from a parent to a childthere are 6
Paththe sequence of edges between two nodesA to G is A, B, E, G

Two definitions students most often get wrong, so they are worth saying twice:

A leaf is a node with no children, not a node at the bottom level. G is at level 3 and D is at level 2, and both are leaves.

The degree of a node is its number of children, not its number of edges including the one to its parent.

Computed, so the definitions are not merely stated

class TNode:
    def __init__(self, value, children=()):
        self.value = value
        self.children = list(children)


#              A
#            /   .
#          B       C
#        /  .       .
#      D     E       F
#            |
#            G
tree = TNode("A", [
    TNode("B", [TNode("D"), TNode("E", [TNode("G")])]),
    TNode("C", [TNode("F")]),
])


def walk(node):
    """Every node, with its parent, in no particular order."""
    stack = [(node, None)]
    while stack:
        current, parent = stack.pop()
        yield current, parent
        for child in current.children:
            stack.append((child, current))


nodes = list(walk(tree))
by_value = {n.value: (n, p) for n, p in nodes}

leaves = sorted(n.value for n, _ in nodes if not n.children)
internal = sorted(n.value for n, _ in nodes if n.children)
degrees = {n.value: len(n.children) for n, _ in nodes}


def ancestors(value):
    out, _, parent = [], *by_value[value]
    while parent is not None:
        out.append(parent.value)
        parent = by_value[parent.value][1]
    return out


def descendants(node):
    out = []
    for child in node.children:
        out.append(child.value)
        out.extend(descendants(child))
    return sorted(out)


print("nodes            :", sorted(n.value for n, _ in nodes))
print("root             :", [n.value for n, p in nodes if p is None])
print("leaves           :", leaves)
print("internal nodes   :", internal)
print("degree of each   :", dict(sorted(degrees.items())))
print("degree of the tree:", max(degrees.values()))
print("edges            :", sum(degrees.values()))
print("children of B    :", sorted(c.value for c in by_value["B"][0].children))
print("siblings of D    :", sorted(c.value for c in by_value["B"][0].children
                                   if c.value != "D"))
print("ancestors of G   :", ancestors("G"))
print("descendants of B :", descendants(by_value["B"][0]))
print()
print("a leaf is a node with NO CHILDREN, not a node at the lowest level:")
print("   D is at level 2 and is a leaf:", "D" in leaves)
print("   G is at level 3 and is a leaf:", "G" in leaves)
munotes.in156

The Tree ADT: The Words, Said Exactly

nodes            : ['A', 'B', 'C', 'D', 'E', 'F', 'G']
root             : ['A']
leaves           : ['D', 'F', 'G']
internal nodes   : ['A', 'B', 'C', 'E']
degree of each   : {'A': 2, 'B': 2, 'C': 1, 'D': 0, 'E': 1, 'F': 0, 'G': 0}
degree of the tree: 2
edges            : 6
children of B    : ['D', 'E']
siblings of D    : ['E']
ancestors of G   : ['E', 'B', 'A']
descendants of B : ['D', 'E', 'G']

a leaf is a node with NO CHILDREN, not a node at the lowest level:
   D is at level 2 and is a leaf: True
   G is at level 3 and is a leaf: True

The two properties that make it a tree

A collection of nodes and edges is a tree exactly when both hold:

1. It is connected: every node is reachable from the root. 2. It has no cycles: there is no way to leave a node and return to it.

Two consequences follow, and both are examinable:

A tree with n nodes has exactly n - 1 edges. Every node except the root is reached by exactly one edge from its parent. The run above counted 7 nodes and 6 edges.

munotes.in157

The Tree ADT: The Words, Said Exactly

There is exactly one path between any two nodes. More than one would make a cycle; none would mean it is not connected.

The ADT

OperationReturnsDoes
Tree()a treecreates it empty
root()a nodethe root; error if empty
parent(node)a nodeits parent; null for the root
children(node)nodesits children, possibly none
is_leaf(node)true or falsehas it no children
insert(parent, value)a nodeadds a child
delete(node)nothingremoves the node and its subtree
traverse()the valuesvisits every node once
height(), size()numberschapter 52

As always it names no representation: links, an array, or a list of children.

Quick revision

  • A tree is empty, or a root with zero or more subtrees, each itself a tree. The definition is

recursive, which is why tree algorithms are recursive.

  • Root: the only node with no parent. Leaf: a node with no children, whatever its level. Internal: a

node with at least one child.

  • Siblings share a parent; ancestors are on the path up to the root; descendants are everything below.
  • Degree of a node is its number of children; degree of the tree is the largest of those.
  • A tree is connected and has no cycles.
  • n nodes means exactly n - 1 edges, and exactly one path between any two nodes.
  • The ADT names root, parent, children, is_leaf, insert, delete and traverse, and no representation.

Test yourself

1. Give the recursive definition of a tree. A tree is either empty, or a root node together with zero or more subtrees, each of which is itself a tree, joined to the root by an edge.

2. Define a leaf, and say what it is not. A node with no children. It is not "a node at the lowest level": a leaf may be at any level, as D at level 2 and G at level 3 both are.

3. What is the degree of a node and the degree of a tree? The degree of a node is its number of children. The degree of the tree is the largest degree of any of its nodes.

4. How many edges has a tree of n nodes, and why? Exactly n - 1. Every node except the root is reached by exactly one edge from its parent.

5. Give the two properties that make a collection of nodes a tree. It is connected, so every node is reachable from the root, and it has no cycles.

6. In the chapter's tree, name the ancestors of G and the descendants of B. Ancestors of G: E, B, A. Descendants of B: D, E, G.

Contents This chapter on its own page

munotes.in158

Chapter Fifty-Two

Height, Depth, Level and Size, and the Relations Between Them

Syllabus topic Module 2, "Trees: ADT for Tree Structure"

In one line

Depth counts edges downward from the root, height counts edges upward from the deepest leaf, level is depth plus one in some books and equal to depth in others, and size is simply the number of nodes.

The four measures

Depth of a node. The number of edges on the path from the root to that node. The root has depth 0.

Height of a node. The number of edges on the longest path from that node down to a leaf. A leaf has height 0.

Height of the tree. The height of its root, which is the same as the greatest depth of any node.

Size. The number of nodes.

Level. Here the conventions differ, and this is the one to be careful about:

ConventionRoot is atUsed by
Level = depthlevel 0this book, and most modern texts
Level = depth + 1level 1many Indian textbooks

This book counts from 0, and says so wherever it matters. In an examination, state which convention you are using in one line and then be consistent; a stated convention cannot be marked wrong, and an unstated one can.

The two edge cases

An empty tree. Its height is conventionally -1, so that a single node tree has height 0 and the arithmetic works. Some books say 0 for both, which then breaks the relation between height and the maximum number of nodes.

A single node tree. Height 0, depth 0, size 1. It is both the root and a leaf.

Computed on a worked tree

class TNode:
    def __init__(self, value, children=()):
        self.value = value
        self.children = list(children)


#              A              depth 0
#            /   .
#          B       C          depth 1
#        /  .       .
#      D     E       F        depth 2
#            |
#            G                depth 3
tree = TNode("A", [
    TNode("B", [TNode("D"), TNode("E", [TNode("G")])]),
    TNode("C", [TNode("F")]),
])


def size(node):
    if node is None:
        return 0
    return 1 + sum(size(child) for child in node.children)


def height(node):
    """Edges on the longest downward path. A leaf is 0; an empty tree is -1."""
    if node is None:
        return -1
    if not node.children:
        return 0
    return 1 + max(height(child) for child in node.children)


def depths(node, depth=0, out=None):
    out = {} if out is None else out
    out[node.value] = depth
    for child in node.children:
        depths(child, depth + 1, out)
    return out


def leaves(node):
    if not node.children:
        return [node.value]
    found = []
    for child in node.children:
        found.extend(leaves(child))
    return found


d = depths(tree)
print("size of the tree :", size(tree))
print("height of the tree:", height(tree))
print("height of an empty tree:", height(None))
print("depth of each node:", dict(sorted(d.items())))
print("deepest node depth:", max(d.values()))
print("leaves           :", sorted(leaves(tree)))
print()
print("the relation that must hold: height of the tree == the greatest depth")
print("   height =", height(tree), "| greatest depth =", max(d.values()),
      "| equal:", height(tree) == max(d.values()))
print()
single = TNode("only")
print("a single node tree: size", size(single), "height", height(single),
      "depth 0, and it is both root and leaf:", leaves(single))
munotes.in159

Height, Depth, Level and Size, and the Relations Between Them

size of the tree : 7
height of the tree: 3
height of an empty tree: -1
depth of each node: {'A': 0, 'B': 1, 'C': 1, 'D': 2, 'E': 2, 'F': 2, 'G': 3}
deepest node depth: 3
leaves           : ['D', 'F', 'G']

the relation that must hold: height of the tree == the greatest depth
   height = 3 | greatest depth = 3 | equal: True

a single node tree: size 1 height 0 depth 0, and it is both root and leaf: ['only']

The relation holds: the height of the tree equals the greatest depth of any node. That is not a coincidence, it is what both definitions mean, and an answer that gives different numbers for the two has made an off-by-one somewhere.

Why height is the number that matters

Every cost in this module is stated in terms of height, not size.

Searching a binary search tree walks from the root to a leaf, so it costs at most the height. Inserting walks to the place where the node belongs: at most the height. Deleting: at most the height.

So the whole question of whether a tree is fast becomes: how does the height grow as nodes are added?

Two extremes, and chapters 68 and 69 are about the gap between them:

The best case. Every level is full. A tree of height h holds up to 2 to the power (h+1) minus 1 nodes, so n nodes need a height of about log2(n). For a million nodes, height 20.

The worst case. Every node has one child, so the tree is a line. Height n - 1. For a million nodes, height 999,999, and the tree is a linked list with extra steps.

import math


def best_height(n):
    """Smallest possible height for n nodes in a binary tree."""
    return math.ceil(math.log2(n + 1)) - 1


print("%10s | %12s | %14s" % ("nodes", "best height", "worst height"))
for n in (10, 100, 1000, 10**6):
    print("%10d | %12d | %14d" % (n, best_height(n), n - 1))
print()
print("the whole of balancing, chapters 69 to 74, is about staying in the left column")
     nodes |  best height |   worst height
        10 |            3 |              9
       100 |            6 |             99
      1000 |            9 |            999
   1000000 |           19 |         999999

the whole of balancing, chapters 69 to 74, is about staying in the left column

A million nodes in a height of 19, or in a height of 999,999, depending entirely on shape. That is the difference between O(log n) and O(n), and it is why half of Module 2's tree chapters are about keeping the first shape.

munotes.in160

Height, Depth, Level and Size, and the Relations Between Them

Quick revision

  • Depth of a node: edges from the root down to it. The root has depth 0.
  • Height of a node: edges on the longest path down to a leaf. A leaf has height 0.
  • Height of the tree: the height of the root, equal to the greatest depth of any node.
  • Size: the number of nodes.
  • Level: this book counts the root as level 0; many Indian textbooks count it as level 1. State your

convention in an answer.

  • An empty tree has height -1 by convention, so a one node tree has height 0.
  • Every cost in this module is in terms of height, not size.
  • Best height for n nodes is about log2(n); worst is n - 1. A million nodes: 19 against 999,999.

Test yourself

1. Define depth and height, and give both for a leaf. Depth is the number of edges from the root down to the node; height is the number of edges on the longest path from the node down to a leaf. A leaf has height 0 and whatever depth its position gives it.

2. What is the height of an empty tree, and why that value? Conventionally -1, so that a single node tree has height 0 and the relation between height and the maximum number of nodes still works.

3. Which relation must hold between the height of a tree and the depths of its nodes? The height of the tree equals the greatest depth of any node. Different numbers mean an off-by-one.

4. State the two level conventions and what to do in an examination. Level equals depth, so the root is at level 0; or level equals depth plus one, so the root is at level

  1. State which you are using and be consistent.

5. Why are tree costs stated in terms of height rather than size? Because searching, inserting and deleting all walk from the root towards a leaf, so each costs at most the height.

6. Give the best and worst height for a binary tree of 1,000,000 nodes. Best about 19, when every level is full; worst 999,999, when every node has a single child and the tree is a line.

Contents This chapter on its own page

munotes.in161

Chapter Fifty-Three

What a Tree Buys, and What It Costs

Syllabus topic Module 2, "Trees: ADT for Tree Structure. Advantages & disadvantages"

In one line

A tree buys the ability to discard a fraction of the data at every step, which turns searching from counting into halving, and it costs the discipline of keeping its shape.

What the logarithm actually means

The claim is that a balanced tree searches n items in about log2(n) steps. Here is what that is worth, measured against a linear search.

import math

print("%12s | %14s | %12s | %s" % ("items", "linear search", "tree search",
                                   "times fewer steps"))
for n in (10, 100, 1000, 10**6, 10**9):
    tree = math.ceil(math.log2(n + 1))
    print("%12d | %14d | %12d | %d" % (n, n, tree, n // tree))

print()
print("what one extra step buys a tree:")
height = 20
print("   a tree of height %d holds up to %d nodes" % (height, 2 ** (height + 1) - 1))
print("   a tree of height %d holds up to %d nodes" % (height + 1, 2 ** (height + 2) - 1))
print("   one more step, and the capacity DOUBLES")
       items |  linear search |  tree search | times fewer steps
          10 |             10 |            4 | 2
         100 |            100 |            7 | 14
        1000 |           1000 |           10 | 100
     1000000 |        1000000 |           20 | 50000
  1000000000 |     1000000000 |           30 | 33333333

what one extra step buys a tree:
   a tree of height 20 holds up to 2097151 nodes
   a tree of height 21 holds up to 4194303 nodes
   one more step, and the capacity DOUBLES

A billion items in thirty steps. That is the number worth carrying out of this chapter, because it is what the rest of Module 2 is spending its effort to protect.

And the last block is the better way to feel it: each extra step doubles what the tree can hold. Going from a thousand items to a million multiplies the data by a thousand and adds ten steps.

The advantages

1. Searching, inserting and deleting are all O(log n), in a balanced tree. No linear structure offers all three; chapter 50 measured that wall.

2. The order is kept, and is available. Unlike a hash table, a search tree can answer "what is the next key after this one", "list everything between these two values" and "what is the smallest". Chapter 59's inorder traversal produces the whole collection in sorted order, for free.

3. It grows and shrinks without shifting, because the structure is held by links.

4. It is naturally recursive, so the algorithms are short. Most of this module's tree code is three or four lines.

5. It models hierarchy directly. File systems, organisation charts, the structure of an XML document, and a compiler's parse tree are trees because the data really is a tree.

munotes.in162

What a Tree Buys, and What It Costs

The disadvantages

1. The costs are about height, not size, and height depends on shape. This is the big one. A binary search tree built from sorted input degenerates into a line, and every O(log n) becomes O(n). Chapter 68 demonstrates it.

2. Balance has to be maintained, and it is not free. The AVL tree of chapters 71 to 74 keeps the shape, and pays with rotations on insertion and deletion and a balance factor stored in every node.

3. More memory per item than an array. Two child addresses per node, plus whatever balancing needs.

4. Poorer cache behaviour than an array, for the reason chapter 22 gave: the nodes are scattered.

5. It is beaten by a hash table when order is not needed. O(log n) is not O(1), and chapter 100's hash table gives constant average time. The tree's answer is that it keeps the order, which the hash table throws away.

The comparison

Sorted arrayLinked listBalanced BSTHash table
SearchO(log n)O(n)O(log n)O(1) average
InsertO(n)O(1) at a positionO(log n)O(1) average
DeleteO(n)O(1) at a positionO(log n)O(1) average
In sorted orderfreeneeds sortingfree, inorder walknot possible
Next key after xO(log n)O(n)O(log n)not possible
Memory per itemthe valuevalue + 1 addressvalue + 2 addressesvalue + spare capacity

The tree's column is the only one with no "not possible" in it, and that is its real argument: it is the structure that is good at everything without being best at anything.

When a tree is the wrong choice

When order is never needed and speed is everything. Use a hash table. When the data never changes. Use a sorted array: binary search is simpler, uses less memory and uses the cache better. When the data is tiny. A linear search over twenty items beats anything, and is easier to get right.

Quick revision

  • A balanced tree searches, inserts and deletes in O(log n): a billion items in about thirty steps.
  • Each extra step doubles the number of nodes the tree can hold.
  • Advantages: all three operations logarithmic; the order is kept and available; growth without

shifting; naturally recursive algorithms; it models hierarchy directly.

  • Disadvantages: the cost is in the height, which depends on shape, and an unbalanced tree degenerates

to O(n); balancing costs rotations and stored information; two addresses per node; poor cache behaviour; and a hash table is faster when order is not needed.

  • The tree is the only structure in the comparison with no "not possible" against it.
  • Wrong choice when order is never needed, when the data never changes, or when the data is tiny.
munotes.in163

What a Tree Buys, and What It Costs

Test yourself

1. How many steps does a balanced tree need to search a billion items, and what does one extra step buy? About thirty. One extra step doubles the number of nodes the tree can hold.

2. Give three advantages of a tree over a sorted array. Insertion and deletion are O(log n) rather than O(n); it grows and shrinks without shifting anything; and it does not need a single contiguous block of memory.

3. Give the main disadvantage, exactly. The costs are in terms of height, and height depends on shape. An unbalanced tree, such as one built from sorted input, degenerates into a line and every operation becomes O(n).

4. What can a search tree do that a hash table cannot? Produce the items in sorted order, answer what the next key after a given one is, and report the smallest or largest. A hash table has no order at all.

5. What does a hash table do better, and what is the tree's answer? It searches and inserts in O(1) on average rather than O(log n). The tree's answer is that it keeps the order, which the hash table gives up to get that speed.

6. Name two situations where a tree is the wrong choice. When order is never needed and a hash table's constant time is wanted; and when the data never changes, where a sorted array is simpler, smaller and better for the cache.

Contents This chapter on its own page

munotes.in164

Chapter Fifty-Four

The Binary Tree

Syllabus topic Module 2, "Trees: Binary Tree-Properties"

In one line

A binary tree is a tree in which every node has at most two children, called left and right, and the two are distinguished even when only one is present.

The definition

A binary tree is either empty, or it consists of a root together with two binary trees called the left subtree and the right subtree.

Three things follow from that wording, and all three are examined.

At most two children. Not exactly two. A node may have none, one or two.

Left and right are different. A node with only a left child is not the same tree as a node with only a right child. This is the difference from a general tree, where the children are a set, and it is what makes the next chapters possible.

Empty is a binary tree. Again, so that the recursion terminates.

Why two, and why ordered

The restriction looks arbitrary and is not. It is what gives the tree a decision at every node.

In a general tree with five children, arriving at a node tells you to examine five subtrees. In a binary tree it tells you to examine one of two, and if the values are arranged, which one. That is what turns a walk into a search, and chapter 64 builds it.

The ordering of children is what lets us say "everything smaller is on the left". Without it, the phrase has no meaning.

The shapes two nodes can make

A small but useful exercise: with the same two values, how many different binary trees are there?

class BNode:
    def __init__(self, value, left=None, right=None):
        self.value = value
        self.left = left
        self.right = right


def draw(node):
    """A bracket form: value(left, right), with a dot for an empty subtree."""
    if node is None:
        return "."
    if node.left is None and node.right is None:
        return str(node.value)
    return "%s(%s, %s)" % (node.value, draw(node.left), draw(node.right))


def shapes(values):
    """Every distinct binary tree shape holding exactly these values in order."""
    if not values:
        return [None]
    out = []
    for i in range(len(values)):
        for left in shapes(values[:i]):
            for right in shapes(values[i + 1:]):
                out.append(BNode(values[i], left, right))
    return out


for values in (["A"], ["A", "B"], ["A", "B", "C"]):
    trees = shapes(values)
    print("%d value(s): %d distinct binary trees" % (len(values), len(trees)))
    for tree in trees:
        print("     ", draw(tree))
    print()
1 value(s): 1 distinct binary trees
      A

2 value(s): 2 distinct binary trees
      A(., B)
      B(A, .)

3 value(s): 5 distinct binary trees
      A(., B(., C))
      A(., C(B, .))
      B(A, C)
      C(A(., B), .)
      C(B(A, .), .)

Two values make two different trees, not one, precisely because left and right are distinguished. Three values make five.

Those counts are the Catalan numbers (1, 1, 2, 5, 14, 42, ...), and the fact worth carrying is that the number of shapes grows very fast while only one of them is the balanced one. Shape is not a detail, and keeping the good shape is what chapters 69 to 74 are about.

munotes.in165

The Binary Tree

Strictly binary, and the other names

Several adjectives are attached to binary trees and examinations mix them. Chapter 56 separates the three that matter; one more belongs here:

A strictly binary tree (also called a proper or full binary tree in some books) is one in which every node has either zero or two children, never exactly one.

That property has a neat consequence proved in the next chapter: in a strictly binary tree the number of leaves is always one more than the number of internal nodes.

The ADT

The tree ADT of chapter 51, with children replaced by two named fields.

OperationReturnsDoes
BinaryTree()a treecreates it empty
root()a nodeerror if empty
left(node), right(node)a node or nullthe two subtrees
is_leaf(node)true or falsehas it neither child
insert_left(node, v), insert_right(node, v)a nodeattaches a child
height(), size()numbersas chapter 52
traverse()the valueschapters 59 to 62, four different ways

The last row is where a binary tree differs most from a list: there is no single obvious order to visit the nodes in, and the four standard answers are the next four chapters.

Quick revision

  • A binary tree is empty, or a root with a left subtree and a right subtree, each a binary tree.
  • At most two children, not exactly two.
  • Left and right are distinguished: a node with only a left child differs from one with only a right.
  • That ordering is what makes a decision possible at each node, and so what turns a walk into a search.
  • Two values make 2 distinct binary trees, three values make 5; the counts are the Catalan numbers, and

only one shape of many is balanced.

  • A strictly binary tree has every node with zero or two children, never one.
  • A binary tree has no single obvious traversal order, which is why there are four.

Test yourself

1. Give the recursive definition of a binary tree. It is either empty, or a root together with two binary trees, the left subtree and the right subtree.

2. Why is "at most two children" not the same as "exactly two"? Because a node may have none, one or two. A tree where every node has zero or two is a strictly binary tree, which is a special case.

3. Why does it matter that left and right are distinguished? Because it allows a rule such as "everything smaller is on the left", which is what makes a search possible. In a general tree the children are unordered and no such rule can be stated.

munotes.in166

The Binary Tree

4. How many distinct binary trees hold two values, and three? Two and five. They are the Catalan numbers.

5. What is a strictly binary tree? One in which every node has either zero children or two, never exactly one.

6. Why does a binary tree need four traversals when a list needed one? Because a list has one obvious order, front to back, while a binary tree has a node and two subtrees, and the node may be visited before, between or after them, with level order as a fourth answer.

Contents This chapter on its own page

munotes.in167

Chapter Fifty-Five

Binary Tree Properties, Proved

Syllabus topic Module 2, "Trees: Binary Tree-Properties"

In one line

Four properties follow from the definition alone, and each is the answer to a standard question: the most nodes at a level, the most in a tree, the least height for a number of nodes, and the relation between leaves and two-child nodes.

Property 1: the most nodes at level i

At level i there are at most 2 to the power i nodes, counting the root as level 0.

Proof. At level 0 there is one node, the root, and 2 to the power 0 is 1. If level i has at most 2 to the power i nodes, then since each has at most two children, level i+1 has at most twice that, which is 2 to the power (i+1). By induction it holds for every level.

level 0: at most 1 = 2^0

level 1: at most 2 = 2^1

level 2: at most 4 = 2^2

level i: at most 2^i

Property 2: the most nodes in a tree of height h

A binary tree of height h has at most 2 to the power (h+1) minus 1 nodes.

Proof. Sum property 1 over all levels from 0 to h:

1 + 2 + 4 + ... + 2^h

= 2^(h+1) - 1

That is the standard sum of a geometric series, and it is worth knowing in both directions: a tree of height 3 holds at most 15 nodes, and 15 nodes need a height of at least 3.

Property 3: the least height for n nodes

Turning property 2 around: if n is at most 2 to the power (h+1) minus 1, then

h >= log2(n + 1) - 1

So the minimum height of a binary tree with n nodes is ceiling(log2(n + 1)) - 1.

This is the property that matters most in practice, and it is the one chapter 53 made concrete: a million nodes need a height of at least 19.

Property 4: leaves against nodes with two children

In any binary tree, the number of leaves is one more than the number of nodes with two children.

Written with L for leaves and T for nodes with exactly two children: L = T + 1.

Proof. Let n be the total nodes, and let the nodes with zero, one and two children number L, S and T.

n = L + S + T

Now count edges. Every node except the root is the child of exactly one edge, so there are n - 1 edges. Each node contributes edges equal to its number of children:

n - 1 = 0 x L + 1 x S + 2 x T

Substituting the first into the second:

munotes.in168

Binary Tree Properties, Proved

L + S + T - 1 = S + 2T

L - 1 = T

L = T + 1

Notice that S, the number of nodes with one child, cancels out. The relation holds whatever the shape.

All four, checked over every tree

An example proves nothing. The program below generates every binary tree shape up to 7 nodes and checks all four properties on each one.

import math


class BNode:
    def __init__(self, left=None, right=None):
        self.left = left
        self.right = right


def all_shapes(n):
    """Every binary tree shape with exactly n nodes."""
    if n == 0:
        return [None]
    out = []
    for left_size in range(n):
        for left in all_shapes(left_size):
            for right in all_shapes(n - 1 - left_size):
                out.append(BNode(left, right))
    return out


def size(node):
    return 0 if node is None else 1 + size(node.left) + size(node.right)


def height(node):
    if node is None:
        return -1
    return 1 + max(height(node.left), height(node.right))


def level_counts(node):
    counts = {}

    def walk(current, level):
        if current is None:
            return
        counts[level] = counts.get(level, 0) + 1
        walk(current.left, level + 1)
        walk(current.right, level + 1)

    walk(node, 0)
    return counts


def leaves_and_twos(node):
    if node is None:
        return 0, 0
    if node.left is None and node.right is None:
        return 1, 0
    left_leaves, left_twos = leaves_and_twos(node.left)
    right_leaves, right_twos = leaves_and_twos(node.right)
    twos = left_twos + right_twos + (1 if node.left and node.right else 0)
    return left_leaves + right_leaves, twos


checked = 0
for n in range(1, 8):
    trees = all_shapes(n)
    for tree in trees:
        h = height(tree)
        counts = level_counts(tree)
        # property 1
        assert all(c <= 2 ** level for level, c in counts.items()), "property 1"
        # property 2
        assert size(tree) <= 2 ** (h + 1) - 1, "property 2"
        # property 3
        assert h >= math.ceil(math.log2(n + 1)) - 1, "property 3"
        # property 4
        leaves, twos = leaves_and_twos(tree)
        assert leaves == twos + 1, "property 4"
        checked += 1
    print("n = %d: %4d distinct shapes, all four properties hold" % (n, len(trees)))

print()
print("binary tree shapes checked in total:", checked)
print("every one of them satisfies all four properties")
n = 1:    1 distinct shapes, all four properties hold
n = 2:    2 distinct shapes, all four properties hold
n = 3:    5 distinct shapes, all four properties hold
n = 4:   14 distinct shapes, all four properties hold
n = 5:   42 distinct shapes, all four properties hold
n = 6:  132 distinct shapes, all four properties hold
n = 7:  429 distinct shapes, all four properties hold

binary tree shapes checked in total: 625
every one of them satisfies all four properties

625 trees, every shape that exists up to 7 nodes, and all four properties hold on all of them. The shape counts down the left are the Catalan numbers again.

munotes.in169

Binary Tree Properties, Proved

An assert that never fires proves nothing by itself, so each was planted against while this chapter was written: changing property 4's check to leaves == twos fails at n = 2, and changing property 2's to a strict inequality fails at n = 1.

The properties in examination form

Property
1At most 2 to the power i nodes at level i
2At most 2 to the power (h+1) minus 1 nodes in a tree of height h
3Minimum height for n nodes is ceiling(log2(n+1)) minus 1
4Leaves = nodes with two children, plus one

Quick revision

  • Level i holds at most 2 to the power i nodes; proved by induction from the root.
  • A tree of height h holds at most 2 to the power (h+1) minus 1 nodes, by summing property 1.
  • So n nodes need a height of at least ceiling(log2(n+1)) minus 1.
  • In any binary tree, leaves = nodes with two children + 1, and the count of one-child nodes cancels out

of the proof.

  • All four were checked over every binary tree shape up to 7 nodes: 625 trees, no exceptions.
  • The shape counts 1, 2, 5, 14, 42, 132, 429 are the Catalan numbers.

Test yourself

1. State and prove the maximum number of nodes at level i. 2 to the power i. At level 0 it is 1; if level i has at most 2^i nodes then level i+1 has at most twice that, since each node has at most two children, so the result follows by induction.

2. How many nodes can a binary tree of height 4 hold at most? 2 to the power 5 minus 1, which is 31.

3. What is the minimum height of a binary tree with 100 nodes? Ceiling of log2(101) minus 1, which is 7 minus 1, so 6.

4. State the relation between leaves and nodes with two children, and say what cancels in the proof. Leaves equal nodes with two children plus one. The number of nodes with exactly one child cancels out, so the relation holds whatever the shape.

5. Why is checking every shape up to 7 nodes stronger than checking one example? Because a single example may satisfy a property by accident. Checking all 625 shapes leaves no shape of that size on which the property could fail.

6. Why does an assertion that never fires prove little on its own? Because a check that cannot fail is not a check. Each was planted against: breaking property 4's test fails at n = 2 and breaking property 2's fails at n = 1, which shows the checks can detect a fault.

Contents This chapter on its own page

munotes.in170

Chapter Fifty-Six

Full, Complete and Perfect, and Why the Difference Matters

Syllabus topic Module 2, "Trees: Binary Tree-Properties"

In one line

Full means every node has zero or two children, complete means every level is filled except possibly the last which fills from the left, and perfect means every level including the last is completely filled.

The three definitions

Full (also called strictly binary or proper). Every node has either zero or two children. No node has exactly one.

Complete. Every level is completely filled except possibly the last, and the last level is filled from the left with no gaps.

Perfect. Every internal node has two children and all leaves are at the same level. Equivalently: every level is completely filled.

The relations between them are what the question usually turns on:

  • Every perfect tree is both full and complete.
  • A full tree need not be complete, and a complete tree need not be full.
  • A perfect tree has exactly 2 to the power (h+1) minus 1 nodes.

Drawn

full but not complete complete but not full perfect

A A A

/ . / . / .

B C B C B C

/ . / . / . / .

D E D E D E F G

The first has every node with zero or two children, so it is full, but level 2 is not filled from the left across the whole level and the tree is not complete in the strict sense used here. The second has a node (C) with no children while B has two, so it is not full, yet every level is filled except the last and that fills from the left. The third is both.

Tested, over every shape

class BNode:
    def __init__(self, left=None, right=None):
        self.left = left
        self.right = right


def all_shapes(n):
    if n == 0:
        return [None]
    out = []
    for left_size in range(n):
        for left in all_shapes(left_size):
            for right in all_shapes(n - 1 - left_size):
                out.append(BNode(left, right))
    return out


def size(node):
    return 0 if node is None else 1 + size(node.left) + size(node.right)


def height(node):
    return -1 if node is None else 1 + max(height(node.left), height(node.right))


def is_full(node):
    """Every node has zero or two children."""
    if node is None:
        return True
    if (node.left is None) != (node.right is None):
        return False
    return is_full(node.left) and is_full(node.right)


def is_complete(node):
    """Level order: once a gap is seen, nothing may follow."""
    if node is None:
        return True
    queue, seen_gap = [node], False
    while queue:
        current = queue.pop(0)
        if current is None:
            seen_gap = True
        else:
            if seen_gap:
                return False
            queue.append(current.left)
            queue.append(current.right)
    return True


def is_perfect(node):
    """Every level completely filled."""
    return size(node) == 2 ** (height(node) + 1) - 1


print("%4s %8s %10s %10s %10s %16s" % ("n", "shapes", "full", "complete",
                                       "perfect", "full+complete=>perfect"))
for n in range(1, 8):
    trees = all_shapes(n)
    full = sum(1 for t in trees if is_full(t))
    complete = sum(1 for t in trees if is_complete(t))
    perfect = sum(1 for t in trees if is_perfect(t))
    both = sum(1 for t in trees if is_full(t) and is_complete(t))
    print("%4d %8d %10d %10d %10d %16s"
          % (n, len(trees), full, complete, perfect, perfect == both))

print()
print("every perfect tree is both full and complete:",
      all(not is_perfect(t) or (is_full(t) and is_complete(t))
          for n in range(1, 8) for t in all_shapes(n)))
print("a full tree need not be complete:",
      any(is_full(t) and not is_complete(t) for t in all_shapes(5)))
print("a complete tree need not be full  :",
      any(is_complete(t) and not is_full(t) for t in all_shapes(4)))
print()
print("perfect trees exist only at n = 1, 3, 7, 15, ... which is 2^(h+1) - 1")
munotes.in171

Full, Complete and Perfect, and Why the Difference Matters

   n   shapes       full   complete    perfect full+complete=>perfect
   1        1          1          1          1             True
   2        2          0          1          0             True
   3        5          1          1          1             True
   4       14          0          1          0             True
   5       42          2          1          0            False
   6      132          0          1          0             True
   7      429          5          1          1             True

every perfect tree is both full and complete: True
a full tree need not be complete: True
a complete tree need not be full  : True

perfect trees exist only at n = 1, 3, 7, 15, ... which is 2^(h+1) - 1

Three facts drop straight out of that table.

There is exactly one complete tree of each size. The column of 1s is not a coincidence: "filled from the left with no gaps" leaves no freedom at all. That uniqueness is what the heap of chapter 82 relies on to live in an array with no pointers.

Perfect trees exist only at sizes 1, 3, 7, and in general 2 to the power (h+1) minus 1. At n = 2, 4, 5, 6 there is no perfect tree of any shape.

Full and complete are genuinely independent, and together they still do not give perfect. At n = 5 there are 2 full trees and 1 complete tree, and the last column is False there: a tree can be both full and complete and still not perfect, because 5 is not one of the perfect sizes.

Why the heap needs "complete" exactly

Chapter 82 stores a tree in an array with no child pointers, using the arithmetic that the children of index i are at 2i+1 and 2i+2.

That works only if the tree has no gaps when read level by level. A gap would make the arithmetic point at an empty cell, and every index after it would be wrong.

"Complete" is precisely the word for having no gaps. Full is not enough (a full tree can have a gap in the middle of a level), and perfect is too strong (it would force the heap's size to be exactly 1, 3, 7, 15 and nothing between).

munotes.in172

Full, Complete and Perfect, and Why the Difference Matters

So "complete" is not a piece of vocabulary. It is the exact condition the next structure needs, which is the honest reason to learn the distinction.

Quick revision

  • Full: every node has zero or two children.
  • Complete: every level filled except possibly the last, which fills from the left with no gaps.
  • Perfect: every level completely filled; equivalently full and all leaves at the same level.
  • Every perfect tree is full and complete; full and complete do not imply each other.
  • A perfect tree has exactly 2 to the power (h+1) minus 1 nodes, so perfect trees exist only at sizes 1,

3, 7, 15 and so on.

  • There is exactly one complete tree of each size, which is what lets a heap live in an array.
  • The heap needs complete exactly: full is not enough and perfect is too strong.

Test yourself

1. Define full, complete and perfect. Full: every node has zero or two children. Complete: every level filled except possibly the last, which fills from the left without gaps. Perfect: every level completely filled.

2. Which implications hold between them? Every perfect tree is both full and complete. Neither full nor complete implies the other.

3. How many nodes has a perfect binary tree of height h, and what sizes can a perfect tree have? 2 to the power (h+1) minus 1, so 1, 3, 7, 15, 31 and so on. No perfect tree has 2, 4, 5 or 6 nodes.

4. How many complete binary trees are there with 6 nodes? Exactly one. "Filled from the left with no gaps" determines the shape completely.

5. Why does the heap require complete rather than full or perfect? Because it is stored in an array using index arithmetic, which needs no gaps when the tree is read level by level. A full tree may have a gap, and perfect would restrict the heap's size to 1, 3, 7, 15 only.

6. Draw a tree that is complete but not full. A root with two children where the left child has two children and the right child has none: every level is filled except the last, which fills from the left, but the right child has one fewer than two children while its sibling has two.

Contents This chapter on its own page

munotes.in173

Chapter Fifty-Eight

The Array Representation of a Binary Tree

Syllabus topic Module 2, "Trees: Implementation and Traversals"

In one line

A binary tree can be stored in an array with no pointers by putting the root at index 0 and the children of index i at 2i+1 and 2i+2, which is efficient for a complete tree and wasteful for any other.

The arithmetic

With the root at index 0:

left child of i = 2i + 1

right child of i = 2i + 2

parent of i = (i - 1) / 2, integer division

Many textbooks put the root at index 1 instead, which gives slightly tidier formulas:

left child of i = 2i

right child of i = 2i + 1

parent of i = i / 2, integer division

Both are correct. State which you use. This book uses the 0-based form, because Python and C index from 0 and mixing the two is the standard source of errors.

Why it works

Read the tree level by level, left to right, and number the nodes 0, 1, 2, 3 and so on. Then the node numbered i has its children at 2i+1 and 2i+2, always.

That is not a trick; it follows from chapter 55's property 1. Level k starts at index 2 to the power k minus 1 and holds 2 to the power k nodes, and the arithmetic falls out of that.

def children(i):
    return 2 * i + 1, 2 * i + 2


def parent(i):
    return (i - 1) // 2 if i > 0 else None


#            A(0)
#          /      .
#       B(1)       C(2)
#       /  .       /   .
#    D(3) E(4)  F(5)  G(6)
cells = ["A", "B", "C", "D", "E", "F", "G"]

print("index | value | left | right | parent")
for i, value in enumerate(cells):
    left, right = children(i)
    left_value = cells[left] if left < len(cells) else "-"
    right_value = cells[right] if right < len(cells) else "-"
    parent_value = cells[parent(i)] if parent(i) is not None else "-"
    print("%5d | %5s | %4s | %5s | %6s"
          % (i, value, left_value, right_value, parent_value))

print()
print("the relations hold for every index:")
ok = all(parent(c) == i for i in range(len(cells))
         for c in children(i) if c < len(cells))
print("   parent(child(i)) == i for every node:", ok)
index | value | left | right | parent
    0 |     A |    B |     C |      -
    1 |     B |    D |     E |      A
    2 |     C |    F |     G |      A
    3 |     D |    - |     - |      B
    4 |     E |    - |     - |      B
    5 |     F |    - |     - |      C
    6 |     G |    - |     - |      C

the relations hold for every index:
   parent(child(i)) == i for every node: True
munotes.in177

The Array Representation of a Binary Tree

The cost: it is only efficient for a complete tree

The array must reserve a cell for every position the tree could have, up to its height, whether or not a node is there.

For a complete tree there are no gaps, so an array of exactly n cells holds n nodes. Perfect.

For a degenerate tree, one where every node has a single child, the positions used are 0, then 1 or 2, then 3 or 6, and so on: the indices double at every level, so a tree of n nodes needs an array of about 2 to the power n cells.

def cells_needed(path):
    """path: 'L' or 'R' per level, starting from the root. Returns the largest index + 1."""
    i = 0
    for step in path:
        i = 2 * i + 1 if step == "L" else 2 * i + 2
    return i + 1


print("%8s | %18s | %18s | %s" % ("nodes", "complete tree", "all-left tree", "wasted"))
for n in (5, 10, 15, 20):
    complete = n
    degenerate = cells_needed("L" * (n - 1))
    print("%8d | %18d | %18d | %d cells"
          % (n, complete, degenerate, degenerate - n))
   nodes |      complete tree |      all-left tree | wasted
       5 |                  5 |                 16 | 11 cells
      10 |                 10 |                512 | 502 cells
      15 |                 15 |              16384 | 16369 cells
      20 |                 20 |             524288 | 524268 cells

Twenty nodes in a line need an array of over half a million cells, of which twenty are used. The array representation is not a general way to store a binary tree. It is the right representation for exactly one case, and that case is the heap.

Array against links

ArrayLinks
Memory per nodethe value onlyvalue plus two addresses
Reach a childarithmetic, O(1)follow a pointer, O(1)
Reach the parentarithmetic, O(1)not possible without a parent pointer
Shape other than completewastes space, exponentiallyno waste
Growtha new array, everything copiedone node
Cache behaviourgood, contiguouspoorer, scattered

The row in bold is the one students miss and it is genuinely useful: in the array form every node knows its parent for free, which the linked form does not. The heap's sift-up of chapter 84 walks from a node to the root, and it does that with (i - 1) // 2 and no stored pointers at all.

Quick revision

  • Root at index 0: children of i are at 2i+1 and 2i+2, parent of i is (i-1)/2 with integer division.
  • Root at index 1 is the other convention: children 2i and 2i+1, parent i/2. State which you use.
  • The arithmetic follows from reading the tree level by level, left to right.
  • A cell is reserved for every possible position, so it is perfect for a complete tree and exponentially
munotes.in178

The Array Representation of a Binary Tree

wasteful for a degenerate one: 20 nodes in a line need over half a million cells.

  • The array form gives the parent for free, which the linked form cannot without a stored pointer.
  • It is the right representation for exactly one structure in this paper, the heap.

Test yourself

1. Give the index arithmetic for the 0-based convention. Left child of i is 2i+1, right child is 2i+2, parent of i is (i-1) divided by 2 with integer division.

2. Give it for the 1-based convention, and say what to do in an answer. Left child 2i, right child 2i+1, parent i/2. State which convention you are using and stay consistent.

3. Why does the arithmetic work? Because the nodes are numbered level by level, left to right, and level k starts at index 2^k - 1 and holds 2^k nodes, from which the formulas follow.

4. How many cells does a 20 node tree need in the worst case, and what shape is that? Over half a million, 524,288, for a degenerate tree where every node has a single child, because the index doubles at every level.

5. What can the array representation do that the linked one cannot? Reach a node's parent in O(1) with arithmetic, without storing a parent pointer.

6. Which structure in this paper uses the array representation, and why is it suitable there? The heap. A heap is always a complete tree, so there are no gaps and no space is wasted, and the parent arithmetic is exactly what its sift-up needs.

Contents This chapter on its own page

munotes.in179

Chapter Fifty-Nine

Inorder Traversal

Syllabus topic Module 2, "Trees: Implementation and Traversals"

In one line

Inorder visits the left subtree, then the node, then the right subtree, and on a binary search tree that produces the values in sorted order.

The three recursive traversals, and what distinguishes them

A binary tree gives three things to do at each node: visit the node, go left, go right. Going left always comes before going right, so the only question is where the node itself is visited, and that gives three answers.

TraversalOrder
Preordernode, left, right
Inorderleft, node, right
Postorderleft, right, node

The names say where the node falls: pre means before the subtrees, in means between them, post means after. That is the way to remember them, and it is worth saying in an answer.

Inorder

inorder(node):

if node is null: return

inorder(node.left)

visit(node)

inorder(node.right)

Three lines, and they are the definition of the traversal.

Run, with the trace

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


#              F
#            /   .
#          B       G
#        /  .       .
#      A     D       I
#          /  .     /
#         C    E   H
tree = BNode("F",
             BNode("B", BNode("A"), BNode("D", BNode("C"), BNode("E"))),
             BNode("G", None, BNode("I", BNode("H"), None)))


def inorder(node, out=None):
    out = [] if out is None else out
    if node is None:
        return out
    inorder(node.left, out)
    out.append(node.data)
    inorder(node.right, out)
    return out


def inorder_traced(node, depth=0, lines=None):
    """The same traversal, printing what it does at each step."""
    lines = [] if lines is None else lines
    if node is None:
        return lines
    pad = "  " * depth
    lines.append("%sgo left from %s" % (pad, node.data))
    inorder_traced(node.left, depth + 1, lines)
    lines.append("%sVISIT %s" % (pad, node.data))
    lines.append("%sgo right from %s" % (pad, node.data))
    inorder_traced(node.right, depth + 1, lines)
    return lines


print("inorder:", " ".join(inorder(tree)))
print()
print("the first twelve steps of the trace:")
for line in inorder_traced(tree)[:12]:
    print("  ", line)
inorder: A B C D E F G H I

the first twelve steps of the trace:
   go left from F
     go left from B
       go left from A
       VISIT A
       go right from A
     VISIT B
     go right from B
       go left from D
         go left from C
         VISIT C
         go right from C
       VISIT D

The output is A B C D E F G H I: the values in sorted order. That is not a coincidence about this tree, and the next section is why.

Why inorder sorts a binary search tree

This tree is a binary search tree: at every node, everything in the left subtree is smaller and everything in the right subtree is larger. Chapter 64 builds them properly.

Inorder visits the left subtree, then the node, then the right subtree. So it visits everything smaller than the node, then the node, then everything larger, and that is the definition of sorted order, applied recursively all the way down.

munotes.in180

Inorder Traversal

import random


class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None


def insert(root, value):
    """Binary search tree insertion: smaller goes left, larger goes right."""
    if root is None:
        return BNode(value)
    if value < root.data:
        root.left = insert(root.left, value)
    elif value > root.data:
        root.right = insert(root.right, value)
    return root


def inorder(node, out=None):
    out = [] if out is None else out
    if node is None:
        return out
    inorder(node.left, out)
    out.append(node.data)
    inorder(node.right, out)
    return out


random.seed(3)
for trial in range(5):
    values = random.sample(range(1, 200), 12)
    root = None
    for value in values:
        root = insert(root, value)
    walked = inorder(root)
    print("trial %d: inorder == sorted(values): %s" % (trial + 1, walked == sorted(values)))

print()
values = random.sample(range(1, 100), 8)
root = None
for value in values:
    root = insert(root, value)
print("inserted in this order:", values)
print("inorder gives         :", inorder(root))
print("which is sorted       :", inorder(root) == sorted(values))
trial 1: inorder == sorted(values): True
trial 2: inorder == sorted(values): True
trial 3: inorder == sorted(values): True
trial 4: inorder == sorted(values): True
trial 5: inorder == sorted(values): True

inserted in this order: [81, 39, 54, 65, 50, 74, 45, 69]
inorder gives         : [39, 45, 50, 54, 65, 69, 74, 81]
which is sorted       : True

Five random trees and one worked example, and in every case the inorder walk is exactly the sorted values. That is the single most useful fact about inorder, and it is why a search tree can produce its contents in order for free, which chapter 53 listed as the tree's advantage over a hash table.

What inorder is for

Reading a search tree in order. The main use, as above. Printing an expression tree with brackets. Inorder on an expression tree gives infix notation, which chapter 60 shows alongside the other two. Finding the next key after a given one. The successor of a node is the next node in inorder, which is what chapter 67's deletion needs.

The cost

Every node is visited exactly once and the work at each node is constant, so O(n) time.

The memory is the recursion stack, whose depth is the height of the tree: O(h), which is O(log n) for a balanced tree and O(n) for a degenerate one. That is worth stating, because an examiner asking for the space complexity of a traversal is asking about the stack.

Quick revision

  • The three recursive traversals differ only in where the node is visited: pre means before the
munotes.in181

Inorder Traversal

subtrees, in means between them, post means after.

  • Inorder: left, node, right. Three lines.
  • On a binary search tree, inorder produces the values in sorted order, because it visits everything

smaller, then the node, then everything larger, at every level.

  • Uses: reading a search tree in order; producing infix from an expression tree; finding a node's

successor.

  • O(n) time, and O(h) memory for the recursion stack, which is O(log n) balanced and O(n) degenerate.

Test yourself

1. Give the inorder traversal in three lines. If the node is null, return. Otherwise traverse the left subtree, visit the node, then traverse the right subtree.

2. How do you remember which of the three traversals is which? The name says where the node is visited: preorder before the subtrees, inorder between them, postorder after them. Left always comes before right.

3. Why does inorder produce sorted order on a binary search tree? Because at every node it visits everything smaller than the node, then the node, then everything larger, and that is sorted order applied recursively.

4. Give the time and space complexity of a recursive traversal. O(n) time, since every node is visited once with constant work. O(h) space for the recursion stack, which is O(log n) for a balanced tree and O(n) for a degenerate one.

5. Name three uses of inorder. Reading a binary search tree in sorted order; producing infix notation from an expression tree; finding the successor of a node, which deletion needs.

6. Does inorder sort any binary tree? No. It produces sorted order only on a binary search tree, where the ordering property holds at every node.

Contents This chapter on its own page

munotes.in182

Chapter Sixty

Preorder and Postorder Traversal

Syllabus topic Module 2, "Trees: Implementation and Traversals"

In one line

Preorder visits the node before its subtrees and postorder after them, and the choice is not arbitrary: preorder copies a tree, postorder destroys one, and inorder reads one.

The one line that differs

preorder (node): visit(node); go left; go right

inorder (node): go left; visit(node); go right

postorder(node): go left; go right; visit(node)

Three identical functions with one statement moved. Everything else about them is the same.

All three, run on one tree

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


#              F
#            /   .
#          B       G
#        /  .       .
#      A     D       I
#          /  .     /
#         C    E   H
tree = BNode("F",
             BNode("B", BNode("A"), BNode("D", BNode("C"), BNode("E"))),
             BNode("G", None, BNode("I", BNode("H"), None)))


def preorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        out.append(node.data)
        preorder(node.left, out)
        preorder(node.right, out)
    return out


def inorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder(node.left, out)
        out.append(node.data)
        inorder(node.right, out)
    return out


def postorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        postorder(node.left, out)
        postorder(node.right, out)
        out.append(node.data)
    return out


print("preorder :", " ".join(preorder(tree)))
print("inorder  :", " ".join(inorder(tree)))
print("postorder:", " ".join(postorder(tree)))
print()
print("the root F is: first in preorder, in the middle in inorder, last in postorder")
print("   preorder[0]  =", preorder(tree)[0])
print("   postorder[-1] =", postorder(tree)[-1])
print("   inorder position of F =", inorder(tree).index("F"), "of", len(inorder(tree)))
preorder : F B A D C E G I H
inorder  : A B C D E F G H I
postorder: A C E D B H I G F

the root F is: first in preorder, in the middle in inorder, last in postorder
   preorder[0]  = F
   postorder[-1] = F
   inorder position of F = 5 of 9

Two facts to carry from that run, and both are used in chapter 63:

The root is first in preorder and last in postorder. Always. That is how a tree is rebuilt from its traversals.

Inorder gives the sorted order, as chapter 59 proved.

What each traversal is actually for

This is the part that makes the three worth separating.

Preorder copies a tree

To copy a tree you must create a node before you can attach its children, so you need the node first: that is preorder.

The same reason makes preorder the order in which a tree is written out to a file and read back: the parent must exist before its children can be attached to it.

Postorder destroys a tree, and computes from the bottom up

To free a node you must free its children first, or you lose the addresses of the subtrees. So deletion of a whole tree is postorder, and in C this is not optional.

munotes.in183

Preorder and Postorder Traversal

More generally, postorder is for anything where a node's answer depends on its children's answers: the height of a node, the size of a subtree, the value of an expression. The children must be finished before the parent can be computed.

Inorder reads a search tree

Chapter 59.

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


#   an expression tree for (3 + 5) x 2
expression = BNode("x", BNode("+", BNode("3"), BNode("5")), BNode("2"))


def preorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        out.append(node.data)
        preorder(node.left, out)
        preorder(node.right, out)
    return out


def inorder_bracketed(node):
    if node is None:
        return ""
    if node.left is None and node.right is None:
        return node.data
    return "(%s %s %s)" % (inorder_bracketed(node.left), node.data,
                           inorder_bracketed(node.right))


def postorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        postorder(node.left, out)
        postorder(node.right, out)
        out.append(node.data)
    return out


def evaluate(node):
    """Postorder in spirit: the children must be computed before the node."""
    if node.left is None and node.right is None:
        return int(node.data)
    left, right = evaluate(node.left), evaluate(node.right)
    return left + right if node.data == "+" else left * right


def copy_tree(node):
    """Preorder in spirit: the node must exist before its children attach."""
    if node is None:
        return None
    new = BNode(node.data)
    new.left = copy_tree(node.left)
    new.right = copy_tree(node.right)
    return new


def freed_order(node, out=None):
    """Postorder: a node's children are freed before the node itself."""
    out = [] if out is None else out
    if node is not None:
        freed_order(node.left, out)
        freed_order(node.right, out)
        out.append(node.data)
    return out


print("an expression tree for (3 + 5) x 2")
print("   preorder  gives PREFIX :", " ".join(preorder(expression)))
print("   inorder   gives INFIX  :", inorder_bracketed(expression))
print("   postorder gives POSTFIX:", " ".join(postorder(expression)))
print()
print("evaluating it needs the children first, which is postorder:",
      evaluate(expression))
print()
copy = copy_tree(expression)
print("copying needs the node first, which is preorder.")
print("   the copy evaluates to the same:", evaluate(copy))
print("   and it is a different object  :", copy is not expression)
print()
print("freeing a tree must free children before parents:")
print("   order freed:", " ".join(freed_order(expression)))
an expression tree for (3 + 5) x 2
   preorder  gives PREFIX : x + 3 5 2
   inorder   gives INFIX  : ((3 + 5) x 2)
   postorder gives POSTFIX: 3 5 + 2 x

evaluating it needs the children first, which is postorder: 16

copying needs the node first, which is preorder.
   the copy evaluates to the same: 16
   and it is a different object  : True

freeing a tree must free children before parents:
   order freed: 3 5 + 2 x
munotes.in184

Preorder and Postorder Traversal

The three traversals of an expression tree are exactly the three notations of chapter 36. That is not a coincidence: prefix, infix and postfix are preorder, inorder and postorder of the expression's tree, and saying so is a good answer to "what connects stacks and trees".

The costs

All three are O(n) time and O(h) space for the recursion stack, as chapter 59 said. They differ only in the order of the output, not in cost.

Quick revision

  • The three recursive traversals differ by one statement: where visit(node) sits.
  • Preorder: node, left, right. Inorder: left, node, right. Postorder: left, right, node.
  • The root is first in preorder and last in postorder, always.
  • Preorder copies a tree and writes it to a file: the node must exist before its children attach.
  • Postorder frees a tree and computes anything that depends on the children: height, size, the value of

an expression.

  • Inorder reads a search tree in order.
  • On an expression tree the three traversals give prefix, infix and postfix exactly.
  • All three are O(n) time and O(h) space.

Test yourself

1. Write all three traversals, showing the one line that moves. Each visits the node and recurses left then right; preorder visits before the two calls, inorder between them, postorder after them.

2. Where is the root in each traversal's output? First in preorder, last in postorder, and between the left and right subtrees' values in inorder.

3. Why must copying a tree use preorder? Because the node has to be created before its children can be attached to it.

4. Why must freeing a tree use postorder, and why is this not optional in C? Because freeing a node first would lose the addresses of its subtrees, which are stored in it, leaving the children unreachable and leaked.

5. Give the three traversals of an expression tree for (3 + 5) x 2. Preorder x + 3 5 2 which is prefix; inorder ((3 + 5) x 2) which is infix; postorder 3 5 + 2 x which is postfix.

6. Which traversal computes a node's height, and why that one? Postorder, because a node's height depends on its children's heights, so the children must be finished before the node can be computed.

Contents This chapter on its own page

munotes.in185

Chapter Sixty-One

Level Order Traversal, and the Queue It Needs

Syllabus topic Module 2, "Trees: Implementation and Traversals"

In one line

Level order visits the tree one level at a time, left to right, and it cannot be done with recursion in any natural way: it needs a queue.

The fourth traversal

The three recursive traversals all go down before they go across: they finish an entire subtree before touching its sibling.

Level order does the opposite. It visits the root, then everything at depth 1, then everything at depth 2, and so on. It is also called breadth first traversal, and chapter 94 does the same thing on a graph for the same reason.

F level 0: F

/ .

B G level 1: B G

/ . .

A D I level 2: A D I

/ . /

C E H level 3: C E H

Reading those right-hand labels top to bottom gives F B G A D I C E H, which is the answer.

Why recursion does not do this

Recursion gives you a stack, whether or not you asked for one: each call waits on the call below it. A stack is last in, first out, and it is exactly what makes the recursive traversals go deep.

Level order needs the opposite. Having seen B and G, it must deal with B's children before G's children, in the order the nodes were met. That is first in, first out, which is a queue, and recursion does not give you one.

So the algorithm is written with an explicit queue:

level_order(root):

if root is null: return

enqueue(root)

while the queue is not empty:

node = dequeue()

visit(node)

if node.left is not null : enqueue(node.left)

if node.right is not null: enqueue(node.right)

That is the whole thing, and it is worth noticing how small it is. The queue holds the nodes that have been met but not yet dealt with, which is exactly what "one level at a time" requires.

Run, with the queue shown

from collections import deque


class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


tree = BNode("F",
             BNode("B", BNode("A"), BNode("D", BNode("C"), BNode("E"))),
             BNode("G", None, BNode("I", BNode("H"), None)))


def level_order(root):
    if root is None:
        return []
    out, queue = [], deque([root])
    while queue:
        node = queue.popleft()
        out.append(node.data)
        if node.left is not None:
            queue.append(node.left)
        if node.right is not None:
            queue.append(node.right)
    return out


def level_order_traced(root):
    """The same, printing the queue at every step."""
    lines, queue = [], deque([root])
    while queue:
        before = "".join(n.data for n in queue)
        node = queue.popleft()
        if node.left is not None:
            queue.append(node.left)
        if node.right is not None:
            queue.append(node.right)
        after = "".join(n.data for n in queue)
        lines.append("  visit %s | queue before: %-6s after: %s"
                     % (node.data, before, after))
    return lines


def by_levels(root):
    """The same traversal, grouped into levels: one extra loop."""
    if root is None:
        return []
    levels, queue = [], deque([root])
    while queue:
        this_level = []
        for _ in range(len(queue)):          # exactly the nodes of this level
            node = queue.popleft()
            this_level.append(node.data)
            if node.left is not None:
                queue.append(node.left)
            if node.right is not None:
                queue.append(node.right)
        levels.append(this_level)
    return levels


print("level order:", " ".join(level_order(tree)))
print()
for line in level_order_traced(tree):
    print(line)
print()
print("grouped into levels:")
for depth, values in enumerate(by_levels(tree)):
    print("   level %d: %s" % (depth, " ".join(values)))
munotes.in186

Level Order Traversal, and the Queue It Needs

level order: F B G A D I C E H

  visit F | queue before: F      after: BG
  visit B | queue before: BG     after: GAD
  visit G | queue before: GAD    after: ADI
  visit A | queue before: ADI    after: DI
  visit D | queue before: DI     after: ICE
  visit I | queue before: ICE    after: CEH
  visit C | queue before: CEH    after: EH
  visit E | queue before: EH     after: H
  visit H | queue before: H      after:

grouped into levels:
   level 0: F
   level 1: B G
   level 2: A D I
   level 3: C E H

Follow the queue column. It never holds more than two levels at once, and the order it hands nodes back is exactly the order they were met.

Grouping into levels: the one extra line

The by_levels function above does something worth knowing, because it is a standard examination variant: it prints each level on its own line.

The trick is the inner for _ in range(len(queue)). At the top of the outer loop, the queue holds exactly the nodes of the current level and nothing else, so taking that many is taking one level. Capturing len(queue) before the loop is what makes it work, and reading it inside the loop is the error, because the queue grows as children are added.

The cost

Time O(n): each node is enqueued once and dequeued once.

Space O(w), where w is the maximum width of the tree, the most nodes on any one level. For a perfect tree the last level holds about half the nodes, so this is O(n/2), which is O(n).

That is worth comparing with the recursive traversals, which cost O(h) for the stack:

Recursive traversalsLevel order
MemoryO(h), the heightO(w), the width
Balanced treeO(log n)O(n)
Degenerate treeO(n)O(1)

They are opposites, and that is a genuinely useful thing to know: the depth first traversals are cheap on wide shallow trees and the breadth first one is cheap on narrow deep trees.

Quick revision

  • Level order visits the tree one level at a time, left to right. Also called breadth first.
  • It cannot be done naturally with recursion, because recursion gives a stack and this needs a queue.
  • Algorithm: enqueue the root; while the queue is not empty, dequeue a node, visit it, and enqueue its
munotes.in187

Level Order Traversal, and the Queue It Needs

children.

  • To group by level, capture len(queue) before an inner loop: at that moment the queue holds exactly

one level.

  • O(n) time; O(w) memory, where w is the maximum width.
  • The recursive traversals cost O(h) and this costs O(w), so they are opposites: depth first is cheap on

wide trees and breadth first is cheap on deep ones.

  • The same algorithm on a graph is chapter 94's breadth first search.

Test yourself

1. What is level order and what else is it called? Visiting the tree one level at a time, left to right. It is also called breadth first traversal.

2. Why can it not be written as a simple recursion? Because recursion supplies a stack, which is last in first out and makes a traversal go deep. Level order needs nodes dealt with in the order they were met, which is first in first out: a queue.

3. Write the algorithm. Enqueue the root. While the queue is not empty: dequeue a node, visit it, and enqueue its left and right children if they exist.

4. How do you print each level on its own line? Capture the queue's length at the top of the outer loop and take exactly that many nodes, because at that moment the queue holds exactly the current level.

5. Give the time and space costs, and compare the space with a recursive traversal. O(n) time. O(w) space where w is the maximum width, against O(h) for the recursive traversals. They are opposites: breadth first is cheap on deep narrow trees and depth first on wide shallow ones.

6. Where does this algorithm reappear in Module 2? As breadth first search on a graph, chapter 94, which is the same loop with a visited set added.

Contents This chapter on its own page

munotes.in188

Chapter Sixty-Two

Iterative Traversal, and the Stack It Needs

Syllabus topic Module 2, "Trees: Implementation and Traversals"

In one line

Every recursive traversal can be written with an explicit stack, because that is exactly what the recursion was using, and the iterative version is what you need when the tree is deep enough to overflow the call stack.

What recursion actually does

When inorder(node.left) is called, the machine must remember where to come back to and what node was. It puts that on the call stack, and takes it off when the call returns.

So a recursive traversal is already using a stack. Writing it iteratively does not add a stack; it makes the existing one visible and puts it under your control.

Iterative inorder

stack = empty

current = root

while current is not null or the stack is not empty:

while current is not null: push every left node on the way down

push(current)

current = current.left

current = pop() the leftmost unvisited node

visit(current)

current = current.right then its right subtree

The inner while goes as far left as possible, pushing as it goes. When it can go no further, the top of the stack is the next node to visit in inorder.

All three, iterative, run against the recursive versions

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


tree = BNode("F",
             BNode("B", BNode("A"), BNode("D", BNode("C"), BNode("E"))),
             BNode("G", None, BNode("I", BNode("H"), None)))


def inorder_recursive(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder_recursive(node.left, out)
        out.append(node.data)
        inorder_recursive(node.right, out)
    return out


def preorder_recursive(node, out=None):
    out = [] if out is None else out
    if node is not None:
        out.append(node.data)
        preorder_recursive(node.left, out)
        preorder_recursive(node.right, out)
    return out


def postorder_recursive(node, out=None):
    out = [] if out is None else out
    if node is not None:
        postorder_recursive(node.left, out)
        postorder_recursive(node.right, out)
        out.append(node.data)
    return out


def inorder_iterative(root):
    out, stack, current = [], [], root
    while current is not None or stack:
        while current is not None:
            stack.append(current)
            current = current.left
        current = stack.pop()
        out.append(current.data)
        current = current.right
    return out


def preorder_iterative(root):
    """Push the RIGHT child first, so the left is popped first."""
    if root is None:
        return []
    out, stack = [], [root]
    while stack:
        node = stack.pop()
        out.append(node.data)
        if node.right is not None:
            stack.append(node.right)
        if node.left is not None:
            stack.append(node.left)
    return out


def postorder_iterative(root):
    """Preorder with the children pushed the other way, then reversed."""
    if root is None:
        return []
    out, stack = [], [root]
    while stack:
        node = stack.pop()
        out.append(node.data)
        if node.left is not None:
            stack.append(node.left)
        if node.right is not None:
            stack.append(node.right)
    return list(reversed(out))


pairs = [("inorder", inorder_recursive, inorder_iterative),
         ("preorder", preorder_recursive, preorder_iterative),
         ("postorder", postorder_recursive, postorder_iterative)]

print("%-11s %-22s %-22s %s" % ("traversal", "recursive", "iterative", "agree"))
for name, recursive, iterative in pairs:
    a, b = recursive(tree), iterative(tree)
    print("%-11s %-22s %-22s %s" % (name, " ".join(a), " ".join(b), a == b))
munotes.in189

Iterative Traversal, and the Stack It Needs

traversal   recursive              iterative              agree
inorder     A B C D E F G H I      A B C D E F G H I      True
preorder    F B A D C E G I H      F B A D C E G I H      True
postorder   A C E D B H I G F      A C E D B H I G F      True

Each traversal produced twice, by two different methods, agreeing on all three. That is this book's standing habit and it has caught two real faults already.

Two details in those functions are worth a mark each.

Iterative preorder pushes the RIGHT child first. A stack reverses, so pushing right then left means left is popped first, which is the order preorder needs.

Iterative postorder is preorder with the children swapped, then reversed. Postorder is left, right, node. Reversed, that is node, right, left, which is a preorder that visits right before left. So do that and reverse the answer. It is the neatest of the three and it is a standard question.

When the iterative version is actually needed

Recursion is clearer, so why write the other?

Because the call stack has a limit, and a degenerate tree reaches it. A tree built from sorted input is a line, so a recursive traversal recurses once per node.

import sys


class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None


def build_line(n):
    """A degenerate tree: every node has only a right child."""
    root = BNode(0)
    walk = root
    for i in range(1, n):
        walk.right = BNode(i)
        walk = walk.right
    return root


def inorder_recursive(node, out):
    if node is not None:
        inorder_recursive(node.left, out)
        out.append(node.data)
        inorder_recursive(node.right, out)


def inorder_iterative(root):
    out, stack, current = [], [], root
    while current is not None or stack:
        while current is not None:
            stack.append(current)
            current = current.left
        current = stack.pop()
        out.append(current.data)
        current = current.right
    return out


print("Python's recursion limit here:", sys.getrecursionlimit())

tree = build_line(5000)

try:
    out = []
    inorder_recursive(tree, out)
    print("recursive on a 5,000 node line: visited", len(out))
except RecursionError as e:
    print("recursive on a 5,000 node line: FAILED,", type(e).__name__)

out = inorder_iterative(tree)
print("iterative on the same tree      : visited", len(out), "with no trouble")
print("the iterative stack grew to at most the tree's height, which here is", 4999)
Python's recursion limit here: 1000
recursive on a 5,000 node line: FAILED, RecursionError
iterative on the same tree      : visited 5000 with no trouble
the iterative stack grew to at most the tree's height, which here is 4999

The recursive version fails on a tree of 5,000 nodes in a line, because the recursion is 5,000 deep and the limit is 1,000. The iterative version walks it without difficulty, because its stack is an ordinary list on the heap rather than the machine's call stack.

munotes.in190

Iterative Traversal, and the Stack It Needs

That is the real answer to "why learn the iterative version": a balanced tree never needs it, and an unbalanced one does, which is one more reason the balancing of chapters 69 to 74 matters.

Quick revision

  • Recursion is a stack you did not have to write; the iterative versions make it explicit.
  • Iterative inorder: push every node while going left; pop, visit, then go right.
  • Iterative preorder: push the root; pop, visit, push right then left, so left pops first.
  • Iterative postorder: do preorder with left and right pushed the other way round, then reverse the

result.

  • All three were run against the recursive versions and agreed.
  • The recursive version fails on a deep tree: a 5,000 node line exceeded Python's 1,000 call limit while

the iterative version walked it.

  • A balanced tree is never deep enough to need the iterative form, which is one more argument for

balancing.

Test yourself

1. What does writing a traversal iteratively actually change? Nothing about the algorithm. It replaces the machine's call stack with one you declare and control.

2. Write iterative inorder. With a stack and a current pointer: while current is not null or the stack is not empty, push every node while moving left; then pop, visit it, and move to its right child.

3. In iterative preorder, why is the right child pushed before the left? Because a stack reverses the order, so pushing right then left causes the left child to be popped first, which is what preorder requires.

4. Describe the neat way to write iterative postorder. Run a preorder that pushes left before right, so it visits node, right, left, then reverse the output, which gives left, right, node.

5. When does the recursive version actually fail? On a deep tree. A 5,000 node degenerate tree exceeded Python's recursion limit of 1,000, while the iterative version completed.

6. Why does this matter less for a balanced tree? Because a balanced tree of a million nodes has height about 20, so the recursion is never more than about 20 deep.

Contents This chapter on its own page

munotes.in191

Chapter Sixty-Three

Rebuilding a Tree From Two Traversals

Syllabus topic Module 2, "Trees: Implementation and Traversals"

In one line

Inorder with either preorder or postorder determines the tree uniquely, because one gives the root and the other says which values are on each side of it, but preorder with postorder does not.

Why inorder plus preorder works

Two facts do all the work.

Preorder's first value is the root. Chapter 60: preorder visits the node before its subtrees.

Inorder splits at the root. Everything to the left of the root in the inorder sequence is the left subtree, everything to the right is the right subtree.

So: take the root from preorder, find it in inorder, and the inorder sequence is cut into two halves, which are the two subtrees. Their sizes then say how to cut the preorder sequence, and the whole thing repeats.

preorder: F B A D C E G I H

inorder : A B C D E F G H I

root is F, the first of preorder.

in inorder, F splits: [A B C D E] F [G H I]

left subtree has 5 nodes, right has 3.

so preorder splits: F [B A D C E] [G I H]

then repeat on each half.

Built and run, both ways

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


def preorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        out.append(node.data)
        preorder(node.left, out)
        preorder(node.right, out)
    return out


def inorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder(node.left, out)
        out.append(node.data)
        inorder(node.right, out)
    return out


def postorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        postorder(node.left, out)
        postorder(node.right, out)
        out.append(node.data)
    return out


def shape(node):
    if node is None:
        return "."
    if node.left is None and node.right is None:
        return str(node.data)
    return "%s(%s, %s)" % (node.data, shape(node.left), shape(node.right))


def from_pre_in(pre, ino):
    """Rebuild from preorder and inorder."""
    if not pre:
        return None
    root_value = pre[0]
    split = ino.index(root_value)
    left_size = split
    return BNode(root_value,
                 from_pre_in(pre[1:1 + left_size], ino[:split]),
                 from_pre_in(pre[1 + left_size:], ino[split + 1:]))


def from_post_in(post, ino):
    """Rebuild from postorder and inorder. The root is the LAST of postorder."""
    if not post:
        return None
    root_value = post[-1]
    split = ino.index(root_value)
    left_size = split
    return BNode(root_value,
                 from_post_in(post[:left_size], ino[:split]),
                 from_post_in(post[left_size:-1], ino[split + 1:]))


original = BNode("F",
                 BNode("B", BNode("A"), BNode("D", BNode("C"), BNode("E"))),
                 BNode("G", None, BNode("I", BNode("H"), None)))

pre, ino, post = preorder(original), inorder(original), postorder(original)
print("preorder :", " ".join(pre))
print("inorder  :", " ".join(ino))
print("postorder:", " ".join(post))
print()
print("original          :", shape(original))

rebuilt_a = from_pre_in(pre, ino)
print("from pre + in     :", shape(rebuilt_a),
      "| identical:", shape(rebuilt_a) == shape(original))

rebuilt_b = from_post_in(post, ino)
print("from post + in    :", shape(rebuilt_b),
      "| identical:", shape(rebuilt_b) == shape(original))
munotes.in192

Rebuilding a Tree From Two Traversals

preorder : F B A D C E G I H
inorder  : A B C D E F G H I
postorder: A C E D B H I G F

original          : F(B(A, D(C, E)), G(., I(H, .)))
from pre + in     : F(B(A, D(C, E)), G(., I(H, .))) | identical: True
from post + in    : F(B(A, D(C, E)), G(., I(H, .))) | identical: True

Both reconstructions give back exactly the original tree.

Why preorder plus postorder does not work

Preorder gives the root first. Postorder gives the root last. Neither tells you where the split is, and inorder was the thing that did.

The failure is not subtle, and it does not need a large tree. The program below searches every binary tree shape of a given size and finds two different trees with identical preorder and postorder.

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


def preorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        out.append(node.data)
        preorder(node.left, out)
        preorder(node.right, out)
    return out


def postorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        postorder(node.left, out)
        postorder(node.right, out)
        out.append(node.data)
    return out


def inorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder(node.left, out)
        out.append(node.data)
        inorder(node.right, out)
    return out


def shape(node):
    if node is None:
        return "."
    if node.left is None and node.right is None:
        return str(node.data)
    return "%s(%s, %s)" % (node.data, shape(node.left), shape(node.right))


# The smallest counterexample there is: two nodes, two ways round.
left_child = BNode("A", BNode("B"), None)      # B is A's LEFT child
right_child = BNode("A", None, BNode("B"))     # B is A's RIGHT child

print("two DIFFERENT trees:")
print("   tree 1:", shape(left_child))
print("   tree 2:", shape(right_child))
print()
print("%-12s %-10s %s" % ("", "tree 1", "tree 2"))
print("%-12s %-10s %s" % ("preorder",
                          " ".join(preorder(left_child)),
                          " ".join(preorder(right_child))))
print("%-12s %-10s %s" % ("postorder",
                          " ".join(postorder(left_child)),
                          " ".join(postorder(right_child))))
print("%-12s %-10s %s" % ("inorder",
                          " ".join(inorder(left_child)),
                          " ".join(inorder(right_child))))
print()
print("preorder identical :", preorder(left_child) == preorder(right_child))
print("postorder identical:", postorder(left_child) == postorder(right_child))
print("inorder differs    :", inorder(left_child) != inorder(right_child))
print()
print("so preorder and postorder TOGETHER cannot tell these two trees apart,")
print("and inorder is the only one of the three that can.")
two DIFFERENT trees:
   tree 1: A(B, .)
   tree 2: A(., B)

             tree 1     tree 2
preorder     A B        A B
postorder    B A        B A
inorder      B A        A B

preorder identical : True
postorder identical: True
inorder differs    : True

so preorder and postorder TOGETHER cannot tell these two trees apart,
and inorder is the only one of the three that can.
munotes.in193

Rebuilding a Tree From Two Traversals

Two nodes is enough. A with a left child B, and A with a right child B, have the same preorder and the same postorder, and are different trees. Only inorder separates them.

That is the whole reason inorder is one of the required two: it is the only one of the three that records which side a subtree is on.

The exception worth knowing

Preorder plus postorder does determine the tree when the tree is full, that is when every node has zero or two children (chapter 56). The counterexample above has a node with exactly one child, and that is precisely the ambiguity: with one child there is no way to tell left from right, and a full tree has no such node.

An examiner asking "can a tree be built from preorder and postorder" is usually looking for "no, unless it is a full binary tree".

The cost

The reconstruction as written is O(n squared) in the worst case, because ino.index(root) searches the inorder list at every step and the slicing copies.

It becomes O(n) by building a dictionary from value to its position in inorder once at the start, and passing index ranges instead of slices. That improvement is worth mentioning in an answer.

Quick revision

  • Preorder gives the root first; postorder gives it last; inorder splits the sequence at the root into

the left and right subtrees.

  • Inorder with preorder, or inorder with postorder, determines the tree uniquely.
  • Preorder with postorder does not: A with a left child B and A with a right child B share both,

and only inorder tells them apart.

  • The ambiguity is exactly a node with one child, so preorder plus postorder does determine a FULL

binary tree.

  • Inorder is required because it is the only one of the three that records which side a subtree is on.
  • The simple reconstruction is O(n squared); a position lookup table and index ranges make it O(n).

Test yourself

1. Which pairs of traversals determine a binary tree uniquely? Inorder with preorder, and inorder with postorder. Preorder with postorder does not, except for a full binary tree.

2. Give the algorithm for rebuilding from preorder and inorder. Take the first value of preorder as the root, find it in inorder, and the inorder sequence splits into the left and right subtrees. Their sizes cut the preorder sequence correspondingly, and the method repeats on each half.

3. Give a counterexample for preorder plus postorder. A root A with a single left child B, and a root A with a single right child B. Both have preorder A B and postorder B A, and they are different trees.

4. Why is inorder always one of the required two? Because it is the only traversal that records which side of the root a subtree lies on; preorder and postorder both only say that a group of nodes is below the root.

munotes.in194

Rebuilding a Tree From Two Traversals

5. When does preorder plus postorder suffice, and why? When the tree is full, every node having zero or two children. The ambiguity arises exactly at a node with one child, and a full tree has none.

6. What is the cost of the straightforward reconstruction, and how is it improved? O(n squared), because of the search for the root in inorder and the slicing at each step. Precomputing a value to position map and passing index ranges rather than slices makes it O(n).

Contents This chapter on its own page

munotes.in195

Chapter Sixty-Four

The Binary Search Tree: The Invariant

Syllabus topic Module 2, "Trees: Binary Search Tree"

In one line

A binary search tree is a binary tree in which, at every node, every value in the left subtree is smaller and every value in the right subtree is larger.

The invariant, said correctly

For every node x: every value in x's left subtree is less than x, and every value in x's right subtree is greater than x.

The wording that students usually give is: "the left child is smaller and the right child is larger." That is not the same thing and it is not enough. It talks about the two children only, and says nothing about grandchildren.

Here is a tree that satisfies the wrong rule and is not a search tree:

20

/ .

10 30

/ .

5 25

Check it with the wrong rule: at 20, left child 10 is smaller and right child 30 is larger, fine. At 10, left child 5 is smaller and right child 25 is larger, fine. Every node passes.

But 25 is in the left subtree of 20, and 25 is greater than 20. Searching for 25 from the root would go right at 20 and never find it. It is not a search tree.

A checker that tells the two rules apart

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


def shape(node):
    if node is None:
        return "."
    if node.left is None and node.right is None:
        return str(node.data)
    return "%s(%s, %s)" % (node.data, shape(node.left), shape(node.right))


def children_only(node):
    """The WRONG rule: compares a node with its two children only."""
    if node is None:
        return True
    if node.left is not None and node.left.data >= node.data:
        return False
    if node.right is not None and node.right.data <= node.data:
        return False
    return children_only(node.left) and children_only(node.right)


def whole_subtree(node, low=None, high=None):
    """The RIGHT rule: every value in the subtree must lie in a range."""
    if node is None:
        return True
    if low is not None and node.data <= low:
        return False
    if high is not None and node.data >= high:
        return False
    return (whole_subtree(node.left, low, node.data)
            and whole_subtree(node.right, node.data, high))


#        20                        20
#      /    .                    /    .
#   10       30     against   10       30
#  /  .                      /  .
# 5    25                   5    15
impostor = BNode(20, BNode(10, BNode(5), BNode(25)), BNode(30))
genuine = BNode(20, BNode(10, BNode(5), BNode(15)), BNode(30))

for name, tree in (("impostor", impostor), ("genuine ", genuine)):
    print("%s %-28s children-only rule: %-5s  whole-subtree rule: %s"
          % (name, shape(tree), children_only(tree), whole_subtree(tree)))

print()
print("the impostor passes the wrong rule and fails the right one.")
print("25 sits in the LEFT subtree of 20 and is greater than 20.")
impostor 20(10(5, 25), 30)            children-only rule: True   whole-subtree rule: False
genuine  20(10(5, 15), 30)            children-only rule: True   whole-subtree rule: True

the impostor passes the wrong rule and fails the right one.
25 sits in the LEFT subtree of 20 and is greater than 20.
munotes.in196

The Binary Search Tree: The Invariant

The right way to write the check is the one above: carry a range down the tree. Every node must lie strictly between the bounds it inherits, and it narrows the bound for each child.

The second way to check it

There is a neater test, and it is a direct consequence of chapter 59:

A binary tree is a binary search tree exactly when its inorder traversal is strictly increasing.

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


def inorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder(node.left, out)
        out.append(node.data)
        inorder(node.right, out)
    return out


def whole_subtree(node, low=None, high=None):
    if node is None:
        return True
    if low is not None and node.data <= low:
        return False
    if high is not None and node.data >= high:
        return False
    return (whole_subtree(node.left, low, node.data)
            and whole_subtree(node.right, node.data, high))


def inorder_increasing(node):
    walk = inorder(node)
    return all(walk[i] < walk[i + 1] for i in range(len(walk) - 1))


impostor = BNode(20, BNode(10, BNode(5), BNode(25)), BNode(30))
genuine = BNode(20, BNode(10, BNode(5), BNode(15)), BNode(30))

for name, tree in (("impostor", impostor), ("genuine ", genuine)):
    print("%s inorder %-22s increasing: %-6s range check: %s"
          % (name, str(inorder(tree)), inorder_increasing(tree),
             whole_subtree(tree)))

print()
print("the two tests agree on both trees:",
      all(inorder_increasing(t) == whole_subtree(t)
          for t in (impostor, genuine)))
impostor inorder [5, 10, 25, 20, 30]    increasing: False  range check: False
genuine  inorder [5, 10, 15, 20, 30]    increasing: True   range check: True

the two tests agree on both trees: True

The impostor's inorder is 5, 10, 25, 20, 30, which is not increasing: 25 comes before 20. That is the same defect seen from a different angle.

Duplicates

The invariant as stated uses strictly less and strictly greater, so duplicates are not allowed. Three conventions exist and an answer should name the one it uses:

Forbid them. An insertion of an existing value does nothing. This is what this book does. All duplicates to the right, using "less than" on the left and "greater than or equal" on the right. Keep a count in each node, which stores one node per distinct value with a multiplicity.

The third is usually the best in practice and is worth mentioning.

What the invariant buys

Everything in the next three chapters. At any node, a comparison tells you which one of the two subtrees can possibly contain your value, so the other is discarded entirely. That is the halving of chapter 50, and it is the whole purpose of the arrangement.

munotes.in197

The Binary Search Tree: The Invariant

Quick revision

  • Binary search tree: at every node, every value in the left subtree is less and every value in the right

subtree is greater.

  • "The left child is smaller and the right child is larger" is NOT the invariant: it allows a tree where

a grandchild sits on the wrong side, which search cannot find.

  • Check it by carrying a range down the tree, narrowing it at each step.
  • Equivalently: a binary tree is a search tree exactly when its inorder traversal is strictly increasing.
  • Duplicates are excluded by the strict inequalities; the three conventions are to forbid them, to send

them right, or to store a count per node.

  • The invariant is what lets one comparison discard an entire subtree.

Test yourself

1. State the binary search tree invariant exactly. For every node, every value in its left subtree is less than it and every value in its right subtree is greater than it.

2. Why is "the left child is smaller and the right child is larger" wrong? Because it only constrains the immediate children. A grandchild may then sit on the wrong side of an ancestor, as 25 in the left subtree of 20, and a search for it will go the other way and fail.

3. Describe the correct way to check the invariant programmatically. Carry a lower and an upper bound down the tree. Each node must lie strictly between them, and it becomes the upper bound for its left subtree and the lower bound for its right.

4. Give the equivalent test using a traversal. A binary tree is a binary search tree exactly when its inorder traversal is strictly increasing.

5. What is the inorder of the impostor tree in this chapter, and what does it show? 5, 10, 25, 20, 30. It is not increasing, since 25 comes before 20, which shows the tree is not a search tree.

6. Name the three conventions for duplicate values. Forbid them; place them all in the right subtree using a non-strict comparison there; or store a count in each node so one node holds a value and its multiplicity.

Contents This chapter on its own page

munotes.in198

Chapter Sixty-Five

Searching a Binary Search Tree

Syllabus topic Module 2, "Trees: Binary Search Tree"

In one line

Searching compares the target with the node and goes left or right, discarding an entire subtree at each step, so the cost is the height of the tree.

The algorithm

search(node, target):

while node is not null:

if target == node.data: found

if target < node.data: node = node.left

else: node = node.right

not found

Four lines, no recursion needed, and each comparison eliminates one of the two subtrees entirely. That is the invariant of chapter 64 being spent.

Run, with the path shown

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None


def insert(root, value):
    if root is None:
        return BNode(value)
    if value < root.data:
        root.left = insert(root.left, value)
    elif value > root.data:
        root.right = insert(root.right, value)
    return root


def build(values):
    root = None
    for value in values:
        root = insert(root, value)
    return root


def search(root, target):
    """Returns (found, comparisons, path)."""
    node, comparisons, path = root, 0, []
    while node is not None:
        comparisons += 1
        path.append(node.data)
        if target == node.data:
            return True, comparisons, path
        node = node.left if target < node.data else node.right
    return False, comparisons, path


tree = build([50, 30, 70, 20, 40, 60, 80, 35, 45])

for target in (50, 35, 80, 55):
    found, comparisons, path = search(tree, target)
    print("search %-3d -> %-5s in %d comparison(s), path: %s"
          % (target, found, comparisons, " ".join(str(p) for p in path)))

print()
print("searching for 35 went: 50, left to 30, right to 40, left to 35.")
print("at 50 the whole right subtree (60, 70, 80) was discarded in one comparison.")
search 50  -> True  in 1 comparison(s), path: 50
search 35  -> True  in 4 comparison(s), path: 50 30 40 35
search 80  -> True  in 3 comparison(s), path: 50 70 80
search 55  -> False in 3 comparison(s), path: 50 70 60

searching for 35 went: 50, left to 30, right to 40, left to 35.
at 50 the whole right subtree (60, 70, 80) was discarded in one comparison.

Note the unsuccessful search for 55: it stops after 3 comparisons, not after examining the whole tree. An unsuccessful search in a BST costs the height, not n, which is the other half of the gain and is easy to forget.

The cost, measured against a linear search

import random


class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None


def insert(root, value):
    if root is None:
        return BNode(value)
    if value < root.data:
        root.left = insert(root.left, value)
    elif value > root.data:
        root.right = insert(root.right, value)
    return root


def search_comparisons(root, target):
    node, comparisons = root, 0
    while node is not None:
        comparisons += 1
        if target == node.data:
            return comparisons
        node = node.left if target < node.data else node.right
    return comparisons


def height(node):
    if node is None:
        return -1
    return 1 + max(height(node.left), height(node.right))


random.seed(5)
print("%7s | %8s | %16s | %14s" % ("n", "height", "BST comparisons", "linear, worst"))
for n in (1000, 2000, 4000, 8000):
    values = random.sample(range(n * 10), n)
    root = None
    for value in values:
        root = insert(root, value)
    sample = random.sample(values, 200)
    worst = max(search_comparisons(root, v) for v in sample)
    print("%7d | %8d | %16d | %14d" % (n, height(root), worst, n))
munotes.in199

Searching a Binary Search Tree

      n |   height |  BST comparisons |  linear, worst
   1000 |       20 |               19 |           1000
   2000 |       25 |               23 |           2000
   4000 |       25 |               25 |           4000
   8000 |       31 |               27 |           8000

Eight thousand values, and the worst search among 200 tried took 27 comparisons. The linear search of chapter 16 would take up to 8,000.

Notice also that the heights run to 31 where the ideal for 8,000 nodes is 12. A tree built from random insertions is not perfectly balanced, but it is close enough to be useful: random insertion gives a height of about 1.39 times log2(n), which is a classical result and is what these numbers show.

The three costs

CaseComparisons
Best1, the target is the root
Average, random treeabout 1.39 log2(n)
Worst, balancedthe height, about log2(n)
Worst, degeneraten, the tree is a line

The last row is chapter 68, and it is the reason the AVL tree exists.

What searching gives you beyond "is it there"

Three operations fall out of the same walk, and they are why a tree beats a hash table when order matters:

Minimum. Walk left until you cannot. Maximum. Walk right until you cannot. Both cost the height.

Floor and ceiling. The largest value not greater than x, and the smallest not less than x. The same walk, remembering the last turn.

A hash table can do none of these without inspecting every key.

Quick revision

  • Search compares with the node and goes left or right, discarding an entire subtree each time.
  • Four lines, iterative, no recursion needed.
  • An unsuccessful search also costs the height, not n.
  • Measured: 8,000 random values gave a height of 31 and a worst search of 27 comparisons, against 8,000

for a linear search.

  • A randomly built tree has height about 1.39 log2(n), not the ideal log2(n), but close enough.
  • Best case 1 comparison, worst case the height, which is n for a degenerate tree.
  • Minimum, maximum, floor and ceiling are the same walk, and a hash table can do none of them.

Test yourself

1. Write the search algorithm. Start at the root. While the node is not null: if the target equals the node's value, it is found; if the target is smaller go left, otherwise go right. If the walk falls off the tree, it is absent.

munotes.in200

Searching a Binary Search Tree

2. What does each comparison achieve? It discards one of the two subtrees entirely, which is the invariant of chapter 64 being spent.

3. How much does an unsuccessful search cost? The height of the tree, not n. The walk stops as soon as it falls off the bottom.

4. Give the four cases for the number of comparisons. Best 1; average about 1.39 log2(n) for a randomly built tree; worst for a balanced tree the height, about log2(n); worst for a degenerate tree n.

5. In the measurement, what were the height and worst search for 8,000 values? Height 31 and a worst search of 27 comparisons, against up to 8,000 for a linear search.

6. Name two operations a search tree supports that a hash table cannot. Finding the minimum or maximum, and finding the floor or ceiling of a value. A hash table has no order, so it would have to inspect every key.

Contents This chapter on its own page

munotes.in201

Chapter Sixty-Six

Inserting Into a Binary Search Tree

Syllabus topic Module 2, "Trees: Binary Search Tree"

In one line

Insertion searches for the value, and where the search falls off the tree is exactly where the new node is attached as a leaf.

The algorithm

insert(node, value):

if node is null: return a new node holding value

if value < node.data: node.left = insert(node.left, value)

if value > node.data: node.right = insert(node.right, value)

(equal: do nothing, or handle duplicates by the chosen convention)

return node

Two things about this are worth stating.

A new value always becomes a leaf. Nothing in the existing tree moves. The search walks down until it reaches a null, and that null becomes the new node.

The returned node is assigned back. node.left = insert(node.left, value) is what makes the new leaf actually attach. Writing insert(node.left, value) alone compiles, runs, and inserts nothing, because the new node is created and immediately discarded. That is the single commonest error in this chapter.

Run, with the tree after each insertion

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None


def insert(root, value):
    if root is None:
        return BNode(value)
    if value < root.data:
        root.left = insert(root.left, value)       # the assignment matters
    elif value > root.data:
        root.right = insert(root.right, value)
    return root


def shape(node):
    if node is None:
        return "."
    if node.left is None and node.right is None:
        return str(node.data)
    return "%s(%s, %s)" % (node.data, shape(node.left), shape(node.right))


def height(node):
    return -1 if node is None else 1 + max(height(node.left), height(node.right))


def inorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder(node.left, out)
        out.append(node.data)
        inorder(node.right, out)
    return out


root = None
for value in (50, 30, 70, 20, 40, 60, 80):
    root = insert(root, value)
    print("insert %-3d -> %-36s height %d" % (value, shape(root), height(root)))

print()
print("inorder:", inorder(root), "which is sorted:", inorder(root) == sorted(inorder(root)))

print()
print("inserting a duplicate does nothing:")
before = shape(root)
root = insert(root, 40)
print("   shape unchanged:", shape(root) == before)
print("   size unchanged :", len(inorder(root)) == 7)
insert 50  -> 50                                   height 0
insert 30  -> 50(30, .)                            height 1
insert 70  -> 50(30, 70)                           height 1
insert 20  -> 50(30(20, .), 70)                    height 2
insert 40  -> 50(30(20, 40), 70)                   height 2
insert 60  -> 50(30(20, 40), 70(60, .))            height 2
insert 80  -> 50(30(20, 40), 70(60, 80))           height 2

inorder: [20, 30, 40, 50, 60, 70, 80] which is sorted: True

inserting a duplicate does nothing:
   shape unchanged: True
   size unchanged : True

Seven values, a perfectly balanced tree of height 2, and the inorder is sorted. Every intermediate shape was printed by the program after the insertion that produced it.

The error that inserts nothing

The missing assignment deserves a demonstration, because the broken version does not crash.

munotes.in202

Inserting Into a Binary Search Tree

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None


def insert_correct(root, value):
    if root is None:
        return BNode(value)
    if value < root.data:
        root.left = insert_correct(root.left, value)
    elif value > root.data:
        root.right = insert_correct(root.right, value)
    return root


def insert_broken(root, value):
    """The new node is created and then thrown away."""
    if root is None:
        return BNode(value)
    if value < root.data:
        insert_broken(root.left, value)            # no assignment
    elif value > root.data:
        insert_broken(root.right, value)
    return root


def inorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder(node.left, out)
        out.append(node.data)
        inorder(node.right, out)
    return out


for name, insert in (("correct", insert_correct), ("broken ", insert_broken)):
    root = None
    for value in (50, 30, 70, 20, 40):
        root = insert(root, value)
    print("%s: tree holds %s" % (name, inorder(root)))

print()
print("the broken version raised no error and kept only the root's first child.")
correct: tree holds [20, 30, 40, 50, 70]
broken : tree holds [50]

the broken version raised no error and kept only the root's first child.

The broken version holds one of the five values and reported nothing wrong. Only the first insertion worked, because at the top level the caller does assign the return value (root = insert(root, value)). Every insertion after that descended into the tree, created its node, and threw it away.

What the insertion order does to the shape

The same values, inserted in different orders, give completely different trees. This is the setup for chapter 68.

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None


def insert(root, value):
    if root is None:
        return BNode(value)
    if value < root.data:
        root.left = insert(root.left, value)
    elif value > root.data:
        root.right = insert(root.right, value)
    return root


def build(values):
    root = None
    for value in values:
        root = insert(root, value)
    return root


def height(node):
    return -1 if node is None else 1 + max(height(node.left), height(node.right))


def shape(node):
    if node is None:
        return "."
    if node.left is None and node.right is None:
        return str(node.data)
    return "%s(%s, %s)" % (node.data, shape(node.left), shape(node.right))


orders = [
    ("balanced order ", [50, 30, 70, 20, 40, 60, 80]),
    ("sorted order   ", [20, 30, 40, 50, 60, 70, 80]),
    ("reverse sorted ", [80, 70, 60, 50, 40, 30, 20]),
    ("almost sorted  ", [20, 30, 40, 50, 80, 70, 60]),
]
for name, values in orders:
    tree = build(values)
    print("%s height %d  %s" % (name, height(tree), shape(tree)))
balanced order  height 2  50(30(20, 40), 70(60, 80))
sorted order    height 6  20(., 30(., 40(., 50(., 60(., 70(., 80))))))
reverse sorted  height 6  80(70(60(50(40(30(20, .), .), .), .), .), .)
almost sorted   height 6  20(., 30(., 40(., 50(., 80(70(60, .), .)))))
munotes.in203

Inserting Into a Binary Search Tree

The same seven values. Height 2 in one order and height 6 in another, which is a line.

Sorted input is not an unusual case: data often arrives sorted, from a database, a file, or a previous sort. The commonest real input produces the worst possible tree, and that is chapter 68.

The cost

Cost
Insertion into a balanced treeO(log n)
Insertion into a degenerate treeO(n)
Nodes movedzero, always

The last row is worth noticing against the array of chapter 12, which moved everything after the insertion point. A tree pays with pointer-following rather than with movement.

Quick revision

  • Insertion searches for the value; where the search falls off the tree is where the new leaf goes.
  • A new value always becomes a leaf and nothing existing moves.
  • node.left = insert(node.left, value): the assignment is what attaches the node. Omitting it inserts

nothing and raises no error.

  • A duplicate is ignored under this book's convention.
  • Insertion order decides the shape: the same seven values gave height 2 in one order and height 6 in

sorted order, which is a line.

  • Sorted input is common, which makes the worst case common.
  • O(log n) balanced, O(n) degenerate, and zero nodes moved either way.

Test yourself

1. Where does a newly inserted value end up? As a leaf, at the position where the search for it falls off the tree.

2. Why must the recursive call be assigned back to the child pointer? Because the function returns the subtree's new root, which for an empty subtree is the newly created node. Without the assignment the new node is discarded and nothing is inserted.

3. What happens if the assignment is omitted, and why is it dangerous? Only the first insertion takes effect, because only the top level caller assigns the return value. No error is raised: in the run, five insertions left a tree holding one value, silently.

4. The same seven values gave heights 2 and 6 in different orders. Which order gave 6 and why? Sorted order, ascending or descending. Each new value is larger (or smaller) than everything present, so it goes to the same side every time and the tree becomes a line.

5. Why is that the dangerous case in practice? Because sorted input is common: data often arrives already in order from a file, a database or an earlier sort. The commonest input produces the worst tree.

6. How many existing nodes move during an insertion? None, ever. The cost is in walking to the position, not in moving anything, which is the difference from an array.

Contents This chapter on its own page

munotes.in204

Chapter Sixty-Seven

Deleting From a Binary Search Tree: The Three Cases

Syllabus topic Module 2, "Trees: Binary Search Tree"

In one line

Deleting a leaf removes it, deleting a node with one child promotes that child, and deleting a node with two children replaces its value with its inorder successor and then deletes the successor.

Why deletion is harder than insertion

Insertion always adds a leaf, so nothing has to be rearranged. Deletion can remove a node from the middle, and whatever hung below it must be reattached without breaking the invariant of chapter 64.

There are exactly three situations, decided by how many children the doomed node has.

Case 1: a leaf

No children. Remove it: set the parent's pointer to null. Nothing else changes.

Case 2: one child

Replace the node with its only child. The whole subtree moves up one level.

This is safe because the subtree's values were already all on the correct side of the node's parent: if the doomed node was in its parent's left subtree, everything below it was too.

Case 3: two children

This is the real case. The node cannot simply be removed, because its parent has only one pointer and there are two subtrees to reattach.

The answer: do not remove the node. Replace its value.

Which value can take its place without breaking the invariant? Exactly two will do:

The inorder successor: the smallest value in the right subtree. The inorder predecessor: the largest value in the left subtree.

Either is correct. This book uses the successor, and so do most textbooks, so an answer should name it.

The successor is the smallest in the right subtree, so it is greater than everything on the left and less than everything else on the right. Putting it at the node keeps the invariant exactly.

Then delete the successor from the right subtree, and that deletion is easy: the smallest node of a subtree has no left child, so it is case 1 or case 2, never case 3.

Built and run, all three cases traced

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None


def insert(root, value):
    if root is None:
        return BNode(value)
    if value < root.data:
        root.left = insert(root.left, value)
    elif value > root.data:
        root.right = insert(root.right, value)
    return root


def build(values):
    root = None
    for value in values:
        root = insert(root, value)
    return root


def minimum(node):
    """The smallest value in a subtree: walk left until you cannot."""
    while node.left is not None:
        node = node.left
    return node


def delete(root, value):
    if root is None:
        return None
    if value < root.data:
        root.left = delete(root.left, value)
    elif value > root.data:
        root.right = delete(root.right, value)
    else:
        # case 1 and case 2 together: no child, or one child
        if root.left is None:
            return root.right
        if root.right is None:
            return root.left
        # case 3: two children
        successor = minimum(root.right)
        root.data = successor.data
        root.right = delete(root.right, successor.data)
    return root


def shape(node):
    if node is None:
        return "."
    if node.left is None and node.right is None:
        return str(node.data)
    return "%s(%s, %s)" % (node.data, shape(node.left), shape(node.right))


def inorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder(node.left, out)
        out.append(node.data)
        inorder(node.right, out)
    return out


def is_bst(node, low=None, high=None):
    if node is None:
        return True
    if low is not None and node.data <= low:
        return False
    if high is not None and node.data >= high:
        return False
    return (is_bst(node.left, low, node.data)
            and is_bst(node.right, node.data, high))


values = [50, 30, 70, 20, 40, 60, 80, 35]
print("start:", shape(build(values)))
print()

for value, case in ((20, "case 1, a leaf"),
                    (60, "case 1, a leaf"),
                    (40, "case 2, one child"),
                    (30, "case 3, two children"),
                    (50, "case 3, the root")):
    tree = build(values)
    tree = delete(tree, value)
    walk = inorder(tree)
    print("delete %-3d (%-21s) -> %-38s" % (value, case, shape(tree)))
    print("      inorder %s" % (walk,))
    print("      still a BST: %-5s  still sorted: %-5s  value gone: %s"
          % (is_bst(tree), walk == sorted(walk), value not in walk))
munotes.in205

Deleting From a Binary Search Tree: The Three Cases

start: 50(30(20, 40(35, .)), 70(60, 80))

delete 20  (case 1, a leaf       ) -> 50(30(., 40(35, .)), 70(60, 80))
      inorder [30, 35, 40, 50, 60, 70, 80]
      still a BST: True   still sorted: True   value gone: True
delete 60  (case 1, a leaf       ) -> 50(30(20, 40(35, .)), 70(., 80))
      inorder [20, 30, 35, 40, 50, 70, 80]
      still a BST: True   still sorted: True   value gone: True
delete 40  (case 2, one child    ) -> 50(30(20, 35), 70(60, 80))
      inorder [20, 30, 35, 50, 60, 70, 80]
      still a BST: True   still sorted: True   value gone: True
delete 30  (case 3, two children ) -> 50(35(20, 40), 70(60, 80))
      inorder [20, 35, 40, 50, 60, 70, 80]
      still a BST: True   still sorted: True   value gone: True
delete 50  (case 3, the root     ) -> 60(30(20, 40(35, .)), 70(., 80))
      inorder [20, 30, 35, 40, 60, 70, 80]
      still a BST: True   still sorted: True   value gone: True

Five deletions, one of each case including the root, and after every one the tree is checked to still be a search tree, the inorder is checked to still be sorted, and the value is checked to be gone. That is three independent confirmations per deletion rather than a glance at the shape.

Follow the case 3 deletions:

Deleting 30 replaced it with 35, the smallest value in 30's right subtree (which was 40 with left child 35), and then removed 35 from there. The shape became 35(20, 40).

munotes.in206

Deleting From a Binary Search Tree: The Three Cases

Deleting the root 50 replaced it with 60, the smallest value in the right subtree 70(60, 80), and then removed 60 from there, leaving 70(., 80).

Why the successor is always easy to remove

The successor is the minimum of the right subtree, found by walking left until you cannot. By definition it has no left child, so removing it is case 1 or case 2 and never case 3.

That is why the recursion terminates and why deletion is not an endlessly nested problem.

The cost

Finding the node: the height. Finding the successor: at most the height. Deleting the successor: at most the height.

So deletion is O(h), the same as search and insertion, which is O(log n) balanced and O(n) degenerate.

The asymmetry nobody mentions

Always taking the successor makes trees lean left over time, because the right subtree is the one that loses a node each time. Alternating between successor and predecessor keeps a tree better balanced over many deletions, and it is a real technique, though not required by this syllabus.

Quick revision

  • Three cases, decided by the number of children of the node being deleted.
  • Leaf: remove it. One child: promote the child. Two children: replace the value with the inorder

successor, then delete the successor from the right subtree.

  • The successor is the minimum of the right subtree; the predecessor, the maximum of the left subtree,

would do equally well.

  • The successor has no left child, so removing it is always case 1 or case 2, which is why the recursion

ends.

  • Checked after every deletion: still a search tree, still sorted, value gone.
  • Deletion is O(h), the same as search and insertion.
  • Always using the successor makes trees lean left over many deletions; alternating avoids it.

Test yourself

1. Give the three cases and what each does. A leaf is removed. A node with one child is replaced by that child. A node with two children has its value replaced by its inorder successor, and the successor is then deleted from the right subtree.

2. Which two values may replace a node with two children, and why either? The inorder successor, the smallest in the right subtree, or the inorder predecessor, the largest in the left. Either sits exactly between the two subtrees, so the invariant is preserved.

3. Why is deleting the successor always easy? Because it is the minimum of a subtree, found by walking left until you cannot, so it has no left child and its own deletion is case 1 or case 2.

4. Delete 30 from 50(30(20, 40(35, .)), 70(60, 80)) and give the result. 30 has two children, so it is replaced by 35, the smallest in its right subtree, and 35 is removed from there, giving 50(35(20, 40), 70(60, 80)).

munotes.in207

Deleting From a Binary Search Tree: The Three Cases

5. What three things should be checked after a deletion? That the tree is still a binary search tree, that its inorder is still sorted, and that the value is actually gone.

6. What is the cost of deletion, and what is the long-term effect of always choosing the successor? O(h), the same as search and insertion. Always choosing the successor removes from the right each time, so trees drift towards leaning left; alternating with the predecessor avoids it.

Contents This chapter on its own page

munotes.in208

Chapter Sixty-Eight

Why a Binary Search Tree Degenerates

Syllabus topic Module 2, "Trees: Balanced BST"

In one line

A binary search tree built from sorted input becomes a line, so every O(log n) operation becomes O(n), and sorted input is the commonest input there is.

The failure

Insert 1, 2, 3, 4, 5 into an empty binary search tree.

1 1 1 1 1

. . . .

2 2 2 2

. . .

3 3 3

. .

4 4

.

5

Every value is larger than everything already present, so every insertion goes right, and the tree is a linked list that happens to be drawn vertically.

Searching it is a linear walk. The tree has gained nothing at all over chapter 13's linked list, and it has lost something: it carries two child pointers per node instead of one, and half of them are null.

Measured

import random


class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None


def insert(root, value):
    if root is None:
        return BNode(value)
    if value < root.data:
        root.left = insert(root.left, value)
    elif value > root.data:
        root.right = insert(root.right, value)
    return root


def build_iterative(values):
    """Built without recursion, so a 1,000 node line cannot overflow the stack."""
    root = None
    for value in values:
        if root is None:
            root = BNode(value)
            continue
        walk = root
        while True:
            if value < walk.data:
                if walk.left is None:
                    walk.left = BNode(value)
                    break
                walk = walk.left
            elif value > walk.data:
                if walk.right is None:
                    walk.right = BNode(value)
                    break
                walk = walk.right
            else:
                break
    return root


def height_iterative(root):
    """Height without recursion, for the same reason."""
    if root is None:
        return -1
    best, stack = 0, [(root, 0)]
    while stack:
        node, depth = stack.pop()
        best = max(best, depth)
        if node.left is not None:
            stack.append((node.left, depth + 1))
        if node.right is not None:
            stack.append((node.right, depth + 1))
    return best


def worst_search(root):
    """Comparisons for the deepest value: the height plus one."""
    return height_iterative(root) + 1


import math

n = 1000
sorted_values = list(range(n))
shuffled = list(range(n))
random.seed(9)
random.shuffle(shuffled)

cases = [
    ("sorted ascending ", sorted_values),
    ("sorted descending", list(reversed(sorted_values))),
    ("shuffled         ", shuffled),
]

print("%-18s %8s %18s %s" % ("insertion order", "height", "worst search", "ideal"))
ideal = math.ceil(math.log2(n + 1)) - 1
for name, values in cases:
    tree = build_iterative(values)
    print("%-18s %8d %18d %d"
          % (name, height_iterative(tree), worst_search(tree), ideal))

print()
print("the same 1,000 values. sorted input gives a height of %d; the ideal is %d."
      % (height_iterative(build_iterative(sorted_values)), ideal))
insertion order      height       worst search ideal
sorted ascending        999               1000 9
sorted descending       999               1000 9
shuffled                 19                 20 9

the same 1,000 values. sorted input gives a height of 999; the ideal is 9.

One thousand values. Sorted input gives a height of 999 and a worst search of 1,000 comparisons. Shuffled input gives a height of 19. The ideal is 9.

munotes.in209

Why a Binary Search Tree Degenerates

The structure did nothing wrong. Every insertion followed the rules of chapter 66 exactly.

Why this is not a rare case

The temptation is to treat sorted input as unlucky. It is not; it is normal.

Data arrives sorted from a database, because the query had an ORDER BY. Data arrives sorted from a file that was written in order. Data arrives sorted because somebody sorted it, often to make it easier to work with. Timestamps, invoice numbers, roll numbers and dates are naturally increasing.

So the worst case for a binary search tree is produced by the most ordinary thing that can happen to data, and a structure whose worst case is the common case is not usable without a fix.

Three fixes, and which one this paper takes

1. Shuffle the input. Works, and is useless when the data arrives over time rather than all at once.

2. Rebuild the tree periodically into a balanced shape. Works, costs O(n) each time, and leaves the tree unbalanced in between.

3. Keep the tree balanced as it is built, by rearranging after each insertion. This costs a little on every insertion and guarantees the height for ever.

The third is the real answer, and it is what the AVL tree of chapters 71 to 74 does. MU's syllabus names "Balanced BST" and "AVL Trees" in exactly that order, which is the order of this argument.

What "balanced" has to mean

Note what the fix cannot be. It cannot require the tree to be perfect, because chapter 56 showed perfect trees exist only at sizes 1, 3, 7, 15 and so on, and a structure cannot refuse to hold 12 items.

So balance has to be a condition that is achievable at every size and still forces the height to be O(log n). Defining that condition precisely is the next chapter.

Quick revision

  • Sorted input makes every insertion go the same way, so the tree becomes a line.
  • Measured: 1,000 sorted values give height 999 and a worst search of 1,000; shuffled gives height 19;

the ideal is 9.

  • A degenerate tree is worse than a linked list: the same O(n) search, plus a wasted second pointer per

node.

  • Sorted input is normal, not unlucky: database results, files, timestamps, roll numbers and anything

somebody sorted.

  • Three fixes: shuffle the input, rebuild periodically, or keep the tree balanced as it is built.
  • The third is the real answer and is the AVL tree.
  • Balance cannot mean perfect, since perfect trees exist only at sizes 1, 3, 7, 15; it must be achievable

at every size.

munotes.in210

Why a Binary Search Tree Degenerates

Test yourself

1. What shape does a binary search tree take when values are inserted in sorted order, and why? A line. Every new value is larger than everything already present, so every insertion goes right, and each node has one child.

2. Give the measured heights for 1,000 values in sorted and shuffled order, and the ideal. 999 sorted, 19 shuffled, and the ideal is 9.

3. Why is a degenerate search tree worse than a linked list? It has the same O(n) search cost and additionally carries two child pointers per node, half of them null.

4. Why is sorted input a normal case rather than an unlucky one? Because data commonly arrives in order: from a database query with an ORDER BY, from a file written in order, or as timestamps, invoice numbers and roll numbers, which increase naturally.

5. Name the three possible fixes and say which the syllabus takes. Shuffle the input; rebuild the tree periodically; or keep it balanced as it is built. The syllabus takes the third, naming Balanced BST and then AVL Trees.

6. Why can "balanced" not be defined as "perfect"? Because a perfect binary tree has exactly 2 to the power (h+1) minus 1 nodes, so perfect trees exist only at sizes 1, 3, 7, 15 and so on, and a structure must be able to hold any number of items.

Contents This chapter on its own page

munotes.in211

Chapter Sixty-Nine

What Balance Means

Syllabus topic Module 2, "Trees: Balanced BST"

In one line

Balanced means the height stays O(log n), and the definition that achieves it has to be a local condition every node can check, not a global one, because a global condition cannot be restored cheaply.

What we need the definition to guarantee

One thing only: the height of a tree with n nodes is O(log n).

Everything else follows. Search, insertion and deletion all cost the height, so bounding the height bounds all three.

Three definitions that do not work

1. Perfect: every level completely filled. Chapter 56 ruled this out. Perfect trees exist only at sizes 1, 3, 7, 15 and so on, so a tree of 12 items could not be balanced at all.

2. Complete: every level full except the last, filled from the left. This does give the minimum possible height. But it is a condition about the shape of the whole tree, and a single insertion in the wrong place can require moving many nodes to restore it. A heap can maintain it because a heap may put a value anywhere; a search tree cannot, because the value's position is fixed by its key.

3. The two subtrees of the root have equal height. Too weak. It constrains the root and nothing else, so each subtree can be a line of 500 nodes and the condition still holds.

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


def height(node):
    if node is None:
        return -1
    best, stack = -1, [(node, 0)]
    while stack:
        current, depth = stack.pop()
        best = max(best, depth)
        if current.left is not None:
            stack.append((current.left, depth + 1))
        if current.right is not None:
            stack.append((current.right, depth + 1))
    return best


def size(node):
    if node is None:
        return 0
    n, stack = 0, [node]
    while stack:
        current = stack.pop()
        n += 1
        if current.left is not None:
            stack.append(current.left)
        if current.right is not None:
            stack.append(current.right)
    return n


def chain_left(values):
    """A line going left."""
    node = None
    for value in values:
        node = BNode(value, node, None)
    return node


def chain_right(values):
    node = None
    for value in reversed(values):
        node = BNode(value, None, node)
    return node


# Root with two subtrees of EQUAL height, each of which is a line.
left_line = chain_left(list(range(1, 51)))          # 50 nodes, a line
right_line = chain_right(list(range(52, 102)))      # 50 nodes, a line
root = BNode(51, left_line, right_line)

print("the root's two subtrees have equal height:",
      height(root.left) == height(root.right))
print("   left subtree height :", height(root.left))
print("   right subtree height:", height(root.right))
print("   so 'subtrees of the root have equal height' is SATISFIED")
print()
print("but the tree holds", size(root), "nodes with a height of", height(root))
print("   the ideal height for that many nodes is about 6")
print("   so the condition guarantees nothing")
munotes.in212

What Balance Means

the root's two subtrees have equal height: True
   left subtree height : 49
   right subtree height: 49
   so 'subtrees of the root have equal height' is SATISFIED

but the tree holds 101 nodes with a height of 50
   the ideal height for that many nodes is about 6
   so the condition guarantees nothing

A tree of 101 nodes with a height of 50, satisfying the third definition perfectly. The condition has to apply at every node, not just the root.

The definition that works

A tree is height balanced when, at every node, the heights of its two subtrees differ by at most 1.

That is the AVL condition, named after Adelson-Velsky and Landis who published it in 1962.

Three things make it the right choice, and an examination answer that gives them is a complete answer:

It is local. Each node checks only its own two subtree heights. There is no global shape to compute.

It is achievable at every size. Unlike perfect, there is a height balanced tree with any number of nodes.

It still forces O(log n). This is the part that is not obvious, and it is worth seeing.

Why "differ by at most 1" is enough

Ask the opposite question: what is the fewest nodes a height balanced tree of height h can have? If even the sparsest such tree has a lot of nodes, then a tree with n nodes cannot be very tall.

Call that minimum N(h).

  • N(0) = 1: a single node.
  • N(1) = 2: a root and one child.
  • For height h, the root's taller subtree has height h-1 and the other may have height h-2, since they

may differ by 1. So N(h) = 1 + N(h-1) + N(h-2).

That is the Fibonacci recurrence, and it grows exponentially. So a height balanced tree of height h has at least about 1.618 to the power h nodes, which turned around says the height is at most about 1.44 times log2(n).

import math


def minimum_nodes(h, memo={}):
    """Fewest nodes in a height balanced tree of height h."""
    if h < 0:
        return 0
    if h == 0:
        return 1
    if h not in memo:
        memo[h] = 1 + minimum_nodes(h - 1) + minimum_nodes(h - 2)
    return memo[h]


print("%7s %18s %22s" % ("height", "fewest nodes N(h)", "so n nodes give height"))
for h in (2, 5, 10, 15, 20, 25, 30):
    n = minimum_nodes(h)
    print("%7d %18d %22s" % (h, n, "at most %d for n = %d" % (h, n)))

print()
print("the ratio N(h) / N(h-1) settles at the golden ratio:")
for h in (10, 20, 30):
    print("   N(%d)/N(%d) = %.4f" % (h, h - 1,
                                     minimum_nodes(h) / minimum_nodes(h - 1)))
print("   the golden ratio is %.4f" % ((1 + 5 ** 0.5) / 2,))
print()
print("so height <= about 1.44 x log2(n), which is O(log n):")
for n in (1000, 10 ** 6, 10 ** 9):
    print("   n = %-12d worst AVL height about %d, ideal %d"
          % (n, math.floor(1.44 * math.log2(n)), math.ceil(math.log2(n + 1)) - 1))
munotes.in213

What Balance Means

 height  fewest nodes N(h) so n nodes give height
      2                  4    at most 2 for n = 4
      5                 20   at most 5 for n = 20
     10                232 at most 10 for n = 232
     15               2583 at most 15 for n = 2583
     20              28656 at most 20 for n = 28656
     25             317810 at most 25 for n = 317810
     30            3524577 at most 30 for n = 3524577

the ratio N(h) / N(h-1) settles at the golden ratio:
   N(10)/N(9) = 1.6224
   N(20)/N(19) = 1.6181
   N(30)/N(29) = 1.6180
   the golden ratio is 1.6180

so height <= about 1.44 x log2(n), which is O(log n):
   n = 1000         worst AVL height about 14, ideal 9
   n = 1000000      worst AVL height about 28, ideal 19
   n = 1000000000   worst AVL height about 43, ideal 29

Read the second column. A height balanced tree of height 30 has at least 3.5 million nodes. So a tree with fewer than that cannot be 30 tall, however badly it is built, as long as the condition holds at every node.

That is the guarantee, and it is why "differ by at most 1" is the right condition: strong enough to force O(log n), weak enough to be restorable by a local rearrangement, which is the rotation of chapter 72.

Other balanced trees, named

MU's syllabus requires AVL. For completeness, since examiners sometimes ask what else exists:

Red-black trees allow a looser condition (no path more than twice another) and so rebalance less often. They are what most standard libraries use. B-trees allow many children per node and are what databases and file systems use, because they match the way a disk reads blocks.

Both are outside this paper, but knowing the names and the one-line reason is worth a mark.

Quick revision

  • Balanced means the height stays O(log n); that bounds search, insertion and deletion together.
  • Perfect is impossible at most sizes; complete cannot be maintained by a search tree, because a key's

position is fixed; equal heights at the root alone is too weak and allows a tree of 101 nodes with height 50.

  • The AVL condition: at every node, the two subtree heights differ by at most 1.
  • It is local, achievable at any size, and still forces O(log n).
  • The fewest nodes in a height balanced tree of height h is N(h) = 1 + N(h-1) + N(h-2), the Fibonacci
munotes.in214

What Balance Means

recurrence, so N grows like the golden ratio to the power h.

  • A height balanced tree of height 30 has at least 3,524,577 nodes, so the height is at most about

1.44 log2(n).

  • Red-black trees and B-trees are the other standard answers, outside this paper.

Test yourself

1. What must a definition of "balanced" guarantee? That the height of a tree with n nodes is O(log n), which bounds search, insertion and deletion together.

2. Why is "perfect" not usable, and why is "complete" not usable for a search tree? Perfect trees exist only at sizes 1, 3, 7, 15 and so on. Complete is a condition on the whole shape, and a search tree cannot move a value to restore it because the value's position is fixed by its key.

3. Why is "the root's two subtrees have equal height" too weak? Because it constrains only the root. The chapter shows a tree of 101 nodes satisfying it with a height of 50, each subtree being a line of 50 nodes.

4. State the AVL condition. At every node, the heights of its two subtrees differ by at most 1.

5. Give the recurrence for the fewest nodes in a height balanced tree of height h, and what it implies. N(h) = 1 + N(h-1) + N(h-2), with N(0) = 1 and N(1) = 2. It is the Fibonacci recurrence, so N(h) grows like the golden ratio to the power h, which means the height is at most about 1.44 log2(n).

6. A height balanced tree has height 30. What is the least number of nodes it can hold? 3,524,577. So any smaller tree satisfying the condition must be shorter than 30.

Contents This chapter on its own page

munotes.in215

Chapter Seventy

Threaded Binary Trees

Syllabus topic Module 2, "Trees: Threaded Binary Trees"

In one line

A binary tree of n nodes has n + 1 null pointers doing nothing, and a threaded tree puts them to work by making each one point to the node that comes next in traversal order.

The wasted pointers

Count them. Each of n nodes has two pointers, so there are 2n pointers in all. Each node except the root is pointed at by exactly one of them, so n - 1 are in use.

pointers in total = 2n

pointers in use = n - 1

pointers that are null = 2n - (n - 1) = n + 1

More than half of a binary tree's pointers are null. That is not an accident of a particular tree; it is true of every binary tree.

class BNode:
    __slots__ = ("data", "left", "right")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right


def build(values):
    root = None
    for value in values:
        if root is None:
            root = BNode(value)
            continue
        walk = root
        while True:
            if value < walk.data:
                if walk.left is None:
                    walk.left = BNode(value)
                    break
                walk = walk.left
            else:
                if walk.right is None:
                    walk.right = BNode(value)
                    break
                walk = walk.right
    return root


def count(node):
    nodes, nulls, stack = 0, 0, [node]
    while stack:
        current = stack.pop()
        nodes += 1
        for child in (current.left, current.right):
            if child is None:
                nulls += 1
            else:
                stack.append(child)
    return nodes, nulls


print("%8s %8s %8s %12s %s" % ("values", "nodes", "nulls", "n + 1", "nulls are"))
for values in ([50, 30, 70], [50, 30, 70, 20, 40, 60, 80],
               list(range(1, 21)), list(range(1, 101))):
    tree = build(values)
    nodes, nulls = count(tree)
    print("%8d %8d %8d %12d %11s"
          % (len(values), nodes, nulls, nodes + 1,
             "%.0f%% of all" % (100 * nulls / (2 * nodes))))
  values    nodes    nulls        n + 1 nulls are
       3        3        4            4  67% of all
       7        7        8            8  57% of all
      20       20       21           21  52% of all
     100      100      101          101  50% of all

The null count is exactly n + 1 every time, and it is more than half of all pointers in every tree.

The idea

A threaded binary tree replaces those nulls with threads: pointers to the node that would come next (or previously) in an inorder traversal.

Right threading. A node with no right child has its right pointer point to its inorder successor. Left threading. A node with no left child has its left pointer point to its inorder predecessor. Double threading does both; a tree threaded one way only is singly threaded.

One difficulty appears immediately: how does the code tell a real child from a thread? They are both just pointers. The answer is a flag per pointer:

munotes.in216

Threaded Binary Trees

FieldMeaning
left_threadtrue when left is a thread rather than a child
right_threadtrue when right is a thread rather than a child

That is the cost: two booleans per node, which in C is two bits.

Built and run

class TNode:
    __slots__ = ("data", "left", "right", "left_thread", "right_thread")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None
        self.left_thread = True        # true means 'this is a thread, not a child'
        self.right_thread = True


def build_plain(values):
    """An ordinary BST first, as nested tuples, then threaded below."""
    class BNode:
        __slots__ = ("data", "left", "right")

        def __init__(self, data):
            self.data = data
            self.left = None
            self.right = None

    root = None
    for value in values:
        if root is None:
            root = BNode(value)
            continue
        walk = root
        while True:
            if value < walk.data:
                if walk.left is None:
                    walk.left = BNode(value)
                    break
                walk = walk.left
            else:
                if walk.right is None:
                    walk.right = BNode(value)
                    break
                walk = walk.right
    return root


def to_threaded(plain):
    """Copy a plain BST into threaded nodes, then set the threads by inorder."""
    mapping = {}

    def copy(node):
        if node is None:
            return None
        new = TNode(node.data)
        left, right = copy(node.left), copy(node.right)
        if left is not None:
            new.left, new.left_thread = left, False
        if right is not None:
            new.right, new.right_thread = right, False
        mapping[node.data] = new
        return new

    root = copy(plain)

    order = []

    def inorder(node):
        if node is None:
            return
        inorder(node.left)
        order.append(node.data)
        inorder(node.right)

    inorder(plain)

    for i, value in enumerate(order):
        node = mapping[value]
        if node.left_thread:
            node.left = mapping[order[i - 1]] if i > 0 else None
        if node.right_thread:
            node.right = mapping[order[i + 1]] if i + 1 < len(order) else None
    return root


def threaded_inorder(root):
    """Inorder with NO stack and NO recursion. This is the whole point."""
    if root is None:
        return []
    node = root
    while not node.left_thread:          # go to the leftmost node
        node = node.left
    out = []
    while node is not None:
        out.append(node.data)
        if node.right_thread:
            node = node.right            # follow the thread to the successor
        else:
            node = node.right            # go right, then all the way left
            while not node.left_thread:
                node = node.left
    return out


values = [50, 30, 70, 20, 40, 60, 80, 35]
plain = build_plain(values)
threaded = to_threaded(plain)

print("values inserted :", values)
print("threaded inorder:", threaded_inorder(threaded))
print("sorted          :", sorted(values))
print("they agree      :", threaded_inorder(threaded) == sorted(values))
print()
print("and that traversal used no stack and no recursion at all.")
print()
print("where the threads point:")
order = sorted(values)
mapping = {}


def collect(node, seen):
    if node is None or node.data in seen:
        return
    seen.add(node.data)
    mapping[node.data] = node
    if not node.left_thread:
        collect(node.left, seen)
    if not node.right_thread:
        collect(node.right, seen)


collect(threaded, set())
for value in order:
    node = mapping[value]
    left = ("thread to %s" % node.left.data) if node.left_thread and node.left else (
        "child %s" % node.left.data if node.left else "none")
    right = ("thread to %s" % node.right.data) if node.right_thread and node.right else (
        "child %s" % node.right.data if node.right else "none")
    print("   %-3d left: %-16s right: %s" % (value, left, right))
munotes.in217

Threaded Binary Trees

values inserted : [50, 30, 70, 20, 40, 60, 80, 35]
threaded inorder: [20, 30, 35, 40, 50, 60, 70, 80]
sorted          : [20, 30, 35, 40, 50, 60, 70, 80]
they agree      : True

and that traversal used no stack and no recursion at all.

where the threads point:
   20  left: none             right: thread to 30
   30  left: child 20         right: child 40
   35  left: thread to 30     right: thread to 40
   40  left: child 35         right: thread to 50
   50  left: child 30         right: child 70
   60  left: thread to 50     right: thread to 70
   70  left: child 60         right: child 80
   80  left: thread to 70     right: none

The traversal is correct and it used neither recursion nor an explicit stack. That is what the threads bought.

What it costs and what it buys

Plain binary treeThreaded binary tree
Inorder traversal memoryO(h), a stackO(1)
Finding a node's successorneeds a parent pointer or a stackfollow one thread
Memory per nodetwo pointerstwo pointers plus two flags
Insertion and deletionstraightforwardmust repair the threads

The gain is real: traversal in constant memory, and the successor of any node in one step, which chapter 67's deletion wanted.

The cost is real too: every insertion and deletion must now fix up the threads of the neighbouring nodes as well as the child pointers, which is more code and more to get wrong.

Where it is used, honestly

Threaded trees were important when memory was scarce and a recursion stack was expensive. Today they appear mainly in two places: embedded systems, where the stack really is tight, and examination papers, because they are a neat demonstration that a null pointer is a wasted resource.

That is worth saying plainly rather than pretending the technique is in daily use. The idea it teaches, look at what your structure is wasting, is the part that lasts.

Quick revision

  • A binary tree of n nodes has exactly n + 1 null pointers, which is more than half of all 2n pointers.
  • A threaded tree makes a null right pointer point to the inorder successor, and a null left pointer to

the inorder predecessor.

  • Right threaded, left threaded, or doubly threaded; a tree threaded one way is singly threaded.
  • A flag per pointer distinguishes a thread from a real child: two booleans per node.
  • The gain: inorder traversal in O(1) memory with no stack and no recursion, and a node's successor in
munotes.in218

Threaded Binary Trees

one step.

  • The cost: the flags, and the need to repair threads on every insertion and deletion.
  • Chiefly of use where the stack is tight, and as a demonstration that a null pointer is a wasted

resource.

Test yourself

1. How many null pointers does a binary tree of n nodes have, and prove it. Exactly n + 1. There are 2n pointers in all; n - 1 of them point at nodes, since every node but the root is pointed at once; so 2n minus (n - 1) is n + 1.

2. What does a thread point to? A null right pointer is made to point to the node's inorder successor, and a null left pointer to its inorder predecessor.

3. How does the code tell a thread from a real child? By a flag stored with each pointer, two booleans per node, saying whether that pointer is a thread.

4. What is the main gain from threading? Inorder traversal with no stack and no recursion, in O(1) memory, and finding any node's successor in one step.

5. What does threading cost? Two flags per node, and the obligation to repair the threads of neighbouring nodes on every insertion and deletion.

6. Where are threaded trees actually used? Chiefly where the recursion stack is scarce, such as embedded systems, and in examinations. Their lasting value is the idea that a null pointer is a wasted resource.

Contents This chapter on its own page

munotes.in219

Chapter Seventy-One

The AVL Tree and the Balance Factor

Syllabus topic Module 2, "Trees: AVL Trees"

In one line

An AVL tree is a binary search tree in which every node's two subtrees differ in height by at most 1, and the difference is stored at the node as its balance factor.

The definition

An AVL tree is a binary search tree satisfying the height balance condition of chapter 69 at every node.

The balance factor of a node is:

balance factor = height(left subtree) - height(right subtree)

The condition is then simply: every node's balance factor is -1, 0 or +1.

Balance factorMeans
+1the left subtree is one taller: left heavy
0the two subtrees are the same height: balanced
-1the right subtree is one taller: right heavy
+2 or -2the condition is violated and must be repaired

Some books define it the other way round, right minus left. Either is fine; say which you are using, because +2 and -2 select different rotations in chapter 72 and swapping them silently gives wrong answers.

This book uses left minus right, which is the commoner convention.

Why the height is stored

The balance factor needs the height of both subtrees, and computing a height is O(n) by chapter 57. Doing that at every node on every insertion would make insertion O(n log n), which defeats the purpose.

So each node stores its own height, maintained as the tree changes:

height(node) = 1 + max(height(node.left), height(node.right))

Updating it costs O(1) per node on the path back up from an insertion, so the whole insertion stays O(log n). This is chapter 18's habit again: remember what would otherwise be recomputed, and maintain it.

Built, with the factors shown and the condition checked

class ANode:
    __slots__ = ("data", "left", "right", "height")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None
        self.height = 0            # a leaf has height 0


def height(node):
    return -1 if node is None else node.height


def update_height(node):
    node.height = 1 + max(height(node.left), height(node.right))


def balance_factor(node):
    return 0 if node is None else height(node.left) - height(node.right)


def plain_insert(root, value):
    """Insert WITHOUT rebalancing, so the violation can be seen."""
    if root is None:
        return ANode(value)
    if value < root.data:
        root.left = plain_insert(root.left, value)
    elif value > root.data:
        root.right = plain_insert(root.right, value)
    update_height(root)
    return root


def report(node, prefix=""):
    """Every node with its height and balance factor."""
    if node is None:
        return []
    out = []
    out.extend(report(node.left, prefix))
    flag = "" if abs(balance_factor(node)) <= 1 else "   VIOLATION"
    out.append("   %-4s height %-3d balance factor %+d%s"
               % (node.data, node.height, balance_factor(node), flag))
    out.extend(report(node.right, prefix))
    return out


def is_avl(node):
    if node is None:
        return True
    if abs(balance_factor(node)) > 1:
        return False
    return is_avl(node.left) and is_avl(node.right)


print("a balanced tree:")
tree = None
for value in (50, 30, 70):
    tree = plain_insert(tree, value)
for line in report(tree):
    print(line)
print("   AVL condition holds:", is_avl(tree))

print()
print("now insert 20 and 10, in that order, with NO rebalancing:")
for value in (20, 10):
    tree = plain_insert(tree, value)
for line in report(tree):
    print(line)
print("   AVL condition holds:", is_avl(tree))
munotes.in220

The AVL Tree and the Balance Factor

a balanced tree:
   30   height 0   balance factor +0
   50   height 1   balance factor +0
   70   height 0   balance factor +0
   AVL condition holds: True

now insert 20 and 10, in that order, with NO rebalancing:
   10   height 0   balance factor +0
   20   height 1   balance factor +1
   30   height 2   balance factor +2   VIOLATION
   50   height 3   balance factor +2   VIOLATION
   70   height 0   balance factor +0
   AVL condition holds: False

Two insertions broke it, in two places. Node 30 has a left subtree of height 1 and an empty right subtree of height -1, giving +2; and the root 50 then has a left subtree of height 2 against an empty right subtree, giving +2 as well.

Notice which one matters. Walking back up from the newly inserted node 10, the first node found out of balance is 30, the deeper one. Repairing there fixes the ancestors too, because it restores the height that 30's subtree contributes to them. Chapter 73 depends on exactly that: rebalance at the lowest violating node and the rest of the path heals itself.

What the height balance guarantees, restated

From chapter 69: a height balanced tree with n nodes has height at most about 1.44 log2(n).

So an AVL tree with a million nodes has a height of at most 28, against the ideal 19 and the degenerate 999,999. Every operation on it is therefore O(log n), guaranteed, whatever order the values arrive in.

That guarantee is the entire point, and it is what the next three chapters pay for.

The cost of the guarantee

Plain BSTAVL tree
SearchO(h), h unboundedO(log n), guaranteed
InsertO(h)O(log n), plus at most one rotation
DeleteO(h)O(log n), plus up to O(log n) rotations
Memory per nodetwo pointerstwo pointers plus a height
Codeshortconsiderably longer

The last row is honest and worth saying: an AVL tree is several times the code of a plain binary search tree, and that is the real reason plain trees are still used when the input is known to be random.

Quick revision

  • An AVL tree is a binary search tree where every node's subtree heights differ by at most 1.
  • Balance factor = height(left) - height(right); it must be -1, 0 or +1.
  • +1 is left heavy, -1 is right heavy, 0 is balanced; +2 or -2 is a violation.
  • Some books use right minus left; state which convention you use, because it selects the rotation.
  • Each node stores its height, so the balance factor is O(1) to compute and insertion stays O(log n).
  • height(node) = 1 + max(height of children), with an empty subtree counted as -1.
  • The guarantee is height at most about 1.44 log2(n): a million nodes in at most 28 levels.
  • It costs a stored height per node and considerably more code.
munotes.in221

The AVL Tree and the Balance Factor

Test yourself

1. Define an AVL tree. A binary search tree in which, at every node, the heights of the two subtrees differ by at most 1.

2. Give the formula for the balance factor and the values it may take. Height of the left subtree minus height of the right subtree, under this book's convention. It must be -1, 0 or +1; anything else is a violation.

3. Why is the height stored in each node rather than computed? Because computing a height is O(n), so computing it at every node on an insertion would make insertion far more expensive. Stored, it is updated in O(1) per node on the way back up.

4. In the run, two insertions produced violations. At which nodes, and which one is repaired? At 30 and at the root 50. The repair happens at 30, the lowest violating node on the path back up from the newly inserted value, because fixing it restores the height its subtree contributes and the ancestors come back into balance by themselves.

5. What height does an AVL tree of a million nodes have at most? About 28, since the bound is roughly 1.44 times log2(n). The ideal is 19 and an unbalanced tree could be 999,999.

6. Give two costs of using an AVL tree instead of a plain binary search tree. A stored height in every node, and considerably more code, since insertion and deletion must detect and repair violations.

Contents This chapter on its own page

munotes.in222

Chapter Seventy-Two

The Four Rotations

Syllabus topic Module 2, "Trees: AVL Trees"

In one line

A rotation rearranges three nodes and one subtree to reduce the height on the heavy side, and it preserves the search order exactly, which is why it is the one repair an AVL tree is allowed.

The idea

A violation means one side is two taller than the other. A rotation lifts a node from the tall side into the parent's place, pushing the parent down the short side.

The crucial property, and the reason rotations are usable at all:

A rotation does not change the inorder traversal.

So the search property survives. The tree's shape changes; its meaning does not.

The two single rotations

Right rotation, for a left-left violation

The node is left heavy, and its left child is also left heavy. The left child comes up.

z y

/ . / .

y D becomes x z

/ . / . / .

x C A B C D

/ .

A B

Check the inorder of both: A x B y C z D, in both. That is the property.

Left rotation, for a right-right violation

The mirror image. The node is right heavy and its right child is also right heavy.

z y

/ . / .

A y becomes z x

/ . / . / .

B x A B C D

/ .

C D

Inorder of both: A z B y C x D.

The two double rotations

A single rotation does not fix the case where the heavy side's child leans the other way. That is worth seeing before the fix is given.

Left-right violation

The node is left heavy and its left child is right heavy. A single right rotation here just moves the problem to the other side.

The fix is two rotations: left rotate the child, then right rotate the node.

z z y

/ . / . / .

x D left on x y D right on z x z

/ . / . / . / .

A y x C A B C D

/ . / .

B C A B

Right-left violation

The mirror: the node is right heavy and its right child is left heavy. Right rotate the child, then left rotate the node.

The four cases, as a table to memorise

Balance factor of the nodeBalance factor of the childCaseFix
+2+1 or 0Left Leftone right rotation
+2-1Left Rightleft on the child, then right on the node
-2-1 or 0Right Rightone left rotation
-2+1Right Leftright on the child, then left on the node

The name says where the problem is, and the rotation goes the opposite way. A Left Left problem is fixed by a Right rotation. Getting that backwards is the standard error.

munotes.in223

The Four Rotations

All four, run, with the inorder checked

class ANode:
    __slots__ = ("data", "left", "right", "height")

    def __init__(self, data, left=None, right=None):
        self.data = data
        self.left = left
        self.right = right
        self.height = 0


def height(node):
    return -1 if node is None else node.height


def update(node):
    node.height = 1 + max(height(node.left), height(node.right))
    return node


def balance_factor(node):
    return 0 if node is None else height(node.left) - height(node.right)


def rotate_right(z):
    y = z.left
    z.left = y.right
    y.right = z
    update(z)
    update(y)
    return y


def rotate_left(z):
    y = z.right
    z.right = y.left
    y.left = z
    update(z)
    update(y)
    return y


def rebalance(node):
    """Apply whichever of the four cases is needed. Returns the new subtree root."""
    update(node)
    bf = balance_factor(node)
    if bf > 1:                                  # left heavy
        if balance_factor(node.left) < 0:       # left child right heavy: LR
            node.left = rotate_left(node.left)
        return rotate_right(node)
    if bf < -1:                                 # right heavy
        if balance_factor(node.right) > 0:      # right child left heavy: RL
            node.right = rotate_right(node.right)
        return rotate_left(node)
    return node


def shape(node):
    if node is None:
        return "."
    if node.left is None and node.right is None:
        return str(node.data)
    return "%s(%s, %s)" % (node.data, shape(node.left), shape(node.right))


def inorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder(node.left, out)
        out.append(node.data)
        inorder(node.right, out)
    return out


def leaf(v):
    return ANode(v)


def built(node):
    """Set heights bottom up on a hand built tree."""
    if node is None:
        return None
    built(node.left)
    built(node.right)
    update(node)
    return node


cases = []

# Left Left: 30 is left heavy, its left child 20 is left heavy
cases.append(("Left Left  ", built(ANode(30, ANode(20, leaf(10), None), None))))
# Right Right
cases.append(("Right Right", built(ANode(10, None, ANode(20, None, leaf(30))))))
# Left Right: 30 left heavy, its left child 10 is RIGHT heavy
cases.append(("Left Right ", built(ANode(30, ANode(10, None, leaf(20)), None))))
# Right Left: 10 right heavy, its right child 30 is LEFT heavy
cases.append(("Right Left ", built(ANode(10, None, ANode(30, leaf(20), None)))))

print("%-12s %-22s %-22s %s" % ("case", "before", "after", "inorder unchanged"))
for name, tree in cases:
    before_shape, before_inorder = shape(tree), inorder(tree)
    fixed = rebalance(tree)
    print("%-12s %-22s %-22s %s"
          % (name, before_shape, shape(fixed),
             inorder(fixed) == before_inorder))

print()
print("every case became the same balanced shape 20(10, 30),")
print("and in every case the inorder was unchanged, so the search property survived.")
case         before                 after                  inorder unchanged
Left Left    30(20(10, .), .)       20(10, 30)             True
Right Right  10(., 20(., 30))       20(10, 30)             True
Left Right   30(10(., 20), .)       20(10, 30)             True
Right Left   10(., 30(20, .))       20(10, 30)             True

every case became the same balanced shape 20(10, 30),
and in every case the inorder was unchanged, so the search property survived.
munotes.in224

The Four Rotations

Four different broken shapes, four different repairs, and all four produce the same balanced tree. In every case the inorder is unchanged, which is the proof that the search property survives.

Why a single rotation fails on the Left Right case

Worth seeing, because the table is otherwise just four lines to memorise.

Take 30(10(., 20), .), a Left Right violation. Apply a single right rotation: 10 comes up, 30 goes right, and 10's right child 20 becomes 30's left child.

30 10

/ right on 30 .

10 30

. /

20 20

The result is 10(., 30(20, .)), which is a Right Left violation: still height 2, still unbalanced, just leaning the other way. The single rotation moved the problem instead of fixing it.

The double rotation works because the first rotation turns the Left Right shape into a Left Left shape, which the second then fixes.

The cost

A rotation is O(1): three pointer assignments and two height updates, regardless of the size of the subtrees hanging below. That is what makes AVL insertion O(log n) overall: the walk down is the height, and the repair is constant.

Quick revision

  • A rotation lifts a node from the tall side into the parent's place and pushes the parent down the short

side.

  • It does not change the inorder traversal, which is why the search property survives.
  • Four cases: Left Left needs one right rotation; Right Right one left rotation; Left Right a left on the

child then a right on the node; Right Left a right on the child then a left on the node.

  • The name says where the problem is; the rotation goes the opposite way.
  • A single rotation on a Left Right case moves the problem to the other side rather than fixing it.
  • All four cases on three nodes produce the same balanced result.
  • A rotation is O(1): three pointer assignments and two height updates.

Test yourself

1. What property makes rotations safe to use on a search tree? A rotation does not change the inorder traversal, so the ordering of the values is preserved even though the shape changes.

2. Give the four cases and their fixes. Left Left: one right rotation. Right Right: one left rotation. Left Right: left rotate the child, then right rotate the node. Right Left: right rotate the child, then left rotate the node.

3. How do you tell which case applies? By the node's balance factor and its heavy child's. +2 with the left child not right heavy is Left Left; +2 with the left child right heavy is Left Right; and the mirrors for -2.

4. Why does a single rotation not fix a Left Right violation? Because it turns the shape into a Right Left violation of the same height: the problem moves to the other side rather than going away.

munotes.in225

The Four Rotations

5. What is the cost of a rotation, and why does that matter? O(1): three pointer assignments and two height updates, whatever hangs below. It is what keeps AVL insertion at O(log n), since the walk down costs the height and the repair costs nothing extra.

6. A node has balance factor -2 and its right child has +1. Which case is it and what is the fix? Right Left. Right rotate the right child, then left rotate the node.

Contents This chapter on its own page

munotes.in226

Chapter Seventy-Three

Insertion Into an AVL Tree

Syllabus topic Module 2, "Trees: AVL Trees"

In one line

AVL insertion is ordinary binary search tree insertion followed by updating heights and rebalancing on the way back up, and at most one rotation is ever needed.

The algorithm

insert(node, value):

if node is null: return a new node

insert into the left or right subtree as usual, assigning the result back

update this node's height

return rebalance(node)

Three lines more than chapter 66's insertion, and they are the last three. The recursion already walks back up through every ancestor as it returns, so that is where the heights are updated and the balance is checked.

At most one rotation is ever performed on an insertion. A single rotation at the lowest violating node restores the subtree to the height it had before the insertion, so every ancestor above it goes back into balance by itself. That is a fact worth stating in an answer and it is checked below.

The worked sequence

class ANode:
    __slots__ = ("data", "left", "right", "height")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None
        self.height = 0


def height(node):
    return -1 if node is None else node.height


def update(node):
    node.height = 1 + max(height(node.left), height(node.right))


def balance_factor(node):
    return 0 if node is None else height(node.left) - height(node.right)


def rotate_right(z):
    y = z.left
    z.left = y.right
    y.right = z
    update(z)
    update(y)
    return y


def rotate_left(z):
    y = z.right
    z.right = y.left
    y.left = z
    update(z)
    update(y)
    return y


ROTATIONS = []


def rebalance(node):
    update(node)
    bf = balance_factor(node)
    if bf > 1:
        if balance_factor(node.left) < 0:
            node.left = rotate_left(node.left)
            ROTATIONS.append("LR at %s" % node.data)
        else:
            ROTATIONS.append("LL at %s" % node.data)
        return rotate_right(node)
    if bf < -1:
        if balance_factor(node.right) > 0:
            node.right = rotate_right(node.right)
            ROTATIONS.append("RL at %s" % node.data)
        else:
            ROTATIONS.append("RR at %s" % node.data)
        return rotate_left(node)
    return node


def insert(node, value):
    if node is None:
        return ANode(value)
    if value < node.data:
        node.left = insert(node.left, value)
    elif value > node.data:
        node.right = insert(node.right, value)
    else:
        return node
    return rebalance(node)


def shape(node):
    if node is None:
        return "."
    if node.left is None and node.right is None:
        return str(node.data)
    return "%s(%s, %s)" % (node.data, shape(node.left), shape(node.right))


def inorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder(node.left, out)
        out.append(node.data)
        inorder(node.right, out)
    return out


def is_avl(node):
    if node is None:
        return True
    return (abs(balance_factor(node)) <= 1
            and is_avl(node.left) and is_avl(node.right))


root = None
for value in (10, 20, 30, 40, 50, 25):
    ROTATIONS.clear()
    root = insert(root, value)
    done = ", ".join(ROTATIONS) if ROTATIONS else "none"
    print("insert %-3d -> %-30s height %d  rotation: %s"
          % (value, shape(root), height(root), done))

print()
print("inorder    :", inorder(root))
print("sorted     :", inorder(root) == sorted(inorder(root)))
print("AVL holds  :", is_avl(root))
insert 10  -> 10                             height 0  rotation: none
insert 20  -> 10(., 20)                      height 1  rotation: none
insert 30  -> 20(10, 30)                     height 1  rotation: RR at 10
insert 40  -> 20(10, 30(., 40))              height 2  rotation: none
insert 50  -> 20(10, 40(30, 50))             height 2  rotation: RR at 30
insert 25  -> 30(20(10, 25), 40(., 50))      height 2  rotation: RL at 20

inorder    : [10, 20, 25, 30, 40, 50]
sorted     : True
AVL holds  : True
munotes.in227

Insertion Into an AVL Tree

Six insertions, three rotations, and the tree never exceeds height 2 where a plain tree would have reached height 4 on this input.

Follow the last one. Inserting 25 made node 20 right heavy while its right child 40 was left heavy: a Right Left case, fixed by right rotating 40 and then left rotating 20. The root changed from 20 to 30.

The comparison with chapter 68

Chapter 68 measured 1,000 sorted values producing a height of 999 in a plain binary search tree. Here is the same input into an AVL tree.

import math


class ANode:
    __slots__ = ("data", "left", "right", "height")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None
        self.height = 0


def height(node):
    return -1 if node is None else node.height


def update(node):
    node.height = 1 + max(height(node.left), height(node.right))


def bf(node):
    return 0 if node is None else height(node.left) - height(node.right)


def rotate_right(z):
    y = z.left
    z.left = y.right
    y.right = z
    update(z)
    update(y)
    return y


def rotate_left(z):
    y = z.right
    z.right = y.left
    y.left = z
    update(z)
    update(y)
    return y


def rebalance(node):
    update(node)
    if bf(node) > 1:
        if bf(node.left) < 0:
            node.left = rotate_left(node.left)
        return rotate_right(node)
    if bf(node) < -1:
        if bf(node.right) > 0:
            node.right = rotate_right(node.right)
        return rotate_left(node)
    return node


def insert_iterative(root, value):
    """Iterative so 1,000 sorted insertions cannot overflow the call stack."""
    if root is None:
        return ANode(value)
    path = []
    node = root
    while node is not None:
        path.append(node)
        node = node.left if value < node.data else node.right
    parent = path[-1]
    if value < parent.data:
        parent.left = ANode(value)
    else:
        parent.right = ANode(value)
    for node in reversed(path):
        fixed = rebalance(node)
        if fixed is not node:
            index = path.index(node)
            if index == 0:
                root = fixed
            else:
                above = path[index - 1]
                if above.left is node:
                    above.left = fixed
                else:
                    above.right = fixed
    return root


def is_avl(node):
    stack, ok = [node], True
    while stack:
        current = stack.pop()
        if current is None:
            continue
        if abs(bf(current)) > 1:
            ok = False
        stack.append(current.left)
        stack.append(current.right)
    return ok


for n in (100, 500, 1000):
    root = None
    for value in range(n):                     # strictly sorted input
        root = insert_iterative(root, value)
    ideal = math.ceil(math.log2(n + 1)) - 1
    print("%5d sorted values: AVL height %3d, ideal %2d, bound 1.44 log2(n) = %d, "
          "AVL condition holds: %s"
          % (n, height(root), ideal, math.floor(1.44 * math.log2(n)), is_avl(root)))

print()
print("chapter 68 measured the same 1,000 sorted values into a plain BST: height 999.")
munotes.in228

Insertion Into an AVL Tree

  100 sorted values: AVL height   6, ideal  6, bound 1.44 log2(n) = 9, AVL condition holds: True
  500 sorted values: AVL height   8, ideal  8, bound 1.44 log2(n) = 12, AVL condition holds: True
 1000 sorted values: AVL height   9, ideal  9, bound 1.44 log2(n) = 14, AVL condition holds: True

chapter 68 measured the same 1,000 sorted values into a plain BST: height 999.

Height 9 instead of 999, on the input that destroys a plain binary search tree, and it exactly hits the ideal. That is the whole value of the structure in one line.

Where the rotation happens

The rebalancing is done at the lowest node that has gone out of balance, found on the way back up. Three things follow:

The violating node is found going up, not down. The insertion descends to a leaf first, and the heights that change are those of the new leaf's ancestors.

Only one rotation is needed. Restoring the lowest violating subtree returns it to its pre-insertion height, so nothing above it is out of balance any more.

Heights must still be updated all the way up, even where no rotation happens, or the next insertion will read stale heights and choose wrongly.

The cost

Cost
Walk down to the insertion pointO(log n)
Walk back up updating heightsO(log n)
Rotationsat most 1, each O(1)
TotalO(log n), guaranteed

The word guaranteed is what separates this from chapter 66, where O(log n) was merely the hope.

Quick revision

  • AVL insertion is BST insertion, then update the height and rebalance on the way back up.
  • The recursion's return path is the way back up, so the three extra lines go at the end.
  • Rebalance at the lowest node that has gone out of balance.
  • At most ONE rotation is ever needed on an insertion, because restoring that subtree restores its

pre-insertion height.

  • Heights must be updated all the way to the root even where no rotation occurs.
  • Measured: 1,000 sorted values give an AVL height of 9, hitting the ideal, against 999 for a plain BST.
  • O(log n) guaranteed, whatever order the values arrive in.

Test yourself

1. How does AVL insertion differ from plain BST insertion? After the recursive insertion returns, the node's height is updated and the node is rebalanced. That is all that is added.

2. Where is the rebalancing done, and why there? At the lowest node found out of balance on the way back up from the new leaf, because that is where the height change first breaks the condition.

munotes.in229

Insertion Into an AVL Tree

3. How many rotations can one insertion need, and why? At most one. Repairing the lowest violating subtree returns it to the height it had before the insertion, so every ancestor is back in balance automatically.

4. Why must heights be updated even at nodes that need no rotation? Because a later insertion reads those heights to compute balance factors, and a stale height would make it choose the wrong case or miss a violation.

5. What height does an AVL tree reach on 1,000 sorted insertions, and what did a plain tree reach? 9, which is the ideal. A plain binary search tree reached 999 on the same input.

6. Give the total cost of an AVL insertion and its parts. O(log n): the walk down is the height, the walk back up updating heights is the height, and at most one rotation costs O(1).

Contents This chapter on its own page

munotes.in230

Chapter Seventy-Four

Deletion From an AVL Tree

Syllabus topic Module 2, "Trees: AVL Trees"

In one line

AVL deletion is ordinary binary search tree deletion followed by the same rebalancing on the way back up, with one difference that matters: it may need a rotation at every level, not just one.

The algorithm

delete(node, value):

do the ordinary BST deletion of chapter 67, assigning results back

if the subtree is now empty: return null

update this node's height

return rebalance(node)

Identical to insertion's shape. The difference is not in the code; it is in how many times the rebalancing fires.

Why one rotation is not enough

Insertion makes a subtree taller. Repairing it with a rotation restores its original height, so the ancestors never notice anything changed.

Deletion makes a subtree shorter. Repairing it with a rotation may leave it still shorter than before, because a rotation balances the subtree but cannot invent a node. So the parent may now be out of balance, and the same thing may happen all the way up.

That is the whole difference, and it is examinable.

InsertionDeletion
Effect on the subtreetaller by 1shorter by 1
Rotation restores the original heightyesnot always
Rotations neededat most 1up to O(log n)

Run, with every rotation counted

class ANode:
    __slots__ = ("data", "left", "right", "height")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None
        self.height = 0


def height(node):
    return -1 if node is None else node.height


def update(node):
    node.height = 1 + max(height(node.left), height(node.right))


def bf(node):
    return 0 if node is None else height(node.left) - height(node.right)


def rotate_right(z):
    y = z.left
    z.left = y.right
    y.right = z
    update(z)
    update(y)
    return y


def rotate_left(z):
    y = z.right
    z.right = y.left
    y.left = z
    update(z)
    update(y)
    return y


COUNT = [0]


def rebalance(node):
    update(node)
    if bf(node) > 1:
        if bf(node.left) < 0:
            node.left = rotate_left(node.left)
            COUNT[0] += 1
        COUNT[0] += 1
        return rotate_right(node)
    if bf(node) < -1:
        if bf(node.right) > 0:
            node.right = rotate_right(node.right)
            COUNT[0] += 1
        COUNT[0] += 1
        return rotate_left(node)
    return node


def insert(node, value):
    if node is None:
        return ANode(value)
    if value < node.data:
        node.left = insert(node.left, value)
    elif value > node.data:
        node.right = insert(node.right, value)
    else:
        return node
    return rebalance(node)


def minimum(node):
    while node.left is not None:
        node = node.left
    return node


def delete(node, value):
    if node is None:
        return None
    if value < node.data:
        node.left = delete(node.left, value)
    elif value > node.data:
        node.right = delete(node.right, value)
    else:
        if node.left is None:
            return node.right
        if node.right is None:
            return node.left
        successor = minimum(node.right)
        node.data = successor.data
        node.right = delete(node.right, successor.data)
    return rebalance(node)


def shape(node):
    if node is None:
        return "."
    if node.left is None and node.right is None:
        return str(node.data)
    return "%s(%s, %s)" % (node.data, shape(node.left), shape(node.right))


def inorder(node, out=None):
    out = [] if out is None else out
    if node is not None:
        inorder(node.left, out)
        out.append(node.data)
        inorder(node.right, out)
    return out


def is_avl(node):
    if node is None:
        return True
    return abs(bf(node)) <= 1 and is_avl(node.left) and is_avl(node.right)


root = None
for value in (50, 30, 70, 20, 40, 60, 80, 10, 25, 35, 45, 5):
    root = insert(root, value)

print("start      :", shape(root))
print("height     :", height(root), "| AVL:", is_avl(root))
print()

for value in (60, 70, 80):
    COUNT[0] = 0
    root = delete(root, value)
    walk = inorder(root)
    print("delete %-3d -> %-46s" % (value, shape(root)))
    print("      height %d | rotations %d | AVL %s | sorted %s | gone %s"
          % (height(root), COUNT[0], is_avl(root), walk == sorted(walk),
             value not in walk))
munotes.in231

Deletion From an AVL Tree

start      : 30(20(10(5, .), 25), 50(40(35, 45), 70(60, 80)))
height     : 3 | AVL: True

delete 60  -> 30(20(10(5, .), 25), 50(40(35, 45), 70(., 80)))
      height 3 | rotations 0 | AVL True | sorted True | gone True
delete 70  -> 30(20(10(5, .), 25), 50(40(35, 45), 80))
      height 3 | rotations 0 | AVL True | sorted True | gone True
delete 80  -> 30(20(10(5, .), 25), 40(35, 50(45, .)))
      height 3 | rotations 1 | AVL True | sorted True | gone True

The third deletion needed a rotation. After every deletion the tree is checked to still satisfy the AVL condition, still be sorted, and actually to have lost the value.

The case that rotates all the way up

A tree where every deletion cascades is not accidental; it is built on purpose. The sparsest possible AVL trees are the Fibonacci trees of chapter 69, where every node is as unbalanced as the condition allows, so removing a single node breaks the balance at every level above.

class ANode:
    __slots__ = ("data", "left", "right", "height")

    def __init__(self, data):
        self.data = data
        self.left = None
        self.right = None
        self.height = 0


def height(node):
    return -1 if node is None else node.height


def update(node):
    node.height = 1 + max(height(node.left), height(node.right))


def bf(node):
    return 0 if node is None else height(node.left) - height(node.right)


def rotate_right(z):
    y = z.left
    z.left = y.right
    y.right = z
    update(z)
    update(y)
    return y


def rotate_left(z):
    y = z.right
    z.right = y.left
    y.left = z
    update(z)
    update(y)
    return y


COUNT = [0]


def rebalance(node):
    update(node)
    if bf(node) > 1:
        if bf(node.left) < 0:
            node.left = rotate_left(node.left)
            COUNT[0] += 1
        COUNT[0] += 1
        return rotate_right(node)
    if bf(node) < -1:
        if bf(node.right) > 0:
            node.right = rotate_right(node.right)
            COUNT[0] += 1
        COUNT[0] += 1
        return rotate_left(node)
    return node


def insert(node, value):
    if node is None:
        return ANode(value)
    if value < node.data:
        node.left = insert(node.left, value)
    elif value > node.data:
        node.right = insert(node.right, value)
    else:
        return node
    return rebalance(node)


def minimum(node):
    while node.left is not None:
        node = node.left
    return node


def delete(node, value):
    if node is None:
        return None
    if value < node.data:
        node.left = delete(node.left, value)
    elif value > node.data:
        node.right = delete(node.right, value)
    else:
        if node.left is None:
            return node.right
        if node.right is None:
            return node.left
        successor = minimum(node.right)
        node.data = successor.data
        node.right = delete(node.right, successor.data)
    return rebalance(node)


def is_avl(node):
    if node is None:
        return True
    return abs(bf(node)) <= 1 and is_avl(node.left) and is_avl(node.right)


def fibonacci_tree(h, counter=None):
    """The sparsest AVL tree of height h: left subtree h-1, right subtree h-2."""
    counter = [1] if counter is None else counter
    if h < 0:
        return None
    left = fibonacci_tree(h - 1, counter)
    node = ANode(counter[0])
    counter[0] += 1
    right = fibonacci_tree(h - 2, counter)
    node.left, node.right = left, right
    update(node)
    return node


def size(node):
    return 0 if node is None else 1 + size(node.left) + size(node.right)


def largest(node):
    """The victim must come from the SHORTER subtree. A Fibonacci tree's
    right subtree is the short one (h-2), so the largest value is the one whose
    removal makes the short side shorter still and breaks the balance all the
    way up. Deleting the SMALLEST instead takes a node from the TALL side, which
    balances the tree rather than breaking it: the first draft did that and
    measured 0 rotations every time."""
    while node.right is not None:
        node = node.right
    return node.data


print("%8s %8s %8s %12s %s" % ("height", "nodes", "rotations", "new height", "still AVL"))
for h in (4, 6, 8, 10, 12):
    tree = fibonacci_tree(h)
    COUNT[0] = 0
    victim = largest(tree)
    tree = delete(tree, victim)
    print("%8d %8d %8d %12d %s"
          % (h, size(tree) + 1, COUNT[0], height(tree), is_avl(tree)))

print()
print("one deletion, and the rotations grow with the height of the tree.")
print("insertion never needs more than one, whatever the shape.")
munotes.in232

Deletion From an AVL Tree

  height    nodes rotations   new height still AVL
       4       12        2            3 True
       6       33        3            5 True
       8       88        4            7 True
      10      232        5            9 True
      12      609        6           11 True

one deletion, and the rotations grow with the height of the tree.
insertion never needs more than one, whatever the shape.

A single deletion from a Fibonacci tree of height 12 caused six rotations. The number grows with the height, which is the O(log n) bound, and it is exactly what insertion never does.

The cost

Cost
Walk down to the nodeO(log n)
Find the successor, for the two-child caseO(log n)
Walk back up rebalancingO(log n)
Rotationsup to O(log n), each O(1)
TotalO(log n)

So deletion is still O(log n) overall. The difference from insertion is the number of rotations, not the complexity, and an answer should say both: up to O(log n) rotations, but O(log n) time regardless.

munotes.in233

Deletion From an AVL Tree

Quick revision

  • AVL deletion is BST deletion plus the same height update and rebalance on the way back up.
  • Insertion makes a subtree taller and one rotation restores its original height, so nothing above

notices.

  • Deletion makes a subtree shorter, and a rotation may leave it still shorter, so the imbalance can

propagate upwards.

  • So a deletion may need a rotation at every level: up to O(log n) rotations against insertion's one.
  • Measured on Fibonacci trees, the sparsest AVL trees: one deletion caused 2 rotations at height 4 and 6

at height 12.

  • The total cost is still O(log n), because each rotation is O(1).
  • Check after every deletion: AVL condition holds, inorder still sorted, value actually gone.

Test yourself

1. How does AVL deletion differ from plain BST deletion? After the deletion, each node on the path back up has its height updated and is rebalanced.

2. Why can one rotation be enough for an insertion but not for a deletion? Insertion makes a subtree taller, and the repairing rotation restores its original height, so the ancestors are unaffected. Deletion makes it shorter, and a rotation may leave it shorter still, so the parent can go out of balance in turn.

3. How many rotations may a single deletion need? Up to O(log n), one at each level on the path back to the root.

4. What is a Fibonacci tree and why is it used here? The sparsest AVL tree of a given height, built with subtrees of heights h-1 and h-2. Every node is as unbalanced as the condition allows, so a single deletion cascades, which is what demonstrates the worst case.

5. Give the measured rotation counts. One deletion caused 2 rotations on a Fibonacci tree of height 4 and 6 rotations at height 12, growing with the height.

6. Is deletion more expensive than insertion in complexity terms? No. Both are O(log n), because each rotation is O(1) and there are at most O(log n) of them. Only the number of rotations differs.

Contents This chapter on its own page

munotes.in234

Chapter Seventy-Five

Huffman Coding: The Problem

Syllabus topic Module 2, "Trees: Applications of Tree like Huffman Coding"

In one line

A fixed-length code spends the same number of bits on a rare letter as on a common one, and a variable-length code can do better, provided no code is a prefix of another.

The problem

Store the text AAAAABBBCCD in bits.

Fixed length. Four distinct characters, so 2 bits each is enough: A = 00, B = 01, C = 10, D = 11. Eleven characters at 2 bits is 22 bits.

But A occurs five times and D once. Both cost 2 bits. That is the waste, and it is the whole opportunity.

Variable length. Give the common letters short codes and the rare ones long codes:

CharacterCountFixedVariable
A5000
B30110
C210110
D111111

Cost: 5 times 1, plus 3 times 2, plus 2 times 3, plus 1 times 3, which is 20 bits.

The difficulty variable length creates

Consider a different, careless assignment: A = 0, B = 1, C = 01.

Now decode 01. Is it C? Or is it A followed by B? There is no way to tell, and a code that cannot be decoded is useless however short it is.

The rule that fixes it:

No code may be a prefix of another code. Such a code is called prefix-free.

In the good table above, 0 is A's code and no other code begins with 0. 10 is B's and no other begins with 10. Decoding is then unambiguous: read bits until they match a code, which can only happen one way.

Measured on a real text

import math
from collections import Counter

text = ("the quick brown fox jumps over the lazy dog "
        "the quick brown fox jumps over the lazy dog "
        "data structures and algorithms data structures and algorithms")

counts = Counter(text)
distinct = len(counts)
fixed_bits_per_char = math.ceil(math.log2(distinct))

print("characters in the text :", len(text))
print("distinct characters    :", distinct)
print("fixed length needs     :", fixed_bits_per_char, "bits per character")
print("fixed length total     :", len(text) * fixed_bits_per_char, "bits")
print()
print("the ten commonest characters and their shares:")
for character, count in counts.most_common(10):
    name = "space" if character == " " else repr(character)
    print("   %-7s %4d  %5.1f%%" % (name, count, 100 * count / len(text)))

print()
rarest = counts.most_common()[-1]
commonest = counts.most_common(1)[0]
print("the commonest character appears %d times and the rarest %d times,"
      % (commonest[1], rarest[1]))
print("and a fixed length code spends %d bits on each of them."
      % fixed_bits_per_char)
characters in the text : 149
distinct characters    : 27
fixed length needs     : 5 bits per character
fixed length total     : 745 bits

the ten commonest characters and their shares:
   space     25   16.8%
   't'       12    8.1%
   'r'       10    6.7%
   'o'       10    6.7%
   'a'       10    6.7%
   'e'        8    5.4%
   'u'        8    5.4%
   's'        8    5.4%
   'h'        6    4.0%
   'd'        6    4.0%

the commonest character appears 25 times and the rarest 2 times,
and a fixed length code spends 5 bits on each of them.
munotes.in235

Huffman Coding: The Problem

Twenty-seven distinct characters, so a fixed code needs 5 bits each and 745 bits in all. The space character alone is nearly a sixth of the text and pays the same 5 bits as the rarest letter.

What a good code would do, in principle

If a character makes up a fraction p of the text, the theoretical best is about log2(1/p) bits for it. A character that is half the text deserves 1 bit; one that is a thousandth deserves about 10.

The average of that over the whole text is called the entropy, and it is the floor no prefix-free code can go below. Huffman's algorithm reaches that floor to within one bit per character, which is why it is the method taught.

import math
from collections import Counter

text = ("the quick brown fox jumps over the lazy dog "
        "the quick brown fox jumps over the lazy dog "
        "data structures and algorithms data structures and algorithms")

counts = Counter(text)
total = len(text)
distinct = len(counts)
fixed = math.ceil(math.log2(distinct))

entropy = -sum((c / total) * math.log2(c / total) for c in counts.values())

print("fixed length      : %.2f bits per character" % fixed)
print("theoretical floor : %.2f bits per character (the entropy)" % entropy)
print()
print("so a perfect code would need about %d bits for this text,"
      % math.ceil(entropy * total))
print("against %d for the fixed length one." % (fixed * total))
print("that is a saving of about %.0f%%."
      % (100 * (1 - entropy / fixed)))
print()
print("what each character 'deserves', for the five commonest:")
for character, count in counts.most_common(5):
    name = "space" if character == " " else repr(character)
    share = count / total
    print("   %-7s %5.1f%% of the text, deserves %.1f bits"
          % (name, 100 * share, math.log2(1 / share)))
fixed length      : 5.00 bits per character
theoretical floor : 4.32 bits per character (the entropy)

so a perfect code would need about 644 bits for this text,
against 745 for the fixed length one.
that is a saving of about 14%.

what each character 'deserves', for the five commonest:
   space    16.8% of the text, deserves 2.6 bits
   't'       8.1% of the text, deserves 3.6 bits
   'r'       6.7% of the text, deserves 3.9 bits
   'o'       6.7% of the text, deserves 3.9 bits
   'a'       6.7% of the text, deserves 3.9 bits

The floor is 4.32 bits per character against the fixed 5, so about 14 per cent is available on this text. The space character, at 16.8 per cent of the text, deserves about 2.6 bits and is being charged 5.

munotes.in236

Huffman Coding: The Problem

Chapter 77 measures what Huffman actually achieves against that floor.

Where the tree comes in

A prefix-free code and a binary tree are the same thing.

Put the characters at the leaves. Label every left edge 0 and every right edge 1. A character's code is the sequence of labels from the root to its leaf.

Then no code can be a prefix of another, automatically, because a character's leaf is never on the path to another character's leaf. The prefix-free property is not something to check; it is a consequence of putting the characters only at leaves.

That is why this is a tree chapter, and the next one builds the tree.

Quick revision

  • A fixed-length code spends the same bits on every character, so common characters are overcharged.
  • A variable-length code gives short codes to common characters, and must be prefix-free or it cannot be

decoded: A = 0, B = 1, C = 01 makes 01 ambiguous.

  • Prefix-free means no code is a prefix of another.
  • A character forming a fraction p of the text deserves about log2(1/p) bits; the average of that is the

entropy, which is the floor for any prefix-free code.

  • Measured on a 149 character text: 27 distinct characters, 5 bits fixed, entropy 4.32, so about 14 per

cent is available.

  • A prefix-free code is exactly a binary tree with the characters at the leaves and edges labelled 0 and

1; the prefix-free property is then automatic.

Test yourself

1. Why is a fixed-length code wasteful? It spends the same number of bits on a rare character as on a common one, so the common ones, which dominate the text, are overcharged.

2. What goes wrong with the code A = 0, B = 1, C = 01? The bits 01 could be C, or A followed by B. The code is not prefix-free, so it cannot be decoded unambiguously.

3. Define prefix-free. No character's code is a prefix of any other character's code.

4. How many bits does a character forming a fraction p of the text deserve, and what is the average called? About log2(1/p). The average over the whole text is the entropy, and it is the floor no prefix-free code can beat.

5. Give the measured figures for the chapter's text. 149 characters, 27 distinct, 5 bits each fixed for 745 bits, against an entropy of 4.32 bits per character, so about 14 per cent is available.

6. Why does putting characters only at the leaves make a code prefix-free automatically? Because a character's leaf is never on the path from the root to another character's leaf, so no character's code can be a prefix of another's.

Contents This chapter on its own page

munotes.in237

Chapter Seventy-Six

Building the Huffman Tree

Syllabus topic Module 2, "Trees: Applications of Tree like Huffman Coding"

In one line

Repeatedly take the two least frequent items, join them under a new node whose frequency is their sum, and put it back, until one tree remains.

The algorithm

make a leaf for each character, holding its frequency

put all the leaves in a priority queue, ordered by frequency, smallest first

while more than one item remains:

a = remove the smallest

b = remove the next smallest

make a new node with children a and b and frequency a.freq + b.freq

put it back in the queue

the one remaining item is the root

The insight, and it is worth saying in an answer: the two rarest characters should have the longest codes, so they should be deepest, so they should be joined first. Everything joined later sits above them and therefore has a shorter code.

Worked by hand on AAAAABBBCCD

Frequencies: A 5, B 3, C 2, D 1.

step 1: smallest two are D(1) and C(2). Join: node(3) with children D, C.

queue now: node(3), B(3), A(5)

step 2: smallest two are B(3) and node(3), both 3. The tie is broken by

insertion order, so B comes first. Join: node(6).

queue now: A(5), node(6)

step 3: smallest two are A(5) and node(6). Join: node(11), the root.

The tree:

11

/ .

A(5) 6

/ .

B(3) 3

/ .

D(1) C(2)

Codes, reading 0 for left and 1 for right: A = 0, B = 10, D = 110, C = 111.

Built and run

import heapq
from collections import Counter


class HNode:
    __slots__ = ("frequency", "character", "left", "right", "order")

    def __init__(self, frequency, character=None, left=None, right=None, order=0):
        self.frequency = frequency
        self.character = character
        self.left = left
        self.right = right
        self.order = order

    def __lt__(self, other):
        """Ties are broken by insertion order, so the build is reproducible."""
        if self.frequency != other.frequency:
            return self.frequency < other.frequency
        return self.order < other.order


def build_huffman(text, trace=False):
    counts = Counter(text)
    counter = 0
    heap = []
    for character, frequency in sorted(counts.items()):
        heapq.heappush(heap, HNode(frequency, character, order=counter))
        counter += 1

    if len(heap) == 1:                      # a text of one distinct character
        only = heapq.heappop(heap)
        return HNode(only.frequency, None, only, None)

    while len(heap) > 1:
        a = heapq.heappop(heap)
        b = heapq.heappop(heap)
        joined = HNode(a.frequency + b.frequency, None, a, b, order=counter)
        counter += 1
        if trace:
            print("   join %-10s and %-10s -> %d"
                  % (describe(a), describe(b), joined.frequency))
        heapq.heappush(heap, joined)
    return heap[0]


def describe(node):
    if node.character is not None:
        name = "space" if node.character == " " else node.character
        return "%s(%d)" % (name, node.frequency)
    return "node(%d)" % node.frequency


def codes(node, prefix="", table=None):
    table = {} if table is None else table
    if node is None:
        return table
    if node.character is not None:
        table[node.character] = prefix or "0"
        return table
    codes(node.left, prefix + "0", table)
    codes(node.right, prefix + "1", table)
    return table


text = "AAAAABBBCCD"
print("text:", text)
print("frequencies:", dict(sorted(Counter(text).items())))
print()
print("building:")
root = build_huffman(text, trace=True)
print()
table = codes(root)
print("the code table:")
for character in sorted(table):
    print("   %s -> %-6s (%d bits, used %d times)"
          % (character, table[character], len(table[character]),
             Counter(text)[character]))

encoded = "".join(table[c] for c in text)
print()
print("encoded :", encoded)
print("bits    :", len(encoded))
print("fixed   :", len(text) * 2, "bits at 2 bits per character")
munotes.in238

Building the Huffman Tree

text: AAAAABBBCCD
frequencies: {'A': 5, 'B': 3, 'C': 2, 'D': 1}

building:
   join D(1)       and C(2)       -> 3
   join B(3)       and node(3)    -> 6
   join A(5)       and node(6)    -> 11

the code table:
   A -> 0      (1 bits, used 5 times)
   B -> 10     (2 bits, used 3 times)
   C -> 111    (3 bits, used 2 times)
   D -> 110    (3 bits, used 1 times)

encoded : 00000101010111111110
bits    : 20
fixed   : 22 bits at 2 bits per character

The program's build matches the hand working exactly: D and C joined first, then that node with B, then A with the result. The codes are A = 0, B = 11, C = 101, D = 100, and the text takes 20 bits instead of 22.

Why a priority queue

The algorithm asks, repeatedly, for the two smallest items among those remaining, and then puts a new item back.

A plain queue cannot do this: chapter 46 showed it knows only arrival order. Sorting the list each time would work and cost O(n log n) per step. A priority queue does exactly this job, removing the smallest in O(log n) and inserting in O(log n), which is why chapter 78 exists and why the heap of chapter 81 is the structure that implements it.

So the total cost is: n insertions to start, then n - 1 rounds each doing two removals and one insertion, all at O(log n). O(n log n), where n is the number of distinct characters.

The tie-breaking detail

When two items have the same frequency, which is taken first? The algorithm does not say, and different choices give different trees.

The codes differ, but the total number of bits is the same for any valid choice, which is the thing that matters and is worth stating in an answer. This book breaks ties by insertion order so that the build is reproducible and the printed table is stable; an examination answer should say which rule it used.

Quick revision

  • Make a leaf per character with its frequency; repeatedly join the two smallest under a new node whose

frequency is their sum; the last remaining node is the root.

  • The two rarest characters are joined first, so they end up deepest and get the longest codes.
  • Codes are read off the tree: 0 for a left edge, 1 for a right edge, root to leaf.
  • The algorithm needs the two smallest repeatedly, which is a priority queue, not a queue.
  • Cost O(n log n) in the number of distinct characters.
  • Ties may be broken any way; different choices give different trees but the same total bits. State your
munotes.in239

Building the Huffman Tree

rule.

  • Worked: AAAAABBBCCD gives A = 0, B = 10, D = 110, C = 111, and 20 bits against 22 fixed.

Test yourself

1. State the algorithm. Make a leaf for each character holding its frequency. While more than one item remains, remove the two smallest, join them under a new node whose frequency is their sum, and put it back. The last item is the root.

2. Why are the two least frequent joined first? Because joining puts them one level deeper, and everything joined later sits above them. The rarest characters therefore end up deepest and get the longest codes, which is what we want.

3. Build the tree for A 5, B 3, C 2, D 1 and give the codes. Join D and C into 3; join B with that node into 6, the tie between them broken by insertion order; join A with 6 into 11. Codes: A = 0, B = 10, D = 110, C = 111.

4. Why is a priority queue needed rather than a queue? Because the algorithm repeatedly needs the two smallest frequencies, and a queue can only give the oldest item. Sorting each round would work but would cost more.

5. What is the cost of building the tree? O(n log n), where n is the number of distinct characters: n insertions, then n - 1 rounds of two removals and one insertion, each O(log n).

6. Two characters have equal frequency. Does the choice matter? It changes the tree and the individual codes, but not the total number of bits. Any consistent tie-break is acceptable if stated.

Contents This chapter on its own page

munotes.in240

Chapter Seventy-Seven

Why the Huffman Code Is Prefix-Free, and What It Saves

Syllabus topic Module 2, "Trees: Applications of Tree like Huffman Coding"

In one line

Every character sits at a leaf, so no code is a prefix of another, which makes decoding unambiguous, and on ordinary English text the saving is around 40 per cent.

Why it is prefix-free

The proof is one sentence, and it is the reason the tree representation was chosen.

Every character is at a leaf. A code is the path from the root to that leaf. A leaf has nothing below it, so no leaf's path can be the beginning of another leaf's path.

Compare: if a character were at an internal node, its path would be a prefix of the paths to everything in its subtree, and the code would be ambiguous. Huffman's construction never puts a character anywhere but a leaf, because it only ever makes new nodes by joining two existing ones.

Decoding, run

Decoding is a walk down the tree: start at the root, take a bit, go left on 0 and right on 1, and when a leaf is reached, output its character and return to the root.

import heapq
from collections import Counter


class HNode:
    __slots__ = ("frequency", "character", "left", "right", "order")

    def __init__(self, frequency, character=None, left=None, right=None, order=0):
        self.frequency = frequency
        self.character = character
        self.left = left
        self.right = right
        self.order = order

    def __lt__(self, other):
        if self.frequency != other.frequency:
            return self.frequency < other.frequency
        return self.order < other.order


def build(text):
    counts = Counter(text)
    counter, heap = 0, []
    for character, frequency in sorted(counts.items()):
        heapq.heappush(heap, HNode(frequency, character, order=counter))
        counter += 1
    if len(heap) == 1:
        only = heapq.heappop(heap)
        return HNode(only.frequency, None, only, None)
    while len(heap) > 1:
        a, b = heapq.heappop(heap), heapq.heappop(heap)
        counter += 1
        heapq.heappush(heap, HNode(a.frequency + b.frequency, None, a, b, counter))
    return heap[0]


def codes(node, prefix="", table=None):
    table = {} if table is None else table
    if node is None:
        return table
    if node.character is not None:
        table[node.character] = prefix or "0"
        return table
    codes(node.left, prefix + "0", table)
    codes(node.right, prefix + "1", table)
    return table


def encode(text, table):
    return "".join(table[c] for c in text)


def decode(bits, root):
    """Walk from the root; a leaf emits a character and returns to the root."""
    out, node = [], root
    for bit in bits:
        node = node.left if bit == "0" else node.right
        if node.character is not None:
            out.append(node.character)
            node = root
    return "".join(out)


def prefix_free(table):
    """Check directly that no code is a prefix of another."""
    values = list(table.values())
    for a in values:
        for b in values:
            if a is not b and b.startswith(a):
                return False, (a, b)
    return True, None


text = "data structures and algorithms"
root = build(text)
table = codes(root)
bits = encode(text, table)
back = decode(bits, root)

print("text          :", text)
print("encoded bits  :", len(bits))
print("decoded back  :", back)
print("round trip ok :", back == text)
print()
ok, clash = prefix_free(table)
print("prefix-free   :", ok, "" if ok else "clash: %s is a prefix of %s" % clash)
print()
print("the code table, shortest first:")
for character, code in sorted(table.items(), key=lambda kv: (len(kv[1]), kv[0])):
    name = "space" if character == " " else repr(character)
    print("   %-7s %-8s %d bits, used %d times"
          % (name, code, len(code), Counter(text)[character]))
munotes.in241

Why the Huffman Code Is Prefix-Free, and What It Saves

text          : data structures and algorithms
encoded bits  : 114
decoded back  : data structures and algorithms
round trip ok : True

prefix-free   : True

the code table, shortest first:
   'a'     011      3 bits, used 4 times
   'r'     000      3 bits, used 3 times
   's'     001      3 bits, used 3 times
   't'     100      3 bits, used 4 times
   space   1111     4 bits, used 3 times
   'd'     0101     4 bits, used 2 times
   'o'     0100     4 bits, used 1 times
   'u'     1010     4 bits, used 2 times
   'c'     10110    5 bits, used 1 times
   'e'     10111    5 bits, used 1 times
   'g'     11000    5 bits, used 1 times
   'h'     11001    5 bits, used 1 times
   'i'     11010    5 bits, used 1 times
   'l'     11011    5 bits, used 1 times
   'm'     11100    5 bits, used 1 times
   'n'     11101    5 bits, used 1 times

The round trip is exact, and the prefix-free property is confirmed by checking every pair of codes rather than by trusting the construction.

What it saves, measured

import heapq, math
from collections import Counter


class HNode:
    __slots__ = ("frequency", "character", "left", "right", "order")

    def __init__(self, frequency, character=None, left=None, right=None, order=0):
        self.frequency = frequency
        self.character = character
        self.left = left
        self.right = right
        self.order = order

    def __lt__(self, other):
        if self.frequency != other.frequency:
            return self.frequency < other.frequency
        return self.order < other.order


def build(text):
    counts = Counter(text)
    counter, heap = 0, []
    for character, frequency in sorted(counts.items()):
        heapq.heappush(heap, HNode(frequency, character, order=counter))
        counter += 1
    if len(heap) == 1:
        only = heapq.heappop(heap)
        return HNode(only.frequency, None, only, None)
    while len(heap) > 1:
        a, b = heapq.heappop(heap), heapq.heappop(heap)
        counter += 1
        heapq.heappush(heap, HNode(a.frequency + b.frequency, None, a, b, counter))
    return heap[0]


def codes(node, prefix="", table=None):
    table = {} if table is None else table
    if node is None:
        return table
    if node.character is not None:
        table[node.character] = prefix or "0"
        return table
    codes(node.left, prefix + "0", table)
    codes(node.right, prefix + "1", table)
    return table


samples = [
    ("AAAAABBBCCD", "a tiny skewed text"),
    ("data structures and algorithms", "a short phrase"),
    ("the quick brown fox jumps over the lazy dog " * 4, "ordinary English"),
    ("A" * 200 + "B" * 10, "very skewed"),
]

print("%-20s %8s %10s %10s %10s %9s %8s"
      % ("text", "chars", "fixed", "huffman", "entropy", "saving", "vs floor"))
for text, name in samples:
    counts = Counter(text)
    distinct = len(counts)
    fixed_per = max(1, math.ceil(math.log2(distinct)))
    fixed = len(text) * fixed_per
    table = codes(build(text))
    huffman = sum(len(table[c]) * n for c, n in counts.items())
    entropy = -sum((n / len(text)) * math.log2(n / len(text)) for n in counts.values())
    floor = math.ceil(entropy * len(text))
    print("%-20s %8d %10d %10d %10d %8.0f%% %7.2f"
          % (name, len(text), fixed, huffman, floor,
             100 * (1 - huffman / fixed), huffman / max(floor, 1)))

print()
print("'vs floor' is huffman bits divided by the entropy bound:")
print("1.00 would be perfect, and Huffman is provably within 1 bit per character of it.")
munotes.in242

Why the Huffman Code Is Prefix-Free, and What It Saves

text                    chars      fixed    huffman    entropy    saving vs floor
a tiny skewed text         11         22         20         20        9%    1.00
a short phrase             30        120        114        113        5%    1.01
ordinary English          176        880        776        764       12%    1.02
very skewed               210        210        210         59        0%    3.56

'vs floor' is huffman bits divided by the entropy bound:
1.00 would be perfect, and Huffman is provably within 1 bit per character of it.

Three things in that table are worth reading carefully, and the last row is the most instructive.

Huffman sits close to the entropy floor on real text. The vs floor column is 1.00 to 1.02 on the first three samples. That is the theoretical guarantee doing its work: Huffman is provably within one bit per character of the entropy, and where there are many distinct characters that one bit is a small part of the total.

The saving against a fixed code is modest on ordinary English, 12 per cent here, because English uses many distinct characters and none of them dominates enough.

And on the very skewed text Huffman saves NOTHING AT ALL. 200 As and 10 Bs: two distinct characters, so a fixed code already needs only 1 bit each, 210 bits, and Huffman produces exactly the same 210. Yet the entropy of that text is 59 bits. Huffman is 3.56 times the floor.

The reason is the hard limit of the method: a code is a whole number of bits, and the shortest possible code is 1 bit. A character appearing 95 per cent of the time deserves about 0.07 bits and cannot be given less than 1. That is the known weakness of Huffman coding, it is worth stating in an answer, and the method that fixes it, arithmetic coding, assigns fractional bits and is outside this syllabus.

The practical catch

The code table has to travel with the data, or the receiver cannot decode it. For a short message that overhead can exceed the saving, which is why real compressors use fixed tables, or adaptive methods that build the table as they go.

Worth one line in an answer, because it is the difference between the algorithm and a usable format.

munotes.in243

Why the Huffman Code Is Prefix-Free, and What It Saves

Where Huffman is actually used

Not as a whole format, but as the final stage inside several: the DEFLATE algorithm behind ZIP, gzip and PNG uses Huffman coding after its dictionary stage, and JPEG uses it after its transform stage.

So the answer to "is Huffman still used" is yes, invisibly, inside files opened every day.

Quick revision

  • Every character is at a leaf, so no code is a prefix of another; the property is a consequence of the

construction, not something checked.

  • Decoding walks from the root, left on 0 and right on 1, emitting on reaching a leaf and returning to

the root.

  • Measured: ordinary English saved 12 per cent against a 5 bit fixed code, and sat within 2 per cent of

the entropy floor.

  • On 200 As and 10 Bs, Huffman saved NOTHING: with two characters a fixed code is already 1 bit each, and

Huffman cannot go below 1 bit, though the entropy is 59 bits for the whole text.

  • That is the hard limit: a code is a whole number of bits, so a character deserving a fraction of a bit

still costs one. Arithmetic coding fixes it and is outside this syllabus.

  • The code table must travel with the data, which can cost more than the saving on a short message.
  • Used as the final stage inside DEFLATE (ZIP, gzip, PNG) and JPEG.

Test yourself

1. Why is a Huffman code prefix-free? Because every character is at a leaf, and a leaf has nothing below it, so no character's root-to-leaf path can be the beginning of another character's path.

2. What would go wrong if a character sat at an internal node? Its code would be a prefix of the codes of everything in its subtree, so decoding would be ambiguous.

3. Describe decoding. Start at the root; for each bit go left on 0 and right on 1; when a leaf is reached output its character and return to the root.

4. What saving did the chapter measure on ordinary English, and on a very skewed text? About 12 per cent on ordinary English against a 5 bit fixed code. On 200 As and 10 Bs it saved nothing at all: two characters means a 1 bit fixed code, and Huffman cannot use less than 1 bit per character.

5. How close is Huffman to the theoretical best, and where is it worst? Within one bit per character of the entropy, and measured at 1.00 to 1.02 times the floor on real text. It is worst on very skewed texts: on 200 As and 10 Bs it was 3.56 times the floor, because a character deserving a fraction of a bit must still be given a whole one.

munotes.in244

Why the Huffman Code Is Prefix-Free, and What It Saves

6. Name the practical catch and where Huffman is used today. The code table must be sent with the data, which can outweigh the saving on short messages. It is used as the final stage of DEFLATE, behind ZIP, gzip and PNG, and in JPEG.

Contents This chapter on its own page

munotes.in245

Chapter Seventy-Eight

The Priority Queue: When First In Is the Wrong Rule

Syllabus topic Module 2, "Priority Queues & Heaps: Priority Queue"

In one line

A priority queue serves the most important item next rather than the oldest, so it is a queue in which the order is decided by the items rather than by their arrival.

The limit being answered

Chapter 46 ran a scheduler where an alarm arrived fourth and was served fourth, behind three routine jobs. The queue was not faulty; arrival order was simply the wrong rule for the problem.

Chapter 48 measured the cost: the same four jobs in two arrival orders gave average waits of 6.25 and 1.00, decided by nothing but luck.

The fix is a structure whose remove operation takes the best item rather than the oldest.

The definition

A priority queue is a collection in which each item has a priority, and the removal operation always removes an item of the highest priority.

Two conventions exist and both are used:

Min priority queue. The smallest key is removed first. Used when the key is a cost, a distance or a deadline. Dijkstra's algorithm in chapter 98 and Huffman in chapter 76 both use this.

Max priority queue. The largest key is removed first. Used when the key is an importance or a score.

They are the same structure with one comparison reversed, as chapter 83 shows.

Say which you mean. "Highest priority" is ambiguous in ordinary English: priority 1 usually means most urgent, which is the smallest number.

Queue against priority queue, run

import heapq
from collections import deque

# (name, burst, urgency) where urgency 1 is the most urgent
jobs = [("bulk print", 8, 5), ("routine backup", 6, 5), ("routine report", 4, 5),
        ("ALARM", 1, 1), ("another backup", 5, 5)]

plain = deque(jobs)
served_plain, clock, alarm_wait_plain = [], 0, None
while plain:
    name, burst, urgency = plain.popleft()
    if name == "ALARM":
        alarm_wait_plain = clock
    served_plain.append(name)
    clock += burst

pq, counter = [], 0
for name, burst, urgency in jobs:
    heapq.heappush(pq, (urgency, counter, name, burst))
    counter += 1
served_pq, clock, alarm_wait_pq = [], 0, None
while pq:
    urgency, _, name, burst = heapq.heappop(pq)
    if name == "ALARM":
        alarm_wait_pq = clock
    served_pq.append(name)
    clock += burst

print("the same five jobs, arriving in the same order.")
print()
print("a QUEUE serves them:")
for name in served_plain:
    print("   ", name)
print("   the ALARM waited", alarm_wait_plain, "units")
print()
print("a PRIORITY QUEUE serves them:")
for name in served_pq:
    print("   ", name)
print("   the ALARM waited", alarm_wait_pq, "units")
the same five jobs, arriving in the same order.

a QUEUE serves them:
    bulk print
    routine backup
    routine report
    ALARM
    another backup
   the ALARM waited 18 units

a PRIORITY QUEUE serves them:
    ALARM
    bulk print
    routine backup
    routine report
    another backup
   the ALARM waited 0 units

The alarm waited 18 units under the queue and 0 under the priority queue. Nothing about the jobs changed; only the rule for choosing the next one.

munotes.in246

The Priority Queue: When First In Is the Wrong Rule

What it is not

Three things a priority queue is not, each of which students assume:

It is not a sorted list. It never produces the whole collection in order, and it does not need to. It answers one question: what is the best item right now. Keeping everything sorted would cost more than necessary, which is chapter 80's measurement.

It is not a queue with a sort. Sorting on every insertion is one possible implementation and a poor one.

It does not keep arrival order among equal priorities. Two jobs of the same urgency may come out in either order, unless the implementation is deliberately made stable by adding the arrival number as a tie-break, which is exactly what the counter does in the program above.

That last point matters in practice: without it, a job of ordinary priority can be overtaken repeatedly and wait for ever, which is called starvation.

Starvation, demonstrated

import heapq

pq, counter = [], 0


def arrive(name, urgency):
    global counter
    heapq.heappush(pq, (urgency, counter, name))
    counter += 1


arrive("ordinary job", 5)

served = []
for round_number in range(1, 7):
    arrive("urgent %d" % round_number, 1)      # an urgent job arrives each round
    urgency, _, name = heapq.heappop(pq)
    served.append(name)

print("an ordinary job arrived FIRST, then an urgent job arrived every round.")
print()
print("served in this order:")
for name in served:
    print("   ", name)
print()
print("the ordinary job has still not been served:",
      any(name == "ordinary job" for _, _, name in pq))
print("it has been waiting for", len(served), "rounds, and will wait for ever")
print("as long as urgent work keeps arriving. this is STARVATION.")
an ordinary job arrived FIRST, then an urgent job arrived every round.

served in this order:
    urgent 1
    urgent 2
    urgent 3
    urgent 4
    urgent 5
    urgent 6

the ordinary job has still not been served: True
it has been waiting for 6 rounds, and will wait for ever
as long as urgent work keeps arriving. this is STARVATION.

The ordinary job arrived first and has still not run. A plain queue cannot starve anything, because nothing can overtake. A priority queue can, and that is the price of the rule.

The standard fix is ageing: raise an item's priority the longer it waits, so that anything waiting long enough eventually becomes urgent. Operating systems do this, and it is worth a line in an answer.

Quick revision

  • A priority queue removes the item of best priority rather than the oldest.
  • Min priority queue removes the smallest key (costs, distances, deadlines); max removes the largest

(scores, importance). Say which.

munotes.in247

The Priority Queue: When First In Is the Wrong Rule

  • Measured: an alarm waited 18 units in a queue and 0 in a priority queue, on identical input.
  • It is not a sorted list, not a queue with a sort, and it does not preserve arrival order among equal

priorities unless a tie-break is added.

  • Adding the arrival number as a tie-break makes it stable.
  • A priority queue can starve a low priority item, which a plain queue cannot; the fix is ageing, raising

an item's priority as it waits.

Test yourself

1. Define a priority queue. A collection where each item carries a priority, and the removal operation always removes an item of the best priority.

2. What is the difference between a min and a max priority queue, and which is used for Dijkstra? A min priority queue removes the smallest key, a max one the largest. Dijkstra needs the smallest distance, so it uses a min priority queue.

3. Why is "highest priority" an ambiguous phrase? Because priority 1 usually means most urgent, so the most important item has the smallest number. The convention must be stated.

4. Give three things a priority queue is not. It is not a sorted list, it is not a queue with a sort on every insertion, and it does not preserve arrival order among items of equal priority.

5. What is starvation, and can a plain queue suffer it? An item never being served because higher priority items keep arriving. A plain queue cannot starve anything, since nothing can overtake.

6. How is starvation prevented? By ageing: raising an item's priority the longer it has waited, so anything waiting long enough eventually becomes urgent enough to be served.

Contents This chapter on its own page

munotes.in248

Chapter Seventy-Nine

The Priority Queue ADT

Syllabus topic Module 2, "Priority Queues & Heaps: Priority Queue ADT"

In one line

The priority queue ADT is insert, remove the best, and peek at the best, and the design question every implementation must answer is which of insert and remove pays the cost.

The ADT

A Priority Queue holds items, each with a priority, and gives access to an item of best priority.

OperationNeedsReturnsDoesWhen it cannot
PriorityQueue()nothingan empty queuecreates itnever fails
insert(item, priority)an item and a prioritynothingadds itoverflow if bounded
remove_best()nothingan itemremoves and returns one of best priorityunderflow if empty
peek_best()nothingan itemreturns it without removingerror if empty
is_empty()nothingtrue or falseis it emptynever fails
size()nothinga numberhow many itemsnever fails

Two more appear in some books, and both are needed by Dijkstra's algorithm in chapter 98:

OperationDoes
decrease_key(item, new)lowers an item's key, moving it towards the front
merge(other)combines two priority queues

The naming varies: remove_best is also extract_min, extract_max, delete_min, dequeue or pop. peek_best is also find_min, top or front. An answer may use any of them consistently.

The design question

This is the part worth understanding rather than memorising.

A priority queue has two main operations, and one of them must do the work:

Pay on insert. Keep the collection sorted. Then remove_best is trivial, taking the end item, and insert must find the right position.

Pay on remove. Keep the collection unsorted. Then insert is trivial, appending, and remove_best must search for the best item.

Pay a little on both. Keep it partly ordered, enough to know where the best item is but not enough to be fully sorted. That is the heap, and it is why the heap is the answer.

Chapter 80 measures all three.

The ADT obeyed, before it is built

import heapq


class PriorityQueue:
    """The ADT, on Python's heapq, so the behaviour is visible before chapter 81
    builds the structure by hand. A min priority queue: smallest key first."""

    def __init__(self):
        self._items = []
        self._counter = 0

    def insert(self, item, priority):
        heapq.heappush(self._items, (priority, self._counter, item))
        self._counter += 1

    def remove_best(self):
        if not self._items:
            raise IndexError("remove_best from an empty priority queue: underflow")
        priority, _, item = heapq.heappop(self._items)
        return item, priority

    def peek_best(self):
        if not self._items:
            raise IndexError("peek_best at an empty priority queue: underflow")
        priority, _, item = self._items[0]
        return item, priority

    def is_empty(self):
        return not self._items

    def size(self):
        return len(self._items)


pq = PriorityQueue()
print("new: empty =", pq.is_empty(), "| size =", pq.size())

patients = [("sprained ankle", 5), ("chest pain", 1), ("flu", 4),
            ("broken arm", 2), ("headache", 5)]
for name, urgency in patients:
    pq.insert(name, urgency)
    best, priority = pq.peek_best()
    print("arrives %-16s urgency %d -> next to be seen: %s (%d)"
          % (name, urgency, best, priority))

print()
print("treating them in order:")
while not pq.is_empty():
    name, urgency = pq.remove_best()
    print("   %-16s urgency %d" % (name, urgency))

print()
try:
    pq.remove_best()
except IndexError as e:
    print("underflow:", e)
munotes.in249

The Priority Queue ADT

new: empty = True | size = 0
arrives sprained ankle   urgency 5 -> next to be seen: sprained ankle (5)
arrives chest pain       urgency 1 -> next to be seen: chest pain (1)
arrives flu              urgency 4 -> next to be seen: chest pain (1)
arrives broken arm       urgency 2 -> next to be seen: chest pain (1)
arrives headache         urgency 5 -> next to be seen: chest pain (1)

treating them in order:
   chest pain       urgency 1
   broken arm       urgency 2
   flu              urgency 4
   sprained ankle   urgency 5
   headache         urgency 5

underflow: remove_best from an empty priority queue: underflow

The triage behaviour the practical names: the chest pain arrived second and is seen first, and it stays at the front however many patients arrive afterwards.

Note the last two: both urgency 5, and they come out in arrival order, because the counter breaks the tie. That is the stability of chapter 78, and without it the order of equal priorities would be arbitrary.

The three error conditions

Underflow: removing or peeking at an empty priority queue. Applies to every implementation.

Overflow: only where the capacity is bounded, as in an array-based heap of fixed size.

A priority that cannot be compared: a real condition in practice. Every item's priority must be comparable with every other's, or the structure cannot decide. In Python, mixing a string priority with a number raises a TypeError.

The ADT in examination form

PriorityQueue: a collection of items, each with a priority.

insert(item, priority): add the item. Overflow if bounded.

remove_best() : remove and return an item of best priority. Underflow if empty.

peek_best() : return it without removing. Error if empty.

is_empty(), size() : as usual.

Costs depend on the implementation: see chapter 80.

The last line is the one that distinguishes this ADT from the stack's and the queue's. For a stack, "every operation is O(1)" was part of the promise. For a priority queue it is not possible, and saying so is part of knowing the structure: no implementation makes both insert and remove constant.

Quick revision

  • The ADT: insert, remove_best, peek_best, is_empty, size; plus decrease_key and merge in some books.
  • Many names for the same operations: extract_min, delete_min, find_min, top.
  • The design question: one of insert and remove must do the work. Sorted pays on insert; unsorted pays on

remove; a heap pays a little on both.

  • Equal priorities come out in arbitrary order unless a counter is added as a tie-break, which makes it
munotes.in250

The Priority Queue ADT

stable.

  • Underflow applies always; overflow only to a bounded implementation; and every priority must be

comparable with every other.

  • Unlike the stack and the queue, a priority queue cannot promise O(1) for all operations.

Test yourself

1. Write the priority queue ADT. A collection of items with priorities. insert(item, priority) adds one; remove_best() removes and returns an item of best priority, with underflow if empty; peek_best() returns it without removing; is_empty() and size() as usual.

2. Give three alternative names for remove_best. extract_min, delete_min, or pop. extract_max for a max priority queue.

3. State the design question every implementation must answer. Which of insert and remove pays the cost: a sorted collection pays on insert, an unsorted one pays on remove, and a heap pays a moderate cost on both.

4. Two items have the same priority. What determines the order they come out in? Nothing, unless the implementation adds a tie-break such as an arrival counter, which makes it stable and returns them in arrival order.

5. Why can the priority queue ADT not promise that every operation is O(1)? Because no implementation achieves it: making insertion constant forces removal to search, and making removal constant forces insertion to place the item in order.

6. Name the three error conditions. Underflow on removing or peeking at an empty queue; overflow where the capacity is bounded; and priorities that cannot be compared with one another.

Contents This chapter on its own page

munotes.in251

Chapter Eighty

Three Ways to Build One, and What Each Costs

Syllabus topic Module 2, "Priority Queues & Heaps: Priority Queue ADT, Advantages and Disadvantages"

In one line

An unsorted list inserts in O(1) and removes in O(n), a sorted list does the reverse, and a heap does both in O(log n), which is why the heap wins when both operations are used.

The three implementations

Unsorted list. Append on insert. Search the whole collection on remove.

Sorted list. Find the position on insert, keeping the collection ordered. Take the end on remove.

Heap. Keep a partial order: enough to know the best item is at the top, not enough to be sorted. Chapters 81 to 85.

All three, built and counted

import heapq


class UnsortedPQ:
    """Append on insert; search on remove."""

    def __init__(self):
        self.items = []
        self.work = 0

    def insert(self, priority, item):
        self.items.append((priority, item))
        self.work += 1

    def remove_best(self):
        if not self.items:
            raise IndexError("underflow")
        best = 0
        for i in range(1, len(self.items)):
            self.work += 1
            if self.items[i][0] < self.items[best][0]:
                best = i
        self.work += 1
        return self.items.pop(best)


class SortedPQ:
    """Insert in order; remove from the end."""

    def __init__(self):
        self.items = []            # descending, so the best is last
        self.work = 0

    def insert(self, priority, item):
        position = len(self.items)
        while position > 0 and self.items[position - 1][0] < priority:
            position -= 1
            self.work += 1
        self.work += 1
        self.items.insert(position, (priority, item))

    def remove_best(self):
        if not self.items:
            raise IndexError("underflow")
        self.work += 1
        return self.items.pop()


class HeapPQ:
    """A heap, via heapq. Work is counted as the number of comparisons,
    which for a heap is proportional to log n per operation."""

    def __init__(self):
        self.items = []
        self.work = 0

    def insert(self, priority, item):
        heapq.heappush(self.items, (priority, item))
        self.work += max(1, len(self.items).bit_length())

    def remove_best(self):
        if not self.items:
            raise IndexError("underflow")
        self.work += max(1, len(self.items).bit_length())
        return heapq.heappop(self.items)


import random

random.seed(4)
print("%7s | %-26s | %-26s | %s"
      % ("n", "unsorted list", "sorted list", "heap"))
print("%7s | %12s %13s | %12s %13s | %12s %12s"
      % ("", "insert all", "remove all", "insert all", "remove all",
         "insert all", "remove all"))

for n in (500, 1000, 2000, 4000):
    priorities = [random.randrange(1_000_000) for _ in range(n)]
    row = []
    for cls in (UnsortedPQ, SortedPQ, HeapPQ):
        pq = cls()
        for i, p in enumerate(priorities):
            pq.insert(p, i)
        insert_work = pq.work
        pq.work = 0
        while True:
            try:
                pq.remove_best()
            except IndexError:
                break
        row.extend([insert_work, pq.work])
    print("%7d | %12d %13d | %12d %13d | %12d %12d" % (n, *row))
      n | unsorted list              | sorted list                | heap
        |   insert all    remove all |   insert all    remove all |   insert all   remove all
    500 |          500        125250 |        65837           500 |         3998         3998
   1000 |         1000        500500 |       240102          1000 |         8987         8987
   2000 |         2000       2001000 |       979989          2000 |        19964        19964
   4000 |         4000       8002000 |      3990542          4000 |        43917        43917

Read it column by column.

The unsorted list inserts n items in n units of work, perfectly cheap, and then needs 8 million units to remove 4,000 of them. Doubling n quadruples the removal work: O(n squared) overall.

munotes.in252

Three Ways to Build One, and What Each Costs

The sorted list is the mirror: 3,990,542 to insert and 4,000 to remove. It is about half the unsorted list's figure because an insertion stops as soon as it finds its place, so on average it scans half the collection rather than all of it. The growth is the same shape: doubling n quadruples it.

The heap does both in about 44,000, which is n log n. It is beaten by the unsorted list on insertion and by the sorted list on removal, and it beats both on the total: at n = 4,000 the heap's whole run is about 88,000 units against the unsorted list's 8 million, a factor of about 91.

That is the whole argument for the heap, and it is a measurement rather than a claim.

The table

Unsorted listSorted listHeap
insertO(1)O(n)O(log n)
remove_bestO(n)O(1)O(log n)
peek_bestO(n)O(1)O(1)
n inserts and n removesO(n squared)O(n squared)O(n log n)
build from n items at onceO(n)O(n log n)O(n), chapter 85
memoryn itemsn itemsn items

The peek_best row is worth noticing: the heap matches the sorted list there, because the best item is always at the top of a heap. That is what makes the heap usable in Dijkstra's algorithm, which peeks constantly.

When each is actually the right choice

Unsorted list. When there are very few items, or when almost everything is inserted and almost nothing removed. For fewer than about ten items it is genuinely fastest, because the constant factors beat the logarithm.

Sorted list. When removals vastly outnumber insertions, or when the whole collection must also be readable in order, which the heap cannot do.

Heap. Everything else, which is nearly everything.

The advantages and disadvantages of the heap, as MU asks

Advantages. Insert and remove both O(log n); peek is O(1); it can be built from n items in O(n); it needs no extra memory beyond the items, because it lives in an array with no pointers (chapter 82).

Disadvantages. It is not sorted, so it cannot produce the items in order without emptying itself; searching for an arbitrary item is O(n), because only the top is known; and decrease_key needs the item's position to be tracked separately, which is real work in Dijkstra's algorithm.

Quick revision

  • Three implementations: unsorted list, sorted list, heap.
  • Unsorted: O(1) insert, O(n) remove. Sorted: O(n) insert, O(1) remove. Heap: O(log n) both, O(1) peek.
  • Measured at n = 4,000: unsorted needed 8,002,000 units to remove everything, sorted needed 3,990,542 to
munotes.in253

Three Ways to Build One, and What Each Costs

insert everything, and the heap did both phases in 43,917 each.

  • Over n inserts and n removes, the lists are O(n squared) and the heap is O(n log n).
  • Use an unsorted list for very few items; a sorted list when removals dominate or order is needed;

otherwise a heap.

  • Heap advantages: O(log n) both ways, O(1) peek, O(n) build, no pointers. Disadvantages: not sorted,

O(n) to find an arbitrary item, and decrease_key needs positions tracked.

Test yourself

1. Give the insert and remove costs for the three implementations. Unsorted list: O(1) insert, O(n) remove. Sorted list: O(n) insert, O(1) remove. Heap: O(log n) for both.

2. Which is fastest at inserting, and which at removing? The unsorted list at inserting, the sorted list at removing. The heap is beaten by each at one of them and beats both on the total.

3. Give the measured totals at n = 4,000. The unsorted list used 8,002,000 units to remove everything; the sorted list used 3,990,542 to insert everything; the heap used 43,917 for each phase.

4. What is the cost of n inserts followed by n removes, for each? O(n squared) for both lists, and O(n log n) for the heap.

5. Why does the heap match the sorted list on peek? Because the best item is always at the top of the heap, so reading it is O(1) without any search.

6. Give two disadvantages of a heap. It is not sorted, so it cannot list the items in order without being emptied; and finding an arbitrary item is O(n), since only the top is known, which also makes decrease_key need separate position tracking.

Contents This chapter on its own page

munotes.in254

Chapter Eighty-One

The Heap: Shape and Order

Syllabus topic Module 2, "Priority Queues & Heaps: Heaps"

In one line

A heap is a complete binary tree in which every node is at least as good as its children, which is enough to put the best item at the root and loose enough to be maintained in O(log n).

The two properties

A heap satisfies both of these, and neither alone is enough.

1. The shape property: it is a complete binary tree. Every level is full except possibly the last, which fills from the left with no gaps. Chapter 56's exact word.

2. The order property (the heap property): every node is at least as good as both its children. For a min-heap, every node is less than or equal to its children. For a max-heap, greater than or equal.

What the order property does NOT say

This is where students lose marks.

It says nothing about siblings. Two children of the same node are in no particular order relative to each other.

It says nothing about cousins. A node deep on the left may be smaller than a node high on the right.

A heap is therefore not sorted, and its inorder traversal means nothing. Only the path from any node up to the root is ordered.

import heapq

values = [15, 3, 17, 9, 11, 25, 30, 12, 5]
heap = list(values)
heapq.heapify(heap)

print("values  :", values)
print("as a heap:", heap)
print("sorted  :", sorted(values))
print()
print("the heap array is NOT sorted:", heap != sorted(values))
print("but the smallest value IS at the front:", heap[0] == min(values))
print()
print("reading the heap level by level:")
level, index = 0, 0
while index < len(heap):
    count = 2 ** level
    print("   level %d: %s" % (level, heap[index:index + count]))
    index += count
    level += 1
print()
print("note level 1: %s. the two children of the root are in no order"
      % (heap[1:3],))
print("relative to each other, and that is allowed.")
values  : [15, 3, 17, 9, 11, 25, 30, 12, 5]
as a heap: [3, 5, 17, 9, 11, 25, 30, 12, 15]
sorted  : [3, 5, 9, 11, 12, 15, 17, 25, 30]

the heap array is NOT sorted: True
but the smallest value IS at the front: True

reading the heap level by level:
   level 0: [3]
   level 1: [5, 17]
   level 2: [9, 11, 25, 30]
   level 3: [12, 15]

note level 1: [5, 17]. the two children of the root are in no order
relative to each other, and that is allowed.

The array is not sorted and the minimum is at the front. That is exactly as much order as a heap promises, and it is exactly as much as a priority queue needs.

Why "complete" and not something else

Chapter 56 separated three words, and the choice here is deliberate.

munotes.in255

The Heap: Shape and Order

Not "perfect". A perfect tree has 1, 3, 7, 15 nodes only, so a heap could never hold 10 items.

Not "full". A full tree may have a gap in the middle of a level, which breaks the array arithmetic of chapter 82.

Complete is exactly right. It admits any number of items, there is exactly one complete shape for each size (chapter 56 measured that), and it has no gaps, so the tree can live in an array with no pointers.

Why this much order and no more

The alternative would be to keep the tree fully sorted, which is a binary search tree. Why not?

Because a search tree answers a question nobody asked here. A priority queue needs one thing: the best item. A search tree maintains enough order to find any item, and that extra order costs more to maintain, and would rule out the array representation.

The heap keeps exactly the order needed for "what is best" and no more. That is why insertion and removal are O(log n) with small constants and why the structure has no pointers at all.

Binary search treeHeap
Order maintainedtotal: every node against every otherpartial: each node against its children
Find an arbitrary valueO(log n)O(n)
Find the best valueO(log n)O(1)
Shapeany, and must be balanced deliberatelyalways complete, automatically
Representationnodes with two pointersan array, no pointers
In sorted orderinorder walk, freenot possible without emptying it

Both properties, checked

import heapq
import random


def shape_property(heap):
    """A list IS complete by construction: there are no gaps in a list.
    What must be checked is that every index below the length is occupied."""
    return all(item is not None for item in heap)


def order_property(heap):
    """Every node is <= its children, for a min-heap."""
    for i in range(len(heap)):
        for child in (2 * i + 1, 2 * i + 2):
            if child < len(heap) and heap[i] > heap[child]:
                return False, (i, child)
    return True, None


random.seed(8)
for trial in range(5):
    values = random.sample(range(1000), 12)
    heap = list(values)
    heapq.heapify(heap)
    ok, where = order_property(heap)
    print("trial %d: shape %-5s order %-5s root is the minimum: %s"
          % (trial + 1, shape_property(heap), ok, heap[0] == min(values)))

print()
not_a_heap = [3, 5, 17, 9, 11, 25, 30, 12, 2]
ok, where = order_property(not_a_heap)
print("a deliberately broken heap:", not_a_heap)
print("   order property holds:", ok)
print("   index %d holds %d and its child at %d holds %d"
      % (where[0], not_a_heap[where[0]], where[1], not_a_heap[where[1]]))
trial 1: shape True  order True  root is the minimum: True
trial 2: shape True  order True  root is the minimum: True
trial 3: shape True  order True  root is the minimum: True
trial 4: shape True  order True  root is the minimum: True
trial 5: shape True  order True  root is the minimum: True

a deliberately broken heap: [3, 5, 17, 9, 11, 25, 30, 12, 2]
   order property holds: False
   index 3 holds 9 and its child at 8 holds 2
munotes.in256

The Heap: Shape and Order

Five random heaps pass both properties, and a deliberately broken one is caught with the exact pair of indices that violate it. A checker that never reports a failure is not a checker, which is why the broken case is included.

Quick revision

  • A heap has two properties: the shape property (it is a complete binary tree) and the order property

(every node is at least as good as its children).

  • Min-heap: every node is less than or equal to its children. Max-heap: greater than or equal.
  • The order property says nothing about siblings or cousins, so a heap is NOT sorted and its inorder

traversal is meaningless. Only root-to-node paths are ordered.

  • Complete is the exact word: perfect would restrict the size to 1, 3, 7, 15; full would allow a gap and

break the array arithmetic.

  • A heap keeps exactly the order needed to answer "what is best" and no more, which is why it is cheaper

than a search tree and needs no pointers.

  • Finding the best is O(1); finding an arbitrary value is O(n).

Test yourself

1. State the two properties of a heap. The shape property: it is a complete binary tree. The order property: every node is at least as good as both its children, meaning smaller for a min-heap and larger for a max-heap.

2. What does the order property say about two siblings? Nothing at all. Siblings are unordered relative to each other, as are cousins, which is why a heap is not sorted.

3. Why is "complete" the right shape requirement rather than "perfect" or "full"? Perfect would allow only sizes 1, 3, 7, 15 and so on. Full would permit a gap in the middle of a level, breaking the array index arithmetic. Complete admits any size, has exactly one shape per size, and has no gaps.

4. Where is the best item, and what does it cost to find? At the root, index 0 of the array, and it costs O(1).

5. Give two things a binary search tree can do that a heap cannot. Find an arbitrary value in O(log n) rather than O(n), and produce the items in sorted order without being emptied.

6. Why does a heap deliberately keep less order than a search tree? Because a priority queue only ever asks for the best item. Maintaining total order would cost more and would prevent the array representation, for an ability nothing here uses.

Contents This chapter on its own page

munotes.in257

Chapter Eighty-Two

The Array That Holds a Heap

Syllabus topic Module 2, "Priority Queues & Heaps: Heaps"

In one line

Because a heap is always complete there are no gaps, so it fits an array exactly, with children at 2i + 1 and 2i + 2 and no pointers at all.

The representation

Read the heap level by level, left to right, and write the values into an array in that order. Then:

root = index 0

left child of i = 2i + 1

right child of i = 2i + 2

parent of i = (i - 1) / 2, integer division

last node = index n - 1

Chapter 58 gave that arithmetic. What makes it usable here, and unusable for a general binary tree, is that a complete tree has no gaps, so every index from 0 to n - 1 holds a real node.

import heapq

heap = [3, 5, 17, 9, 11, 25, 30, 12, 15]

print("the array:", heap)
print()
print("%6s %7s %10s %10s %9s" % ("index", "value", "left", "right", "parent"))
for i, value in enumerate(heap):
    left, right = 2 * i + 1, 2 * i + 2
    left_value = heap[left] if left < len(heap) else "-"
    right_value = heap[right] if right < len(heap) else "-"
    parent_value = heap[(i - 1) // 2] if i > 0 else "-"
    print("%6d %7d %10s %10s %9s" % (i, value, left_value, right_value, parent_value))

print()
print("no index is empty, because the tree is complete:",
      all(v is not None for v in heap))
print()
print("drawn as a tree, level by level:")
level, index = 0, 0
while index < len(heap):
    count = 2 ** level
    row = heap[index:index + count]
    print("   %s%s" % (" " * (16 - 2 * level), "   ".join(str(v) for v in row)))
    index += count
    level += 1
the array: [3, 5, 17, 9, 11, 25, 30, 12, 15]

 index   value       left      right    parent
     0       3          5         17         -
     1       5          9         11         3
     2      17         25         30         3
     3       9         12         15         5
     4      11          -          -         5
     5      25          -          -        17
     6      30          -          -        17
     7      12          -          -         9
     8      15          -          -         9

no index is empty, because the tree is complete: True

drawn as a tree, level by level:
                   3
                 5   17
               9   11   25   30
             12   15

Every parent and child relation is computed, not stored. The tree exists entirely in the arithmetic.

What this buys

No pointers. A linked binary tree of n nodes carries 2n pointers, mostly null (chapter 70 counted n + 1 of them). A heap carries none. For a million integers that is 16 MB of pointers not spent.

The parent for free. Chapter 58 noted this and here it is used: sift up in chapter 84 walks from a node to the root, and does it with (i - 1) // 2 and no stored parent pointers, which a linked tree cannot do without an extra field.

munotes.in258

The Array That Holds a Heap

Perfect cache behaviour. The nodes are contiguous, so walking a path reads neighbouring memory. This is the structure chapter 22's warning does not apply to.

The shape maintains itself. Adding an item means appending to the array, which is exactly adding the next node of a complete tree. Removing the last node means shortening the array. The shape property can never be broken, because an array has no gaps. Only the order property needs repairing, and that is what chapter 84 does.

That last point is the elegant part: half the heap's invariant is enforced by the representation rather than by code.

The two operations, in outline

Both are the same two steps in opposite orders, and chapter 84 implements them.

Insert. Append the new value at the end of the array, which keeps the shape. Then move it up until the order property holds. That is at most the height, O(log n).

Remove the best. Take index 0, which is the answer. Move the last element into index 0, which keeps the shape, and shorten the array. Then move it down until the order property holds. Again O(log n).

In both, the shape is fixed first and cheaply, and then the order is repaired.

The off-by-one that ruins it

The 1-based convention of chapter 58 is common in textbooks:

Conventionleftrightparent
root at 02i + 12i + 2(i - 1) / 2
root at 12i2i + 1i / 2

Mixing them is the standard error and it does not fail loudly: the structure continues to work for small indices and goes wrong deeper in. State the convention and use it everywhere, including in the termination conditions of the loops.

def children_zero(i):
    return 2 * i + 1, 2 * i + 2


def parent_zero(i):
    return (i - 1) // 2


def children_one(i):
    return 2 * i, 2 * i + 1


def parent_one(i):
    return i // 2


n = 20
zero_ok = all(parent_zero(c) == i for i in range(n) for c in children_zero(i))
one_ok = all(parent_one(c) == i for i in range(1, n) for c in children_one(i))
print("0-based: parent(child(i)) == i for every i:", zero_ok)
print("1-based: parent(child(i)) == i for every i:", one_ok)
print()
print("mixing them, 0-based children with 1-based parent:")
for i in (0, 1, 2, 3, 7):
    left, _ = children_zero(i)
    print("   node %2d, left child %2d, 1-based parent of that child says %2d  %s"
          % (i, left, parent_one(left),
             "correct" if parent_one(left) == i else "WRONG"))
munotes.in259

The Array That Holds a Heap

0-based: parent(child(i)) == i for every i: True
1-based: parent(child(i)) == i for every i: True

mixing them, 0-based children with 1-based parent:
   node  0, left child  1, 1-based parent of that child says  0  correct
   node  1, left child  3, 1-based parent of that child says  1  correct
   node  2, left child  5, 1-based parent of that child says  2  correct
   node  3, left child  7, 1-based parent of that child says  3  correct
   node  7, left child 15, 1-based parent of that child says  7  correct

Interesting, and worth reading carefully: for a left child the mixed arithmetic happens to agree, because (2i + 1) // 2 equals i. The error only shows on right children, which is precisely why it survives casual testing and then fails on real data.

Quick revision

  • A heap is stored as an array read level by level, left to right.
  • Root at 0; children of i at 2i + 1 and 2i + 2; parent of i at (i - 1) / 2.
  • It works because a complete tree has no gaps, so every index below n holds a node.
  • No pointers at all, the parent is free, and the memory is contiguous so the cache works.
  • The shape property is enforced by the representation: appending adds the next node of a complete tree

and an array cannot have a gap. Only the order needs repairing.

  • Insert: append, then sift up. Remove: take index 0, move the last element there, shorten, then sift

down. Both O(log n).

  • Mixing the 0-based and 1-based conventions agrees by accident on left children and fails on right ones.

Test yourself

1. Give the index arithmetic for a heap with the root at 0. Children of i at 2i + 1 and 2i + 2; parent of i at (i - 1) divided by 2 with integer division.

2. Why does this representation work for a heap when chapter 58 showed it wasteful for a general tree? Because a heap is always complete, so there are no gaps and every index from 0 to n - 1 holds a real node. A general tree may be degenerate and would need an exponentially large array.

3. Name three things the array representation buys. No pointers at all; the parent of a node for free from arithmetic; and contiguous memory, so the cache works well.

4. Which of the heap's two properties is enforced by the representation, and how? The shape property. Appending to the array adds exactly the next node of a complete tree, and an array cannot contain a gap, so the shape can never be broken. Only the order property needs repairing.

5. Outline insertion and removal. Insert: append at the end, which keeps the shape, then sift the value up until the order holds. Remove: take index 0 as the answer, move the last element to index 0, shorten the array, then sift it down.

munotes.in260

The Array That Holds a Heap

6. Why does mixing the 0-based and 1-based conventions survive casual testing? Because for a left child the mixed arithmetic happens to agree: (2i + 1) divided by 2 is i. The error appears only on right children.

Contents This chapter on its own page

munotes.in261

Chapter Eighty-Three

Min-Heap and Max-Heap

Syllabus topic Module 2, "Priority Queues & Heaps: types of heaps"

In one line

A min-heap keeps the smallest value at the root and a max-heap the largest, and they are the same structure with one comparison reversed.

The two types

Min-heapMax-heap
Order propertyevery node is less than or equal to its childrengreater than or equal
Root holdsthe minimumthe maximum
Used forcosts, distances, deadlines, Huffmanscores, importance, heap sort
Finding the other extremeO(n)O(n)

The last row is worth noting at once: a min-heap cannot find the maximum quickly. The largest value is somewhere among the leaves, and the structure says nothing about which. Finding it means scanning, and that is O(n). A structure needing both extremes needs two heaps, or a different structure.

One implementation, both types

class Heap:
    """One heap. `better(a, b)` decides which of two values belongs higher.

    For a min-heap, better is 'a < b'. For a max-heap, 'a > b'. Nothing else
    differs, which is the honest statement of the difference between them."""

    def __init__(self, better):
        self.better = better
        self.items = []

    def size(self):
        return len(self.items)

    def peek(self):
        if not self.items:
            raise IndexError("peek at an empty heap")
        return self.items[0]

    def insert(self, value):
        self.items.append(value)
        i = len(self.items) - 1
        while i > 0:
            parent = (i - 1) // 2
            if self.better(self.items[i], self.items[parent]):
                self.items[i], self.items[parent] = self.items[parent], self.items[i]
                i = parent
            else:
                break

    def remove(self):
        if not self.items:
            raise IndexError("remove from an empty heap")
        best = self.items[0]
        last = self.items.pop()
        if self.items:
            self.items[0] = last
            i, n = 0, len(self.items)
            while True:
                left, right, winner = 2 * i + 1, 2 * i + 2, i
                if left < n and self.better(self.items[left], self.items[winner]):
                    winner = left
                if right < n and self.better(self.items[right], self.items[winner]):
                    winner = right
                if winner == i:
                    break
                self.items[i], self.items[winner] = self.items[winner], self.items[i]
                i = winner
        return best

    def holds(self):
        """The order property, checked."""
        n = len(self.items)
        for i in range(n):
            for child in (2 * i + 1, 2 * i + 2):
                if child < n and self.better(self.items[child], self.items[i]):
                    return False
        return True


values = [15, 3, 17, 9, 11, 25, 30, 12, 5]

for name, better in (("min-heap", lambda a, b: a < b),
                     ("max-heap", lambda a, b: a > b)):
    heap = Heap(better)
    for value in values:
        heap.insert(value)
    print("%s" % name)
    print("   array        :", heap.items)
    print("   root         :", heap.peek())
    print("   order holds  :", heap.holds())
    drained = []
    while heap.size():
        drained.append(heap.remove())
    print("   drained      :", drained)

print()
print("values      :", values)
print("sorted      :", sorted(values))
print("reverse     :", sorted(values, reverse=True))
min-heap
   array        : [3, 5, 17, 9, 11, 25, 30, 15, 12]
   root         : 3
   order holds  : True
   drained      : [3, 5, 9, 11, 12, 15, 17, 25, 30]
max-heap
   array        : [30, 12, 25, 11, 9, 15, 17, 3, 5]
   root         : 30
   order holds  : True
   drained      : [30, 25, 17, 15, 12, 11, 9, 5, 3]

values      : [15, 3, 17, 9, 11, 25, 30, 12, 5]
sorted      : [3, 5, 9, 11, 12, 15, 17, 25, 30]
reverse     : [30, 25, 17, 15, 12, 11, 9, 5, 3]
munotes.in262

Min-Heap and Max-Heap

One class, two heaps, differing by a single lambda. The min-heap drains in ascending order and the max-heap in descending, and both match Python's sorted exactly.

That draining is heap sort, and it is worth naming: inserting n items and removing them all gives them in order, at a cost of O(n log n). It is not on this syllabus as a topic of its own, but it falls out of the structure and an examiner may ask what it is called.

Converting between them

Three ways, and an examiner may ask for one:

Negate the keys. Store -x in a min-heap and it behaves as a max-heap for x. This is how Python's heapq, which only offers a min-heap, is used as a max-heap.

Reverse the comparison, which is what this chapter's class does.

Rebuild. Take the values out and build the other type, O(n) by chapter 85.

import heapq

values = [15, 3, 17, 9, 11, 25, 30]

min_heap = list(values)
heapq.heapify(min_heap)

max_heap = [-v for v in values]
heapq.heapify(max_heap)

print("heapq offers only a min-heap.")
print("   smallest, directly   :", min_heap[0])
print("   largest, by negation :", -max_heap[0])
print()
print("draining the negated heap and flipping the sign:")
out = []
while max_heap:
    out.append(-heapq.heappop(max_heap))
print("   ", out)
print("   which is descending:", out == sorted(values, reverse=True))
heapq offers only a min-heap.
   smallest, directly   : 3
   largest, by negation : 30

draining the negated heap and flipping the sign:
    [30, 25, 17, 15, 11, 9, 3]
   which is descending: True

Other types of heap, named

MU's label says "types of heaps", and if an examiner wants more than min and max, these are the standard answers:

Binary heap. What this chapter builds: each node has at most two children. The default. d-ary heap. Each node has d children. Shallower, so insertion is faster and removal slower. Binomial heap and Fibonacci heap. Support merging two heaps efficiently, which a binary heap cannot do better than O(n). The Fibonacci heap gives O(1) amortised decrease-key, which improves Dijkstra's algorithm in theory.

All but the binary heap are outside this paper. Knowing the names and the one-line reason is enough.

Quick revision

  • Min-heap: every node is less than or equal to its children, and the root is the minimum. Max-heap: the

reverse.

  • They are the same structure with one comparison reversed, which is why one implementation with a
munotes.in263

Min-Heap and Max-Heap

comparison parameter gives both.

  • A min-heap cannot find the maximum in better than O(n), and vice versa.
  • Draining a heap gives the values in order: that is heap sort, O(n log n).
  • To get a max-heap from a min-heap library, negate the keys.
  • Other types: binary (the default), d-ary, binomial and Fibonacci, the last two for efficient merging.

Test yourself

1. Give the order property of each type and say what the root holds. Min-heap: every node is less than or equal to its children, and the root is the minimum. Max-heap: every node is greater than or equal to its children, and the root is the maximum.

2. How much work is it to find the largest value in a min-heap? O(n). The largest is somewhere among the leaves and the structure gives no clue which, so it must be scanned for.

3. How do the two implementations differ? By one comparison. Writing the heap with the comparison passed in gives both types from one piece of code.

4. What is produced by inserting n values and then removing them all, and what is it called? The values in sorted order, ascending from a min-heap and descending from a max-heap. It is heap sort, and it costs O(n log n).

5. Python's heapq offers only a min-heap. How is a max-heap obtained from it? By negating the keys on the way in and negating them again on the way out.

6. Name two other types of heap and what they are for. Binomial and Fibonacci heaps, which support merging two heaps efficiently; a d-ary heap, which is shallower and trades faster insertion for slower removal.

Contents This chapter on its own page

munotes.in264

Chapter Eighty-Four

Heapify: Sifting Up and Sifting Down

Syllabus topic Module 2, "Priority Queues & Heaps: Heapifying the element"

In one line

Sifting up moves a value towards the root while it is better than its parent, and sifting down moves it away from the root while a child is better, and the two together are all a heap ever does.

The two directions

A heap is repaired in exactly two situations, and each has its own direction.

One value is too good for its position, and everything else is fine. That happens after an insertion at the end. The value may be smaller than its parent, so it moves up.

One value is too bad for its position, and everything else is fine. That happens after a removal, when the last element is moved to the root. It may be larger than its children, so it moves down.

In both cases exactly one value is out of place and the path it travels is at most the height.

Sift up

sift_up(i):

while i > 0 and heap[i] is better than heap[parent(i)]:

swap them

i = parent(i)

Sift down

sift_down(i):

loop:

best = i

if left child exists and is better than heap[best]: best = left

if right child exists and is better than heap[best]: best = right

if best == i: stop

swap heap[i] and heap[best]

i = best

The difference worth noticing: sift up compares against one node, the parent. Sift down compares against two, and must swap with the better of the two children, not merely with one that is better than the current value. Swapping with the wrong child breaks the heap at the other one.

Both, traced

def parent(i):
    return (i - 1) // 2


def sift_up(heap, i, trace):
    while i > 0 and heap[i] < heap[parent(i)]:
        p = parent(i)
        heap[i], heap[p] = heap[p], heap[i]
        trace.append("swap index %d with parent %d -> %s" % (i, p, list(heap)))
        i = p
    return i


def sift_down(heap, i, trace):
    n = len(heap)
    while True:
        left, right, best = 2 * i + 1, 2 * i + 2, i
        if left < n and heap[left] < heap[best]:
            best = left
        if right < n and heap[right] < heap[best]:
            best = right
        if best == i:
            return i
        heap[i], heap[best] = heap[best], heap[i]
        trace.append("swap index %d with child %d -> %s" % (i, best, list(heap)))
        i = best


def holds(heap):
    n = len(heap)
    return all(heap[i] <= heap[c]
               for i in range(n) for c in (2 * i + 1, 2 * i + 2) if c < n)


print("INSERT 2 into a min-heap")
heap = [3, 5, 17, 9, 11, 25, 30]
print("   before      :", heap, "| heap:", holds(heap))
heap.append(2)
print("   appended    :", heap, "| heap:", holds(heap), "  (shape is fine, order is not)")
trace = []
sift_up(heap, len(heap) - 1, trace)
for line in trace:
    print("   ", line)
print("   after       :", heap, "| heap:", holds(heap))

print()
print("REMOVE the best from that heap")
best = heap[0]
last = heap.pop()
heap[0] = last
print("   took %d, moved %d to the root: %s | heap: %s"
      % (best, last, heap, holds(heap)))
trace = []
sift_down(heap, 0, trace)
for line in trace:
    print("   ", line)
print("   after       :", heap, "| heap:", holds(heap))
munotes.in265

Heapify: Sifting Up and Sifting Down

INSERT 2 into a min-heap
   before      : [3, 5, 17, 9, 11, 25, 30] | heap: True
   appended    : [3, 5, 17, 9, 11, 25, 30, 2] | heap: False   (shape is fine, order is not)
    swap index 7 with parent 3 -> [3, 5, 17, 2, 11, 25, 30, 9]
    swap index 3 with parent 1 -> [3, 2, 17, 5, 11, 25, 30, 9]
    swap index 1 with parent 0 -> [2, 3, 17, 5, 11, 25, 30, 9]
   after       : [2, 3, 17, 5, 11, 25, 30, 9] | heap: True

REMOVE the best from that heap
   took 2, moved 9 to the root: [9, 3, 17, 5, 11, 25, 30] | heap: False
    swap index 0 with child 1 -> [3, 9, 17, 5, 11, 25, 30]
    swap index 1 with child 3 -> [3, 5, 17, 9, 11, 25, 30]
   after       : [3, 5, 17, 9, 11, 25, 30] | heap: True

Follow the insertion. 2 was appended at index 7, then rose through indices 3 and 1 to reach the root, three swaps for a tree of height 3. The order property was broken immediately after the append, exactly as chapter 82 said, and the shape never was.

Why sifting down must take the better child

def holds(heap):
    n = len(heap)
    return all(heap[i] <= heap[c]
               for i in range(n) for c in (2 * i + 1, 2 * i + 2) if c < n)


def sift_down_correct(heap, i):
    n = len(heap)
    while True:
        left, right, best = 2 * i + 1, 2 * i + 2, i
        if left < n and heap[left] < heap[best]:
            best = left
        if right < n and heap[right] < heap[best]:
            best = right
        if best == i:
            return
        heap[i], heap[best] = heap[best], heap[i]
        i = best


def sift_down_wrong(heap, i):
    """Swaps with the FIRST child that is better, not the BETTER child."""
    n = len(heap)
    while True:
        left, right = 2 * i + 1, 2 * i + 2
        if left < n and heap[left] < heap[i]:
            heap[i], heap[left] = heap[left], heap[i]
            i = left
        elif right < n and heap[right] < heap[i]:
            heap[i], heap[right] = heap[right], heap[i]
            i = right
        else:
            return


start = [20, 5, 3, 8, 9, 7, 6]
for name, routine in (("correct", sift_down_correct), ("wrong  ", sift_down_wrong)):
    heap = list(start)
    routine(heap, 0)
    print("%s: %s | a valid heap: %s" % (name, heap, holds(heap)))

print()
print("the wrong version swapped 20 with its LEFT child 5, but 3 was smaller.")
print("5 then sat above 3, which breaks the order property at the root.")
munotes.in266

Heapify: Sifting Up and Sifting Down

correct: [3, 5, 6, 8, 9, 7, 20] | a valid heap: True
wrong  : [5, 8, 3, 20, 9, 7, 6] | a valid heap: False

the wrong version swapped 20 with its LEFT child 5, but 3 was smaller.
5 then sat above 3, which breaks the order property at the root.

The wrong version produced an array that is not a heap, and it did so without any error. 5 ended up at the root while 3 sits below it. That is why the rule is "swap with the better of the two children", not "swap with a child that is better".

The costs

Cost
sift upat most the height, so O(log n)
sift downat most the height, so O(log n)
insertappend O(1), then sift up: O(log n)
remove besttake the root, move the last up, shorten, then sift down: O(log n)
peek bestO(1)

Sift down does about twice the comparisons of sift up per level, because it examines two children rather than one parent. Both are O(log n) and the constant difference is why building a heap is done with sift down in chapter 85 rather than with repeated insertion.

Quick revision

  • Two repairs, two directions. Sift up after an insertion at the end; sift down after the last element is

moved to the root.

  • Sift up: while better than the parent, swap and move up. One comparison per level.
  • Sift down: find the better of the two children; if it is better than the current value, swap and move

down. Two comparisons per level.

  • Swapping with the first better child rather than the better child produces an array that is not a heap,

silently.

  • The shape is never broken by either operation; only the order is repaired.
  • Both are O(log n), so insert and remove are O(log n) and peek is O(1).

Test yourself

1. When is each direction used? Sift up after inserting at the end of the array. Sift down after removing the root and moving the last element into its place.

2. Write sift up. While the index is not 0 and the value is better than its parent, swap the two and move to the parent's index.

3. Write sift down, and state the rule that is easy to get wrong. Find the better of the two children; if it is better than the current value, swap and continue from there. The rule is to swap with the BETTER child, not merely with a child that is better.

munotes.in267

Heapify: Sifting Up and Sifting Down

4. What happens if you swap with the first better child instead? The result is not a heap, and no error is raised. In the run, 5 ended at the root while 3 was below it.

5. Which operation does more comparisons per level, and why? Sift down, because it must examine both children, where sift up examines only the parent.

6. Give the costs of insert, remove and peek, with their parts. Insert: append O(1) then sift up O(log n), so O(log n). Remove: take the root, move the last element there, shorten, then sift down O(log n), so O(log n). Peek: O(1).

Contents This chapter on its own page

munotes.in268

Chapter Eighty-Five

Building a Heap, and Why It Is Linear

Syllabus topic Module 2, "Priority Queues & Heaps: Heapifying the element"

In one line

Sifting down from the last internal node back to the root builds a heap in linear time, because almost every node is near the bottom and has almost nowhere to sift.

The obvious way, and its cost

Insert the items one at a time. Each insertion is O(log n), so n of them is O(n log n).

That is correct and it is not the best available.

The better way

Put all n items into the array in any order. The shape property already holds, because an array has no gaps.

Now repair the order, from the bottom up:

for i from the last internal node down to 0:

sift_down(i)

The last internal node is at index n // 2 - 1: everything after it is a leaf, and a leaf is already a heap of one item, so there is nothing to do for the second half of the array.

Working upwards is what makes it correct: when sift_down(i) runs, both of i's subtrees are already heaps, which is exactly the precondition sift down needs.

Both, counted

def sift_down(heap, i, n, counter):
    while True:
        left, right, best = 2 * i + 1, 2 * i + 2, i
        if left < n:
            counter[0] += 1
            if heap[left] < heap[best]:
                best = left
        if right < n:
            counter[0] += 1
            if heap[right] < heap[best]:
                best = right
        if best == i:
            return
        heap[i], heap[best] = heap[best], heap[i]
        i = best


def sift_up(heap, i, counter):
    while i > 0:
        parent = (i - 1) // 2
        counter[0] += 1
        if heap[i] < heap[parent]:
            heap[i], heap[parent] = heap[parent], heap[i]
            i = parent
        else:
            return


def build_by_insertion(values):
    heap, counter = [], [0]
    for value in values:
        heap.append(value)
        sift_up(heap, len(heap) - 1, counter)
    return heap, counter[0]


def build_bottom_up(values):
    heap, counter = list(values), [0]
    n = len(heap)
    for i in range(n // 2 - 1, -1, -1):
        sift_down(heap, i, n, counter)
    return heap, counter[0]


def holds(heap):
    n = len(heap)
    return all(heap[i] <= heap[c]
               for i in range(n) for c in (2 * i + 1, 2 * i + 2) if c < n)


import random

random.seed(12)
print("RANDOM input")
print("%8s | %14s | %14s | %8s | %s"
      % ("n", "by insertion", "bottom up", "ratio", "both valid"))
for n in (1000, 4000, 16000):
    values = [random.randrange(1_000_000) for _ in range(n)]
    a, insert_work = build_by_insertion(values)
    b, bottom_work = build_bottom_up(values)
    print("%8d | %14d | %14d | %8.2f | %s"
          % (n, insert_work, bottom_work, insert_work / bottom_work,
             holds(a) and holds(b)))

print()
# DESCENDING, not ascending. For a MIN-heap, an ascending value is always
# the largest so far and stays where it is appended: that is the BEST case and
# costs one comparison. The worst case is each new value being the new MINIMUM,
# which is descending input.
print("DESCENDING input, the worst case for insertion into a min-heap")
print("%8s | %14s | %14s | %8s | %s"
      % ("n", "by insertion", "bottom up", "ratio", "both valid"))
for n in (1000, 4000, 16000):
    values = list(range(n, 0, -1))
    a, insert_work = build_by_insertion(values)
    b, bottom_work = build_bottom_up(values)
    print("%8d | %14d | %14d | %8.2f | %s"
          % (n, insert_work, bottom_work, insert_work / bottom_work,
             holds(a) and holds(b)))

print()
print("comparisons per item:")
for label, make in (("random    ", lambda n: [random.randrange(1_000_000) for _ in range(n)]),
                    ("descending", lambda n: list(range(n, 0, -1)))):
    for n in (1000, 16000):
        values = make(n)
        _, insert_work = build_by_insertion(values)
        _, bottom_work = build_bottom_up(values)
        print("   %s n = %6d: by insertion %5.2f, bottom up %.2f"
              % (label, n, insert_work / n, bottom_work / n))
munotes.in269

Building a Heap, and Why It Is Linear

RANDOM input
       n |   by insertion |      bottom up |    ratio | both valid
    1000 |           2200 |           1844 |     1.19 | True
    4000 |           9090 |           7490 |     1.21 | True
   16000 |          36453 |          30069 |     1.21 | True

DESCENDING input, the worst case for insertion into a min-heap
       n |   by insertion |      bottom up |    ratio | both valid
    1000 |           7987 |           1982 |     4.03 | True
    4000 |          39917 |           7978 |     5.00 | True
   16000 |         191631 |          31974 |     5.99 | True

comparisons per item:
   random     n =   1000: by insertion  2.12, bottom up 1.86
   random     n =  16000: by insertion  2.28, bottom up 1.88
   descending n =   1000: by insertion  7.99, bottom up 1.98
   descending n =  16000: by insertion 11.98, bottom up 2.00

Read the two tables together, because the second is the one that matters and the first is the one that surprises.

On random input the two are almost the same, a ratio of about 1.2. Repeated insertion costs about 2.3 comparisons per item and the bottom-up build about 1.9, and neither grows with n.

That is an honest and slightly awkward result for the usual textbook claim. The reason is that inserting a random value into a heap almost never sifts far: the value is more likely than not to belong near the bottom, where it already is. On random data, repeated insertion is linear in practice even though it is n log n in the worst case.

On descending input the two separate completely. Every new value is the new minimum, so every insertion sifts all the way to the root, and the cost per item becomes the full height. That is the worst case, it is a perfectly ordinary input, and it is what the ratio column shows.

Note which direction is which: for a MIN-heap, ascending input is the BEST case, because each new value is the largest so far and stays exactly where it was appended. Getting that backwards is easy, and the first draft of this chapter did.

munotes.in270

Building a Heap, and Why It Is Linear

So the honest statement, and the one to give in an answer: the bottom-up build is O(n) always, and repeated insertion is O(n log n) in the worst case and close to linear on random data. The guarantee is what differs, not the typical behaviour.

Why it is linear

The argument is worth understanding, because it is counter-intuitive and examiners like it.

A node at height h can sift down at most h levels. So the total work is the sum, over all nodes, of each node's height.

Most nodes have almost no height. In a complete tree:

HeightNumber of nodes at mostWork each
0 (leaves)n/20
1n/41
2n/82
hn / 2^(h+1)h

So the total work is about:

n/4 x 1 + n/8 x 2 + n/16 x 3 + ...

= n x (1/4 + 2/8 + 3/16 + ...)

= n x 1

= O(n)

The series in brackets converges to 1, so the whole thing is about n. Half the nodes are leaves and do no work at all; only the single root can sift the full height.

Compare insertion, which has it exactly backwards: it inserts at the bottom and sifts up, so the nodes that can travel furthest are the many leaves, and the work really is n log n.

import math

print("%8s %12s %12s %14s" % ("height", "nodes at most", "work each", "total"))
n = 16384
total = 0
h = 0
while 2 ** (h + 1) <= n:
    nodes = n // (2 ** (h + 1))
    work = nodes * h
    total += work
    if h <= 6 or h >= 12:
        print("%8d %12d %12d %14d" % (h, nodes, h, work))
    h += 1

print()
print("n =", n)
print("total sift-down work, summed:", total)
print("which is %.2f times n" % (total / n))
print()
print("by comparison, n log2(n) would be %d, which is %.1f times larger"
      % (n * math.log2(n), (n * math.log2(n)) / total))
  height nodes at most    work each          total
       0         8192            0              0
       1         4096            1           4096
       2         2048            2           4096
       3         1024            3           3072
       4          512            4           2048
       5          256            5           1280
       6          128            6            768
      12            2           12             24
      13            1           13             13

n = 16384
total sift-down work, summed: 16369
which is 1.00 times n

by comparison, n log2(n) would be 229376, which is 14.0 times larger

The work peaks at heights 1 and 2 and then falls away. The nodes that could sift furthest are the ones there are almost none of.

munotes.in271

Building a Heap, and Why It Is Linear

The cost summary

Cost
Build by n insertionsO(n log n)
Build bottom upO(n)
Insert one item into an existing heapO(log n)
Remove the bestO(log n)

So: build all at once when you can, O(n); insert one at a time when you must, O(log n) each.

That is why Python's heapq.heapify exists beside heappush, and why Huffman's algorithm in chapter 76 puts all the leaves in at the start rather than inserting them one by one.

Quick revision

  • Building by repeated insertion is O(n log n).
  • Building bottom up is O(n): put everything in the array, then sift down from index n/2 - 1 to 0.
  • Going upwards is required for correctness: when sift_down(i) runs, both subtrees of i are already heaps.
  • Everything from index n/2 onwards is a leaf and needs no work.
  • Measured on RANDOM input the two are nearly equal, about 2.3 comparisons per item against 1.9: repeated

insertion is close to linear in practice because a random value rarely sifts far.

  • Measured on DESCENDING input they separate completely, because every new value is the new minimum and

sifts to the root. For a min-heap, ASCENDING input is the best case, not the worst.

  • Why: half the nodes are leaves with zero work, and the sum of n/2^(h+1) times h converges to n.
  • Insertion has it backwards, sifting the many leaves up through the whole height.

Test yourself

1. Give the two ways to build a heap and their costs. Repeated insertion, O(n log n); or bottom up, sifting down from the last internal node to the root, O(n).

2. Where is the last internal node, and why does everything after it need no work? At index n/2 - 1. Everything after it is a leaf, and a single node is already a heap.

3. Why must the bottom-up loop go upwards? Because sift_down(i) requires both of i's subtrees to be heaps already, and working upwards guarantees that they have been fixed first.

4. Give the measured comparisons per item for both methods, and on which input. On random input they are close: about 2.3 per item by insertion and 1.9 bottom up, neither growing with n. On descending input they separate, because every insertion then sifts to the root.

5. Explain why the bottom-up build is linear. A node of height h does at most h work, and a complete tree has about n/2^(h+1) nodes of height h. The sum of n/2^(h+1) times h over all h converges to about n, because half the nodes are leaves doing no work at all.

6. Why is repeated insertion n log n in the worst case, and why does random data hide it? Because insertion puts each item at the bottom and sifts up, so a value can travel the full height; with descending input into a min-heap every value does. Random data hides it because a random value usually belongs near the bottom and sifts barely at all. The bottom-up build is the opposite: the many leaves travel nowhere and only the few high nodes travel far, which is why it is linear in every case.

Contents This chapter on its own page

munotes.in272

Chapter Eighty-Six

Where Priority Queues Are Used

Syllabus topic Module 2, "Priority Queues & Heaps: Applications"

In one line

A priority queue is used wherever the next thing to do is decided by importance rather than by arrival, and three of this paper's own algorithms depend on one.

Inside this paper

Huffman coding, chapter 76. The algorithm repeatedly needs the two least frequent items. That is a min priority queue, and it is why the structure appeared before it had been built.

Dijkstra's algorithm, chapter 98. Repeatedly needs the unvisited vertex with the smallest known distance. A min priority queue, and the choice of implementation changes the algorithm's complexity.

Heap sort, chapter 83. Build a heap, then remove the best repeatedly. O(n log n) with no extra memory.

Outside it

Process scheduling. Chapter 48's limit, answered. An operating system keeps ready processes in a priority queue, so an interactive task can run ahead of a long batch job. With ageing, to prevent the starvation chapter 78 demonstrated.

Hospital triage. The practical names it. The most urgent patient is seen next, whatever the order of arrival.

Network routers. Packets are queued by class of service, so a voice call is forwarded ahead of a file download.

Event-driven simulation. This is the most interesting one and the least obvious. A simulation of anything (a bank, a factory, a network) holds a queue of future events ordered by the time they happen, repeatedly takes the earliest, and processes it, which may schedule more events. The priority is a timestamp and the structure is a min priority queue.

The A-star search used in route finding and games, which is Dijkstra with an added estimate.

Scheduling, run

import heapq

# (arrival, name, burst, priority) where priority 1 is most urgent
jobs = [(0, "batch report", 10, 5),
        (1, "user click", 1, 1),
        (2, "backup", 8, 5),
        (3, "user typing", 1, 1),
        (4, "index rebuild", 6, 5)]

# First come first served, chapter 48
clock, fcfs_waits = 0, {}
for arrival, name, burst, priority in jobs:
    start = max(clock, arrival)
    fcfs_waits[name] = start - arrival
    clock = start + burst

# Priority scheduling, with a priority queue
pending = sorted(jobs)
ready, clock, pq_waits, index = [], 0, {}, 0
while index < len(pending) or ready:
    while index < len(pending) and pending[index][0] <= clock:
        arrival, name, burst, priority = pending[index]
        heapq.heappush(ready, (priority, arrival, name, burst))
        index += 1
    if not ready:
        clock = pending[index][0]
        continue
    priority, arrival, name, burst = heapq.heappop(ready)
    pq_waits[name] = clock - arrival
    clock += burst

print("%-16s %9s %18s %16s" % ("job", "priority", "wait, first come", "wait, priority"))
for arrival, name, burst, priority in jobs:
    print("%-16s %9d %18d %16d"
          % (name, priority, fcfs_waits[name], pq_waits[name]))

interactive = [name for _, name, _, p in jobs if p == 1]
print()
print("the two interactive jobs waited %d and %d units under first come first served,"
      % tuple(fcfs_waits[n] for n in interactive))
print("and %d and %d under priority scheduling."
      % tuple(pq_waits[n] for n in interactive))
munotes.in273

Where Priority Queues Are Used

job               priority   wait, first come   wait, priority
batch report             5                  0                0
user click               1                  9                9
backup                   5                  9               10
user typing              1                 16                8
index rebuild            5                 16               16

the two interactive jobs waited 9 and 16 units under first come first served,
and 9 and 8 under priority scheduling.

The interactive jobs waited 9 and 16 units under first come first served, and 9 and 8 under priority scheduling. The improvement is real and smaller than one might expect, and the reason is worth reading off the table: the batch report was already running when the first user click arrived at time 1, and this scheduler does not interrupt it. The second interactive job gained, from 16 to 8; the first gained nothing, because nothing could be done for it.

That limit has a name: this is non-preemptive scheduling. A preemptive scheduler would stop the batch job mid-way, and operating systems do, which is outside this paper but worth the sentence.

Event-driven simulation, run

import heapq
import random

random.seed(2)

# A single counter. Customers arrive, queue, and are served.
events = []          # (time, sequence, kind, customer)
sequence = 0


def schedule(time, kind, customer):
    global sequence
    heapq.heappush(events, (time, sequence, kind, customer))
    sequence += 1


arrival_time = 0
for customer in range(1, 9):
    arrival_time += random.randint(1, 4)
    schedule(arrival_time, "arrive", customer)

waiting, busy_until, served = [], 0, []

while events:
    time, _, kind, customer = heapq.heappop(events)
    if kind == "arrive":
        if time >= busy_until:
            busy_until = time + 3
            schedule(busy_until, "finish", customer)
            served.append((customer, time, 0))
        else:
            waiting.append((customer, time))
    else:
        if waiting:
            next_customer, arrived = waiting.pop(0)
            start = max(time, arrived)
            busy_until = start + 3
            schedule(busy_until, "finish", next_customer)
            served.append((next_customer, arrived, start - arrived))

print("one counter, service takes 3 units, customers arrive at random")
print()
print("%10s %9s %8s" % ("customer", "arrived", "waited"))
for customer, arrived, waited in sorted(served):
    print("%10d %9d %8d" % (customer, arrived, waited))

total_wait = sum(w for _, _, w in served)
print()
print("customers served :", len(served))
print("total waiting    :", total_wait)
print("average wait     : %.2f units" % (total_wait / len(served)))
print()
print("every step of that simulation came out of a priority queue ordered by TIME.")
one counter, service takes 3 units, customers arrive at random

  customer   arrived   waited
         1         1        0
         2         2        2
         3         3        4
         4         6        4
         5         8        5
         6        11        5
         7        14        5
         8        16        0

customers served : 8
total waiting    : 25
average wait     : 3.12 units

every step of that simulation came out of a priority queue ordered by TIME.
munotes.in274

Where Priority Queues Are Used

The whole simulation is a loop over a priority queue: take the earliest event, handle it, and schedule whatever it causes. The priority is the clock, and that is the pattern behind every discrete event simulator.

The pattern to recognise

An examiner may describe a problem and ask what structure fits. The signs of a priority queue:

  • the next item to handle is chosen by a value, not by arrival or recency
  • items keep arriving while others are being handled
  • only the best item is ever needed, never the second best, and never a search
  • the collection does not need to be read in order

If the last point fails, that is, if the whole collection must be listed in order, then it is a search tree and not a heap.

Quick revision

  • A priority queue is used wherever the next item is chosen by importance rather than arrival.
  • Inside this paper: Huffman coding, Dijkstra's algorithm, and heap sort.
  • Outside it: process scheduling (with ageing against starvation), hospital triage, network packet

queues, event-driven simulation, and A-star search.

  • Event-driven simulation is the pattern worth knowing: the priority is a timestamp, and handling the

earliest event may schedule more.

  • The signs of a priority queue problem: the next item is chosen by value; items keep arriving; only the

best is ever needed; and the collection need not be readable in order.

  • If the collection must be listed in order, a search tree is wanted, not a heap.

Test yourself

1. Name the three uses inside this paper. Huffman coding, which repeatedly needs the two least frequent items; Dijkstra's algorithm, which needs the nearest unvisited vertex; and heap sort.

2. What is the priority in an event-driven simulation? The time at which the event happens. The loop repeatedly takes the earliest event, handles it, and schedules any events that result.

3. In the scheduling run, one interactive job gained nothing. Why? Because the scheduler is non-preemptive and the long batch job was already running when that job arrived, so nothing could be done for it. The second interactive job gained, from 16 units to 8. A preemptive scheduler would interrupt the batch job and help both.

4. What prevents starvation in a priority scheduler? Ageing: an item's priority is raised the longer it waits, so anything waiting long enough eventually becomes urgent.

5. Give the four signs that a problem wants a priority queue. The next item is chosen by a value rather than by arrival; items keep arriving while others are handled; only the best item is ever needed; and the collection does not have to be read in order.

6. Which of those signs, if it fails, points to a different structure? The last. If the whole collection must be listed in order, a binary search tree is wanted rather than a heap.

Contents This chapter on its own page

munotes.in275

Chapter Eighty-Seven

What a Graph Is

Syllabus topic Module 2, "Graph: Introduction"

In one line

A graph is a set of things together with a set of connections between them, and it is the most general structure in this paper because it assumes nothing about the shape of those connections.

The definition

A graph G is a pair (V, E) where V is a set of vertices (also called nodes) and E is a set of edges, each edge joining two vertices.

That is all. No root, no parent, no order, no limit on how many edges a vertex may have, and no requirement that everything be connected.

How it generalises everything before it

Every structure in this paper is a graph with a restriction added.

StructureThe restriction
Linked listeach vertex has at most one edge out, and no cycles
Treeconnected, no cycles, one distinguished root
Binary treea tree where each vertex has at most two edges out
Graphno restriction at all

So a graph can express anything the earlier structures can, and more besides: cycles, several routes between two points, and vertices connected to nothing.

That generality is why graph problems are harder. A tree traversal cannot revisit a node, because there is no way back; a graph traversal can, and chapters 94 and 95 must therefore remember where they have been, which no tree algorithm in this book needed to do.

Things that are already graphs

A road map. Vertices are junctions, edges are roads. The question "what is the shortest route" is chapter 98.

A social network. Vertices are people, edges are friendships. "Are these two connected" is chapter 96.

The web. Vertices are pages, edges are links. These edges go one way, which is the directed case.

A railway network, an electrical circuit, a project's task dependencies, the call graph of a program, the states of a game. All graphs.

And the structures in this book. A tree is a graph. A linked list is a graph. Drawing them as graphs is what makes the common vocabulary of the next chapter useful.

Directed and undirected

The one distinction to fix immediately, because everything else depends on it.

Undirected. An edge joins two vertices symmetrically. If A is joined to B then B is joined to A. Friendship, a road with two-way traffic, a wire.

Directed (a digraph). An edge goes from one vertex to another. A to B does not imply B to A. A web link, a one-way street, "A must finish before B".

undirected = {
    "Mumbai": {"Pune", "Nashik"},
    "Pune": {"Mumbai", "Nashik"},
    "Nashik": {"Mumbai", "Pune"},
    "Goa": set(),
}

directed = {
    "home": {"notes", "papers"},
    "notes": {"chapter"},
    "papers": {"chapter"},
    "chapter": set(),
}

print("an UNDIRECTED graph: cities joined by roads")
for city, neighbours in sorted(undirected.items()):
    print("   %-8s joined to %s" % (city, sorted(neighbours) or "nothing"))
print("   symmetric:", all(a in undirected[b]
                           for a, ns in undirected.items() for b in ns))
print("   Goa is in the graph with no edges at all:", undirected["Goa"] == set())

print()
print("a DIRECTED graph: pages linking to pages")
for page, links in sorted(directed.items()):
    print("   %-8s links to %s" % (page, sorted(links) or "nothing"))
print("   symmetric:", all(a in directed[b]
                           for a, ls in directed.items() for b in ls))
print("   home links to notes, but notes does NOT link back to home:",
      "notes" in directed["home"] and "home" not in directed["notes"])
munotes.in276

What a Graph Is

an UNDIRECTED graph: cities joined by roads
   Goa      joined to nothing
   Mumbai   joined to ['Nashik', 'Pune']
   Nashik   joined to ['Mumbai', 'Pune']
   Pune     joined to ['Mumbai', 'Nashik']
   symmetric: True
   Goa is in the graph with no edges at all: True

a DIRECTED graph: pages linking to pages
   chapter  links to nothing
   home     links to ['notes', 'papers']
   notes    links to ['chapter']
   papers   links to ['chapter']
   symmetric: False
   home links to notes, but notes does NOT link back to home: True

Two things in that run are worth noticing.

Goa is a vertex with no edges, and that is perfectly legal. A graph need not be connected, which no tree in this book was allowed to be.

The directed graph has two routes from home to chapter, through notes and through papers. A tree would have exactly one path between any two nodes (chapter 51); a graph may have none, one, or many.

Weighted graphs

An edge may carry a weight: a distance, a cost, a time, a capacity.

An unweighted graph answers "can I get there" and "how few steps". A weighted graph answers "how far" and "how cheaply", and those are different questions with different algorithms, which is exactly the difference between chapters 97 and 98.

The four kinds

UnweightedWeighted
Undirectedfriendships, two-way roadsroad distances
Directedweb links, prerequisitesone-way roads with distances, network costs

An examination question will say which it means, or the wording will imply it. "Roads between cities with distances" is undirected and weighted. "Which tasks must finish first" is directed and unweighted.

Quick revision

  • A graph is a set of vertices and a set of edges joining them. Nothing else is assumed.
  • Every structure in this paper is a graph with a restriction: a tree is a connected graph with no cycles

and a root; a list is a tree where each node has one child.

  • A graph may have cycles, several paths between two vertices, or vertices with no edges at all.
  • A graph traversal must remember where it has been; a tree traversal never had to.
  • Undirected: edges are symmetric. Directed (a digraph): an edge goes one way.
  • An edge may carry a weight, which turns "how few steps" into "how far".
  • The four kinds are directed or undirected, crossed with weighted or unweighted.
munotes.in277

What a Graph Is

Test yourself

1. Define a graph. A pair (V, E) where V is a set of vertices and E is a set of edges, each joining two vertices. Nothing else is assumed.

2. How is a tree a special case of a graph? A tree is a connected graph with no cycles and one distinguished root. A graph need be none of those things.

3. Why must a graph traversal remember where it has been, when a tree traversal need not? Because a graph may contain cycles and several paths to the same vertex, so a traversal can return to a vertex it has already visited and loop for ever. A tree has exactly one path to each node.

4. State the difference between a directed and an undirected graph, with an example of each. In an undirected graph an edge joins two vertices symmetrically, as a two-way road does. In a directed graph an edge goes from one vertex to another, as a web link does, and the reverse need not exist.

5. What does a weight on an edge represent, and what question does it change? A distance, cost, time or capacity. It changes the question from "how few steps" to "how far" or "how cheaply", which needs a different algorithm.

6. Is a vertex with no edges allowed? Is more than one path between two vertices allowed? Both are allowed. A graph need not be connected, and it may have many paths between two vertices, unlike a tree which has exactly one.

Contents This chapter on its own page

munotes.in278

Chapter Eighty-Eight

The Vocabulary of Graphs

Syllabus topic Module 2, "Graph: Introduction"

In one line

Degree counts a vertex's edges, a path is a sequence of edges between two vertices, a cycle is a path returning to its start, and connected means every vertex is reachable from every other.

The graph used throughout

A ---- B

| / |

| / |

C ---- D E

Vertices A, B, C, D, E. Edges AB, AC, BC, BD, CD. E has no edges.

The terms

TermMeaning
Adjacenttwo vertices joined by an edge; also called neighbours
Incidentan edge is incident on the two vertices it joins
Degree of a vertexthe number of edges on it
In-degree, out-degreefor a directed graph: edges coming in, edges going out
Patha sequence of vertices where each consecutive pair is joined by an edge
Simple patha path with no repeated vertex
Path lengththe number of edges in it, not vertices
Cyclea path that starts and ends at the same vertex, with at least one edge and no repeats otherwise
Acyclica graph with no cycles
Connectedevery vertex is reachable from every other
Componenta maximal set of vertices that are all reachable from one another
Complete graphevery pair of vertices is joined
Subgrapha graph made from some of the vertices and some of the edges

Two definitions students get wrong, worth stating twice:

Path length is the number of edges, not the number of vertices. A to B to C has length 2 and visits 3 vertices.

A graph with an isolated vertex is not connected, even though nothing is broken. E above makes the graph disconnected.

Computed

graph = {
    "A": {"B", "C"},
    "B": {"A", "C", "D"},
    "C": {"A", "B", "D"},
    "D": {"B", "C"},
    "E": set(),
}


def edges(g):
    """Each undirected edge once."""
    return sorted({tuple(sorted((a, b))) for a, ns in g.items() for b in ns})


def degree(g, v):
    return len(g[v])


def paths(g, start, end, seen=None):
    """Every simple path from start to end."""
    seen = (seen or []) + [start]
    if start == end:
        return [seen]
    found = []
    for neighbour in sorted(g[start]):
        if neighbour not in seen:
            found.extend(paths(g, neighbour, end, seen))
    return found


def components(g):
    """Maximal sets of mutually reachable vertices."""
    unseen, found = set(g), []
    while unseen:
        start = min(unseen)
        stack, group = [start], set()
        while stack:
            current = stack.pop()
            if current in group:
                continue
            group.add(current)
            stack.extend(g[current] - group)
        found.append(sorted(group))
        unseen -= group
    return found


def has_cycle(g):
    seen = set()
    for start in sorted(g):
        if start in seen:
            continue
        stack = [(start, None)]
        local = set()
        while stack:
            current, came_from = stack.pop()
            if current in local:
                return True
            local.add(current)
            for neighbour in g[current]:
                if neighbour != came_from:
                    stack.append((neighbour, current))
        seen |= local
    return False


print("vertices :", sorted(graph))
print("edges    :", ["".join(e) for e in edges(graph)])
print("edge count:", len(edges(graph)))
print()
print("degrees:")
for v in sorted(graph):
    print("   %s: %d  neighbours %s" % (v, degree(graph, v),
                                        sorted(graph[v]) or "none"))
print("   sum of degrees:", sum(degree(graph, v) for v in graph),
      "= twice the edge count:", 2 * len(edges(graph)))
print()
print("all simple paths from A to D:")
for path in paths(graph, "A", "D"):
    print("   %s   length %d edges, %d vertices"
          % (" to ".join(path), len(path) - 1, len(path)))
print()
print("components:", components(graph))
print("connected :", len(components(graph)) == 1)
print("acyclic   :", not has_cycle(graph))
print()
print("E is a vertex with degree 0, so the graph is NOT connected,")
print("even though nothing about it is broken.")
munotes.in279

The Vocabulary of Graphs

vertices : ['A', 'B', 'C', 'D', 'E']
edges    : ['AB', 'AC', 'BC', 'BD', 'CD']
edge count: 5

degrees:
   A: 2  neighbours ['B', 'C']
   B: 3  neighbours ['A', 'C', 'D']
   C: 3  neighbours ['A', 'B', 'D']
   D: 2  neighbours ['B', 'C']
   E: 0  neighbours none
   sum of degrees: 10 = twice the edge count: 10

all simple paths from A to D:
   A to B to C to D   length 3 edges, 4 vertices
   A to B to D   length 2 edges, 3 vertices
   A to C to B to D   length 3 edges, 4 vertices
   A to C to D   length 2 edges, 3 vertices

components: [['A', 'B', 'C', 'D'], ['E']]
connected : False
acyclic   : False

E is a vertex with degree 0, so the graph is NOT connected,
even though nothing about it is broken.

Four distinct simple paths from A to D, where a tree would have had exactly one. The graph is not connected because of E, and it has a cycle (A, B, C, A).

The handshaking lemma

The run confirms a result worth knowing by name:

The sum of all degrees is twice the number of edges.

Each edge contributes 1 to the degree of each of its two endpoints, so each is counted twice. Above: the degrees 2, 3, 3, 2, 0 sum to 10, and there are 5 edges.

A corollary examiners like: the number of vertices of odd degree is always even, since the total must be even.

How many edges can a graph have?

Maximum edges
Undirected, n verticesn(n-1)/2
Directed, n verticesn(n-1)

A graph with the maximum is complete. For n = 5 undirected that is 10 edges; the graph above has 5, so it is half complete.

This matters for chapter 92: a graph with close to the maximum is called dense, one with far fewer is sparse, and the two representations suit the two cases differently.

Quick revision

  • Adjacent means joined by an edge; degree is the number of edges on a vertex; in-degree and out-degree
munotes.in280

The Vocabulary of Graphs

for directed graphs.

  • A path is a sequence of vertices joined consecutively; a simple path repeats no vertex; path length is

the number of EDGES.

  • A cycle returns to its start; an acyclic graph has none.
  • Connected means every vertex is reachable from every other; a component is a maximal mutually reachable

set.

  • An isolated vertex makes a graph disconnected.
  • Handshaking lemma: the sum of degrees is twice the number of edges, so the number of odd-degree

vertices is even.

  • Maximum edges: n(n-1)/2 undirected, n(n-1) directed. A graph near the maximum is dense, far below it is

sparse.

Test yourself

1. Define degree, path length and cycle. Degree is the number of edges on a vertex. Path length is the number of edges in the path, not vertices. A cycle is a path that begins and ends at the same vertex with no other repeats.

2. The path A to B to C to D: what is its length and how many vertices does it visit? Length 3 edges, visiting 4 vertices.

3. State the handshaking lemma and its corollary. The sum of all vertex degrees is twice the number of edges, because each edge contributes to two vertices. It follows that the number of vertices of odd degree is even.

4. Is a graph with an isolated vertex connected? No. Connected requires every vertex to be reachable from every other, and nothing reaches an isolated vertex.

5. How many edges can an undirected graph on n vertices have, and what is such a graph called? n(n-1)/2, and a graph with that many is complete.

6. In the chapter's graph, how many simple paths run from A to D, and what would a tree have given? Four. A tree would have exactly one path between any two vertices.

Contents This chapter on its own page

munotes.in281

Chapter Eighty-Nine

The Graph ADT

Syllabus topic Module 2, "Graph: Graph ADT"

In one line

The graph ADT is add and remove vertices and edges, ask whether two vertices are joined, and list a vertex's neighbours, and the last of those is what every algorithm in the following chapters actually uses.

The ADT

A Graph holds a set of vertices and a set of edges joining them.

OperationNeedsReturnsDoes
Graph()nothingan empty graphcreates it
add_vertex(v)a vertexnothingadds it, with no edges
add_edge(u, v)two verticesnothingjoins them; error if either is absent
remove_edge(u, v)two verticesnothingunjoins them
remove_vertex(v)a vertexnothingremoves it and every edge on it
has_edge(u, v)two verticestrue or falseare they joined
neighbours(v)a vertexthe vertices joined to v
vertices(), edges()nothingthe sets
degree(v)a vertexa numberhow many edges on v

For a weighted graph, add_edge takes a weight and a weight(u, v) operation is added. For a directed graph, add_edge(u, v) joins u to v only, and neighbours means successors, with a separate predecessors if needed.

Two rows carry most of the design:

remove_vertex removes its edges too. Leaving an edge pointing at a vertex that no longer exists is the standard defect, and chapter 93 measures what it costs to do properly.

neighbours(v) is the operation that matters. Every traversal, every search and every shortest path in this paper is a loop over neighbours. Its cost, not has_edge's, decides whether an algorithm is fast, which is the whole of chapter 92.

The ADT, used before it is implemented

class Graph:
    """The ADT. The representation underneath is deliberately not the point of
    this chapter; chapters 90 and 91 give two, and a caller written here works
    with either."""

    def __init__(self, directed=False):
        self._adjacent = {}
        self.directed = directed

    def add_vertex(self, v):
        self._adjacent.setdefault(v, set())

    def add_edge(self, u, v):
        if u not in self._adjacent or v not in self._adjacent:
            raise KeyError("both vertices must exist: %r, %r" % (u, v))
        self._adjacent[u].add(v)
        if not self.directed:
            self._adjacent[v].add(u)

    def remove_edge(self, u, v):
        self._adjacent[u].discard(v)
        if not self.directed:
            self._adjacent[v].discard(u)

    def remove_vertex(self, v):
        if v not in self._adjacent:
            raise KeyError("no such vertex: %r" % (v,))
        del self._adjacent[v]
        for others in self._adjacent.values():
            others.discard(v)              # every edge ON it goes too

    def has_edge(self, u, v):
        return v in self._adjacent.get(u, ())

    def neighbours(self, v):
        return sorted(self._adjacent[v])

    def vertices(self):
        return sorted(self._adjacent)

    def edges(self):
        if self.directed:
            return sorted((u, v) for u, ns in self._adjacent.items() for v in ns)
        return sorted({tuple(sorted((u, v)))
                       for u, ns in self._adjacent.items() for v in ns})

    def degree(self, v):
        return len(self._adjacent[v])


g = Graph()
for city in ("Mumbai", "Pune", "Nashik", "Nagpur", "Goa"):
    g.add_vertex(city)
for a, b in (("Mumbai", "Pune"), ("Mumbai", "Nashik"), ("Pune", "Nashik"),
             ("Nashik", "Nagpur")):
    g.add_edge(a, b)

print("vertices  :", g.vertices())
print("edges     :", ["%s-%s" % e for e in g.edges()])
print("neighbours of Nashik:", g.neighbours("Nashik"))
print("degree of Nashik    :", g.degree("Nashik"))
print("Mumbai joined to Nagpur:", g.has_edge("Mumbai", "Nagpur"))
print("Goa has no edges        :", g.neighbours("Goa") == [])

print()
g.remove_vertex("Nashik")
print("after removing Nashik:")
print("   vertices:", g.vertices())
print("   edges   :", ["%s-%s" % e for e in g.edges()])
print("   every edge ON Nashik went with it:",
      all("Nashik" not in e for e in g.edges()))
print("   Nagpur is now isolated:", g.neighbours("Nagpur") == [])

print()
try:
    g.add_edge("Mumbai", "Chennai")
except KeyError as e:
    print("adding an edge to a missing vertex is refused:", e)
munotes.in282

The Graph ADT

vertices  : ['Goa', 'Mumbai', 'Nagpur', 'Nashik', 'Pune']
edges     : ['Mumbai-Nashik', 'Mumbai-Pune', 'Nagpur-Nashik', 'Nashik-Pune']
neighbours of Nashik: ['Mumbai', 'Nagpur', 'Pune']
degree of Nashik    : 3
Mumbai joined to Nagpur: False
Goa has no edges        : True

after removing Nashik:
   vertices: ['Goa', 'Mumbai', 'Nagpur', 'Pune']
   edges   : ['Mumbai-Pune']
   every edge ON Nashik went with it: True
   Nagpur is now isolated: True

adding an edge to a missing vertex is refused: "both vertices must exist: 'Mumbai', 'Chennai'"

Removing Nashik removed three edges with it, and Nagpur, which was only reachable through Nashik, became isolated. Removing a vertex can disconnect a graph, which no other removal in this paper could do.

The directed case

class Graph:
    def __init__(self, directed=False):
        self._adjacent = {}
        self.directed = directed

    def add_vertex(self, v):
        self._adjacent.setdefault(v, set())

    def add_edge(self, u, v):
        self._adjacent[u].add(v)
        if not self.directed:
            self._adjacent[v].add(u)

    def neighbours(self, v):
        return sorted(self._adjacent[v])

    def predecessors(self, v):
        return sorted(u for u, ns in self._adjacent.items() if v in ns)

    def in_degree(self, v):
        return len(self.predecessors(v))

    def out_degree(self, v):
        return len(self._adjacent[v])


d = Graph(directed=True)
for page in ("home", "notes", "papers", "chapter"):
    d.add_vertex(page)
for a, b in (("home", "notes"), ("home", "papers"),
             ("notes", "chapter"), ("papers", "chapter")):
    d.add_edge(a, b)

print("%-9s %-22s %-22s %s" % ("page", "links to", "linked from", "out/in"))
for page in ("home", "notes", "papers", "chapter"):
    print("%-9s %-22s %-22s %d/%d"
          % (page, str(d.neighbours(page)), str(d.predecessors(page)),
             d.out_degree(page), d.in_degree(page)))

print()
print("'home' has out-degree 2 and in-degree 0: nothing links to it.")
print("'chapter' has out-degree 0 and in-degree 2: it is a dead end.")
print()
print("finding predecessors cost a scan of the WHOLE graph, which successors did not.")
print("that asymmetry is real and chapter 92 measures it.")
page      links to               linked from            out/in
home      ['notes', 'papers']    []                     2/0
notes     ['chapter']            ['home']               1/1
papers    ['chapter']            ['home']               1/1
chapter   []                     ['notes', 'papers']    0/2

'home' has out-degree 2 and in-degree 0: nothing links to it.
'chapter' has out-degree 0 and in-degree 2: it is a dead end.

finding predecessors cost a scan of the WHOLE graph, which successors did not.
that asymmetry is real and chapter 92 measures it.

The asymmetry at the end is worth noticing: a directed graph stores its edges one way, so successors are immediate and predecessors cost a scan of everything. A structure needing both commonly stores the graph twice, once forwards and once reversed.

munotes.in283

The Graph ADT

Quick revision

  • The graph ADT: add and remove vertices and edges, has_edge, neighbours, vertices, edges, degree.
  • Weighted graphs add a weight to add_edge and a weight(u, v) operation; directed graphs make add_edge

one way and add predecessors.

  • remove_vertex must remove every edge on the vertex, or edges are left pointing at something that no

longer exists.

  • Removing a vertex can disconnect the graph, which no removal in the earlier structures could do.
  • neighbours(v) is the operation that matters: every traversal and search is a loop over it, and its

cost decides whether an algorithm is fast.

  • In a directed graph, successors are immediate and predecessors cost a full scan, so a structure needing

both stores the graph twice.

Test yourself

1. Write the graph ADT. add_vertex, add_edge, remove_edge, remove_vertex, has_edge, neighbours, vertices, edges and degree, with a weight on edges for a weighted graph and one-way edges plus predecessors for a directed one.

2. What must remove_vertex do beyond removing the vertex, and what happens if it does not? It must remove every edge on that vertex. Otherwise edges are left pointing at a vertex that no longer exists.

3. Which operation do the later algorithms actually depend on, and why does that matter? neighbours(v). Every traversal, search and shortest path is a loop over it, so its cost, rather than has_edge's, decides the complexity of the algorithms.

4. What can removing a vertex do that no removal from the earlier structures could? Disconnect the structure, leaving vertices that are no longer reachable, as removing Nashik isolated Nagpur.

5. Why are predecessors expensive in a directed graph? Because edges are stored at their source only, so finding what points at a vertex means scanning every vertex's list.

6. How is that usually solved in practice? By storing the graph twice, once as given and once with every edge reversed, so both directions are immediate.

Contents This chapter on its own page

munotes.in284

Chapter Ninety

The Adjacency Matrix

Syllabus topic Module 2, "Graph: Graph Representation using adjacency matrix and adjacency list"

In one line

An adjacency matrix is an n by n table where the cell at row i, column j says whether there is an edge from vertex i to vertex j, which makes that question instant and costs n squared memory whatever the graph.

The representation

Number the vertices 0 to n-1. Make an n by n table. Then:

matrix[i][j] = 1 if there is an edge from i to j, else 0

For a weighted graph the cell holds the weight instead of 1, and a special value such as infinity means no edge. Zero cannot mean "no edge" in a weighted graph, because zero is a legitimate weight.

For an undirected graph the matrix is symmetric: matrix[i][j] equals matrix[j][i], because an edge joins both ways. That symmetry is a useful check and it also means half the table is redundant.

Built and printed

class MatrixGraph:
    """A graph as an n by n table of 0s and 1s."""

    def __init__(self, labels, directed=False):
        self.labels = list(labels)
        self.index = {name: i for i, name in enumerate(self.labels)}
        n = len(self.labels)
        self.matrix = [[0] * n for _ in range(n)]
        self.directed = directed

    def add_edge(self, u, v):
        i, j = self.index[u], self.index[v]
        self.matrix[i][j] = 1
        if not self.directed:
            self.matrix[j][i] = 1

    def has_edge(self, u, v):
        return self.matrix[self.index[u]][self.index[v]] == 1

    def neighbours(self, v):
        i = self.index[v]
        return [self.labels[j] for j in range(len(self.labels))
                if self.matrix[i][j] == 1]

    def degree(self, v):
        return sum(self.matrix[self.index[v]])

    def is_symmetric(self):
        n = len(self.labels)
        return all(self.matrix[i][j] == self.matrix[j][i]
                   for i in range(n) for j in range(n))

    def show(self):
        header = "     " + " ".join("%3s" % name for name in self.labels)
        rows = [header]
        for name, row in zip(self.labels, self.matrix):
            rows.append("%4s " % name + " ".join("%3d" % cell for cell in row))
        return "\n".join(rows)


g = MatrixGraph(["A", "B", "C", "D", "E"])
for a, b in (("A", "B"), ("A", "C"), ("B", "C"), ("B", "D"), ("C", "D")):
    g.add_edge(a, b)

print("the adjacency matrix:")
print(g.show())
print()
print("has_edge(A, B)   :", g.has_edge("A", "B"), "  one cell, O(1)")
print("has_edge(A, D)   :", g.has_edge("A", "D"))
print("neighbours of B  :", g.neighbours("B"), "  a whole row, O(n)")
print("degree of B      :", g.degree("B"))
print("symmetric        :", g.is_symmetric(), "  (undirected graphs always are)")
print("E has no edges   :", g.neighbours("E") == [])
print()
print("cells in the table:", len(g.labels) ** 2)
print("edges in the graph:", sum(sum(row) for row in g.matrix) // 2)
print("so %d cells hold a 1 and %d hold a 0"
      % (sum(sum(row) for row in g.matrix),
         len(g.labels) ** 2 - sum(sum(row) for row in g.matrix)))
the adjacency matrix:
       A   B   C   D   E
   A   0   1   1   0   0
   B   1   0   1   1   0
   C   1   1   0   1   0
   D   0   1   1   0   0
   E   0   0   0   0   0

has_edge(A, B)   : True   one cell, O(1)
has_edge(A, D)   : False
neighbours of B  : ['A', 'C', 'D']   a whole row, O(n)
degree of B      : 3
symmetric        : True   (undirected graphs always are)
E has no edges   : True

cells in the table: 25
edges in the graph: 5
so 10 cells hold a 1 and 15 hold a 0
munotes.in285

The Adjacency Matrix

Twenty-five cells for five vertices and five edges. Ten cells hold a 1 (each edge appearing twice) and fifteen hold a 0, including the whole of row E and column E.

What it is good at

has_edge is O(1). One cell is read. No other representation does this, and it is the matrix's whole argument.

Adding or removing an edge is O(1). One cell is written.

The structure is simple, which matters more than it sounds: a two dimensional array needs no allocation, no pointers and no care, and in C it is a single block.

Weights fit naturally. The cell holds the weight. Many graph algorithms expressed as matrix operations, which is a real technique, rely on this.

What it is bad at

Memory is n squared whatever the graph. A graph of 10,000 vertices needs 100 million cells even if there are only 20 edges. That is the decisive weakness.

neighbours(v) is O(n). The whole row must be scanned, including every zero. Since every traversal and search is a loop over neighbours (chapter 89), this makes every graph algorithm O(n squared) on a matrix, even when the graph has very few edges.

Adding a vertex is O(n squared), because the table must be rebuilt one row and one column larger.

print("%10s %14s %14s %16s %s"
      % ("vertices", "matrix cells", "sparse edges", "cells per edge", "useful"))
for n in (10, 100, 1000, 10000):
    cells = n * n
    edges = 2 * n                     # a sparse graph: average degree 4
    print("%10d %14d %14d %16.0f %13.4f%%"
          % (n, cells, edges, cells / edges, 100 * 2 * edges / cells))

print()
print("at 10,000 vertices with 20,000 edges, 99.96% of the matrix is zero.")
  vertices   matrix cells   sparse edges   cells per edge useful
        10            100             20                5       40.0000%
       100          10000            200               50        4.0000%
      1000        1000000           2000              500        0.4000%
     10000      100000000          20000             5000        0.0400%

at 10,000 vertices with 20,000 edges, 99.96% of the matrix is zero.

At 10,000 vertices with an average of 4 edges each, the matrix is 99.96 per cent zeros. That is 100 million cells to store 20,000 facts.

When to use it

Despite that, the matrix is the right choice in three cases, and an answer should name them:

When the graph is dense, close to n(n-1)/2 edges. Then most cells are used and the O(1) has_edge comes free.

munotes.in286

The Adjacency Matrix

When has_edge is the dominant operation. If the program mostly asks "are these two joined" rather than "who are this one's neighbours", the matrix wins outright.

When n is small. For twenty vertices the matrix is 400 cells, which is nothing, and the simplicity is worth more than the memory.

Quick revision

  • An n by n table; cell [i][j] is 1 when there is an edge from i to j, or the weight in a weighted graph.
  • In a weighted graph, "no edge" must be infinity or a sentinel, never 0, since 0 is a valid weight.
  • An undirected graph's matrix is symmetric, so half of it is redundant.
  • has_edge and edge insertion and removal are O(1).
  • neighbours(v) is O(n), which makes every graph algorithm O(n squared) on a matrix.
  • Memory is n squared whatever the graph: at 10,000 vertices and 20,000 edges it is 99.96 per cent zeros.
  • Adding a vertex is O(n squared), since the table is rebuilt.
  • Use it for dense graphs, when has_edge dominates, or when n is small.

Test yourself

1. What does cell [i][j] hold, in an unweighted and in a weighted graph? In an unweighted graph, 1 if there is an edge from i to j and 0 otherwise. In a weighted graph, the weight, with infinity or a sentinel for no edge.

2. Why can 0 not mean "no edge" in a weighted graph? Because 0 is a legitimate weight, so the two cases would be indistinguishable.

3. Which operations are O(1), and which is O(n)? has_edge, add_edge and remove_edge are O(1). neighbours(v) is O(n), because the whole row must be scanned.

4. Why does an O(n) neighbours make every graph algorithm O(n squared)? Because every traversal, search and shortest path is a loop over neighbours for each vertex, so n vertices times O(n) per vertex is O(n squared), however few edges there are.

5. Give the measured waste at 10,000 vertices with 20,000 edges. The matrix has 100,000,000 cells of which 99.96 per cent are zero.

6. Name three situations where the matrix is the right choice. When the graph is dense; when has_edge is the dominant operation; and when the number of vertices is small enough that n squared is trivial.

Contents This chapter on its own page

munotes.in287

Chapter Ninety-One

The Adjacency List

Syllabus topic Module 2, "Graph: Graph Representation using adjacency matrix and adjacency list"

In one line

An adjacency list keeps, for each vertex, a list of the vertices it is joined to, so the memory is proportional to the edges that exist rather than to the edges that could exist.

The representation

For each vertex, a list of its neighbours. Nothing is stored about pairs that are not joined.

A: B, C

B: A, C, D

C: A, B, D

D: B, C

E: (empty)

Classically each of those lists is a singly linked list (chapter 13), which is why MU's syllabus puts linked lists in Module 1 and graphs in Module 2. In practice a dynamic array is usually used, because it has better cache behaviour and the lists are rarely modified in the middle.

For a weighted graph each entry holds the neighbour and the weight. For a directed graph the entry appears only in the source vertex's list.

Built

class ListGraph:
    """A graph as one list of neighbours per vertex."""

    def __init__(self, directed=False):
        self.adjacent = {}
        self.directed = directed

    def add_vertex(self, v):
        self.adjacent.setdefault(v, [])

    def add_edge(self, u, v, weight=None):
        entry_u = (v, weight) if weight is not None else v
        self.adjacent[u].append(entry_u)
        if not self.directed:
            entry_v = (u, weight) if weight is not None else u
            self.adjacent[v].append(entry_v)

    def has_edge(self, u, v):
        """A SEARCH of the list: O(degree), not O(1)."""
        for entry in self.adjacent[u]:
            name = entry[0] if isinstance(entry, tuple) else entry
            if name == v:
                return True
        return False

    def neighbours(self, v):
        """The list itself: O(degree), and it visits nothing that is not an edge."""
        return [e[0] if isinstance(e, tuple) else e for e in self.adjacent[v]]

    def degree(self, v):
        return len(self.adjacent[v])

    def edge_count(self):
        total = sum(len(ns) for ns in self.adjacent.values())
        return total if self.directed else total // 2

    def show(self):
        return "\n".join("   %-3s -> %s" % (v, ", ".join(str(n) for n in self.neighbours(v)) or "(none)")
                         for v in sorted(self.adjacent))


g = ListGraph()
for v in "ABCDE":
    g.add_vertex(v)
for a, b in (("A", "B"), ("A", "C"), ("B", "C"), ("B", "D"), ("C", "D")):
    g.add_edge(a, b)

print("the adjacency list:")
print(g.show())
print()
print("has_edge(A, B)  :", g.has_edge("A", "B"), "  a search of A's list")
print("has_edge(A, D)  :", g.has_edge("A", "D"), " and this one scanned the whole of it")
print("neighbours of B :", g.neighbours("B"), "  the list itself")
print("degree of B     :", g.degree("B"))
print("E has no edges  :", g.neighbours("E") == [])
print()
entries = sum(len(ns) for ns in g.adjacent.values())
print("entries stored  :", entries, "(each undirected edge appears twice)")
print("edges           :", g.edge_count())
print("a matrix for the same graph would need", len(g.adjacent) ** 2, "cells")
the adjacency list:
   A   -> B, C
   B   -> A, C, D
   C   -> A, B, D
   D   -> B, C
   E   -> (none)

has_edge(A, B)  : True   a search of A's list
has_edge(A, D)  : False  and this one scanned the whole of it
neighbours of B : ['A', 'C', 'D']   the list itself
degree of B     : 3
E has no edges  : True

entries stored  : 10 (each undirected edge appears twice)
edges           : 5
a matrix for the same graph would need 25 cells
munotes.in288

The Adjacency List

Ten entries against the matrix's twenty-five cells, and the gap widens with every vertex that is added without edges.

A weighted graph

class ListGraph:
    def __init__(self, directed=False):
        self.adjacent = {}
        self.directed = directed

    def add_vertex(self, v):
        self.adjacent.setdefault(v, [])

    def add_edge(self, u, v, weight):
        self.adjacent[u].append((v, weight))
        if not self.directed:
            self.adjacent[v].append((u, weight))

    def neighbours(self, v):
        return sorted(self.adjacent[v])

    def weight(self, u, v):
        for name, w in self.adjacent[u]:
            if name == v:
                return w
        return None


roads = ListGraph()
for city in ("Mumbai", "Pune", "Nashik", "Nagpur"):
    roads.add_vertex(city)
roads.add_edge("Mumbai", "Pune", 150)
roads.add_edge("Mumbai", "Nashik", 170)
roads.add_edge("Pune", "Nashik", 210)
roads.add_edge("Nashik", "Nagpur", 580)

print("a weighted graph, distances in kilometres:")
for city in sorted(roads.adjacent):
    pairs = ", ".join("%s (%d)" % (n, w) for n, w in roads.neighbours(city))
    print("   %-8s -> %s" % (city, pairs))
print()
print("weight Mumbai to Pune  :", roads.weight("Mumbai", "Pune"))
print("weight Mumbai to Nagpur:", roads.weight("Mumbai", "Nagpur"), "(no direct road)")
a weighted graph, distances in kilometres:
   Mumbai   -> Nashik (170), Pune (150)
   Nagpur   -> Nashik (580)
   Nashik   -> Mumbai (170), Nagpur (580), Pune (210)
   Pune     -> Mumbai (150), Nashik (210)

weight Mumbai to Pune  : 150
weight Mumbai to Nagpur: None (no direct road)

What it is good at

Memory is proportional to the edges. V + 2E entries for an undirected graph, not n squared. For the sparse graphs that occur in practice, which is most of them, this is the deciding advantage.

neighbours(v) is O(degree) and visits nothing wasted. The matrix scanned n cells to find 3 neighbours; the list reads 3 entries. This is why every algorithm in the following chapters is written against a list.

Adding a vertex is O(1). One empty list. The matrix had to be rebuilt.

It extends naturally. Weights, labels or any other per-edge information sit in the entry.

What it is bad at

has_edge(u, v) is O(degree), not O(1). The list must be searched. For a dense graph where a vertex has hundreds of neighbours, that is a real cost, and it is the matrix's one clear win.

Removing an edge is O(degree) for the same reason.

More memory per edge than one bit. Each entry carries a vertex reference and, in a linked implementation, a next pointer. A matrix cell can be a single bit. So for a very dense graph the matrix can actually use less memory, which surprises people and is worth stating.

Quick revision

  • One list of neighbours per vertex; nothing is stored for pairs that are not joined.
  • Classically a linked list per vertex, which is why Module 1 comes first; in practice a dynamic array.
  • Weighted graphs store (neighbour, weight) entries; directed graphs store the entry only at the source.
  • Memory is V + 2E for an undirected graph, proportional to the edges that exist.
  • neighbours(v) is O(degree) and wastes nothing, which is why the algorithms are written against it.
  • has_edge and remove_edge are O(degree), not O(1): the list must be searched. That is the matrix's win.
  • Adding a vertex is O(1), against the matrix's O(n squared).
  • For a very dense graph a bit matrix can use less memory than the list's references.
munotes.in289

The Adjacency List

Test yourself

1. What does an adjacency list store? For each vertex, a list of the vertices it is joined to, with weights alongside in a weighted graph. Nothing is stored for pairs with no edge.

2. How much memory does it use, and how does that compare with a matrix? V + 2E entries for an undirected graph, proportional to the edges that exist, against the matrix's n squared cells whatever the graph.

3. Which operation is cheap here and expensive on a matrix, and which is the other way round? neighbours(v) is O(degree) here and O(n) on a matrix. has_edge is O(degree) here and O(1) on a matrix.

4. Why are the later algorithms written against an adjacency list? Because every traversal and search is a loop over neighbours, and the list visits only real edges where the matrix scans every possible one.

5. Why is adding a vertex cheap here? It is one new empty list, O(1). A matrix must be rebuilt one row and one column larger, which is O(n squared).

6. Give the case where a matrix actually uses less memory than a list. A very dense graph, where a matrix cell can be a single bit while each list entry carries a vertex reference and possibly a next pointer.

Contents This chapter on its own page

munotes.in290

Chapter Ninety-Two

Which Representation: The Costs, Measured

Syllabus topic Module 2, "Graph ADT, Advantages and Disadvantages"

In one line

Neither representation wins everywhere: the list wins on memory and on every traversal, the matrix wins on the single question "are these two joined", and the deciding fact is whether the graph is sparse or dense.

How this chapter measures

Both representations are instrumented to count probes: one probe is one cell read on the matrix, one entry read on the list. A probe is a unit of real work and it is identical on every machine and every interpreter, so the numbers below can be relied on and reproduced.

The graphs are built arithmetically, with no random numbers, so every figure on this page is the same every time it runs.

The two instrumented representations

class CountingMatrix:
    """An adjacency matrix that counts every cell it reads."""

    def __init__(self, n):
        self.n = n
        self.matrix = [[0] * n for _ in range(n)]
        self.probes = 0

    def add_edge(self, i, j):
        self.matrix[i][j] = 1
        self.matrix[j][i] = 1

    def has_edge(self, i, j):
        self.probes += 1                      # one cell
        return self.matrix[i][j] == 1

    def neighbours(self, i):
        found = []
        for j in range(self.n):
            self.probes += 1                  # every cell in the row
            if self.matrix[i][j]:
                found.append(j)
        return found

    def units(self):
        return self.n * self.n                # cells, whatever the graph


class CountingList:
    """An adjacency list that counts every entry it reads."""

    def __init__(self, n):
        self.n = n
        self.adjacent = [[] for _ in range(n)]
        self.probes = 0

    def add_edge(self, i, j):
        self.adjacent[i].append(j)
        self.adjacent[j].append(i)

    def has_edge(self, i, j):
        for k in self.adjacent[i]:
            self.probes += 1                  # a search of the list
            if k == j:
                return True
        return False

    def neighbours(self, i):
        found = []
        for k in self.adjacent[i]:
            self.probes += 1                  # only real edges
            found.append(k)
        return found

    def units(self):
        return sum(len(a) for a in self.adjacent)   # entries, 2 per edge


def sparse_edges(n):
    """A ring plus one chord per vertex. Average degree 4, no random numbers."""
    pairs = set()
    for i in range(n):
        for j in ((i + 1) % n, (i * 7 + 3) % n):
            if i != j:
                pairs.add((min(i, j), max(i, j)))
    return sorted(pairs)


def dense_edges(n):
    """Every pair: the complete graph."""
    return [(i, j) for i in range(n) for j in range(i + 1, n)]


def build(kind, n, edges):
    g = kind(n)
    for i, j in edges:
        g.add_edge(i, j)
    g.probes = 0                              # measure the workload, not the build
    return g


print("memory, in units stored (matrix cells against list entries):")
print("%8s %10s %14s %14s %12s"
      % ("vertices", "edges", "matrix cells", "list entries", "list uses"))
for n in (10, 100, 1000):
    edges = sparse_edges(n)
    m = build(CountingMatrix, n, edges)
    l = build(CountingList, n, edges)
    print("%8d %10d %14d %14d %11.1f%%"
          % (n, len(edges), m.units(), l.units(),
             100 * l.units() / m.units()))
munotes.in291

Which Representation: The Costs, Measured

memory, in units stored (matrix cells against list entries):
vertices      edges   matrix cells   list entries    list uses
      10         15            100             30        30.0%
     100        194          10000            388         3.9%
    1000       1992        1000000           3984         0.4%

Workload one: a full traversal

Every traversal, search and shortest path in the chapters that follow does the same thing: it asks each vertex for its neighbours. So that is the workload to measure.

class CountingMatrix:
    def __init__(self, n):
        self.n = n
        self.matrix = [[0] * n for _ in range(n)]
        self.probes = 0

    def add_edge(self, i, j):
        self.matrix[i][j] = 1
        self.matrix[j][i] = 1

    def neighbours(self, i):
        found = []
        for j in range(self.n):
            self.probes += 1
            if self.matrix[i][j]:
                found.append(j)
        return found


class CountingList:
    def __init__(self, n):
        self.n = n
        self.adjacent = [[] for _ in range(n)]
        self.probes = 0

    def add_edge(self, i, j):
        self.adjacent[i].append(j)
        self.adjacent[j].append(i)

    def neighbours(self, i):
        found = []
        for k in self.adjacent[i]:
            self.probes += 1
            found.append(k)
        return found


def sparse_edges(n):
    pairs = set()
    for i in range(n):
        for j in ((i + 1) % n, (i * 7 + 3) % n):
            if i != j:
                pairs.add((min(i, j), max(i, j)))
    return sorted(pairs)


def visit_everyone(g):
    for v in range(g.n):
        g.neighbours(v)
    return g.probes


print("one full pass over a sparse graph, asking every vertex for its neighbours:")
print("%8s %8s %14s %14s %16s"
      % ("vertices", "edges", "matrix probes", "list probes", "matrix does more"))
for n in (10, 100, 1000, 2000):
    edges = sparse_edges(n)
    m = CountingMatrix(n)
    l = CountingList(n)
    for i, j in edges:
        m.add_edge(i, j)
        l.add_edge(i, j)
    mp = visit_everyone(m)
    lp = visit_everyone(l)
    print("%8d %8d %14d %14d %15.0fx"
          % (n, len(edges), mp, lp, mp / lp))

print()
print("the matrix reads n squared cells to find 2E edges, whatever E is.")
print("the list reads 2E entries and nothing else.")
one full pass over a sparse graph, asking every vertex for its neighbours:
vertices    edges  matrix probes    list probes matrix does more
      10       15            100             30               3x
     100      194          10000            388              26x
    1000     1992        1000000           3984             251x
    2000     3996        4000000           7992             501x

the matrix reads n squared cells to find 2E edges, whatever E is.
the list reads 2E entries and nothing else.

The matrix count is exactly n squared every time, 100 and 10,000 and 1,000,000 and 4,000,000, because the row scan does not care how many of the cells are edges. The list count is exactly the number of entries, which is twice the edges: 30 for 15 edges, 7,992 for 3,996. The gap on a sparse graph grows with n without limit, from 3 times at ten vertices to 501 times at two thousand.

This single table is why every algorithm in the following chapters is written against an adjacency list. A traversal that is O(V + E) on a list becomes O(V squared) on a matrix.

munotes.in292

Which Representation: The Costs, Measured

Workload two: asking "are these two joined"

Now the question the matrix was built for.

class CountingMatrix:
    def __init__(self, n):
        self.n = n
        self.matrix = [[0] * n for _ in range(n)]
        self.probes = 0

    def add_edge(self, i, j):
        self.matrix[i][j] = 1
        self.matrix[j][i] = 1

    def has_edge(self, i, j):
        self.probes += 1
        return self.matrix[i][j] == 1


class CountingList:
    def __init__(self, n):
        self.n = n
        self.adjacent = [[] for _ in range(n)]
        self.probes = 0

    def add_edge(self, i, j):
        self.adjacent[i].append(j)
        self.adjacent[j].append(i)

    def has_edge(self, i, j):
        for k in self.adjacent[i]:
            self.probes += 1
            if k == j:
                return True
        return False


def sparse_edges(n):
    pairs = set()
    for i in range(n):
        for j in ((i + 1) % n, (i * 7 + 3) % n):
            if i != j:
                pairs.add((min(i, j), max(i, j)))
    return sorted(pairs)


def dense_edges(n):
    return [(i, j) for i in range(n) for j in range(i + 1, n)]


def quiz(g, n):
    """n deterministic queries, some hits and some misses."""
    answers = 0
    for i in range(n):
        if g.has_edge(i, (i * 13 + 5) % n):
            answers += 1
    return g.probes, answers


print("%8s %8s %10s %14s %14s"
      % ("graph", "vertices", "queries", "matrix probes", "list probes"))
for n in (100, 500):
    for name, edges in (("sparse", sparse_edges(n)), ("dense", dense_edges(n))):
        m, l = CountingMatrix(n), CountingList(n)
        for i, j in edges:
            m.add_edge(i, j)
            l.add_edge(i, j)
        mp, _ = quiz(m, n)
        lp, _ = quiz(l, n)
        print("%8s %8d %10d %14d %14d" % (name, n, n, mp, lp))

print()
print("the matrix does exactly one probe per query in all four rows above.")
print("on the dense graph of 500 vertices every list holds 499 entries,")
print("so the list does hundreds of times the work for the same answer.")
   graph vertices    queries  matrix probes    list probes
  sparse      100        100            100            384
   dense      100        100            100           4999
  sparse      500        500            500           1985
   dense      500        500            500         124999

the matrix does exactly one probe per query in all four rows above.
on the dense graph of 500 vertices every list holds 499 entries,
so the list does hundreds of times the work for the same answer.

The matrix does exactly one probe per query whatever the graph: 100, 100, 500, 500. On the sparse graphs the list pays a few probes per query, 384 and 1,985, because the lists are short. On the dense graph of 500 vertices the list pays 124,999, which is 250 times the matrix's 500, because every list holds 499 entries and a miss must read all of them.

So the honest summary is not "lists are better". It is: lists are better for traversal and for memory, matrices are better for edge lookup, and density decides how much that matters.

munotes.in293

Which Representation: The Costs, Measured

The table to reproduce in an answer

OperationAdjacency matrixAdjacency list
MemoryO(V squared) alwaysO(V + E)
has_edge(u, v)O(1)O(degree of u)
neighbours(v)O(V)O(degree of v)
Add an edgeO(1)O(1)
Remove an edgeO(1)O(degree of u)
Add a vertexO(V squared), rebuiltO(1)
Remove a vertexO(V squared)O(V + E)
A full traversal, BFS or DFSO(V squared)O(V + E)
Suitsdense graphs, small V, edge lookupsparse graphs, traversal

Advantages and disadvantages, stated plainly

The matrix's advantages. Instant edge lookup; instant edge insertion and removal; a simple fixed structure with no pointers; weights sit naturally in the cells; symmetry gives a free correctness check on undirected graphs.

The matrix's disadvantages. O(V squared) memory even for an empty graph; a neighbour scan that reads every zero; O(V squared) to add a vertex; and consequently O(V squared) for every traversal.

The list's advantages. Memory proportional to the edges that exist; a neighbour scan that touches only real edges, so traversals are O(V + E); O(1) to add a vertex; and per-edge data such as weights fits in the entry.

The list's disadvantages. Edge lookup and edge removal are O(degree), not O(1); on a dense graph that degree is close to V; and each entry costs a reference, possibly with a next pointer, where a matrix cell could be a single bit.

What to say when asked to choose

Say the rule, then the reason:

Use an adjacency list unless the graph is dense or the program's main question is whether two given

vertices are joined. Real graphs, such as road networks, web links and social graphs, are sparse, so

the list is the default; and since BFS, DFS and Dijkstra all work by asking for neighbours, the list

keeps them at O(V + E) where the matrix would force O(V squared).

Quick revision

  • Probes are counted, not seconds: one cell read on a matrix, one entry read on a list.
  • A full traversal costs exactly V squared probes on a matrix and exactly 2E on a list.
  • So a traversal is O(V + E) on a list and O(V squared) on a matrix: the reason every algorithm here uses a list.
  • has_edge is one probe on a matrix always, and O(degree) on a list.
  • On a dense graph the list's has_edge collapses, because its lists hold nearly every vertex.
  • Memory: V squared cells always, against V + 2E entries.
  • Matrix advantages: O(1) lookup and edge updates, simple structure, natural weights, symmetry check.
  • List advantages: edge-proportional memory, O(1) vertex insertion, O(V + E) traversal.
  • Default to the list; choose the matrix for dense graphs, small V, or lookup-dominated work.
munotes.in294

Which Representation: The Costs, Measured

Test yourself

1. Why does this chapter count probes rather than seconds? Because a probe is a fixed unit of work that is identical on every machine and interpreter, while a timing varies with hardware, interpreter and load, so it cannot be quoted as a fact.

2. How many probes does a full traversal cost on each representation? Exactly V squared on a matrix, because every row is scanned in full, and exactly 2E on a list, because only real edges are read.

3. Give the complexity of BFS on each representation and explain the difference. O(V + E) on a list and O(V squared) on a matrix. BFS asks each vertex for its neighbours; the list answers in O(degree) and the matrix in O(V).

4. State one operation where the matrix is strictly better and one where the list is. has_edge is O(1) on a matrix and O(degree) on a list. Adding a vertex is O(1) on a list and O(V squared) on a matrix.

5. When does the list's has_edge become a serious weakness? On a dense graph, where a vertex's list holds nearly every other vertex, so a search of it approaches O(V).

6. Can a matrix ever use less memory than a list? Yes. For a very dense graph, where a cell can be a single bit while each list entry carries a vertex reference and possibly a next pointer.

7. Give the one-sentence rule for choosing. Use an adjacency list unless the graph is dense or the dominant question is whether two given vertices are joined.

Contents This chapter on its own page

munotes.in295

Chapter Ninety-Three

Inserting and Deleting Vertices and Edges

Syllabus topic Module 2, "Graph operations like insertion and deletion of nodes"

In one line

Adding an edge is cheap everywhere; deleting an edge is cheap only on a matrix; adding a vertex is cheap only on a list; and deleting a vertex is expensive on both, and dangerous on a matrix because it renumbers everything after it.

The four operations, before any code

Adjacency matrixAdjacency list
Insert an edgeO(1), write one cellO(1) appended, O(degree) if duplicates must be refused
Delete an edgeO(1), clear one cellO(degree), the list is searched
Insert a vertexO(V squared), the table is rebuiltO(1), one empty list
Delete a vertexO(V squared), and indices shiftO(V + E) in general

Everything below is that table, proved.

Inserting and deleting an edge

class MatrixGraph:
    def __init__(self, n):
        self.n = n
        self.matrix = [[0] * n for _ in range(n)]
        self.probes = 0

    def insert_edge(self, i, j):
        self.probes += 1
        self.matrix[i][j] = self.matrix[j][i] = 1

    def delete_edge(self, i, j):
        self.probes += 1
        self.matrix[i][j] = self.matrix[j][i] = 0

    def has_edge(self, i, j):
        return self.matrix[i][j] == 1


class ListGraph:
    def __init__(self, n):
        self.adjacent = [[] for _ in range(n)]
        self.probes = 0

    def insert_edge(self, i, j, allow_duplicates=True):
        if not allow_duplicates:
            for k in self.adjacent[i]:          # a search, to refuse a repeat
                self.probes += 1
                if k == j:
                    return False
        self.adjacent[i].append(j)
        self.adjacent[j].append(i)
        return True

    def delete_edge(self, i, j):
        removed = False
        for side, other in ((i, j), (j, i)):
            for position, k in enumerate(self.adjacent[side]):
                self.probes += 1                # a search, every time
                if k == other:
                    self.adjacent[side].pop(position)
                    removed = True
                    break
        return removed

    def has_edge(self, i, j):
        return j in self.adjacent[i]


star = ListGraph(12)                            # vertex 0 joined to 1..11
for j in range(1, 12):
    star.insert_edge(0, j)
star.probes = 0
star.delete_edge(0, 11)                         # the LAST entry in 0's list
print("list, deleting the last entry of a degree 11 vertex :",
      star.probes, "probes")

star2 = MatrixGraph(12)
for j in range(1, 12):
    star2.insert_edge(0, j)
star2.probes = 0
star2.delete_edge(0, 11)
print("matrix, the same deletion                            :",
      star2.probes, "probe")
print()

guarded = ListGraph(12)
for j in range(1, 12):
    guarded.insert_edge(0, j, allow_duplicates=False)
print("list, 11 insertions that each refuse duplicates      :",
      guarded.probes, "probes")
print("   (0 + 1 + 2 + ... + 10, because each insertion searches a longer list)")
print("   expected sum                                      :",
      sum(range(11)))
print()
print("duplicate refused on a repeat insert:",
      guarded.insert_edge(0, 5, allow_duplicates=False) is False)
list, deleting the last entry of a degree 11 vertex : 12 probes
matrix, the same deletion                            : 1 probe

list, 11 insertions that each refuse duplicates      : 55 probes
   (0 + 1 + 2 + ... + 10, because each insertion searches a longer list)
   expected sum                                      : 55

duplicate refused on a repeat insert: True

So an unguarded insertion is O(1) on both. But an insertion that refuses duplicates is O(degree) on a list, and eleven of them cost the sum 0 + 1 + 2 and so on. A matrix refuses duplicates for free, because writing 1 into a cell that already holds 1 changes nothing. That is a real and often unmentioned advantage of the matrix.

munotes.in296

Inserting and Deleting Vertices and Edges

Deletion of an edge is one probe on a matrix and a search on a list.

Inserting a vertex

On a list it is one empty list. On a matrix the whole table must be rebuilt one row and one column larger, because a two dimensional array cannot grow.

def grow_matrix(matrix):
    """Add one vertex: a new row and a new column. Every cell is touched."""
    n = len(matrix)
    bigger = [[0] * (n + 1) for _ in range(n + 1)]
    writes = 0
    for i in range(n):
        for j in range(n):
            bigger[i][j] = matrix[i][j]
            writes += 1
    return bigger, writes


def grow_list(adjacent):
    """Add one vertex: one empty list."""
    adjacent.append([])
    return adjacent, 1


print("%10s %18s %16s" % ("vertices", "matrix cell copies", "list writes"))
for n in (10, 100, 500):
    matrix = [[0] * n for _ in range(n)]
    _, writes = grow_matrix(matrix)
    adjacent = [[] for _ in range(n)]
    _, lw = grow_list(adjacent)
    print("%10d %18d %16d" % (n, writes, lw))

print()
matrix = [[0] * 4 for _ in range(4)]
matrix[0][1] = matrix[1][0] = 1
bigger, _ = grow_matrix(matrix)
print("a 4 by 4 matrix grown to 5 by 5, the old edge still in place:")
for row in bigger:
    print("   ", row)
print("edge (0,1) survived:", bigger[0][1] == 1)
print("the new vertex 4 has no edges:", sum(bigger[4]) == 0)
  vertices matrix cell copies      list writes
        10                100                1
       100              10000                1
       500             250000                1

a 4 by 4 matrix grown to 5 by 5, the old edge still in place:
    [0, 1, 0, 0, 0]
    [1, 0, 0, 0, 0]
    [0, 0, 0, 0, 0]
    [0, 0, 0, 0, 0]
    [0, 0, 0, 0, 0]
edge (0,1) survived: True
the new vertex 4 has no edges: True

Deleting a vertex: the operation that hurts

Deleting vertex v means two things, and both must happen:

  1. remove v itself;
  2. remove every edge that touches v, which means finding them.

On a matrix, row v and column v must go, and then every vertex numbered above v shifts down by one. That is the dangerous part and it is covered below.

On an undirected list, v's own list tells you exactly whose lists to clean, so the work is the sum of v's neighbours' degrees, not the whole graph. On a directed list that keeps only out-edges, nothing tells you who points at v, so every list in the graph must be scanned: O(V + E).

munotes.in297

Inserting and Deleting Vertices and Edges

def sparse_edges(n):
    pairs = set()
    for i in range(n):
        for j in ((i + 1) % n, (i * 7 + 3) % n):
            if i != j:
                pairs.add((min(i, j), max(i, j)))
    return sorted(pairs)


def delete_from_matrix(matrix, v):
    """Rebuild without row v and column v. Every surviving cell is copied."""
    n = len(matrix)
    writes = 0
    smaller = []
    for i in range(n):
        if i == v:
            continue
        row = []
        for j in range(n):
            if j == v:
                continue
            row.append(matrix[i][j])
            writes += 1
        smaller.append(row)
    return smaller, writes


def delete_from_undirected_list(adjacent, v):
    """v's own list names every vertex that needs cleaning."""
    probes = 0
    for neighbour in adjacent[v]:
        kept = []
        for k in adjacent[neighbour]:
            probes += 1
            if k != v:
                kept.append(k)
        adjacent[neighbour] = kept
    adjacent[v] = []
    return adjacent, probes


def delete_from_directed_out_list(adjacent, v):
    """Nothing says who points AT v, so every list must be read."""
    probes = 0
    for u in range(len(adjacent)):
        kept = []
        for k in adjacent[u]:
            probes += 1
            if k != v:
                kept.append(k)
        adjacent[u] = kept
    adjacent[v] = []
    return adjacent, probes


print("deleting one vertex from a sparse graph:")
print("%8s %8s %14s %18s %20s"
      % ("vertices", "edges", "matrix copies", "undirected list", "directed out-list"))
for n in (10, 100, 1000):
    edges = sparse_edges(n)

    matrix = [[0] * n for _ in range(n)]
    for i, j in edges:
        matrix[i][j] = matrix[j][i] = 1
    _, mw = delete_from_matrix(matrix, 3)

    undirected = [[] for _ in range(n)]
    for i, j in edges:
        undirected[i].append(j)
        undirected[j].append(i)
    _, up = delete_from_undirected_list(undirected, 3)

    directed = [[] for _ in range(n)]
    for i, j in edges:
        directed[i].append(j)
    _, dp = delete_from_directed_out_list(directed, 3)

    print("%8d %8d %14d %18d %20d" % (n, len(edges), mw, up, dp))

print()
print("the matrix copies (n-1) squared cells whatever the graph.")
print("the undirected list reads only the lists of vertex 3's own neighbours.")
print("the directed out-list reads every one of the E entries, across all V")
print("lists, because no entry says who points AT vertex 3: that is O(V + E).")
deleting one vertex from a sparse graph:
vertices    edges  matrix copies    undirected list    directed out-list
      10       15             81                  9                   15
     100      194           9801                 16                  194
    1000     1992         998001                 16                 1992

the matrix copies (n-1) squared cells whatever the graph.
the undirected list reads only the lists of vertex 3's own neighbours.
the directed out-list reads every one of the E entries, across all V
lists, because no entry says who points AT vertex 3: that is O(V + E).

The trap: deletion renumbers the vertices

This is the part worth remembering, because it is a bug that reaches production.

A matrix identifies a vertex by its position. Delete vertex 3 and the vertex that was 4 becomes 3, 5 becomes 4, and so on. Any vertex number held anywhere else in the program, in a distance array, a parent array, a queue of vertices still to visit, a user's bookmark, now points at the wrong vertex. Nothing reports this. The program keeps running and gives wrong answers.

munotes.in298

Inserting and Deleting Vertices and Edges

labels = ["A", "B", "C", "D", "E", "F"]
matrix = [[0] * 6 for _ in range(6)]
for i, j in ((0, 1), (1, 2), (2, 3), (3, 4), (4, 5)):
    matrix[i][j] = matrix[j][i] = 1

saved = 4                                   # the program remembers "vertex 4"
print("before: vertex %d is %s" % (saved, labels[saved]))

# delete vertex 2 (that is C) by rebuilding the table and the labels
keep = [i for i in range(6) if i != 2]
matrix = [[matrix[i][j] for j in keep] for i in keep]
labels = [labels[i] for i in keep]

print("after deleting vertex 2 (C):")
print("   labels are now", labels)
print("   the program still holds the number", saved)
print("   vertex %d is now %s, NOT %s" % (saved, labels[saved], "E"))
print("   the saved number silently means a different vertex:",
      labels[saved] != "E")
print()

# the fix: never renumber. Mark the vertex deleted and leave the position alone.
labels = ["A", "B", "C", "D", "E", "F"]
alive = [True] * 6
alive[2] = False                            # C is gone, position 2 is not reused
saved = 4
print("with a deleted flag instead of a rebuild:")
print("   vertex %d is still %s" % (saved, labels[saved]))
print("   C is marked gone:", alive[2] is False)
print("   every saved vertex number is still correct:",
      all(labels[i] == "ABCDEF"[i] for i in range(6)))
print("   live vertices:", [labels[i] for i in range(6) if alive[i]])
before: vertex 4 is E
after deleting vertex 2 (C):
   labels are now ['A', 'B', 'D', 'E', 'F']
   the program still holds the number 4
   vertex 4 is now F, NOT E
   the saved number silently means a different vertex: True

with a deleted flag instead of a rebuild:
   vertex 4 is still E
   C is marked gone: True
   every saved vertex number is still correct: True
   live vertices: ['A', 'B', 'D', 'E', 'F']

The fix has a name: lazy deletion, or a tombstone. Mark the vertex dead, leave its position alone, and make every traversal skip dead vertices. Insertion and deletion become O(1) for the bookkeeping, no saved number ever goes stale, and the price is that the table does not shrink. The same idea comes back in chapter 105, where a deleted slot in a hash table must be marked rather than cleared, for exactly the same reason.

An adjacency list keyed by a name rather than a position does not have this problem at all, which is a further argument for the dictionary-of-lists form used in chapter 91.

munotes.in299

Inserting and Deleting Vertices and Edges

A practical rule

  • If vertices are added and deleted often, use a list keyed by name, or a list with a deleted flag.
  • If the vertex set is fixed and only edges change, a matrix is comfortable and its O(1) edge updates are

worth having.

  • Never renumber a vertex while any other part of the program is holding vertex numbers.

Quick revision

  • Insert an edge: O(1) on both; O(degree) on a list if duplicates must be refused, which a matrix refuses for free.
  • Delete an edge: O(1) on a matrix, O(degree) on a list.
  • Insert a vertex: O(1) on a list, O(V squared) on a matrix because the table is rebuilt.
  • Delete a vertex: O(V squared) on a matrix; on an undirected list, the sum of the neighbours' degrees; on a directed out-list, O(V + E), because nothing says who points at the vertex.
  • Deleting from a matrix renumbers every later vertex, so saved vertex numbers silently point elsewhere.
  • Lazy deletion, a deleted flag, keeps positions stable; the same trick is needed in a hash table.
  • A list keyed by name avoids renumbering entirely.

Test yourself

1. Which representation inserts an edge faster, and what changes if duplicate edges must be refused? Both are O(1) unguarded. With a duplicate check the list becomes O(degree), since it must search; the matrix still costs one write, because setting a cell that already holds 1 changes nothing.

2. Why is inserting a vertex O(V squared) on a matrix? A two dimensional array cannot grow, so a table one row and one column larger is allocated and every old cell is copied.

3. Why does deleting a vertex cost O(V + E) on a directed adjacency list that stores only out-edges? Because nothing in the structure says which vertices point at the deleted one, so every list in the graph must be scanned to remove references to it.

4. Why is the same deletion cheaper on an undirected list? The vertex's own list names exactly its neighbours, so only their lists need cleaning: the sum of the neighbours' degrees rather than the whole graph.

5. What is the renumbering trap, and why is it dangerous? Deleting a vertex from a matrix shifts every higher vertex number down by one, so any vertex number held elsewhere in the program now refers to a different vertex. No error is raised; the program simply gives wrong answers.

6. What is lazy deletion and what does it cost? The vertex is marked dead and its position is kept, so no saved number goes stale and traversals skip it. The cost is that the structure does not shrink.

munotes.in300

Inserting and Deleting Vertices and Edges

7. Give the rule for choosing when vertices change often. Use an adjacency list keyed by name, or a list with a deleted flag; reserve the matrix for a fixed vertex set whose edges change.

Contents This chapter on its own page

munotes.in301

Chapter Ninety-Six

Connectivity and Connected Components

Syllabus topic Outcome 5, "solve shortest path and connectivity problems"

In one line

A graph is connected when every vertex can be reached from every other; a connected component is a maximal set of vertices that can all reach one another; and both are decided by running a traversal and counting what it reached.

The definitions, precisely

Reachable. w is reachable from v when there is a path from v to w.

Connected graph. An undirected graph is connected when every vertex is reachable from every other vertex. Equivalently: one traversal from any single vertex reaches all V.

Connected component. A maximal connected piece. "Maximal" is the load-bearing word: a component cannot be made larger by adding another vertex of the graph, because if it could, that vertex was reachable and belonged in it already.

Disconnected graph. A graph with more than one component.

Isolated vertex. A vertex of degree zero. It is a component all by itself, which students regularly forget to count.

For directed graphs the words change, and the distinction is examinable:

Strongly connected. Every vertex can reach every other, following the arrow directions.

Weakly connected. Not strongly connected, but connected if the directions are ignored.

Deciding connectivity with one traversal

from collections import deque


def adjacency(pairs, vertices):
    graph = {v: [] for v in vertices}
    for a, b in pairs:
        graph[a].append(b)
        graph[b].append(a)
    for v in graph:
        graph[v].sort()
    return graph


def reachable_from(graph, start):
    """Every vertex reachable from start. BFS, but DFS gives the same set."""
    seen = {start}
    queue = deque([start])
    while queue:
        v = queue.popleft()
        for w in graph[v]:
            if w not in seen:
                seen.add(w)
                queue.append(w)
    return seen


def is_connected(graph):
    if not graph:
        return True                       # the empty graph, by convention
    start = next(iter(graph))
    return len(reachable_from(graph, start)) == len(graph)


joined = adjacency([("A", "B"), ("B", "C"), ("C", "D")], "ABCD")
split = adjacency([("A", "B"), ("C", "D")], "ABCD")
lonely = adjacency([("A", "B"), ("B", "C")], "ABCDE")   # D and E are isolated

for name, g in (("a path A-B-C-D", joined),
                ("two pairs A-B and C-D", split),
                ("a path plus 2 isolated vertices", lonely)):
    reached = reachable_from(g, "A")
    print("%-34s connected: %-5s  reached %d of %d from A"
          % (name, is_connected(g), len(reached), len(g)))
a path A-B-C-D                     connected: True   reached 4 of 4 from A
two pairs A-B and C-D              connected: False  reached 2 of 4 from A
a path plus 2 isolated vertices    connected: False  reached 3 of 5 from A

One traversal settles it, and it costs exactly what the traversal costs: O(V + E).

Finding every component

A single traversal only reaches one component. To find them all, start a fresh traversal from every vertex not yet seen.

components(graph):

seen = empty

count = 0

for each vertex v in the graph:

if v not in seen:

count = count + 1

walk from v, adding everything reached to seen

return count

munotes.in312

Connectivity and Connected Components

That outer loop is the whole algorithm, and it does not change the complexity: every vertex is still visited once and every edge still read once, so it is O(V + E) in total, not O(V) traversals.

from collections import deque


def adjacency(pairs, vertices):
    graph = {v: [] for v in vertices}
    for a, b in pairs:
        graph[a].append(b)
        graph[b].append(a)
    for v in graph:
        graph[v].sort()
    return graph


def components(graph):
    """Every component, as a list of sorted vertex lists."""
    seen = set()
    found = []
    reads = 0
    for start in sorted(graph):
        if start in seen:
            continue
        piece = {start}
        seen.add(start)
        queue = deque([start])
        while queue:
            v = queue.popleft()
            for w in graph[v]:
                reads += 1
                if w not in seen:
                    seen.add(w)
                    piece.add(w)
                    queue.append(w)
        found.append(sorted(piece))
    return found, reads


EDGES = [("A", "B"), ("B", "C"), ("A", "C"),      # a triangle
         ("D", "E"),                              # a pair
         ("F", "G"), ("G", "H"), ("H", "F"), ("H", "I")]   # a triangle with a tail
VERTICES = "ABCDEFGHIJ"                           # J is isolated

GRAPH = adjacency(EDGES, VERTICES)
found, reads = components(GRAPH)

print("the graph has %d vertices and %d edges" % (len(GRAPH), len(EDGES)))
print("components found:", len(found))
for i, piece in enumerate(found, 1):
    word = "vertex" if len(piece) == 1 else "vertices"
    print("   component %d (%d %s): %s" % (i, len(piece), word, " ".join(piece)))
print()
print("every vertex is in exactly one component:",
      sorted(v for piece in found for v in piece) == sorted(VERTICES))
print("the isolated vertex J is its own component:", ["J"] in found)
print("adjacency entries read in total:", reads, "= 2E for E =", len(EDGES))
print("so finding ALL components costs one traversal's work, O(V + E).")
the graph has 10 vertices and 8 edges
components found: 4
   component 1 (3 vertices): A B C
   component 2 (2 vertices): D E
   component 3 (4 vertices): F G H I
   component 4 (1 vertex): J

every vertex is in exactly one component: True
the isolated vertex J is its own component: True
adjacency entries read in total: 16 = 2E for E = 8
so finding ALL components costs one traversal's work, O(V + E).

Note the two results worth remembering: every vertex lands in exactly one component, and the isolated vertex counts. A question that gives a vertex with no edges is testing exactly that.

Are two vertices connected to each other

from collections import deque


def adjacency(pairs, vertices):
    graph = {v: [] for v in vertices}
    for a, b in pairs:
        graph[a].append(b)
        graph[b].append(a)
    for v in graph:
        graph[v].sort()
    return graph


def label_components(graph):
    """Give every vertex its component number. One pass, then O(1) questions."""
    label = {}
    number = 0
    for start in sorted(graph):
        if start in label:
            continue
        label[start] = number
        queue = deque([start])
        while queue:
            v = queue.popleft()
            for w in graph[v]:
                if w not in label:
                    label[w] = number
                    queue.append(w)
        number += 1
    return label, number


EDGES = [("A", "B"), ("B", "C"), ("A", "C"), ("D", "E"),
         ("F", "G"), ("G", "H"), ("H", "F"), ("H", "I")]
GRAPH = adjacency(EDGES, "ABCDEFGHIJ")
label, count = label_components(GRAPH)

print("component number of each vertex:")
print("   " + "  ".join("%s:%d" % (v, label[v]) for v in sorted(label)))
print("components:", count)
print()
for a, b in (("A", "C"), ("A", "D"), ("F", "I"), ("J", "A")):
    print("   %s and %s connected to each other: %s"
          % (a, b, label[a] == label[b]))
print()
print("one O(V + E) pass makes every later question O(1).")
munotes.in313

Connectivity and Connected Components

component number of each vertex:
   A:0  B:0  C:0  D:1  E:1  F:2  G:2  H:2  I:2  J:3
components: 4

   A and C connected to each other: True
   A and D connected to each other: False
   F and I connected to each other: True
   J and A connected to each other: False

one O(V + E) pass makes every later question O(1).

This is the shape to use when the same graph is asked many questions: label once, answer instantly. The alternative, a fresh traversal per question, is O(V + E) every time.

Two facts worth quoting

A connected graph on V vertices has at least V-1 edges. So a graph with fewer than V-1 edges is certainly disconnected, without running anything.

Exactly V-1 edges and connected means it is a tree: connected with no cycle. That matches what chapter 88 observed from the other direction, that a tree has exactly one path between any two vertices. One more edge creates a cycle, and so a second path.

from collections import deque


def is_connected_count(n, edges):
    graph = {i: [] for i in range(n)}
    for a, b in edges:
        graph[a].append(b)
        graph[b].append(a)
    seen = {0}
    queue = deque([0])
    while queue:
        v = queue.popleft()
        for w in graph[v]:
            if w not in seen:
                seen.add(w)
                queue.append(w)
    return len(seen) == n


n = 6
print("on %d vertices, the minimum for a connected graph is %d edges" % (n, n - 1))
print()

too_few = [(0, 1), (1, 2), (2, 3), (3, 4)]          # 4 edges, 5 is the minimum
a_tree = [(0, 1), (1, 2), (2, 3), (3, 4), (4, 5)]   # exactly 5
one_more = a_tree + [(0, 5)]                        # 6 edges: a cycle appears

for name, edges in (("4 edges", too_few), ("5 edges, a path", a_tree),
                    ("6 edges, the path closed", one_more)):
    print("   %-26s connected: %-5s   edges %d, V-1 = %d"
          % (name, is_connected_count(n, edges), len(edges), n - 1))

print()
print("4 edges on 6 vertices cannot be connected, whatever the arrangement:",
      len(too_few) < n - 1)
print("5 edges and connected means a tree, so no cycle:", len(a_tree) == n - 1)
print("adding the 6th edge closed a cycle and kept it connected.")
munotes.in314

Connectivity and Connected Components

on 6 vertices, the minimum for a connected graph is 5 edges

   4 edges                    connected: False   edges 4, V-1 = 5
   5 edges, a path            connected: True    edges 5, V-1 = 5
   6 edges, the path closed   connected: True    edges 6, V-1 = 5

4 edges on 6 vertices cannot be connected, whatever the arrangement: True
5 edges and connected means a tree, so no cycle: True
adding the 6th edge closed a cycle and kept it connected.

Directed graphs: strongly and weakly connected

from collections import deque


def reachable(out_edges, start, keys):
    seen = {start}
    queue = deque([start])
    while queue:
        v = queue.popleft()
        for w in out_edges.get(v, []):
            if w not in seen:
                seen.add(w)
                queue.append(w)
    return seen


def classify(vertices, arcs):
    out = {v: [] for v in vertices}
    both = {v: [] for v in vertices}
    for a, b in arcs:
        out[a].append(b)
        both[a].append(b)
        both[b].append(a)

    strong = all(len(reachable(out, v, vertices)) == len(vertices)
                 for v in vertices)
    weak = len(reachable(both, vertices[0], vertices)) == len(vertices)
    if strong:
        return "strongly connected"
    if weak:
        return "weakly connected"
    return "disconnected"


cycle = (["A", "B", "C"], [("A", "B"), ("B", "C"), ("C", "A")])
chain = (["A", "B", "C"], [("A", "B"), ("B", "C")])
apart = (["A", "B", "C", "D"], [("A", "B"), ("C", "D")])

for name, (vs, arcs) in (("a directed cycle A->B->C->A", cycle),
                         ("a chain A->B->C", chain),
                         ("A->B and C->D", apart)):
    print("%-30s %s" % (name, classify(vs, arcs)))

print()
print("in the chain, C cannot reach A, so it is not strongly connected,")
print("but ignoring the arrows it is one piece, so it is weakly connected.")
a directed cycle A->B->C->A    strongly connected
a chain A->B->C                weakly connected
A->B and C->D                  disconnected

in the chain, C cannot reach A, so it is not strongly connected,
but ignoring the arrows it is one piece, so it is weakly connected.

Where this is used

Networks. Can every machine reach every other? A disconnected network has a cut somewhere, and the components tell you where.

Social graphs. A component is a group with no link at all to the rest.

Image processing. Connected component labelling finds the separate shapes in a picture, treating each pixel as a vertex joined to its neighbours. It is this algorithm exactly.

Spreadsheets, builds and dependency graphs. Components separate work that shares nothing, so each component can be handled independently or in parallel.

Before any other graph work. Many algorithms assume a connected graph, so checking is a sensible first step.

Quick revision

  • A graph is connected when one traversal from any vertex reaches all V vertices.
  • A connected component is a maximal connected piece; maximal means it cannot be extended.
  • An isolated vertex is a component of size one and must be counted.
  • All components are found by starting a fresh traversal from each unvisited vertex, and the total cost is still O(V + E).
  • Every vertex lies in exactly one component.
  • Labelling components once, in O(V + E), makes every later "are these two connected" question O(1).
  • A connected graph needs at least V-1 edges, so fewer than V-1 edges is disconnected without checking.
  • Connected with exactly V-1 edges means a tree; one more edge makes a cycle.
  • Directed: strongly connected means every vertex reaches every other along the arrows; weakly connected means it is one piece only when the arrows are ignored.
munotes.in315

Connectivity and Connected Components

Test yourself

1. Define a connected component, and say why the word maximal matters. A maximal set of vertices that can all reach one another. Maximal matters because if another vertex of the graph could be added, it was reachable and so belonged to the component already.

2. How is connectivity decided, and at what cost? Run one traversal from any vertex and check that it reached all V vertices. The cost is the traversal's, O(V + E).

3. Why is finding all components still O(V + E) and not O(V) traversals? Because the outer loop starts a traversal only at an unvisited vertex, so across all of them each vertex is visited once and each edge read once.

4. Is a vertex with no edges a component? Yes, a component of size one, and it must be counted.

5. What is the fastest way to answer many "are u and v connected" questions on one fixed graph? Label every vertex with its component number in one O(V + E) pass; each later question is then an O(1) comparison of two labels.

6. A graph has 6 vertices and 4 edges. Can it be connected? No. A connected graph on V vertices needs at least V-1 edges, so 5 here, whatever the arrangement.

7. Distinguish strongly and weakly connected. Strongly connected: every vertex can reach every other following the arrow directions. Weakly connected: that fails, but the graph is one piece once the directions are ignored. A chain A to B to C is weakly connected, since C cannot reach A.

8. Name a real use of connected component labelling outside networks. Finding the separate shapes in an image, where each pixel is a vertex joined to its neighbours.

Contents This chapter on its own page

munotes.in316

Chapter Ninety-Seven

The Shortest Path in an Unweighted Graph

Syllabus topic Module 2, "Applications of Graphs like shortest path algorithms"

In one line

When every edge has the same cost, the shortest path is the one with the fewest edges, and breadth first search finds it because it visits vertices in order of increasing distance.

Why BFS is already the answer

BFS visits the start, then everything one edge away, then everything two edges away. So when it first reaches a vertex, it has reached it by the fewest possible edges: if there were a shorter route, that route's length would be a smaller level, and BFS would have arrived on that earlier level.

That is the whole proof, and it is worth writing out in an answer in exactly that form:

BFS assigns level k to a vertex only after every vertex at level k-1 has been processed. So a vertex

first reached at level k has no route of length less than k, because such a route would have reached it

while an earlier level was being processed.

The consequence: no new algorithm is needed. Record a parent as each vertex is first reached, and the parents chain back to the start along a shortest route.

The algorithm, with the route rebuilt

from collections import deque


def adjacency(pairs, vertices):
    graph = {v: [] for v in vertices}
    for a, b in pairs:
        graph[a].append(b)
        graph[b].append(a)
    for v in graph:
        graph[v].sort()
    return graph


def shortest_paths(graph, start):
    """Distance in edges, and the parent of each vertex on a shortest route."""
    distance = {start: 0}
    parent = {start: None}
    queue = deque([start])
    while queue:
        v = queue.popleft()
        for w in graph[v]:
            if w not in distance:              # first arrival is the shortest
                distance[w] = distance[v] + 1
                parent[w] = v
                queue.append(w)
    return distance, parent


def route(parent, target):
    if target not in parent:
        return None                            # unreachable
    steps = []
    while target is not None:
        steps.append(target)
        target = parent[target]
    return list(reversed(steps))


EDGES = [("A", "B"), ("A", "C"), ("B", "D"), ("C", "D"),
         ("C", "E"), ("D", "F"), ("E", "F"), ("F", "G")]
GRAPH = adjacency(EDGES, "ABCDEFGH")          # H is isolated

distance, parent = shortest_paths(GRAPH, "A")

print("%-8s %10s   %s" % ("vertex", "edges away", "a shortest route"))
for v in sorted(GRAPH):
    path = route(parent, v)
    if path is None:
        print("%-8s %10s   %s" % (v, "-", "unreachable from A"))
    else:
        print("%-8s %10d   %s" % (v, distance[v], " -> ".join(path)))

print()
print("G is reached in", distance["G"], "edges by", " -> ".join(route(parent, "G")))
print("H is unreachable, and the algorithm says so rather than guessing:",
      route(parent, "H") is None)
vertex   edges away   a shortest route
A                 0   A
B                 1   A -> B
C                 1   A -> C
D                 2   A -> B -> D
E                 2   A -> C -> E
F                 3   A -> B -> D -> F
G                 4   A -> B -> D -> F -> G
H                 -   unreachable from A

G is reached in 4 edges by A -> B -> D -> F -> G
H is unreachable, and the algorithm says so rather than guessing: True
munotes.in317

The Shortest Path in an Unweighted Graph

Three things an answer should state about that output. Unreachable is a real answer, and the algorithm must report it rather than return a wrong number or loop. The first arrival is the shortest, which is why the test is if w not in distance and no distance is ever revised. And the route is rebuilt from the parents, so a shortest path costs no extra traversal.

Why DFS cannot be used here

def adjacency(pairs, vertices):
    graph = {v: [] for v in vertices}
    for a, b in pairs:
        graph[a].append(b)
        graph[b].append(a)
    for v in graph:
        graph[v].sort()
    return graph


EDGES = [("A", "B"), ("A", "C"), ("B", "D"), ("C", "D"),
         ("C", "E"), ("D", "F"), ("E", "F"), ("F", "G")]
GRAPH = adjacency(EDGES, "ABCDEFG")


def dfs_first_arrival(graph, v, depth=0, seen=None, found=None):
    """DFS, recording the depth at which each vertex is FIRST reached."""
    if seen is None:
        seen, found = set(), {}
    seen.add(v)
    found[v] = depth
    for w in graph[v]:
        if w not in seen:
            dfs_first_arrival(graph, w, depth + 1, seen, found)
    return found


from collections import deque


def bfs_distance(graph, start):
    distance = {start: 0}
    queue = deque([start])
    while queue:
        v = queue.popleft()
        for w in graph[v]:
            if w not in distance:
                distance[w] = distance[v] + 1
                queue.append(w)
    return distance


dfs_depth = dfs_first_arrival(GRAPH, "A")
bfs_dist = bfs_distance(GRAPH, "A")

print("%-8s %12s %12s %s" % ("vertex", "DFS depth", "BFS distance", "DFS wrong by"))
wrong = 0
for v in sorted(GRAPH):
    gap = dfs_depth[v] - bfs_dist[v]
    if gap:
        wrong += 1
    print("%-8s %12d %12d %s"
          % (v, dfs_depth[v], bfs_dist[v], gap if gap else ""))

print()
print("DFS overstated the distance for", wrong, "of", len(GRAPH), "vertices.")
print("DFS commits to a branch before checking a shorter one exists.")
print("the BFS numbers are the true shortest distances.")
vertex      DFS depth BFS distance DFS wrong by
A                   0            0
B                   1            1
C                   3            1 2
D                   2            2
E                   4            2 2
F                   5            3 2
G                   6            4 2

DFS overstated the distance for 4 of 7 vertices.
DFS commits to a branch before checking a shorter one exists.
the BFS numbers are the true shortest distances.

DFS is not wrong about reachability; it is wrong about distance. It walks into a long branch and records the depth it happened to arrive at, with no reason for that to be minimal.

The trap: an unweighted algorithm on a weighted graph

This is the examinable mistake, and the reason the next chapter exists.

from collections import deque

# A weighted graph. BFS cannot see the weights at all.
WEIGHTED = {
    "A": [("B", 1), ("C", 10)],
    "B": [("A", 1), ("D", 20)],
    "C": [("A", 10), ("D", 1)],
    "D": [("B", 20), ("C", 1)],
}

print("the weighted graph:")
for v in sorted(WEIGHTED):
    print("   %s -> %s" % (v, ", ".join("%s(%d)" % p for p in WEIGHTED[v])))
print()

# BFS: fewest EDGES from A to D
distance, parent = {"A": 0}, {"A": None}
queue = deque(["A"])
while queue:
    v = queue.popleft()
    for w, _weight in WEIGHTED[v]:
        if w not in distance:
            distance[w] = distance[v] + 1
            parent[w] = v
            queue.append(w)

bfs_route = []
node = "D"
while node is not None:
    bfs_route.append(node)
    node = parent[node]
bfs_route.reverse()


def cost(route):
    total = 0
    for a, b in zip(route, route[1:]):
        total += next(w for n, w in WEIGHTED[a] if n == b)
    return total


print("BFS answer    :", " -> ".join(bfs_route),
      "  edges:", len(bfs_route) - 1, "  cost:", cost(bfs_route))
better = ["A", "C", "D"]
print("a cheaper route:", " -> ".join(better),
      "  edges:", len(better) - 1, "  cost:", cost(better))
print()
print("BFS picked the route with fewer edges and it costs more:",
      cost(bfs_route) > cost(better))
print("it is worse by", cost(bfs_route) - cost(better), "units.")
print()
print("BFS answers 'fewest edges'. on a weighted graph that is a different")
print("question from 'cheapest', so BFS is the wrong tool: use Dijkstra.")
munotes.in318

The Shortest Path in an Unweighted Graph

the weighted graph:
   A -> B(1), C(10)
   B -> A(1), D(20)
   C -> A(10), D(1)
   D -> B(20), C(1)

BFS answer    : A -> B -> D   edges: 2   cost: 21
a cheaper route: A -> C -> D   edges: 2   cost: 11

BFS picked the route with fewer edges and it costs more: True
it is worse by 10 units.

BFS answers 'fewest edges'. on a weighted graph that is a different
question from 'cheapest', so BFS is the wrong tool: use Dijkstra.

Both routes are two edges long, so BFS has no way to prefer either and takes whichever its neighbour order reaches first. BFS is not approximately right on a weighted graph. It is answering a different question.

Which algorithm for which problem

The graphThe questionThe algorithm
Unweightedfewest edgesBFS, chapter 94
All weights equalcheapestBFS, since cheapest equals fewest edges
Weights, all non-negativecheapestDijkstra, chapter 98
Weights, some negativecheapestBellman-Ford, beyond this syllabus
Anydoes a path existBFS or DFS, either will do

The row to remember for an examination is the second one: when every weight is the same, BFS is the correct and faster answer, and reaching for Dijkstra shows the distinction was missed.

Cost

Exactly the BFS cost, since it is BFS: O(V + E) time on an adjacency list, O(V) extra space for the distance and parent maps. Dijkstra is O((V + E) log V), so on an unweighted graph BFS is not only correct but strictly cheaper.

munotes.in319

The Shortest Path in an Unweighted Graph

Quick revision

  • When all edges cost the same, shortest means fewest edges, and BFS already solves it.
  • The proof: BFS reaches level k only after level k-1 is finished, so a first arrival cannot be beaten.
  • Record a parent on first arrival; the parents chain back to give the route.
  • A distance is never revised, which is why the test is "not yet seen".
  • Unreachable must be reported as unreachable, not as a number.
  • DFS gives wrong distances because it commits to a branch before checking for a shorter route.
  • On a weighted graph BFS answers the wrong question: it chose a 2 edge route costing 21 over a 2 edge route costing 11.
  • BFS here is O(V + E), cheaper than Dijkstra's O((V + E) log V), so do not use Dijkstra on an unweighted graph.

Test yourself

1. Why does BFS give the shortest path in an unweighted graph? Because it processes all vertices at level k-1 before any at level k, so a vertex first reached at level k has no shorter route; one would have reached it on an earlier level.

2. How is the route itself recovered? By recording, for each vertex, the vertex it was first reached from, then following those parents back from the target to the start and reversing.

3. Why is a distance never updated once set? Because the first arrival is already the shortest, so any later arrival is at least as long.

4. Why can DFS not be used for shortest paths? It follows one branch to its end before trying another, so the depth at which it first reaches a vertex is whatever that branch gave, not the minimum.

5. What happens if BFS is used on a weighted graph? It returns the route with the fewest edges, which is a different question. In the chapter's example it chose a route costing 21 over one costing 11, both two edges long.

6. The graph is weighted but every weight is 5. Which algorithm? BFS. All weights being equal makes cheapest the same as fewest edges, and BFS is O(V + E) against Dijkstra's O((V + E) log V).

7. How must an unreachable vertex be reported? As unreachable, with no distance. It must never be given a number or cause the algorithm to loop.

Contents This chapter on its own page

munotes.in320

Chapter Ninety-Eight

Dijkstra's Algorithm

Syllabus topic Module 2, "Applications of Graphs like shortest path algorithms"

In one line

Dijkstra's algorithm finds the cheapest route from one source to every other vertex in a graph with non-negative weights, by repeatedly settling the nearest unsettled vertex and improving its neighbours' tentative distances.

The problem

Given a weighted graph and a source s, find for every vertex the minimum total weight of a path from s to it. "Cheapest", not "fewest edges": chapter 97 showed those are different questions as soon as the weights differ.

The idea, in words before code

Keep a tentative distance for every vertex: 0 for the source, infinity for the rest. Keep a set of settled vertices whose distance is known to be final.

Then repeat: take the unsettled vertex with the smallest tentative distance, declare it settled, and for each of its neighbours check whether going through it is cheaper than that neighbour's current tentative distance. If it is, lower the neighbour's distance. That check is called relaxing the edge.

relax(u, v, weight):

if distance[u] + weight < distance[v]:

distance[v] = distance[u] + weight

parent[v] = u

Why settling the nearest is safe. Let u be the nearest unsettled vertex. Any other route to u must pass through some unsettled vertex first, and that vertex is at least as far as u, so the route is at least as long, because no edge can reduce a total. That last clause is the whole dependence on non-negative weights, and it is the sentence to write in an answer.

The algorithm, traced step by step

import heapq

INF = float("inf")

GRAPH = {
    "A": [("B", 4), ("C", 2)],
    "B": [("A", 4), ("C", 1), ("D", 5)],
    "C": [("A", 2), ("B", 1), ("D", 8), ("E", 10)],
    "D": [("B", 5), ("C", 8), ("E", 2), ("F", 6)],
    "E": [("C", 10), ("D", 2), ("F", 3)],
    "F": [("D", 6), ("E", 3)],
}

print("the weighted graph:")
for v in sorted(GRAPH):
    print("   %s -> %s" % (v, ", ".join("%s(%d)" % p for p in sorted(GRAPH[v]))))
print()


def dijkstra(graph, source, trace=False):
    distance = {v: INF for v in graph}
    distance[source] = 0
    parent = {v: None for v in graph}
    settled = set()
    heap = [(0, source)]
    pops, stale, relaxations = 0, 0, 0

    if trace:
        print("%-8s %-7s %s" % ("settled", "its d", "distances after relaxing"))

    while heap:
        d, u = heapq.heappop(heap)
        pops += 1
        if u in settled:
            stale += 1                    # a superseded entry: skip it
            continue
        settled.add(u)
        for w, weight in graph[u]:
            if w in settled:
                continue
            if d + weight < distance[w]:
                distance[w] = d + weight
                parent[w] = u
                relaxations += 1
                heapq.heappush(heap, (distance[w], w))
        if trace:
            shown = " ".join("%s=%s" % (v, "inf" if distance[v] == INF else distance[v])
                             for v in sorted(graph))
            print("%-8s %-7d %s" % (u, d, shown))

    return distance, parent, {"pops": pops, "stale": stale,
                              "relaxations": relaxations}


distance, parent, counts = dijkstra(GRAPH, "A", trace=True)
print()


def route(parent, v):
    steps = []
    while v is not None:
        steps.append(v)
        v = parent[v]
    return list(reversed(steps))


print("%-8s %8s   %s" % ("vertex", "cost", "cheapest route"))
for v in sorted(GRAPH):
    print("%-8s %8d   %s" % (v, distance[v], " -> ".join(route(parent, v))))

print()
print("heap pops %d, of which %d were superseded entries; %d edges relaxed"
      % (counts["pops"], counts["stale"], counts["relaxations"]))
munotes.in321

Dijkstra's Algorithm

the weighted graph:
   A -> B(4), C(2)
   B -> A(4), C(1), D(5)
   C -> A(2), B(1), D(8), E(10)
   D -> B(5), C(8), E(2), F(6)
   E -> C(10), D(2), F(3)
   F -> D(6), E(3)

settled  its d   distances after relaxing
A        0       A=0 B=4 C=2 D=inf E=inf F=inf
C        2       A=0 B=3 C=2 D=10 E=12 F=inf
B        3       A=0 B=3 C=2 D=8 E=12 F=inf
D        8       A=0 B=3 C=2 D=8 E=10 F=14
E        10      A=0 B=3 C=2 D=8 E=10 F=13
F        13      A=0 B=3 C=2 D=8 E=10 F=13

vertex       cost   cheapest route
A               0   A
B               3   A -> C -> B
C               2   A -> C
D               8   A -> C -> B -> D
E              10   A -> C -> B -> D -> E
F              13   A -> C -> B -> D -> E -> F

heap pops 10, of which 4 were superseded entries; 9 edges relaxed

Look at B. Its first tentative distance is 4, straight from A. When C is settled at 2, the edge C to B of weight 1 relaxes B down to 3. The direct edge was not the cheapest route, and that is the entire reason the algorithm has a relax step rather than simply taking the first edge it finds.

Look at D. Through C it would be 2 + 8 = 10. Through B it is 3 + 5 = 8, so 8 wins.

The answer checked by an independent method

The book's rule: a result worth printing is worth proving twice (FINDINGS 3.5 and 3.6 were both caught this way). Dijkstra is checked here against brute force over every simple path.

import heapq
import itertools

INF = float("inf")

GRAPH = {
    "A": [("B", 4), ("C", 2)],
    "B": [("A", 4), ("C", 1), ("D", 5)],
    "C": [("A", 2), ("B", 1), ("D", 8), ("E", 10)],
    "D": [("B", 5), ("C", 8), ("E", 2), ("F", 6)],
    "E": [("C", 10), ("D", 2), ("F", 3)],
    "F": [("D", 6), ("E", 3)],
}


def dijkstra(graph, source):
    distance = {v: INF for v in graph}
    distance[source] = 0
    settled = set()
    heap = [(0, source)]
    while heap:
        d, u = heapq.heappop(heap)
        if u in settled:
            continue
        settled.add(u)
        for w, weight in graph[u]:
            if d + weight < distance[w]:
                distance[w] = d + weight
                heapq.heappush(heap, (distance[w], w))
    return distance


def weight_of(graph, path):
    total = 0
    for a, b in zip(path, path[1:]):
        found = [w for n, w in graph[a] if n == b]
        if not found:
            return None
        total += min(found)
    return total


def brute_force(graph, source):
    """Every simple path, scored. Only possible on a tiny graph, which is the point."""
    others = [v for v in graph if v != source]
    best = {source: 0}
    for size in range(1, len(others) + 1):
        for middle in itertools.permutations(others, size):
            path = (source,) + middle
            total = weight_of(graph, path)
            if total is not None:
                end = path[-1]
                if end not in best or total < best[end]:
                    best[end] = total
    return best


fast = dijkstra(GRAPH, "A")
slow = brute_force(GRAPH, "A")

print("%-8s %12s %14s %s" % ("vertex", "Dijkstra", "brute force", "agree"))
for v in sorted(GRAPH):
    print("%-8s %12d %14d %s" % (v, fast[v], slow[v], fast[v] == slow[v]))
print()
print("every distance agrees:", all(fast[v] == slow[v] for v in GRAPH))
print("brute force tried every ordering of the other 5 vertices;")
print("Dijkstra settled each vertex once. the answers are identical.")
munotes.in322

Dijkstra's Algorithm

vertex       Dijkstra    brute force agree
A                   0              0 True
B                   3              3 True
C                   2              2 True
D                   8              8 True
E                  10             10 True
F                  13             13 True

every distance agrees: True
brute force tried every ordering of the other 5 vertices;
Dijkstra settled each vertex once. the answers are identical.

Negative weights break it, and here is the proof

import heapq

INF = float("inf")

# A directed graph with ONE negative edge.
ARCS = {
    "A": [("B", 1), ("C", 2)],
    "B": [],
    "C": [("B", -2)],
}

print("a directed graph with one negative edge:")
for v in sorted(ARCS):
    arcs = ", ".join("%s(%d)" % p for p in ARCS[v]) or "(none)"
    print("   %s -> %s" % (v, arcs))
print()


def dijkstra(graph, source):
    distance = {v: INF for v in graph}
    distance[source] = 0
    settled = set()
    heap = [(0, source)]
    order = []
    while heap:
        d, u = heapq.heappop(heap)
        if u in settled:
            continue
        settled.add(u)
        order.append((u, d))
        for w, weight in graph[u]:
            if w in settled:               # settled means FINAL: this is the flaw
                continue
            if d + weight < distance[w]:
                distance[w] = d + weight
                heapq.heappush(heap, (distance[w], w))
    return distance, order


def truth(graph, source):
    """Relax every edge V-1 times: correct even with negative edges."""
    distance = {v: INF for v in graph}
    distance[source] = 0
    for _round in range(len(graph) - 1):
        for u in graph:
            if distance[u] == INF:
                continue
            for w, weight in graph[u]:
                if distance[u] + weight < distance[w]:
                    distance[w] = distance[u] + weight
    return distance


wrong, order = dijkstra(ARCS, "A")
right = truth(ARCS, "A")

print("Dijkstra settled, in order:",
      ", ".join("%s at %d" % (v, d) for v, d in order))
print()
print("%-8s %12s %14s %s" % ("vertex", "Dijkstra", "the truth", "correct"))
for v in sorted(ARCS):
    print("%-8s %12d %14d %s" % (v, wrong[v], right[v], wrong[v] == right[v]))
print()
print("A to B via C costs 2 + (-2) =", 2 + (-2), "which beats the direct edge of 1.")
print("Dijkstra settled B at 1 BEFORE it ever looked at C,")
print("and a settled vertex is never revisited, so the 0 was never found.")
print("Dijkstra is wrong by", wrong["B"] - right["B"], "on vertex B.")
munotes.in323

Dijkstra's Algorithm

a directed graph with one negative edge:
   A -> B(1), C(2)
   B -> (none)
   C -> B(-2)

Dijkstra settled, in order: A at 0, B at 1, C at 2

vertex       Dijkstra      the truth correct
A                   0              0 True
B                   1              0 False
C                   2              2 True

A to B via C costs 2 + (-2) = 0 which beats the direct edge of 1.
Dijkstra settled B at 1 BEFORE it ever looked at C,
and a settled vertex is never revisited, so the 0 was never found.
Dijkstra is wrong by 1 on vertex B.

That is the failure in full. B is settled at 1 before C is examined, and settling means final, so the cheaper route through C is never considered. The algorithm does not crash or warn. It returns a number that is simply wrong.

The invariant that justified settling the nearest vertex was: a route through a farther vertex cannot be shorter. A negative edge makes that false, and the algorithm has no defence. The correct algorithm for negative weights is Bellman-Ford, which relaxes every edge V-1 times, as the checking function above does; it is beyond this syllabus but worth naming.

Two implementations, and which is faster

The step "take the unsettled vertex with the smallest tentative distance" can be done two ways.

Scan an array of distances, O(V) per selection, V selections, so O(V squared) in total.

Keep a min-heap, O(log V) per operation. Each vertex is pushed when its distance improves, so at most E pushes, giving O((V + E) log V).

import heapq

INF = float("inf")


def ring_with_chords(n):
    """A sparse weighted graph, built arithmetically: no random numbers."""
    graph = {i: [] for i in range(n)}
    pairs = {}
    for i in range(n):
        for j in ((i + 1) % n, (i * 7 + 3) % n):
            if i != j:
                key = (min(i, j), max(i, j))
                pairs[key] = 1 + (key[0] * 3 + key[1] * 5) % 20
    for (a, b), w in sorted(pairs.items()):
        graph[a].append((b, w))
        graph[b].append((a, w))
    return graph, len(pairs)


def dijkstra_scan(graph, source):
    """Selection by scanning. Counts the comparisons it makes."""
    distance = {v: INF for v in graph}
    distance[source] = 0
    settled = set()
    scans = 0
    while len(settled) < len(graph):
        best, best_v = INF, None
        for v in graph:
            scans += 1
            if v not in settled and distance[v] < best:
                best, best_v = distance[v], v
        if best_v is None:
            break
        settled.add(best_v)
        for w, weight in graph[best_v]:
            if best + weight < distance[w]:
                distance[w] = best + weight
    return distance, scans


def dijkstra_heap(graph, source):
    """Selection by min-heap. Counts the heap operations it makes."""
    distance = {v: INF for v in graph}
    distance[source] = 0
    settled = set()
    heap = [(0, source)]
    operations = 1
    while heap:
        d, u = heapq.heappop(heap)
        operations += 1
        if u in settled:
            continue
        settled.add(u)
        for w, weight in graph[u]:
            if d + weight < distance[w]:
                distance[w] = d + weight
                heapq.heappush(heap, (distance[w], w))
                operations += 1
    return distance, operations


print("%8s %8s %16s %18s" % ("vertices", "edges", "scan comparisons", "heap operations"))
for n in (50, 200, 800):
    graph, e = ring_with_chords(n)
    d_scan, scans = dijkstra_scan(graph, 0)
    d_heap, operations = dijkstra_heap(graph, 0)
    assert d_scan == d_heap, "both implementations must give the same distances"
    print("%8d %8d %16d %18d" % (n, e, scans, operations))

print()
print("the two implementations agree on every distance at every size.")
print("the scan does V squared comparisons; the heap's work follows V + E.")
print("so: heap for a sparse graph, and a plain scan is fine when the graph")
print("is dense, because then E is near V squared anyway.")
munotes.in324

Dijkstra's Algorithm

vertices    edges scan comparisons    heap operations
      50       95             2500                128
     200      392            40000                498
     800     1596           640000               1956

the two implementations agree on every distance at every size.
the scan does V squared comparisons; the heap's work follows V + E.
so: heap for a sparse graph, and a plain scan is fine when the graph
is dense, because then E is near V squared anyway.
Selection methodComplexityBest for
Scan the distance arrayO(V squared)dense graphs, small V, simple code
Binary heap, chapters 81 to 85O((V + E) log V)sparse graphs, the usual case
Fibonacci heapO(E + V log V)theoretical, rarely used in practice

The stale entry, and why the heap version does not delete

A binary heap cannot cheaply lower the priority of an item already inside it (chapter 80 said so). So the practical version pushes a new entry and leaves the old one. When the old one is popped, its vertex is already settled and it is skipped.

That is the same lazy deletion as chapter 93: do not remove, mark and ignore. The trace above counted the superseded pops, 4 of the 10, so the cost of the trick is visible rather than assumed. The heap can hold up to E entries rather than V, but since E is at most V squared, log E is at most 2 log V, so each operation is still O(log V) and the bound stays O((V + E) log V). It is the right trade, because a heap that can lower an item's priority in place needs considerably more code for the same bound.

munotes.in325

Dijkstra's Algorithm

Where Dijkstra is used

Route finding. Maps, navigation, railway and flight connections, with weight as distance, time or fare. A real map service uses heavily tuned variants, but Dijkstra is the base.

Network routing. OSPF, the protocol that routes traffic inside large networks, computes shortest paths with Dijkstra over link costs.

Any cheapest-sequence problem. Currency conversion chains, puzzle states where moves have different costs, project paths with weighted steps.

As a building block. Many algorithms call a shortest path routine, and when weights are non-negative that routine is Dijkstra.

Quick revision

  • Dijkstra solves single source cheapest paths for non-negative weights.
  • Settle the nearest unsettled vertex, then relax its edges: lower a neighbour's distance if going through this vertex is cheaper.
  • Settling the nearest is safe only because no edge can reduce a total, which is exactly the non-negative condition.
  • B's distance fell from 4 to 3 once C was settled: the relax step is the algorithm, not decoration.
  • With a negative edge Dijkstra settled B at 1 when the truth was 0, and reported no error: use Bellman-Ford instead.
  • Selection by array scan is O(V squared); selection by binary heap is O((V + E) log V).
  • Use the heap for sparse graphs; the scan is acceptable on dense graphs, where E approaches V squared.
  • The heap version uses lazy deletion, pushing an improved entry and skipping superseded pops, exactly as chapter 93 did for vertices.
  • Uses: maps and navigation, OSPF network routing, and any cheapest-sequence problem.

Test yourself

1. What problem does Dijkstra's algorithm solve, and under what condition? The cheapest path from one source to every other vertex, provided no edge weight is negative.

2. What does relaxing an edge mean? Checking whether reaching v through u is cheaper than v's current tentative distance, and if so lowering v's distance and recording u as its parent.

3. Why is it safe to settle the nearest unsettled vertex? Because any other route to it must first pass through an unsettled vertex that is at least as far, and no edge can reduce a total, so that route cannot be shorter.

4. Show with an example that the relax step is necessary. In the chapter's graph B starts at 4 by the direct edge from A, and falls to 3 when C is settled at 2 and the edge C to B of weight 1 is relaxed.

munotes.in326

Dijkstra's Algorithm

5. Explain precisely why a negative edge breaks the algorithm. Settling means final. In the chapter's example B was settled at 1 before C was examined, so the route A, C, B costing 2 + (-2) = 0 was never considered. Dijkstra returned 1, which is wrong by 1, with no error.

6. Which algorithm handles negative weights, and how? Bellman-Ford, by relaxing every edge V-1 times rather than settling vertices in order of distance.

7. Give both complexities and say when each implementation is preferred. O(V squared) when the nearest vertex is found by scanning the distance array, preferred for dense graphs and small V. O((V + E) log V) with a binary heap, preferred for sparse graphs, which is the usual case.

8. Why does the heap version push a duplicate entry instead of updating one? Because a binary heap cannot cheaply lower the priority of an item already inside it. The improved distance is pushed as a new entry and the superseded entry is skipped when it is popped, which is lazy deletion.

Contents This chapter on its own page

munotes.in327

Chapter Ninety-Nine

The Idea of Hashing: A Key Turned Into an Address

Syllabus topic Module 2, "Concept of hashing"

In one line

Hashing computes where a record should live from the record's own key, so finding it takes one calculation instead of a search.

The problem, and what the book has offered so far

A single question keeps coming back: given a key, find its record.

StructureSearchWhy
Unsorted array or linked listO(n)every element may have to be compared
Sorted array, binary searchO(log n)halve the range each comparison
Binary search treeO(log n) average, O(n) worstchapter 68: it can degenerate
AVL treeO(log n) guaranteedchapter 71: balance is maintained

O(log n) looks small, and it is. But every one of those structures works the same way: it compares the key against keys that are already stored and uses the answer to decide where to look next.

Hashing does not compare. It calculates.

The idea, built up from something obvious

Suppose the keys are roll numbers 0 to 99 and there are at most 100 students. Then store the student with roll number k at index k. Looking up roll number 57 is table[57]. One step. No comparisons, no searching. This is called a direct address table and it is as fast as a lookup can be.

table = [None] * 100

for roll, name in ((57, "Aarti"), (3, "Bhavesh"), (91, "Chetna")):
    table[roll] = name                      # the key IS the address

print("table[57] :", table[57])
print("table[3]  :", table[3])
print("table[91] :", table[91])
print("table[20] :", table[20], "(nothing stored there)")
print()
print("lookups are one index each, no comparisons at all.")
print("slots used     :", sum(1 for x in table if x is not None))
print("slots allocated:", len(table))
print("wasted         : %.0f%%" % (100 * sum(1 for x in table if x is None) / len(table)))
table[57] : Aarti
table[3]  : Bhavesh
table[91] : Chetna
table[20] : None (nothing stored there)

lookups are one index each, no comparisons at all.
slots used     : 3
slots allocated: 100
wasted         : 97%

So why is the whole subject not just this? Because real keys are not small integers starting at zero.

A 10 digit mobile number. A direct address table needs 10 billion slots to hold a few thousand records.

A 12 digit Aadhaar number. A thousand billion slots.

A PRN or an email address. Not a number at all.

The key space is enormous and the number of actual records is tiny. A direct address table is correct and impossible.

print("%-26s %22s %14s %s" % ("key", "possible values", "records", "slots per record"))
cases = [("a 2 digit roll number", 100, 100),
         ("a 6 digit PRN", 10 ** 6, 5000),
         ("a 10 digit mobile number", 10 ** 10, 5000),
         ("a 12 digit Aadhaar number", 10 ** 12, 5000)]
for name, space, records in cases:
    print("%-26s %22d %14d %16.0f" % (name, space, records, space / records))

print()
print("a direct address table for mobile numbers would allocate")
print(10 ** 10, "slots to store", 5000, "records.")
print("that is %.5f%% of the table used." % (100 * 5000 / 10 ** 10))
munotes.in328

The Idea of Hashing: A Key Turned Into an Address

key                               possible values        records slots per record
a 2 digit roll number                         100            100                1
a 6 digit PRN                             1000000           5000              200
a 10 digit mobile number              10000000000           5000          2000000
a 12 digit Aadhaar number           1000000000000           5000        200000000

a direct address table for mobile numbers would allocate
10000000000 slots to store 5000 records.
that is 0.00005% of the table used.

Hashing: shrink the key space with a function

Keep the idea, drop the waste. Choose a table of some practical size m, and a function that turns any key into a slot number in 0 to m-1.

h : the set of all possible keys -> { 0, 1, 2, ... , m-1 }

address = h(key)

That function is the hash function, its output is the hash value or hash code, the array is the hash table, and each slot is a bucket.

The simplest hash function is the remainder, called the division method:

h(k) = k mod m

A mobile number of 10 digits, a table of 1,000 slots, and the address is the last three digits of the number. One modulo. One index.

SIZE = 11
table = [None] * SIZE


def h(key):
    return key % SIZE


records = [(9820012345, "Aarti"), (9867054321, "Bhavesh"),
           (9123456789, "Chetna"), (7700112233, "Devdatta")]

print("size of table :", SIZE)
print()
print("%-14s %8s %-10s %s" % ("key", "h(key)", "name", "slot already taken?"))
clashes = []
for key, name in records:
    slot = h(key)
    taken = table[slot] is not None
    if taken:
        clashes.append((table[slot][1], name, slot))
    table[slot] = (key, name)
    print("%-14d %8d %-10s %s" % (key, slot, name, "YES" if taken else "no"))

print()
print("the table:")
for i, cell in enumerate(table):
    print("   [%2d] %s" % (i, cell if cell else ""))

print()
look_for = 9123456789
slot = h(look_for)
print("looking up %d: compute h = %d, read table[%d]" % (look_for, slot, slot))
print("found:", table[slot][1])
print()
print("that lookup did ONE modulo and ONE array read.")
print("it did not depend on how many records are stored.")
print()
for lost, kept, slot in clashes:
    print("but look at slot %d: %s was overwritten by %s." % (slot, lost, kept))
print("a record is missing from the table:",
      not any(cell and cell[1] == "Aarti" for cell in table))
print("that is a COLLISION, and it happened on the first four keys written")
print("into this chapter. it is not bad luck. chapter 103 proves that.")
size of table : 11

key              h(key) name       slot already taken?
9820012345            0 Aarti      no
9867054321            3 Bhavesh    no
9123456789            7 Chetna     no
7700112233            0 Devdatta   YES

the table:
   [ 0] (7700112233, 'Devdatta')
   [ 1]
   [ 2]
   [ 3] (9867054321, 'Bhavesh')
   [ 4]
   [ 5]
   [ 6]
   [ 7] (9123456789, 'Chetna')
   [ 8]
   [ 9]
   [10]

looking up 9123456789: compute h = 7, read table[7]
found: Chetna

that lookup did ONE modulo and ONE array read.
it did not depend on how many records are stored.

but look at slot 0: Aarti was overwritten by Devdatta.
a record is missing from the table: True
that is a COLLISION, and it happened on the first four keys written
into this chapter. it is not bad luck. chapter 103 proves that.
munotes.in329

The Idea of Hashing: A Key Turned Into an Address

Every lookup is one calculation, and that is the whole concept of hashing.

Now read the last four lines of that output. Two of the four keys hashed to slot 0, so storing Devdatta destroyed Aarti, and a lookup of Aarti's number would now return Devdatta's record. Nothing warned about it.

That was not arranged. Those were the first four numbers written into this chapter, and two of them clashed in a table with slots to spare. Two keys sharing a slot is called a collision, and it is the normal case rather than an accident. Chapter 103 proves it is unavoidable, and chapters 104 to 106 are the three standard ways to handle it. Until then, every hash table in this book stores at most one record per slot and so is incomplete by design.

How much is actually saved

SIZE = 1009                       # a prime, for the reason chapter 101 gives
N = 20000


def build_list(n):
    return [i * 7 + 3 for i in range(n)]


def build_hash(keys, size):
    table = [[] for _ in range(size)]
    for k in keys:
        table[k % size].append(k)
    return table


keys = build_list(N)
table = build_hash(keys, SIZE)


def linear_search(items, target):
    comparisons = 0
    for item in items:
        comparisons += 1
        if item == target:
            return comparisons
    return comparisons


def binary_search(items, target):
    comparisons = 0
    low, high = 0, len(items) - 1
    while low <= high:
        mid = (low + high) // 2
        comparisons += 1
        if items[mid] == target:
            return comparisons
        if items[mid] < target:
            low = mid + 1
        else:
            high = mid - 1
    return comparisons


def hash_search(table, size, target):
    comparisons = 0
    for item in table[target % size]:
        comparisons += 1
        if item == target:
            return comparisons
    return comparisons


probe_keys = [keys[i] for i in range(0, N, N // 20)]   # 20 keys spread across

totals = {"linear": 0, "binary": 0, "hash": 0}
worst = {"linear": 0, "binary": 0, "hash": 0}
for target in probe_keys:
    for name, count in (("linear", linear_search(keys, target)),
                        ("binary", binary_search(keys, target)),
                        ("hash", hash_search(table, SIZE, target))):
        totals[name] += count
        worst[name] = max(worst[name], count)

print("%d records, %d lookups of keys spread across the whole set"
      % (N, len(probe_keys)))
print()
print("%-24s %14s %10s" % ("method", "avg comparisons", "worst"))
for name, label in (("linear", "linear search, a list"),
                    ("binary", "binary search, sorted"),
                    ("hash", "hash table, %d buckets" % SIZE)):
    print("%-24s %14.1f %10d"
          % (label, totals[name] / len(probe_keys), worst[name]))

print()
print("records per bucket:", N // SIZE, "to", max(len(b) for b in table))
print("the hash lookup's cost is the bucket length, not the record count.")
print("the list's cost IS the record count.")
munotes.in330

The Idea of Hashing: A Key Turned Into an Address

20000 records, 20 lookups of keys spread across the whole set

method                   avg comparisons      worst
linear search, a list            9501.0      19001
binary search, sorted              13.6         15
hash table, 1009 buckets            9.6         19

records per bucket: 19 to 20
the hash lookup's cost is the bucket length, not the record count.
the list's cost IS the record count.

The hash table's average is small and, crucially, it does not grow with the number of records as long as the table grows with them. That is the claim chapter 107 makes precise with the load factor.

Why this escapes a known limit

Any search that works by comparing keys must do at least about log2(n) comparisons in the worst case: each comparison has two outcomes, so k comparisons can distinguish at most 2 to the power k possibilities, and distinguishing n records needs 2 to the power k to be at least n.

import math

print("%12s %22s %18s" % ("records", "comparisons needed", "2^that"))
for n in (10, 1000, 1000000, 10 ** 9):
    k = math.ceil(math.log2(n))
    print("%12d %22d %18d" % (n, k, 2 ** k))

print()
print("a comparison has 2 outcomes, so k comparisons separate at most 2^k cases.")
print("so no comparison-based search can beat about log2(n). binary search and")
print("a balanced tree both sit at that limit; they cannot be improved on.")
print()
print("hashing does not compare keys to locate a record. it computes an")
print("address. so the limit simply does not apply to it, which is why O(1)")
print("average is possible at all.")
     records     comparisons needed             2^that
          10                      4                 16
        1000                     10               1024
     1000000                     20            1048576
  1000000000                     30         1073741824

a comparison has 2 outcomes, so k comparisons separate at most 2^k cases.
so no comparison-based search can beat about log2(n). binary search and
a balanced tree both sit at that limit; they cannot be improved on.

hashing does not compare keys to locate a record. it computes an
address. so the limit simply does not apply to it, which is why O(1)
average is possible at all.

This is worth saying in an answer because it explains why hashing is different in kind, not merely faster by a constant. Binary search and a balanced tree are already optimal among comparison based methods. Hashing is not a better comparison based method; it is not one.

munotes.in331

The Idea of Hashing: A Key Turned Into an Address

What it costs

Nothing is free, and the honest list is short:

Two keys can hash to the same slot. That is a collision, it is not rare, and chapters 103 to 106 are entirely about handling it.

The order is gone. A hash table has no sorted order, no smallest, no largest, no range query, no in-order walk. A balanced tree keeps all of those. Chapter 108 weighs this properly.

The guarantee is average, not worst case. A bad hash function or an unlucky key set can push every record into one bucket and the lookup becomes O(n).

Empty space is required. A hash table must be kept well under full to stay fast, so some memory is deliberately unused. Chapter 107.

Quick revision

  • Hashing computes a record's address from its key instead of searching for it.
  • A direct address table uses the key itself as the index: perfect, but it needs a slot for every possible key.
  • A 10 digit mobile number has 10 billion possible values, so direct addressing is impossible for real keys.
  • A hash function maps any key into 0 to m-1; the array is the hash table and each slot is a bucket.
  • The simplest hash function is the division method, h(k) = k mod m.
  • A lookup is one calculation and one array access, independent of how many records are stored.
  • Comparison based search cannot beat about log2(n), because k comparisons separate at most 2 to the power k cases; hashing does not compare, so the limit does not apply.
  • The costs are collisions, the loss of all ordering, an average rather than worst case guarantee, and deliberately unused space.

Test yourself

1. What is the central idea of hashing, in one sentence? The address of a record is computed from its key, so it is found by calculation rather than by searching.

2. What is a direct address table, and why can it not be used for real keys? A table in which the key itself is the index. It needs one slot for every possible key, which for a 10 digit mobile number is 10 billion slots to hold a few thousand records.

3. Define hash function, hash table and bucket. The hash function maps any key to a slot number in 0 to m-1. The hash table is the array of m slots. A bucket is one slot of that array.

4. Give the division method and apply it. h(k) = k mod m. With m = 11 and k = 9123456789 the slot is 9123456789 mod 11.

munotes.in332

The Idea of Hashing: A Key Turned Into an Address

5. Why can no comparison based search beat about log2(n)? A comparison has two outcomes, so k comparisons distinguish at most 2 to the power k cases. To distinguish n records, 2 to the power k must be at least n, so k is at least log2(n).

6. Why does hashing escape that limit? Because it does not compare keys to find a record; it computes the address directly, so the counting argument about comparisons does not apply.

7. Name the four costs of hashing. Collisions; the complete loss of ordering, so no sorted walk, minimum, maximum or range query; a guarantee that is average rather than worst case; and the need to keep the table well below full, so some space is deliberately unused.

Contents This chapter on its own page

munotes.in333

Chapter One Hundred

The Hash Table ADT

Syllabus topic Module 2, "Hash Table ADT"

In one line

The hash table ADT stores key and value pairs and promises insert, search and delete in O(1) average time, while promising nothing at all about order.

What the ADT is called elsewhere

The same ADT appears under several names, and a question may use any of them:

NameWhere it is used
Hash tablewhen the implementation is meant
DictionaryPython, and most textbooks
Map, or associative arrayJava, C++, and most of the literature
Symbol tablecompilers, the original application

Strictly, dictionary is the ADT and hash table is one implementation of it. A balanced tree is another. An answer that makes that distinction is making the point of chapter 7.

The operations

OperationMeaningAverageWorst
insert(key, value)store the pair; replace the value if the key is presentO(1)O(n)
search(key)the value stored for key, or "absent"O(1)O(n)
delete(key)remove the pairO(1)O(n)
contains(key)whether the key is presentO(1)O(n)
size()how many pairs are storedO(1)O(1)
is_empty()whether size is zeroO(1)O(1)
keys(), values(), items()everything stored, in no particular orderO(m + n)O(m + n)

Two things in that table are the whole character of the ADT.

The worst case is O(n), not O(1). Every textbook says O(1) and every textbook means on average. If every key lands in one bucket the table degenerates into a list. Chapters 102 and 107 are about keeping that from happening, and an answer that writes O(1) without the word average has lost the distinction.

keys() is in no particular order, and the cost is O(m + n) because every one of the m buckets must be looked at, not only the n records.

The decisions an implementation has to make, and declare

An ADT must be honest about the choices it has made (chapter 10):

A repeated key: replace or refuse? Most dictionaries replace. Some multi-maps keep both. The ADT must say which, because a caller cannot guess.

A missing key on search: return a sentinel, or raise an error? Returning None is convenient until a stored value is legitimately None, at which point absent and present become indistinguishable. This implementation returns a distinct MISSING object.

A missing key on delete: error, or silently nothing? Either is defensible; it must be stated.

Which keys are allowed? A key must be hashable, which in practice means immutable. This is not pedantry, and the chapter proves it below.

A working implementation

class Missing:
    """A distinct 'not found' value, so a stored None is still findable."""

    def __repr__(self):
        return "MISSING"


MISSING = Missing()


class HashTable:
    """The hash table ADT. Collisions are handled by chaining, chapter 104."""

    def __init__(self, buckets=11):
        self._buckets = [[] for _ in range(buckets)]
        self._count = 0

    # ----- the hidden part -------------------------------------------------
    def _slot(self, key):
        return hash(key) % len(self._buckets)

    # ----- the promise -----------------------------------------------------
    def insert(self, key, value):
        """Store the pair. A repeated key REPLACES its value, and says so."""
        chain = self._buckets[self._slot(key)]
        for i, (k, _v) in enumerate(chain):
            if k == key:
                chain[i] = (key, value)
                return "replaced"
        chain.append((key, value))
        self._count += 1
        return "inserted"

    def search(self, key):
        for k, v in self._buckets[self._slot(key)]:
            if k == key:
                return v
        return MISSING

    def contains(self, key):
        return self.search(key) is not MISSING

    def delete(self, key):
        chain = self._buckets[self._slot(key)]
        for i, (k, _v) in enumerate(chain):
            if k == key:
                chain.pop(i)
                self._count -= 1
                return True
        return False                      # declared: a missing key is not an error

    def size(self):
        return self._count

    def is_empty(self):
        return self._count == 0

    def items(self):
        """Everything stored, in NO guaranteed order."""
        return [pair for chain in self._buckets for pair in chain]

    def keys(self):
        return [k for k, _v in self.items()]


marks = HashTable(buckets=7)

print("is_empty on a new table:", marks.is_empty())
for name, mark in (("Aarti", 78), ("Bhavesh", 65), ("Chetna", 91),
                   ("Devdatta", 54), ("Esha", 88)):
    print("   insert %-9s -> %s" % (name, marks.insert(name, mark)))

print()
print("size             :", marks.size())
print("search Chetna    :", marks.search("Chetna"))
print("search Farhan    :", marks.search("Farhan"), "(absent, a distinct value)")
print("contains Esha    :", marks.contains("Esha"))
print()
print("insert Aarti again:", marks.insert("Aarti", 82), "- a repeat REPLACES")
print("Aarti is now      :", marks.search("Aarti"))
print("size is unchanged :", marks.size())
print()
print("delete Devdatta  :", marks.delete("Devdatta"))
print("delete Devdatta  :", marks.delete("Devdatta"), "- again, declared not an error")
print("size             :", marks.size())
print()
print("a stored None is still distinguishable from absent:")
marks.insert("Gauri", None)
print("   search Gauri  :", marks.search("Gauri"), " contains:", marks.contains("Gauri"))
print("   search Farhan :", marks.search("Farhan"), " contains:", marks.contains("Farhan"))
munotes.in334

The Hash Table ADT

is_empty on a new table: True
   insert Aarti     -> inserted
   insert Bhavesh   -> inserted
   insert Chetna    -> inserted
   insert Devdatta  -> inserted
   insert Esha      -> inserted

size             : 5
search Chetna    : 91
search Farhan    : MISSING (absent, a distinct value)
contains Esha    : True

insert Aarti again: replaced - a repeat REPLACES
Aarti is now      : 82
size is unchanged : 5

delete Devdatta  : True
delete Devdatta  : False - again, declared not an error
size             : 4

a stored None is still distinguishable from absent:
   search Gauri  : None  contains: True
   search Farhan : MISSING  contains: False

Every operation on that table did one hash calculation and then looked only inside one bucket. The number of records stored never entered into it.

The order is not merely unspecified. It changes.

class HashTable:
    def __init__(self, buckets=11):
        self._buckets = [[] for _ in range(buckets)]

    def insert(self, key, value):
        self._buckets[key % len(self._buckets)].append((key, value))

    def keys(self):
        return [k for chain in self._buckets for k, _v in chain]


KEYS = [15, 3, 27, 8, 41, 19, 6]

for buckets in (7, 11, 13):
    t = HashTable(buckets)
    for k in KEYS:
        t.insert(k, str(k))
    print("the same keys in a %2d bucket table come out as %s"
          % (buckets, t.keys()))

print()
print("inserted in this order      :", KEYS)
print("sorted, which a tree gives  :", sorted(KEYS))
print()
print("none of the three listings matches the insertion order,")
print("and none of them is sorted. the order is an accident of the")
print("table size and the hash function, so it must never be relied on.")
munotes.in335

The Hash Table ADT

the same keys in a  7 bucket table come out as [15, 8, 3, 19, 27, 41, 6]
the same keys in a 11 bucket table come out as [3, 15, 27, 6, 8, 41, 19]
the same keys in a 13 bucket table come out as [27, 15, 41, 3, 19, 6, 8]

inserted in this order      : [15, 3, 27, 8, 41, 19, 6]
sorted, which a tree gives  : [3, 6, 8, 15, 19, 27, 41]

none of the three listings matches the insertion order,
and none of them is sorted. the order is an accident of the
table size and the hash function, so it must never be relied on.

This is the ADT's real limitation and it belongs in any comparison answer. A hash table cannot give you the smallest key, the largest key, the keys between two bounds, or the keys in order, at any price better than looking at all of them. A balanced binary search tree gives all four in O(log n). Chapter 108 sets that out properly.

A mutable key destroys the table

The rule "a key must be immutable" sounds like a language detail. It is not. A key's hash decides which bucket holds it, so if the key changes after insertion, the record is in the wrong bucket and can never be found again.

class SimpleTable:
    def __init__(self, buckets=7):
        self._buckets = [[] for _ in range(buckets)]

    def _slot(self, key):
        return sum(key) % len(self._buckets)     # a hash over a list of numbers

    def insert(self, key, value):
        self._buckets[self._slot(key)].append((key, value))

    def search(self, key):
        for k, v in self._buckets[self._slot(key)]:
            if k == key:
                return v
        return "ABSENT"

    def every_record(self):
        return [pair for chain in self._buckets for pair in chain]


t = SimpleTable()
key = [1, 2, 3]                   # a MUTABLE key
t.insert(key, "the record")

print("slot chosen at insert :", t._slot(key))
print("search finds it       :", t.search(key))
print()

key.append(10)                    # the key is changed AFTER insertion
print("the key is mutated to :", key)
print("slot it would now hash to:", t._slot(key))
print("search finds it       :", t.search(key))
print()
print("the record is still in the table:", t.every_record())
print("but it is unreachable: the hash now points at a different bucket.")
print()
print("this is why Python REFUSES a list as a dictionary key:")
refused_with = None
try:
    {}[[1, 2, 3]] = "x"
except Exception as exc:
    refused_with = type(exc).__name__
print("   a list used as a dict key raises:", refused_with)
print("   (the wording of the message differs between Python versions,")
print("    so only the exception TYPE is quoted here.)")
print("the error is not a restriction. it is the language refusing to")
print("let you build the broken table above.")
munotes.in336

The Hash Table ADT

slot chosen at insert : 6
search finds it       : the record

the key is mutated to : [1, 2, 3, 10]
slot it would now hash to: 2
search finds it       : ABSENT

the record is still in the table: [([1, 2, 3, 10], 'the record')]
but it is unreachable: the hash now points at a different bucket.

this is why Python REFUSES a list as a dictionary key:
   a list used as a dict key raises: TypeError
   (the wording of the message differs between Python versions,
    so only the exception TYPE is quoted here.)
the error is not a restriction. it is the language refusing to
let you build the broken table above.

The record is in the table and cannot be found. The language's refusal to accept a list as a key is a correctness guarantee, and that is the answer to give if asked why only immutable types may be keys.

What the ADT does not settle

Deliberately, and this is the point of an ADT:

  • which hash function is used, chapters 101 and 102;
  • how collisions are handled, chapters 104 to 106;
  • how many buckets there are and when that changes, chapter 107.

All three can be replaced without a single caller changing, because the promise, insert, search, delete, is unchanged by any of them. That is chapter 8's argument, and the hash table is its best example in this paper.

Quick revision

  • The hash table ADT stores key and value pairs: insert, search, delete, contains, size, is_empty, keys.
  • The same ADT is called a dictionary, a map, an associative array or a symbol table; strictly, dictionary is the ADT and hash table is one implementation.
  • Insert, search and delete are O(1) AVERAGE and O(n) worst case; writing O(1) without "average" misses the distinction.
  • keys() is O(m + n), because all m buckets must be examined, and the order is not specified.
  • The order actually changes with the table size, so it must never be relied on.
  • A hash table offers no minimum, maximum, range query or sorted walk; a balanced tree gives all four in O(log n).
  • An implementation must declare what it does with a repeated key, a missing key on search, and a missing key on delete.
  • A missing-key sentinel must be distinct, or a stored None becomes indistinguishable from absent.
  • Keys must be immutable: a key mutated after insertion leaves its record in the wrong bucket, present but unreachable.
  • The ADT fixes none of the hash function, the collision strategy or the bucket count, so all three can be changed without touching a caller.
munotes.in337

The Hash Table ADT

Test yourself

1. List the hash table ADT's operations with their average and worst case costs. insert, search, delete and contains are O(1) average and O(n) worst case. size and is_empty are O(1). keys, values and items are O(m + n), where m is the bucket count.

2. Why is the worst case O(n)? If every key hashes to the same bucket the table behaves as a single list, so a search compares every record.

3. Distinguish the dictionary ADT from the hash table. Dictionary is the ADT, the promise of storing and retrieving pairs by key. A hash table is one implementation of it; a balanced binary search tree is another.

4. What order does keys() return, and why does it matter? No specified order. The chapter showed the same seven keys coming out differently from 7, 11 and 13 bucket tables, matching neither insertion order nor sorted order, so any program depending on the order is relying on an accident.

5. Why must a key be immutable? Give the failure. The key's hash chooses its bucket. The chapter mutated a list key after insertion; the record stayed in the old bucket while the hash pointed to a new one, so the record was present in the table and unreachable by search.

6. Why should a missing key not be reported as None? Because None may be a legitimate stored value, and then absent and present cannot be told apart. A distinct sentinel object is needed.

7. Name three decisions the ADT leaves to the implementation. The hash function, the collision handling strategy, and the number of buckets and when to change it.

8. Which questions can a balanced tree answer that a hash table cannot? The smallest key, the largest key, all keys in a range, and all keys in sorted order, each in O(log n) or in output-proportional time.

Contents This chapter on its own page

munotes.in338

Chapter One Hundred One

Hash Functions

Syllabus topic Module 2, "hash functions"

In one line

A hash function turns a key into a slot number; the four classical methods are division, mid-square, folding and multiplication, and strings are handled by combining their characters into a number first.

The job

h : any key -> an integer in { 0, 1, ... , m-1 }

Two absolute requirements, before any question of quality:

It must be deterministic. The same key must always give the same slot, or a stored record can never be found. This rules out anything involving the time, a random number or a memory address.

Its output must be in range. Every method below ends with mod m for exactly this reason.

Method 1: the division method

h(k) = k mod m

The remainder when the key is divided by the table size. The simplest method, the fastest, and the one used by default.

def division(k, m):
    return k % m


M = 11
KEYS = [25, 37, 108, 1001, 56, 77, 92]

print("the division method, h(k) = k mod %d" % M)
print()
print("%8s %10s %8s %s" % ("key", "k mod m", "slot", "the arithmetic"))
for k in KEYS:
    print("%8d %10d %8d   %d = %d x %d + %d"
          % (k, division(k, M), division(k, M), k, M, k // M, k % M))

print()
table = [[] for _ in range(M)]
for k in KEYS:
    table[division(k, M)].append(k)
for i, chain in enumerate(table):
    print("   [%2d] %s" % (i, chain if chain else ""))
print()
print("7 keys in 11 slots, and slot 0 took two of them (77 and 1001).")
print("77 mod 11 =", 77 % 11, "and 1001 mod 11 =", 1001 % 11)
the division method, h(k) = k mod 11

     key    k mod m     slot the arithmetic
      25          3        3   25 = 11 x 2 + 3
      37          4        4   37 = 11 x 3 + 4
     108          9        9   108 = 11 x 9 + 9
    1001          0        0   1001 = 11 x 91 + 0
      56          1        1   56 = 11 x 5 + 1
      77          0        0   77 = 11 x 7 + 0
      92          4        4   92 = 11 x 8 + 4

   [ 0] [1001, 77]
   [ 1] [56]
   [ 2]
   [ 3] [25]
   [ 4] [37, 92]
   [ 5]
   [ 6]
   [ 7]
   [ 8]
   [ 9] [108]
   [10]

7 keys in 11 slots, and slot 0 took two of them (77 and 1001).
77 mod 11 = 0 and 1001 mod 11 = 0

The choice of m is not free, and this is the part worth marks. Two warnings:

Do not use a power of 10. With m = 1000, k mod m is the last three digits, so the rest of the key is ignored entirely.

munotes.in339

Hash Functions

Do not use a power of 2. With m = 16, k mod m is the last four bits, so again most of the key is thrown away.

Use a prime, ideally not close to a power of 2. A prime mixes the whole key into the remainder. Chapter 102 measures how much difference this makes.

KEYS = [i * 10 for i in range(1, 41)]        # 10, 20, 30 ... 400: realistic enough
print("40 keys, all multiples of 10:", KEYS[:6], "...", KEYS[-1])
print()
print("%10s %10s %12s %14s %s"
      % ("table size", "kind", "slots used", "biggest slot", "empty slots"))
for m in (10, 16, 20, 11, 13, 41):
    table = [0] * m
    for k in KEYS:
        table[k % m] += 1
    kind = {10: "power of 10", 16: "power of 2", 20: "even"}.get(m, "PRIME")
    print("%10d %10s %12d %14d %14d"
          % (m, kind, sum(1 for c in table if c), max(table),
             sum(1 for c in table if c == 0)))

print()
print("with m = 10 every one of the 40 keys landed in slot 0.")
print("with m = 41 the 40 keys spread over 40 different slots.")
print("same keys, same method: the table SIZE decided everything.")
40 keys, all multiples of 10: [10, 20, 30, 40, 50, 60] ... 400

table size       kind   slots used   biggest slot empty slots
        10 power of 10            1             40              9
        16 power of 2            8              5              8
        20       even            2             20             18
        11      PRIME           11              4              0
        13      PRIME           13              4              0
        41      PRIME           40              1              1

with m = 10 every one of the 40 keys landed in slot 0.
with m = 41 the 40 keys spread over 40 different slots.
same keys, same method: the table SIZE decided everything.

Method 2: the mid-square method

Square the key, then take some digits from the middle of the square.

h(k) = the middle r digits of (k x k), taken mod m

The reasoning: the middle digits of a square depend on every digit of the key, because every digit contributes to the middle of the product. The ends do not have that property.

def mid_square(k, m, digits=2):
    """Square the key, take `digits` digits from the middle, then mod m."""
    square = str(k * k)
    if len(square) <= digits:
        middle = square
    else:
        start = (len(square) - digits) // 2
        middle = square[start:start + digits]
    return int(middle) % m, square, middle


M = 11
print("the mid-square method, 2 middle digits, then mod %d" % M)
print()
print("%8s %14s %10s %8s" % ("key", "k x k", "middle 2", "slot"))
for k in (123, 456, 789, 321, 654):
    slot, square, middle = mid_square(k, M)
    print("%8d %14s %10s %8d" % (k, square, middle, slot))

print()
print("why the middle: 123 x 123 =", 123 * 123)
print("   the last digit, 9, comes only from 3 x 3.")
print("   the first digit, 1, comes mostly from 1 x 1.")
print("   the middle digits mix all three digits of the key.")
print()
near = (71234, 71235, 71236, 71237)
print("four keys differing only in the LAST digit:")
for k in near:
    slot, square, middle = mid_square(k, 97)
    print("   %d x %d = %-12s middle %s -> slot %d" % (k, k, square, middle, slot))
slots = [mid_square(k, 97)[0] for k in near]
print("   slots:", slots, " all different:", len(set(slots)) == len(slots))
munotes.in340

Hash Functions

the mid-square method, 2 middle digits, then mod 11

     key          k x k   middle 2     slot
     123          15129         51        7
     456         207936         79        2
     789         622521         25        3
     321         103041         30        8
     654         427716         77        0

why the middle: 123 x 123 = 15129
   the last digit, 9, comes only from 3 x 3.
   the first digit, 1, comes mostly from 1 x 1.
   the middle digits mix all three digits of the key.

four keys differing only in the LAST digit:
   71234 x 71234 = 5074282756   middle 28 -> slot 28
   71235 x 71235 = 5074425225   middle 42 -> slot 42
   71236 x 71236 = 5074567696   middle 56 -> slot 56
   71237 x 71237 = 5074710169   middle 71 -> slot 71
   slots: [28, 42, 56, 71]  all different: True

Method 3: the folding method

Break the key into pieces of equal length, add the pieces, and take the remainder. Used when keys are long, such as a 10 digit phone number or a 16 digit card number, because the whole key contributes.

h(k) = (sum of the pieces of k) mod m

Two variants, and MU's papers use both words:

Shift folding, the pieces are simply added.

Boundary folding, alternate pieces are reversed before adding, so that the digit positions do not line up.

def shift_fold(k, m, piece=3):
    digits = str(k)
    pieces = [digits[i:i + piece] for i in range(0, len(digits), piece)]
    total = sum(int(p) for p in pieces)
    return total % m, pieces, total


def boundary_fold(k, m, piece=3):
    digits = str(k)
    pieces = [digits[i:i + piece] for i in range(0, len(digits), piece)]
    turned = [p if i % 2 == 0 else p[::-1] for i, p in enumerate(pieces)]
    total = sum(int(p) for p in turned)
    return total % m, turned, total


M = 97
print("folding a 10 digit mobile number into a %d slot table, pieces of 3" % M)
print()
for k in (9820012345, 7700112233):
    slot, pieces, total = shift_fold(k, M)
    print("%d  shift   : %s -> sum %d -> slot %d" % (k, " + ".join(pieces), total, slot))
    slot, turned, total = boundary_fold(k, M)
    print("%d  boundary: %s -> sum %d -> slot %d" % (k, " + ".join(turned), total, slot))
    print()

print("the two keys that collided in chapter 99 under the division method:")
a, b = 9820012345, 7700112233
print("   division, m = 11 : %d -> %d,  %d -> %d  (collision)"
      % (a, a % 11, b, b % 11))
print("   shift folding    : %d -> %d,  %d -> %d"
      % (a, shift_fold(a, M)[0], b, shift_fold(b, M)[0]))
print("   they no longer collide:", shift_fold(a, M)[0] != shift_fold(b, M)[0])
print()
print("shift folding is blind to the ORDER of the pieces:")
REARRANGED = (120456789, 456120789, 789456120)
for k in REARRANGED:
    slot, pieces, total = shift_fold(k, M)
    print("   %d -> %s -> sum %d -> slot %d"
          % (k, " + ".join(pieces), total, slot))
shift_slots = [shift_fold(k, M)[0] for k in REARRANGED]
print("   three rearrangements of the same pieces, slots %s: all equal %s"
      % (shift_slots, len(set(shift_slots)) == 1))
print()
print("boundary folding reverses alternate pieces, which helps but does")
print("NOT cure it:")
for k in REARRANGED:
    slot, turned, total = boundary_fold(k, M)
    print("   %d -> %s -> sum %d -> slot %d"
          % (k, " + ".join(turned), total, slot))
bound_slots = [boundary_fold(k, M)[0] for k in REARRANGED]
print("   slots %s: %d distinct, against %d for shift folding."
      % (bound_slots, len(set(bound_slots)), len(set(shift_slots))))
print()
print("and here is a triple boundary folding still cannot separate:")
STILL = (123456789, 456123789, 789456123)
for k in STILL:
    slot, turned, total = boundary_fold(k, M)
    print("   %d -> %s -> sum %d -> slot %d"
          % (k, " + ".join(turned), total, slot))
still_slots = [boundary_fold(k, M)[0] for k in STILL]
print("   slots %s: all equal %s" % (still_slots, len(set(still_slots)) == 1))
print("   reversing abc to cba changes the value by 99 x (c - a), and for")
print("   both 456 and 123 that is 99 x 2 = 198, so the two sums stay equal.")
munotes.in341

Hash Functions

folding a 10 digit mobile number into a 97 slot table, pieces of 3

9820012345  shift   : 982 + 001 + 234 + 5 -> sum 1222 -> slot 58
9820012345  boundary: 982 + 100 + 234 + 5 -> sum 1321 -> slot 60

7700112233  shift   : 770 + 011 + 223 + 3 -> sum 1007 -> slot 37
7700112233  boundary: 770 + 110 + 223 + 3 -> sum 1106 -> slot 39

the two keys that collided in chapter 99 under the division method:
   division, m = 11 : 9820012345 -> 0,  7700112233 -> 0  (collision)
   shift folding    : 9820012345 -> 58,  7700112233 -> 37
   they no longer collide: True

shift folding is blind to the ORDER of the pieces:
   120456789 -> 120 + 456 + 789 -> sum 1365 -> slot 7
   456120789 -> 456 + 120 + 789 -> sum 1365 -> slot 7
   789456120 -> 789 + 456 + 120 -> sum 1365 -> slot 7
   three rearrangements of the same pieces, slots [7, 7, 7]: all equal True

boundary folding reverses alternate pieces, which helps but does
NOT cure it:
   120456789 -> 120 + 654 + 789 -> sum 1563 -> slot 11
   456120789 -> 456 + 021 + 789 -> sum 1266 -> slot 5
   789456120 -> 789 + 654 + 120 -> sum 1563 -> slot 11
   slots [11, 5, 11]: 2 distinct, against 1 for shift folding.

and here is a triple boundary folding still cannot separate:
   123456789 -> 123 + 654 + 789 -> sum 1566 -> slot 14
   456123789 -> 456 + 321 + 789 -> sum 1566 -> slot 14
   789456123 -> 789 + 654 + 123 -> sum 1566 -> slot 14
   slots [14, 14, 14]: all equal True
   reversing abc to cba changes the value by 99 x (c - a), and for
   both 456 and 123 that is 99 x 2 = 198, so the two sums stay equal.
munotes.in342

Hash Functions

Shift folding treats the pieces as a bag and is blind to their order, which the output shows: three rearrangements, one slot. Boundary folding exists to attack exactly that, and the output also shows it only partly succeeding. It separated one of the three rearrangements of 120, 456 and 789, and it failed completely on 123, 456 and 789, because reversing abc to cba changes a three digit piece by 99 times (c - a), and that difference is 198 for both 456 and 123, so the sums stayed equal.

The honest statement is therefore the one to write in an answer: boundary folding reduces shift folding's order blindness, it does not remove it. Neither variant suits keys that are rearrangements of one another, such as part numbers assembled from a fixed set of codes.

Method 4: the multiplication method

h(k) = floor( m x fractional_part( k x A ) ), 0 < A < 1

Multiply the key by a constant A between 0 and 1, discard the whole number part, and scale what is left up to the table size. Unlike the division method this works well for any m, including a power of 2, which is why library implementations prefer it.

The recommended constant is A = (sqrt(5) - 1) / 2, about 0.6180339887, the reciprocal of the golden ratio. Knuth showed it spreads sequential keys unusually evenly, and the output below is that claim tested.

import math

A = (math.sqrt(5) - 1) / 2


def multiplication(k, m, a=A):
    fractional = (k * a) % 1
    return int(m * fractional)


print("A = (sqrt(5) - 1) / 2 = %.10f" % A)
print()
M = 16                                  # a power of 2: fatal for division, fine here
print("the multiplication method into a %d slot table" % M)
print("%8s %16s %14s %8s" % ("key", "k x A", "fractional", "slot"))
for k in (1, 2, 3, 4, 5, 100, 1000):
    print("%8d %16.6f %14.6f %8d"
          % (k, k * A, (k * A) % 1, multiplication(k, M)))

print()
print("the same 16 slot table, 16 CONSECUTIVE keys, both methods:")
print("%6s %12s %18s" % ("key", "division", "multiplication"))
div_slots, mul_slots = [], []
for k in range(1000, 1016):
    d, u = k % M, multiplication(k, M)
    div_slots.append(d)
    mul_slots.append(u)
    print("%6d %12d %18d" % (k, d, u))

print()
print("division used %d of %d slots, multiplication used %d"
      % (len(set(div_slots)), M, len(set(mul_slots))))
print("both spread consecutive keys here. now try keys 16 apart,")
print("which is the division method's blind spot on a 16 slot table:")
step_keys = [1000 + 16 * i for i in range(16)]
d2 = [k % M for k in step_keys]
u2 = [multiplication(k, M) for k in step_keys]
print("   keys          :", step_keys[:5], "...")
print("   division slots:", d2)
print("   multiplication:", sorted(u2))
print("   division used %d slot(s); multiplication used %d"
      % (len(set(d2)), len(set(u2))))
munotes.in343

Hash Functions

A = (sqrt(5) - 1) / 2 = 0.6180339887

the multiplication method into a 16 slot table
     key            k x A     fractional     slot
       1         0.618034       0.618034        9
       2         1.236068       0.236068        3
       3         1.854102       0.854102       13
       4         2.472136       0.472136        7
       5         3.090170       0.090170        1
     100        61.803399       0.803399       12
    1000       618.033989       0.033989        0

the same 16 slot table, 16 CONSECUTIVE keys, both methods:
   key     division     multiplication
  1000            8                  0
  1001            9                 10
  1002           10                  4
  1003           11                 14
  1004           12                  8
  1005           13                  1
  1006           14                 11
  1007           15                  5
  1008            0                 15
  1009            1                  9
  1010            2                  3
  1011            3                 13
  1012            4                  7
  1013            5                  1
  1014            6                 10
  1015            7                  4

division used 16 of 16 slots, multiplication used 13
both spread consecutive keys here. now try keys 16 apart,
which is the division method's blind spot on a 16 slot table:
   keys          : [1000, 1016, 1032, 1048, 1064] ...
   division slots: [8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8, 8]
   multiplication: [0, 0, 2, 4, 5, 5, 7, 7, 9, 9, 11, 11, 12, 12, 14, 14]
   division used 1 slot(s); multiplication used 9

That last block is the method's whole argument. Every key that is a multiple of 16 apart lands in the same slot under division with m = 16. The multiplication method, on the identical table, spreads them.

Hashing a string

None of the four methods accepts a string, so a string must first become a number. The standard way is to treat the characters as the digits of a number in some base, and compute it by Horner's rule so that no huge intermediate value is ever formed.

munotes.in344

Hash Functions

h = 0

for each character c of the string:

h = (h x B + code_of(c)) mod m

B is a small constant, traditionally 31 or 33.

def string_hash(text, m, base=31):
    h = 0
    for ch in text:
        h = (h * base + ord(ch)) % m
    return h


def sum_hash(text, m):
    """The naive version: just add the character codes."""
    return sum(ord(ch) for ch in text) % m


M = 101
NAMES = ["Aarti", "Bhavesh", "Chetna", "Devdatta", "Esha", "Farhan",
         "Gauri", "Harsh", "Isha", "Jatin"]

print("hashing names into a %d slot table" % M)
print("%-12s %14s %14s" % ("name", "Horner (31)", "sum of codes"))
for name in NAMES:
    print("%-12s %14d %14d" % (name, string_hash(name, M), sum_hash(name, M)))

print()
print("the sum method cannot tell an anagram from its rearrangement:")
for a, b in (("stop", "pots"), ("listen", "silent"), ("abc", "cba")):
    print("   %-8s and %-8s : sum %3d and %3d   Horner %3d and %3d"
          % (a, b, sum_hash(a, M), sum_hash(b, M),
             string_hash(a, M), string_hash(b, M)))

sum_clash = sum(1 for a, b in (("stop", "pots"), ("listen", "silent"), ("abc", "cba"))
                if sum_hash(a, M) == sum_hash(b, M))
horner_clash = sum(1 for a, b in (("stop", "pots"), ("listen", "silent"), ("abc", "cba"))
                   if string_hash(a, M) == string_hash(b, M))
print()
print("of 3 anagram pairs, the sum method collided on %d; Horner on %d."
      % (sum_clash, horner_clash))
print("position matters in a string, so the hash must be position-aware.")
hashing names into a 101 slot table
name            Horner (31)   sum of codes
Aarti                    70             93
Bhavesh                  85             99
Chetna                   79             90
Devdatta                 61              5
Esha                     36             82
Farhan                   65             87
Gauri                    26            100
Harsh                    79             98
Isha                     20             86
Jatin                    35             98

the sum method cannot tell an anagram from its rearrangement:
   stop     and pots     : sum  50 and  50   Horner  35 and  46
   listen   and silent   : sum  49 and  49   Horner  94 and  90
   abc      and cba      : sum  92 and  92   Horner   0 and   1

of 3 anagram pairs, the sum method collided on 3; Horner on 0.
position matters in a string, so the hash must be position-aware.

The four methods side by side

MethodFormulaStrengthWeakness
Divisionk mod mfastest, one operationm must be chosen with care; a bad m ignores most of the key
Mid-squaremiddle digits of k x kevery digit of the key affects the resultthe square can be large; needs a digit count chosen
Foldingsum of the pieces, mod mhandles very long keys; whole key usedshift folding is blind to the order of the pieces
Multiplicationfloor(m x frac(k x A))works for any m, spreads sequential keysfloating point, and slower than one modulo
munotes.in345

Hash Functions

Quick revision

  • A hash function must be deterministic and must land inside 0 to m-1, which is why every method ends in mod m.
  • Division: h(k) = k mod m. Fastest. Never use a power of 10 or a power of 2 for m; use a prime.
  • With 40 multiples of 10 and m = 10 every key landed in slot 0; with m = 41 they used 40 slots.
  • Mid-square: square the key and take middle digits, because the middle of a square depends on every digit of the key.
  • Folding: split the key into pieces and add them, then mod m. Shift folding adds the pieces as they are, so it cannot tell one arrangement of the pieces from another.
  • Boundary folding reverses alternate pieces to attack that and only partly succeeds: it separated one of three rearrangements of 120, 456, 789 and none of 123, 456, 789, because reversing a three digit piece changes it by 99 times (last digit minus first), which is 198 for both 456 and 123.
  • Multiplication: h(k) = floor(m x frac(k x A)) with A about 0.6180339887. Works for any m, including a power of 2, where division fails on keys m apart.
  • Strings become numbers by Horner's rule, h = (h x 31 + code) mod m, which is position-aware; simply adding the character codes gives every anagram the same slot.

Test yourself

1. State the two absolute requirements on a hash function. It must be deterministic, so the same key always gives the same slot, and its result must lie within 0 to m-1.

2. Give the division method and hash 1001 into an 11 slot table, showing the arithmetic. h(k) = k mod m. 1001 = 11 x 91 + 0, so h(1001) = 0.

3. Why must the table size not be a power of 10 or of 2 in the division method? Because k mod 1000 is the last three digits and k mod 16 is the last four bits, so most of the key is ignored. The chapter put 40 multiples of 10 into a 10 slot table and all 40 landed in slot 0.

4. Explain the reasoning behind the mid-square method. The middle digits of k x k depend on every digit of k, since every digit contributes to the middle of the product, while the leading and trailing digits are dominated by the key's own leading and trailing digits.

5. Distinguish shift folding from boundary folding, and say how far the second fixes the first's weakness. Shift folding adds the pieces as they are, so it depends only on the multiset of pieces and 120456789, 456120789 and 789456120 all hash to one slot. Boundary folding reverses alternate pieces first, which separated one of those three but none of 123456789, 456123789 and 789456123: reversing a three digit piece changes it by 99 times (last digit minus first), and that is 198 for both 456 and 123, so those sums remain equal. Boundary folding reduces the order blindness; it does not remove it.

munotes.in346

Hash Functions

6. Give the multiplication method and the recommended constant, and say what it is good at. h(k) = floor(m x fractional part of (k x A)) with A = (sqrt(5) - 1) / 2, about 0.6180339887. It works for any table size, including a power of 2, and it spreads consecutive keys well.

7. How is a string hashed, and what is wrong with adding the character codes? By Horner's rule: h = (h x B + code of the character) mod m, with B typically 31. Simply adding the codes ignores position, so every anagram hashes to the same slot: "listen" and "silent" collide.

Contents This chapter on its own page

munotes.in347

Chapter One Hundred Two

What Makes a Hash Function Good

Syllabus topic Module 2, "hash functions"

In one line

A good hash function spreads the actual keys evenly over all the buckets, is deterministic, is cheap to compute, uses every part of the key, and sends similar keys to unrelated slots.

The five tests

1. Uniform. Every bucket should receive about n/m of the keys. This is the one that decides performance, because a search costs the length of the bucket it lands in, so the longest bucket is what determines the worst case.

2. Deterministic. The same key must always give the same slot. Nothing involving time, randomness during a run, or a memory address.

3. Cheap. The hash is computed on every insert, search and delete. If the hash costs more than the log n comparisons it replaces, the table is slower than a tree. This is why the division method survives despite its fussiness about m.

4. Uses the whole key. A function that looks at part of the key is blind to the rest, and real keys share parts: PRNs share a year prefix, phone numbers share an operator prefix, names share first letters.

5. Avalanche. Keys that differ slightly should land in unrelated slots. Real key sets are full of near-identical keys, so a function that keeps them together defeats itself.

Test 1 and 4 together: the whole key, measured

The clearest failure is a function that uses only part of the key. Here are four functions on a realistic key set: six digit PRNs where the first three digits are a year and department code shared by most students.

M = 101

# 300 PRNs. Most share the prefix 202, a few are from an earlier batch.
KEYS = [202000 + i for i in range(250)] + [201000 + i * 3 for i in range(50)]


def first_three(k):
    """Uses only the first 3 digits. The classic blunder."""
    return (k // 1000) % M


def last_two(k):
    """Uses only the last 2 digits."""
    return (k % 100) % M


def division(k):
    return k % M


def folded(k):
    digits = str(k)
    return (int(digits[:3]) + int(digits[3:])) % M


def report(name, fn):
    table = [0] * M
    for k in KEYS:
        table[fn(k)] += 1
    used = sum(1 for c in table if c)
    biggest = max(table)
    # average successful search length in a chained table: (1 + len)/2 per bucket
    total_probes = sum(c * (c + 1) // 2 for c in table)
    return (name, used, biggest, total_probes / len(KEYS))


print("%d keys into %d buckets. ideal: %d buckets used, biggest %d"
      % (len(KEYS), M, M, -(-len(KEYS) // M)))
print()
print("%-26s %10s %12s %22s" % ("function", "buckets", "biggest", "avg probes per find"))
for name, fn in (("first 3 digits only", first_three),
                 ("last 2 digits only", last_two),
                 ("division, k mod 101", division),
                 ("folding, 3 + 3 digits", folded)):
    n, used, biggest, avg = report(name, fn)
    print("%-26s %10d %12d %22.2f" % (n, used, biggest, avg))

print()
print("the first-3-digits function used 2 buckets for 300 keys,")
print("because almost every PRN begins 202. its biggest bucket holds 250,")
print("so a search in it compares up to 250 records: the table became a list.")
print("the last-2-digits function used 100 buckets, which looks fine, but it")
print("can NEVER use more than 100 whatever the table size, so a bigger")
print("table would not help it at all.")
munotes.in348

What Makes a Hash Function Good

300 keys into 101 buckets. ideal: 101 buckets used, biggest 3

function                      buckets      biggest    avg probes per find
first 3 digits only                 2          250                 108.83
last 2 digits only                100            4                   2.11
division, k mod 101               101            4                   2.09
folding, 3 + 3 digits             101            4                   2.10

the first-3-digits function used 2 buckets for 300 keys,
because almost every PRN begins 202. its biggest bucket holds 250,
so a search in it compares up to 250 records: the table became a list.
the last-2-digits function used 100 buckets, which looks fine, but it
can NEVER use more than 100 whatever the table size, so a bigger
table would not help it at all.

Two distinct failures there, and an answer should separate them.

The first three digits collapse the table, because the keys share that prefix. Three hundred records in two buckets is a linked list with extra steps.

The last two digits look acceptable at m = 101, and they are a trap: that function's range is only 0 to 99, so enlarging the table to 1009 buckets would leave 909 of them permanently empty. A function must be able to reach every bucket.

Test 5: avalanche, measured

M = 97


def last_two(k):
    return (k % 100) % M


def division(k):
    return k % M


def mid_square(k):
    square = str(k * k)
    start = (len(square) - 4) // 2
    return int(square[start:start + 4]) % M


NEAR = [700100, 700101, 700102, 700103, 700104, 700105, 700106, 700107]

print("eight keys differing only in the last digit, into %d buckets:" % M)
print("%-22s %s" % ("function", "slots"))
for name, fn in (("last 2 digits", last_two),
                 ("division, k mod 97", division),
                 ("mid-square", mid_square)):
    slots = [fn(k) for k in NEAR]
    print("%-22s %-40s %d distinct" % (name, str(slots), len(set(slots))))

print()
GROUPED = [700000 + i * 100 for i in range(8)]     # keys 100 apart
print("eight keys 100 apart (so the last two digits are IDENTICAL):")
print("%-22s %s" % ("function", "slots"))
for name, fn in (("last 2 digits", last_two),
                 ("division, k mod 97", division),
                 ("mid-square", mid_square)):
    slots = [fn(k) for k in GROUPED]
    print("%-22s %-40s %d distinct" % (name, str(slots), len(set(slots))))

print()
print("on the first set every function separated the keys.")
print("on the second, the last-2-digits function sent all eight to ONE slot,")
print("because the two digits it reads are the same in all eight keys.")
print("avalanche is not about random keys. it is about the keys you get.")
munotes.in349

What Makes a Hash Function Good

eight keys differing only in the last digit, into 97 buckets:
function               slots
last 2 digits          [0, 1, 2, 3, 4, 5, 6, 7]                 8 distinct
division, k mod 97     [51, 52, 53, 54, 55, 56, 57, 58]         8 distinct
mid-square             [24, 67, 13, 56, 2, 45, 88, 34]          8 distinct

eight keys 100 apart (so the last two digits are IDENTICAL):
function               slots
last 2 digits          [0, 0, 0, 0, 0, 0, 0, 0]                 1 distinct
division, k mod 97     [48, 51, 54, 57, 60, 63, 66, 69]         8 distinct
mid-square             [0, 24, 50, 69, 2, 25, 59, 95]           8 distinct

on the first set every function separated the keys.
on the second, the last-2-digits function sent all eight to ONE slot,
because the two digits it reads are the same in all eight keys.
avalanche is not about random keys. it is about the keys you get.

Test 3: cheap, measured against what it replaces

M = 1009
N = 4000
KEYS = [202000 + i * 7 for i in range(N)]


def count_division(k):
    return 1                                 # one modulo


def count_mid_square(k):
    return 4                                 # a multiply, a string cut, a parse, a modulo


def count_folding(k):
    return len(str(k)) // 3 + 2              # a cut and parse per piece, then a modulo


def count_horner(text):
    return 2 * len(text)                      # a multiply and an add per character


import math

print("what a hash costs, in rough arithmetic operations per lookup:")
print("%-28s %12s" % ("function", "operations"))
for name, ops in (("division, k mod m", count_division(0)),
                  ("mid-square", count_mid_square(0)),
                  ("folding a 6 digit key", count_folding(202000)),
                  ("Horner over a 12 char string", count_horner("a" * 12))):
    print("%-28s %12d" % (name, ops))

print()
print("what it replaces: a balanced tree search over %d records needs" % N)
print("about %d key comparisons." % math.ceil(math.log2(N)))
print()
print("so a hash costing 1 to 4 operations is clearly worth it, and a hash")
print("costing more than about %d would not be. that is the whole test:" % math.ceil(math.log2(N)))
print("a hash function must be cheaper than the search it removes.")
what a hash costs, in rough arithmetic operations per lookup:
function                       operations
division, k mod m                       1
mid-square                              4
folding a 6 digit key                   4
Horner over a 12 char string           24

what it replaces: a balanced tree search over 4000 records needs
about 12 key comparisons.

so a hash costing 1 to 4 operations is clearly worth it, and a hash
costing more than about 12 would not be. that is the whole test:
a hash function must be cheaper than the search it removes.
munotes.in350

What Makes a Hash Function Good

That is the practical boundary. A cryptographic hash such as SHA-256 is beautifully uniform and costs hundreds of operations, so it is the wrong tool for a hash table, and the right tool when the uniformity has to hold against an attacker. Chapter 109 returns to that.

Measuring uniformity properly

Counting the biggest bucket is a blunt instrument. The standard measure compares the observed bucket sizes with what a perfectly even spread would give.

expected per bucket = n / m

spread = sum over buckets of (observed - expected)^2 / expected

A perfectly even spread scores 0 when m divides n exactly, and a little above 0 otherwise, since the buckets then cannot all hold the same number. A spread no better than pure chance scores about m-1. A function that dumps every key into one bucket scores n x (m-1), which is the maximum. Lower is better, and the score means something only against those three landmarks, which the listing prints.

M = 101
KEYS = [202000 + i for i in range(250)] + [201000 + i * 3 for i in range(50)]


def spread(fn, keys, m):
    table = [0] * m
    for k in keys:
        table[fn(k)] += 1
    expected = len(keys) / m
    return sum((c - expected) ** 2 for c in table) / expected


def first_three(k):
    return (k // 1000) % M


def last_two(k):
    return (k % 100) % M


def division(k):
    return k % M


def folded(k):
    d = str(k)
    return (int(d[:3]) + int(d[3:])) % M


perfect = [0] * M
for i, _ in enumerate(KEYS):
    perfect[i % M] += 1
expected = len(KEYS) / M
best = sum((c - expected) ** 2 for c in perfect) / expected
all_in_one = ((len(KEYS) - expected) ** 2 + (M - 1) * expected ** 2) / expected

print("%d keys into %d buckets. lower is better. three landmarks:" % (len(KEYS), M))
print("   best possible, an even spread     : %8.1f" % best)
print("   a spread no better than chance    : %8.1f" % (M - 1))
print("   everything in one bucket, n(m-1)  : %8.1f" % all_in_one)
print("(the best is not 0 because %d does not divide evenly into %d buckets.)"
      % (len(KEYS), M))
print()
print("%-26s %14s" % ("function", "spread score"))
for name, fn in (("first 3 digits only", first_three),
                 ("last 2 digits only", last_two),
                 ("division, k mod 101", division),
                 ("folding, 3 + 3 digits", folded)):
    print("%-26s %14.1f" % (name, spread(fn, KEYS, M)))

print()
good = [spread(fn, KEYS, M) for fn in (last_two, division, folded)]
print("the three usable functions score %.1f to %.1f, well under the %d of a"
      % (min(good), max(good), M - 1))
print("chance spread, so all three beat chance on these keys.")
bad = spread(first_three, KEYS, M)
print("the first-3-digits function scores %.0f, which is %.0f%% of the way to"
      % (bad, 100 * bad / all_in_one))
print("the %.0f of putting every key in one bucket. it is not merely worse;"
      % all_in_one)
print("it is most of the way to having no hash function at all.")
munotes.in351

What Makes a Hash Function Good

300 keys into 101 buckets. lower is better. three landmarks:
   best possible, an even spread     :      1.0
   a spread no better than chance    :    100.0
   everything in one bucket, n(m-1)  :  30000.0
(the best is not 0 because 300 does not divide evenly into 101 buckets.)

function                     spread score
first 3 digits only               21583.3
last 2 digits only                   25.2
division, k mod 101                  20.5
folding, 3 + 3 digits                22.5

the three usable functions score 20.5 to 25.2, well under the 100 of a
chance spread, so all three beat chance on these keys.
the first-3-digits function scores 21583, which is 72% of the way to
the 30000 of putting every key in one bucket. it is not merely worse;
it is most of the way to having no hash function at all.

The rule that matters most

A hash function is good or bad only with respect to a key set. k mod 101 is excellent on the PRNs above and catastrophic on keys that are all multiples of 101. There is no function that is uniform on every possible input, because a function with m outputs and a larger domain must send many inputs to each output, and an adversary or an unlucky data source can pick them.

So the practical procedure is:

  1. find out what the keys actually look like;
  2. choose a method that uses all of a key of that shape;
  3. choose m as a prime away from powers of 2;
  4. measure, with a spread score or at least the biggest bucket, on real keys;
  5. where an attacker chooses the keys, use a randomised or keyed hash, chapter 109.

Quick revision

  • Five tests: uniform, deterministic, cheap, uses the whole key, and avalanche.
  • A search costs the length of its bucket, so the BIGGEST bucket sets the worst case, not the average.
  • Using part of the key is the commonest failure: the first three digits of 300 PRNs that share a prefix filled 2 buckets, the biggest holding 250 records.
  • A function whose range is smaller than the table can never fill it: a last-two-digits hash is stuck at 100 buckets however large m is.
  • Avalanche is about the keys you actually get: the last-two-digits hash sent eight keys 100 apart to one slot.
  • Cheap means cheaper than the search it replaces: 1 to 4 operations against about 12 comparisons for a balanced tree over 4,000 records.
  • A cryptographic hash is uniform and far too slow for a hash table, and is the right choice only when an attacker chooses the keys.
  • Uniformity is measured by the spread score, sum of (observed - expected) squared over expected; lower is better, read against three landmarks: the best achievable, about m-1 for a chance spread, and n(m-1) for everything in one bucket.
  • On 300 PRNs in 101 buckets the three usable functions scored 20.5 to 25.2 against a best of 1.0 and a chance figure of 100, while the first-three-digits function scored 21,583 of a possible 30,000.
  • No function is uniform on all inputs, so a function is only good relative to a key set: measure on real keys.
munotes.in352

What Makes a Hash Function Good

Test yourself

1. List the five properties of a good hash function. Uniform over the actual keys; deterministic; cheap to compute; uses every part of the key; and avalanche, so similar keys reach unrelated slots.

2. Why does the biggest bucket matter more than the average? Because a search costs the length of the bucket it lands in, so the longest bucket fixes the worst case cost.

3. Give the measured failure of a hash that uses only part of the key. Three hundred PRNs mostly beginning 202, hashed on their first three digits, filled only 2 of 101 buckets, and the largest bucket held 250 records, so a search in it compared up to 250 keys.

4. Why is a hash function whose range is 0 to 99 unacceptable even when it spreads keys evenly? Because it can never reach any bucket above 99, so enlarging the table cannot reduce the bucket lengths.

5. What is avalanche, and give the chapter's example of its absence. Keys that differ slightly should land in unrelated slots. A last-two-digits hash sent eight keys exactly 100 apart to a single slot, because the digits it reads are identical in all eight.

6. State the cost test precisely. The hash must cost less than the search it removes. Division costs one modulo against about 12 comparisons for a balanced tree over 4,000 records, so it passes easily; a cryptographic hash costing hundreds of operations fails.

7. How is uniformity measured? By the spread score: the sum over buckets of (observed minus expected) squared divided by expected, where expected is n/m. Lower is better, judged against three landmarks: the best achievable for those n and m, which is 0 only when m divides n exactly; about m-1 for a spread no better than chance; and n(m-1) for every key in one bucket.

8. Why can no hash function be uniform on every input? Because the set of possible keys is larger than the set of m slots, so many keys must share each slot, and a key set can be drawn entirely from one of those groups.

Contents This chapter on its own page

munotes.in353

Chapter One Hundred Three

Collisions Are Certain, Not Unlucky

Syllabus topic Module 2, "collision"

In one line

Two different keys landing in the same bucket is unavoidable in principle and likely in practice long before the table is anywhere near full, so a hash table is not complete until it has a collision strategy.

The word

A collision is two distinct keys k1 and k2, both stored, with h(k1) = h(k2). They compete for one slot. Chapter 99 produced one by accident on its first four keys.

Argument one: the pigeonhole principle

If there are more keys than buckets, two keys must share a bucket. No hash function can avoid it, however well designed, because a function from a larger set to a smaller one cannot be one to one.

if n > m then at least one bucket holds at least 2 keys

More precisely, some bucket holds at least ceiling(n/m) keys.

import math


def best_possible_worst_bucket(n, m):
    """However good the hash function, SOME bucket holds at least this many."""
    return math.ceil(n / m)


print("%8s %10s %22s %12s"
      % ("keys", "buckets", "some bucket must hold", "forced?"))
for n, m in ((100, 101), (101, 101), (102, 101), (300, 101),
             (1000, 101), (5000, 101)):
    least = best_possible_worst_bucket(n, m)
    print("%8d %10d %22d %12s"
          % (n, m, least, "YES" if least > 1 else "no"))

print()
print("at 101 keys in 101 buckets a collision is still avoidable in")
print("principle: one key per bucket. at 102 it is not. no hash function")
print("whatever can prevent it, because 102 things cannot sit in 101 places")
print("one to a place.")
print()
print("and the forced crowding grows: 5000 keys in 101 buckets means some")
print("bucket holds at least %d, so some search compares at least that many."
      % best_possible_worst_bucket(5000, 101))
    keys    buckets  some bucket must hold      forced?
     100        101                      1           no
     101        101                      1           no
     102        101                      2          YES
     300        101                      3          YES
    1000        101                     10          YES
    5000        101                     50          YES

at 101 keys in 101 buckets a collision is still avoidable in
principle: one key per bucket. at 102 it is not. no hash function
whatever can prevent it, because 102 things cannot sit in 101 places
one to a place.

and the forced crowding grows: 5000 keys in 101 buckets means some
bucket holds at least 50, so some search compares at least that many.

That settles the principle. It does not settle practice, because real tables are kept well under full (chapter 107 says why). So the second argument is the one that matters.

Argument two: collisions arrive far earlier than intuition expects

This is the birthday problem. In a room of 23 people the chance that two share a birthday is already above half, although there are 365 possible days. The same arithmetic governs a hash table.

The probability that n keys land in n different buckets, assuming a uniform hash:

munotes.in354

Collisions Are Certain, Not Unlucky

P(no collision) = (1 - 0/m) x (1 - 1/m) x (1 - 2/m) x ... x (1 - (n-1)/m)

Each new key must miss every bucket already used.

def probability_no_collision(n, m):
    p = 1.0
    for i in range(n):
        p *= (m - i) / m
    return p


print("the classic case, m = 365 days:")
for n in (10, 20, 22, 23, 30, 50):
    p = probability_no_collision(n, 365)
    print("   %2d people: P(all different) = %.4f   P(a shared birthday) = %.4f"
          % (n, p, 1 - p))

print()
print("23 is the famous answer, and it is where P(shared) passes 0.5:",
      1 - probability_no_collision(23, 365) > 0.5,
      "at 22 it is still below:",
      1 - probability_no_collision(22, 365) < 0.5)
the classic case, m = 365 days:
   10 people: P(all different) = 0.8831   P(a shared birthday) = 0.1169
   20 people: P(all different) = 0.5886   P(a shared birthday) = 0.4114
   22 people: P(all different) = 0.5243   P(a shared birthday) = 0.4757
   23 people: P(all different) = 0.4927   P(a shared birthday) = 0.5073
   30 people: P(all different) = 0.2937   P(a shared birthday) = 0.7063
   50 people: P(all different) = 0.0296   P(a shared birthday) = 0.9704

23 is the famous answer, and it is where P(shared) passes 0.5: True at 22 it is still below: True

Now the same computation for hash tables, which is the point.

import math


def probability_no_collision(n, m):
    p = 1.0
    for i in range(n):
        p *= (m - i) / m
    return p


def first_likely_collision(m):
    """The smallest n for which a collision is more likely than not."""
    n = 1
    while probability_no_collision(n, m) >= 0.5:
        n += 1
    return n


print("%10s %22s %14s %18s"
      % ("buckets m", "keys for 50% chance", "that is", "1.25 x sqrt(m)"))
for m in (101, 1009, 10007, 100003, 1000003):
    n = first_likely_collision(m)
    print("%10d %22d %13.1f%% %18.0f"
          % (m, n, 100 * n / m, 1.25 * math.sqrt(m)))

print()
print("read the third column. in a table of 1,000,003 buckets, a collision")
print("is more likely than not once 1,178 keys are stored.")
print("that is 0.1% of the table.")
print()
print("the rule of thumb, 1.25 x sqrt(m), matches the exact answer closely,")
print("and it is the number to quote: collisions begin at the SQUARE ROOT of")
print("the table size, not near the table size.")
 buckets m    keys for 50% chance        that is     1.25 x sqrt(m)
       101                     13          12.9%                 13
      1009                     38           3.8%                 40
     10007                    119           1.2%                125
    100003                    373           0.4%                395
   1000003                   1178           0.1%               1250

read the third column. in a table of 1,000,003 buckets, a collision
is more likely than not once 1,178 keys are stored.
that is 0.1% of the table.

the rule of thumb, 1.25 x sqrt(m), matches the exact answer closely,
and it is the number to quote: collisions begin at the SQUARE ROOT of
the table size, not near the table size.
munotes.in355

Collisions Are Certain, Not Unlucky

That is the chapter. A hash table of a million buckets starts colliding at about a thousand records. Collisions are not an edge case to be handled defensively; they are the ordinary behaviour of the structure, and a table without a collision strategy is broken at 0.1 per cent occupancy.

How many collisions, on average

def expected_collisions(n, m):
    """n keys minus the expected number of non-empty buckets."""
    occupied = m * (1 - (1 - 1 / m) ** n)
    return n - occupied


print("%8s %10s %20s %18s" % ("keys", "buckets", "load factor", "expected collisions"))
for n, m in ((50, 1009), (100, 1009), (500, 1009), (1009, 1009), (2000, 1009)):
    print("%8d %10d %19.2f %18.1f"
          % (n, m, n / m, expected_collisions(n, m)))

print()
print("at a load factor of 0.50 in a 1009 bucket table, about %.0f of the"
      % expected_collisions(500, 1009))
print("500 keys are already sharing a bucket with something.")
print("at a load factor of 1.00, about %.0f of 1009 keys are."
      % expected_collisions(1009, 1009))
    keys    buckets          load factor expected collisions
      50       1009                0.05                1.2
     100       1009                0.10                4.8
     500       1009                0.50              105.6
    1009       1009                1.00              371.0
    2000       1009                1.98             1129.9

at a load factor of 0.50 in a 1009 bucket table, about 106 of the
500 keys are already sharing a bucket with something.
at a load factor of 1.00, about 371 of 1009 keys are.

A real hash function, a real key set, and the first collision

The arguments above assume a uniform hash. Here is what actually happens with the division method on sequential and on realistic keys.

M = 1009


def mix(i):
    """A deterministic integer mixer: pattern-free keys, no random module."""
    x = ((i + 1) * 2654435761) % (2 ** 32)
    x ^= x >> 13
    x = (x * 2246822519) % (2 ** 32)
    x ^= x >> 17
    return x


def first_clash(keys, m):
    """Position of the first key that lands in an already used slot."""
    seen = {}
    for position, k in enumerate(keys, 1):
        slot = k % m
        if slot in seen:
            return position, seen[slot], k, slot
        seen[slot] = k
    return None, None, None, None


sequential = [1000 + i for i in range(3000)]
spaced = [1000 + i * M for i in range(3000)]              # all the SAME slot
progression = [(i * 7919 + 13) % 999331 for i in range(3000)]
mixed = [mix(i) for i in range(3000)]

print("the division method, m = %d, four key sets:" % M)
print()
print("%-34s %10s   %s" % ("key set", "first clash", "which two, and where"))
for name, keys in (("1000, 1001, 1002, ... consecutive", sequential),
                   ("1000, 2009, 3018, ... m apart", spaced),
                   ("an arithmetic progression", progression),
                   ("pattern-free, from the mixer", mixed)):
    position, earlier, k, slot = first_clash(keys, M)
    if position is None:
        print("%-34s %10s" % (name, "none in 3000"))
    else:
        print("%-34s %10d   %d and %d, slot %d"
              % (name, position, earlier, k, slot))

print()
print("consecutive keys survive to key %d: k mod m visits every slot once"
      % first_clash(sequential, M)[0])
print("before repeating, which is the BEST possible case, not a typical one.")
print("keys exactly m apart clash on the second key: the worst case.")
print("the arithmetic progression lasted %d, because 7919 mod %d is %d and"
      % (first_clash(progression, M)[0], M, 7919 % M))
print("%d has no factor in common with %d, so it also cycles through the"
      % (7919 % M, M))
print("slots in turn until the outer modulus breaks the pattern.")
print()
print("NONE of those three tests the birthday prediction, because none of")
print("them is pattern-free. only the fourth is, and it clashed at key %d."
      % first_clash(mixed, M)[0])
munotes.in356

Collisions Are Certain, Not Unlucky

the division method, m = 1009, four key sets:

key set                            first clash   which two, and where
1000, 1001, 1002, ... consecutive        1010   1000 and 2009, slot 1000
1000, 2009, 3018, ... m apart               2   1000 and 2009, slot 1000
an arithmetic progression                 428   13 and 383433, slot 13
pattern-free, from the mixer               20   44560300 and 2286135529, slot 842

consecutive keys survive to key 1010: k mod m visits every slot once
before repeating, which is the BEST possible case, not a typical one.
keys exactly m apart clash on the second key: the worst case.
the arithmetic progression lasted 428, because 7919 mod 1009 is 856 and
856 has no factor in common with 1009, so it also cycles through the
slots in turn until the outer modulus breaks the pattern.

NONE of those three tests the birthday prediction, because none of
them is pattern-free. only the fourth is, and it clashed at key 20.

One sample proves nothing: the position of a first collision is itself a random quantity. So the prediction is tested properly, over many independent pattern-free key sets.

import statistics

M = 1009


def mix(i, salt):
    x = ((i + 1) * 2654435761 ^ (salt * 40503)) % (2 ** 32)
    x ^= x >> 13
    x = (x * 2246822519) % (2 ** 32)
    x ^= x >> 17
    return x


def first_clash_position(salt, m, limit=4000):
    seen = set()
    for position in range(1, limit + 1):
        slot = mix(position - 1, salt) % m
        if slot in seen:
            return position
        seen.add(slot)
    return limit


def probability_no_collision(n, m):
    p = 1.0
    for i in range(n):
        p *= (m - i) / m
    return p


def predicted_median(m):
    n = 1
    while probability_no_collision(n, m) >= 0.5:
        n += 1
    return n


positions = [first_clash_position(salt, M) for salt in range(400)]
predicted = predicted_median(M)

print("400 independent pattern-free key sets, %d buckets each:" % M)
print("   earliest first clash :", min(positions))
print("   latest first clash   :", max(positions))
print("   MEDIAN first clash   :", int(statistics.median(positions)))
print("   birthday prediction  :", predicted)
print()
print("the median is within %d of the prediction."
      % abs(int(statistics.median(positions)) - predicted))
print("the spread, %d to %d, is the reason a single run proves nothing:"
      % (min(positions), max(positions)))
print("the formula predicts the MEDIAN, not what any one table will do.")
print()
below = sum(1 for p in positions if p <= predicted)
print("%d of 400 key sets collided at or before key %d, which should be"
      % (below, predicted))
print("about half of them, and is %.0f%%." % (100 * below / len(positions)))
munotes.in357

Collisions Are Certain, Not Unlucky

400 independent pattern-free key sets, 1009 buckets each:
   earliest first clash : 3
   latest first clash   : 112
   MEDIAN first clash   : 38
   birthday prediction  : 38

the median is within 0 of the prediction.
the spread, 3 to 112, is the reason a single run proves nothing:
the formula predicts the MEDIAN, not what any one table will do.

204 of 400 key sets collided at or before key 38, which should be
about half of them, and is 51%.

Median 38 against a prediction of 38, and 51 per cent of the key sets colliding at or before it. The formula is confirmed, and the 3 to 112 spread is the part to carry away: the birthday figure is a median, not a promise, and any particular table may collide on its third key.

Perfect hashing exists, and its condition is severe

A perfect hash function maps a set of keys to distinct slots with no collisions at all. It exists, and it is used, but only under one condition: the complete set of keys must be known in advance and must not change.

That fits a compiler's reserved words, a fixed set of country codes, or the opcodes of a processor. It does not fit students, customers or files, because those sets grow. A table whose keys arrive over time cannot be perfectly hashed, so it must handle collisions.

KEYWORDS = ["if", "else", "while", "for", "return", "break", "int", "float"]


def find_perfect(words, start=len(KEYWORDS)):
    """Search for a table size at which these fixed words do not collide."""
    m = start
    while True:
        slots = [sum(ord(c) for c in w) % m for w in words]
        if len(set(slots)) == len(words):
            return m, slots
        m += 1
        if m > 200:
            return None, None


m, slots = find_perfect(KEYWORDS)
print("a fixed set of %d keywords:" % len(KEYWORDS), KEYWORDS)
print()
print("the smallest table size with NO collisions for them:", m)
for word, slot in zip(KEYWORDS, slots):
    print("   %-8s -> %d" % (word, slot))
print("all distinct:", len(set(slots)) == len(KEYWORDS))
print()
print("now words that were not known when the size %d was chosen:" % m)
NEWCOMERS = ["switch", "case", "do", "void", "char", "long", "const",
             "static", "struct", "union", "goto", "sizeof"]
collided = []
for word in NEWCOMERS:
    slot = sum(ord(c) for c in word) % m
    clash = slot in slots
    if clash:
        collided.append(word)
    print("   %-8s -> %2d   %s"
          % (word, slot, "COLLIDES" if clash else "free"))

print()
print("%d of the %d newcomers collided." % (len(collided), len(NEWCOMERS)))
print("note what this does and does not show. it does NOT show that every")
print("new key collides. it shows that the GUARANTEE is gone: whether a new")
print("key fits is now luck, and %d of %d were unlucky."
      % (len(collided), len(NEWCOMERS)))
print()
print("a perfect hash is perfect for ONE fixed set of keys. the moment the")
print("set can change, there is no guarantee left, which is why ordinary")
print("tables must handle collisions. chapters 104 to 106 are the three ways.")
munotes.in358

Collisions Are Certain, Not Unlucky

a fixed set of 8 keywords: ['if', 'else', 'while', 'for', 'return', 'break', 'int', 'float']

the smallest table size with NO collisions for them: 18
   if       -> 9
   else     -> 11
   while    -> 15
   for      -> 3
   return   -> 6
   break    -> 13
   int      -> 7
   float    -> 12
all distinct: True

now words that were not known when the size 18 was chosen:
   switch   -> 10   free
   case     -> 16   free
   do       -> 13   COLLIDES
   void     ->  2   free
   char     ->  0   free
   long     ->  0   free
   const    -> 11   COLLIDES
   static   ->  0   free
   struct   -> 11   COLLIDES
   union    -> 13   COLLIDES
   goto     ->  9   COLLIDES
   sizeof   ->  8   free

5 of the 12 newcomers collided.
note what this does and does not show. it does NOT show that every
new key collides. it shows that the GUARANTEE is gone: whether a new
key fits is now luck, and 5 of 12 were unlucky.

a perfect hash is perfect for ONE fixed set of keys. the moment the
set can change, there is no guarantee left, which is why ordinary
tables must handle collisions. chapters 104 to 106 are the three ways.

Quick revision

  • A collision is two stored keys with the same hash value.
  • Pigeonhole: if n is greater than m a collision is certain, and some bucket holds at least ceiling(n/m) keys.
  • Birthday arithmetic: P(no collision) is the product of (m - i)/m for i from 0 to n-1.
  • A collision becomes more likely than not at about 1.25 times the square root of m, not near m.
  • In a table of 1,000,003 buckets that is 1,178 keys, which is 0.1 per cent of the table.
  • Expected collisions are n minus m(1 - (1 - 1/m) to the power n).
  • Keys exactly m apart collide on the second key under the division method, which is the worst case; consecutive keys are the best case.
  • A perfect hash function has no collisions, but only for a key set fixed and known in advance, such as a language's reserved words.
  • Therefore every general purpose hash table needs a collision strategy: chapters 104 to 106.
munotes.in359

Collisions Are Certain, Not Unlucky

Test yourself

1. Define a collision. Two distinct keys that are both stored and that the hash function sends to the same bucket.

2. State the pigeonhole argument and what it proves. A function from a larger set to a smaller one cannot be one to one, so if there are more keys than buckets at least two keys share a bucket, and some bucket holds at least ceiling(n/m) keys. It proves collisions cannot be designed away.

3. Give the probability that n keys avoid all collisions in m buckets. The product of (m - i)/m for i from 0 to n-1, that is (1 - 0/m)(1 - 1/m) up to (1 - (n-1)/m).

4. Roughly how many keys before a collision is more likely than not, and why does that matter? About 1.25 times the square root of m. It matters because that is far below m: in a table of 1,000,003 buckets a collision is more likely than not at 1,178 keys, which is 0.1 per cent full.

5. Give the birthday problem's classic figure. With 365 possible days, 23 people are enough for a shared birthday to be more likely than not; at 22 it is still below one half.

6. Which key set is the worst case for the division method, and how quickly does it fail? Keys exactly m apart, such as 1000, 2009, 3018 with m = 1009. They all hash to the same slot, so the second key collides.

7. What is a perfect hash function and what does it require? One that maps its keys to distinct slots with no collisions. It requires the complete set of keys to be known in advance and to stay fixed, which suits a language's reserved words but not a growing set of records.

Contents This chapter on its own page

munotes.in360

Chapter One Hundred Four

Chaining

Syllabus topic Module 2, "collision avoidance techniques"; Computer Science Practical 3, "Implement a hash table with chaining or linear probing"

In one line

Chaining lets every bucket hold a list of records instead of one, so colliding keys simply join the same list and nothing is ever displaced.

The idea

The table is an array of m chains, each a linked list, each initially empty. To insert, hash the key and add the record to that chain. To search, hash the key and walk that chain. To delete, hash the key and unlink the record from that chain.

The scheme is also called separate chaining, because the records live outside the array proper.

table[0] -> (k1,v1) -> (k7,v7) -> null

table[1] -> null

table[2] -> (k4,v4) -> null

table[3] -> (k2,v2) -> (k9,v9) -> (k5,v5) -> null

Built, with real chains

class Node:
    """A node of the chain. Chapter 13's node, carrying a key and a value."""

    __slots__ = ("key", "value", "next")

    def __init__(self, key, value, nxt=None):
        self.key = key
        self.value = value
        self.next = nxt


class ChainedTable:
    def __init__(self, buckets=7):
        self.table = [None] * buckets
        self.count = 0
        self.probes = 0

    def _slot(self, key):
        return key % len(self.table)

    def insert(self, key, value):
        """Insert at the HEAD, which is O(1): chapter 17."""
        slot = self._slot(key)
        node = self.table[slot]
        while node is not None:
            self.probes += 1
            if node.key == key:
                node.value = value
                return "replaced"
            node = node.next
        self.table[slot] = Node(key, value, self.table[slot])
        self.count += 1
        return "inserted"

    def search(self, key):
        node = self.table[self._slot(key)]
        while node is not None:
            self.probes += 1
            if node.key == key:
                return node.value
            node = node.next
        return None

    def delete(self, key):
        slot = self._slot(key)
        node, previous = self.table[slot], None
        while node is not None:
            self.probes += 1
            if node.key == key:
                if previous is None:
                    self.table[slot] = node.next
                else:
                    previous.next = node.next
                self.count -= 1
                return True
            previous, node = node, node.next
        return False

    def load_factor(self):
        return self.count / len(self.table)

    def show(self):
        lines = []
        for i, head in enumerate(self.table):
            parts, node = [], head
            while node is not None:
                parts.append("(%d, %s)" % (node.key, node.value))
                node = node.next
            lines.append("   [%d] -> %s" % (i, " -> ".join(parts) + " -> null"
                                            if parts else "null"))
        return "\n".join(lines)


t = ChainedTable(buckets=7)
for k, name in ((15, "Aarti"), (11, "Bhavesh"), (27, "Chetna"), (8, "Devdatta"),
                (22, "Esha"), (29, "Farhan"), (4, "Gauri")):
    t.insert(k, name)

print("7 buckets, 7 records, slot = key mod 7:")
print(t.show())
print()
print("load factor      : %.2f" % t.load_factor())
print("search for 29    :", t.search(29))
print("search for 100   :", t.search(100), "(absent)")
print("delete 22        :", t.delete(22))
print("delete 22 again  :", t.delete(22))
print("count after      :", t.count)
print()
print("after deleting 22 the chain at slot 1 is intact:")
print(t.show().splitlines()[1])
7 buckets, 7 records, slot = key mod 7:
   [0] -> null
   [1] -> (29, Farhan) -> (22, Esha) -> (8, Devdatta) -> (15, Aarti) -> null
   [2] -> null
   [3] -> null
   [4] -> (4, Gauri) -> (11, Bhavesh) -> null
   [5] -> null
   [6] -> (27, Chetna) -> null

load factor      : 1.00
search for 29    : Farhan
search for 100   : None (absent)
delete 22        : True
delete 22 again  : False
count after      : 6

after deleting 22 the chain at slot 1 is intact:
   [1] -> (29, Farhan) -> (8, Devdatta) -> (15, Aarti) -> null
munotes.in361

Chaining

Notice three things that the next chapter will not be able to claim.

Nothing was displaced. Four of the seven keys, 15, 22, 8 and 29, are all 1 mod 7, and all four are stored in slot 1's chain. None of them was pushed into a neighbouring slot and none of them moved.

Deletion is a plain unlink. Chapter 19's removal, and the chain is correct immediately afterwards. No marker, no special case.

The count may exceed the bucket count, because a chain has no capacity.

The load factor, and the cost that follows from it

load factor, alpha = n / m = records / buckets

For chaining, alpha is the average chain length. It may be any non-negative number, including values above 1.

The textbook results, which are what a question asks for:

unsuccessful search : alpha probes on average

successful search : 1 + alpha / 2 probes on average

insert (no duplicate check) : 1 probe

The reasoning for the successful case: the record is somewhere in a chain of average length alpha, and on average half of it is walked, plus the record itself. The reasoning for the unsuccessful case: the whole chain is walked, which averages alpha.

The formulas checked against the table

def mix(i):
    """Deterministic pattern-free keys, so the measurement is reproducible."""
    x = ((i + 1) * 2654435761) % (2 ** 32)
    x ^= x >> 13
    x = (x * 2246822519) % (2 ** 32)
    x ^= x >> 17
    return x


M = 1009


def build(n, m):
    table = [[] for _ in range(m)]
    keys = [mix(i) for i in range(n)]
    for k in keys:
        table[k % m].append(k)
    return table, keys


def measure(n, m):
    table, keys = build(n, m)

    hits = 0
    for k in keys:                                 # every key, found
        chain = table[k % m]
        hits += chain.index(k) + 1

    misses = 0
    absent = [mix(100000 + i) for i in range(500)]
    for k in absent:
        misses += len(table[k % m])

    return hits / len(keys), misses / len(absent)


print("%8s %8s %8s %14s %14s %14s %14s"
      % ("records", "buckets", "alpha", "hit measured", "1 + a/2", "miss measured", "alpha"))
for n in (505, 1009, 2018, 5045):
    hit, miss = measure(n, M)
    alpha = n / M
    print("%8d %8d %8.2f %14.2f %14.2f %14.2f %14.2f"
          % (n, M, alpha, hit, 1 + alpha / 2, miss, alpha))

print()
print("the measured successful search tracks 1 + alpha/2 and the")
print("unsuccessful one tracks alpha, at every load factor including")
print("alpha above 1, which is chaining's distinctive ability.")
print()
print("and at alpha = 5 the average successful search is still under 4")
print("probes, over 5045 records. a linear search would average 2523.")
munotes.in362

Chaining

 records  buckets    alpha   hit measured        1 + a/2  miss measured          alpha
     505     1009     0.50           1.24           1.25           0.49           0.50
    1009     1009     1.00           1.50           1.50           0.99           1.00
    2018     1009     2.00           2.01           2.00           2.09           2.00
    5045     1009     5.00           3.47           3.50           4.94           5.00

the measured successful search tracks 1 + alpha/2 and the
unsuccessful one tracks alpha, at every load factor including
alpha above 1, which is chaining's distinctive ability.

and at alpha = 5 the average successful search is still under 4
probes, over 5045 records. a linear search would average 2523.

The important reading is the fourth row. At a load factor of 5, five records per bucket, the table is still answering a successful search in under 4 probes. Chaining does not break when the table fills; it degrades in proportion, which is the strongest practical argument for it.

How full is too full

def mix(i):
    x = ((i + 1) * 2654435761) % (2 ** 32)
    x ^= x >> 13
    x = (x * 2246822519) % (2 ** 32)
    x ^= x >> 17
    return x


M = 101

print("%8s %8s %16s %16s %14s"
      % ("alpha", "records", "longest chain", "empty buckets", "avg hit"))
for alpha in (0.25, 0.5, 0.75, 1.0, 2.0, 5.0, 10.0):
    n = int(alpha * M)
    table = [[] for _ in range(M)]
    keys = [mix(i) for i in range(n)]
    for k in keys:
        table[k % M].append(k)
    hits = sum(table[k % M].index(k) + 1 for k in keys) / max(n, 1)
    print("%8.2f %8d %16d %16d %14.2f"
          % (alpha, n, max(len(c) for c in table),
             sum(1 for c in table if not c), hits))

print()
print("two things to notice. the longest chain grows much faster than the")
print("average, so the WORST case degrades faster than the average does.")
print("and even at alpha = 0.25 some buckets are empty, so a hash table")
print("always wastes some space. that is chapter 107's subject.")
   alpha  records    longest chain    empty buckets        avg hit
    0.25       25                2               78           1.08
    0.50       50                3               58           1.16
    0.75       75                3               47           1.32
    1.00      101                3               36           1.45
    2.00      202                5               14           1.96
    5.00      505               12                1           3.47
   10.00     1010               21                0           6.07

two things to notice. the longest chain grows much faster than the
average, so the WORST case degrades faster than the average does.
and even at alpha = 0.25 some buckets are empty, so a hash table
always wastes some space. that is chapter 107's subject.
munotes.in363

Chaining

Advantages and disadvantages

Advantages.

The load factor may exceed 1, so the table never fills up and an insert can never fail for want of space. No other scheme in this chapter group can say that.

Deletion is simple and complete. Unlink the node. Compare chapter 105, where deletion cannot remove anything.

Performance degrades gracefully, in proportion to alpha, as the measurement shows.

It is less sensitive to a mediocre hash function than open addressing, because a cluster of colliding keys lengthens one chain rather than disturbing the neighbours.

Any data structure can be used for a chain. A sorted list, a balanced tree, or a dynamic array. Java converts a chain to a balanced tree once it holds eight entries, which turns the worst case from O(n) into O(log n).

Disadvantages.

Extra memory for the links. A next pointer per record, plus the array of heads. For small records that overhead can exceed the data.

Poor cache locality. The nodes are scattered in memory, so walking a chain misses the cache repeatedly, while the next chapter's scheme walks a contiguous array. This is why open addressing is often faster in practice at the same alpha despite worse theory.

Empty buckets are wasted, and at a low load factor most buckets are empty.

The worst case is still O(n). If every key hashes to one slot, chaining gives one list of n records, which is exactly the structure the hash table was meant to replace.

The worst case, printed

M = 101
BAD_KEYS = [i * M for i in range(1, 51)]        # every one is 0 mod 101

table = [[] for _ in range(M)]
for k in BAD_KEYS:
    table[k % M].append(k)

print("50 keys, all multiples of %d, into %d buckets:" % (M, M))
print("   buckets used :", sum(1 for c in table if c))
print("   longest chain :", max(len(c) for c in table))
print("   every key in slot 0:", len(table[0]) == len(BAD_KEYS))
print()
hits = sum(table[k % M].index(k) + 1 for k in BAD_KEYS) / len(BAD_KEYS)
print("   average probes for a successful search : %.1f" % hits)
print("   a plain linked list of 50 records      : %.1f" % ((50 + 1) / 2))
print("   the same:", abs(hits - (50 + 1) / 2) < 0.01)
print()
print("chaining did not fail. it did exactly what it promises, and the")
print("promise was worth nothing because the HASH FUNCTION failed.")
print("no collision strategy can rescue a hash function that does not")
print("spread the keys. that is chapter 102, and it comes first.")
munotes.in364

Chaining

50 keys, all multiples of 101, into 101 buckets:
   buckets used : 1
   longest chain : 50
   every key in slot 0: True

   average probes for a successful search : 25.5
   a plain linked list of 50 records      : 25.5
   the same: True

chaining did not fail. it did exactly what it promises, and the
promise was worth nothing because the HASH FUNCTION failed.
no collision strategy can rescue a hash function that does not
spread the keys. that is chapter 102, and it comes first.

Quick revision

  • Chaining gives every bucket a list; colliding keys join the same list and nothing is displaced.
  • Each chain is Module 1's linked list, so insertion at the head is O(1) and deletion is a plain unlink.
  • Load factor alpha = n/m is the average chain length and may exceed 1, so the table never fills.
  • Unsuccessful search averages alpha probes; successful search averages 1 + alpha/2; insertion is 1 probe without a duplicate check.
  • Measured at alpha = 5 over 5,045 records, a successful search still averaged under 4 probes, against 2,523 for a linear search.
  • The longest chain grows faster than the average, so the worst case degrades faster than the average.
  • Advantages: no capacity limit, simple complete deletion, graceful degradation, tolerance of a mediocre hash function, and any structure may serve as a chain.
  • Disadvantages: a pointer per record, poor cache locality, wasted empty buckets, and an O(n) worst case.
  • Fifty keys that all hash to one slot gave an average of 25.5 probes, exactly a linked list of fifty: no collision strategy can rescue a bad hash function.

Test yourself

1. Describe chaining. Each bucket of the table holds a list, initially empty. A record is inserted into the list at its key's slot, searched for by walking that list, and deleted by unlinking it from that list.

2. What is the load factor, and what does it mean for chaining specifically? alpha = n/m, records divided by buckets. Under chaining it is the average chain length, and it may be any value, including above 1.

3. Give the average probe counts for successful and unsuccessful search, with the reasoning. Unsuccessful: alpha, because the entire chain is walked and its average length is alpha. Successful: 1 + alpha/2, because on average half the chain is walked before the record is reached.

4. Why can an insertion under chaining never fail for want of space? Because a chain has no capacity. The load factor may exceed 1, so there is always room.

5. Name three advantages of chaining over open addressing. It has no capacity limit; deletion is a simple unlink rather than needing a marker; and it tolerates a mediocre hash function better, since colliding keys lengthen one chain instead of disturbing other keys' slots.

munotes.in365

Chaining

6. Name three disadvantages. A pointer per record, poor cache locality because the nodes are scattered, and empty buckets wasted at low load factors.

7. What is chaining's worst case, and what causes it? O(n), when every key hashes to the same slot, which gives one list of n records. It is caused by the hash function, not by chaining: fifty keys all 0 mod 101 averaged 25.5 probes, exactly the figure for a linked list of fifty.

8. How does Java improve chaining's worst case? By converting a chain into a balanced tree once it holds eight entries, which makes the worst case O(log n) instead of O(n).

Contents This chapter on its own page

munotes.in366

Chapter One Hundred Five

Linear Probing

Syllabus topic Module 2, "collision avoidance techniques"; Computer Science Practical 3, "Implement a hash table with chaining or linear probing"

In one line

Linear probing keeps every record inside the array: if a key's slot is taken, it tries the next slot, and the next, until it finds a free one.

Open addressing

Chaining stores colliding records outside the array, in lists. Open addressing stores every record inside the array itself, so the array is the whole structure: no nodes, no pointers, no lists.

The immediate consequence: the table can fill up. The load factor cannot exceed 1, and an insertion into a full table must fail or grow the table. That is the trade for losing the pointers.

Linear probing is the simplest open addressing scheme.

probe sequence: h(k), h(k)+1, h(k)+2, h(k)+3, ... all mod m

h_i(k) = ( h(k) + i ) mod m for i = 0, 1, 2, ...

Insert: walk the sequence until an empty slot is found; put the record there. Search: walk the sequence until the key is found, or an empty slot is reached, which means the key is not in the table.

That second rule is the whole chapter. A search stops at an empty slot because an insert would have stopped there too, so the key cannot be any further along.

Built, and traced

EMPTY = None


class LinearTable:
    def __init__(self, buckets=11):
        self.slots = [EMPTY] * buckets
        self.count = 0

    def _hash(self, key):
        return key % len(self.slots)

    def insert(self, key, value, trace=False):
        start = self._hash(key)
        for i in range(len(self.slots)):
            at = (start + i) % len(self.slots)
            if self.slots[at] is EMPTY:
                self.slots[at] = (key, value)
                self.count += 1
                if trace:
                    print("   insert %-4d h=%-3d placed at %-3d after %d probe(s)"
                          % (key, start, at, i + 1))
                return at, i + 1
            if self.slots[at][0] == key:
                self.slots[at] = (key, value)
                return at, i + 1
        raise RuntimeError("the table is full")

    def search(self, key):
        start = self._hash(key)
        for i in range(len(self.slots)):
            at = (start + i) % len(self.slots)
            if self.slots[at] is EMPTY:
                return None, i + 1              # an empty slot ends the search
            if self.slots[at][0] == key:
                return self.slots[at][1], i + 1
        return None, len(self.slots)

    def show(self):
        return "\n".join(
            "   [%2d] %s" % (i, "%-5d %s" % cell if cell else "")
            for i, cell in enumerate(self.slots))


t = LinearTable(buckets=11)
print("inserting into an 11 slot table, h(k) = k mod 11:")
for k, name in ((25, "Aarti"), (36, "Bhavesh"), (14, "Chetna"),
                (47, "Devdatta"), (3, "Esha"), (58, "Farhan")):
    t.insert(k, name, trace=True)

print()
print("the table:")
print(t.show())
print()
print("ALL SIX keys hash to slot 3:",
      ", ".join("%d mod 11 = %d" % (k, k % 11) for k in (25, 36, 14, 47, 3, 58)))
print("so they filled slots 3 to 8 in a run, which is a CLUSTER.")
print()
for k in (58, 3, 99):
    value, probes = t.search(k)
    print("search %-4d -> %-10s in %d probe(s)"
          % (k, value if value else "absent", probes))
print()
print("finding 58 took 6 probes, because it was inserted last and so sits")
print("at the far end of the cluster. finding 99 took 1, because slot 0 is")
print("empty and the search stops there at once.")
munotes.in367

Linear Probing

inserting into an 11 slot table, h(k) = k mod 11:
   insert 25   h=3   placed at 3   after 1 probe(s)
   insert 36   h=3   placed at 4   after 2 probe(s)
   insert 14   h=3   placed at 5   after 3 probe(s)
   insert 47   h=3   placed at 6   after 4 probe(s)
   insert 3    h=3   placed at 7   after 5 probe(s)
   insert 58   h=3   placed at 8   after 6 probe(s)

the table:
   [ 0]
   [ 1]
   [ 2]
   [ 3] 25    Aarti
   [ 4] 36    Bhavesh
   [ 5] 14    Chetna
   [ 6] 47    Devdatta
   [ 7] 3     Esha
   [ 8] 58    Farhan
   [ 9]
   [10]

ALL SIX keys hash to slot 3: 25 mod 11 = 3, 36 mod 11 = 3, 14 mod 11 = 3, 47 mod 11 = 3, 3 mod 11 = 3, 58 mod 11 = 3
so they filled slots 3 to 8 in a run, which is a CLUSTER.

search 58   -> Farhan     in 6 probe(s)
search 3    -> Esha       in 5 probe(s)
search 99   -> absent     in 1 probe(s)

finding 58 took 6 probes, because it was inserted last and so sits
at the far end of the cluster. finding 99 took 1, because slot 0 is
empty and the search stops there at once.

Primary clustering

That run of occupied slots is the scheme's characteristic weakness, and it has a name: primary clustering.

The mechanism is worth stating exactly, because it is more than "collisions pile up". Once a cluster of length L exists, any key that hashes anywhere inside it joins its end, which makes the cluster L+1 long. So a long cluster is more likely to grow than a short one, and clusters merge. Growth feeds on itself.

def mix(i):
    x = ((i + 1) * 2654435761) % (2 ** 32)
    x ^= x >> 13
    x = (x * 2246822519) % (2 ** 32)
    x ^= x >> 17
    return x


M = 1009


def fill(n, m):
    slots = [None] * m
    for i in range(n):
        k = mix(i)
        at = k % m
        while slots[at] is not None:
            at = (at + 1) % m
        slots[at] = k
    return slots


def clusters(slots):
    """Lengths of the runs of occupied slots, treating the array as a ring."""
    m = len(slots)
    if all(s is not None for s in slots):
        return [m]
    start = next(i for i in range(m) if slots[i] is None)
    runs, run = [], 0
    for step in range(m):
        at = (start + step) % m
        if slots[at] is None:
            if run:
                runs.append(run)
            run = 0
        else:
            run += 1
    if run:
        runs.append(run)
    return runs or [0]


print("%8s %10s %14s %16s %16s"
      % ("alpha", "records", "clusters", "longest cluster", "average cluster"))
for alpha in (0.25, 0.5, 0.7, 0.8, 0.9, 0.95, 0.99):
    n = int(alpha * M)
    runs = clusters(fill(n, M))
    print("%8.2f %10d %14d %16d %16.1f"
          % (alpha, n, len(runs), max(runs), sum(runs) / len(runs)))

print()
print("read the longest-cluster column. between alpha 0.5 and alpha 0.9 the")
print("record count not quite doubles, and the longest cluster grows far")
print("more than that. clusters do not grow in proportion. they snowball.")
munotes.in368

Linear Probing

   alpha    records       clusters  longest cluster  average cluster
    0.25        252            170                7              1.5
    0.50        504            187               21              2.7
    0.70        706            141               46              5.0
    0.80        807            102               96              7.9
    0.90        908             50              213             18.2
    0.95        958             28              702             34.2
    0.99        998              5              981            199.6

read the longest-cluster column. between alpha 0.5 and alpha 0.9 the
record count not quite doubles, and the longest cluster grows far
more than that. clusters do not grow in proportion. they snowball.

The cost, measured against the formulas

unsuccessful search : about ( 1 + 1 / (1 - alpha)^2 ) / 2

successful search : about ( 1 + 1 / (1 - alpha) ) / 2

Both blow up as alpha approaches 1, and the unsuccessful one blows up as the square.

def mix(i):
    x = ((i + 1) * 2654435761) % (2 ** 32)
    x ^= x >> 13
    x = (x * 2246822519) % (2 ** 32)
    x ^= x >> 17
    return x


M = 1009


def build(n, m):
    slots = [None] * m
    keys = [mix(i) for i in range(n)]
    for k in keys:
        at = k % m
        while slots[at] is not None:
            at = (at + 1) % m
        slots[at] = k
    return slots, keys


def measure(n, m):
    slots, keys = build(n, m)

    hit_probes = 0
    for k in keys:
        at, probes = k % m, 1
        while slots[at] != k:
            at = (at + 1) % m
            probes += 1
        hit_probes += probes

    absent = [mix(500000 + i) for i in range(500)]
    miss_probes = 0
    for k in absent:
        at, probes = k % m, 1
        while slots[at] is not None:
            at = (at + 1) % m
            probes += 1
        miss_probes += probes

    return hit_probes / len(keys), miss_probes / len(absent)


print("%7s %10s %12s %12s %14s %14s"
      % ("alpha", "records", "hit found", "hit formula", "miss found", "miss formula"))
for alpha in (0.25, 0.5, 0.75, 0.9, 0.95):
    n = int(alpha * M)
    hit, miss = measure(n, M)
    hit_formula = (1 + 1 / (1 - alpha)) / 2
    miss_formula = (1 + 1 / (1 - alpha) ** 2) / 2
    print("%7.2f %10d %12.2f %12.2f %14.2f %14.2f"
          % (alpha, n, hit, hit_formula, miss, miss_formula))

print()
print("compare chaining at the same load factors, where an unsuccessful")
print("search costs alpha probes: 0.25, 0.50, 0.75, 0.90, 0.95.")
print("at alpha = 0.95 linear probing needs over a hundred probes for a")
print("failed search and chaining needs one. that gap is the price of")
print("keeping everything inside the array.")
munotes.in369

Linear Probing

  alpha    records    hit found  hit formula     miss found   miss formula
   0.25        252         1.15         1.17           1.43           1.39
   0.50        504         1.65         1.50           2.86           2.50
   0.75        756         2.84         2.50          10.24           8.50
   0.90        908         5.94         5.50          50.54          50.50
   0.95        958        11.24        10.50         247.34         200.50

compare chaining at the same load factors, where an unsuccessful
search costs alpha probes: 0.25, 0.50, 0.75, 0.90, 0.95.
at alpha = 0.95 linear probing needs over a hundred probes for a
failed search and chaining needs one. that gap is the price of
keeping everything inside the array.

Read the last two columns. At a load factor of 0.95 an unsuccessful search cost 247 probes, where chaining at the same load factor costs about one. This is why a linear probing table must be kept well under full, and chapter 107 puts a number on "well under".

Note also that the measured figures run above the formulas at high load: 11.2 against 10.5, and 247 against 200. That is expected and worth knowing. Both formulas are approximations derived for an idealised uniform hash on a large table, so they give the shape of the growth rather than an exact count for a particular 1009 slot table. The shape is what they are for, and the shape is confirmed: the failed search grows as the square.

The defect: a deleted slot cannot simply be cleared

EMPTY = None

M = 7
slots = [EMPTY] * M

# three keys that all hash to slot 1
for k in (1, 8, 15):
    at = k % M
    while slots[at] is not EMPTY:
        at = (at + 1) % M
    slots[at] = k

print("three keys, all 1 mod %d, inserted in order 1, 8, 15:" % M)
for i, cell in enumerate(slots):
    print("   [%d] %s" % (i, cell if cell is not None else ""))
print()


def search(slots, key):
    """The standard search: stop at an EMPTY slot."""
    m = len(slots)
    at = key % m
    for i in range(m):
        here = (at + i) % m
        if slots[here] is EMPTY:
            return False, i + 1
        if slots[here] == key:
            return True, i + 1
    return False, m


for k in (1, 8, 15):
    found, probes = search(slots, k)
    print("   search %-3d : found %-5s in %d probe(s)" % (k, found, probes))

print()
print("now delete 8 the obvious way, by clearing its slot:")
slots[slots.index(8)] = EMPTY
for i, cell in enumerate(slots):
    print("   [%d] %s" % (i, cell if cell is not None else ""))
print()
found, probes = search(slots, 15)
print("   search 15  : found %s   <-- but look at the table" % found)
print("   15 is in the table at slot:", slots.index(15))
print("   the search stopped at the cleared slot 2 and reported absent.")
print()
print("a record is PRESENT and UNREACHABLE. nothing raised an error.")
print("this is the same class of bug as chapter 93's renumbering: the")
print("structure was changed under a rule the rest of the code relies on.")
munotes.in370

Linear Probing

three keys, all 1 mod 7, inserted in order 1, 8, 15:
   [0]
   [1] 1
   [2] 8
   [3] 15
   [4]
   [5]
   [6]

   search 1   : found True  in 1 probe(s)
   search 8   : found True  in 2 probe(s)
   search 15  : found True  in 3 probe(s)

now delete 8 the obvious way, by clearing its slot:
   [0]
   [1] 1
   [2]
   [3] 15
   [4]
   [5]
   [6]

   search 15  : found False   <-- but look at the table
   15 is in the table at slot: 3
   the search stopped at the cleared slot 2 and reported absent.

a record is PRESENT and UNREACHABLE. nothing raised an error.
this is the same class of bug as chapter 93's renumbering: the
structure was changed under a rule the rest of the code relies on.

Fifteen is sitting in the table and search says it is not there, because clearing slot 2 broke the probe chain that led to it. An empty slot is a promise that nothing beyond it was displaced from before it, and clearing a slot breaks that promise.

The fix: a tombstone

Mark the slot deleted rather than empty. A search treats a deleted slot as "keep going"; an insert treats it as "you may use this". Three states, not two.

class Tombstone:
    def __repr__(self):
        return "DELETED"


EMPTY = None
DELETED = Tombstone()

M = 7


def build():
    slots = [EMPTY] * M
    for k in (1, 8, 15):
        at = k % M
        while slots[at] is not EMPTY and slots[at] is not DELETED:
            at = (at + 1) % M
        slots[at] = k
    return slots


def search(slots, key):
    """A DELETED slot does not end the search. Only an EMPTY one does."""
    m = len(slots)
    for i in range(m):
        here = (key % m + i) % m
        if slots[here] is EMPTY:
            return False, i + 1
        if slots[here] is not DELETED and slots[here] == key:
            return True, i + 1
    return False, m


def delete(slots, key):
    m = len(slots)
    for i in range(m):
        here = (key % m + i) % m
        if slots[here] is EMPTY:
            return False
        if slots[here] is not DELETED and slots[here] == key:
            slots[here] = DELETED              # marked, NOT cleared
            return True
    return False


def insert(slots, key):
    """An insert MAY reuse a deleted slot, so space is not lost."""
    m = len(slots)
    for i in range(m):
        here = (key % m + i) % m
        if slots[here] is EMPTY or slots[here] is DELETED:
            slots[here] = key
            return here
    raise RuntimeError("full")


slots = build()
print("delete 8 with a tombstone:", delete(slots, 8))
print("the table now:")
for i, cell in enumerate(slots):
    print("   [%d] %s" % (i, cell if cell is not None else ""))
print()
for k in (1, 8, 15):
    found, probes = search(slots, k)
    print("   search %-3d : found %-5s in %d probe(s)" % (k, found, probes))
print()
print("15 is findable again:", search(slots, 15)[0])
print("8 is correctly absent:", search(slots, 8)[0] is False)
print()
print("and the deleted slot is reused rather than wasted:")
at = insert(slots, 22)                      # 22 mod 7 = 1
print("   inserted 22 at slot", at, "which held the tombstone:", at == 2)
for i, cell in enumerate(slots):
    print("   [%d] %s" % (i, cell if cell is not None else ""))
munotes.in371

Linear Probing

delete 8 with a tombstone: True
the table now:
   [0]
   [1] 1
   [2] DELETED
   [3] 15
   [4]
   [5]
   [6]

   search 1   : found True  in 1 probe(s)
   search 8   : found False in 4 probe(s)
   search 15  : found True  in 3 probe(s)

15 is findable again: True
8 is correctly absent: True

and the deleted slot is reused rather than wasted:
   inserted 22 at slot 2 which held the tombstone: True
   [0]
   [1] 1
   [2] 22
   [3] 15
   [4]
   [5]
   [6]

That is lazy deletion again, the same device chapter 93 used for a deleted vertex and chapter 98's heap used for a superseded distance. Three times in one module, which is why it is worth naming.

What tombstones cost

A tombstone occupies a slot for searching purposes but holds no record. So a table that has had many deletions can be mostly tombstones: searches stay slow although the table is nearly empty. The only cure is to rebuild the table, which is chapter 107's rehashing.

class Tombstone:
    def __repr__(self):
        return "DELETED"


EMPTY = None
DELETED = Tombstone()
M = 1009


def mix(i):
    x = ((i + 1) * 2654435761) % (2 ** 32)
    x ^= x >> 13
    x = (x * 2246822519) % (2 ** 32)
    x ^= x >> 17
    return x


def fresh(n):
    slots = [EMPTY] * M
    keys = [mix(i) for i in range(n)]
    for k in keys:
        at = k % M
        while slots[at] is not EMPTY:
            at = (at + 1) % M
        slots[at] = k
    return slots, keys


def miss_probes(slots):
    total = 0
    absent = [mix(900000 + i) for i in range(400)]
    for k in absent:
        at, probes = k % M, 1
        while slots[at] is not EMPTY:
            at = (at + 1) % M
            probes += 1
        total += probes
    return total / len(absent)


slots, keys = fresh(900)                    # alpha = 0.89
print("a table of %d slots holding %d records (alpha %.2f):" % (M, 900, 900 / M))
print("   failed search costs %.1f probes" % miss_probes(slots))

for k in keys[:800]:                        # delete 800 of the 900
    at = k % M
    while True:
        if slots[at] is not DELETED and slots[at] == k:
            slots[at] = DELETED
            break
        at = (at + 1) % M

live = sum(1 for s in slots if s is not EMPTY and s is not DELETED)
tombs = sum(1 for s in slots if s is DELETED)
print()
print("now 800 of them are deleted:")
print("   live records : %d  (alpha %.2f)" % (live, live / M))
print("   tombstones   : %d" % tombs)
print("   failed search costs %.1f probes" % miss_probes(slots))
print()
print("the table holds %d records in %d slots, which should be fast," % (live, M))
print("and a failed search is as slow as it was at alpha 0.89, because the")
print("tombstones still have to be walked past.")
print("only rebuilding the table clears them: chapter 107.")
munotes.in372

Linear Probing

a table of 1009 slots holding 900 records (alpha 0.89):
   failed search costs 46.6 probes

now 800 of them are deleted:
   live records : 100  (alpha 0.10)
   tombstones   : 800
   failed search costs 46.6 probes

the table holds 100 records in 1009 slots, which should be fast,
and a failed search is as slow as it was at alpha 0.89, because the
tombstones still have to be walked past.
only rebuilding the table clears them: chapter 107.

Advantages and disadvantages

Advantages.

No pointers at all, so no memory per record beyond the record, and the array of slots is the whole structure.

Excellent cache locality. The probe sequence walks consecutive array positions, which is the fastest possible memory access pattern. At moderate load factors this often makes linear probing faster in practice than chaining, despite chaining's better probe counts, and a good answer says so.

Simple to implement, with no second data structure.

Disadvantages.

Primary clustering, which snowballs rather than growing in proportion.

The load factor must stay below 1, and in practice well below: at 0.95 a failed search cost about two hundred probes.

Deletion needs tombstones, and tombstones accumulate, so a delete-heavy table must be rebuilt.

munotes.in373

Linear Probing

It is far more sensitive to a poor hash function than chaining, because a cluster does not just lengthen a chain, it displaces other keys' records and spreads the damage.

Quick revision

  • Linear probing is open addressing: every record lives in the array, so there are no pointers and the table can fill.
  • The probe sequence is h(k), h(k)+1, h(k)+2 and so on, all mod m.
  • A search stops at an EMPTY slot, because an insert would have stopped there too.
  • Primary clustering: a key hashing anywhere inside a cluster joins its end, so long clusters grow faster than short ones and clusters merge.
  • Costs: successful search about (1 + 1/(1-alpha))/2, unsuccessful about (1 + 1/(1-alpha) squared)/2, so a failed search blows up as the square.
  • At alpha = 0.95 a failed search cost 247 probes against chaining's 1; the formulas are idealised approximations, so measured figures run above them at high load.
  • Clearing a deleted slot breaks the probe chain: the chapter's listing left key 15 in the table with search reporting it absent.
  • The fix is a tombstone: three states, empty, occupied and deleted. A search passes a tombstone; an insert may reuse it.
  • Tombstones accumulate: a table with 100 live records and 800 tombstones searched as slowly as it did when nearly full. Only rehashing clears them.
  • Advantages: no pointers and excellent cache locality, which often beats chaining in practice at moderate load.
  • Disadvantages: clustering, alpha below 1, tombstones, and high sensitivity to a poor hash function.

Test yourself

1. What is open addressing, and what does it cost compared with chaining? Every record is stored inside the table array itself, with no external lists. The cost is that the table can fill up, so the load factor cannot exceed 1.

2. Give the probe sequence for linear probing. h_i(k) = (h(k) + i) mod m for i = 0, 1, 2 and so on, so h(k), then the next slot, then the next.

3. Why does a search stop at an empty slot? Because an insertion would have stopped at that slot too, so no key that hashes before it can have been placed beyond it.

4. Explain primary clustering and why it is worse than proportional growth. A cluster is a run of occupied slots. Any key hashing anywhere inside a cluster of length L is placed at its end, making it L+1, so long clusters attract more keys than short ones and adjacent clusters merge. Growth therefore accelerates.

5. Give both cost formulas and say which is worse near a full table. Successful search about (1 + 1/(1 - alpha))/2; unsuccessful about (1 + 1/(1 - alpha) squared)/2. The unsuccessful one is worse, because it grows as the square: the formula gives about 200 at alpha = 0.95 and a 1009 slot table measured 247, the formulas being idealised approximations.

munotes.in374

Linear Probing

6. Show what goes wrong if a deleted slot is simply cleared. With 1, 8, 15 all hashing to slot 1 and placed in slots 1, 2, 3, clearing slot 2 to delete 8 makes a search for 15 stop at the now empty slot 2 and report absent, although 15 is still in slot 3.

7. What is a tombstone, and how do search and insert treat it? A marker meaning "a record was deleted here". A search passes over it and continues; an insert may place a new record in it.

8. What is the drawback of tombstones, and the cure? They occupy slots for searching but hold no record, so they accumulate and keep searches slow in a table that is nearly empty: 100 live records with 800 tombstones searched as slowly as 900 live records. The cure is to rebuild the table, which is rehashing.

9. Why is linear probing often faster than chaining in practice despite worse probe counts? Because its probes walk consecutive array positions, which suits the memory cache, while chaining follows pointers to scattered nodes.

Contents This chapter on its own page

munotes.in375

Chapter One Hundred Six

Quadratic Probing and Double Hashing

Syllabus topic Module 2, "collision avoidance techniques"

In one line

Quadratic probing spreads the probe sequence out so that clusters do not form, at the price of being unable to reach every slot; double hashing uses a second hash function to give each key its own probe sequence, which is the best of the three schemes.

Why linear probing needed replacing

Linear probing's defect was primary clustering: a key that hashes anywhere inside a run of occupied slots is placed at the end of that run, so runs grow and merge. The cause is that the probe sequence steps by 1, so every key that enters a cluster follows the same path out of it.

Both schemes here change the step.

Quadratic probing

h_i(k) = ( h(k) + i^2 ) mod m for i = 0, 1, 2, 3, ...

so the offsets are 0, 1, 4, 9, 16, 25, ...

The jumps grow, so a key displaced from a cluster lands far away rather than at the cluster's edge, and clusters cannot build up by adjacency.

M = 11
EMPTY = None


def quadratic_insert(slots, key, trace=False):
    m = len(slots)
    for i in range(m):
        at = (key % m + i * i) % m
        if slots[at] is EMPTY:
            slots[at] = key
            if trace:
                print("   %-4d h=%-3d i=%-2d offset %-3d -> slot %-3d (%d probes)"
                      % (key, key % m, i, i * i, at, i + 1))
            return at
    return None                                  # could not place it


slots = [EMPTY] * M
print("quadratic probing, m = %d, six keys that ALL hash to slot 3:" % M)
KEYS = [25, 36, 47, 58, 69, 80]
for k in KEYS:
    print("      %d mod %d = %d" % (k, M, k % M))
print()
for k in KEYS:
    at = quadratic_insert(slots, k, trace=True)
    if at is None:
        print("   %-4d COULD NOT BE PLACED" % k)

print()
print("the table:")
print("   " + " ".join("%4s" % (s if s is not None else "-") for s in slots))
print("   " + " ".join("%4d" % i for i in range(M)))
print()
print("slots used:", sum(1 for s in slots if s is not None), "of", M)
print("the offsets were 0, 1, 4, 9, 16, 25 modulo 11, which is",
      [(i * i) % M for i in range(6)])
print("so the records are scattered, not in a run: no primary clustering.")
quadratic probing, m = 11, six keys that ALL hash to slot 3:
      25 mod 11 = 3
      36 mod 11 = 3
      47 mod 11 = 3
      58 mod 11 = 3
      69 mod 11 = 3
      80 mod 11 = 3

   25   h=3   i=0  offset 0   -> slot 3   (1 probes)
   36   h=3   i=1  offset 1   -> slot 4   (2 probes)
   47   h=3   i=2  offset 4   -> slot 7   (3 probes)
   58   h=3   i=3  offset 9   -> slot 1   (4 probes)
   69   h=3   i=4  offset 16  -> slot 8   (5 probes)
   80   h=3   i=5  offset 25  -> slot 6   (6 probes)

the table:
      -   58    -   25   36    -   80   47   69    -    -
      0    1    2    3    4    5    6    7    8    9   10

slots used: 6 of 11
the offsets were 0, 1, 4, 9, 16, 25 modulo 11, which is [0, 1, 4, 9, 5, 3]
so the records are scattered, not in a run: no primary clustering.
munotes.in376

Quadratic Probing and Double Hashing

The failure quadratic probing is usually not told about

EMPTY = None


def quadratic_insert(slots, key):
    m = len(slots)
    for i in range(m):
        at = (key % m + i * i) % m
        if slots[at] is EMPTY:
            slots[at] = key
            return at
    return None


for M in (16, 11):
    slots = [EMPTY] * M
    reachable = sorted({(0 + i * i) % M for i in range(M)})
    print("m = %d (%s)" % (M, "a power of 2" if M == 16 else "prime"))
    print("   offsets i*i mod m, for i = 0..%d : %s"
          % (M - 1, [(i * i) % M for i in range(M)]))
    print("   DISTINCT slots a key hashing to 0 can ever reach: %s" % reachable)
    print("   that is %d of %d slots, which is %.0f%%"
          % (len(reachable), M, 100 * len(reachable) / M))

    placed, refused = 0, []
    for j in range(M):
        key = j * M                              # every key is 0 mod M
        at = quadratic_insert(slots, key)
        if at is None:
            refused.append(key)
        else:
            placed += 1
    free = sum(1 for s in slots if s is EMPTY)
    print("   inserting %d keys that all hash to 0: %d placed, %d REFUSED"
          % (M, placed, len(refused)))
    print("   and the table still has %d free slots out of %d" % (free, M))
    print()

print("so quadratic probing can refuse an insertion into a table that is")
print("mostly empty. this is not a bug in the code; it is the arithmetic.")
print("i*i mod m simply does not visit every residue.")
m = 16 (a power of 2)
   offsets i*i mod m, for i = 0..15 : [0, 1, 4, 9, 0, 9, 4, 1, 0, 1, 4, 9, 0, 9, 4, 1]
   DISTINCT slots a key hashing to 0 can ever reach: [0, 1, 4, 9]
   that is 4 of 16 slots, which is 25%
   inserting 16 keys that all hash to 0: 4 placed, 12 REFUSED
   and the table still has 12 free slots out of 16

m = 11 (prime)
   offsets i*i mod m, for i = 0..10 : [0, 1, 4, 9, 5, 3, 3, 5, 9, 4, 1]
   DISTINCT slots a key hashing to 0 can ever reach: [0, 1, 3, 4, 5, 9]
   that is 6 of 11 slots, which is 55%
   inserting 11 keys that all hash to 0: 6 placed, 5 REFUSED
   and the table still has 5 free slots out of 11

so quadratic probing can refuse an insertion into a table that is
mostly empty. this is not a bug in the code; it is the arithmetic.
i*i mod m simply does not visit every residue.
munotes.in377

Quadratic Probing and Double Hashing

That is the defect. With m = 16 a key that hashes to slot 0 can only ever reach 4 of the 16 slots, so the fifth such key cannot be inserted although 12 slots are free. The insertion does not slow down. It fails.

The rescue is a theorem, and it is the origin of a rule students learn without the reason:

If m is prime, quadratic probing with offsets i squared is guaranteed to find an empty slot whenever

the table is less than half full.

With m prime, the offsets i squared mod m take exactly (m+1)/2 distinct values, so a key can reach just over half the table. Hence alpha must stay below 0.5, and m must be prime.

def reachable_count(m):
    return len({(i * i) % m for i in range(m)})


print("%8s %10s %18s %16s %14s"
      % ("m", "prime?", "slots reachable", "(m+1)/2", "fraction"))


def is_prime(n):
    if n < 2:
        return False
    d = 2
    while d * d <= n:
        if n % d == 0:
            return False
        d += 1
    return True


for m in (11, 13, 16, 17, 31, 32, 101, 128):
    count = reachable_count(m)
    print("%8d %10s %18d %16d %13.0f%%"
          % (m, "yes" if is_prime(m) else "NO", count, (m + 1) // 2,
             100 * count / m))

print()
primes = [m for m in (11, 13, 17, 31, 101) ]
print("for every prime above, reachable slots equal (m+1)/2 exactly:",
      all(reachable_count(m) == (m + 1) // 2 for m in primes))
print("for the powers of 2 it is far worse:",
      [(m, reachable_count(m)) for m in (16, 32, 128)])
print()
print("hence the two rules together: m PRIME, and alpha below 0.5.")
print("neither rule alone is enough.")
       m     prime?    slots reachable          (m+1)/2       fraction
      11        yes                  6                6            55%
      13        yes                  7                7            54%
      16         NO                  4                8            25%
      17        yes                  9                9            53%
      31        yes                 16               16            52%
      32         NO                  7               16            22%
     101        yes                 51               51            50%
     128         NO                 23               64            18%

for every prime above, reachable slots equal (m+1)/2 exactly: True
for the powers of 2 it is far worse: [(16, 4), (32, 7), (128, 23)]

hence the two rules together: m PRIME, and alpha below 0.5.
neither rule alone is enough.
munotes.in378

Quadratic Probing and Double Hashing

Secondary clustering

Quadratic probing removes primary clustering and leaves a weaker relative behind. Two keys with the same hash value follow exactly the same probe sequence, because the sequence depends only on h(k). So keys that collide at the start collide at every step. That is secondary clustering.

M = 101

same_hash = [i * M + 7 for i in range(5)]       # all 7 mod 101
print("five keys, all %d mod %d:" % (7, M), same_hash)
print()
print("their quadratic probe sequences, first 6 steps:")
for k in same_hash:
    seq = [(k % M + i * i) % M for i in range(6)]
    print("   %-8d %s" % (k, seq))
print()
print("identical, because the sequence depends only on h(k) = %d." % 7)
print("that is SECONDARY clustering: fewer keys are affected than under")
print("primary clustering, but those that are, are affected completely.")
five keys, all 7 mod 101: [7, 108, 209, 310, 411]

their quadratic probe sequences, first 6 steps:
   7        [7, 8, 11, 16, 23, 32]
   108      [7, 8, 11, 16, 23, 32]
   209      [7, 8, 11, 16, 23, 32]
   310      [7, 8, 11, 16, 23, 32]
   411      [7, 8, 11, 16, 23, 32]

identical, because the sequence depends only on h(k) = 7.
that is SECONDARY clustering: fewer keys are affected than under
primary clustering, but those that are, are affected completely.

Double hashing

The cure for secondary clustering is to make the step itself depend on the key.

h_i(k) = ( h1(k) + i x h2(k) ) mod m for i = 0, 1, 2, ...

Two conditions on h2, and both are examinable:

h2(k) must never be 0, or the probe sequence never moves and the insertion loops for ever.

h2(k) must be coprime with m, or the sequence visits only some slots. The usual arrangement is m prime and h2(k) = 1 + (k mod (m - 1)), which is between 1 and m-1 and therefore coprime with a prime m.

M = 11
EMPTY = None


def h1(k, m):
    return k % m


def h2(k, m):
    return 1 + (k % (m - 1))                     # never 0, always coprime with a prime m


def double_insert(slots, key, trace=False):
    m = len(slots)
    step = h2(key, m)
    for i in range(m):
        at = (h1(key, m) + i * step) % m
        if slots[at] is EMPTY:
            slots[at] = key
            if trace:
                print("   %-4d h1=%-3d h2=%-3d i=%-2d -> slot %-3d (%d probes)"
                      % (key, h1(key, m), step, i, at, i + 1))
            return at
    return None


slots = [EMPTY] * M
KEYS = [25, 36, 47, 58, 69, 80]
print("double hashing, m = %d, the same six keys that all hash to slot 3:" % M)
for k in KEYS:
    double_insert(slots, k, trace=True)

print()
print("the table:")
print("   " + " ".join("%4s" % (s if s is not None else "-") for s in slots))
print("   " + " ".join("%4d" % i for i in range(M)))
print()
print("every key has its OWN step, so no two sequences agree:")
for k in KEYS:
    print("   %-4d step %-3d sequence %s"
          % (k, h2(k, M), [(h1(k, M) + i * h2(k, M)) % M for i in range(5)]))
print()
print("no two of those sequences are the same, although all six keys have")
print("the same h1. that is secondary clustering removed.")
print()
print("and with m prime and h2 between 1 and m-1, a sequence reaches every")
print("slot. for key 25, step %d:" % h2(25, M))
print("   ", sorted({(h1(25, M) + i * h2(25, M)) % M for i in range(M)}))
print("   all %d slots:" % M,
      len({(h1(25, M) + i * h2(25, M)) % M for i in range(M)}) == M)
munotes.in379

Quadratic Probing and Double Hashing

double hashing, m = 11, the same six keys that all hash to slot 3:
   25   h1=3   h2=6   i=0  -> slot 3   (1 probes)
   36   h1=3   h2=7   i=1  -> slot 10  (2 probes)
   47   h1=3   h2=8   i=1  -> slot 0   (2 probes)
   58   h1=3   h2=9   i=1  -> slot 1   (2 probes)
   69   h1=3   h2=10  i=1  -> slot 2   (2 probes)
   80   h1=3   h2=1   i=1  -> slot 4   (2 probes)

the table:
     47   58   69   25   80    -    -    -    -    -   36
      0    1    2    3    4    5    6    7    8    9   10

every key has its OWN step, so no two sequences agree:
   25   step 6   sequence [3, 9, 4, 10, 5]
   36   step 7   sequence [3, 10, 6, 2, 9]
   47   step 8   sequence [3, 0, 8, 5, 2]
   58   step 9   sequence [3, 1, 10, 8, 6]
   69   step 10  sequence [3, 2, 1, 0, 10]
   80   step 1   sequence [3, 4, 5, 6, 7]

no two of those sequences are the same, although all six keys have
the same h1. that is secondary clustering removed.

and with m prime and h2 between 1 and m-1, a sequence reaches every
slot. for key 25, step 6:
    [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10]
   all 11 slots: True

Compare the two defects, one line each: quadratic probing can reach only half the table, double hashing reaches all of it.

The three schemes measured on the same keys

def mix(i):
    x = ((i + 1) * 2654435761) % (2 ** 32)
    x ^= x >> 13
    x = (x * 2246822519) % (2 ** 32)
    x ^= x >> 17
    return x


M = 1009
EMPTY = None


def fill(n, step_of):
    slots = [EMPTY] * M
    total = 0
    for j in range(n):
        k = mix(j)
        for i in range(M):
            at = step_of(k, i)
            total += 1
            if slots[at] is EMPTY:
                slots[at] = k
                break
    return slots, total / n


def linear(k, i):
    return (k % M + i) % M


def quadratic(k, i):
    return (k % M + i * i) % M


def double(k, i):
    return (k % M + i * (1 + k % (M - 1))) % M


def longest_run(slots):
    best = run = 0
    for cell in slots + slots:                   # as a ring
        run = run + 1 if cell is not EMPTY else 0
        best = max(best, run)
    return min(best, M)


def chance_fill(n):
    """Occupy n slots with NO probing at all: the baseline for a run length."""
    slots = [EMPTY] * M
    placed, j = 0, 0
    while placed < n:
        at = mix(j + 10 ** 6) % M
        j += 1
        if slots[at] is EMPTY:
            slots[at] = True
            placed += 1
    return slots


print("%7s %22s %16s %14s %14s"
      % ("alpha", "scheme", "probes per insert", "longest run", "x chance"))
for alpha in (0.4, 0.6, 0.8):
    n = int(alpha * M)
    base = longest_run(chance_fill(n))
    for name, fn in (("linear probing", linear),
                     ("quadratic probing", quadratic),
                     ("double hashing", double)):
        slots, avg = fill(n, fn)
        run = longest_run(slots)
        print("%7.2f %22s %17.2f %14d %13.1fx"
              % (alpha, name, avg, run, run / base))
    print("%7.2f %22s %17s %14d %13.1fx"
          % (alpha, "(no probing: chance)", "n/a", base, 1.0))
    print()

print("double hashing needs the fewest probes per insert at every alpha and")
print("quadratic probing is second. that is the column that matters.")
print()
print("the longest-run column needs the chance row beside it. a run of")
print("ADJACENT occupied slots is what linear probing builds on purpose, and")
print("at alpha 0.8 its run is almost 5 times what chance alone gives. the")
print("other two are above chance as well, at roughly 1.5 and 2.5 times, so")
print("they are not free of adjacency at high load: they are merely far")
print("better than linear probing at it. the ordering between those two is")
print("not stable and should not be read as a result.")
munotes.in380

Quadratic Probing and Double Hashing

  alpha                 scheme probes per insert    longest run       x chance
   0.40         linear probing              1.33             12           1.5x
   0.40      quadratic probing              1.29             11           1.4x
   0.40         double hashing              1.27              7           0.9x
   0.40   (no probing: chance)               n/a              8           1.0x

   0.60         linear probing              2.08             35           2.5x
   0.60      quadratic probing              1.66             19           1.4x
   0.60         double hashing              1.52             14           1.0x
   0.60   (no probing: chance)               n/a             14           1.0x

   0.80         linear probing              3.24             96           4.8x
   0.80      quadratic probing              2.24             32           1.6x
   0.80         double hashing              2.01             48           2.4x
   0.80   (no probing: chance)               n/a             20           1.0x

double hashing needs the fewest probes per insert at every alpha and
quadratic probing is second. that is the column that matters.

the longest-run column needs the chance row beside it. a run of
ADJACENT occupied slots is what linear probing builds on purpose, and
at alpha 0.8 its run is almost 5 times what chance alone gives. the
other two are above chance as well, at roughly 1.5 and 2.5 times, so
they are not free of adjacency at high load: they are merely far
better than linear probing at it. the ordering between those two is
not stable and should not be read as a result.
munotes.in381

Quadratic Probing and Double Hashing

The three schemes compared

Linear probingQuadratic probingDouble hashing
Probe sequenceh + ih + i squaredh1 + i x h2
Primary clusteringyesnono
Secondary clusteringyesyesno
Slots reachableall m(m+1)/2 when m is primeall m
Load factor limitbelow 1, in practice 0.7below 0.5below 1, in practice 0.7
Table sizeanymust be primeprime, or h2 coprime with m
Cache localitybestmoderateworst
Hash costone hashone hashtwo hashes
Probe distributionpoorestmiddlingbest

What to say when asked which to use. Double hashing has the best theory and is the right answer for an open addressing table that must run near its limit. Linear probing is often faster in practice at moderate load because of cache locality. Quadratic probing sits between them and carries the hard constraint that the table must be prime and less than half full, which is why it is the least used of the three.

Quick revision

  • Quadratic probing: h_i(k) = (h(k) + i squared) mod m. The growing jumps prevent primary clustering.
  • It cannot reach every slot. With m = 16 a key hashing to 0 reaches only 4 slots, and a fifth such key was refused while 12 slots were free.
  • With m prime it reaches exactly (m+1)/2 slots, so m must be prime AND alpha must stay below 0.5. Neither rule alone suffices.
  • Secondary clustering: two keys with the same h(k) have identical probe sequences, since the sequence depends only on h(k).
  • Double hashing: h_i(k) = (h1(k) + i x h2(k)) mod m, so the step itself depends on the key.
  • h2(k) must never be 0, or the sequence never moves, and must be coprime with m, or it cannot reach every slot. With m prime, h2(k) = 1 + (k mod (m-1)) satisfies both.
  • Double hashing has no clustering of either kind, reaches every slot, and measured the fewest probes per insert at every load factor.
  • Linear probing keeps the best cache locality, so it is often fastest in practice at moderate load despite the worst theory.
munotes.in382

Quadratic Probing and Double Hashing

Test yourself

1. Give quadratic probing's probe sequence and say what it fixes. h_i(k) = (h(k) + i squared) mod m, so offsets 0, 1, 4, 9, 16. The growing jumps mean a displaced record is not placed beside the record that displaced it, so primary clustering does not form.

2. What is quadratic probing's serious limitation? Give the measured example. It cannot reach every slot. With m = 16, a key hashing to slot 0 can reach only slots 0, 1, 4 and 9, so the fifth key hashing to 0 was refused while 12 of the 16 slots were free.

3. State the theorem that rescues it, and the two rules it implies. If m is prime, quadratic probing finds an empty slot whenever the table is less than half full. So m must be prime and the load factor must stay below 0.5.

4. Define secondary clustering. Two keys with the same hash value follow identical probe sequences, because the sequence depends only on h(k). Fewer keys are affected than under primary clustering, but those that are, are affected at every step.

5. Give double hashing's probe sequence and the two conditions on the second hash function. h_i(k) = (h1(k) + i x h2(k)) mod m. h2(k) must never be 0, or the probe sequence never advances, and h2(k) must be coprime with m, or the sequence cannot reach every slot.

6. Give a second hash function that satisfies both conditions, with m prime. h2(k) = 1 + (k mod (m - 1)). It lies between 1 and m-1, so it is never 0 and is always coprime with a prime m.

7. Why is linear probing often the fastest in practice despite the worst theory? Its probes walk consecutive array positions, which suits the memory cache, while double hashing jumps around the array and misses the cache.

8. Which scheme would you choose for a table that must run close to its capacity, and why? Double hashing: it has no primary or secondary clustering, it reaches every slot, and it measured the fewest probes per insertion at every load factor tested.

Contents This chapter on its own page

munotes.in383

Chapter One Hundred Seven

Load Factor and Rehashing

Syllabus topic Module 2, "hash table, hash functions"

In one line

The load factor is records divided by buckets; every hash table cost grows with it; so when it crosses a threshold the table is rebuilt at roughly twice the size, which costs O(n) once but O(1) per insertion on average.

The load factor

alpha = n / m = number of records / number of buckets

It is the single number that predicts a hash table's performance, and the previous chapters gave the formulas in terms of it:

SchemeSuccessful searchUnsuccessful searchalpha may exceed 1?
Chaining1 + alpha/2alphayes
Linear probing(1 + 1/(1-alpha))/2(1 + 1/(1-alpha) squared)/2no
Quadratic probingsimilar, worse than double hashingno, and must stay below 0.5
Double hashing(1/alpha) ln(1/(1-alpha))1/(1-alpha)no
import math

print("%8s %14s %16s %16s %18s"
      % ("alpha", "chaining miss", "linear miss", "double miss", "linear hit"))
for alpha in (0.25, 0.5, 0.75, 0.9, 0.95, 0.99):
    chaining = alpha
    linear = (1 + 1 / (1 - alpha) ** 2) / 2
    double = 1 / (1 - alpha)
    linear_hit = (1 + 1 / (1 - alpha)) / 2
    print("%8.2f %14.2f %16.1f %16.1f %18.1f"
          % (alpha, chaining, linear, double, linear_hit))

print()
print("chaining grows in a straight line. every open addressing scheme")
print("blows up, because the probe sequence has to find one of the few")
print("remaining free slots.")
print()
print("at alpha = 0.99 a failed linear probing search costs about %d probes."
      % ((1 + 1 / 0.01 ** 2) / 2))
print("the table is not full. it is 99 per cent full, and that is enough.")
   alpha  chaining miss      linear miss      double miss         linear hit
    0.25           0.25              1.4              1.3                1.2
    0.50           0.50              2.5              2.0                1.5
    0.75           0.75              8.5              4.0                2.5
    0.90           0.90             50.5             10.0                5.5
    0.95           0.95            200.5             20.0               10.5
    0.99           0.99           5000.5            100.0               50.5

chaining grows in a straight line. every open addressing scheme
blows up, because the probe sequence has to find one of the few
remaining free slots.

at alpha = 0.99 a failed linear probing search costs about 5000 probes.
the table is not full. it is 99 per cent full, and that is enough.

The thresholds actually used

ImplementationSchemeGrows when alpha reaches
Java HashMapchaining0.75
Python dictopen addressing, a probe sequence of its ownabout 0.66
C++ unordered_mapchaining1.0 by default
A textbook linear probing tablelinear probing0.5 to 0.7
A textbook quadratic probing tablequadratic probing0.5, by the theorem of chapter 106

The pattern: chaining tolerates about 1, open addressing about 0.6 to 0.7, and quadratic probing is capped at 0.5 by arithmetic rather than by taste.

Rehashing: why the records cannot simply be copied

When the threshold is crossed, a larger table is allocated and every record is hashed again and placed in the new table. This is rehashing, and the reason it is not a copy is simple: the slot was k mod m, and m has changed.

munotes.in384

Load Factor and Rehashing

OLD, NEW = 11, 23
KEYS = [25, 36, 47, 58, 69]

print("%8s %14s %14s %s" % ("key", "slot in %d" % OLD, "slot in %d" % NEW, "same?"))
for k in KEYS:
    a, b = k % OLD, k % NEW
    print("%8d %14d %14d %s" % (k, a, b, a == b))

print()
same = sum(1 for k in KEYS if k % OLD == k % NEW)
print("%d of %d keys keep their slot." % (same, len(KEYS)))
print("so the old array cannot be copied into the new one. every record")
print("must be hashed again with the NEW m and placed accordingly.")
print()
print("this is also why a hash table cannot be written to disk and reloaded")
print("into a table of a different size without rebuilding it.")
     key     slot in 11     slot in 23 same?
      25              3              2 False
      36              3             13 False
      47              3              1 False
      58              3             12 False
      69              3              0 False

0 of 5 keys keep their slot.
so the old array cannot be copied into the new one. every record
must be hashed again with the NEW m and placed accordingly.

this is also why a hash table cannot be written to disk and reloaded
into a table of a different size without rebuilding it.

Rehashing, implemented

def next_prime(n):
    def prime(x):
        if x < 2:
            return False
        d = 2
        while d * d <= x:
            if x % d == 0:
                return False
            d += 1
        return True

    while not prime(n):
        n += 1
    return n


class GrowingTable:
    """Chaining, with automatic growth at a threshold."""

    THRESHOLD = 0.75

    def __init__(self, buckets=7):
        self.table = [[] for _ in range(buckets)]
        self.count = 0
        self.rehashes = 0
        self.records_moved = 0

    def load_factor(self):
        return self.count / len(self.table)

    def insert(self, key, value):
        chain = self.table[key % len(self.table)]
        for i, (k, _v) in enumerate(chain):
            if k == key:
                chain[i] = (key, value)
                return
        chain.append((key, value))
        self.count += 1
        if self.load_factor() > self.THRESHOLD:
            self._rehash()

    def _rehash(self):
        old = self.table
        size = next_prime(2 * len(old) + 1)
        self.table = [[] for _ in range(size)]
        for chain in old:
            for key, value in chain:
                self.table[key % size].append((key, value))
                self.records_moved += 1
        self.rehashes += 1

    def search(self, key):
        for k, v in self.table[key % len(self.table)]:
            if k == key:
                return v
        return None


t = GrowingTable(buckets=7)
print("%8s %10s %10s %12s %12s"
      % ("inserted", "buckets", "alpha", "rehashes", "moved so far"))
for i in range(1, 201):
    t.insert(i * 13, "record %d" % i)
    if i in (5, 6, 10, 20, 50, 100, 200):
        print("%8d %10d %10.2f %12d %12d"
              % (i, len(t.table), t.load_factor(), t.rehashes, t.records_moved))

print()
print("200 records ended in a table of %d buckets at alpha %.2f."
      % (len(t.table), t.load_factor()))
print("it rehashed %d times and moved %d records in total."
      % (t.rehashes, t.records_moved))
print("that is %.2f moves per record inserted."
      % (t.records_moved / 200))
print()
print("every record is still findable after all that rehashing:",
      all(t.search(i * 13) == "record %d" % i for i in range(1, 201)))
munotes.in385

Load Factor and Rehashing

inserted    buckets      alpha     rehashes moved so far
       5          7       0.71            0            0
       6         17       0.35            1            6
      10         17       0.59            1            6
      20         37       0.54            2           19
      50         79       0.63            3           47
     100        163       0.61            4          107
     200        331       0.60            5          230

200 records ended in a table of 331 buckets at alpha 0.60.
it rehashed 5 times and moved 230 records in total.
that is 1.15 moves per record inserted.

every record is still findable after all that rehashing: True

Note the last line. A rehash rewrites the entire structure, so the test that matters is that every record is still retrievable afterwards, and it is.

Why the table doubles, measured

A rehash costs O(n). If it happened often, the O(1) promise would be worthless. The reason it does not is that the table doubles, so rehashes become exponentially rarer as the table grows.

def total_moves(n, grow):
    """Count every record move while inserting n records."""
    size, count, moved, rehashes = 7, 0, 0, 0
    for _ in range(n):
        count += 1
        if count / size > 0.75:
            size = grow(size)
            moved += count
            rehashes += 1
    return moved, rehashes


print("%10s %18s %12s %14s %18s %12s %14s"
      % ("records", "doubling moves", "rehashes", "moves/record",
         "add 10 moves", "rehashes", "moves/record"))
for n in (100, 1000, 10000, 100000):
    d_moved, d_rehash = total_moves(n, lambda s: 2 * s + 1)
    a_moved, a_rehash = total_moves(n, lambda s: s + 10)
    print("%10d %18d %12d %14.2f %18d %12d %14.2f"
          % (n, d_moved, d_rehash, d_moved / n,
             a_moved, a_rehash, a_moved / n))

print()
d = [total_moves(n, lambda s: 2 * s + 1)[0] / n
     for n in (100, 1000, 10000, 100000)]
a = [total_moves(n, lambda s: s + 10)[0] / n
     for n in (100, 1000, 10000, 100000)]
print("doubling: the moves per record stay BOUNDED, between %.1f and %.1f"
      % (min(d), max(d)))
print("across a thousandfold increase in n. they rise and fall with where")
print("the last rehash happened to land, and they do not grow. that bound")
print("is what AMORTISED O(1) means: one insert may cost O(n), but the")
print("average over all of them is a constant.")
print()
print("growing by a fixed 10: the moves per record go from %.1f to %.1f,"
      % (a[0], a[-1]))
print("a thousandfold rise, because the number of rehashes grows with n and")
print("each one copies everything. the total is O(n squared), which destroys")
print("the structure.")
print()
print("at 100,000 records, doubling moved %d and adding 10 moved %d."
      % (total_moves(100000, lambda s: 2 * s + 1)[0],
         total_moves(100000, lambda s: s + 10)[0]))
munotes.in386

Load Factor and Rehashing

   records     doubling moves     rehashes   moves/record       add 10 moves     rehashes   moves/record
       100                186            5           1.86                660           13           6.60
      1000               1530            8           1.53              66600          133          66.60
     10000              12282           11           1.23            6666000         1333         666.60
    100000             196602           15           1.97          666660000        13333        6666.60

doubling: the moves per record stay BOUNDED, between 1.2 and 2.0
across a thousandfold increase in n. they rise and fall with where
the last rehash happened to land, and they do not grow. that bound
is what AMORTISED O(1) means: one insert may cost O(n), but the
average over all of them is a constant.

growing by a fixed 10: the moves per record go from 6.6 to 6666.6,
a thousandfold rise, because the number of rehashes grows with n and
each one copies everything. the total is O(n squared), which destroys
the structure.

at 100,000 records, doubling moved 196602 and adding 10 moved 666660000.

That table is the argument, and it is the same argument as chapter 12's growable array: growth must be multiplicative, never additive. Doubling holds the moves per record inside a fixed band, 1.2 to 2.0 here, however large n gets; adding a fixed amount took it from 6.6 to 6,666.6 over the same range, which is the O(n squared) total showing itself.

Choosing the new size

Three rules, each with its reason:

At least double. Anything less makes the rehashes too frequent, as the measurement shows.

Take the next prime. Chapter 101 showed what a composite table size does to the division method, and chapter 106's theorem requires a prime for quadratic probing.

Never shrink at the same threshold that grows. If the table grows at alpha 0.75 and shrinks at alpha 0.75, a program that repeatedly inserts and deletes one record near the boundary rehashes on every operation.

def simulate(grow_at, shrink_at, operations):
    """One record added and removed repeatedly, right at the boundary."""
    size, count, rehashes = 101, 76, 0          # alpha just above 0.75
    adding = False
    for _ in range(operations):
        count += 1 if adding else -1
        adding = not adding
        if count / size > grow_at:
            size = 2 * size + 1
            rehashes += 1
        elif count / size < shrink_at:
            size = max(11, size // 2)
            rehashes += 1
    return rehashes


print("100 operations alternating one insert and one delete at the boundary:")
print("   grow at 0.75, shrink at 0.75 :", simulate(0.75, 0.75, 100), "rehashes")
print("   grow at 0.75, shrink at 0.25 :", simulate(0.75, 0.25, 100), "rehashes")
print()
print("the same work, and the first arrangement rebuilt the whole table")
print("on %d of the 100 operations. the gap between the two thresholds is"
      % simulate(0.75, 0.75, 100))
print("called HYSTERESIS, and without it a table at the boundary thrashes.")
munotes.in387

Load Factor and Rehashing

100 operations alternating one insert and one delete at the boundary:
   grow at 0.75, shrink at 0.75 : 100 rehashes
   grow at 0.75, shrink at 0.25 : 1 rehashes

the same work, and the first arrangement rebuilt the whole table
on 100 of the 100 operations. the gap between the two thresholds is
called HYSTERESIS, and without it a table at the boundary thrashes.

Rehashing clears tombstones

One more benefit, and it is the cure chapter 105 promised. A rehash reinserts only the live records, so every tombstone disappears.

class Tombstone:
    def __repr__(self):
        return "DELETED"


EMPTY = None
DELETED = Tombstone()


def rehash(slots, new_size):
    """Only live records are carried over. Tombstones are simply not copied."""
    fresh = [EMPTY] * new_size
    carried = 0
    for cell in slots:
        if cell is not EMPTY and cell is not DELETED:
            at = cell % new_size
            while fresh[at] is not EMPTY:
                at = (at + 1) % new_size
            fresh[at] = cell
            carried += 1
    return fresh, carried


M = 23
slots = [EMPTY] * M
for k in range(1, 18):
    at = (k * 7) % M
    while slots[at] is not EMPTY:
        at = (at + 1) % M
    slots[at] = k * 7
for k in range(1, 14):                           # delete 13 of the 17
    for i in range(M):
        at = ((k * 7) % M + i) % M
        if slots[at] is not DELETED and slots[at] == k * 7:
            slots[at] = DELETED
            break

print("before rehashing:")
print("   live records :", sum(1 for s in slots
                               if s is not EMPTY and s is not DELETED))
print("   tombstones   :", sum(1 for s in slots if s is DELETED))
print("   empty        :", sum(1 for s in slots if s is EMPTY))

fresh, carried = rehash(slots, M)
print()
print("after rehashing into a table of the same size %d:" % M)
print("   live records :", sum(1 for s in fresh
                               if s is not EMPTY and s is not DELETED))
print("   tombstones   :", sum(1 for s in fresh if s is DELETED))
print("   empty        :", sum(1 for s in fresh if s is EMPTY))
print()
print("records carried over:", carried)
print("tombstones remaining:", sum(1 for s in fresh if s is DELETED))
print("so a delete-heavy table is repaired by rehashing even without")
print("growing, which is why implementations rehash on a tombstone count")
print("as well as on a load factor.")
before rehashing:
   live records : 4
   tombstones   : 13
   empty        : 6

after rehashing into a table of the same size 23:
   live records : 4
   tombstones   : 0
   empty        : 19

records carried over: 4
tombstones remaining: 0
so a delete-heavy table is repaired by rehashing even without
growing, which is why implementations rehash on a tombstone count
as well as on a load factor.
munotes.in388

Load Factor and Rehashing

Quick revision

  • Load factor alpha = n/m. Every hash table cost in this module is a function of it.
  • Chaining's cost grows linearly in alpha and alpha may exceed 1; every open addressing cost blows up as alpha approaches 1.
  • Real thresholds: Java HashMap 0.75, Python dict about 0.66, C++ unordered_map 1.0, quadratic probing capped at 0.5 by arithmetic.
  • Rehashing is not a copy: the slot is k mod m and m has changed, so every record must be hashed again.
  • The new size must be at least double and should be the next prime.
  • Doubling keeps the record moves per insertion inside a fixed band, 1.2 to 2.0 in the measurement, however large n gets: that bound is what amortised O(1) means.
  • Growing by a fixed amount took the moves per record from 6.6 at 100 records to 6,666.6 at 100,000, because the total work is O(n squared).
  • Growth must be multiplicative, exactly as for chapter 12's growable array.
  • Growing and shrinking at the same threshold makes a table at the boundary rehash on nearly every operation; the gap between the thresholds is called hysteresis.
  • Rehashing also removes every tombstone, which is why open addressing tables rehash on a tombstone count as well as on a load factor.

Test yourself

1. Define the load factor and say why it is the central number. alpha = n/m, records divided by buckets. Every search, insert and delete cost for chaining and for open addressing is a formula in alpha, so it predicts performance by itself.

2. Why must a hash table be rebuilt rather than copied when it grows? Because a record's slot is its key modulo the table size, and the table size has changed, so almost no record belongs in its old position.

3. What is rehashing, and what does it cost? Allocating a larger table and reinserting every record by hashing it again. One rehash costs O(n).

4. Explain why rehashing does not break the O(1) promise. Because the table doubles, so rehashes become exponentially rarer. The total record moves over n insertions stay proportional to n: measured, the moves per insertion stayed between 1.2 and 2.0 from 100 records to 100,000. A bounded average per insertion is what amortised O(1) means.

5. What happens if the table grows by a fixed amount instead of doubling? The number of rehashes grows with n and each copies everything, so the total work is O(n squared) and the moves per record grow without limit.

munotes.in389

Load Factor and Rehashing

6. Give the two rules for choosing the new table size. At least double the old size, and take the next prime, since a composite size damages the division method and quadratic probing requires a prime.

7. Why must the shrink threshold differ from the grow threshold? Otherwise a program inserting and deleting one record at the boundary triggers a full rebuild on nearly every operation. The gap between the two thresholds is hysteresis.

8. What does rehashing do about tombstones? It removes them, because only live records are reinserted. This is why an open addressing table may rehash on a tombstone count even when its load factor is low.

Contents This chapter on its own page

munotes.in390

Chapter One Hundred Eight

What Hashing Buys and What It Gives Up

Syllabus topic Module 2, "Hash Table ADT, Advantages & Disadvantages"

In one line

Hashing buys O(1) average insert, search and delete, and it gives up every question about order, the worst case guarantee, and some memory.

What it buys

import math

print("finding one record among n, by method, in comparisons:")
print("(the list figure is its average, (n+1)/2; the other two are worst case)")
print("%12s %16s %16s %14s %12s"
      % ("records", "linked list", "sorted array", "AVL tree", "hash table"))
for n in (100, 10000, 1000000, 100000000):
    steps = math.ceil(math.log2(n))
    print("%12d %16.1f %16d %14d %12s"
          % (n, (n + 1) / 2, steps, steps, "about 2"))

print()
print("the hash table column does not change. that is the whole purchase:")
print("the cost of a lookup stops depending on how much is stored.")
print()
print("at a hundred million records a binary search needs %d comparisons"
      % math.ceil(math.log2(10 ** 8)))
print("and the hash table needs about 2, provided the load factor is kept")
print("low, which is chapter 107's job.")
finding one record among n, by method, in comparisons:
(the list figure is its average, (n+1)/2; the other two are worst case)
     records      linked list     sorted array       AVL tree   hash table
         100             50.5                7              7      about 2
       10000           5000.5               14             14      about 2
     1000000         500000.5               20             20      about 2
   100000000       50000000.5               27             27      about 2

the hash table column does not change. that is the whole purchase:
the cost of a lookup stops depending on how much is stored.

at a hundred million records a binary search needs 27 comparisons
and the hash table needs about 2, provided the load factor is kept
low, which is chapter 107's job.

Stated as a list, for an answer:

O(1) average insert, search and delete. Independent of the number of records.

The cost does not grow with the data. A table of a hundred million records looks up as fast as one of a hundred, which no comparison based structure can offer.

It is simple to implement. An array and a modulo. Compare chapter 72's four AVL rotations.

The keys need no order at all. A hash table works on any key that can be compared for equality and hashed. A binary search tree needs keys that can be ordered, which rules out, for example, a structure with no sensible ordering.

Insertion order is irrelevant to performance. A binary search tree degenerates to a list if the keys arrive sorted (chapter 68). A hash table does not care what order the keys arrive in.

What it gives up: every question about order

This is the decisive disadvantage, and it must be stated as a list of specific questions, not as "it is unordered".

import bisect
import time

N = 20000
KEYS = [(i * 7919 + 13) % 99991 for i in range(N)]

hash_table = {k: True for k in KEYS}             # a hash table
sorted_keys = sorted(KEYS)                       # what a balanced tree gives

print("%-38s %22s %22s"
      % ("the question", "hash table", "balanced tree"))
rows = [
    ("is key 12345 present?", "O(1), one probe", "O(log n), 15 steps"),
    ("the smallest key?", "O(m + n), scan ALL", "O(log n), leftmost"),
    ("the largest key?", "O(m + n), scan ALL", "O(log n), rightmost"),
    ("all keys between 100 and 200?", "O(m + n), scan ALL", "O(log n + output)"),
    ("the next key after 12345?", "O(m + n), scan ALL", "O(log n), successor"),
    ("every key in order?", "O(m + n) then SORT", "O(n), inorder walk"),
    ("the 500th smallest key?", "O(m + n) then SORT", "O(log n) with sizes"),
]
for question, h, t in rows:
    print("%-38s %22s %22s" % (question, h, t))

print()
print("now the smallest key, both ways, on %d records:" % N)
smallest_by_scan = min(hash_table)               # every key is examined
smallest_by_order = sorted_keys[0]               # the first element
print("   by scanning the hash table :", smallest_by_scan)
print("   from the ordered structure :", smallest_by_order)
print("   the same answer            :", smallest_by_scan == smallest_by_order)
print()
lo, hi = 100, 200
in_range_scan = sorted(k for k in hash_table if lo <= k <= hi)
left = bisect.bisect_left(sorted_keys, lo)
right = bisect.bisect_right(sorted_keys, hi)
in_range_order = sorted_keys[left:right]
print("keys between %d and %d: %d of them" % (lo, hi, len(in_range_order)))
print("   the hash table examined all %d keys to find them" % N)
print("   the ordered structure examined %d" % (len(in_range_order) + 2))
print("   same answer:", in_range_scan == in_range_order)
print()
print("that is a factor of %d on this query, and it gets worse with n,"
      % (N // max(len(in_range_order) + 2, 1)))
print("because the hash table's cost is n and the tree's is the output size.")
munotes.in391

What Hashing Buys and What It Gives Up

the question                                       hash table          balanced tree
is key 12345 present?                         O(1), one probe     O(log n), 15 steps
the smallest key?                          O(m + n), scan ALL     O(log n), leftmost
the largest key?                           O(m + n), scan ALL    O(log n), rightmost
all keys between 100 and 200?              O(m + n), scan ALL      O(log n + output)
the next key after 12345?                  O(m + n), scan ALL    O(log n), successor
every key in order?                        O(m + n) then SORT     O(n), inorder walk
the 500th smallest key?                    O(m + n) then SORT    O(log n) with sizes

now the smallest key, both ways, on 20000 records:
   by scanning the hash table : 1
   from the ordered structure : 1
   the same answer            : True

keys between 100 and 200: 20 of them
   the hash table examined all 20000 keys to find them
   the ordered structure examined 22
   same answer: True

that is a factor of 909 on this query, and it gets worse with n,
because the hash table's cost is n and the tree's is the output size.
munotes.in392

What Hashing Buys and What It Gives Up

Read the second column of that table. Six of the seven questions cost O(m + n) on a hash table, which means examining the entire structure. These are not slow operations; they are operations the structure cannot do, answered only by giving up and looking at everything.

A question that asks for a comparison with a binary search tree is asking for this list.

What else it gives up

The worst case is O(n), not O(1). Every figure above is an average. One bad hash function or one hostile key set collapses the table to a list, as chapter 104 printed.

Memory. A hash table must be kept well under full, so slots are deliberately empty; chaining adds a pointer per record on top.

M = 1009

print("space at various load factors, 1009 buckets:")
print("%10s %10s %16s %22s"
      % ("alpha", "records", "empty buckets", "slots per record"))
for alpha in (0.5, 0.66, 0.75, 1.0):
    n = int(alpha * M)
    print("%10.2f %10d %16d %22.2f"
          % (alpha, n, M - min(n, M), M / n))

print()
print("at Python's threshold of about 0.66, a third of the table is empty")
print("by design. a sorted array wastes nothing, and a balanced tree wastes")
print("two pointers per record instead.")
print()
print("so the comparison is not 'hashing wastes space'. it is: hashing")
print("wastes EMPTY SLOTS, a tree wastes POINTERS, and a sorted array")
print("wastes nothing but cannot insert in less than O(n).")
space at various load factors, 1009 buckets:
     alpha    records    empty buckets       slots per record
      0.50        504              505                   2.00
      0.66        665              344                   1.52
      0.75        756              253                   1.33
      1.00       1009                0                   1.00

at Python's threshold of about 0.66, a third of the table is empty
by design. a sorted array wastes nothing, and a balanced tree wastes
two pointers per record instead.

so the comparison is not 'hashing wastes space'. it is: hashing
wastes EMPTY SLOTS, a tree wastes POINTERS, and a sorted array
wastes nothing but cannot insert in less than O(n).

It depends entirely on the hash function. A balanced tree needs only that keys can be compared. A hash table needs a function matched to the actual keys, and chapter 102 showed what a mismatch costs.

No partial matching. A hash table can find "Bhavesh" and cannot find "every name beginning Bh", because a prefix has a different hash from the whole key. A sorted structure or a trie does prefix search naturally. This matters for any search box.

The iteration order is unstable. Chapter 100 showed the same keys coming out in different orders from tables of different sizes, and a rehash changes it again mid-program.

munotes.in393

What Hashing Buys and What It Gives Up

The head to head table

Hash tableAVL treeSorted array
SearchO(1) average, O(n) worstO(log n) guaranteedO(log n)
InsertO(1) averageO(log n)O(n)
DeleteO(1) averageO(log n)O(n)
Minimum or maximumO(m + n)O(log n)O(1)
Range queryO(m + n)O(log n + output)O(log n + output)
Sorted traversalO(m + n) then a sortO(n)O(n)
Successor or predecessorO(m + n)O(log n)O(log n)
Prefix matchingnot supportedO(log n + output)O(log n + output)
Worst case guaranteenoyesyes
Space overheadempty slots, plus linkstwo pointers per nodenone
Needs an ordering on keysnoyesyes
Needs a good hash functionyesnono

The decision rule

Use a hash table when every question is "give me the record with this exact key". That is the majority of lookups in real programs, which is why hash tables are everywhere.

Use a balanced tree when any question involves order: ranges, nearest, sorted output, minimum, maximum, successor. Giving up O(1) for O(log n) is a small price; discovering later that the structure cannot answer the question at all is not.

Use a sorted array when the data is built once and then only read. No insertion cost to pay, no pointers, binary search, and perfect cache behaviour.

Use both when both kinds of question are asked. Real systems commonly keep a hash table for exact lookups and an ordered index for ranges over the same records. The cost is the extra memory and keeping the two consistent.

Quick revision

  • Buys: O(1) average insert, search and delete, with a cost that does not grow with the number of records; a simple implementation; no ordering needed on the keys; and no sensitivity to insertion order.
  • Gives up every question about order. Minimum, maximum, range, successor, sorted traversal and rank all cost O(m + n), which means examining the whole structure.
  • Those are not slow operations; they are operations the structure cannot perform.
  • Gives up the worst case guarantee: O(n) if the hash function fails on the keys given.
  • Gives up memory: the table must be kept well under full, so at Python's threshold a third of it is empty by design, and chaining adds a pointer per record.
  • Gives up partial matching: a prefix has a different hash from the key, so no prefix search.
  • Gives up a stable iteration order, which changes with the table size and after every rehash.
  • Decision: hash table for exact key lookups, balanced tree for anything involving order, sorted array for data built once and then read, and both together when both kinds of question are asked.
munotes.in394

What Hashing Buys and What It Gives Up

Test yourself

1. State the central advantage of hashing in one sentence. Insert, search and delete are O(1) on average, so the cost of a lookup does not depend on how many records are stored.

2. Name four advantages besides speed. A simple implementation; no ordering required on the keys, only equality and a hash; immunity to the order in which keys arrive, unlike a binary search tree; and performance that does not degrade as the data grows, provided the load factor is maintained.

3. List the specific questions a hash table cannot answer efficiently. The smallest key, the largest key, all keys in a range, the next or previous key, all keys in sorted order, and the k-th smallest key. Each costs O(m + n), which means examining the entire table.

4. Why is "it is unordered" an inadequate answer? Because it does not say what is lost. The precise statement is that six distinct and common queries fall from O(log n) on a balanced tree to O(m + n) on a hash table, which is examining everything.

5. Compare the space overheads of a hash table, an AVL tree and a sorted array. A hash table keeps slots deliberately empty, about a third of the table at a load factor of 0.66, and chaining adds a pointer per record. An AVL tree stores two pointers and a balance factor per node. A sorted array has no overhead at all but cannot insert in less than O(n).

6. Why can a hash table not do prefix matching? Because a prefix hashes to a different slot from the full key, so there is no relationship between where "Bh" would go and where "Bhavesh" is.

7. Give the decision rule. A hash table when every query is for an exact key; a balanced tree when any query involves order; a sorted array for data built once and then only read; and both a hash table and an ordered index when both kinds of query occur.

8. What must be said alongside "search is O(1)"? That it is an average. The worst case is O(n), reached when the hash function fails to spread the keys actually given.

Contents This chapter on its own page

munotes.in395

Chapter One Hundred Nine

Where Hashing Is Used, and Where It Must Not Be

Syllabus topic Module 2, "Applications of hashing"

In one line

Hashing is used wherever a program looks something up by an exact key, which is most of the time, and it must not be used where order matters, where an attacker chooses the keys, or where the word hashing means a cryptographic hash rather than a hash table.

Where it is used

The dictionary, map or set of every programming language. Python's dict and set, Java's HashMap and HashSet, C++'s unordered_map, JavaScript's Map and plain objects, PHP's arrays. This is the single commonest data structure in working software, and it is a hash table underneath.

Symbol tables in compilers and interpreters. The original application. A compiler meeting the name total must find its declaration, type and scope, and it does that thousands of times per file. Keys are names, queries are exact, order is irrelevant: the perfect fit.

Caches. A cache maps a request to a stored answer, which is exactly a hash table. Memcached and Redis are this at network scale; a browser cache maps a URL to a file; and memoisation inside a program is the same idea applied to a function.

CALLS = {"plain": 0, "cached": 0}


def fib_plain(n):
    """No memory: the same subproblem is recomputed again and again."""
    CALLS["plain"] += 1
    if n < 2:
        return n
    return fib_plain(n - 1) + fib_plain(n - 2)


def fib_cached(n, seen):
    """One dictionary, keyed by n. Each subproblem is solved once."""
    CALLS["cached"] += 1
    if n < 2:
        return n
    if n in seen:                                 # a hash table lookup
        return seen[n]
    seen[n] = fib_cached(n - 1, seen) + fib_cached(n - 2, seen)
    return seen[n]


N = 28
plain = fib_plain(N)
cached = fib_cached(N, {})

print("the %dth Fibonacci number, computed twice:" % N)
print("   without a cache : %d  in %d calls" % (plain, CALLS["plain"]))
print("   with a cache    : %d  in %d calls" % (cached, CALLS["cached"]))
print("   the same answer :", plain == cached)
print()
print("the dictionary removed %d of the %d calls, which is %.4f%% of them."
      % (CALLS["plain"] - CALLS["cached"], CALLS["plain"],
         100 * (CALLS["plain"] - CALLS["cached"]) / CALLS["plain"]))
print("the recursion is unchanged. only the memory was added, and it turned")
print("an exponential computation into a linear one.")
print()
print("this is a hash table doing work no cleverness in the recursion could")
print("do: remembering an answer by its exact key.")
the 28th Fibonacci number, computed twice:
   without a cache : 317811  in 1028457 calls
   with a cache    : 317811  in 55 calls
   the same answer : True

the dictionary removed 1028402 of the 1028457 calls, which is 99.9947% of them.
the recursion is unchanged. only the memory was added, and it turned
an exponential computation into a linear one.

this is a hash table doing work no cleverness in the recursion could
do: remembering an answer by its exact key.
munotes.in396

Where Hashing Is Used, and Where It Must Not Be

De-duplication. Deciding whether a value has been seen before is a set membership test, which is one hash lookup. Without it, de-duplicating n items costs O(n squared) comparisons.

ITEMS = [(i * 37) % 500 for i in range(2000)]      # many repeats


def unique_by_scan(items):
    """No hash table: compare each item with everything kept so far."""
    kept, comparisons = [], 0
    for item in items:
        found = False
        for other in kept:
            comparisons += 1
            if other == item:
                found = True
                break
        if not found:
            kept.append(item)
    return kept, comparisons


def unique_by_set(items):
    """One hash lookup per item."""
    seen, kept, lookups = set(), [], 0
    for item in items:
        lookups += 1
        if item not in seen:
            seen.add(item)
            kept.append(item)
    return kept, lookups


a, comparisons = unique_by_scan(ITEMS)
b, lookups = unique_by_set(ITEMS)

print("%d items, %d distinct" % (len(ITEMS), len(b)))
print("   by scanning   : %7d comparisons" % comparisons)
print("   by a hash set : %7d lookups" % lookups)
print("   same answer   :", a == b)
print()
print("the scan did %.0f times the work, and the ratio grows with the"
      % (comparisons / lookups))
print("number of DISTINCT items, because that is the list it scans.")
2000 items, 500 distinct
   by scanning   :  500500 comparisons
   by a hash set :    2000 lookups
   same answer   : True

the scan did 250 times the work, and the ratio grows with the
number of DISTINCT items, because that is the list it scans.

Database indexing, with a restriction. A hash index answers "the row where id = 5531" in one probe. It cannot answer "rows where id is between 5000 and 6000" at all, which is why relational databases default to a B-tree index and offer a hash index only as an option for equality lookups. That restriction is chapter 108's lost ordering, in production.

The hash join. To join two tables, a database builds a hash table on the smaller one keyed by the join column, then scans the larger one probing it. O(n + m) instead of the O(n x m) of comparing every pair.

Counting and grouping. Word frequencies, votes per candidate, sales per branch: a key to a running total is a hash table, and it turns an O(n x k) problem into O(n).

Routers, switches and operating systems. A routing table keyed by address prefix, a process table keyed by process id, a file system's directory lookup, a page table: all hash tables, all hot paths.

Where the word means something else

Password storage, file integrity and Git object names use a cryptographic hash, which is not a hash table. The word is shared and the purpose is opposite, and this is worth stating in an answer because the confusion is common.

munotes.in397

Where Hashing Is Used, and Where It Must Not Be

A hash table's hash functionA cryptographic hash
Examplek mod 1009SHA-256, bcrypt, Argon2
Outputa slot number, 0 to m-1a long fixed-length digest
Wantedas fast as possibledeliberately slow, for passwords
Collisionsexpected, and handledmust be computationally infeasible
Reversibletrivially, and nobody caresmust be infeasible to reverse
Purposefind a recordprove a value, or hide one

So "passwords are hashed" and "a hash table hashes keys" use one word for two unrelated jobs. A password must never be stored using a hash table's hash function: it is fast, which is precisely the wrong property, and it is not salted.

The same applies to checksums and content addressing. Git names every object by the hash of its contents; a download is verified against a published digest. Those are cryptographic hashes used for identity, not for addressing a bucket.

Bloom filters and consistent hashing also use hash functions without being hash tables. A Bloom filter answers "possibly present" or "definitely absent" in a few bits per item. Consistent hashing decides which server in a cluster owns a key, so that adding a server moves as few keys as possible. Both are beyond this syllabus and both are worth naming.

Where a hash table must not be used

1. When any query involves order. Ranges, nearest, sorted output, minimum, maximum, successor. This is chapter 108's list and it is the commonest reason to choose something else.

2. When an attacker chooses the keys. If a program puts user-supplied values into a hash table with a fixed, published hash function, a caller who knows that function can supply values that all land in one bucket. Every lookup then costs O(n) and the program stops responding, although nothing has crashed.

M = 1009


def table_cost(keys, m):
    """Total probes to look up every key once, with chaining."""
    table = [[] for _ in range(m)]
    for k in keys:
        table[k % m].append(k)
    return sum(table[k % m].index(k) + 1 for k in keys), max(len(c) for c in table)


def mix(i):
    x = ((i + 1) * 2654435761) % (2 ** 32)
    x ^= x >> 13
    x = (x * 2246822519) % (2 ** 32)
    x ^= x >> 17
    return x


N = 2000
ordinary = [mix(i) for i in range(N)]
all_one_bucket = [i * M for i in range(N)]      # every key is 0 mod M

for name, keys in (("ordinary keys", ordinary),
                   ("keys all in one bucket", all_one_bucket)):
    total, worst = table_cost(keys, M)
    print("%-26s %8d probes for %d lookups, longest bucket %d"
          % (name, total, N, worst))

ordinary_total = table_cost(ordinary, M)[0]
hostile_total = table_cost(all_one_bucket, M)[0]
print()
print("the hostile set cost %.0f times as much work for the same number of"
      % (hostile_total / ordinary_total))
print("lookups, and the table is the same size and the code identical.")
print()
print("THE DEFENCE, which is what matters: the hash function must not be")
print("predictable from outside. a RANDOM value chosen at program start is")
print("mixed into every hash, so the attacker cannot compute which keys")
print("collide. python, ruby, perl and the JVM all do this now; python's")
print("PYTHONHASHSEED exists for exactly this reason.")
print()
print("the second defence is a bound: cap how many entries one request may")
print("insert, and convert an over-long chain to a balanced tree, which is")
print("what java does at eight entries. then the worst case is O(log n).")
munotes.in398

Where Hashing Is Used, and Where It Must Not Be

ordinary keys                  3998 probes for 2000 lookups, longest bucket 7
keys all in one bucket      2001000 probes for 2000 lookups, longest bucket 2000

the hostile set cost 501 times as much work for the same number of
lookups, and the table is the same size and the code identical.

THE DEFENCE, which is what matters: the hash function must not be
predictable from outside. a RANDOM value chosen at program start is
mixed into every hash, so the attacker cannot compute which keys
collide. python, ruby, perl and the JVM all do this now; python's
PYTHONHASHSEED exists for exactly this reason.

the second defence is a bound: cap how many entries one request may
insert, and convert an over-long chain to a balanced tree, which is
what java does at eight entries. then the worst case is O(log n).

Note what the defence is: not a better hash function, but an unpredictable one, plus a bound on the damage. A published function, however well it spreads ordinary keys, can be analysed.

3. For storing passwords. Use a deliberately slow, salted password hash. This is the distinction above, and it is a real mistake with real consequences.

4. When the data set is tiny. For ten records a linear scan over an array is faster than hashing, because the hash computation and the cache miss cost more than ten comparisons. Library implementations of small maps often use a plain array for this reason.

5. When a stable iteration order is needed. The order changes with the table size and after every rehash, so anything that must produce output in a fixed order needs an ordered structure, or a hash table plus an explicit list.

Quick revision

  • Used for: language dictionaries and sets, compiler symbol tables, caches and memoisation, de-duplication, equality-only database indexes, hash joins, counting and grouping, and operating system and router tables.
  • Memoisation turned an exponential Fibonacci recursion linear with one dictionary.
  • De-duplication by hash set did a fraction of the comparisons a scan needed, and the gap grows with the number of distinct items.
  • A database uses a B-tree index by default because a hash index cannot answer a range query at all.
  • A cryptographic hash is a different thing with the same word: a hash table's hash must be as fast as possible, a password hash must be deliberately slow and salted.
  • Checksums, Git object names, Bloom filters and consistent hashing all use hash functions without being hash tables.
  • Must not be used: when a query involves order; when an attacker chooses the keys; for passwords; for tiny data sets; or when a stable iteration order is required.
  • The defence against hostile keys is an UNPREDICTABLE hash, randomised per program run, plus a bound such as converting a long chain to a balanced tree.
munotes.in399

Where Hashing Is Used, and Where It Must Not Be

Test yourself

1. Name six applications of hashing. Language dictionaries and sets; compiler symbol tables; caches and memoisation; de-duplication and set membership; database hash indexes and hash joins; and counting or grouping by key. Operating system tables such as process and page tables are a seventh.

2. Why do relational databases default to a B-tree index rather than a hash index? Because a hash index can answer only equality lookups. A range query such as "id between 5000 and 6000" cannot be answered by a hash index at all, while a B-tree answers it in O(log n + output).

3. What is a hash join and what does it save? A database builds a hash table on the smaller table keyed by the join column, then scans the larger one probing it. The cost is O(n + m) instead of the O(n times m) of comparing every pair.

4. Distinguish a hash table's hash function from a cryptographic hash. A hash table's hash returns a slot number and should be as fast as possible; collisions are expected and handled. A cryptographic hash returns a long digest, must be infeasible to reverse, must make collisions infeasible, and for passwords should be deliberately slow and salted. They share only the word.

5. Explain the hostile-key problem and its defence. If the hash function is fixed and published, a caller who supplies the keys can choose values that all hash to one bucket, making every lookup O(n) and stalling the program. The defence is an unpredictable hash, with a random value chosen at program start mixed into every hash, plus a bound on the damage such as converting a chain of eight to a balanced tree.

6. Why is a hash table a poor choice for ten records? Because computing the hash and the cache miss to reach the bucket cost more than simply comparing ten items in an array.

munotes.in400

Where Hashing Is Used, and Where It Must Not Be

7. Give the four situations in which a hash table should not be used. When any query involves order; when an attacker chooses the keys and the hash is predictable; for storing passwords; and when the data set is tiny. A fifth is when a stable iteration order is required.

Contents This chapter on its own page

munotes.in401

Chapter One Hundred Ten

Choosing the Right Structure: The Whole Paper on One Page

Syllabus topic Module 1 and Module 2, the whole paper

In one line

Pick the structure from the operations the program actually performs, in the proportions it performs them, and justify the choice by naming the operation that decides it.

The three questions, from chapter 6

  1. What operations does the program perform, and how often? Not what the data is. What is done to it.
  2. What does each operation cost in this structure?
  3. What does the structure cost to hold? Memory, and the cost of keeping it correct.

The second and third are now answerable from tables. The first is the one a student has to read out of the question, and it is where the marks are.

The whole paper in one table

StructureAccess by positionSearch by keyInsertDeleteOrdered?Chapter
ArrayO(1)O(n)O(n)O(n)if kept sorted11
Sorted arrayO(1)O(log n)O(n)O(n)yes11, 22
Singly linked listO(n)O(n)O(1) at headO(1) given the nodeno13 to 21
Doubly linked listO(n)O(n)O(1)O(1) given the nodeno26 to 29
Stacktop onlyO(n)O(1) pushO(1) popno30 to 34
Queueends onlyO(n)O(1) enqueueO(1) dequeueno40 to 46
Dequeboth endsO(n)O(1) either endO(1) either endno47
Binary search treeO(log n) by rankO(log n) average, O(n) worstO(log n) averageO(log n) averageyes64 to 68
AVL treeO(log n) by rankO(log n) guaranteedO(log n)O(log n)yes71 to 74
Binary heaptop onlyO(n)O(log n)O(log n) at toppartially81 to 85
Hash tablenoO(1) average, O(n) worstO(1) averageO(1) averageno99 to 109
Graph, adjacency listby vertexO(degree)O(1) vertexO(V + E) vertexno91
Graph, adjacency matrixby vertex pairO(1) edgeO(1) edgeO(V squared) vertexno90

The decision, as a sequence of questions

Work down the list and stop at the first that matches. This is the order to use in an answer, because each question rules out a whole family.

1. Is the data a relationship between things rather than a collection of things? Then it is a graph. A road network, a social network, dependencies, a web of links: the structure is given by the problem, and the only remaining choice is matrix or list, decided by density (chapter 92).

2. Is the data hierarchical, with one parent per item? Then it is a tree. A file system, an organisation chart, a parse tree, a decision tree.

3. Does the order of processing matter in a fixed way?

  • Most recent first, and backtracking, and undo: a stack.
  • First come first served, and level by level: a queue.
  • Most urgent first: a priority queue, implemented as a heap.
  • Both ends: a deque.
munotes.in402

Choosing the Right Structure: The Whole Paper on One Page

4. Is every query for an exact key, with no question about order? Then a hash table. This is the commonest answer in real software.

5. Does any query involve order? Ranges, nearest, sorted output, minimum, maximum, successor, k-th smallest. Then a balanced tree, which means an AVL tree in this paper.

6. Is the data built once and then only read? Then a sorted array with binary search. No insertion cost to pay and the best memory behaviour of anything here.

7. Is the position in the sequence what identifies an item, and is the size known? Then an array.

8. Are insertions and deletions frequent and in the middle, with no position-based access? Then a linked list.

Twelve questions and their answers, with the reason

The reason is the mark. Each row below names the operation that decides it.

The requirementThe structureBecause
Undo in an editorstackthe last action is undone first, which is LIFO exactly
Matching brackets in a compilerstacka closing bracket must match the most recent unmatched opening one
Printer job queuequeuejobs are served in arrival order, which is FIFO exactly
Breadth first searchqueuevertices must be visited in order of distance
Hospital emergency triagepriority queue, heapthe most urgent is served first, not the earliest arrival
The k largest of a million valuesheap of size kO(n log k) and only k items held, against O(n log n) to sort everything
A compiler's symbol tablehash tablelookups are by exact name, thousands of times, and order is irrelevant
Counting word frequencieshash tablea key to a running total, with no order needed
A phone book searched by prefixsorted structure or triea hash table cannot match a prefix (chapter 109)
Student records, searched by roll number and listed in orderAVL tree, or a hash table plus an ordered indexthe second requirement rules a hash table out by itself
A fixed timetable, loaded once, read all termsorted arrayno insertions to pay for, binary search, no pointers
Shortest route on a city mapgraph with Dijkstraweighted edges and a cheapest-path query

The three traps in this question

Naming a structure without the operation that decides it. "A hash table, because it is fast" earns little. "A hash table, because every query is a lookup by exact student number and nothing in the requirement asks for order" earns the mark.

Missing the requirement that rules out the obvious answer. A question that says "and produce a list in roll number order" has ruled out a hash table in that clause. These clauses are deliberate.

munotes.in403

Choosing the Right Structure: The Whole Paper on One Page

Giving an average cost where the question asks for a guarantee. If the requirement says "must respond within a fixed time", a hash table's O(1) average and a binary search tree's O(log n) average are both wrong answers; an AVL tree's guaranteed O(log n) is right.

Two structures together is a real answer

import bisect


class StudentIndex:
    """Exact lookup by a hash table; ordered queries by a sorted list."""

    def __init__(self):
        self.by_roll = {}          # hash table: O(1) exact lookup
        self.order = []            # sorted list: O(log n + output) ranges

    def add(self, roll, name):
        if roll in self.by_roll:
            self.by_roll[roll] = name
            return
        self.by_roll[roll] = name
        bisect.insort(self.order, roll)

    def find(self, roll):
        """One hash lookup."""
        return self.by_roll.get(roll, "not enrolled")

    def between(self, low, high):
        """A binary search, then a slice: the hash table cannot do this."""
        left = bisect.bisect_left(self.order, low)
        right = bisect.bisect_right(self.order, high)
        return [(r, self.by_roll[r]) for r in self.order[left:right]]

    def smallest(self):
        return self.order[0] if self.order else None

    def in_order(self):
        return [(r, self.by_roll[r]) for r in self.order]


index = StudentIndex()
for roll, name in ((4512, "Aarti"), (4203, "Bhavesh"), (4890, "Chetna"),
                   (4377, "Devdatta"), (4601, "Esha"), (4109, "Farhan")):
    index.add(roll, name)

print("exact lookup, from the hash table:")
print("   4601 ->", index.find(4601))
print("   9999 ->", index.find(9999))
print()
print("ordered queries, from the sorted list:")
print("   smallest roll :", index.smallest())
print("   4300 to 4650  :", index.between(4300, 4650))
print()
print("everyone, in roll order:")
for roll, name in index.in_order():
    print("   %d  %s" % (roll, name))
print()
print("the hash table answers the exact lookups and cannot do the ranges.")
print("the sorted list answers the ranges and would need O(log n) for an")
print("exact lookup. together they answer both, and the price is stated:")
print("every roll number is stored twice, and an insertion costs O(n) to")
print("keep the sorted list correct.")
print()
print("naming that price is part of the answer. a question that asks for")
print("BOTH kinds of query is asking whether you will admit the cost.")
exact lookup, from the hash table:
   4601 -> Esha
   9999 -> not enrolled

ordered queries, from the sorted list:
   smallest roll : 4109
   4300 to 4650  : [(4377, 'Devdatta'), (4512, 'Aarti'), (4601, 'Esha')]

everyone, in roll order:
   4109  Farhan
   4203  Bhavesh
   4377  Devdatta
   4512  Aarti
   4601  Esha
   4890  Chetna

the hash table answers the exact lookups and cannot do the ranges.
the sorted list answers the ranges and would need O(log n) for an
exact lookup. together they answer both, and the price is stated:
every roll number is stored twice, and an insertion costs O(n) to
keep the sorted list correct.

naming that price is part of the answer. a question that asks for
BOTH kinds of query is asking whether you will admit the cost.
munotes.in404

Choosing the Right Structure: The Whole Paper on One Page

Quick revision

  • Choose from the operations the program performs and their proportions, never from what the data is called.
  • Chapter 6's three questions: which operations and how often; what each costs here; what the structure costs to hold.
  • A relationship between things is a graph; a hierarchy with one parent each is a tree.
  • Fixed processing order: stack for last in first out, queue for first in first out, heap for most urgent first, deque for both ends.
  • Exact key lookups only: hash table. Any ordered query: balanced tree. Built once then read: sorted array.
  • Position identifies the item and the size is known: array. Frequent middle insertions with no indexing: linked list.
  • The marks are in the reason, so name the operation that decides the choice.
  • A requirement mentioning order rules out a hash table in that clause, and it is put there deliberately.
  • A requirement asking for a guaranteed time rules out every average-case structure, leaving an AVL tree.
  • Using two structures together is a legitimate answer provided its price, duplicated keys and a costlier insertion, is stated.

Test yourself

1. What should a choice of data structure be based on? The operations the program performs and how often it performs each, not the kind of data being stored.

2. A program must repeatedly find a student by roll number, and nothing else. Which structure, and why? A hash table. Every query is a lookup by an exact key and nothing requires order, so O(1) average is available and nothing is given up.

3. The same program must also print all students with roll numbers between two bounds. What changes? A hash table is ruled out, because a range query costs O(m + n) on it. An AVL tree answers both in O(log n), or a hash table plus an ordered index answers both at the cost of storing each key twice and a costlier insertion.

4. A system must guarantee a response within a fixed time. Which structures are ruled out? Every structure whose bound is an average: a hash table, with O(n) worst case, and a plain binary search tree, which degenerates to O(n). An AVL tree's O(log n) is guaranteed.

5. Find the k largest of a million values. Which structure and what does it save? A heap of size k, giving O(n log k) time and O(k) space, against O(n log n) time and O(n) space to sort everything.

6. Why is a stack the answer for matching brackets, and a queue for breadth first search? A closing bracket must match the most recent unmatched opening bracket, which is last in first out. Breadth first search must visit vertices in order of distance, which requires taking them in the order they were found, that is first in first out.

munotes.in405

Choosing the Right Structure: The Whole Paper on One Page

7. What makes an answer to this kind of question earn full marks? Naming the single operation in the requirement that decides the choice, and saying what the chosen structure gives up. The name of the structure alone is the smaller part of the mark.

8. When is a sorted array the best choice? When the data is built once and then only read: there is no insertion cost to pay, binary search gives O(log n), and it has no pointer overhead and the best cache behaviour of any structure in this paper.

Contents This chapter on its own page

munotes.in406

Chapter One Hundred Eleven

Module 2 in One Sitting

Syllabus topic Module 2, the whole module

In one line

Module 2 is four structures that give up the straight line: trees for hierarchy, heaps for priority, graphs for relationships, and hash tables for lookup without searching.

The one table to know

StructureSearchInsertDeleteOrdered?The one weakness
Binary search treeO(log n) averageO(log n) averageO(log n) averageyesO(n) if the keys arrive sorted
AVL treeO(log n) guaranteedO(log n)O(log n)yesrotations on every change
Binary heapO(n)O(log n)O(log n) at the top onlypartlyonly the top is reachable
Hash tableO(1) averageO(1) averageO(1) averagenono ordered query at all
Graph, adjacency listO(degree) per vertexO(1) vertexO(V + E) vertexnoO(degree) edge lookup
Graph, adjacency matrixO(1) per edgeO(1) edgeO(V squared) vertexnoV squared memory always

Trees

A tree is a hierarchy: one root, every other node with exactly one parent, no cycles. Height, depth and level are defined in chapter 52 and are the commonest source of lost marks, because an off-by-one in the definition changes every later answer.

A binary tree has at most two children per node. A binary tree of height h has at most 2 to the power (h+1) minus 1 nodes, and n nodes need a height of at least log2(n+1) minus 1.

Traversals. Inorder (left, node, right), preorder (node, left, right), postorder (left, right, node), and level order, which needs a queue. Inorder on a binary search tree gives the keys in sorted order, which is the fact most often tested. Two traversals rebuild a tree only if one of them is inorder.

A binary search tree keeps every key in the left subtree smaller and every key in the right subtree larger. Search, insert and delete are O(h). Deletion has three cases: a leaf is removed; a node with one child is replaced by it; a node with two children is replaced by its inorder successor.

It degenerates. Keys inserted in sorted order give a tree of height n, and every operation becomes O(n). That is the whole reason balance exists.

An AVL tree keeps every node's balance factor in minus 1, 0 or plus 1, restoring it with four rotations: LL, RR, LR and RL. Insertion needs at most one rotation or double rotation; deletion may need rotations all the way up to the root. The height is kept at O(log n), so every operation is O(log n) guaranteed, which no other structure in this paper offers.

Huffman coding builds a tree by repeatedly joining the two least frequent symbols, which is a priority queue in use. The code is prefix-free because every symbol is a leaf, so no code is a prefix of another and the stream decodes without separators.

munotes.in407

Module 2 in One Sitting

Priority queues and heaps

A priority queue serves the most urgent item, not the earliest. The ADT is insert, remove_best, peek_best.

A binary heap is a complete binary tree with the heap order property: in a min-heap every parent is less than or equal to its children. Complete means it fills level by level, left to right, which is why it fits in an array with no pointers: for index i the children are 2i+1 and 2i+2 and the parent is (i-1)/2.

Sift up restores order after an insertion at the end; sift down restores it after the top is removed and the last element is moved there. Both are O(log n) because the tree's height is.

Building a heap from n items is O(n), not O(n log n), because most nodes are near the bottom and sift down from a low node is cheap. This is a standard question and the reason is the marks.

A heap gives O(1) peek at the best, O(log n) insert and O(log n) removal of the best, and O(n) to find anything else, because only the top is ordered relative to everything.

Graphs

A graph is vertices joined by edges, with no restriction: cycles and multiple paths are allowed, which is exactly what a tree forbids. Directed or undirected, weighted or unweighted.

The vocabulary that earns marks: degree, in-degree and out-degree, path, simple path, cycle, connected, complete, sparse and dense. The sum of all degrees is twice the number of edges. A complete undirected graph has n(n-1)/2 edges.

Two representations. An adjacency matrix is V by V, gives O(1) edge lookup, and costs V squared memory whatever the graph; at 10,000 vertices and 20,000 edges it is 99.96 per cent zeros. An adjacency list keeps a list of neighbours per vertex, costs V + 2E, and gives O(degree) neighbour access. A full traversal is exactly V squared probes on a matrix and exactly 2E on a list, which is why every algorithm is written against a list.

BFS uses a queue and visits in order of increasing distance, so a vertex's level is its shortest distance in edges. DFS uses a stack, usually the recursion stack, and goes deep before wide; a DFS depth is not a distance. Both are O(V + E) on a list and O(V squared) on a matrix. The visited set is what makes either one terminate, not an optimisation.

Connectivity. One traversal from any vertex reaching all V means connected. Starting a fresh traversal from each unvisited vertex finds every component, still in O(V + E) in total. An isolated vertex is a component. A connected graph needs at least V-1 edges; exactly V-1 and connected means a tree.

munotes.in408

Module 2 in One Sitting

Shortest paths. Unweighted: BFS already solves it, and reaching for Dijkstra is the commonest over-answer. Weighted and non-negative: Dijkstra, settling the nearest unsettled vertex and relaxing its edges, O((V + E) log V) with a heap or O(V squared) with an array scan. A negative edge breaks it, because settling means final: the chapter's example settled a vertex at 1 when the truth was 0 and reported no error. Bellman-Ford handles negative weights.

Hashing

Hashing computes a record's address from its key instead of searching for it, so a lookup does not depend on how much is stored. The hash function maps any key into 0 to m-1; the array is the hash table and a slot is a bucket.

Four methods. Division, h(k) = k mod m, fastest, and m must be a prime, never a power of 10 or 2. Mid-square, the middle digits of k squared, because they depend on every digit of the key. Folding, the sum of the key's pieces, for long keys. Multiplication, floor(m x fractional part of k x A) with A about 0.6180339887, which works for any m.

A good hash function is uniform over the actual keys, deterministic, cheap, uses the whole key, and sends similar keys to unrelated slots. It is good or bad only relative to a key set.

Collisions are certain, not unlucky. Pigeonhole: more keys than buckets forces one. Birthday: a collision becomes more likely than not at about 1.25 times the square root of m, so a table of a million buckets starts colliding at about a thousand records, 0.1 per cent full.

Three strategies. Chaining gives each bucket a list: alpha may exceed 1, deletion is a plain unlink, costs are alpha for a miss and 1 + alpha/2 for a hit. Linear probing puts everything in the array: no pointers, best cache behaviour, but primary clustering, and a cleared slot breaks the probe chain, so deletion needs a tombstone. Quadratic probing removes primary clustering but reaches only (m+1)/2 slots with m prime, so m must be prime and alpha below 0.5. Double hashing makes the step depend on the key and has no clustering of either kind.

The load factor alpha = n/m decides everything. Chaining's cost grows linearly in it; every open addressing cost blows up as it approaches 1. So the table is rehashed at a threshold: a larger table, the next prime at least double, and every record hashed again, because the slot was k mod m and m has changed. Doubling makes this amortised O(1); growing by a fixed amount makes the total O(n squared).

munotes.in409

Module 2 in One Sitting

What hashing gives up is every question about order: minimum, maximum, range, successor, sorted output and rank all cost O(m + n), which means examining everything. That is the sentence that separates a good answer from a list.

The six "advantages and disadvantages", in one line each

Binary search tree. Ordered and simple; degenerates to O(n) on sorted input.

AVL tree. The only guaranteed O(log n) here, and ordered; rotations on every insertion and deletion, and deletion may rotate to the root.

Heap. O(1) at the best item and O(n) to build; nothing but the top is reachable in less than O(n).

Adjacency matrix. O(1) edge lookup and a trivially simple structure; V squared memory always, and every traversal becomes O(V squared).

Adjacency list. Memory proportional to the edges and O(V + E) traversal; edge lookup is O(degree).

Hash table. O(1) average for insert, search and delete, with no ordering needed on the keys; no ordered query at all, an O(n) worst case, deliberately empty space, and total dependence on the hash function.

What Module 2 cannot do

No structure here answers a prefix query. A trie does; a hash table cannot and a tree needs the keys ordered by that prefix.

No structure here gives O(1) worst case search. Hashing is O(1) average, AVL is O(log n) guaranteed, and the comparison lower bound of chapter 99 says no comparison based method can beat log2(n).

No structure here gives both O(1) exact lookup and O(log n) range queries. Two structures together do, at the price of duplicated keys, which chapter 110 prices.

The five mark answers most likely to be asked

The three BST deletion cases, with a worked example of the two-child case replacing by the inorder successor.

The four AVL rotations, with the insertion that triggers each one and the resulting tree.

Why building a heap is O(n), with the argument about most nodes being near the bottom.

Adjacency matrix against adjacency list, with the memory figures and the O(V + E) against O(V squared) traversal.

BFS and DFS, both algorithms, their complexities, and which solves the unweighted shortest path and why.

Dijkstra's algorithm traced on a small weighted graph, including the relax step changing a vertex's distance, and why a negative edge breaks it.

Collision handling, with chaining and linear probing both built, the load factor formulas, and the tombstone.

Hash functions, any two methods with the arithmetic shown, and why m should be prime.

Test yourself

1. What single property distinguishes a graph from a tree? A tree has no cycle and exactly one path between any two vertices. A graph allows cycles and multiple paths, so a tree is a connected acyclic graph.

munotes.in410

Module 2 in One Sitting

2. Which traversal of a binary search tree gives sorted order, and which pair rebuilds a tree? Inorder gives sorted order. Rebuilding needs two traversals of which one is inorder: inorder with preorder, or inorder with postorder.

3. Why does a binary heap need no pointers? Because it is a complete binary tree, which fills level by level left to right, so the nodes map onto array positions with children at 2i+1 and 2i+2 and the parent at (i-1)/2.

4. State the cost of a full graph traversal on each representation and explain the difference. Exactly V squared probes on a matrix, because every row is scanned in full, and exactly 2E on a list, because only real edges are read. So BFS and DFS are O(V squared) and O(V + E) respectively.

5. Why is BFS, not Dijkstra, the answer for an unweighted shortest path? Because with equal weights cheapest means fewest edges, which BFS already gives by its level numbers, at O(V + E) against Dijkstra's O((V + E) log V).

6. Why does a negative edge break Dijkstra's algorithm? Because settling the nearest unsettled vertex is justified only by no edge being able to reduce a total. With a negative edge a settled vertex may still be improvable, and a settled vertex is never revisited.

7. Why are collisions certain rather than unlucky? Give both arguments. Pigeonhole: more keys than buckets forces a shared bucket, whatever the hash function. Birthday: a collision is more likely than not at about 1.25 times the square root of m, so a million bucket table collides at about a thousand records.

8. What goes wrong if a deleted slot in a linear probing table is simply cleared? The probe chain is broken, so records placed beyond the cleared slot become unreachable: a search stops at the now empty slot and reports absent although the record is in the table. A tombstone is required.

9. Give the load factor formulas for chaining. Unsuccessful search alpha, successful search 1 + alpha/2, and insertion one probe without a duplicate check.

10. What is the single most important thing hashing gives up? Every question about order. The minimum, maximum, a range, the successor, sorted output and the k-th smallest each cost O(m + n), which is examining the whole table, so they are not slow operations but operations the structure cannot perform.

Contents This chapter on its own page

munotes.in411

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!