munotes®

Principles of Operating Systems Notes | B.Sc. (Computer Science) Semester 3 | Mumbai University | munotes

Get access to whole semester resourcesSemester Pass

Official Notes munotes.in

Principles of Operating Systems

B.SC. (COMPUTER SCIENCE) · SEMESTER 3

Strictly as per the University of Mumbai NEP syllabus in force for B.Sc. (Computer Science)

For B.Sc. (Computer Science) students of the University of Mumbai and all its affiliated colleges

Open the book ↓

munotes.in Second Year

Principles of Operating Systems

Copyright © 2026 munotes.in. All rights reserved.

Written and first published by munotes.in, 2026.

This book is free for individual students to read at munotes.in. No part of it may be reproduced, distributed, stored, translated or used for institutional or classroom purposes in any form without a prior written licence from munotes.in.

Licensing and permissions: contact@munotes.in

The text of statutes and of judgments reproduced in this book is in the public domain under section 52(1)(q) of the Copyright Act 1957. The commentary, arrangement, examples and questions are the original work of munotes.in.

munotes.in is an independent study resource for MU students. It is not affiliated with, endorsed by, or officially connected to the University of Mumbai. Course names and university references describe the students and syllabus the material relates to.

munotes.in

Contents

Module I Fundamentals of Operating Systems, Processes and Threads, Process Synchronization, and CPU Scheduling

  1. What an Operating System Is 1
  2. The Two Jobs: Handing Out the Machine, and Hiding It 5
  3. Interrupts, Traps and the Two Modes 9
  4. The Timer, and Why the Operating System Always Gets the Processor Back 14
  5. The Functions of an Operating System 19
  6. Where Operating Systems Run 23
  7. The Services an Operating System Offers 27
  8. The Command Line and the Desktop 31
  9. What a System Call Is 35
  10. Watching System Calls Happen 39
  11. The Six Families of System Call 43
  12. How an Operating System Is Built 46
  13. Microkernels, Modules and What Linux Actually Is 49
  14. What a Process Is 53
  15. The Five States of a Process 57
  16. The Process Control Block 61
  17. The Queues, the Schedulers and the Context Switch 65
  18. Creating a Process: fork 69
  19. Running a Different Program: exec 73
  20. Waiting, Exiting, the Zombie and the Orphan 77
  21. Inter-process Communication: The Two Models 81
  22. Shared Memory, in Code That Runs 85
  23. Pipes 89
  24. Message Queues 93
  25. Blocking and Non-blocking Communication 97
  26. What a Thread Is 101
  27. Making Threads, and Waiting for Them 105
  28. Multicore Programming, and the Limit on It 109
  29. The Three Multithreading Models 113
  30. Thread Pools, and Handing Work Out 117
  31. Measuring It: Sequential Against Threaded 121
  32. The Shape of Every Concurrent Program 126
  33. A Race Condition, Made to Happen 130
  34. The Critical Section Problem, and the Three Conditions 135
  35. Peterson's Solution 139
  36. Hardware Help: Test and Set, and Compare and Swap 143
  37. The Mutex Lock 147
  38. The Semaphore 151
  39. The Bounded Buffer, Solved 155
  40. The Readers and the Writers 160
  41. The Dining Philosophers 164
  42. Monitors and Condition Variables 169
  43. Why Scheduling Exists: The Burst Cycle 174
  44. The Dispatcher, Preemption, and What a Switch Costs 178
  45. The Five Criteria, and the Arithmetic of Each 182
  46. Reading and Drawing a Gantt Chart 185
  47. First Come First Served 188
  48. Shortest Job First 191
  49. Shortest Remaining Time First 194
  50. Priority Scheduling, Starvation and Ageing 197
  51. Round Robin, and Choosing the Quantum 200
  52. Multilevel Queue Scheduling 203
  53. Multilevel Feedback Queue Scheduling 206
  54. All Seven Algorithms on One Problem 210
  55. Thread Scheduling, and What Linux Actually Does 214

Module II Deadlocks, Memory Management, Virtual Memory and Mass Storage, and the File System

  1. What a Deadlock Is, and the System Model 218
  2. The Four Conditions 222
  3. The Resource Allocation Graph 226
  4. The Four Ways to Handle a Deadlock 230
  5. Deadlock Prevention 233
  6. Safe States, and Avoidance 237
  7. The Banker's Algorithm 241
  8. The Resource Request Algorithm 245
  9. Deadlock Detection 249
  10. Recovery from Deadlock 254
  11. Why Memory Needs Managing 258
  12. Binding an Address: Compile, Load, Run 262
  13. The Memory Management Unit 266
  14. Swapping 270
  15. Contiguous Allocation, and the Holes It Leaves 274
  16. First Fit, Best Fit and Worst Fit 278
  17. Fragmentation, Internal and External 282
  18. Segmentation 286
  19. Paging 290
  20. Splitting an Address, and Translating One 295
  21. The TLB, and the Effective Access Time 299
  22. Protection and Sharing in a Paged System 304
  23. The Page Table Is Too Big: Hierarchical Paging 309
  24. Hashed and Inverted Page Tables 314
  25. Virtual Memory: Running What Will Not Fit 318
  26. Demand Paging, and the Page Fault 323
  27. The Effective Access Time Under Demand Paging 328
  28. Copy on Write 332
  29. Page Replacement: The Problem 337
  30. FIFO Replacement, and Belady's Anomaly 341
  31. Optimal Replacement 345
  32. LRU Replacement 348
  33. Second Chance, and the Counting Algorithms 352
  34. Comparing the Algorithms, and the Hit Ratio 357
  35. Allocation of Frames 361
  36. Thrashing, and the Working Set 365
  37. What a Disk Is, and What It Costs 370
  38. Disk Structure, and the Logical Block 374
  39. Disk Scheduling: FCFS and SSTF 378
  40. SCAN, C-SCAN, LOOK and C-LOOK 382
  41. Random Scheduling, and All Six Compared 386
  42. Disk Management 389
  43. What a File Is 394
  44. Opening a File, and What the Kernel Keeps 398
  45. Access Methods 402
  46. Directories, and the Shapes They Take 406
  47. Mounting 412
  48. File Sharing, and Locking 416
  49. The Layers a Read Passes Through 421
  50. On the Disk: Superblock, Inode, Data Blocks 425
  51. Directory Implementation 429
  52. Contiguous and Linked Allocation 433
  53. Indexed Allocation, and What a Real File System Does 438
  54. Free Space Management 443
  55. Designing a Small File System 448
munotes.in

Module I

Fundamentals of Operating Systems, Processes and Threads, Process Synchronization, and CPU Scheduling

munotes.in

Chapter One

What an Operating System Is

Syllabus topic Module 1, "Fundamentals of Operating systems - Definition of Operating System"

In one line

An operating system is the program that starts first, stays running, and stands between every other program and the machine.

In the words you can write in an examination: an operating system is the software that manages the computer's hardware, provides a platform on which application programs can run, and acts as an intermediary between the user and the hardware.

Both of those say the same thing. The first is what it means. The second is what to write.

Why there is one at all

Imagine a computer with no operating system on it. It has a processor, some memory, a disk, a keyboard and a screen. Now try to write a program on it that adds two numbers and prints the answer.

You cannot print the answer, because printing means putting characters into the right place in the screen's memory, and the position depends on which screen is fitted. You cannot read the two numbers, because reading means noticing that a key has been pressed, working out which key, and waiting while the person types. You cannot store the answer, because storing means asking the disk controller for a particular block, on a particular surface, and then waiting for the disk to turn. And if a second program were running, nothing would stop it from writing over your numbers, because memory is just memory.

Every one of those jobs has two things in common. They are the same for every program, and they depend on the hardware fitted rather than on the problem being solved. Writing them once and letting every program share them is the whole idea of an operating system. The program that adds two numbers gets to be a program about adding two numbers.

A useful test: if a job is needed by every program and nobody wants to write it twice, it belongs in the operating system.

The definition, in three parts

A definition of anything this large has to say what it does, what it hides, and who it serves. An operating system:

  1. Manages the hardware. It decides which program gets the processor, which gets memory,

which gets the disk next, and it hands each one out and takes it back.

  1. Provides a platform. It offers every program the same set of services, asked for in the

same way, whatever hardware is underneath. A program that reads a file does not know whether the file is on a hard disk, a solid state drive, a memory card or another computer.

  1. Acts as an intermediary. The person at the keyboard and the machine do not speak to each

other directly. Everything one says to the other passes through the operating system.

munotes.in1

What an Operating System Is

Two words that are easy to swap. A resource is anything a program needs and cannot have all of: the processor, memory, disk space, a printer. A service is something the operating system will do for a program if it asks: open a file, start another program, send a message. Resources are handed out. Services are performed. The next chapter is about the first and Chapter seven is about the second.

What it is, on a real machine

Four things are commonly called "the operating system" in ordinary speech, and only one of them is the operating system in the sense this paper means.

NameWhat it really is
The kernelthe operating system proper: the part that is always in memory and always in control
The shellan ordinary program that reads your commands and asks the kernel to carry them out
The utilitiesordinary programs, ls, cp, grep, shipped alongside
The distributionthe kernel, the shell, the utilities and thousands of other programs, packaged together and given a name such as Ubuntu

The lab machine can show that these are genuinely different things, because they have different version numbers and come from different projects.

$ uname -s
Linux
$ uname -r
6.8.0-117-generic
$ ls --version | head -1
ls (GNU coreutils) 9.4
$ bash --version | head -1
GNU bash, version 5.2.21(1)-release (aarch64-unknown-linux-gnu)
$ grep PRETTY_NAME /etc/os-release
PRETTY_NAME="Ubuntu 24.04.5 LTS"

Four different version numbers for four different things. uname -r is the kernel, which is the operating system. ls belongs to GNU coreutils, a separate project, and is a program that the kernel runs like any other. bash is another such program. Ubuntu 24.04 is the distribution, which is the box all of it came in.

That the shell is an ordinary program is worth proving, because it is the single most common misunderstanding at this stage.

$ which bash
/usr/bin/bash
$ ls -l /usr/bin/bash
-rwxr-xr-x 1 root root 1543048 Sep  1 14:30 /usr/bin/bash

It is a file, of a size, with an owner and permissions, sitting in a directory. The kernel is not a file you can list like that while it is running: it was loaded before any file system existed to list it from.

Worked example: what happens when you type a command

Anita sits at the machine and types ls and presses Enter. Eight things happen, and every one of them is a topic later in this book.

  1. The keyboard raises an interrupt. The processor stops what it was doing and the kernel's

keyboard handler runs (Chapter three).

  1. The kernel puts the characters where the shell can read them, and wakes the shell up, because

the shell had asked to be woken when a line arrived (Chapter seventeen).

munotes.in2

What an Operating System Is

  1. The shell reads the line and works out that ls is not one of its own built in commands, so

it must be a program on the disk.

  1. The shell asks the kernel to make a copy of itself, with fork (Chapter nineteen).
  2. The copy asks the kernel to replace it with the program /usr/bin/ls, with exec (Chapter

twenty).

  1. The kernel finds the file on the disk, gives the new program some memory, loads it, and puts

it in the queue of programs waiting for the processor (Chapters fifteen and seventeen).

  1. ls asks the kernel to read the directory, and asks the kernel to write its output. It never

touches the disk or the screen itself (Chapter nine).

  1. ls finishes. The kernel tells the shell, which prints a new prompt (Chapter twenty one).

Notice what Anita did and what the operating system did. Anita named a program. The operating system found it, loaded it, ran it, read a directory for it, printed for it, and cleaned up after it. That is the intermediary in the definition.

Distinctions that carry marks

Operating systemApplication program
Startswhen the machine is switched onwhen somebody or something starts it
Runs inkernel mode, with full control of the hardwareuser mode, with no direct access to hardware
Purposeto run other programsto do a job for a person
Knows aboutthe hardware fittedthe problem being solved
ExampleLinux, Windows, Android, iOSa browser, a compiler, ls
KernelOperating system, loosely used
Meaningthe part always resident and always in controlthe kernel plus the programs shipped with it
Sizeone programthousands of programs
In this paperthis is what "operating system" meansthis is what a shop means by it

What it does not mean

It is not the part you can see. The desktop, the windows and the mouse pointer are programs running on the operating system, not the operating system. A machine with no screen attached still has one, fully working.

It is not a program you run. Every other program is started by something. The operating system is started by the hardware, which is a topic of its own: a small program in the machine's own memory loads a bootstrap program from the disk, which loads the kernel.

It is not only for big machines. A washing machine, a car's engine controller and a smart watch all run one. What changes is how much of it there is, not whether there is one.

It is not optional in a machine that runs more than one program. Two programs sharing one processor and one memory need somebody to decide who gets what, and that somebody must be a program with more power than either of them.

munotes.in3

What an Operating System Is

Quick revision

  • An operating system manages the hardware, provides a platform for programs, and acts as the

intermediary between the user and the machine.

  • It is the program that starts first and stays in control.
  • A resource is handed out (processor, memory, disk); a service is performed (open a

file, start a program).

  • The kernel is the operating system proper. The shell, the utilities and the

distribution are not.

  • The kernel runs in kernel mode; every other program runs in user mode.
  • On the lab machine the kernel is Linux 6.8, the shell is bash 5.2, ls is GNU coreutils 9.4,

and the distribution is Ubuntu 24.04. Four version numbers, four different things.

  • A job belongs in the operating system when every program needs it and it depends on the

hardware rather than on the problem.

Test yourself

  1. Define an operating system in one sentence. Software that manages the computer's hardware,

provides a platform on which application programs run, and acts as an intermediary between the user and the hardware.

  1. Why can a program not print to the screen by itself? Because printing means writing to the

particular hardware fitted, which differs from machine to machine, and because two programs writing at once would interfere. The operating system does it once, for everybody.

  1. Is the shell part of the operating system? No. It is an ordinary program, in a file, with

permissions and a version of its own. It asks the kernel to do things, exactly as any other program does.

  1. Give one example each of a resource and a service. Resource: the processor, or memory.

Service: opening a file, or starting another program.

  1. What is the difference between Linux and Ubuntu? Linux is the kernel, which is the

operating system. Ubuntu is a distribution: the kernel plus a shell, utilities and thousands of programs, packaged and named. 6. A microwave oven has a timer, a display and three buttons. Does it have an operating system? It may or may not, and the test is whether more than one thing has to be managed at once. If a single program can poll the buttons and drive the display in a loop, none is needed. As soon as there are several activities to be scheduled, something must schedule them, and that something is an operating system however small.

Contents This chapter on its own page

munotes.in4

Chapter Two

The Two Jobs: Handing Out the Machine, and Hiding It

Syllabus topic Module 1, "Fundamentals of Operating systems - Operating System's role"

In one line

The operating system has two jobs: it is a resource allocator, which hands out the machine fairly, and it is a control program, which stops programs doing damage.

In the words to write: the operating system acts as a resource allocator, deciding which program gets which resource and for how long, and as a control program, controlling the execution of programs to prevent errors and improper use of the computer.

Why the role has to be split in two

The two jobs pull in opposite directions, and that is why they are named separately.

A resource allocator wants to say yes. Its measure of success is that the machine is busy: the processor is working, memory is full of useful things, the disk is not idle. A machine that is mostly idle has been wasted, and the person who paid for it has been let down.

A control program wants to say no. Its measure of success is that nothing went wrong: no program read another program's memory, no program held the processor for ever, no program wrote to a file it had no right to. A machine that crashes once a day is useless however busy it was.

Every design decision in this paper is a trade between those two. When you reach round robin scheduling in Chapter fifty two, the quantum is exactly this trade: a long quantum wastes less time switching, which serves the allocator, and a short quantum stops one program monopolising the machine, which serves the control program. When you reach the banker's algorithm in Chapter sixty two, you will watch the operating system refuse a request for resources that are sitting there free, because granting it might lead to deadlock. That refusal is the control program overruling the allocator, on purpose.

Job one: the resource allocator

A resource is anything a program needs and cannot have all of. The four that matter in this paper are the processor, main memory, the disk, and the devices.

ResourceThe question the operating system answersWhere in this book
Processorwhich program runs next, and for how longChapters forty three to fifty five
Main memorywhich program sits where, and how much it getsChapters sixty six to seventy nine
Diskwhich request the head serves next, and where a file's blocks goChapters ninety two to one hundred nine
Deviceswho gets the printer, and in what orderthroughout

For each one, the operating system must do four things: know what it has, know who is using it, decide who gets it next, and take it back when they are finished. That is the whole of resource allocation, and a request for a resource that is never taken back is a resource leak.

munotes.in5

The Two Jobs: Handing Out the Machine, and Hiding It

The measures the allocator is judged by are worth meeting now, because they come back as examinable definitions in Chapter forty six.

  • Utilisation: the fraction of the time a resource is busy. Higher is better.
  • Throughput: how much work is finished per unit of time. Higher is better.
  • Fairness: no program is starved while others are served.

Job two: the control program

The control program's work is easiest to see by asking what a program could do if nobody stopped it. On a machine with no control at all, any program could:

  1. read or write any part of memory, including another program's data and the operating system's

own tables;

  1. read or write any file, including other people's;
  2. talk to any device directly, and confuse it;
  3. never give the processor up, so that nothing else ever ran again;
  4. execute an instruction that halts the machine.

Every one of those is prevented, and the mechanism is the same in each case: the hardware refuses, and the refusal becomes the operating system's business. That mechanism is the next chapter. The five protections it buys are:

DangerWhat stops it
Reading another program's memoryaddress translation in the memory management unit (Chapter sixty eight)
Reading a file you may not readthe permission bits, checked by the kernel on every open (Chapter ninety nine)
Talking to a device directlydevice instructions are privileged and trap in user mode (Chapter three)
Never giving the processor upthe timer interrupt (Chapter four)
Halting the machinethe halt instruction is privileged too (Chapter three)

Two of those five can be watched on the lab machine straight away. A program that tries to read memory it does not own is stopped by the hardware, and the operating system then kills it.

#include <stdio.h>

int main(void)
{
    int *not_mine = (int *)1;          /* an address this program does not own */
    printf("about to read address 1\n");
    printf("it holds %d\n", *not_mine);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o trespass trespass.c
$ sh -c ./trespass
about to read address 1
Segmentation fault (core dumped)
$ echo $?
139
$ kill -l 11
SEGV

The first line printed. The second never did. The program was not allowed to finish, and the message is the operating system telling you why: the program referred to a segment of memory it had no right to, so it was stopped. A signal is a short message the kernel sends to a process, and a process killed by signal number n reports the status 128 + n, so 139 means signal

  1. The last command asks the shell what signal 11 is called: SEGV, for segmentation violation.
munotes.in6

The Two Jobs: Handing Out the Machine, and Hiding It

This is the control program working, not a bug in the program. A machine with no protection would have printed whatever happened to be at address 1, which might have been part of the operating system.

The file permission check is just as visible.

$ echo secret > private.txt
$ chmod 000 private.txt
$ ls -l private.txt
---------- 1 student student 7 Sep 30  2026 private.txt
$ cat private.txt
cat: private.txt: Permission denied
$ echo $?
1

cat is a perfectly ordinary program and it asked politely. The kernel said no. Notice that the refusal came from the kernel and not from cat: cat only reported it.

Worked example: two students, one machine

A college laboratory machine has one processor, four gigabytes of memory, and two students logged in. Ravi is compiling a large program. Sneha is editing a file.

The allocator gives Ravi's compiler the processor in long stretches, because it has a great deal of computing to do and nothing to wait for. When Sneha presses a key, the allocator takes the processor away from the compiler within milliseconds and gives it to the editor, because the editor has almost nothing to do but must do it now. The compiler gets most of the processor and the editor gets it promptly. Both measures are served: the machine is busy, and neither student is kept waiting.

The control program makes sure that the compiler, which uses a great deal of memory, cannot read the file Sneha is editing unless she has given it permission; that a bug in the compiler cannot crash the editor; and that if the compiler goes into an infinite loop, the timer still fires and Sneha still gets her keystrokes serviced.

Now change one thing. Suppose the allocator were removed and each program simply ran until it finished. Sneha would wait for the whole compile before her keystroke appeared. Suppose instead the control program were removed. The compiler's bug would take the machine down and both students would lose their work. Neither job is optional.

Distinctions that carry marks

Resource allocatorControl program
Wants tosay yessay no
Measured byutilisation, throughput, fairnesscorrectness, protection, no crash
Fails byleaving the machine idleletting one program harm another
Example decisiongive the processor to P2 nextrefuse the write, the file is read only
Example chapterCPU schedulingmemory protection
A program asking for a resourceA program asking for a service
It wantssomething to hold for a whilesomething done now
Examplefour megabytes of memorywrite these bytes to a file
Given back?yes, it must be releasedno, the work is simply finished
The operating system's worrywho else wants itmay this program have this done
munotes.in7

The Two Jobs: Handing Out the Machine, and Hiding It

What it does not mean

The operating system does not do the program's work. It does not compile, edit or calculate. It decides who computes, and it performs the jobs that need the hardware.

"Resource allocator" does not mean sharing everything equally. Fair does not mean equal. A program that needs the processor for a hundredth of a second and one that needs it for an hour are not treated the same, and should not be.

The control program is not an antivirus. It is not looking for malicious programs. It enforces rules on every program, good or bad, and it enforces them with the hardware's help rather than by inspection.

Quick revision

  • The role of an operating system is two jobs: resource allocator and control program.
  • The allocator hands out the processor, memory, the disk and the devices, and takes them back.
  • It is judged by utilisation, throughput and fairness.
  • The control program prevents errors and improper use, and it does it with the hardware's help:

address translation, permission bits, privileged instructions, and the timer.

  • The two jobs conflict on purpose. The scheduling quantum and the banker's algorithm are both

the conflict being settled.

  • A program that reads memory it does not own is killed with a segmentation fault, exit status

139, which is 128 plus signal 11.

  • A program denied a file is denied by the kernel, not by itself.

Test yourself

  1. State the operating system's role in one sentence. It acts as a resource allocator,

deciding which program gets which resource and for how long, and as a control program, controlling the execution of programs to prevent errors and improper use of the computer.

  1. Name the four resources this paper is about. The processor, main memory, the disk, and the

devices.

  1. Give one decision where the two jobs conflict, and say which wins. The banker's algorithm

refuses a request for free resources because granting it could lead to deadlock. The control program wins, deliberately, at a cost in utilisation.

  1. Why is a segmentation fault evidence of good design rather than a failure? Because the

hardware noticed a program touching memory it did not own and the operating system stopped it. Without that, the program would have read or damaged something belonging to somebody else.

  1. A program runs in an infinite loop. Which of the two jobs deals with it, and how? The

control program, using the timer interrupt: the processor is taken back whether the program co-operates or not.

  1. Is "fair" the same as "equal"? No. Fairness means nobody is starved. Two programs with

completely different needs are correctly given different amounts.

Contents This chapter on its own page

munotes.in8

Chapter Three

Interrupts, Traps and the Two Modes

Syllabus topic Module 1, "Fundamentals of Operating systems - Operating-System Operations"

In one line

The operating system does not run all the time. It runs when something happens, and the hardware is what tells it something has happened.

The precise form: a modern operating system is interrupt driven. The processor executes a user program until an interrupt or a trap transfers control to the operating system, which services it and returns.

Why a machine has to work this way

Suppose the kernel wanted to know whether a key had been pressed. Without help from the hardware it would have to keep asking, forever, in a loop. That is called polling, and it wastes the whole machine: the processor is busy asking about a keyboard instead of computing.

The alternative is to let the hardware interrupt. The keyboard is given a wire to the processor. When a key is pressed, the keyboard pulls the wire, the processor stops in the middle of whatever it was doing, and the kernel's keyboard routine runs. When that routine finishes, the interrupted program continues from exactly where it stopped and never knows it was paused.

So the kernel is not a program that runs alongside yours. It is a collection of routines that run when something calls for them, and then get out of the way.

The vocabulary, which is examined

Three words are used for three different things and students reliably swap them.

WordCaused byExampleIs it expected?
Interrupthardware, outside the processora key pressed, a disk finishing, the timeryes, but not when
Trapsoftware, on purposea program asking for a service with a system callyes, exactly then
Exceptionsoftware, by accidentdividing by zero, touching memory you do not ownno

A trap and an exception are both software interrupts: the processor raises them itself, rather than being poked from outside. The difference is intent. A trap is a program asking for something. An exception is a program doing something wrong.

Some books call all three "interrupts" and then say "hardware interrupt" and "software interrupt". That is the same idea with different labels. If an examination question says "software interrupt", it means a trap or an exception.

What actually happens, step by step

  1. The device raises the interrupt on its wire, or the processor raises the trap itself.
  2. The processor finishes the instruction it is on. It does not stop halfway.
  3. The processor saves enough state to get back: at least the address of the next instruction and

the flags.

  1. The processor switches to kernel mode (the next section).
  2. The processor reads the interrupt number and looks it up in the interrupt vector, a

table in memory that holds the address of the routine for each number.

munotes.in9

Interrupts, Traps and the Two Modes

  1. It jumps to that routine, the interrupt service routine or handler.
  2. The handler does the work, quickly.
  3. The handler executes a return from interrupt instruction. The saved state comes back, the mode

bit goes back to user mode, and the interrupted program continues.

The interrupt vector is why the table is indexed rather than searched. There may be a hundred devices and an interrupt must be dispatched in nanoseconds, so the number is an index into an array, never a key to look up in a list.

The kernel keeps a running count of every interrupt it has taken, and of every context switch it has made, and publishes both in one file.

$ awk '/^intr/{print $1, $2} /^ctxt/{print $1, $2}' /proc/stat
intr 75282970
ctxt 74892941

intr is the total number of interrupts since the machine started and ctxt the total number of context switches. Those two numbers are the machine's whole history of being interrupted and of changing hands, and they never stop climbing.

$ a=$(awk '/^intr/{print $2}' /proc/stat)
$ sleep 1
$ b=$(awk '/^intr/{print $2}' /proc/stat)
$ echo "$((b - a)) interrupts in one second, with nothing happening"
4559 interrupts in one second, with nothing happening

Nobody touched the keyboard during that second. The machine was interrupted thousands of times anyway, and most of those were the timer of the next chapter.

The two modes, which are what make protection possible

The processor has a mode bit. It has two values.

User modeKernel mode
Mode bit10
Also calleduser spacesupervisor, system, monitor or privileged mode
What runs thereevery ordinary programthe operating system
Privileged instructionsrefusedallowed
Memory it can touchonly its ownall of it

A privileged instruction is one the processor will only execute in kernel mode. The set differs between machines, but it always includes:

  • switching the mode bit itself (or the machine could simply promote itself);
  • loading the register that controls address translation;
  • talking to a device directly;
  • turning interrupts off;
  • halting the machine.

The first one on that list is the whole design. If a user program could set the mode bit, protection would be a suggestion. It cannot, so the only way into kernel mode is an interrupt, a trap or an exception, and every one of those lands on a handler the operating system wrote.

This answers a question beginners ask: if the operating system is just software, what stops my program from doing whatever the operating system does? The mode bit stops it, and the mode bit is hardware.

Seeing an exception happen, and learning who decides

Chapter two showed one exception already: a program read an address it did not own, the processor refused, and the kernel killed it with signal 11. Here is the other example every textbook gives, dividing by zero, and it does something worth the whole chapter.

munotes.in10

Interrupts, Traps and the Two Modes

#include <stdio.h>
#include <stdlib.h>

int main(int argc, char **argv)
{
    (void)argc;
    int top = atoi(argv[1]);
    int bottom = atoi(argv[2]);
    printf("dividing %d by %d\n", top, bottom);
    printf("the answer is %d\n", top / bottom);
    return 0;
}

The two numbers are read from the command line so that the compiler cannot work the answer out in advance. The division really happens while the program runs.

$ gcc -std=c17 -Wall -Wextra -o divzero divzero.c
$ sh -c './divzero 10 0'
dividing 10 by 0
the answer is 0
$ echo $?
0
$ uname -m
aarch64

No exception was raised. The program divided ten by zero, was told the answer is zero, and exited normally. That is not a mistake in the program and it is not a mistake in the book: it is this processor's rule. The machine is an ARM processor, and its integer divide instruction is defined to produce zero when the divisor is zero rather than to raise anything.

So an exception is not a property of the operation. It is a property of the hardware. Dividing by zero is meaningless arithmetic in either case; whether the processor notices is a design decision taken by whoever built the processor. On a machine whose divide instruction does raise, the same program dies and the kernel reports signal 8:

$ kill -l 8
FPE

SIGFPE exists, the kernel will deliver it, and nothing on this machine caused it. A student whose own laptop has a different processor may well see this program die, and both results are correct for the machine they were run on.

Write this carefully in an examination. "Division by zero raises an exception" is the answer a paper expects and is true of the processors the textbooks describe. It is more accurate to say that an arithmetic error raises an exception where the processor defines one, and to give invalid memory access as the example that holds everywhere.

Compare the two cases. The invalid address was refused by hardware on this machine and on every machine, because memory protection is the point of the memory management unit. The division was not, because this processor chose to define an answer instead. Neither program asked the kernel for anything: where the kernel arrived, it arrived because the hardware called it.

Seeing a trap happen

A trap is the deliberate version, and it is what Chapter nine is about. A single system call is enough to show one. strace asks the kernel to report every trap a program makes.

munotes.in11

Interrupts, Traps and the Two Modes

$ strace -e trace=write -o trace.txt /bin/echo hello
hello
$ cat trace.txt
write(1, "hello\n", 6)                  = 6
+++ exited with 0 +++

echo printed its word by trapping into the kernel once, with the system call named write. The program did not touch the screen. It asked.

Distinctions that carry marks

InterruptTrap
Raised bya device, outside the processorthe running program itself
Timingunpredictableexactly where the instruction is
Purposeto report that something happenedto ask the operating system for something
Examplethe disk has finished readingopen a file
Also calledhardware interruptsoftware interrupt
TrapException
Intentdeliberateaccidental
Raised bya special instruction the program executedthe processor noticing an illegal operation
Usual outcomethe service is performed and the program continuesthe program is killed
Examplea system calldivide by zero, invalid address
PollingInterrupt driven
The processorkeeps asking the deviceis told by the device
Cost when nothing happensthe whole processornothing
Latencyup to one polling intervalalmost none
Used forvery fast devices where the wait is shorter than an interrupteverything else

What it does not mean

An interrupt does not stop the machine halfway through an instruction. The current instruction finishes first. This matters: it is why a handler can safely assume the machine is in a consistent state.

The operating system is not "in the background". Between interrupts it is not running at all. Nothing of it is executing while your program computes.

Kernel mode is not a user account. Being the administrator, or root, is a matter of file permissions and is checked by software. Kernel mode is a bit in the processor. A root program still runs in user mode.

Turning interrupts off is not a way for a program to get more speed. It is privileged, and for good reason: a program that turned them off and looped would own the machine for ever.

Quick revision

  • A modern operating system is interrupt driven: it runs when an interrupt, a trap or an

exception transfers control to it, and returns.

  • Interrupt: from hardware, unexpected timing. Trap: from software, on purpose, a system

call. Exception: from software, by accident, an error.

  • The steps: finish the instruction, save state, switch to kernel mode, look the number up in the

interrupt vector, run the handler, return from interrupt.

  • The interrupt vector is an array indexed by interrupt number, so dispatch is one memory

reference and not a search.

  • The processor has a mode bit: user mode and kernel mode. Privileged instructions run

only in kernel mode.

  • Setting the mode bit is itself privileged, so the only way into the kernel is through a handler
munotes.in12

Interrupts, Traps and the Two Modes

the operating system wrote.

  • An invalid memory access gives signal 11, SEGV, status 139, on every machine.
  • Whether an arithmetic error raises an exception is decided by the processor. On the ARM

processor these programs were run on, integer division by zero produces zero and raises nothing; the signal for an arithmetic exception, where one is raised, is 8, FPE, status 136.

  • Polling wastes the processor when nothing is happening; interrupts cost nothing until

something does.

Test yourself

  1. What does "interrupt driven" mean? That the operating system is not continuously running:

control passes to it when an interrupt, trap or exception occurs, and returns to the program afterwards.

  1. Distinguish an interrupt from a trap. An interrupt comes from hardware outside the

processor at an unpredictable moment; a trap is raised by the running program on purpose, to ask the operating system for a service.

  1. What is the interrupt vector and why is it a table? A table holding the address of the

handler for each interrupt number. It is indexed by the number so that dispatch takes one memory reference, which matters when interrupts arrive thousands of times a second.

  1. Name three privileged instructions. Setting the mode bit, loading the address translation

register, and halting the machine. Turning interrupts off and talking to a device directly are also privileged.

  1. Why must setting the mode bit be privileged? Otherwise any program could put itself into

kernel mode and all protection would be voluntary.

  1. A program prints one line and then dies with status 136. What happened? It was killed by

signal 8, an arithmetic exception. The processor raised the exception and the kernel killed the program. 7. Your friend's laptop kills a program that divides by zero and yours prints zero. Who is right? Both. Raising an exception on integer division by zero is a decision taken in the processor, not in the operating system and not in C. The ARM processor used in this book defines the result as zero; others raise.

  1. Give one case where polling is the right choice. A device so fast that waiting for it

takes less time than taking an interrupt would, for example a high speed network card under heavy load.

Contents This chapter on its own page

munotes.in13

Chapter Four

The Timer, and Why the Operating System Always Gets the Processor Back

Syllabus topic Module 1, "Fundamentals of Operating systems - Operating-System Operations"

In one line

A device called the timer interrupts the processor many times a second, so the operating system gets control back whether the running program co-operates or not.

The precise form: the operating system sets a timer to interrupt after a specified period. A timer interrupt returns control to the operating system, which may then take the processor away from the running program. Setting the timer is a privileged instruction.

Why nothing works without it

Consider an operating system with no timer. A program is given the processor. What brings the operating system back?

  • If the program asks for something, a trap does. A program that reads a file gives up the

processor while it waits.

  • If the program makes a mistake, an exception does.
  • If a device finishes, an interrupt does, but that only borrows the processor for a moment and

returns it to the same program.

So control comes back only when the program lets it, and a program that computes in a loop without asking for anything never lets it. One line of C is enough to own the machine for ever:

int main(void)
{
    while (1) {
        /* nothing at all */
    }
}

On an operating system that relies on programs to give the processor up, that program is the end of the session. This is not a hypothetical: early personal computers behaved exactly like that, and it is why a single misbehaving program could hang the whole machine.

The word for the two designs is examined.

Non-preemptive, or co-operativePreemptive
The processor is taken backonly when the program gives it upat any moment, by the operating system
Needsprograms to behavea timer
One bad programstops everythingis stopped
Used byearly systems, some embedded controllersevery general purpose operating system today

The timer, and what a tick is

The timer is an ordinary device with one job: count down, and raise an interrupt when it reaches zero. Two ways of using it are both in service.

  1. A periodic tick. The timer is set to interrupt at a fixed rate, say a hundred or a

thousand times a second, for ever. Every tick the kernel updates the clock, charges the running program for the time it has used, and asks the scheduler whether somebody else should run.

  1. A one shot timer. The timer is set for exactly as long as the running program is allowed,

and it fires once. This is what modern kernels prefer, because a machine with nothing to do should not be woken a thousand times a second to be told there is nothing to do.

The tick is the unit in which the kernel charges a program for the processor. The rate is published by the system, and the lab machine reports it like this.

munotes.in14

The Timer, and Why the Operating System Always Gets the Processor Back

$ getconf CLK_TCK
100

A hundred ticks to the second, so one tick is ten milliseconds. Every figure the kernel reports about how long a program has computed is a whole number of those ticks. That is why the processor time of a short program is often reported as zero: it did not run for a whole tick.

A program that computes is charged; a program that waits is not. That is the whole meaning of the tick, and it can be measured. Here are two programs that both take about two seconds of the world's time.

#include <stdio.h>

int main(void)
{
    volatile long i;
    long total = 0;

    for (i = 0; i < 200000000L; i++) {
        total += i & 1;
    }
    printf("total %ld\n", total);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o count count.c
$ /usr/bin/time -f "%U" -o busy.txt ./count
total 100000000
$ /usr/bin/time -f "%U" -o idle.txt sleep 2
$ awk '{print ($1 > 0.2) ? "the counting program was charged real processor time" : "charged almost nothing"}' busy.txt
the counting program was charged real processor time
$ awk '{print ($1 < 0.1) ? "the sleeping program was charged none at all" : "charged something"}' idle.txt
the sleeping program was charged none at all

%U is the processor time a program spent in user mode. The counting program used the processor and was billed for it, in ticks. The sleeping program was not running at all, so there was no tick during which it held the processor, and its bill is zero.

Preemption, watched happening

Here is the proof that the operating system gets the processor back. The program above is run, another command is run at the same time, and the second one finishes.

$ gcc -std=c17 -Wall -Wextra -o spin spin.c
$ set +m
$ ./spin & spin=$!
$ date
Tue Sep 29 10:30:17 IST 2026
$ echo 2 + 2 | bc
4
$ kill $spin 2>/dev/null
$ wait 2>/dev/null

./spin & started the loop in the background, where it does nothing but compute. The two commands after it ran perfectly, in their turn, and answered. Nothing asked the loop politely to step aside: it was interrupted by the timer, hundreds of times, and the scheduler let the other work through.

& puts a command in the background and kill %1 ends it. The loop would otherwise have run until the machine was switched off, which is the point of the chapter.

The other half of the proof is that the kernel knows how long the loop ran without the loop ever telling it.

munotes.in15

The Timer, and Why the Operating System Always Gets the Processor Back

$ set +m
$ ./spin & spin=$!
$ sleep 2
$ awk '{print "ticks charged: " $14 + $15}' /proc/$spin/stat
ticks charged: 40
$ kill $spin 2>/dev/null
$ wait 2>/dev/null

Fields 14 and 15 of that file are the ticks the program has spent in user mode and in kernel mode. The program printed nothing, asked for nothing and reported nothing. The count exists because the kernel was interrupted while the program was running, looked at whose program it was, and added one to the tally.

Thirty four ticks is a third of a second, not two, and that is honest: the lab machine shares its processors with other work, so the loop was ready far more often than it was running. On an idle machine the two figures would be much closer. What the chapter proves is that the count exists at all and that nothing the program did produced it.

Worked example: how a quantum becomes a number

A round robin scheduler wants to give each program ten milliseconds at a time. The tick is one millisecond.

  1. The kernel sets the timer for ten ticks and dispatches program P1.
  2. Ticks one to nine arrive. At each one the kernel adds a tick to P1's account and sees that the

quantum has not run out, so it returns to P1 immediately. Nine interrupts, nine returns.

  1. Tick ten arrives. The quantum has run out. The kernel saves P1's registers into its process

control block, puts P1 at the back of the ready queue, chooses P2, loads P2's registers, and returns into P2.

  1. The timer is set for ten ticks again.

Two things in that list are examinable. The first is that the quantum is counted in ticks, so a quantum cannot be shorter than a tick, and asking for a four and a half millisecond quantum on a machine with a one millisecond tick gets four or five. The second is that step 3 is the context switch of Chapter seventeen, and steps 1 to 4 happen between ten and a thousand times a second on the machine you are reading this on.

Count the cost. If a context switch takes five microseconds and the quantum is ten milliseconds, the machine spends a fraction 5 / 10005 of its time switching, which is a twentieth of a per cent. Make the quantum one hundred microseconds instead and it spends 5 / 105, about 4.8 per cent. That is the trade of Chapter fifty two, and it is arithmetic rather than opinion.

Why setting the timer is privileged

If a program could set the timer, it could set it to never fire. Everything in this chapter would then be voluntary again, and one line of C would still be the end of the session. So loading the timer is a privileged instruction, exactly like setting the mode bit, and only the kernel does it.

munotes.in16

The Timer, and Why the Operating System Always Gets the Processor Back

The same argument covers turning interrupts off. A program that could do that could ignore the timer even while it fired.

Distinctions that carry marks

Timer interruptSystem call
Who caused itthe hardware, on a schedule the kernel setthe running program
Was the program asking?noyes
Kindinterrupttrap
Usual resultthe clock advances, perhaps a context switcha service is performed
A tickA quantum
What it isthe interval between two timer interruptshow long a program is allowed to run
Set bythe kernel's build, or the kernel at run timethe scheduler
Typical sizeone to ten millisecondsten to a hundred milliseconds
Relationshipthe quantum is a whole number of ticksit cannot be less than one tick

What it does not mean

The timer does not cause a context switch every time it fires. Most ticks the kernel simply notes the time and returns to the same program. A switch happens when the quantum has run out or somebody more urgent is ready.

Preemption is not unfair to the program that is stopped. It loses nothing: its registers are saved and restored, and it cannot tell. The only thing it loses is the illusion that it owns the machine.

The timer is not the clock. The clock is the date and time, kept in a battery backed chip and in the kernel. The timer is a countdown device used for scheduling. They are related only in that the kernel uses timer interrupts to advance its idea of the time.

A one second sleep is not one second of processor time. A sleeping program uses none at all. TIME in the ps output above was two seconds because the loop was computing, not waiting.

Quick revision

  • The operating system sets a timer to interrupt after a period. The interrupt returns

control to the kernel, which may take the processor away. Setting the timer is privileged.

  • Without a timer, control returns only when the program gives it up: that is non-preemptive

or co-operative scheduling, and one infinite loop ends the session.

  • A tick is the interval between timer interrupts. getconf CLK_TCK reports 100 on the lab

machine, so a tick is ten milliseconds and processor time is charged in those units.

  • A quantum is how long a program may run, counted in whole ticks.
  • Fields 14 and 15 of /proc/<pid>/stat are the ticks a program has been charged in user mode

and in kernel mode. A computing program's count climbs; a sleeping program's does not.

munotes.in17

The Timer, and Why the Operating System Always Gets the Processor Back

  • Two designs: a periodic tick at a fixed rate, or a one shot timer set to the quantum.
  • Most ticks do not cause a context switch. The kernel notes the time and returns.
  • The switching overhead is the switch time divided by the quantum plus the switch time, and it

is a sum, not an opinion.

Test yourself

  1. Why must the timer exist? Because without it the operating system regains control only

when the running program asks for something, so a program that computes in a loop would keep the processor for ever.

  1. Why is setting the timer a privileged instruction? A program that could set it could stop

it firing, and then preemption would be voluntary again.

  1. What is a tick, and what is the tick on the lab machine? The interval between two timer

interrupts. getconf CLK_TCK gives 100 ticks per second, so ten milliseconds.

  1. Does every timer interrupt cause a context switch? No. Usually the kernel advances the

time, charges the running program, finds the quantum has not expired, and returns to the same program. 5. A machine has a five microsecond context switch. Compare the overhead at a ten millisecond quantum and at a one hundred microsecond quantum. the fraction 5 / 10005, a twentieth of a per cent, against 5 / 105, about 4.8 per cent. A short quantum responds faster and wastes far more.

  1. A program sleeps for five seconds. How much processor time is it charged? Almost none. It

is not running while it sleeps, and processor time is charged only for ticks during which it held the processor.

Contents This chapter on its own page

munotes.in18

Chapter Five

The Functions of an Operating System

Syllabus topic Module 1, "Fundamentals of Operating systems - Functions of Operating System"

In one line

An operating system manages six things: processes, memory, files, devices, secondary storage, and who is allowed to do what. Everything else it does is one of those six in disguise.

Why this is a list and not a definition

Chapter one said what an operating system is and Chapter two said what its role is. Neither says what work it actually does, and an examination question that asks for the functions wants the work.

The six functions below are the standard division, and they are not arbitrary: each one is a resource with a manager. Recognising that is worth more than memorising the list, because it tells you what every manager must do. A manager of anything must know what it has, know who has it, decide who gets it next, and take it back.

The six functions

1. Process management

A process is a program in execution. The operating system creates processes, ends them, suspends and resumes them, gives each one the processor in turn, and provides the machinery for two processes to talk to each other and to stay out of each other's way.

The workWhere in this book
Create and end a processChapters nineteen to twenty one
Suspend and resumeChapters fifteen and seventeen
Schedule the processorChapters forty three to fifty five
Let processes communicateChapters twenty two to twenty six
Keep them out of each other's wayChapters thirty three to forty two
Get them out of a deadlockChapters fifty six to sixty five

2. Memory management

Main memory is the only storage the processor can use directly, and there is never enough of it. The operating system keeps track of which parts are in use and by whom, decides which processes to bring in when space is free, and allocates and frees space as processes ask.

The workWhere in this book
Keep track of what is used and by whomChapters sixty six to seventy three
Decide who is in memoryChapters sixty nine and ninety one
Allocate and free spaceChapters seventy to seventy nine
Run a program bigger than the memoryChapters eighty to ninety one

3. File management

A file is a named collection of related information. The operating system creates and deletes files and directories, supports the operations on them, maps files onto the disk, and backs them up.

The workWhere in this book
Create, delete, read, writeChapters ninety eight to one hundred
DirectoriesChapters one hundred one and one hundred six
Map a file onto blocksChapters one hundred seven and one hundred eight

4. Device, or input and output, management

Every device is different and no program should have to know how. The operating system provides one uniform interface, drives the hardware through device drivers, and buffers, caches and spools so that slow devices do not hold fast ones up.

munotes.in19

The Functions of an Operating System

Three words that are examined together and are constantly confused:

WordWhat it isWhy
Bufferingkeeping data in memory while it is in transitthe device and the program work at different speeds
Cachingkeeping a copy of data that is likely to be wanted againa second read of the same thing should not cost a second disk access
Spoolingqueueing whole jobs for a device that cannot interleave thema printer cannot print two documents at once

A buffer holds data on the way through. A cache holds a copy of data that is somewhere else. A spool is a queue of whole jobs.

5. Secondary storage management

The disk is where everything lives when it is not in memory. The operating system manages free space, allocates storage, and schedules the requests that reach the disk.

The workWhere in this book
Free spaceChapter one hundred nine
AllocationChapters one hundred seven and one hundred eight
Disk schedulingChapters ninety four to ninety six

6. Protection and security

Protection is controlling what the programs and users on this machine may do to each other. Security is defending the machine from outside it. Both are the control program of Chapter two.

Those two words are not synonyms and a question that names one does not accept the other. Protection is internal and is enforced by the kernel on every access. Security is external: passwords, encryption, keeping the software patched.

And two that are usually added

Networking is really device management plus communication, but it is large enough that most books list it separately. Command interpretation is the shell of Chapter eight, which is not part of the kernel at all but is always shipped with it.

Seeing all six on one machine

One command each, and every one of them is the operating system reporting on a function.

$ ps -o pid,stat,comm --no-headers | head -3
      8 Ss   sh
      9 S    bash
     12 R+   ps
$ free -m | head -2
               total        used        free      shared  buff/cache   available
Mem:            3905         512         591           0        2985        3393
$ ls -l /etc/hostname
-rw-r--r-- 1 root root 7 Sep 30  2026 /etc/hostname
$ df -h / | awk 'NR==2 {print "the root file system is of type", $1}'
the root file system is of type overlay
$ id
uid=1000(student) gid=1000(student) groups=1000(student),27(sudo)

ps is process management, free is memory management, ls is file management, df is secondary storage management, and id is the identity every protection decision is made against. Not one of those commands touched the hardware. Each asked the kernel and printed the answer.

munotes.in20

The Functions of an Operating System

Worked example: one keystroke, all six functions

Priya presses Ctrl and S in a text editor to save her file. Follow it.

  1. Device management. The keyboard interrupts. The kernel's driver reads the key and puts it

where the editor can find it.

  1. Process management. The editor was blocked waiting for input. The kernel moves it from the

waiting state to the ready state, and the scheduler eventually gives it the processor.

  1. File management. The editor asks the kernel to write the file. The kernel checks the file

exists, checks Priya may write it, and finds where its blocks are.

  1. Protection. That permission check. If the file belonged to somebody else and was not

writable by her, the write is refused here and nothing below happens.

  1. Memory management. The bytes to be written are in the editor's memory. The kernel copies

them into its own buffers, because the editor's memory may be paged out before the disk is ready.

  1. Secondary storage management. The kernel decides which blocks on the disk the file will

occupy, marks them used in the free space map, and queues the write.

  1. Device management again. The disk driver takes the queued request. Later the disk

interrupts to say it is done.

Every function in the list did something, for one keystroke, in a few milliseconds. That is why they are listed together.

Distinctions that carry marks

ProtectionSecurity
Againstother users and programs on this machineattackers outside it
Enforced bythe kernel, on every accesspasswords, encryption, patching, policy
Exampleyou may not read my filesomebody guessed my password
BufferingCachingSpooling
Holdsdata in transita copy of data kept elsewherewhole jobs waiting
Reasonspeed mismatchavoid fetching twicethe device takes one job at a time
Exampledisk block on its way to a programthe page cachethe print queue

What it does not mean

The six are not six separate programs. They are six kinds of work done by one kernel, and they constantly call on each other, as the worked example shows.

"File management" does not mean an application. A file manager with icons is a program. File management is the kernel's work of turning a name into blocks on a disk.

The shell is not a function of the operating system. Command interpretation is a service that is shipped with it and runs as an ordinary program.

Quick revision

  • Six functions: process management, memory management, file management, device

management, secondary storage management, protection and security. Networking and command interpretation are usually added.

  • Every function is a resource with a manager, and every manager must know what it has, know who
munotes.in21

The Functions of an Operating System

has it, decide who is next, and take it back.

  • Buffering is data in transit. Caching is a copy of data held elsewhere. Spooling is

a queue of whole jobs.

  • Protection is internal and enforced by the kernel. Security is external.
  • A single keystroke that saves a file uses every one of the six.

Test yourself

  1. List the functions of an operating system. Process management, memory management, file

management, device or input and output management, secondary storage management, and protection and security. Networking and command interpretation are commonly added.

  1. Distinguish buffering, caching and spooling. A buffer holds data in transit between a

device and a program working at different speeds. A cache holds a copy of data that lives elsewhere, so a second use is cheap. A spool queues whole jobs for a device that can handle only one at a time.

  1. Distinguish protection from security. Protection controls what users and programs on this

machine may do to each other and is enforced by the kernel. Security defends the machine from outside.

  1. What must any resource manager do? Know what it has, know who holds what, decide who gets

it next, and take it back when they are done.

  1. Which function does the print queue belong to? Device management, as spooling.
  2. Name the function that a permission check belongs to, and say when it happens. Protection,

and it happens on every access, inside the kernel, before the work is done.

Contents This chapter on its own page

munotes.in22

Chapter Six

Where Operating Systems Run

Syllabus topic Module 1, "Fundamentals of Operating systems - Computing Environments"

In one line

The same ideas run on very different machines, and what changes between them is which of the ideas matters most.

Why the environment changes the operating system

Every environment in this chapter needs process management, memory management and a file system. What differs is the constraint that dominates. On a phone it is the battery. On a server it is throughput. In an engine controller it is the deadline. In the cloud it is that the machine you are given is not a machine at all.

That is the useful way to revise this topic: for each environment, name the constraint, and then say what the operating system does differently because of it.

The environments

Traditional computing

A desktop or laptop with one person using it. One or a few programs are interactive and the rest are waiting. The constraint is response time: the person notices a tenth of a second.

What follows: preemptive scheduling with a short quantum, a scheduler that favours interactive programs over computing ones, and a great deal of memory given to caching so that the machine feels quick.

Mobile computing

A phone or tablet. The constraints are battery and memory, and there is no fan, so heat is a third.

What follows: the operating system suspends programs the moment they leave the screen rather than scheduling them; it will kill a background program to give memory to the one in front; and it does not use a swap area on flash storage the way a desktop uses one on a disk, because writing to flash wears it out. Android is a Linux kernel with all of that on top; iOS is a different kernel with the same constraints.

Client server computing

One machine provides a service and many ask for it. The constraint is throughput: how many requests are finished per second.

What follows: scheduling for fairness between many equal jobs rather than for one person's response time, long quanta because there is nobody watching a cursor blink, and a file system tuned for many readers at once.

Client and server are roles, not machines. The same computer is a server to the browser on your phone and a client to the database behind it.

Peer to peer computing

No machine is in charge. Every machine is both client and server, and a machine joining the network announces itself or is found by a broadcast. The constraint is that there is no single place to ask.

Virtualisation

One physical machine pretends to be several. A hypervisor runs on the real hardware and each guest operating system believes it has a machine of its own. The constraint is that the guest must not be able to tell, and must not be able to escape.

munotes.in23

Where Operating Systems Run

This one is worth attention because it is how almost all server computing is now done, and because it is a beautiful use of Chapter three: the guest kernel runs in user mode. When it executes a privileged instruction, the processor traps, the hypervisor catches the trap, does the equivalent thing to the guest's simulated hardware, and returns. The guest's own protection still works, because the guest's programs are one further level down.

A container is the lighter relative: one kernel, but each group of processes given its own view of the file system, the process table and the network. It is not a second operating system, and that is the difference examiners look for.

Cloud computing

Computing bought by the hour over a network. Three shapes are named in every syllabus:

NameYou are givenYou manage
Infrastructure as a servicea machine, or something that behaves like onethe operating system and everything above
Platform as a servicesomewhere to put your programyour program
Software as a servicea working applicationnothing

Cloud computing is virtualisation plus a network plus somebody else's electricity bill. The operating system ideas do not change; the question of who runs them does.

The older kinds, which a question still names

MU's own textbook came up through these, and a question asks for them by name, so they are worth a table even though the machine on your desk is none of them.

KindHow it workedWhat the operating system had to do
batchjobs on cards or tape, collected and run one after another with no user presentqueue the jobs, run one, print its output, run the next. No interaction at all
multiprogrammed batchseveral jobs in memory, and when one waits for input or output another runskeep several jobs resident, and switch when one blocks. This is where scheduling begins
time sharing, also called multitaskingmany users at terminals, each given the processor for a short sliceswitch often enough that each user believes the machine is theirs: a short quantum, which is Chapter fifty one's arithmetic
single processorone processing coreschedule one queue
multiprocessor, or parallelseveral processors sharing memory, either symmetric, where every processor runs the same kernel, or asymmetric, where one is the masterbalance the load, and protect shared kernel data, which is Chapter thirty four's problem

The three properties a question asks for. A batch system maximises throughput and nobody waits at a terminal; a time sharing system minimises response time and accepts more overhead to get it; a real time system guarantees a deadline and gives up average speed for predictability. Chapter forty five's five criteria are the same three ideas, measured.

munotes.in24

Where Operating Systems Run

The advantage of a multiprocessor system is also a stock question: more work done in the same time, better economy than that many separate machines because they share memory and devices, and more reliability, because the work of a failed processor can be taken by the others, which is called graceful degradation.

Real time embedded systems

A controller inside something else: a washing machine, an anti-lock brake, a pacemaker, a satellite. The constraint is the deadline, and this is the one environment where being late is being wrong.

Hard real timeSoft real time
A missed deadlineis a failuredegrades the result
Examplean airbag, a brake controllervideo playback, audio
Needsguaranteed worst case timingbest effort with priority

A hard real time system is not a fast system. It is a predictable one. A processor twice as fast with unpredictable timing is worse, not better. That is why hard real time kernels give up features this book teaches: virtual memory is often switched off, because a page fault takes an unpredictable time, and secondary storage may be absent altogether.

The lab machine, honestly

The machine every program in this book was run on is worth describing, because it is an example of two of the environments above at once.

$ uname -s -r -m
Linux 6.8.0-117-generic aarch64
$ nproc
2
$ free -m | awk 'NR==2 {print "memory " $2 " megabytes"}'
memory 3905 megabytes
$ grep -c processor /proc/cpuinfo
2

Two processors, four gigabytes of memory, an ARM instruction set, and a Linux kernel. It is a container inside a virtual machine: a single kernel shared with its host, with this book's programs given their own view of the file system and the process table. Everything about processes, memory, synchronisation and files behaves exactly as it does on a physical machine, which is why it can be used at all.

Distinctions that carry marks

Virtual machineContainer
Kernels runningone per guest, plus the hostone, shared
Isolationa whole simulated machinea separate view of the file system, processes and network
Starts intens of secondsmilliseconds
Can runa different operating systemonly programs for the one kernel
Hard real timeTraditional
Judged byworst caseaverage case
Virtual memoryoften switched offessential
A late answeris wrongis merely slow

What it does not mean

"Mobile" is not a cut down desktop. The constraints are different, so the decisions are different, and a phone operating system does things a desktop would never do, like killing a running program to free memory.

munotes.in25

Where Operating Systems Run

The cloud is not a kind of operating system. It is a way of buying computing. Underneath it is virtualisation and the same operating systems.

Real time does not mean quick. It means on time, every time, with a known worst case.

Quick revision

  • Environments: traditional, mobile, client server, peer to peer,

virtualisation, cloud, real time embedded.

  • Revise each one by its dominating constraint: response time, battery, throughput, no central

authority, the guest must not tell, who pays, the deadline.

  • A hypervisor runs guests whose kernels run in user mode; their privileged instructions trap

to the hypervisor.

  • A container shares one kernel and is given a separate view of the system. It is not a

second operating system.

  • Cloud service shapes: infrastructure, platform and software as a service.
  • Hard real time means a missed deadline is a failure, and it buys predictability by giving

up features such as virtual memory. Soft real time degrades gracefully.

Test yourself

  1. Name the computing environments. Traditional, mobile, client server, peer to peer,

virtualised, cloud, and real time embedded.

  1. Why does a phone kill background programs when a desktop does not? Memory is limited and

there is no disk to swap to without wearing the flash storage out, so the operating system reclaims memory by ending programs the user is not looking at.

  1. How does a hypervisor keep a guest kernel under control? The guest kernel runs in user

mode, so every privileged instruction it executes traps to the hypervisor, which performs the equivalent operation on the guest's simulated hardware.

  1. Distinguish a virtual machine from a container. A virtual machine runs its own kernel and

can run a different operating system; a container shares the host's kernel and only gets a separate view of the file system, processes and network.

  1. Is a hard real time system a fast system? No. It is a predictable one. A faster but

unpredictable machine is worse for the job.

  1. Which cloud shape leaves you managing the operating system? Infrastructure as a service.

Contents This chapter on its own page

munotes.in26

Chapter Seven

The Services an Operating System Offers

Syllabus topic Module 1, "Fundamentals of Operating systems - Operating-System Services"

In one line

An operating system offers one set of services to programs, so that a program need not know how the machine works, and a second set to the machine's owner, so that the machine can be shared.

Why they are in two groups

The grouping is the examinable part and it is easy to state. Some services are useful to the program that asks for them. Others are useful to everybody except the program that they are applied to.

A program wants its file opened: that is a service for the program. A program does not want to be accounted for, limited, or made to share: those are services for the person who owns the machine. Putting them in one list hides the difference.

Group one: services for the program

Program execution

The operating system will load a program into memory and run it, and will end it, normally or otherwise. A program can ask for this about another program, which is what a shell does every time you type a command.

Input and output operations

A running program cannot drive a device. It asks. This is not only convenience: a program that could drive the disk directly could read any file on it, so the refusal is protection as well as abstraction.

File system manipulation

Create, delete, read, write, search, list, and manage permissions. Notice that permission is part of the service: the operating system does not simply do as it is told.

Communications

Two processes talking, on one machine through shared memory or messages, or across a network. The whole of Chapters twenty two to twenty six is this one service.

Error detection

The operating system must notice errors and act. Errors in the hardware (a memory parity error, a disk that will not read), in devices (paper out), and in programs (an arithmetic exception, an invalid address). For each, it must take an action that keeps the rest of the machine correct.

Chapter two's segmentation fault was this service working. The kernel noticed, and acted, and the rest of the machine was unaffected.

Group two: services for the machine and its users

Resource allocation

Which program gets the processor, how much memory, which disk blocks, who gets the printer. All of Chapter two's first job.

Accounting

Keeping a record of who used what and for how long. On a college machine this is how a fair usage policy is enforced, and on a cloud machine it is the bill.

Protection and security

Every access checked; every user identified; the machine defended from outside.

System programs, which are not services

A question asks what a system program is, and the answer is the distinction this chapter turns on. A service is offered by the kernel through a system call. A system program is an ordinary program that ships with the operating system and uses those services like any other program.

munotes.in27

The Services an Operating System Offers

CategoryExamples
file managementcopy, move, delete, list, rename
status informationthe date, the free space, who is logged in, how busy the machine is
file modificationeditors, and tools that search or transform a file's text
programming language supportcompilers, assemblers, interpreters, debuggers
program loading and executionthe loader, the linker, and the shell that starts a program
communicationsmail, remote login, file transfer, and the tools that make a connection
background servicesthe daemons that run from boot to shutdown: printing, logging, the clock

Most of what a user calls the operating system is system programs. The kernel is small and invisible; the shell, the file manager and the compiler are what a person actually touches, and replacing any of them changes nothing about the kernel underneath.

Seeing them, one at a time

Each command below is one service being asked for and given.

$ /usr/bin/time -f "the kernel accounted: %U user, %S system, %M kilobytes peak" ls > /dev/null
the kernel accounted: 0.00 user, 0.00 system, 2176 kilobytes peak
$ ipcs -l | head -4

------ Messages Limits --------
max queues system wide = 32000
max size of message (bytes) = 8192
$ id -u
1000
$ ulimit -n
1024

/usr/bin/time printed the kernel's own accounting record for a program that had just finished: processor time in both modes, and the most memory it ever held. ipcs -l listed the limits of the communication facilities the kernel offers. id -u is the identity every protection decision is made against. ulimit -u is a resource allocation limit: the most processes this user may have at once, which is what stops one account bringing the machine down by making processes for ever.

Worked example: a service refused, three ways

The same program asks for three things and is refused each time, for three different reasons. Each refusal is a service, not a failure.

$ cat /etc/shadow
cat: /etc/shadow: Permission denied
$ mkdir /usr/newdir
mkdir: cannot create directory ‘/usr/newdir’: Permission denied
$ ulimit -v 2000
$ ls
ls: error while loading shared libraries: libc.so.6: failed to map segment from shared object
$ ulimit -v unlimited
bash: ulimit: virtual memory: cannot modify limit: Operation not permitted

The first two refusals are protection: the file and the directory belong to root, and the check was made by the kernel rather than by cat or mkdir.

The third is resource allocation. The shell was told to allow itself and its children only two megabytes of virtual memory. ls could not even map the C library into that, so it never started. The kernel did not crash, did not let the program run without memory, and reported exactly what went wrong.

munotes.in28

The Services an Operating System Offers

The fourth line is the same service again, seen from the other side. Putting the limit back was refused. Every limit of this kind has a soft value, which a program may lower and then raise again up to a hard ceiling, and a hard value, which may only ever be lowered. ulimit -v 2000 lowered both, so there was no way back. That is deliberate: a program that could raise its own ceiling would not be limited at all.

This is worth knowing before you try it on your own machine. A shell whose hard limit you have lowered cannot undo it. Close that shell and open another.

Distinctions that carry marks

A serviceA resource
Iswork the operating system does on requestsomething held for a while
Asked for witha system calla system call
Given back?nothing to give backmust be released
Examplewrite these bytesfour megabytes of memory
Services for the programServices for the system
Who benefitsthe program that askedeverybody else
Examplesprogram execution, input and output, files, communication, error detectionresource allocation, accounting, protection and security
Would a program ask for it?yesno: it is applied to the program

What it does not mean

A service is not a program. The shell, the compiler and the file manager are programs that use services. The service is what the kernel does when asked.

Error detection does not mean the operating system fixes your bug. It means it notices, reports, and keeps the rest of the machine correct. Your program still dies.

Accounting is not surveillance of what you typed. It is a count of resources: processor time, memory, disk, network. The kernel does not read your file to account for it.

Quick revision

  • Services for the program: program execution, input and output operations, file system

manipulation, communications, error detection.

  • Services for the system and its users: resource allocation, accounting, protection and

security.

  • The test for which group a service is in: would the program itself ask for it?
  • Input and output through the operating system is protection as well as convenience: a program

that could drive the disk could read every file on it.

  • Error detection means notice, report, and keep the rest of the machine correct.
  • /usr/bin/time shows accounting, ipcs -l the communication facilities, id the identity

protection is checked against, ulimit an allocation limit.

  • A limit has a soft value a program may lower and raise again, and a hard value that may
munotes.in29

The Services an Operating System Offers

only be lowered. Lowering the hard value cannot be undone.

Test yourself

  1. List the services an operating system provides. For programs: program execution, input and

output operations, file system manipulation, communications, error detection. For the system: resource allocation, accounting, protection and security.

  1. Why is input and output a service rather than something a program does itself? Because a

program cannot be trusted or expected to drive the hardware: it would have to know which device is fitted, and it could read anything on it.

  1. Which service was at work when a program died with a segmentation fault? Error detection.

The kernel noticed an invalid access, ended that program, and left the rest of the machine correct.

  1. Give a service that a program would never ask for, and say who wants it. Accounting. The

owner of the machine wants it, to enforce fair use or to bill.

  1. What does ulimit -n control, and which service is that? The most files a program may

have open at once. Resource allocation.

  1. Why can a program lower its own resource limit and not raise it again? Because the hard

value may only be lowered. If a program could raise its own ceiling the limit would mean nothing.

  1. Distinguish a service from a resource. A service is work done on request and there is

nothing to give back. A resource is held for a while and must be released.

Contents This chapter on its own page

munotes.in30

Chapter Eight

The Command Line and the Desktop

Syllabus topic Module 1, "Fundamentals of Operating systems - User and Operating-System Interface"

In one line

There are two ways for a person to ask an operating system for something: type a command, or point at a picture of it. Neither is part of the operating system.

Why there are two, and why both survive

The command line came first because a teletype was what there was. It survived because it has three properties the desktop does not:

  1. It composes. The output of one command can be the input of the next, so two simple

programs make a third thing nobody wrote.

  1. It repeats. A command can be saved in a file and run again, or run a thousand times on a

thousand files.

  1. It travels. A command line works over a slow network connection to a machine in another

country.

The desktop came second and survived because it has one property the command line does not: you do not have to know the name of the thing you want. You can see it.

A student is examined on both, and the honest comparison is that they are for different jobs. A server room is run from the command line. A phone has no command line at all.

The shell is a program

This is the part that matters for the rest of the book. The shell, or command interpreter, is an ordinary program. It has a name, a file, a size and a version. Several are available and you may run whichever you like.

$ readlink /proc/$$/exe
/usr/bin/bash
$ which bash sh dash
/usr/bin/bash
/usr/bin/sh
/usr/bin/dash
$ ls -l /usr/bin/sh
lrwxrwxrwx 1 root root 4 Sep  1 14:30 /usr/bin/sh -> dash
$ sh -c 'echo I am a different shell, and $0 is my name'
I am a different shell, and sh is my name

The first command asks the kernel which executable file the running shell came from, and the answer is a path. Three shells are installed. sh is a link to dash, a small fast one used for scripts. sh -c runs one command in a fresh shell and exits, which is why it appears throughout this book whenever a program's death has to be reported plainly.

The loop the shell runs is four lines long, and it explains everything a shell does:

  1. Print a prompt.
  2. Read a line.
  3. If the line names one of the shell's own built in commands, do it. Otherwise find the program,

start it, and wait for it to finish.

  1. Go back to 1.

Step 3 is where fork and exec live, and Chapters nineteen and twenty build exactly that.

Built in against external, which is examined

Some commands are inside the shell and some are programs on the disk. The shell will tell you which is which.

munotes.in31

The Command Line and the Desktop

$ type cd
cd is a shell builtin
$ type ls
ls is hashed (/usr/bin/ls)
$ type -a echo
echo is a shell builtin
echo is /usr/bin/echo
echo is /bin/echo
$ type grep
grep is /usr/bin/grep

cd has to be built in. Changing directory changes a property of the running process, and a separate program could only change its own, then exit and take the change with it. That is a favourite examination question and the reason is the whole answer.

ls is a program in a file, and "hashed" means the shell has remembered where it found it so that it need not search the next time.

echo is both: the shell has its own built in and there is also a program on the disk. The built in wins unless you ask for the other by its path. It is listed twice because /bin on this system is a link to /usr/bin, so both names reach the same file.

What the shell adds, beyond running a command

None of the following is the operating system's work. All of it is the shell doing arithmetic on file descriptors and process ids, using the same system calls any program may use.

WrittenWhat the shell does
command > fileopens the file and makes it the program's output before starting it
command < filethe same for input
the vertical bar between two commandsmakes a pipe, starts both, wires one to the other (Chapter twenty four)
command &starts it and does not wait (Chapter twenty one)
*.txtexpands the pattern to the matching file names before the program is started
$HOMEreplaces the variable with its value

The program never sees the pattern. By the time ls *.txt runs, ls has been given a list of names. This is worth proving, because it explains a great deal of otherwise confusing behaviour.

$ touch a.txt b.txt c.log
$ echo *.txt
a.txt b.txt
$ ls -1 *.txt
a.txt
b.txt

echo has no idea about files, and it printed two file names. The shell had already replaced the pattern.

The graphical interface

A desktop is a program too, or rather several: a display server that owns the screen and the mouse, a window manager that decides where windows go and which is in front, and a file manager that draws the icons. On Linux the display server is X11 or Wayland; on Windows and macOS the equivalent is shipped as part of the product and is not separable.

The vocabulary MU's text book uses is worth knowing: the desktop metaphor presents icons for objects, a pointer for selection, windows for the programs, and menus for the available operations. Selecting an icon and choosing an operation is the same request as typing a command; the route from there into the kernel is identical, because the file manager makes the same system calls.

munotes.in32

The Command Line and the Desktop

$ startx

The lab machine has no screen, so nothing graphical can be run or shown here. What can be said is what the interface is made of, and that its requests reach the kernel by exactly the system calls of the next chapter.

Distinctions that carry marks

Command line interfaceGraphical interface
You must knowthe name of the commandnothing: you can look
Composesyes, with pipes and redirectionrarely
Repeatableyes, in a scriptwith difficulty
Over a slow networkworks wellworks badly
Memory and processor neededvery littlea great deal more
Learningslow at first, fast laterfast at first
Shell built inExternal command
Livesinside the shell processin a file on the disk
Costsnothing to starta fork and an exec
Can change the shell itselfyesno
Examplescd, export, exitls, grep, cp

What it does not mean

The shell is not the operating system. It is a program that asks the operating system for things, and it can be replaced without touching the kernel.

A shell script is not a different language from the command line. It is the same commands, in a file.

The graphical interface is not "easier" in every sense. It is easier to start and harder to automate, which is why both survive.

Quick revision

  • Two interfaces: the command line, or command interpreter or shell, and the

graphical one. Neither is part of the kernel.

  • The shell's loop: print a prompt, read a line, run a built in or find and start a program and

wait, repeat. Step three is fork and exec.

  • cd must be built in, because it changes the running process's own directory, and a

separate program could change only its own.

  • The shell does redirection, pipes, background running, pattern expansion and variable

substitution itself, with ordinary system calls. The program is handed the result, never the pattern.

  • type reports whether a command is a built in, an alias or a file. type -a shows all of

them.

  • The graphical interface is a display server plus a window manager plus a file manager, and it

makes the same system calls.

Test yourself

1. Name the two kinds of user interface and say whether either is part of the operating system. The command line interpreter, or shell, and the graphical user interface. Neither is: both are programs that use the operating system.

  1. Write the shell's main loop in four steps. Print a prompt; read a line; if it is a built
munotes.in33

The Command Line and the Desktop

in do it, otherwise find the program, start it and wait; repeat.

  1. Why must cd be a shell built in? It changes the current directory of the process. A

separate program could only change its own, and would then exit, taking the change with it.

  1. What does ls *.txt hand to ls? A list of the matching file names. The shell expands

the pattern before ls starts, so ls never sees the asterisk.

  1. Give two advantages of the command line over the desktop, and one the other way. It

composes with pipes and it can be repeated in a script; and it works over a slow connection. The desktop lets you find things without knowing their names. 6. A command is both a built in and a program on the disk. Which runs, and how would you get the other? The built in runs. Ask for the other by its full path, for example /usr/bin/echo.

Contents This chapter on its own page

munotes.in34

Chapter Nine

What a System Call Is

Syllabus topic Module 1, "Fundamentals of Operating systems - System Calls"

In one line

A system call is a request from a program to the kernel, made by executing a special instruction that traps into kernel mode.

The precise form: a system call is the programming interface to the services of the operating system. It is invoked by a trap instruction, which switches the processor to kernel mode and transfers control to the kernel's system call handler.

Why a program cannot simply call the kernel

A program calls one of its own functions by jumping to its address. Why not jump to the kernel's function the same way?

Because of Chapter three. The kernel's code sits in memory the program cannot touch, and it runs in kernel mode, which the program cannot enter. A jump would be an invalid access and the program would die with a segmentation fault. The mode bit cannot be set by the program, because setting it is privileged.

So there has to be one controlled door, and the door is a trap. The program executes an instruction whose whole purpose is to say "kernel, I want something", the processor switches mode, and the kernel's own handler decides what to do. The program never gets to choose where in the kernel it lands.

That single sentence is why the system call interface is the security boundary of the whole machine. Everything a program is allowed to do, it does through one of a few hundred numbered requests, each of which the kernel checks.

The steps, in order

  1. The program puts the system call number where the kernel will look for it, and the

arguments where the kernel will look for those.

  1. The program executes the trap instruction.
  2. The processor switches to kernel mode and jumps to the kernel's system call handler.
  3. The handler reads the number and uses it as an index into the system call table, which

holds the address of the routine for each number.

  1. That routine checks the arguments, does the work, and puts a return value where the program

will find it.

  1. The handler returns from the trap. The mode bit goes back to user mode and the program

continues at the instruction after the trap.

Step 5's check is not optional and is not a formality. The arguments came from a program that may be hostile. A pointer the program supplies might point into the kernel; a length might be negative; a file descriptor might not be open. Every one of those has to be verified before it is used, and a kernel bug in that checking is a security hole.

The three ways arguments are passed

More arguments are often needed than there are registers, and the trap instruction cannot carry them. Three methods are in use and all three are examined.

munotes.in35

What a System Call Is

MethodHowLimit
In registerseach argument in a named registerthe number of registers, typically six
In a block or table in memorythe program builds a block, and passes its address in one registernone
On the stackthe program pushes them, and the kernel pops themnone

Linux on this machine uses registers, which is why write takes exactly three arguments and why a call needing more, such as mmap with six, is at the practical limit.

A library call is not a system call

This is the distinction the chapter exists for. printf is a function in the C library. write is a system call. printf eventually calls write, but it is not the same thing, and the difference can be measured.

#include <stdio.h>
#include <unistd.h>
#include <string.h>

int main(void)
{
    const char *a = "written by the library, with printf\n";
    const char *b = "written by the system call, with write\n";

    printf("%s", a);
    write(STDOUT_FILENO, b, strlen(b));
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o twoways twoways.c
$ ./twoways
written by the library, with printf
written by the system call, with write

On the screen they look identical. Now send the output to a file instead, and the order changes.

$ ./twoways > out.txt
$ cat out.txt
written by the system call, with write
written by the library, with printf

The lines came out in the wrong order, and that is the proof. printf did not write anything when it was called: it put the text in a buffer inside the C library and returned. The write call went straight to the kernel. Only when the program ended did the library flush its buffer, which is why the library's line is last.

When the output was the screen, the library flushed at every newline, because a terminal is treated as interactive. When the output was a file, it flushed only when the buffer was full or the program ended. Nothing about the kernel changed. The library changed its mind.

This is a real bug students meet. A program that prints with printf and then crashes loses the printed lines, because they were never given to the kernel. A program that prints with write does not.

The count of system calls confirms it: printf called write once, at the end, with both lines buffered together or separately depending on where the output went.

What a system call looks like in the manual

Every system call has a manual page in section 2, and the lab machine also carries the POSIX standard's own page for the same call in section 3p. Reading both is how this book knows what a call promises.

munotes.in36

What a System Call Is

$ man 2 write 2>/dev/null | sed -n '/^SYNOPSIS/,/^DESCRIPTION/p' | head -6
SYNOPSIS
       #include <unistd.h>

       ssize_t write(int fd, const void buf[.count], size_t count);

DESCRIPTION
$ man 3p write 2>/dev/null | sed -n '2,4p'

PROLOG
       This  manual  page is part of the POSIX Programmer's Manual.  The Linux

The first is what Linux does. The second is what every Unix must do. Where they differ, the Linux page says so, and this book says which one it is quoting.

Worked example: counting the doors a simple program uses

/bin/true does nothing at all. It is the simplest program on the machine. Count the system calls it makes.

$ strace -c /bin/true 2>&1 | tail -2
------ ----------- ----------- --------- --------- ----------------
100.00    0.000000           0        27         1 total

Twenty seven system calls to do nothing at all. They are the program being loaded: the kernel finding and opening the C library, mapping it into memory, arranging the stack, and then exiting. Every program pays that, and it is why starting a process is not free, which is the point Chapter thirty one makes about thread pools.

Distinctions that carry marks

System callLibrary function
Runs inkernel modeuser mode
Entered bya trap instructionan ordinary function call
Costsa mode switch each timea jump
Can be avoided?no, for anything needing the hardwareyes, it is just code
Examplewrite, open, forkprintf, strlen, malloc
Where documentedmanual section 2manual section 3
System callOrdinary function call
Destinationchosen by the kernel from a tablechosen by the caller
Modechangesdoes not change
Arguments checkedalways, by the kernelnot usually

What it does not mean

A system call is not slow because it does a lot. It is expensive because of the mode switch and the checking, so a program that makes a million small calls is slower than one that makes a thousand large ones. This is why buffering exists at all.

printf is not "the system call for printing". There is no system call for printing. There is write, which puts bytes somewhere, and printf is a formatting function that ends up calling it.

A system call is not a function in your program that the kernel calls back. The traffic is one way: your program asks, the kernel answers.

The number of system calls is not large. Linux has a few hundred. Everything any program does goes through them.

Quick revision

  • A system call is the programming interface to the operating system's services, invoked by a

trap that switches to kernel mode.

  • There is one controlled door because the program cannot enter kernel mode itself and cannot
munotes.in37

What a System Call Is

choose where in the kernel it lands.

  • The steps: number and arguments in place, trap, mode switch, handler, system call table

lookup, argument check, work, return.

  • Arguments are passed in registers, in a block whose address is passed, or on the

stack.

  • printf is a library function, write is a system call. Redirecting a program's

output to a file changes the order in which they appear, because the library buffers and the system call does not.

  • A system call costs a mode switch, so few large calls beat many small ones.
  • Manual section 2 is Linux's system calls; section 3p is the POSIX standard's own page.

Test yourself

  1. Define a system call. The programming interface to the services of the operating system,

invoked by a trap instruction that switches the processor into kernel mode and enters the kernel's handler.

  1. Why can a program not just call a kernel function directly? The kernel's memory is not

accessible to it and the kernel runs in kernel mode, which the program cannot enter because setting the mode bit is privileged.

  1. Name the three ways arguments are passed to a system call. In registers, in a block of

memory whose address is passed in a register, and on the stack. 4. A program prints two lines, one with printf and one with write, and the output is redirected to a file. What order do they appear in and why? The write line first. printf buffered its text in the C library and the buffer was not flushed until the program ended.

  1. Why is a system call more expensive than a function call? It switches processor mode twice

and the kernel must check every argument, because the arguments came from a program that may be hostile.

  1. /bin/true does nothing, yet makes twenty seven system calls. What are they? The work of

starting a program: opening and mapping the C library, arranging memory, and exiting.

Contents This chapter on its own page

munotes.in38

Chapter Ten

Watching System Calls Happen

Syllabus topic Module 1, "Fundamentals of Operating systems - System Calls"

In one line

strace asks the kernel to report every system call a program makes, with its arguments and its answer, in order.

Why this is worth a chapter

Everything in this paper is invisible. You cannot see a process being scheduled, a page being replaced or a lock being taken. The system call interface is the one place where the whole of the operating system's work becomes a list of lines you can read.

It is also the best debugging tool a student will meet. A program that reports it cannot find a file becomes a trace showing the exact name it asked for and the exact answer it got.

How it works, in one paragraph

The kernel offers a facility by which one process may watch another: it stops the watched process at every entry to and exit from a system call and lets the watcher inspect it. strace uses that facility. The consequence worth knowing is that a program under strace runs far more slowly, because every call now stops it twice, so timings taken under it mean nothing.

The whole life of the simplest program

/bin/true returns success and does nothing else. Here is everything it asks the kernel for.

$ strace -o t.txt /bin/true
$ wc -l < t.txt
29
$ head -4 t.txt | sed 's/0x[0-9a-f]*/ADDRESS/g'
execve("/bin/true", ["/bin/true"], ADDRESS /* 14 vars */) = 0
brk(NULL)                               = ADDRESS
mmap(NULL, 8192, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = ADDRESS
faccessat(AT_FDCWD, "/etc/ld.so.preload", R_OK) = -1 ENOENT (No such file or directory)
$ grep -c mmap t.txt
6
$ tail -2 t.txt
exit_group(0)                           = ?
+++ exited with 0 +++

Read it as a story. The program was started with execve. A cache of library locations was looked for and opened. The C library was found, opened and mapped into memory six times over, once for each part of it that needs different permissions. Then the program exited.

Twenty nine lines and twenty seven calls: the last two lines are the exit and the note that the program finished, so they are not calls. Not one of those calls is the program's own work. It is the cost of being a program at all.

Two lines in that list are worth naming because they come back later in the book.

  • execve is the system call of Chapter twenty, the one that replaces a process's program.
  • mmap is the call of Chapter eighty, the one that maps something into a process's address

space. The C library arrives in memory by being mapped, not by being read.

A program that actually does something

Now a program that reads a file, so that the interesting calls are not buried.

#include <fcntl.h>
#include <unistd.h>
#include <stdio.h>

int main(void)
{
    char buf[32];
    int fd = open("greeting.txt", O_RDONLY);

    if (fd < 0) {
        perror("open");
        return 1;
    }
    ssize_t n = read(fd, buf, sizeof buf);
    write(STDOUT_FILENO, buf, (size_t)n);
    close(fd);
    return 0;
}
munotes.in39

Watching System Calls Happen

hello from a file
$ gcc -std=c17 -Wall -Wextra -o readone readone.c
$ strace -e trace=openat,read,write,close -o r.txt ./readone
hello from a file
$ tail -5 r.txt
openat(AT_FDCWD, "greeting.txt", O_RDONLY) = 3
read(3, "hello from a file\n", 32)      = 18
write(1, "hello from a file\n", 18)     = 18
close(3)                                = 0
+++ exited with 0 +++

Those are the last five lines: the ones above them are the C library being loaded, exactly as in the trace before. Four calls are the program's own, and every one of them is readable.

  • openat returned 3. That is the file descriptor, a small integer the kernel gives the

program to refer to the open file. Descriptors 0, 1 and 2 are already taken: standard input, standard output and standard error. So the first file a program opens is almost always 3.

  • read(3, ..., 32) asked for up to 32 bytes and got 18.

The return value is what was actually read, and a program that assumes it got everything it asked for is a broken program.

  • write(1, ..., 18) wrote to descriptor 1, standard output, which is why the text appeared on

the screen.

  • close(3) gave the descriptor back.

open appears in the trace as openat. The library calls the newer system call, which takes a directory to be relative to, and AT_FDCWD means the current directory. This is a real and common surprise: the name in your program need not be the name of the system call underneath.

Counting, rather than listing

For a large program the list is too long to read. Counting is more useful.

$ strace -c -o c.txt ./readone
hello from a file
$ awk 'NR<3 || $NF ~ /^(openat|read|close|mmap)$/ {print $(NF-1), $NF}' c.txt | sort
--------- ----------------
2 read
3 close
3 openat
6 mmap
errors syscall

The columns are the share of time, the total time, the average per call, the number of calls, the number that returned an error, and the name. The errors column is the one to look at first when a program misbehaves, because a call that failed and was ignored is the commonest cause of mysterious behaviour.

Worked example: why did it not find the file?

This is the use a student will get most value from. A program is given a name that does not exist.

$ rm -f greeting.txt
$ strace -e trace=openat -o miss.txt ./readone
open: No such file or directory
$ grep greeting miss.txt
openat(AT_FDCWD, "greeting.txt", O_RDONLY) = -1 ENOENT (No such file or directory)
munotes.in40

Watching System Calls Happen

The answer is -1 and the reason is ENOENT. A system call reports failure by returning -1 and setting an error number, and perror is the library function that turns that number into the sentence the program printed. Every system call in this book follows that rule, and Chapter twenty one shows the one famous exception, fork, which has three possible answers rather than two.

Distinctions that carry marks

straceA debugger
Showsthe boundary between program and kernelthe inside of the program
Needs the source?nousually
Slows the programgreatlygreatly
Answerswhat did it ask the system forwhat is it doing and why
Return value of a system callError number
On successthe useful answer: a descriptor, a countnot set meaningfully
On failure-1says which failure
Read withthe value itselferrno, printed by perror

What it does not mean

A trace is not the program's source code. It shows only what crossed the boundary. Everything the program computed between two calls is invisible here, and that is the point.

Timings under strace are not real timings. Every call stops the program twice.

Not every function in your program appears. strlen, malloc and most of printf are library code and make no system call at all, so they leave no trace.

Quick revision

  • strace lists every system call a program makes, with arguments and return value. -e trace=

selects calls, -o writes to a file, -c counts instead of listing.

  • A program under strace is much slower, so timings taken under it are worthless.
  • The first file a program opens usually gets descriptor 3: 0, 1 and 2 are standard input,

output and error.

  • read returns how much it actually read, which may be less than was asked for.
  • open in a program appears as openat in the trace, with AT_FDCWD for the current

directory.

  • A system call reports failure as -1 plus an error number, such as ENOENT. perror prints

it.

  • Even /bin/true makes about twenty seven calls: that is the cost of starting any program.

Test yourself

  1. What does strace show, and what does it not show? Every system call a program makes,

with arguments and results. It shows nothing that happens inside the program between calls.

  1. Why is descriptor 3 the first one a program usually gets? Because 0, 1 and 2 are already

open as standard input, standard output and standard error, and the kernel gives out the lowest free number.

  1. read(3, buf, 32) returns 18. What happened, and what must the program do? Eighteen bytes

were available and were read. The program must use the return value rather than assume it got 32, and must call read again if it needs more.

munotes.in41

Watching System Calls Happen

  1. A program says it cannot find a file. How would you find out which name it looked for?

Trace it and look at the openat line: the trace shows the exact string and the error, for example -1 ENOENT.

  1. How does a system call report failure? It returns -1 and sets an error number, which

perror or strerror turns into a message.

  1. Why does strlen never appear in a trace? It is library code that computes in user mode

and makes no request of the kernel.

Contents This chapter on its own page

munotes.in42

Chapter Eleven

The Six Families of System Call

Syllabus topic Module 1, "Fundamentals of Operating systems - Types of System Calls"

In one line

System calls fall into six families: process control, file management, device management, information maintenance, communications, and protection.

Why this grouping and not another

The six families are the six functions of Chapter five seen from the program's side. That is the way to remember them: for each thing the operating system manages, there is a family of calls that asks it to manage that thing. Process management becomes process control; file management becomes file management; and so on.

The six families

1. Process control

Create and end processes, load and run a program, wait for one to finish, get and set a process's attributes, allocate and free memory.

CallWhat it does
forkmake a copy of this process (Chapter nineteen)
execvereplace this process's program (Chapter twenty)
waitwait for a child to finish (Chapter twenty one)
exitend this process
killsend a signal to a process
brk, mmapget more memory

2. File management

Create, delete, open, close, read, write and reposition a file, and get or set its attributes.

CallWhat it does
open, openatopen a file, giving a descriptor
read, writemove bytes
lseekmove the file position
closegive the descriptor back
statread a file's attributes
unlinkremove a name

3. Device management

Ask for a device, release it, read and write it, and set its attributes. On a Unix system this family looks like the last one, and that is deliberate: a device is reached through a file name, so /dev/null is opened, read and written exactly as an ordinary file is. The one call that is peculiar to devices is ioctl, which carries requests that do not fit reading and writing, such as setting the speed of a serial line.

4. Information maintenance

Ask the system the time, the date, the number of users, the version of the kernel, or a process's own attributes; and set those that may be set.

CallWhat it does
getpidwhich process am I
unamewhat kernel is this
clock_gettimewhat is the time
getrusagehow much have I used

5. Communications

Two models, and Chapters twenty two to twenty six are this family.

CallModel
pipemessage passing, between related processes
mq_open, mq_sendmessage passing, by queue
shmget, shmatshared memory
socket, connectmessage passing, over a network

6. Protection

Set and get the permissions on a file, and ask who the running program is acting as.

CallWhat it does
chmodchange a file's permissions
umaskset the default permissions for new files
getuid, setuidwhich user is this program acting as
accessmay I do this to this file
munotes.in43

The Six Families of System Call

One from each family, run

Six programs would be six pages. One program makes one call from each family, and the trace shows all six.

#include <stdio.h>
#include <string.h>
#include <sys/utsname.h>
#include <sys/wait.h>
#include <unistd.h>
#include <fcntl.h>
#include <sys/stat.h>

int main(void)
{
    /* 1. process control */
    pid_t child = fork();
    if (child == 0) {
        _exit(0);
    }
    wait(NULL);

    /* 2. file management */
    int fd = open("six.txt", O_CREAT | O_WRONLY | O_TRUNC, 0644);
    write(fd, "six\n", 4);
    close(fd);

    /* 3. device management: a device is opened like a file */
    int null = open("/dev/null", O_WRONLY);
    write(null, "into the void\n", 14);
    close(null);

    /* 4. information maintenance */
    struct utsname u;
    uname(&u);

    /* 5. communications */
    int ends[2];
    pipe(ends);
    write(ends[1], "ping", 4);
    char got[8] = {0};
    read(ends[0], got, 4);
    close(ends[0]);
    close(ends[1]);

    /* 6. protection */
    chmod("six.txt", 0600);

    printf("kernel %s, pipe carried %s, all six families used\n", u.sysname, got);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o sixfamilies sixfamilies.c
$ ./sixfamilies
kernel Linux, pipe carried ping, all six families used
$ strace -e trace=clone,wait4,openat,uname,pipe2,fchmodat -o trace6.txt ./sixfamilies > /dev/null
$ grep -E 'clone|wait4|uname|pipe2|chmod|/dev/null' trace6.txt | sed 's/0x[0-9a-f]*/ADDRESS/g'
clone(child_stack=NULL, flags=CLONE_CHILD_CLEARTID|CLONE_CHILD_SETTID|SIGCHLD, child_tidptr=ADDRESS) = 27
wait4(-1, NULL, 0, NULL)                = 27
openat(AT_FDCWD, "/dev/null", O_WRONLY) = 3
uname({sysname="Linux", nodename="ubuntu", ...}) = 0
pipe2([3, 4], 0)                        = 0
fchmodat(AT_FDCWD, "six.txt", 0600)     = 0
$ ls -l six.txt
-rw------- 1 student student 4 Sep 30 09:37 six.txt

Four names in the program are different names in the trace, and that is the lesson to take from this chapter beside the list. It is also a warning about strace: asking it to trace chmod shows nothing at all, because the call the machine makes is fchmodat.

Written in the programAppears in the traceWhy
forkcloneLinux has one call that makes both processes and threads, and fork is a particular set of its flags
waitwait4the library calls the more general version
pipepipe2the newer call takes flags; the old one is still there
chmodfchmodatthe newer call takes a directory to be relative to

The permissions at the end are the protection family's work: the file was made with 0644 and ended as 0600, which ls -l shows as -rw-------.

Distinctions that carry marks

FamilyThe question it answersOne call
Process controlwho is running, and whatfork
File managementwhere is the dataopen
Device managementhow do I reach the hardwareioctl
Information maintenancewhat is the state of the systemgetpid
Communicationshow do two processes talkpipe
Protectionwho is allowedchmod
File managementDevice management
On Unixopen, read, write, closethe same calls, on a name under /dev
The peculiar callnoneioctl, for requests that are not reading or writing
Why they look alikea device is presented as a file, so one interface serves both
munotes.in44

The Six Families of System Call

What it does not mean

Six families does not mean six hundred calls. Linux has a few hundred in all, and most programs use twenty.

"Device management" does not mean writing a driver. The driver is inside the kernel. This family is how a program asks the driver for something.

Information maintenance is not a small family. It is where a great deal of a modern system's work lives, because so much of what a program wants to know about its environment is a call.

Quick revision

  • Six families: process control, file management, device management,

information maintenance, communications, protection.

  • They mirror the six functions of Chapter five, seen from the program's side.
  • On Unix a device is opened by name under /dev and read and written like a file; ioctl

carries what does not fit.

  • The name in your program may not be the name of the system call: fork is clone, wait is

wait4, pipe is pipe2, open is openat.

  • One call from each family: fork, open, ioctl, getpid, pipe, chmod.

Test yourself

  1. Name the six types of system call, with one example of each. Process control (fork),

file management (open), device management (ioctl), information maintenance (getpid), communications (pipe), protection (chmod).

  1. Why do file and device calls look the same on Unix? Because a device is presented as a

name in the file system, so one interface serves both and a program need not know which it has.

  1. What is ioctl for? Requests to a device that are not reading or writing, such as setting

the speed of a serial line or the size of a terminal.

  1. You wrote fork and the trace says clone. Is that an error? No. Linux implements

process and thread creation with one call, clone, and fork is a particular set of its flags. The same happens to wait, which is wait4, to pipe, which is pipe2, to open, which is openat, and to chmod, which is fchmodat.

  1. Which family does getrusage belong to and what does it report? Information maintenance.

It reports the resources a process has used, which is the accounting service of Chapter seven.

  1. A program wants to know whether it may read a file before trying. access, in the

protection family. It answers the question without opening the file.

Contents This chapter on its own page

munotes.in45

Chapter Twelve

How an Operating System Is Built

Syllabus topic Module 1, "Fundamentals of Operating systems - Operating-System Structure"

In one line

An operating system is a large program, and how its inside is divided decides how easy it is to change, how fast it runs, and how badly one mistake hurts.

Why structure is a question at all

A kernel is hundreds of thousands of lines long, written by many people over decades, and it may not crash. Those three facts together make its internal division a real engineering problem, and the structures below are four answers to it, in the order they were tried.

Every one of them is a trade between two things that cannot both be had: a clean division costs a call across the division, and a call across the division costs time. Hold that sentence and the whole topic follows.

Simple structure

No division at all. Everything can call everything, and there is no boundary between the parts, nor always between the operating system and the programs.

MS-DOS is the standard example. It was written for a machine with no mode bit and very little memory, so there was nowhere to put a boundary and nothing to enforce it with. A program could call the disk routines directly, and could also overwrite them.

Gainedvery little memory used, very fast
Lostno protection at all; one bad program could take the machine down
Fit fora machine with one program and no hardware support

Monolithic structure

The whole operating system is one program, running in kernel mode, in one address space. There is a system call boundary between it and user programs, and no boundary inside it.

The original UNIX is the example, and it is the example MU's text book uses. Ritchie and Thompson's own account of it describes a system of two parts, the kernel and everything else, and its size is worth repeating: on the PDP-11/45 the kernel occupied 42 kilobytes.

Gainedspeed: a call from one part of the kernel to another is an ordinary function call
Losta mistake anywhere in the kernel can damage anything in the kernel
Hard tochange, because everything may depend on everything

"Monolithic" does not mean badly organised. A monolithic kernel usually has very clear internal structure. What it lacks is an enforced boundary: nothing stops one part reaching into another, so the structure is a convention.

Layered structure

The operating system is cut into layers, numbered from the hardware upwards. Layer n may use only layer n minus one. The bottom layer is the hardware and the top layer is the user interface.

The gain is that each layer can be built and debugged with everything below it already known to work. If layer 3 misbehaves, the fault is in layer 3, because layers 0 to 2 were already proved.

munotes.in46

How an Operating System Is Built

Gainedeach layer is built on proved ground; a fault is localised
Lostspeed: a request from the top passes through every layer, and each one copies and checks
Lostit is genuinely hard to decide what goes in which layer

The second loss is why pure layering was abandoned, and the reason is worth knowing because it is a favourite question. Consider which layer holds the disk driver and which holds the memory manager. The memory manager needs the disk, to swap pages out. The disk driver needs memory, for its buffers. Each needs the other, so neither can be below the other, and a pure hierarchy is impossible.

Later systems kept a small number of layers rather than many, which is why you will read that layering "influenced" modern designs rather than that it is used.

The size of the boundary, measured

Whatever the inside looks like, the boundary a program sees is the system call interface, and it is worth knowing how wide that door is.

$ man 2 syscalls 2>/dev/null | grep -cE '^ +[a-z_0-9]+\(2\)'
484
$ ls /usr/include/asm-generic/unistd.h
/usr/include/asm-generic/unistd.h
$ grep -c '^#define __NR_' /usr/include/asm-generic/unistd.h
355

A few hundred numbered requests, listed in one header file, is the whole interface between every program on the machine and the operating system. Everything in the rest of this book happens on the far side of it.

Distinctions that carry marks

SimpleMonolithicLayered
Boundary inside the kernelnonenonebetween every pair of layers
Boundary at the system callnoneyesyes
Speedfastestfastslower
A fault in one partcan damage anythingcan damage the kernelis localised
ExampleMS-DOSoriginal UNIXTHE, and parts of modern systems

What it does not mean

"Structure" is not about source files. Every kernel is organised into files and directories. The question is which parts may call which, and whether anything enforces it.

Layering is not dead. Almost every modern kernel is layered in places, for example the file system sitting on the block layer sitting on the driver. What was abandoned is the rule that the whole kernel must be a strict hierarchy.

A monolithic kernel is not a small kernel. It is one address space, and Linux's is millions of lines.

Quick revision

  • Simple structure: no boundaries at all, MS-DOS. Fast, tiny, no protection.
  • Monolithic: one program in one address space in kernel mode, original UNIX. Fast

internally, a fault anywhere can damage anything, hard to change.

  • Layered: layer n uses only layer n minus one. Each layer built on proved ground, a fault is

localised, but every request passes through every layer and the ordering problem is real: the memory manager needs the disk and the disk driver needs memory.

munotes.in47

How an Operating System Is Built

  • Every structure is a trade: a clean division costs a call across the division.
  • The system call interface is a few hundred numbered requests, and that is the whole door.

Test yourself

  1. What is a monolithic kernel? One in which the whole operating system is a single program

running in kernel mode in one address space, with a system call boundary to user programs and no enforced boundary inside.

  1. Give the advantage and the disadvantage of a layered structure. Each layer is built and

debugged on layers already known to work, so a fault is localised. Every request passes through every layer, which costs time, and deciding which layer a service belongs in is genuinely hard.

  1. Why can a kernel not be strictly layered? Because some pairs of services need each other:

the memory manager needs the disk driver to swap, and the disk driver needs memory for its buffers, so neither can be below the other.

  1. What did MS-DOS lack, and why? Any protection boundary. The hardware it was written for

had no mode bit, so there was nothing to enforce one with.

  1. Does "monolithic" mean disorganised? No. It means there is no enforced boundary inside the

kernel. The organisation may be excellent, but it is a convention rather than a rule.

Contents This chapter on its own page

munotes.in48

Chapter Thirteen

Microkernels, Modules and What Linux Actually Is

Syllabus topic Module 1, "Fundamentals of Operating systems - Operating-System Structure"

In one line

A microkernel keeps almost nothing in the kernel and puts the rest in ordinary processes; a modular kernel keeps one kernel but loads and unloads parts of it while running; and every real system is a mixture.

The microkernel

Take the monolithic kernel of the last chapter and remove everything that does not absolutely have to be in kernel mode. What is left is the microkernel, and it usually keeps only three things:

  1. minimal process management,
  2. minimal memory management,
  3. communication between processes.

Everything else, the file system, the device drivers, the network stack, becomes an ordinary server process running in user mode. A program that wants a file sends a message to the file server. The microkernel's job is to carry the message.

This is why communication is the one thing a microkernel must do well. In a monolithic kernel, a request for a file is a function call. In a microkernel it is a message from the program to the kernel, then from the kernel to the file server, then an answer back the same way: four crossings of the boundary instead of two. That is the whole cost, and it is also the whole benefit.

Gaineda driver that crashes takes down one process, not the machine
Gaineda new service is added without touching the kernel
Gainedthe kernel is small enough to be reasoned about, and in some cases proved correct
Gainedthe same kernel ports to new hardware more easily
Lostspeed, because message passing replaces function calls

Mach is the example every book gives, and macOS's kernel is built on it. QNX is the one used where reliability matters more than speed, in cars and medical equipment. MINIX 3 was written to show that a microkernel can be a real operating system.

The performance loss is real but is often overstated in old books. Later microkernels closed much of the gap by making message passing extremely cheap. What has not changed is the shape of the argument.

The modular kernel

The idea that won, and it is the structure of Linux, Solaris, macOS and Windows alike. Keep a single kernel, in one address space, for speed. But build it out of modules that can be loaded and unloaded while the machine is running, and let each module talk to the others through a defined interface.

A module is not a process. It is kernel code, running in kernel mode, in the kernel's own address space. What makes it a module is that it is not present until it is needed and can be removed.

Microkernel serverKernel module
Runs inuser modekernel mode
Isan ordinary processpart of the kernel
Talks bymessagesfunction calls
If it crashesone process diesthe machine may go down
Loaded whenstarted like any programneeded, and unloaded after
munotes.in49

Microkernels, Modules and What Linux Actually Is

The gain over pure layering is that a module may call any other module directly, so there is no hierarchy to get wrong. The gain over a fixed monolithic kernel is that a machine only carries the drivers for the hardware it has.

$ head -5 /proc/filesystems
nodev	sysfs
nodev	tmpfs
nodev	bdev
nodev	proc
nodev	cgroup
$ cat /proc/sys/kernel/osrelease
6.8.0-117-generic

The kernel publishes the list of file system types it is able to mount, and every entry on it is a part of the kernel that is either built in or was loaded as a module. Dozens of kinds of file system from one kernel is what modularity buys: a machine carries the ones it uses and no more.

The lab machine is a container, so it shares its host's kernel and cannot load or unload a module of its own. What it can show is the mechanism's own rule, which is that a module is built for one kernel version: the version is published in that last file, and modules for a kernel are kept in a directory named after it.

The hybrid, which is what everything actually is

No shipped operating system is purely any of the five structures. Each one picks what it wants:

SystemWhat it really is
Linuxmonolithic and modular: one address space, loadable modules
Windowsa hybrid: a microkernel-influenced core, with subsystems, and drivers in kernel mode
macOS and iOSa Mach microkernel and a BSD monolithic kernel in one, plus loadable extensions
Androidthe Linux kernel, with its own libraries and its own virtual machine above it

So the honest answer to "what structure is Linux" is "monolithic with loadable modules", and to "what structure is macOS" is "hybrid". An examination answer that says any real system is a pure microkernel is wrong.

Worked example: adding support for a new device, three ways

A new kind of camera is to be supported.

In a monolithic kernel with no modules: the driver is added to the kernel's source, the kernel is rebuilt, and the machine is restarted. Everybody who wants the camera runs a kernel containing the driver, whether they have a camera or not.

In a microkernel: the driver is written as an ordinary program. It is started like any other program. If it is wrong, it crashes, and nothing else does. The kernel is untouched.

In a modular kernel: the driver is built as a module against this kernel's version. It is loaded when the camera is plugged in and unloaded when it is removed. It runs in kernel mode, so a serious fault in it can still bring the machine down, but no machine carries it unless it has the hardware.

munotes.in50

Microkernels, Modules and What Linux Actually Is

The third is the compromise the industry chose: the speed of the second chapter's monolithic kernel with most of the flexibility of the microkernel, and none of its protection.

Distinctions that carry marks

MonolithicMicrokernelModular
In the kerneleverythingprocess, memory, communication onlya core plus loaded modules
Drivers run inkernel modeuser modekernel mode
A driver faultmay crash the machinekills one processmay crash the machine
A request costsa function callfour boundary crossingsa function call
Add a servicerebuild the kernelstart a programload a module
Exampleoriginal UNIXMach, QNX, MINIX 3Linux, Solaris

What it does not mean

A microkernel is not a small monolithic kernel. The difference is not size, it is where the services run: in user mode as processes, or in kernel mode.

A module is not a user program. It is kernel code and has all the kernel's power.

"Hybrid" is not a way of avoiding the question. It is the accurate answer, and naming what each system took from where is the marks.

Quick revision

  • A microkernel keeps only minimal process management, minimal memory management and

communication; everything else runs as a user mode server process.

  • The cost is four boundary crossings instead of two; the benefit is that a driver fault kills

one process. Mach, QNX, MINIX 3.

  • A modular kernel is one address space with parts that load and unload while running. A

module runs in kernel mode and is not a process. Linux, Solaris.

  • A module is built for one kernel version, which is why modules live in a directory named after

the version.

  • Every real system is a hybrid. Linux is monolithic with loadable modules; macOS is Mach

plus BSD plus extensions; Windows is a layered hybrid with drivers in kernel mode.

Test yourself

  1. What does a microkernel keep, and what does it move out? It keeps minimal process

management, minimal memory management and inter-process communication. File systems, drivers and the network stack move out into user mode server processes.

  1. Why must a microkernel's message passing be fast? Because every service request becomes

messages through the kernel: four crossings of the user and kernel boundary instead of two.

  1. Give the main advantage and the main disadvantage of a microkernel. A failing driver kills

only its own process, and services can be added without touching the kernel. The cost is performance.

  1. Distinguish a kernel module from a microkernel server. A module is kernel code in kernel
munotes.in51

Microkernels, Modules and What Linux Actually Is

mode in the kernel's address space, called as a function. A server is an ordinary user mode process reached by messages.

  1. What structure is Linux? Monolithic with loadable modules. One kernel address space for

speed, with parts loaded and unloaded while the machine runs.

  1. Where does a kernel publish what it can do? In /proc. The list of file system types it

can mount is one such list, and every entry is either built in or was loaded as a module.

  1. Why is a module built for one particular kernel version? Because it becomes part of that

kernel and uses its internal interfaces, which are not promised to stay the same. That is why modules are stored in a directory named after the kernel's version.

Contents This chapter on its own page

munotes.in52

Chapter Fourteen

What a Process Is

Syllabus topic Module 1, "Processes - Process Concept"

In one line

A program is a file on the disk. A process is that program actually running, with memory of its own and a place in the operating system's tables.

The form to write: a process is a program in execution. It is an active entity, with a program counter, a stack, and its own data, while a program is a passive entity: a file of instructions stored on disk.

Why the two words are not the same thing

The distinction is the first examinable idea in this module and it is worth being precise about, with three consequences that make it concrete.

  1. One program can be many processes. Two students logged in and both running the editor are

two processes from one file. They have different memory, different positions in their file, different everything except the instructions.

  1. A process can outlive the file it came from. Delete the program while it runs and it keeps

running: the kernel holds what it needs.

  1. A process has a state that a program does not have. Where it has got to, what it holds,

what it is waiting for. None of that exists in the file.

The useful test: if you can copy it, e-mail it or delete it, it is a program. If it has a number, an owner and a parent, it is a process.

What a process is made of

A process has four sections of memory and one number, and a question about it will ask for the four.

SectionHoldsFixed size?
Textthe instructions themselvesyes, and usually read only
Datavariables that exist for the whole runyes
Heapmemory asked for while runninggrows and shrinks
Stackone frame per function call: arguments, local variables, the return addressgrows and shrinks

Besides the memory, a process has the contents of the processor's registers, and the most important of those is the program counter, which holds the address of the next instruction. That single register is what "where it has got to" means.

The heap and the stack grow towards each other. They are put at opposite ends of the space the process has, precisely so that each can grow without a fixed boundary limiting it. If they meet, the process is out of memory.

The four sections, on a real process

The kernel publishes the map of every process's memory. Here is a running program's own map, shortened to the lines that matter.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>

int global_variable = 42;                 /* data */

int main(void)
{
    int local_variable = 7;               /* stack */
    char *asked_for = malloc(1000);       /* heap */

    printf("code   below global : %s\n",
        (void *)main < (void *)&global_variable ? "yes" : "no");
    printf("global below heap   : %s\n",
        (void *)&global_variable < (void *)asked_for ? "yes" : "no");
    printf("heap   below stack  : %s\n",
        (void *)asked_for < (void *)&local_variable ? "yes" : "no");
    printf("stack  %p\n", (void *)&local_variable);
    free(asked_for);
    return 0;
}
munotes.in53

What a Process Is

$ gcc -std=c17 -Wall -Wextra -o showmap showmap.c
$ ./showmap | head -3
code   below global : yes
global below heap   : yes
heap   below stack  : yes

Those three answers are yes on every machine and on every run. The code is lowest, the global just above it, the heap higher still, and the local variable on the stack at a completely different, much larger address. That gap is the room the heap and the stack have to grow into, and it is why the two are at opposite ends.

The addresses themselves are not a fact about every run, and the program can be made to prove it.

$ a=$(./showmap | awk '/^stack/ {print $2}')
$ b=$(./showmap | awk '/^stack/ {print $2}')
$ [ "$a" != "$b" ] && echo "the stack was at a different address on the two runs"
the stack was at a different address on the two runs

The program was run twice and its stack was somewhere else each time. The kernel deliberately puts a program's pieces at random addresses, a defence called address space layout randomisation, so that an attacker who knows a program's code cannot know where anything is. That is why this book prints the order of the addresses rather than the addresses: the order is a fact about every run, and the numbers are a fact about one.

The kernel's own map of a process says the same thing with names instead of numbers. Every line is one region of memory: its address range, its permissions, and what it is.

$ awk '/\[heap\]|\[stack\]/ || (/r-xp/ && !seen++) {print $2, $6}' /proc/self/maps
r-xp /usr/bin/gawk
rw-p [heap]
rw-p [stack]

r-xp is readable and executable but not writable: that is the text section, and the kernel enforces the read only rule with the memory management unit of Chapter sixty eight. [heap] and [stack] are named by the kernel itself, and their addresses are as far apart as the four numbers above suggested.

That map belongs to the awk that printed it, because a process can only read its own. The shape is the same for every process on the machine.

Worked example: one program, three processes

Sunita opens three windows and runs the same calculator in each.

There is one file on the disk, /usr/bin/calculator. There are three processes. Each has:

  • its own process id, a number the kernel gives it;
  • its own data, heap and stack, so the number typed in one window is not in the
munotes.in54

What a Process Is

others;

  • its own program counter, so one may be halfway through a multiplication while another waits

for a keystroke;

  • its own entry in the kernel's tables.

They share one thing: the text. The instructions are identical and are read only, so the kernel maps the same physical memory into all three. That is a saving of two thirds and it is done without the processes knowing; Chapter eighty three is the general mechanism.

A favourite examination trap: whether two processes running the same program share their variables. They do not. They share the instructions. Every variable is private.

Distinctions that carry marks

ProgramProcess
Ispassive: a fileactive: a file being executed
Lives onthe diskin memory, with kernel bookkeeping
Has a program counternoyes
Has a statenoyes: new, ready, running, waiting, terminated
How many of itone fileas many processes as you start
Can be deleted while in useyes, the process continuesending it is killing it
SectionGrows?Shared between two processes of the same program?
Textnoyes, read only
Datanono
Heapyesno
Stackyesno

What it does not mean

A process is not a program with extra features. It is a different kind of thing: the program is the recipe and the process is the cooking.

The stack is not "the memory". It is one of four sections, and it holds only the frames of the function calls currently in progress. A local variable disappears when its function returns, which is exactly why returning a pointer to one is a bug.

A process id is not a program's identity. It is given out by the kernel when the process starts and is reused after it ends. The same program gets a different number every time.

Quick revision

  • A program is a passive file; a process is a program in execution, an active entity with

a program counter, a stack and its own data.

  • Four sections of memory: text (instructions, read only), data (variables for the whole

run), heap (asked for while running), stack (one frame per call).

  • The heap and the stack grow towards each other from opposite ends, so neither has a fixed

boundary.

  • One program can be many processes; they share the text and share nothing else.
  • A process has a process id, an owner and a parent. A program has none of those.
  • Deleting the file does not stop the process.

Test yourself

  1. Define a process, and distinguish it from a program. A process is a program in execution:

an active entity with a program counter, a stack and its own data. A program is passive, a file of instructions on disk.

munotes.in55

What a Process Is

  1. Name the four sections of a process's memory and say what each holds. Text, the

instructions; data, variables that last the whole run; heap, memory asked for while running; stack, a frame per function call holding arguments, locals and the return address.

  1. Why are the heap and the stack at opposite ends of the address space? So that each can

grow without a fixed boundary between them. They meet only when the process is out of memory.

  1. Three windows run the same program. What is shared and what is not? The text is shared,

read only, as one copy in physical memory. Everything else, data, heap, stack, registers, process id, is separate.

  1. A local variable's address is much larger than a global's. Why? The local is on the stack,

which is placed at the high end of the address space, and the global is in the data section, just above the code at the low end.

  1. Is a process id a permanent name for a program? No. The kernel gives it out when the

process starts and may reuse it after the process ends.

Contents This chapter on its own page

munotes.in56

Chapter Fifteen

The Five States of a Process

Syllabus topic Module 1, "Processes - Process Concept"

In one line

A process is in exactly one of five states at a time: new, ready, running, waiting, or terminated.

Why five and not two

The obvious division is running and not running. It is not enough, because not running covers two completely different situations and the operating system must treat them differently.

  • A process that is ready could run this instant. All it needs is the processor.
  • A process that is waiting could not use the processor if you gave it to one. It is waiting

for a disk, a keystroke, another process.

Giving the processor to a waiting process would waste it, so the scheduler chooses only from the ready ones. That single fact is why the two states are separate, and it is the answer to the question "why distinguish ready from waiting".

The remaining three are the beginning, the end, and the moment of being on the processor.

The five states

StateMeaning
Newthe process is being created: the kernel is building its records and its memory
Readyit can run, and is waiting for the processor
Runningits instructions are being executed
Waiting, also called blockedit is waiting for an event, usually input or output
Terminatedit has finished; its records are being cleaned up

On a machine with one processor, exactly one process is running. Many may be ready and many may be waiting. That is the whole of what "multiprogramming" means and it is worth saying out loud, because a diagram with three boxes tempts a reader into thinking three things run at once.

The transitions, each with its event

There are exactly six ways to move between the states, and naming the event is where the marks are.

FromToThe event that causes itWho acts
NewReadythe kernel has finished making the process, and admits itthe long term scheduler
ReadyRunningthe process is chosen and given the processorthe short term scheduler, through the dispatcher
RunningReadyits quantum expired, or something more urgent became readythe timer interrupt, then the scheduler
RunningWaitingit asked for something that is not ready: a file, a keystroke, a lockthe process itself, by a system call
WaitingReadythe thing it waited for happeneda device interrupt
RunningTerminatedit finished, or was killeditself, or another process

Notice two transitions that DO NOT EXIST, and this is a standard two marks.

  • Waiting to Running is impossible. A process whose event has happened becomes ready, and

must then be chosen like everybody else. It does not jump onto the processor.

  • Ready to Waiting is impossible. A process cannot begin waiting for something without
munotes.in57

The Five States of a Process

running long enough to ask for it.

Written as a diagram, in words

The diagram to draw has five boxes. New at the top left with an arrow marked admit to Ready. Ready and Running side by side in the middle, with dispatch from Ready to Running and interrupt back the other way. Waiting below them, with input or output wait from Running down to Waiting and input or output complete from Waiting back up to Ready. Terminated at the top right, with exit from Running.

If you can label all six arrows you have the question. The two arrows that do not exist are worth a sentence underneath.

The states a real machine reports

Linux does not use those five words. It reports a letter per process, and the letters are worth knowing because a practical question will show you ps output.

$ set +m
$ sleep 30 & s=$!
$ yes > /dev/null & y=$!
$ sleep 0.3
$ ps -o pid,stat,comm -p $s -p $y --no-headers
     13 S+   sleep
     15 R+   yes
$ kill -STOP $y 2>/dev/null
$ sleep 0.2
$ ps -o pid,stat,comm -p $y --no-headers
     15 T+   yes
$ kill -CONT $y 2>/dev/null
$ kill $s $y 2>/dev/null
$ wait 2>/dev/null

Two processes were started on purpose. sleep is waiting for time to pass, so it prints S. yes computes without stopping, so it prints R. Then yes was sent a stop signal and printed T, and was let go again.

LetterLinux calls itWhich of the five
Rrunning or runnableRunning or Ready: Linux does not distinguish them in this column
Sinterruptible sleepWaiting
Duninterruptible sleep, usually a diskWaiting
TstoppedWaiting, in effect: it was suspended
ZzombieTerminated but not yet cleaned up (Chapter twenty one)

The suffixes matter less but are asked: s means the process leads a session, + means it is in the foreground, l means it has more than one thread.

R covers two of the five states. A process on the processor and a process merely able to run both print R, because from the kernel's point of view they are in one queue and only one of them happens to be executing. A student who expects a separate letter for ready will not find one.

D is the state to recognise in real life. A process in uninterruptible sleep cannot even be killed, because it is inside the kernel waiting for hardware. If a machine has a process stuck in D, the disk or the network is the problem, not the program.

Worked example: following one process through all five

Ramesh types sort bigfile.txt at the prompt.

munotes.in58

The Five States of a Process

  1. New. The shell asks the kernel to make a process. The kernel allocates a process id,

builds its records, and sets up its memory. The process exists but has not executed an instruction.

  1. Ready. The kernel admits it. It joins the queue of processes that could run.
  2. Running. The scheduler chooses it. It executes: it asks to open the file.
  3. Waiting. The file's first block is not in memory, so the kernel starts a disk read and the

process can do nothing until it arrives. It moves to waiting, and the scheduler gives the processor to somebody else. This is the step students skip, and it is the reason the state exists.

  1. Ready. The disk interrupts. The block has arrived. The process becomes ready.
  2. Running. Chosen again, it sorts what it has.
  3. Steps 4 to 6 repeat for every block of the file, perhaps thousands of times.
  4. Running to Ready whenever its quantum expires, perhaps hundreds of times.
  5. Terminated. It has printed the sorted lines and calls exit. The kernel keeps its exit

status until the shell collects it, then removes its records.

Count the transitions in that story: a program that looks like one thing happening moved between states thousands of times.

Distinctions that carry marks

ReadyWaiting
Needsthe processoran event
Would the processor help?yes, immediatelyno, it could not use it
Chosen by the scheduler?yesnever
Leaves whenit is dispatchedits event happens, and then it becomes ready
Running to ReadyRunning to Waiting
Caused bythe timer, or a more urgent processthe process's own request
Voluntary?noyes
Calledpreemptionblocking

What it does not mean

"Waiting" does not mean waiting for the processor. That is ready. Waiting means waiting for something else.

A terminated process is not immediately gone. Its exit status has to be collected by its parent first, which is why the zombie state exists.

New is not a long state. It lasts as long as it takes to build the records, unless the system is so loaded that the long term scheduler refuses to admit more work.

Quick revision

  • Five states: new, ready, running, waiting (or blocked), terminated.
  • Ready needs only the processor; waiting needs an event and could not use the processor.

That is why they are separate: the scheduler chooses only from the ready.

  • Six transitions: admit, dispatch, interrupt, input or output wait, input or output complete,

exit.

  • Waiting to Running does not exist (it goes to ready first), and

Ready to Waiting does not exist (it must run to ask).

  • On one processor exactly one process is running.
  • Linux prints R for both running and ready, S for interruptible sleep, D for
munotes.in59

The Five States of a Process

uninterruptible sleep, T for stopped, Z for zombie.

  • A process in D cannot be killed: the hardware it is waiting for is the problem.

Test yourself

  1. Name the five states of a process. New, ready, running, waiting or blocked, and

terminated.

  1. Why are ready and waiting separate states? A ready process could use the processor at

once; a waiting one could not, because it is waiting for an event. The scheduler chooses only from the ready ones, so giving the processor to a waiting process would waste it.

  1. Name the six transitions with their events. New to ready on admit; ready to running on

dispatch; running to ready on interrupt or quantum expiry; running to waiting on an input or output request; waiting to ready on input or output completion; running to terminated on exit.

  1. Can a process go from waiting straight to running? No. It becomes ready and must be chosen

by the scheduler like any other ready process.

  1. ps shows R for a process. Which of the five states is it in? Either running or ready.

Linux does not distinguish them in that column.

  1. A process is stuck in state D and kill does nothing. What does that tell you? It is in

uninterruptible sleep inside the kernel, waiting for hardware. The device, not the program, is the problem.

Contents This chapter on its own page

munotes.in60

Chapter Sixteen

The Process Control Block

Syllabus topic Module 1, "Processes - Process Concept"

In one line

The process control block is the record the kernel keeps about one process, and it holds everything the kernel needs to take the processor away and give it back later without the process noticing.

Why it exists, and why that is the definition

A process is suspended hundreds of times a second. Each time, the processor is handed to somebody else and every register is overwritten. When the process runs again it must be exactly as it was: the same program counter, the same registers, the same open files, the same memory.

So the test of what belongs in the process control block is one question: would the process notice if this were lost? If yes, it is in the block. That single question generates the whole list and is better than memorising it.

The block is also called the task control block or, in Linux's own source, the task struct.

What is in it

FieldWhy it must be kept
Process statenew, ready, running, waiting, terminated (Chapter fifteen)
Process idthe number everything else refers to it by
Program counterthe address of the next instruction. Lose this and the process is destroyed
Registersaccumulators, index registers, stack pointers, general purpose registers
Scheduling informationits priority, which queue it is in, how much processor it has had
Memory management informationthe base and limit registers, or the page table (Chapter seventy four)
Accounting informationprocessor time used, real time elapsed, limits, the user it belongs to
Input and output statusthe list of open files, the devices it holds

The program counter is listed separately from the registers on purpose, in every textbook, because it is the one register whose loss cannot be recovered from. Everything else could in principle be recomputed; where you were cannot.

Reading one on a real machine

Linux publishes most of a process's control block as a file. Here is a real one, shortened to the fields in the table above.

$ set +m
$ sleep 60 & job=$!
$ sleep 0.2
$ grep -E '^(Name|Pid|PPid|State|Uid|Threads|VmSize|VmRSS|voluntary)' /proc/$job/status
Name:	sleep
State:	S (sleeping)
Pid:	13
PPid:	9
Uid:	1000	1000	1000	1000
VmSize:	    2696 kB
VmRSS:	    1696 kB
Threads:	1
voluntary_ctxt_switches:	1
$ ls /proc/$job/fd
0  1  2
$ cat /proc/$job/cmdline | tr '\0' ' '
sleep 60
$ kill $job 2>/dev/null
$ wait 2>/dev/null

Read it against the table.

  • State: S (sleeping) is the process state, and S is the waiting state of Chapter

fifteen.

  • Pid and PPid are the process id and the parent's.
  • Uid is the accounting and protection information: which user this process acts as.
  • VmSize and VmRSS are memory management information: how much address space it has, and
munotes.in61

The Process Control Block

how much of it is actually in physical memory.

  • Threads is 1, so this process has one thread (Chapter twenty six).
  • voluntary_ctxt_switches is the count of times it gave the processor up by itself, which is

scheduling information.

  • The directory of numbered files is its input and output status: the open descriptors 0, 1

and 2, which are standard input, standard output and standard error.

The program counter and the registers are not in that file, and they cannot be: they are the process's private state and publishing them would be a security hole. They live in the kernel's own copy of the block.

Where the blocks live

The kernel keeps every process control block in a table, and in Linux the table is a doubly linked list, so that the scheduler can walk it. The size of the table is the limit on how many processes the machine can have.

$ cat /proc/sys/kernel/pid_max
4194304
$ ps -e --no-headers | wc -l
5

The first number is the largest process id this kernel will give out, and the second is how many processes exist at this moment. On a quiet container there are a handful; on a desktop there are several hundred, and every one of them has a block.

Worked example: what a context switch actually copies

P1 is running and the timer fires. The kernel must switch to P2. Exactly this happens, and every step touches a control block.

  1. The processor traps into the kernel, which saves the program counter and the flags

automatically as part of taking the interrupt.

  1. The kernel's switch routine copies every remaining register into P1's control block.
  2. It copies P1's state from running to ready, and appends P1 to the ready queue.
  3. It updates P1's accounting: add the ticks P1 just used.
  4. It chooses P2 from the ready queue, using P2's scheduling information.
  5. It loads P2's memory management information into the hardware, which on a paged machine

means pointing the memory management unit at P2's page table. This is usually the most expensive step, because the translation cache must be emptied (Chapter seventy six).

  1. It copies every register out of P2's control block into the processor.
  2. It sets P2's state to running and returns from the interrupt, which restores P2's program

counter. P2 resumes at the instruction it was interrupted at, perhaps a second ago, and cannot tell.

Count what that cost: two blocks read or written in full, one queue operation, one page table switch, and pure overhead throughout. No user work is done during a context switch, which is why Chapter four's arithmetic about the quantum matters and why Chapter twenty seven's threads are attractive: two threads of one process share step 6, so switching between them is much cheaper.

munotes.in62

The Process Control Block

Distinctions that carry marks

Process control blockThe process's own memory
Belongs tothe kernelthe process
Readable by the processnoyes
Holdsstate, id, registers, scheduling, memory and accounting information, open filestext, data, heap, stack
Survives the processonly until the parent collects the exit statusno
The program counterThe other registers
Holdswhere the process iswhat the process was computing with
If lostthe process cannot be resumed at allthe computation is wrong
Saved bythe interrupt hardware, firstthe kernel's switch routine

What it does not mean

The control block is not the process. It is the kernel's record of it. The process is the memory and the execution.

A process cannot read its own control block directly. It can ask the kernel for parts of it, which is what getpid and getrusage do, and on Linux it can read the published parts through /proc.

Switching is not free because the block is small. Most of the cost is the memory management information in step 6, not the registers.

Quick revision

  • The process control block, also task control block, is the kernel's record of one process.
  • The test of what belongs in it: would the process notice if this were lost?
  • Fields: process state, process id, program counter, registers, scheduling

information, memory management information, accounting information, input and output status.

  • The program counter is listed separately because its loss is unrecoverable.
  • Linux publishes much of it as /proc/<pid>/status, and the open descriptors as

/proc/<pid>/fd. The registers are not published.

  • A context switch saves one block, chooses from the queue, loads another block, and switches the

memory mapping. All of it is overhead, and the memory mapping is usually the most expensive part.

Test yourself

  1. What is a process control block? The record the operating system keeps about a process,

holding everything needed to suspend it and resume it without the process noticing.

  1. List its contents. Process state, process id, program counter, registers, scheduling

information such as priority, memory management information such as the page table, accounting information, and input and output status including open files.

  1. Why is the program counter always named separately from the registers? Because it is the

one whose loss cannot be recovered from: without it there is no way to know where to resume.

  1. Give the test for whether something belongs in the block. Would the process notice if it

were lost? If so, it belongs.

  1. Which step of a context switch is usually the most expensive, and why? Switching the

memory management information, because the hardware's translation cache has to be emptied and refilled.

munotes.in63

The Process Control Block

  1. Can a process read its own control block? Not directly. It can ask the kernel for parts of

it with system calls, and on Linux read the published parts under /proc.

Contents This chapter on its own page

munotes.in64

Chapter Seventeen

The Queues, the Schedulers and the Context Switch

Syllabus topic Module 1, "Processes - Process Scheduling"

In one line

The kernel keeps its processes in queues, three schedulers move them between the queues, and moving the processor from one process to another is called a context switch.

Why there are queues at all

Chapter fifteen said a process is ready or waiting. That is a state, not a place. The kernel needs a place, because when the processor becomes free it must find a ready process in a few microseconds, and searching a table of four hundred processes is far too slow.

So each state has a queue, and a process's control block is linked into the queue for its state. Choosing the next process becomes taking the head of a list.

QueueHoldsHow many
Job queueevery process in the systemone
Ready queueevery process that is readyone, or one per priority, or one per processor
Device queueprocesses waiting for one particular deviceone per device

There is no single "waiting queue". A process waiting for the disk is in the disk's queue and a process waiting for the keyboard is in the keyboard's queue, because when the disk interrupts, the kernel must wake the processes waiting for the disk and nobody else.

The three schedulers

Three different decisions, on three different time scales, and a question will ask you to distinguish them.

Long term, or job schedulerShort term, or CPU schedulerMedium term
Decideswhich processes are admitted into memory at allwhich ready process gets the processor nextwhich process to swap out of memory for a while
Runsseldom: seconds or minutes apartvery often: every few millisecondswhen memory is short
Must bethoughtful; it has timeextremely fast; it is pure overheadthoughtful
Controlsthe degree of multiprogrammingnothing but the next few millisecondsthe degree of multiprogramming, downwards

The long term scheduler controls the degree of multiprogramming, which is the number of processes in memory at once. That phrase is examined, and the reason it matters is the next paragraph.

On a modern interactive system the long term scheduler barely exists. When you type a command, the process is admitted at once; nobody queues your editor for later. It is important in batch systems, where hundreds of jobs are submitted and the system decides how many to run together, and the idea still matters because admitting too many processes causes thrashing (Chapter ninety one).

The medium term scheduler is the one that removes a process from memory entirely and brings it back later. It is swapping (Chapter sixty nine), used as a scheduling decision rather than as a memory one.

Input and output bound against processor bound

The long term scheduler's real job is to keep a mixture, and the vocabulary is examined.

munotes.in65

The Queues, the Schedulers and the Context Switch

Input and output boundProcessor bound
Spends its timewaiting for devicescomputing
Its processor bursts areshort and frequentlong
Examplean editor, a database servera compiler, a video encoder

A system holding only one kind is badly used. All processor bound, and the disk is idle and the ready queue is long. All input and output bound, and the processor is idle while everybody waits. The long term scheduler tries for a mixture so that both are busy, and that sentence is a complete answer to a standard question.

The context switch

Switching the processor from one process to another means saving the old process's state into its control block and loading the new one's, exactly as Chapter sixteen's worked example sets out.

A context switch is pure overhead. During it the machine does no user work at all. Its cost is a few microseconds, and it depends on the hardware: a machine with several sets of registers can switch by changing which set is in use, and a machine with one set must copy.

The machine keeps its own count of switches, and so does every process.

$ awk '/^ctxt/ {print "context switches since boot:", $2}' /proc/stat
context switches since boot: 75038804
$ a=$(awk '/^ctxt/{print $2}' /proc/stat); sleep 1; b=$(awk '/^ctxt/{print $2}' /proc/stat)
$ echo "$((b - a)) context switches in one second on an idle machine"
4607 context switches in one second on an idle machine
$ vmstat 1 2 | tail -1 | awk '{print "vmstat agrees, cs =", $12}'
vmstat agrees, cs = 1326

Thousands a second with nothing happening. On a busy machine it is hundreds of thousands, and at some point the overhead becomes visible, which is why the number is worth watching.

A single process's own counts are published too, and they are worth reading because they separate the two reasons for switching.

$ set +m
$ sleep 5 & job=$!
$ sleep 0.2
$ grep ctxt /proc/$job/status
voluntary_ctxt_switches:	1
nonvoluntary_ctxt_switches:	1
$ wait 2>/dev/null
  • Voluntary switches are the process blocking on its own: Chapter fifteen's running to

waiting.

  • Nonvoluntary switches are preemption: running to ready, done to it by the timer.

Those two counts diagnose a program. A program with a huge voluntary count is waiting for devices and is input and output bound. A program with a huge nonvoluntary count is computing and is being preempted, so it is processor bound. The vocabulary of the previous section becomes two numbers you can read.

Worked example: the queues, followed

A machine has one processor. P1 is computing. P2 is waiting for the disk. P3 is ready. P4 has just been typed at a terminal.

munotes.in66

The Queues, the Schedulers and the Context Switch

  • Job queue: P1, P2, P3, P4. Everything.
  • Ready queue: P3.
  • Disk's device queue: P2.
  • Running: P1.

Now three things happen in turn.

  1. The timer fires and P1's quantum is over. P1 moves running to ready and joins the ready

queue behind P3. The short term scheduler takes P3. P3 runs. This is a nonvoluntary switch for P1.

  1. The disk finishes P2's read and interrupts. The kernel takes P2 out of the disk's device

queue and puts it in the ready queue. P2 does not start running: Chapter fifteen's rule.

  1. P3 asks to read a file. P3 moves running to waiting and joins the disk's device queue.

This is a voluntary switch. The scheduler chooses from the ready queue, which now holds P1 and P2.

Three events, four queue operations, two context switches, and no user work done during either switch.

Distinctions that carry marks

Long term schedulerShort term scheduler
Also calledjob schedulerCPU scheduler
Chooses fromnew processes on diskthe ready queue
Frequencyseconds or minutesmilliseconds
Controlsthe degree of multiprogrammingwhich process runs now
May take its timeyesno: it is overhead
Voluntary context switchNonvoluntary context switch
Caused bythe process blockingthe timer or a more urgent process
The process wasasking for somethingcomputing
Suggests the process isinput and output boundprocessor bound

What it does not mean

"Process scheduling" in MU's label is not the algorithms. It is the queues and the switching. The algorithms are her CPU Scheduling row.

The ready queue is not necessarily first in first out. It is a queue in the sense of a place to wait. Chapters forty seven to fifty three are seven different ways of choosing from it.

A context switch is not a system call. It is something the kernel does to a process, usually because of an interrupt.

Degree of multiprogramming is not the number of processors. It is the number of processes in memory.

Quick revision

  • Three queues: the job queue (everything), the ready queue (ready processes), and a

device queue per device. There is no single waiting queue.

  • Three schedulers: long term or job scheduler, admits processes and controls the

degree of multiprogramming; short term or CPU scheduler, chooses from the ready queue every few milliseconds; medium term, swaps a process out of memory for a while.

  • The long term scheduler wants a mixture of input and output bound and

processor bound processes, so that neither the devices nor the processor is idle.

  • A context switch saves one control block and loads another. It is pure overhead.
  • /proc/stat's ctxt counts every switch on the machine; /proc/<pid>/status separates a
munotes.in67

The Queues, the Schedulers and the Context Switch

process's voluntary switches (it blocked) from its nonvoluntary ones (it was preempted).

Test yourself

  1. Name the three kinds of queue and say what each holds. The job queue holds every process

in the system; the ready queue holds the processes ready to run; a device queue holds the processes waiting for one particular device.

  1. Why is there no single waiting queue? Because when a device interrupts, the kernel must

wake the processes waiting for that device and nobody else.

  1. Distinguish the long term from the short term scheduler. The long term, or job, scheduler

decides which processes are admitted into memory and so controls the degree of multiprogramming; it runs seldom and may take its time. The short term, or CPU, scheduler chooses which ready process runs next; it runs every few milliseconds and must be very fast, because it is overhead.

  1. What is the degree of multiprogramming and who controls it? The number of processes in

memory at once. The long term scheduler, admitting them, and the medium term scheduler, swapping them out. 5. Why does the long term scheduler want a mixture of processor bound and input and output bound processes? So that both the processor and the devices stay busy. All of one kind leaves the other resource idle. 6. A process shows 40,000 voluntary and 12 nonvoluntary context switches. What kind of process is it? Input and output bound: it is giving the processor up by itself, over and over, because it is waiting for devices rather than computing.

Contents This chapter on its own page

munotes.in68

Chapter Eighteen

Creating a Process: fork

Syllabus topic Module 1, "Processes - Operations on Processes"

In one line

fork makes a copy of the calling process, and then both the original and the copy return from it.

Why it is strange, and why it is designed that way

Every other function you have written returns once. fork returns twice: once in the process that called it, and once in a brand new process that did not exist when the call was made.

That is not a trick. It is the honest consequence of what the call does. The kernel copies the whole process, including the program counter, which was pointing at the instruction after the call. So the copy is a process that is exactly in the middle of returning from fork, and there is nothing else it could sensibly do but return.

The vocabulary: the process that called is the parent, the new one is the child. A child has one parent. A parent may have many children.

What is copied and what is shared

Copied for the childShared with the parent
Text (the instructions)not copied, shared read onlyyes
Data, heap, stackcopiedno
Program counter and registerscopiedno
Open file descriptorscopied, and they refer to the same open filesthe file position is shared
Process idno: the child gets a new oneno
Parent's process idthe child's parent id is the caller's idno

The file descriptor row is the one that surprises people and it is examined. The descriptors are copied, so both processes have descriptor 1. But they point at the same open file, so if both write, the writes interleave in one file, and if one moves the file position the other sees it move. Chapter ninety nine shows the table this happens in.

The return value, which is the whole call

fork returnsin which processmeaning
0in the childyou are the child
a positive numberin the parentthat is your child's process id
-1in the parentthe fork failed and there is no child

fork is the one call in this book with three possible answers rather than the usual two. The reason is that the child needs no id (it can ask with getpid) while the parent must be told its child's, or it could never wait for it.

The program that shows both sides

The first line of the program is not decoration. pid_t, the type a process id has, is a POSIX name and not part of standard C. Compiled with -std=c17, which asks for standard C and nothing else, the C library hides it, and the compiler says unknown type name 'pid_t'. #define _POSIX_C_SOURCE 200809L before the first include asks for the POSIX names as well. Every listing in this book that uses a POSIX type begins with it.

munotes.in69

Creating a Process: fork

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <sys/types.h>
#include <sys/wait.h>
#include <unistd.h>

int main(void)
{
    printf("before the fork, I am process %d\n", getpid());

    pid_t r = fork();

    if (r < 0) {
        perror("fork");
        return 1;
    }
    if (r == 0) {
        printf("child : fork returned %d, I am %d, my parent is %d\n",
            r, getpid(), getppid());
        printf("child : this line is printed by process %d\n", getpid());
        return 0;
    }
    wait(NULL);                           /* let the child finish and speak first */
    printf("parent: fork returned %d, I am %d\n", r, getpid());
    printf("parent: this line is printed by process %d\n", getpid());
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o twosides twosides.c
$ ./twosides
before the fork, I am process 21
child : fork returned 0, I am 22, my parent is 21
child : this line is printed by process 22
parent: fork returned 22, I am 21
parent: this line is printed by process 21

Read it carefully, because four things are in those four lines.

  1. "Before the fork" was printed once. There was one process then.
  2. The last line was printed twice, once by each process, because both of them reached it.
  3. The child's fork returned 0 and the parent's returned the child's id, 21.
  4. The child's parent id is the parent's id. The family relationship is recorded in the

kernel.

The wait at the end of the parent is not optional, and leaving it out was tried. Without it the parent often finished first, and the child's getppid() then answered 1 instead of 20, because a child whose parent has gone is given to the process the kernel keeps for the purpose. That is the orphan of the next chapter, and it appeared here by accident, which is the honest way to meet it.

The order of the last two lines is not guaranteed. Both processes are ready; which the scheduler runs first is its business. Run it twenty times and you will see both orders. A program that depends on the order is wrong, and Chapters thirty three to forty two are about that problem.

Counting: how many processes does a loop of forks make?

A favourite examination question, and the answer is a power of two.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <unistd.h>
#include <sys/wait.h>

int main(void)
{
    for (int i = 0; i < 3; i++) {
        if (fork() < 0) {
            perror("fork");
            return 1;
        }
    }
    printf("hello from %d\n", getpid());
    while (wait(NULL) > 0) {
        /* collect every child before finishing */
    }
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o howmany howmany.c
$ ./howmany | wc -l
8
munotes.in70

Creating a Process: fork

Eight lines, so eight processes, from three forks.

The reason, and the general rule. After the first fork there are 2 processes. Both reach the second fork, so there are 4. All four reach the third, so there are 8. n forks in a loop give 2 to the power n processes, of which 2 to the power n minus 1 are new children. Three forks give eight processes and seven children.

The mistake to avoid: answering "three children". The children fork too.

Worked example: how a shell runs a command

This is what step 3 of Chapter eight's shell loop really does.

  1. The shell reads ls.
  2. The shell calls fork. There are now two shells, identical.
  3. In the child (fork returned 0), the child calls exec to replace itself with ls. That

is the next chapter.

  1. In the parent (fork returned the child's id), the shell calls wait on that id, and

blocks. It moves to the waiting state of Chapter fifteen.

  1. ls runs and exits.
  2. The kernel wakes the shell. wait returns. The shell prints a new prompt.

Why fork and exec are two calls rather than one is the deepest question in this chapter, and it is asked. Because between step 2 and step 3 the child is a copy of the shell and can change things about itself before becoming ls: it can open a file and make it descriptor 1, so that the output goes to a file; it can close descriptors; it can change directory. All of shell redirection lives in that gap. One combined call to "run a program with these settings" would need a parameter for every setting anyone might ever want.

Distinctions that carry marks

ParentChild
fork returnsthe child's process id, a positive number0
getpid() givesits own idits own, different, id
getppid() givesits own parent's idthe parent's id
Memoryits owna copy
Open filesits ownthe same open files, through copied descriptors

What it does not mean

fork does not start the program again from the beginning. The child begins where the parent was, at the return from fork. Nothing above the call runs twice.

fork does not make the child run first, or second. Which runs first is the scheduler's business.

fork does not copy the memory immediately in a real system. Copying megabytes that are usually thrown away at once by exec would be waste, so the kernel marks it copied and copies only what is written to. That is copy on write, Chapter eighty three, and it is why fork is fast.

munotes.in71

Creating a Process: fork

A failed fork is not impossible. The system may be out of process slots or out of memory, and fork then returns -1. Every one of these programs checks.

Quick revision

  • fork creates a new process by copying the caller. It returns twice.
  • It returns 0 in the child, the child's process id in the parent, and -1 on failure.
  • Copied: data, heap, stack, registers, descriptors. Shared: the text, read only, and the

open files the copied descriptors refer to, including the file position.

  • The child gets a new process id; its parent id is the caller's id.
  • n forks in a loop give 2 to the power n processes. Three forks give eight.
  • The order in which parent and child run is not defined.
  • fork and exec are separate so that the child can change its own settings, which is where

all shell redirection happens.

  • Real systems do not copy the memory at once: they use copy on write.

Test yourself

  1. What does fork return, and where? 0 in the child, the child's process id in the parent,

and -1 in the parent if it failed.

  1. Why does it return twice? Because it copies the whole process including the program

counter, so the copy is also in the middle of returning from the call.

  1. A program calls fork four times in a loop. How many processes exist at the end? Sixteen,

of which fifteen are children. 4. Parent and child both write to standard output, which is redirected to a file. What happens? The descriptors were copied but refer to one open file, so both writes go to the same file and the file position is shared. The output interleaves.

  1. Why are fork and exec two calls? So that between them the child can change its own

environment, for example opening a file as descriptor 1 to redirect the output, before turning itself into the new program.

  1. Which runs first after a fork? Undefined. Both are ready and the scheduler chooses.

Contents This chapter on its own page

munotes.in72

Chapter Nineteen

Running a Different Program: exec

Syllabus topic Module 1, "Processes - Operations on Processes"

In one line

exec throws away the program a process is running and loads a different one in its place, keeping the same process.

Why it is not "start a program"

The name misleads every beginner, so put it plainly: exec does not create anything. There is one process before and one process after, with the same process id, the same parent, and the same open files. What changed is the program inside it.

A successful exec never returns. There is nothing to return to: the instructions that would have received the return value have been replaced. So the line after an exec runs only if the exec failed, and that is why every correct program treats reaching it as an error.

What survives an exec and what does not

SurvivesReplaced
Process idyes
Parent, and the parent's knowledge of ityes
Open file descriptorsyes, unless marked close on exec
Current directory, user, groupyes
Text, data, heap, stackall replaced by the new program's
Registers and program counterset to the new program's start
Signal handlersreset: the new program has none of the old one's

The surviving descriptors are the point. They are how redirection works: the child opens a file as descriptor 1, then execs, and the new program, which knows nothing about any of this, writes to descriptor 1 and the bytes land in the file.

The six forms, and how to remember them

There is one system call, execve, and the C library offers six front doors to it. The names look arbitrary until you learn that the letters after exec are the answer to three questions.

LetterMeans
lthe arguments are a list, written out one by one, ending with a null pointer
vthe arguments are a vector, an array of strings you built
pthe program is looked for on the path, so a bare name such as ls works
eyou supply the environment yourself
FormArgumentsSearches the path?Environment
execllistnoinherited
execlplistyesinherited
execlelistnoyou supply it
execvvectornoinherited
execvpvectoryesinherited
execvevectornoyou supply it

So execlp("ls", "ls", "-l", NULL) is a list, searched on the path. execv("/bin/ls", args) is a vector, with a full path. execve is the real system call, which is why the trace in Chapter ten showed that name whatever the program wrote.

The first argument is the program's own name and must be given twice. execlp("ls", "ls", NULL) looks like a mistake and is not: the first is what to run, the second becomes the new program's argument zero. Leaving the second out gives a program that does not know its own name, and some programs behave differently depending on it.

munotes.in73

Running a Different Program: exec

fork then exec, in one program

#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <sys/wait.h>

int main(void)
{
    printf("shell : I am %d, about to run wc\n", getpid());
    fflush(stdout);                       /* see the note below */

    pid_t child = fork();
    if (child < 0) {
        perror("fork");
        return 1;
    }
    if (child == 0) {
        printf("child : I am %d, and I am about to become wc\n", getpid());
        fflush(stdout);
        execlp("wc", "wc", "-l", "lines.txt", NULL);
        perror("execlp");                 /* reached only if exec failed */
        _exit(127);
    }
    int status = 0;
    wait(&status);
    printf("shell : child %d finished with status %d\n", child,
        WEXITSTATUS(status));
    return 0;
}
one
two
three
$ gcc -std=c17 -Wall -Wextra -o forkexec forkexec.c
$ ./forkexec
shell : I am 24, about to run wc
child : I am 25, and I am about to become wc
3 lines.txt
shell : child 25 finished with status 0

Three observations, and each is examinable.

  1. The child printed a line and then stopped being itself. The printf after execlp never

ran, and neither did perror, because the exec succeeded and the program was gone.

  1. wc printed as process 21, the same process the child was. The parent waited for 21 and

got it.

  1. fflush is not decoration. Chapter nine showed that printf buffers. A buffer is part of

the program's data, so exec throws it away unflushed, and the line would have been lost. The same happens across fork, where the buffer is copied and the line is printed twice.

That third point is the bug this chapter exists to prevent. A printf before a fork or an exec, with output going to a file, is printed twice or not at all. Flush, or use write.

Redirection, which is the whole reason for the gap

Now the thing Chapter eight promised. The child changes descriptor 1 before exec, and the new program never knows.

#include <stdio.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/wait.h>

int main(void)
{
    pid_t child = fork();

    if (child == 0) {
        int fd = open("count.txt", O_CREAT | O_WRONLY | O_TRUNC, 0644);
        dup2(fd, STDOUT_FILENO);          /* make the file descriptor 1 */
        close(fd);
        execlp("wc", "wc", "-l", "lines.txt", NULL);
        _exit(127);
    }
    wait(NULL);
    printf("wc wrote nothing to the screen. the file holds:\n");
    fflush(stdout);
    execlp("cat", "cat", "count.txt", NULL);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o redirect redirect.c
$ ./redirect
wc wrote nothing to the screen. the file holds:
3 lines.txt

wc was given no file to write to and no flag. It wrote to descriptor 1, as it always does. The child had quietly made descriptor 1 be the file, with dup2, before becoming wc. That is exactly what the shell does when you type wc -l lines.txt > count.txt, and it is why fork and exec are two calls.

munotes.in74

Running a Different Program: exec

Notice the last line of the parent: it execs cat, so the parent becomes cat and never returns. There is no return 0 reached. The process ends as cat.

Distinctions that carry marks

forkexec
Creates a processyesno
Processes afterwardstwoone
Returnstwicenever, if it succeeds
The programunchanged in bothreplaced
Process idthe child gets a new oneunchanged
execl familyexecv family
Arguments written asa list in the call, ending in NULLan array you built
Use whenyou know the arguments while writingthe arguments are computed

What it does not mean

exec is not slow because it loads a program. On a paged system it maps the file rather than reading it, and pages arrive as they are touched (Chapter eighty one).

A failed exec is not a crash. It returns -1 like any call, and the program is still the old one, which is why the error must be handled: usually by exiting, because the child has nothing useful left to do.

exec does not close your files. They survive, on purpose. A descriptor that should not survive is marked close on exec when it is opened.

Quick revision

  • exec replaces the program running in a process. It creates nothing, and

a successful exec never returns.

  • Survives: process id, parent, open descriptors, current directory, user. Replaced: text, data,

heap, stack, registers, signal handlers.

  • Six forms, from three letters: l list, v vector, p search the path, e supply the

environment. execve is the real system call.

  • The first argument is given twice: what to run, and what the program's own argument zero should

be.

  • Flush before fork or exec, or a buffered printf is lost or printed twice.
  • Redirection is done in the gap between fork and exec, with dup2, and the new program

never knows.

Test yourself

  1. What does exec do, and how many processes exist afterwards? It replaces the program

running in the calling process. One process exists, with the same id.

  1. Why does a successful exec never return? Because the code that would receive the return

value has been replaced by the new program.

  1. What does the line after exec mean? That the exec failed. It is an error path, and

should report and exit.

  1. Explain the letters in execlp. l for a list of arguments written out in the call, p

for searching the path so that a bare program name works.

  1. How does the shell implement command > file? It forks; in the child it opens the file
munotes.in75

Running a Different Program: exec

and makes it descriptor 1 with dup2; then it execs the command, which writes to descriptor 1 as usual and knows nothing about the redirection.

  1. Why must you flush before forking? Because the C library's output buffer is part of the

process's data. After a fork both copies hold it and both print it; after an exec it is thrown away unprinted.

Contents This chapter on its own page

munotes.in76

Chapter Twenty

Waiting, Exiting, the Zombie and the Orphan

Syllabus topic Module 1, "Processes - Operations on Processes"

In one line

A process ends by calling exit; its parent collects the exit status with wait; and if the parent does not, the dead child stays in the tables as a zombie, while a child whose parent dies first becomes an orphan and is given to another process.

Why a dead process does not disappear at once

When a process exits it has one thing left that somebody may want: how it ended. Did it succeed? Did it fail, and with what code? Was it killed, and by which signal?

That answer cannot be thrown away with the process, because the parent may not have asked for it yet. So the kernel frees the process's memory, closes its files, and keeps only its exit status and its process id, until the parent collects them.

A zombie is therefore not a mistake in the kernel and not a process that is still running. It is a small record, a few hundred bytes, holding an answer nobody has asked for. It uses no processor and no memory beyond the record. What it does use is a process id, and process ids are finite, so a program that makes zombies for hours eventually cannot make any more processes at all.

Exit, and what the status carries

exit(n) ends a process with status n. Only the low eight bits of n survive, so the useful range is 0 to 255, and the convention is that 0 means success.

The parent does not receive that number directly. It receives an integer that packs several things together, and it is taken apart with macros.

MacroAnswers
WIFEXITED(status)did it exit normally?
WEXITSTATUS(status)if so, with what number
WIFSIGNALED(status)was it killed by a signal?
WTERMSIG(status)if so, which signal

Do not read the raw integer. A status of 139 at the shell means signal 11, but the integer the parent receives is not 139: the shell computes that. Use the macros.

There are two calls for ending a process and the difference is examined.

exit_exit
Flushes the C library's buffersyesno
Runs functions registered with atexityesno
Isa library functionthe system call
Use in a child after a failed execnoyes

The last row is the rule worth remembering. A child that could not exec must not flush buffers it inherited from its parent, or the parent's un-printed output appears twice. Chapter nineteen's listing uses _exit for exactly that reason.

Wait, and how a parent collects

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <stdlib.h>
#include <sys/wait.h>
#include <unistd.h>

int main(void)
{
    for (int which = 0; which < 2; which++) {
        pid_t child = fork();
        if (child == 0) {
            if (which == 0) {
                exit(42);                         /* ends normally */
            }
            abort();                              /* killed by a signal */
        }
        int status = 0;
        pid_t got = wait(&status);
        if (WIFEXITED(status)) {
            printf("child %d exited normally with status %d\n",
                got, WEXITSTATUS(status));
        } else if (WIFSIGNALED(status)) {
            printf("child %d was killed by signal %d\n",
                got, WTERMSIG(status));
        }
    }
    return 0;
}
munotes.in77

Waiting, Exiting, the Zombie and the Orphan

$ gcc -std=c17 -Wall -Wextra -o collect collect.c
$ ./collect
child 22 exited normally with status 42
child 23 was killed by signal 6

Signal 6 is SIGABRT, which is what abort raises. The parent learned how each child ended, not merely that it had.

wait waits for any child and blocks until one ends. waitpid waits for a particular child, and with the flag WNOHANG it asks without blocking, which is how a program collects children while getting on with its own work.

A zombie, made and then cleaned up

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <stdlib.h>
#include <sys/wait.h>
#include <unistd.h>

int main(void)
{
    pid_t child = fork();

    if (child == 0) {
        exit(0);                          /* the child ends at once */
    }
    FILE *note = fopen("child.pid", "w");   /* so the shell can find it below */
    fprintf(note, "%d\n", child);
    fclose(note);

    printf("child %d has exited and I am NOT collecting it yet\n", child);
    fflush(stdout);
    sleep(2);                             /* the parent does something else */

    wait(NULL);                           /* now collect */
    printf("collected. it is gone from the table\n");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o zombie zombie.c
$ set +m
$ ./zombie &
$ sleep 1
child 37 has exited and I am NOT collecting it yet
$ ps -o pid,ppid,stat,comm -p "$(cat child.pid)" --no-headers
     37      35 Z+   zombie
$ wait 2>/dev/null
collected. it is gone from the table
$ ps -o pid,stat,comm -p "$(cat child.pid)" --no-headers
$ echo "the table no longer has it"
the table no longer has it

Z in the state column, and ps shows it as still named. Two seconds later the parent called wait, the record was released, and ps found nothing at all.

A zombie is cleaned up by its PARENT calling wait, and by nothing else. You cannot kill a zombie: it is already dead, and a signal to a dead process does nothing. The only two ways it goes are the parent collecting it, or the parent dying, after which the kernel's own process adopts it and collects it immediately. So a machine full of zombies is fixed by fixing or restarting the parent.

An orphan, and who adopts it

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <unistd.h>

int main(void)
{
    pid_t child = fork();

    if (child == 0) {
        FILE *log = fopen("orphan.log", "w");
        fprintf(log, "child : while my parent lived it was %d\n", getppid());
        fflush(log);
        sleep(2);                         /* outlive the parent */
        fprintf(log, "child : now my parent is %d\n", getppid());
        fclose(log);
        return 0;
    }
    printf("parent: I am %d and I am leaving now\n", getpid());
    return 0;
}
munotes.in78

Waiting, Exiting, the Zombie and the Orphan

$ gcc -std=c17 -Wall -Wextra -o orphan orphan.c
$ ./orphan
parent: I am 57 and I am leaving now
$ sleep 3
$ cat orphan.log
child : while my parent lived it was 57
child : now my parent is 1

The child wrote to a file rather than to the screen, because it wakes up two seconds after the session has moved on and a line arriving in the middle of somebody else's output proves nothing. Read its two lines: the child's parent changed under it. A process whose parent dies is re-parented, on Linux to the nearest ancestor that has asked for the job, or otherwise to process 1, the first process the kernel started. That process's whole remaining duty is to wait for whatever it inherits, which is why orphans never become permanent zombies.

An orphan is not an error. A program deliberately made an orphan of itself is how a daemon is written: fork, let the parent exit, and the child runs on with no terminal and no parent.

Cascading termination, which is the other policy

A question asks for the term, and the answer is a comparison with what this chapter has just measured. When a parent dies, a system must decide what happens to its children, and there are two answers.

PolicyWhat happens to the childrenWhere it is used
reparenting, which is what the run above showedeach child is adopted, by init or by a subreaper, and goes on runningUNIX and Linux
cascading terminationevery child is killed as well, and their children after them, all the way downsystems that do not allow a process to outlive its parent, and it is what a process group kill does on request

Cascading termination is initiated by the operating system, not by the child, and its point is that no process is left running with nobody to collect it. Linux does not do it by default, which is why the orphan above kept running and printed its line after its parent had gone.

Distinctions that carry marks

ZombieOrphan
The process isdeadalive
Who has gonenobodythe parent
What remainsits exit status and its process idthe whole process
Caused bya parent that has not called waita parent that exited first
Cleaned up bythe parent calling wait, or the parent dyingbeing re-parented, to process 1
Can it be killed?no, it is already deadyes, like any process
Harmit uses a process id, and ids run outnone, usually deliberate
munotes.in79

Waiting, Exiting, the Zombie and the Orphan

waitwaitpid
Waits forany childone named child, or a group
Blocksalways, until a child endsunless WNOHANG is given
Use whenthere is one childthere are several, or you must not block

What it does not mean

A zombie is not using the processor. It is a record. The name is theatrical; the cost is one process id.

Killing a zombie is not possible and is not the fix. Fix the parent.

An orphan is not a zombie. The orphan is alive and running; the zombie is dead and uncollected. Students swap these two constantly and it is worth a line in every answer.

exit(0) and return 0 from main are the same thing, but exit from anywhere else is not the same as returning: it ends the process rather than the function.

Quick revision

  • exit(n) ends a process; only the low eight bits of n survive; 0 means success.
  • The parent collects with wait or waitpid and takes the status apart with WIFEXITED,

WEXITSTATUS, WIFSIGNALED, WTERMSIG.

  • exit flushes the C library's buffers; _exit does not, which is why a child that failed

to exec must use _exit.

  • A zombie is a dead child whose exit status has not been collected. It holds a process id

and nothing else, shows Z in ps, and cannot be killed. Only the parent's wait, or the parent's death, removes it.

  • An orphan is a live child whose parent has died. It is re-parented, ultimately to process

1, whose job is to collect it.

  • Making an orphan on purpose is how a daemon is started.

Test yourself

  1. What is a zombie process and why does it exist? A process that has ended but whose exit

status has not yet been collected by its parent. It exists because the status cannot be thrown away before somebody asks for it.

  1. How is a zombie removed? The parent calls wait or waitpid. If the parent dies, the

process the kernel keeps for the purpose adopts it and collects it at once. It cannot be killed.

  1. What is an orphan, and what happens to it? A running process whose parent has exited. The

kernel re-parents it, ultimately to process 1.

  1. Distinguish a zombie from an orphan in one sentence each. A zombie is dead and

uncollected. An orphan is alive and has lost its parent.

  1. Why must a child use _exit rather than exit after a failed exec? Because exit

flushes the C library buffers, which were copied from the parent, so the parent's unwritten output would be printed a second time.

munotes.in80

Waiting, Exiting, the Zombie and the Orphan

  1. A server has accumulated 30,000 zombies. What is wrong and what fixes it? The parent is

not calling wait. Fix the parent, or restart it: its zombies are then adopted and collected. Killing the zombies does nothing.

  1. Why is a machine full of zombies eventually unable to start any process? Each zombie holds

a process id, and the supply of ids is finite.

Contents This chapter on its own page

munotes.in81

Chapter Twenty-One

Inter-process Communication: The Two Models

Syllabus topic Module 1, "Processes - Inter-process Communication"

In one line

Two processes can communicate in exactly two ways: they can share a piece of memory and both look at it, or they can send each other messages through the kernel.

Why processes need to communicate at all

Chapter fourteen said that every process has its own memory and shares nothing but read only instructions. That is protection, and it is the point. But it means that two processes which are co-operating have no way to say anything to each other, and co-operating processes are the normal case:

  • a shell and the command it started;
  • a browser and the process that draws one of its tabs;
  • a database server and the twenty programs asking it questions;
  • the producer of some data and the consumer of it, which is the pattern behind almost all of the

rest of this module.

MU's text book gives four reasons for wanting co-operating processes, and they are worth a line each: information sharing (two programs want the same file), computation speedup (split the work across processors), modularity (build the system out of separate processes), and convenience (one person editing, printing and compiling at once).

The two models

Shared memory

The kernel is asked, once, to make a region of memory that appears in both processes' address spaces. After that the kernel is not involved. Each process reads and writes it with ordinary instructions, at memory speed.

Message passing

Neither process can see the other's memory at all. One calls send, the kernel copies the data out of that process, and the other calls receive, and the kernel copies it in. Every message crosses the boundary twice.

The comparison, which is the examined part

Shared memoryMessage passing
Set up bysystem calls, oncesystem calls, once
Each transfer costsnothing: ordinary memory accesstwo system calls and two copies
Speedas fast as memorymuch slower per message
Good forlarge amounts of data, frequent exchangesmall amounts, or a network
Synchronisationthe programmer's problemhandled by the kernel
Between machinesimpossibleworks unchanged
Dangera race condition (Chapter thirty four)a full mailbox, a lost message, deadlock
Easier to get rightnoyes

The synchronisation row is the whole trade and is worth stating as a sentence in any answer. Shared memory is faster because the kernel is not involved in each transfer. The kernel not being involved is exactly why nothing stops both processes writing the same place at the same moment. So shared memory buys speed and hands the programmer a problem that Chapters thirty three to forty two exist to solve. Message passing is slower and the problem never arises, because a message is delivered once, whole, by somebody in charge.

munotes.in81

Inter-process Communication: The Two Models

The questions a message passing scheme must answer

Any message passing system has to settle five things, and a question may ask for them.

  1. Direct or indirect. Direct: the sender names the receiving process. Indirect: both use a

mailbox or port, and the sender does not know who will take the message. Indirect is more flexible, and a mailbox may be owned by a process or by the operating system.

  1. Blocking or non-blocking. Chapter twenty five is this question alone.
  2. Buffered or not. Zero capacity, so the sender must wait for a receiver; bounded capacity,

so the sender waits only when the buffer is full; or unbounded, where the sender never waits.

  1. Fixed or variable length messages. Fixed is simple for the kernel and awkward for the

programmer; variable is the other way round.

  1. What happens when things fail. A message lost, a process that ends while a message is in

flight, a message that arrives twice.

What is actually available, on a real machine

There are more than two facilities, and they sort into the two models. All of them are in the next four chapters.

FacilityModelBetweenSurvives the process?
Pipemessage passingrelated processes, usually parent and childno
Named pipe, or FIFOmessage passingany two processes on the machinethe name does
System V message queuemessage passingany two on the machineyes
POSIX message queuemessage passingany two on the machineyes
System V shared memoryshared memoryany two on the machineyes
POSIX shared memoryshared memoryany two on the machineyes
Socketmessage passingany two, on one machine or across a networkno
Signalmessage passing, of one bitany two, with permissionno
$ ipcs -q -m | head -8

------ Message Queues --------
key        msqid      owner      perms      used-bytes   messages

------ Shared Memory Segments --------
key        shmid      owner      perms      bytes      nattch     status
$ ls -l /dev/mqueue 2>/dev/null | head -2 || echo "no POSIX queues exist yet"
total 0

Both tables are empty because nothing on this machine is using them yet. That the tables exist at all, and are separate from any process, is the point: a System V queue or segment belongs to the system, not to the process that made it, which is why Chapter twenty four's queue outlives the program that created it.

Worked example: choosing a model twice

A video editor and its rendering helper exchange whole frames, several megabytes each, sixty times a second. Copying that through the kernel twice per frame would cost more than the rendering. Shared memory, with a lock, is the only sensible choice, and the programmer accepts the synchronisation problem.

munotes.in82

Inter-process Communication: The Two Models

A print spooler and the programs that submit jobs exchange a file name and a few options, perhaps once a minute. The data is tiny, the programs are unrelated and were written by different people, and correctness matters far more than microseconds. Message passing: a queue, so the spooler need not be running when a job is submitted.

The rule the two examples give: large and frequent, share memory; small, occasional or across a network, pass messages.

What it does not mean

Shared memory is not shared variables. The processes agree on a region and must agree on what is in it, byte by byte. Nothing checks that agreement.

Message passing is not slow in absolute terms. It is slower per byte than memory. A few thousand messages a second is nothing to a modern machine.

A pipe is not shared memory even though it feels like a shared thing. Every byte through a pipe is copied by the kernel, twice.

Using shared memory does not remove the need for the kernel. It is asked for the region and asked to tear it down, and the locks the processes need are often kernel objects too.

Quick revision

  • Two models: shared memory and message passing.
  • Reasons to co-operate: information sharing, computation speedup, modularity, convenience.
  • Shared memory: fast, no kernel involvement per access, and

synchronisation is the programmer's problem.

  • Message passing: two system calls and two copies per message, slower, and the kernel

delivers each message whole, so no race arises. It works across a network.

  • A message passing scheme must settle: direct or indirect (a mailbox or port), blocking

or non-blocking, buffered with zero, bounded or unbounded capacity, fixed or variable length, and what happens on failure.

  • Facilities: pipes and named pipes, System V and POSIX message queues, System V and POSIX shared

memory, sockets, signals.

  • Large and frequent data, share memory. Small, occasional or remote, pass messages.

Test yourself

  1. Name the two models of inter-process communication and the essential difference. Shared

memory and message passing. In shared memory the kernel sets up a region once and is not involved again; in message passing every transfer goes through the kernel.

  1. Why is shared memory faster, and what does that cost? Because reads and writes are

ordinary memory accesses with no system call. The cost is that nothing co-ordinates the two processes, so the programmer must prevent race conditions.

  1. Give the four reasons for co-operating processes. Information sharing, computation

speedup, modularity, convenience.

  1. Distinguish direct from indirect communication. In direct communication the sender names

the receiving process. In indirect communication both use a mailbox or port, and the sender need not know who receives.

munotes.in83

Inter-process Communication: The Two Models

  1. What are the three buffering capacities? Zero, so the sender waits for a receiver;

bounded, so the sender waits only when the buffer is full; unbounded, so the sender never waits.

  1. Two processes exchange five megabyte images sixty times a second. Which model, and why?

Shared memory: copying that much data through the kernel twice per image would cost more than the work itself. The programmer accepts responsibility for locking.

Contents This chapter on its own page

munotes.in84

Chapter Twenty-Two

Shared Memory, in Code That Runs

Syllabus topic Module 1, "Processes - Inter-process Communication"; Computer Science Practical 3, Module 1, "Process Communication using Shared Memory"

In one line

Shared memory is one piece of physical memory that appears in two processes' address spaces at once, so that a write by one is immediately visible to the other.

Why it looks the way it does

Chapter fourteen said each process has its own memory. The memory management unit of Chapter sixty eight is what makes that true: it maps each process's addresses to different physical memory. Shared memory is that same machinery used deliberately the other way: the kernel maps one piece of physical memory into two processes' page tables, at whatever address each of them happens to have free.

The two processes need not see it at the same address, and usually do not. That is the single most important practical consequence, and it is examined: you may not store a pointer in shared memory, because a pointer is an address and the other process's addresses are different. Store offsets, or plain data.

The four calls

There are two families. System V shared memory is what MU's text book describes and what ipcs reports, so it is taught here; POSIX shared memory is mentioned at the end.

CallWhat it does
shmget(key, size, flags)make or find a segment, and return its id
shmat(id, NULL, 0)attach it: map it into this process, and return the address
shmdt(address)detach it: unmap it from this process
shmctl(id, IPC_RMID, NULL)remove it from the system

A key is a name and an id is a handle. The key is a number two unrelated programs agree on in advance, so that both can find the same segment. IPC_PRIVATE as the key means that no name is needed and the kernel should give out a fresh segment, which is enough when the other process is a child and inherits the id.

shmdt and shmctl(IPC_RMID) are not the same thing, and this is where marks are lost. Detaching removes it from one process. Removing destroys it for everybody. A program that detaches and forgets to remove leaves the segment in the system for ever, and the next chapter's ipcs output is how you find them.

Seen from outside, with no program at all

The shell can make a segment, list it and remove it, which is the quickest way to see that a segment is a thing the system owns rather than a thing a process owns.

$ ipcmk -M 4096
Shared memory id: 0
$ ipcs -m | awk '$3 == "student" {print "id", $2, "owner", $3, "perms", $4, "bytes", $5, "attached", $6}'
id 0 owner student perms 644 bytes 4096 attached 0
$ ipcrm -m 0
$ ipcs -m | awk '$3 == "student" {print "a segment is still there"} END {print "the table has been read"}'
the table has been read
munotes.in85

Shared Memory, in Code That Runs

A segment of 4096 bytes existed, owned by student, with nattch of 0 because no process had attached it. Then it was removed and the table was empty again. No program was running at any point. The segment outlived the command that made it, which is exactly the property that makes System V shared memory useful and dangerous.

Two processes, one segment

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <string.h>
#include <sys/ipc.h>
#include <sys/shm.h>
#include <sys/wait.h>
#include <unistd.h>

struct board {
    int ready;
    int count;
    char message[64];
};

int main(void)
{
    int id = shmget(IPC_PRIVATE, sizeof(struct board), IPC_CREAT | 0600);
    if (id < 0) {
        perror("shmget");
        return 1;
    }
    printf("the segment id is %d, and it is %zu bytes\n", id, sizeof(struct board));

    pid_t child = fork();
    if (child == 0) {
        struct board *b = shmat(id, NULL, 0);
        strcpy(b->message, "the child wrote this straight into memory");
        b->count = 7;
        b->ready = 1;                     /* last, so the parent sees a whole message */
        printf("child : written, and detaching\n");
        shmdt(b);
        _exit(0);
    }
    wait(NULL);

    struct board *b = shmat(id, NULL, 0);
    printf("parent: ready is %d, count is %d\n", b->ready, b->count);
    printf("parent: message is \"%s\"\n", b->message);
    shmdt(b);
    shmctl(id, IPC_RMID, NULL);           /*  without this the segment stays */
    printf("parent: segment removed\n");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o shareit shareit.c
$ ./shareit
the segment id is 1, and it is 72 bytes
child : written, and detaching
parent: ready is 1, count is 7
parent: message is "the child wrote this straight into memory"
parent: segment removed
$ ipcs -m | awk '$3 == "student" {print "a segment is still there"} END {print "the table has been read"}'
the table has been read

Notice what is not in that program: no send, no receive, no system call for the transfer. The child assigned to a structure and the parent read the structure. Between those two lines the kernel did nothing at all, and that is the whole reason shared memory exists.

Notice also the order of the child's three writes: the message first, the count next, the ready flag last. That is not an accident and it is the right habit: the flag says "everything else is valid", so it must be written after everything else. Chapters thirty three to forty two are about doing this properly rather than carefully.

The danger, named now and solved later

The program above is safe only because the parent called wait, so the two processes never touched the segment at the same moment. Take that away and the program is broken in a way that usually works, which is the worst kind of broken.

munotes.in86

Shared Memory, in Code That Runs

Shared memory with no synchronisation is a race condition waiting to be reported as an intermittent bug. Chapter thirty four makes one happen on purpose; Chapters thirty seven to forty two are the tools that fix it. The practical's own instruction is "Explore issues of race conditions and how to avoid them", and that is the honest order: see the mechanism here, see the danger in Chapter thirty four, see the fix in Chapter thirty nine.

The POSIX family, in one paragraph

POSIX shared memory does the same job with the file system as its naming scheme. shm_open("/name", ...) gives a descriptor, ftruncate sets the size, and mmap maps it. The name looks like a file name and appears under /dev/shm, so an ordinary ls finds leftovers. It is newer and easier to clean up; the System V family is what an examination question will name, because it is what MU's text book uses.

Distinctions that carry marks

shmdtshmctl(id, IPC_RMID, NULL)
Effectdetaches from this processdestroys the segment for everybody
Other processesunaffectedlose it
Forget it andthis process's mapping goes when it exitsthe segment stays in the system for ever
Shared memoryMessage passing
Kernel involved per transfernoyes
Speedmemory speedtwo copies and two system calls
Synchronisationyours to arrangethe kernel's
May contain pointersno: the addresses differ per processnot applicable
KeyId
Isa name two programs agree ona handle the kernel gives back
Chosen bythe programmer, or IPC_PRIVATEthe kernel
Used byshmget to find the segmentevery other call

What it does not mean

Shared memory is not a shared variable. There is no declaration that makes two processes share a variable. There is a region of bytes, and both programs must agree on its layout.

Attaching is not copying. shmat maps; nothing is copied, which is why it is fast however large the segment is.

Removing it is not immediate if somebody is attached. IPC_RMID marks it for destruction; it goes when the last process detaches. So a segment can be removed and still be usable by the processes already holding it, which is the safe way to clean up.

Quick revision

  • Shared memory maps one piece of physical memory into two address spaces. No kernel

involvement per access, so it is the fastest form of communication.

  • shmget to make or find (returns an id), shmat to attach (returns an address), shmdt

to detach, shmctl with IPC_RMID to remove.

  • A key is an agreed name; IPC_PRIVATE asks for an unnamed one, enough when a child

inherits the id.

  • The two processes may see the segment at different addresses, so never store a pointer in
munotes.in87

Shared Memory, in Code That Runs

it.

  • shmdt is not IPC_RMID. Detach affects one process; remove destroys it for all. A

forgotten remove leaves the segment in the system, and ipcs -m finds it.

  • Write the data first and the ready flag last.
  • Shared memory with no synchronisation is a race condition. Chapters thirty four and thirty

nine.

  • The POSIX family is shm_open, ftruncate, mmap, and its names appear under /dev/shm.

Test yourself

  1. Name the four System V shared memory calls and what each does. shmget makes or finds a

segment and returns its id; shmat attaches it and returns an address; shmdt detaches it from this process; shmctl with IPC_RMID removes it from the system.

  1. Why must a pointer never be stored in shared memory? Because each process may map the

segment at a different address, so an address valid in one is meaningless in the other. Store offsets.

  1. Distinguish detaching from removing. Detaching unmaps it from the calling process only.

Removing destroys it for everybody, once the last process detaches.

  1. What happens if a program forgets to remove a segment? It stays in the system after the

program ends. ipcs -m lists it and ipcrm -m removes it.

  1. Why is the ready flag written last? Because it tells the reader that everything else is

valid, so it must not become true before the data it vouches for is there.

  1. Why is shared memory faster than message passing? Because after the one time setup the

kernel is not involved: a transfer is an ordinary memory access rather than two system calls and two copies.

  1. What is the price of that speed? Nothing co-ordinates the two processes, so preventing

race conditions becomes the programmer's job.

Contents This chapter on its own page

munotes.in88

Chapter Twenty-Three

Pipes

Syllabus topic Module 1, "Processes - Inter-process Communication"; Computer Science Practical 3, Module 1, "Process Communication using Message Passing"

In one line

A pipe is a one way channel with two ends: whatever is written into one end can be read out of the other, in order.

Why a pipe is the shape it is

A pipe is the oldest form of inter-process communication on Unix and it was designed for one purpose: to let the output of one program be the input of another without either program knowing. That purpose explains every property it has.

  • It is one way, because a pipeline flows one way.
  • It carries a stream of bytes with no message boundaries, because that is what a program's

output is.

  • It is reached through file descriptors, so a program that can write to a file can write to

a pipe without being changed at all. That is the whole trick: wc has no idea whether descriptor 0 is a file, a terminal or a pipe.

  • It has a fixed size buffer in the kernel, so a fast writer is slowed to the reader's pace

automatically.

The call, and the two descriptors

pipe(int fd[2]) fills an array with two descriptors.

WhichFor
fd[0]the read endreading. Remember: 0 is like standard input
fd[1]the write endwriting. Remember: 1 is like standard output

Each process must close the end it does not use, and forgetting is the commonest pipe bug there is. The reason is exactly this: a read on a pipe returns 0, meaning end of file, only when every write end everywhere is closed. If the reading process still holds its own copy of the write end, the kernel can see a write end open, so the read waits for ever for data that will never come. The program hangs and nothing says why.

A pipe between parent and child

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <string.h>
#include <sys/wait.h>
#include <unistd.h>

int main(void)
{
    int ends[2];

    if (pipe(ends) < 0) {
        perror("pipe");
        return 1;
    }
    pid_t child = fork();
    if (child == 0) {
        close(ends[0]);                   /* the child only writes */
        const char *lines[] = {"first\n", "second\n", "third\n"};
        for (int i = 0; i < 3; i++) {
            write(ends[1], lines[i], strlen(lines[i]));
        }
        close(ends[1]);                   /*  so the parent sees end of file */
        _exit(0);
    }
    close(ends[1]);                       /* the parent only reads */

    char buf[64];
    ssize_t n;
    int reads = 0;
    while ((n = read(ends[0], buf, sizeof buf)) > 0) {
        reads++;
        printf("parent: read %zd bytes: %.*s", n, (int)n, buf);
    }
    printf("parent: read returned %zd after %d read(s), so the pipe is finished\n",
        n, reads);
    close(ends[0]);
    wait(NULL);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o pipeit pipeit.c
$ ./pipeit
parent: read 19 bytes: first
second
third
parent: read returned 0 after 1 read(s), so the pipe is finished
munotes.in89

Pipes

Three writes came out as one read, and that is the lesson. The child wrote "first", "second" and "third" separately. The parent asked for up to 64 bytes and got all 19 at once. A pipe has no message boundaries: it is a stream. A program that needs to know where one message ends must put the boundary in the bytes itself, for example a newline, which is exactly why Unix tools are line based.

Then read returned 0. That is end of file, and it happened because the child closed its write end and the parent had closed its copy. Had the parent not closed ends[1], that read would still be waiting.

Named pipes, which any two programs can use

An ordinary pipe can only be shared by inheritance, so the two processes must be related. A named pipe, also called a FIFO, is a pipe with a name in the file system, so any two programs can open it.

$ mkfifo channel
$ ls -l channel
prw-r--r-- 1 student student 0 Sep 30  2026 channel
$ set +m
$ ( echo "sent through a named pipe" > channel ) &
$ cat channel
sent through a named pipe
$ wait 2>/dev/null
$ rm channel

Three things in that transcript.

  1. ls -l shows the type letter p, for pipe. It is a file system entry, with permissions,

and it is not a file: nothing is stored in it.

  1. Its size is 0 and always will be. The data lives in the kernel's buffer, not on the disk.
  2. The writer was put in the background on purpose.

Opening a FIFO for writing blocks until somebody opens it for reading, and the other way round, so running both in the foreground in the wrong order simply waits.

What the shell's vertical bar is

a | b is exactly this chapter plus Chapter nineteen.

  1. The shell calls pipe.
  2. It forks twice, once for a and once for b.
  3. In the process that will be a: dup2(ends[1], 1), so the pipe's write end becomes standard

output; close both original ends; exec a.

  1. In the process that will be b: dup2(ends[0], 0), so the pipe's read end becomes standard

input; close both original ends; exec b.

  1. The shell closes both ends itself and waits for both children.

Neither a nor b knows a pipe exists. They write to 1 and read from 0 as always. That is why almost every Unix program composes with every other, and it is the strongest single argument for the whole design.

The capacity is real and can be measured. A writer that fills the buffer is made to wait.

munotes.in90

Pipes

$ mkfifo slow
$ set +m
$ ( yes 0123456789 | head -20000 > slow ) &
$ sleep 1
$ pgrep -c yes
1
$ ps -o stat= -p "$(pgrep -n yes)"
S+
$ head -c 20 slow
0123456789
012345678
$ kill $(jobs -p) 2>/dev/null
$ wait 2>/dev/null
$ rm -f slow

One yes is running and its state is S, sleeping: it is waiting for the write to be accepted. The + after it only means it is in the foreground of its own job. That is flow control for free, and it is why yes | head does not fill your disk.

Distinctions that carry marks

Ordinary pipeNamed pipe, FIFO
Has a namenoyes, in the file system
Made bypipe()mkfifo, the command or the call
Usable byrelated processes only, by inheritanceany two processes with permission
Survives the processesnothe name does; the data does not
Listed by ls -l asnothingtype p, size 0
PipeMessage queue (Chapter twenty four)
Boundariesnone: a byte streamyes: a message is delivered whole
Directionone wayone way, but many senders and readers
Survives its processesnoyes
Prioritynone: strictly in orderPOSIX queues have priorities

What it does not mean

A pipe is not a file. A named pipe has a name and permissions and no contents. Its size is always zero.

A pipe is not two way. Two way needs two pipes. A socket pair is the two way relative.

A pipe does not preserve write boundaries. Three writes may be one read, and one write may be several reads.

A pipe is not shared memory. Every byte is copied by the kernel, twice.

Quick revision

  • A pipe is a one way byte stream with a read end, fd[0], and a write end, fd[1].
  • Close the end you do not use. read returns 0 only when every write end is closed, so a

forgotten copy makes the reader wait for ever.

  • There are no message boundaries. Three writes may arrive as one read; put the boundary in

the bytes, which is why Unix tools are line based.

  • An ordinary pipe is shared by inheritance, so the processes must be related. A

named pipe or FIFO has a name in the file system and can be used by any two processes; ls -l shows type p and size 0.

  • Opening a FIFO for writing blocks until a reader opens it, and the other way round.
  • The shell's a | b is one pipe, two forks, a dup2 in each child, and exec. Neither
munotes.in91

Pipes

program knows.

  • The kernel's buffer is finite, so a fast writer blocks when it is full: flow control for free.

Test yourself

  1. What is a pipe and which end is which? A one way byte stream between processes. fd[0] is

the read end and fd[1] is the write end.

  1. Why must each process close the end it does not use? Because a read returns end of file

only when all write ends are closed. A reader holding its own copy of the write end will wait for ever.

  1. Three writes of six bytes go into a pipe. How many reads take them out? Any number from

one to eighteen. A pipe is a stream with no message boundaries.

  1. Distinguish an ordinary pipe from a named pipe. An ordinary pipe has no name and can only

be shared by inheritance, so the processes must be related. A named pipe has a name in the file system and any two processes with permission can open it.

  1. Explain how the shell implements sort | uniq. It creates a pipe, forks twice, in one

child makes the write end descriptor 1 and execs sort, in the other makes the read end descriptor 0 and execs uniq, closes both ends itself, and waits for both.

  1. What happens when a writer fills a pipe faster than the reader empties it? The write

blocks until space is free. The kernel's finite buffer gives flow control without either program asking for it.

  1. Why is a named pipe's size always zero? Because it stores nothing. The data is in the

kernel's buffer, and the file system entry is only a name.

Contents This chapter on its own page

munotes.in92

Chapter Twenty-Four

Message Queues

Syllabus topic Module 1, "Processes - Inter-process Communication"; Computer Science Practical 3, Module 1, "Use message queues/pipes to solve the producer-consumer problem"

In one line

A message queue is a list of whole messages, held by the kernel, that one process puts into and another takes out of, and it exists whether or not either process is running.

Why it exists when pipes already do

A pipe has three limitations, and a message queue removes all three.

A pipeA message queue
is a byte stream with no boundariesdelivers whole messages: one send, one receive
dies with its processessurvives them, until it is removed or the machine restarts
is strictly in orderPOSIX queues deliver the highest priority first

The second is the one to hold on to. A print spooler can be started and stopped without the programs that submit jobs knowing, because the jobs wait in the queue. With a pipe, a submitter with nobody at the other end has nowhere to put anything.

The POSIX calls

CallWhat it does
mq_open("/name", flags, mode, attr)make or open a queue, and return a descriptor
mq_send(q, buf, len, priority)put one message in
mq_receive(q, buf, len, &priority)take the highest priority message out
mq_getattr(q, &attr)how many messages are in it, and what the limits are
mq_close(q)this process is finished with it
mq_unlink("/name")remove the queue from the system

The name begins with a slash and contains no other slash. /jobs is right, jobs and /a/b are not.

mq_close is not mq_unlink, and it is the same distinction as shmdt against IPC_RMID in Chapter twenty two. Closing ends this process's use. Unlinking removes the queue for everybody. A program that closes and never unlinks leaves the queue behind, which is sometimes exactly what you want and sometimes a leak.

Programs using these calls must be linked with -lrt, the real time library. A missing -lrt gives an undefined reference at link time and is the commonest first failure.

A message left behind, and picked up later

Two programs. The first sends and exits. The second, run afterwards, receives. Nothing is running in between, and the message survives.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <string.h>
#include <fcntl.h>
#include <mqueue.h>
#include <sys/stat.h>

int main(void)
{
    struct mq_attr want = {0};
    want.mq_maxmsg = 10;
    want.mq_msgsize = 128;

    mqd_t q = mq_open("/college", O_CREAT | O_WRONLY, 0600, &want);
    if (q == (mqd_t)-1) {
        perror("mq_open");
        return 1;
    }
    mq_send(q, "ordinary notice", 15, 1);
    mq_send(q, "URGENT notice", 13, 9);
    mq_send(q, "another ordinary", 16, 1);

    struct mq_attr now;
    mq_getattr(q, &now);
    printf("poster : %ld message(s) are waiting in the queue\n", now.mq_curmsgs);
    mq_close(q);
    printf("poster : closed, and exiting. the queue stays.\n");
    return 0;
}
#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <fcntl.h>
#include <mqueue.h>

int main(void)
{
    mqd_t q = mq_open("/college", O_RDONLY);
    if (q == (mqd_t)-1) {
        perror("mq_open");
        return 1;
    }
    char buf[129];
    unsigned prio;
    for (int i = 0; i < 3; i++) {
        ssize_t n = mq_receive(q, buf, sizeof buf - 1, &prio);
        if (n < 0) {
            perror("mq_receive");
            return 1;
        }
        buf[n] = '\0';
        printf("reader : priority %u, %zd bytes: %s\n", prio, n, buf);
    }
    mq_close(q);
    mq_unlink("/college");
    printf("reader : queue removed\n");
    return 0;
}
munotes.in93

Message Queues

$ gcc -std=c17 -Wall -Wextra -o poster poster.c -lrt
$ gcc -std=c17 -Wall -Wextra -o reader reader.c -lrt
$ ./poster
poster : 3 message(s) are waiting in the queue
poster : closed, and exiting. the queue stays.
$ ls -l /dev/mqueue/college
-rw------- 1 student student 80 Sep 30  2026 /dev/mqueue/college
$ cat /dev/mqueue/college
QSIZE:44         NOTIFY:0     SIGNO:0     NOTIFY_PID:0
$ ./reader
reader : priority 9, 13 bytes: URGENT notice
reader : priority 1, 15 bytes: ordinary notice
reader : priority 1, 16 bytes: another ordinary
reader : queue removed
$ ls /dev/mqueue/

Four things worth naming.

  1. The poster exited and the messages stayed. Between the two programs nothing of either was

running. The queue is the kernel's, not the program's.

  1. The queue appears in the file system, under /dev/mqueue, so an ordinary ls finds

leftovers and cat reports how many bytes are waiting. That is a convenience Linux adds; the standard does not require it.

  1. Priority 9 came out first, although it was sent second. mq_receive always takes the

highest priority waiting, and among equal priorities the oldest. That is the difference from a pipe that an examination question is most likely to want.

  1. Each receive returned exactly one message, with its own length. Fifteen, thirteen and

sixteen bytes came out as three separate messages. Compare Chapter twenty three, where three writes came out as one read.

The System V family, which MU's text book also describes

The older family is msgget, msgsnd, msgrcv and msgctl. Its differences are examinable.

System V queuePOSIX queue
Named bya numeric keya name like /college
Selects bya message type, a long integer chosen by the receiverpriority, highest first
Listed byipcs -qls /dev/mqueue on Linux
Removed bymsgctl(id, IPC_RMID, NULL) or ipcrm -qmq_unlink
Notification when a message arrivesnoyes, mq_notify

The message type is the interesting difference. In a System V queue the receiver says "give me a message of type 4", so one queue can carry several independent conversations. In a POSIX queue the receiver takes whatever is most urgent.

$ ipcmk -Q
Message queue id: 0
$ ipcs -q | awk '$3 == "student" {print "id", $2, "owner", $3, "perms", $4}'
id 0 owner student perms 644
$ ipcrm -q 0
munotes.in94

Message Queues

Worked example: the producer and the consumer, with a queue

The practical asks for the producer and consumer problem solved by message passing, and the point is how much simpler it is than the shared memory version of Chapter thirty nine.

  • The producer loops: make an item, mq_send it. If the queue is full, mq_send blocks

until there is room.

  • The consumer loops: mq_receive, use the item. If the queue is empty, mq_receive blocks

until something arrives.

That is the whole solution. There are no semaphores, no mutex, no shared counter and no race condition, because the queue has a size and the kernel does the waiting. Compare Chapter thirty nine, where the same problem in shared memory needs a mutex and two counting semaphores and can be got wrong in four different ways.

The cost is the one from Chapter twenty one: every item is copied twice and costs two system calls. For three items a second that is nothing. For three million it is everything.

Distinctions that carry marks

PipeMessage queue
Boundariesnone: a streameach message whole
Orderstrictly first in, first outby priority, then by age
Survives its processesnoyes
Namedonly a FIFO isyes
Needs the processes to be relatedan ordinary pipe doesno
mq_closemq_unlink
Effectthis process stops using the queuethe queue is removed from the system
Othersunaffectedlose it
Leftover if forgottennothing seriousthe queue stays after the program ends

What it does not mean

A message queue is not shared memory. Every message is copied into the kernel and out again.

"Priority" is not a process priority. It is a number carried by the message, chosen by the sender, and it affects only the order of delivery.

The queue is not unbounded. mq_maxmsg and mq_msgsize are fixed when it is created, and both are capped by the system. A full queue blocks the sender, which is the next chapter.

A System V message type is not a priority. The receiver asks for a type it wants; it does not get the most urgent.

Quick revision

  • A message queue holds whole messages, belongs to the kernel, and survives the

processes that use it.

  • POSIX: mq_open, mq_send, mq_receive, mq_getattr, mq_close, mq_unlink. Names begin

with one slash. Link with -lrt.

  • mq_close ends this process's use; mq_unlink removes the queue for everybody.
  • mq_receive returns the highest priority message, and among equals the oldest.
  • On Linux a POSIX queue appears under /dev/mqueue, so leftovers can be found with ls.
  • System V: msgget, msgsnd, msgrcv, msgctl, named by a numeric key, selected by

message type, listed by ipcs -q.

munotes.in95

Message Queues

  • The producer and consumer problem with a queue needs no semaphores at all: the queue's size

and the kernel's blocking do the work.

Test yourself

  1. Give two things a message queue does that a pipe cannot. It preserves message boundaries,

and it outlives the processes that use it, so a sender and a receiver need never run at the same time.

  1. Distinguish mq_close from mq_unlink. mq_close ends the calling process's use of the

queue. mq_unlink removes the queue from the system.

  1. Three messages are sent with priorities 1, 9 and 1. In what order are they received? The

priority 9 message first, then the two priority 1 messages in the order they were sent.

  1. A program using mq_send fails to link. What is missing? -lrt, the real time library.
  2. How is a System V queue's message type different from a POSIX priority? The receiver names

the type it wants, so one queue can carry several independent conversations. A priority only decides which waiting message is delivered first.

  1. Solve producer and consumer with a message queue. The producer sends each item and blocks

if the queue is full; the consumer receives each item and blocks if it is empty. No semaphores and no mutex are needed, because the queue is bounded and the kernel does the waiting.

  1. What is the cost of that simplicity? Two system calls and two copies per item, against a

memory write in the shared memory solution.

Contents This chapter on its own page

munotes.in96

Chapter Twenty-Five

Blocking and Non-blocking Communication

Syllabus topic Computer Science Practical 3, Module 1, "Analyze blocking vs. non-blocking communication."

In one line

A blocking call waits until it can do its job; a non-blocking call gives up immediately and tells you it could not.

Why there are four combinations and not two

Sending and receiving are separate decisions, so there are four.

Blocking sendNon-blocking send
Blocking receivethe simplest to write; both sides wait as neededthe sender never waits; the receiver does
Non-blocking receivethe sender waits; the receiver pollsneither waits; both must handle failure

A blocking call is also called synchronous, and a non-blocking one asynchronous. Those are the words a question is likely to use.

The right way to think about it is whose time is being spent. A blocking call spends the calling process's time doing nothing, which is free, because the process is in the waiting state and the processor is given to somebody else. A non-blocking call spends the calling process's time asking again, which is not free at all.

What a blocking call actually does

It is worth being exact, because a beginner imagines a loop.

  1. The process calls read, and there is nothing to read.
  2. The kernel moves the process from running to waiting and puts it in the queue for that

pipe or that queue (Chapter seventeen).

  1. The scheduler gives the processor to somebody else.

The blocked process uses no processor at all.

  1. Data arrives. The kernel moves the process to ready.
  2. Eventually it is chosen, read returns, and the process continues.

So "blocking" does not mean "wasting the processor". It means the opposite: it is the cheapest possible way to wait. The mistake is to think a program must poll to be efficient; almost always the reverse is true.

Making a call non-blocking

One flag, O_NONBLOCK, given when the channel is opened or set afterwards with fcntl. Then a call that would have waited returns -1 at once and sets the error number to EAGAIN, which means "nothing now, try again".

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <string.h>
#include <errno.h>
#include <fcntl.h>
#include <unistd.h>

int main(void)
{
    int ends[2];
    if (pipe(ends) < 0) {
        perror("pipe");
        return 1;
    }
    char buf[32];

    /* 1. non-blocking read of an empty pipe */
    fcntl(ends[0], F_SETFL, O_NONBLOCK);
    ssize_t n = read(ends[0], buf, sizeof buf);
    printf("empty pipe, non-blocking read returned %zd, errno says %s\n",
        n, strerror(errno));

    /* 2. the same read once there is something */
    write(ends[1], "data", 4);
    n = read(ends[0], buf, sizeof buf);
    printf("after a write, the same read returned %zd\n", n);

    /* 3. a non-blocking write into a pipe with no room left */
    fcntl(ends[1], F_SETFL, O_NONBLOCK);
    long long written = 0;
    char block[4096];
    memset(block, 'x', sizeof block);
    while ((n = write(ends[1], block, sizeof block)) > 0) {
        written += n;
    }
    printf("the pipe took %lld bytes and then said %s\n", written,
        strerror(errno));
    close(ends[0]);
    close(ends[1]);
    return 0;
}
munotes.in97

Blocking and Non-blocking Communication

$ gcc -std=c17 -Wall -Wextra -o nonblock nonblock.c
$ ./nonblock
empty pipe, non-blocking read returned -1, errno says Resource temporarily unavailable
after a write, the same read returned 4
the pipe took 65536 bytes and then said Resource temporarily unavailable

Three facts from one program.

  1. An empty pipe read -1 with EAGAIN, whose message is "Resource temporarily unavailable".

Had the pipe been blocking, the program would still be sitting there.

  1. The same call succeeded the moment there was data. Non-blocking changes when the call

returns, not what it does.

  1. The pipe's capacity is a real, measurable number: 65,536 bytes on this machine. A blocking

write would have waited at that point; the non-blocking one refused. That number is the flow control of Chapter twenty three, measured.

Which to choose, and the rule that answers the viva question

Use blocking unless you have another job to do. That is the whole answer, and the reasons are:

  • a blocked process costs nothing, while a polling process costs a processor;
  • blocking code is shorter and has fewer states, so it has fewer bugs;
  • the kernel wakes you at exactly the right moment, and a poll cannot.

Use non-blocking when one process must serve many channels, which is the case blocking cannot handle: a server with a hundred connections cannot block on the first one, because the other ninety nine would wait. And that is where the third answer comes in, because polling a hundred channels in a loop is also wrong.

The needThe right tool
One channel, nothing else to doblocking
Many channels, one processselect or poll: block on all of them at once and be told which is ready
Something else to do meanwhile, on a timernon-blocking, checked when convenient
Nevera tight loop of non-blocking calls

A tight loop of non-blocking calls is called busy waiting and it is the wrong answer to every question in this chapter. It burns a whole processor to learn nothing. The same mistake appears in Chapter thirty seven as the spinlock, where it is occasionally right, and there the reason it is right is precisely measured.

Worked example: two designs for one job

A program must read from a pipe and also update a clock on the screen every second.

With blocking alone: impossible in one process. Blocked on the pipe, it cannot touch the clock; busy waiting on the pipe, it wastes the machine. Two processes, or two threads (Chapter twenty seven), or the next answer.

With select: one call waits on the pipe and a one second timeout at the same time. It returns when either happens. The process uses no processor while waiting and never misses a second. This is the design a real program uses, and it is blocking, on several things at once.

munotes.in98

Blocking and Non-blocking Communication

With non-blocking and a loop: the program reads, gets EAGAIN, sleeps a little, tries again. It works, it wastes a little, and it may be up to the sleep interval late. Acceptable for a small tool and wrong for a server.

Distinctions that carry marks

Blocking, synchronousNon-blocking, asynchronous
Returnswhen the work can be doneimmediately
On failuredoes not fail: it waits-1 with EAGAIN
Processor used while waitingnonenone, if you wait properly; a whole one if you poll
Codeshortmust handle "not yet" everywhere
Right whenthere is one channel and nothing else to domany channels, or other work
Non-blockingBusy waiting
Isa property of the calla way of using it
Costsnothing by itselfa whole processor
Acceptableoftenalmost never

What it does not mean

Blocking does not mean the machine stops. One process waits; everything else runs. A blocked process is in the waiting state of Chapter fifteen.

Non-blocking is not faster. It returns sooner. The work happens at the same time either way.

EAGAIN is not an error in the ordinary sense. It means "not now". A program that reports it to the user as a failure is wrong.

Asynchronous input and output is a third thing. Strictly, asynchronous means the kernel does the work and tells you when it is finished, which is aio_read and its family. A non-blocking call does no work at all if it cannot finish at once. An examination that uses "asynchronous" loosely means non-blocking.

Quick revision

  • Blocking, or synchronous: the call waits. Non-blocking, or asynchronous: it returns -1

with EAGAIN at once.

  • Four combinations, because send and receive choose separately.
  • A blocked process is in the waiting state and uses no processor. Blocking is the cheap

way to wait, not the wasteful one.

  • O_NONBLOCK, set at open or with fcntl, makes a channel non-blocking.
  • A tight loop of non-blocking calls is busy waiting and wastes a whole processor.
  • Many channels in one process: select or poll, which blocks on all of them at once.
  • The measured capacity of a pipe on the lab machine is 65,536 bytes; a blocking write waits

there, a non-blocking one refuses.

Test yourself

  1. Define blocking and non-blocking communication. A blocking call does not return until it

can do its job. A non-blocking call returns immediately, with -1 and EAGAIN, if it cannot.

munotes.in99

Blocking and Non-blocking Communication

  1. How many combinations are there, and why? Four, because sending and receiving are

independent choices.

  1. Does a blocking read waste the processor? No. The process is moved to the waiting state

and the processor is given to another process. It is the cheapest way to wait.

  1. What does EAGAIN mean and what should a program do about it? Nothing is available now.

Try again later, ideally when told to by select or poll rather than by a loop. 5. A server handles a hundred connections in one process. Which model, and why not the others? Neither plain blocking nor a poll loop: blocking on one connection ignores the rest, and polling all of them burns a processor. Use select or poll, which blocks on all of them at once and reports which are ready.

  1. Why is busy waiting almost always wrong? It uses a whole processor to discover, over and

over, that there is nothing to do, and the kernel could have woken the process at exactly the right moment for nothing.

Contents This chapter on its own page

munotes.in100

Chapter Twenty-Six

What a Thread Is

Syllabus topic Module 1, "Processes - Threads - Overview"

In one line

A thread is one flow of execution inside a process, and a process may have many of them, all sharing its memory.

The form to write: a thread is the basic unit of processor utilisation. It has a thread id, a program counter, a register set and a stack of its own, and it shares its code, its data and its operating system resources with the other threads of the same process.

Why threads exist at all

Chapter twenty two showed two processes sharing memory, and it took four system calls and a great deal of care. Now ask the opposite question: what if two flows of execution wanted to share everything, all the time, by default?

That is a thread. And the reason to want it is that most real programs have several things to do at once on the same data:

  • a browser drawing a page while downloading the next image into the same document;
  • a word processor checking spelling while you type into the same text;
  • a web server answering two hundred requests out of the same cache.

Doing each of those with separate processes means copying or sharing the data explicitly. Doing them with threads means they are simply looking at the same variables.

What is shared and what is private

This is the table the topic reduces to and it is asked directly.

Shared by every thread of a processPrivate to each thread
the code, the text sectionthe program counter
the data section, the global variablesthe registers
the heapthe stack
open files and descriptorsthe thread id
the current directory, the userits signal mask, its errno
the process idits own place in a queue

Every thread has its own stack, and that is not optional. A stack holds the frames of the function calls in progress, and two threads are in different functions at the same moment, so they cannot share one. The stack is what makes a thread a thread.

A global variable is shared and a local variable is not, and that one sentence is where half the marks in this whole row are. A local variable lives on the thread's own stack; a global lives in the data section, which is shared. Chapter thirty four's race condition is exactly a shared global being changed by two threads, and Chapter thirty seven's fix is a lock around it.

The four benefits

MU's text book gives four and an examination asks for them by name.

  1. Responsiveness. A program can keep answering the user while part of it is busy or blocked.

A browser whose download thread is waiting for the network still scrolls.

munotes.in101

What a Thread Is

  1. Resource sharing. Threads share memory by default, with no system call and no setup. Two

processes must ask for shared memory; two threads already have it.

  1. Economy. Making a thread is far cheaper than making a process, and switching between two

threads of one process is cheaper than switching between two processes. Chapter sixteen's worked example says why: the memory mapping does not change.

  1. Scalability, also called utilisation of multiprocessor architectures. A process with

one thread can use one processor however many the machine has. Threads can run on all of them at once.

The fourth is the one that has changed in importance. When the four were first written, machines with more than one processor were rare. Every phone now has six or eight cores, and a single threaded program uses one of them.

Seeing threads on a real machine

A thread is not a process, but Linux shows both in the same table if you ask.

$ ps -o pid,nlwp,comm -p $$ --no-headers
      9    1 bash
$ grep Threads /proc/self/status
Threads:	1
$ ls -d /proc/$$/task/*
/proc/9/task/9

nlwp is the number of light weight processes, which is what Linux calls a thread, and Threads in the status file says the same thing. Each one has a directory of its own under task. A program with one thread has one of each, and Chapter twenty seven's program has ten.

Worked example: one job, three designs

A program must read a hundred files and count the words in each.

One process, one thread. Open a file, read it, count, next. While the disk fetches a block the program does nothing at all. On a two processor machine it uses one processor, and badly.

A hundred processes. Each counts one file. They genuinely run at once and use every processor. But each costs a full process creation, each has its own address space, and collecting the hundred answers needs inter-process communication: shared memory or a pipe per child, and Chapter twenty two or twenty three of work.

One process, several threads. Each thread takes a file from a shared list and adds its count to a shared total. They run at once, they use every processor, creation is cheap, and the answers are simply added into a shared variable with no communication machinery at all.

The third design is best and it is also the one that can go wrong in a way the other two cannot. "Adds its count to a shared total" is a race condition (Chapter thirty four) unless the addition is protected (Chapter thirty seven). Threads buy sharing and sell safety, and the whole of MU's process synchronisation row is the price.

munotes.in102

What a Thread Is

Distinctions that carry marks

ProcessThread
Address spaceits ownshared with its siblings
Creation costhigh: a whole address spacelow
Switching costhigh: the memory mapping changeslow: it does not
Communicationneeds shared memory or messagesordinary variables
Protection between themfull: one cannot touch the other's memorynone at all
If one crashes badlythe others survivethe whole process usually dies
Also calleda heavyweight processa lightweight process
Its own stackThe shared heap
Holdslocal variables, call frames, the return addressanything malloc gave out
Visible to other threadsno, in practiceyes
Growsper threadfor the whole process

What it does not mean

Threads do not make a program faster by themselves. They let it use more processors and stop waiting. A program that is limited by the disk gets nothing from more threads.

A thread is not a lightweight process in the sense of being a small process. It has no address space of its own at all. Linux calls it a light weight process because it implements both with one mechanism.

Threads do not remove the need for synchronisation. They create it. Two processes are protected from each other by the hardware. Two threads are protected from each other by nothing.

A single threaded program is not obsolete. It is simpler, it cannot race, and for most tasks it is the right answer.

Quick revision

  • A thread is the basic unit of processor utilisation: a thread id, a program counter,

registers and a stack of its own, sharing code, data, the heap and open files with its siblings.

  • Shared: text, data, globals, heap, open files, the process id. Private: program

counter, registers, stack, thread id.

  • A global variable is shared; a local variable is on the thread's own stack and is not.
  • Four benefits: responsiveness, resource sharing, economy, scalability on

several processors.

  • Threads are cheaper to make and cheaper to switch between than processes, because the address

space does not change.

  • Linux calls a thread a light weight process; ps -o nlwp, /proc/<pid>/status and

/proc/<pid>/task all count them.

  • Threads buy sharing and sell safety: everything in MU's synchronisation row is the price.

Test yourself

  1. Define a thread. The basic unit of processor utilisation, with its own thread id, program

counter, registers and stack, sharing its code, data and operating system resources with the other threads of the same process.

  1. What does a thread have of its own, and what does it share? Its own program counter,

registers, stack and thread id. It shares the text, the data section and globals, the heap, the open files and the process id.

  1. Why must each thread have its own stack? Because the stack holds the frames of the calls
munotes.in103

What a Thread Is

in progress, and two threads are inside different functions at the same moment.

  1. Name the four benefits of threads. Responsiveness, resource sharing, economy, and

scalability on a machine with several processors.

  1. Why is switching between two threads cheaper than between two processes? The address space

does not change, so the memory management hardware and its translation cache are untouched. 6. A program counts words in a hundred files using ten threads that add to one total. What is the danger? The shared total is changed by ten threads at once, which is a race condition. The addition must be protected by a lock.

  1. Is a local variable shared between threads? No. It lives on the calling thread's own

stack. A global variable is shared, because it lives in the shared data section.

Contents This chapter on its own page

munotes.in104

Chapter Twenty-Seven

Making Threads, and Waiting for Them

Syllabus topic Module 1, "Processes - Threads - Overview"; Computer Science Practical 3, Module 1, "Practice thread creation and basic thread lifecycle using standard libraries"

In one line

pthread_create starts a new flow of execution inside this process, and pthread_join waits for one to finish.

The life cycle, which is what the practical asks for

A thread has four states in its life and each has a call or an event that moves it on.

StateReached byLeft by
Created, or readypthread_create returnsthe scheduler choosing it
Runningbeing chosenblocking, being preempted, or finishing
Blockedwaiting for a lock, a condition, input or outputthe thing it waited for
Terminatedreturning from its function, or pthread_exitbeing joined, which releases its record

A terminated thread that nobody joins is the thread version of Chapter twenty one's zombie. Its stack and its record stay allocated until somebody joins it. A program that creates threads in a loop and never joins them runs out of memory, and the two ways to avoid that are to join every thread or to detach it, which tells the library that nobody will ever join it and it may be cleaned up the moment it ends.

The five calls

CallWhat it does
pthread_create(&id, attr, function, argument)start a thread running function(argument)
pthread_join(id, &result)wait for that thread, and collect what it returned
pthread_exit(value)end this thread with that value
pthread_detach(id)nobody will join this one; clean it up when it ends
pthread_self()which thread am I

The function takes one void and returns one void . That is the whole interface, and it is why every real program passes a pointer to a structure when it needs more than one argument.

Compile with -pthread. Without it the program may link and then behave strangely, because -pthread sets a compiler flag as well as adding the library.

Ten threads, made and joined

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>

#define HOW_MANY 10

struct job {
    int  number;
    long answer;
};

static void *work(void *argument)
{
    struct job *j = argument;

    j->answer = (long)j->number * j->number;
    printf("thread %d computed %ld\n", j->number, j->answer);
    return NULL;
}

int main(void)
{
    pthread_t id[HOW_MANY];
    struct job jobs[HOW_MANY];

    for (int i = 0; i < HOW_MANY; i++) {
        jobs[i].number = i + 1;
        if (pthread_create(&id[i], NULL, work, &jobs[i]) != 0) {
            perror("pthread_create");
            return 1;
        }
    }
    long total = 0;
    for (int i = 0; i < HOW_MANY; i++) {
        pthread_join(id[i], NULL);        /* wait, in order */
        total += jobs[i].answer;
    }
    printf("every thread joined, and the total is %ld\n", total);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o ten ten.c
$ ./ten | grep thread | sort | tail -2
thread 8 computed 64
thread 9 computed 81
$ ./ten | grep total
every thread joined, and the total is 385
$ ./ten | grep total
every thread joined, and the total is 385
munotes.in105

Making Threads, and Waiting for Them

The total is the same every run and the order of the lines is not. The sum of the squares from one to ten is 385 whatever order the threads ran in, because each thread wrote to its own job. That is what makes this program correct, and Chapter thirty four is the same program written so that it is not.

The order really does change

$ for i in $(seq 20); do ./ten | head -1; done | sort -u > firsts.txt
$ n=$(wc -l < firsts.txt)
$ [ "$n" -gt 1 ] && echo "over twenty runs, $n different threads reached the screen first"
over twenty runs, 3 different threads reached the screen first

Three runs of the same program, and the thread that reached the screen first was not always the same. Nothing in the program decides that. It is the scheduler, and a program that depends on which thread runs first is wrong even when it works.

Counting the threads from outside

The kernel can be asked how many threads a process has while it is running.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <string.h>
#include <pthread.h>
#include <unistd.h>

static void *wait_a_while(void *unused)
{
    (void)unused;
    sleep(2);
    return NULL;
}

int main(void)
{
    pthread_t id[4];

    for (int i = 0; i < 4; i++) {
        pthread_create(&id[i], NULL, wait_a_while, NULL);
    }
    FILE *f = fopen("/proc/self/status", "r");
    char line[128];
    while (fgets(line, sizeof line, f)) {
        if (strncmp(line, "Threads:", 8) == 0) {
            printf("%s", line);
        }
    }
    fclose(f);
    for (int i = 0; i < 4; i++) {
        pthread_join(id[i], NULL);
    }
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o count-threads count-threads.c
$ ./count-threads
Threads:	5

Five: the four that were created and the one that created them. The first thread of a process is a thread like any other, and main runs on it. There is no such thing as a process with no threads.

Passing an argument, and the trap in it

Passing the address of a loop variable to every thread is the commonest thread bug there is.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>

static void *show(void *argument)
{
    int *n = argument;

    printf("%d ", *n);
    return NULL;
}

int main(void)
{
    pthread_t id[4];

    for (int i = 0; i < 4; i++) {
        pthread_create(&id[i], NULL, show, &i);   /* the SAME address, four times */
    }
    for (int i = 0; i < 4; i++) {
        pthread_join(id[i], NULL);
    }
    printf("\n");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o wrong wrong.c
$ ./wrong
4 4 4 4

Four threads, and every one of them printed the same number. They were all given the address of i, which by the time they ran had reached 4. The fix is what ten.c did: give each thread a pointer to something of its own.

munotes.in106

Making Threads, and Waiting for Them

Distinctions that carry marks

forkpthread_create
Createsa process with its own address spacea thread sharing this one
The new one startsat the return from forkat the function given to it
Returnstwiceonce
Costhighlow
Collected withwaitpthread_join
Uncollected leavesa zombie processa thread record and its stack
pthread_joinpthread_detach
Waitsyesno
Gets the return valueyesno
After itthe record is releasedthe record is released when the thread ends
Use whenyou need the answer or the orderingyou do not care when it finishes

What it does not mean

A thread function is not main. It takes one void and returns one void , and returning from it ends only that thread.

pthread_exit in main does not end the process. It ends the first thread and lets the others run on, which is occasionally what you want and usually a surprise.

Joining is not synchronisation between running threads. It waits for one to finish. Two threads that must co-ordinate while both are running need Chapters thirty seven to forty two.

Quick revision

  • pthread_create(&id, attr, function, argument) starts a thread; pthread_join(id, &result)

waits for it and collects its value.

  • The function takes one void and returns one void . Pass a structure for more.
  • Compile with -pthread.
  • Life cycle: created, running, blocked, terminated. A terminated thread's record survives until

it is joined or it was detached: the thread version of a zombie.

  • The first thread of a process runs main, so a process with four created threads reports five.
  • The order in which threads run is the scheduler's business and changes between runs. A

program that depends on it is wrong.

  • Never pass the address of the loop variable to every thread: they all see its final value.

Test yourself

  1. Which calls create a thread and wait for one? pthread_create and pthread_join.
  2. What is the signature of a thread function? It takes one void * and returns one

void *.

  1. What happens to a thread nobody joins? Its record and stack stay allocated, like a zombie

process. Join it, or detach it so the library cleans it up when it ends.

  1. A process creates four threads. How many does the kernel report and why? Five. The thread

running main is a thread like any other. 5. Four threads are given &i, the address of the loop variable, and all print 4. Why, and what is the fix? They share one variable, which had reached its final value before they ran. Give each thread a pointer to its own data.

munotes.in107

Making Threads, and Waiting for Them

  1. Why does the total in ten.c never change while the order of the lines does? Each thread

wrote to its own structure, so nothing was shared and no result depended on the order. Only the moment each thread reached the screen depended on the scheduler.

Contents This chapter on its own page

munotes.in108

Chapter Twenty-Eight

Multicore Programming, and the Limit on It

Syllabus topic Module 1, "Processes - Multicore Programming"

In one line

A machine with several processors can run several threads at once, and the speed a program gains is limited by the part of it that cannot be split.

The words, which are examined separately

WordMeans
Concurrencymore than one task is in progress. They may take turns on one processor
Parallelismmore than one task is executing at this instant. This needs more than one processor

Concurrency does not require parallelism, and this distinction is asked directly. A single processor running twenty threads by taking turns is concurrent and not parallel. Parallelism is a property of the machine; concurrency is a property of the program.

WordMeans
Coreone processor on one chip. A chip with eight of them is an eight core processor
Multicoreseveral cores on one chip, sharing memory
Multiprocessorseveral separate processors, historically several chips

The two kinds of parallelism

Data parallelismTask parallelism
Split upthe datathe work
Each thread doesthe same operation on a different parta different operation
Exampleeight threads each add up one eighth of a million numbersone thread draws, one plays sound, one saves
Scales withthe size of the datathe number of different jobs, which is fixed

Data parallelism scales and task parallelism does not. A program with four different jobs cannot use sixteen cores by task parallelism however hard it tries; a program adding up a million numbers can use as many as you give it. Real programs use both.

The five challenges

MU's text book lists five and they are asked as a list.

  1. Identifying tasks. Finding the parts that are genuinely independent. This is the hard one

and it cannot be automated in general.

  1. Balance. Each thread should get about the same amount of work. A thread that finishes

early leaves a core idle while the others work.

  1. Data splitting. The data has to be divided so that each thread's part is separate. Where

two threads need the same item, a lock is needed and the gain shrinks.

  1. Data dependency. Where one task needs another's result, they cannot run at the same time

and must be synchronised.

  1. Testing and debugging. A parallel program has an enormous number of possible orderings,

and a bug may appear in one of them. A test that passes says far less than it does for a single threaded program.

The fifth is the one students underrate and the one professional programmers fear. Chapter thirty four's race condition appears on some runs and not others, on some machines and not others.

Amdahl's law, worked

The limit on what parallelism can buy. If a fraction S of a program is strictly serial, meaning it cannot be split at any price, and the rest is perfectly parallel over N cores, then

munotes.in109

Multicore Programming, and the Limit on It

speedup <= 1 / (S + (1 - S) / N)

Work it for a program that is 25 per cent serial, so S = 0.25, on eight cores.

S + (1 - S) / N = 0.25 + 0.75 / 8 = 0.25 + 0.09375 = 0.34375

speedup <= 1 / 0.34375

One divided by 0.34375 is 32 / 11, which is about 2.91.

Eight cores, and under three times the speed. That is the whole point of the law, and it is why a question asks for the number rather than the principle.

Now the same program on more and more cores, to see where it stops paying.

CoresSerial partParallel partDivisorSpeedup
10.250.7511
20.250.3750.6251.6
40.250.18750.4375about 2.29
80.250.093750.34375about 2.91
160.250.0468750.296875about 3.37
10000.250.000750.25075about 3.99

The last row is the answer to the question this law exists for. A thousand cores give not quite four times the speed, because the limit as N grows without bound is 1 / S, and 1 / 0.25 = 4. A program that is 25 per cent serial can never be more than four times faster, on any machine that will ever be built.

So the lever that matters is not the number of cores. It is S. Halving the serial fraction from 0.25 to 0.125 raises the ceiling from 4 to 8.

The machine this book runs on

$ nproc
2
$ grep -c ^processor /proc/cpuinfo
2
$ getconf _NPROCESSORS_ONLN
2

Two cores, so no program on this machine can be more than twice as fast by parallelism, whatever its serial fraction. That is worth knowing before Chapter thirty one measures a threaded program here: the ceiling is 2, and the measurement comes out below it.

Worked example: is it worth threading?

A report takes 100 seconds: 20 seconds reading a file, which cannot be split because the disk is one device, and 80 seconds of arithmetic, which can. The office machine has four cores.

S = 20 / 100 = 0.2.

speedup <= 1 / (0.2 + 0.8 / 4) = 1 / (0.2 + 0.2) = 1 / 0.4 = 2.5

So 100 seconds becomes about 40. That is worth doing.

Now the same report where the file is bigger: 60 seconds reading, 40 computing. S = 0.6.

speedup <= 1 / (0.6 + 0.4 / 4) = 1 / (0.6 + 0.1) = 1 / 0.7

One divided by 0.7 is 10 / 7, which is about 1.43.

100 seconds becomes 70, for all the risk of Chapter thirty four. The right decision here is to make the reading faster, not to add threads, and the law is what says so.

munotes.in110

Multicore Programming, and the Limit on It

Distinctions that carry marks

ConcurrencyParallelism
Needs several processorsnoyes
Meansseveral tasks in progressseveral tasks executing now
A property ofthe programthe machine
On one core, twenty threadsconcurrentnot parallel
Data parallelismTask parallelism
Dividedthe datathe work
Threads dothe same thing to different datadifferent things
Scales with more coresyesonly up to the number of distinct jobs

What it does not mean

Amdahl's law is not a reason not to use threads. It is a way of deciding how much they will buy before writing them.

A core is not a processor in the old sense. Eight cores on one chip share memory and often share caches, which is why two threads on one chip can interfere with each other in ways two separate machines could not.

Perfectly parallel does not mean free. Creating threads, dividing the data and joining at the end all cost, and on small inputs they cost more than they save. Chapter thirty one measures exactly that.

Quick revision

  • Concurrency: several tasks in progress, possibly taking turns on one processor.

Parallelism: several executing at this instant, which needs several cores.

  • Data parallelism splits the data and scales; task parallelism splits the work and does

not.

  • Five challenges: identifying tasks, balance, data splitting, data dependency, testing and

debugging.

  • Amdahl's law: speedup is at most 1 / (S + (1 - S) / N), where S is the strictly serial

fraction.

  • 25 per cent serial on eight cores gives 2.909090909090909, and on any number of cores at all

the limit is 1 / S = 4.

  • The lever that matters is S, not N.
  • The lab machine has two cores, so nothing measured on it can beat a speedup of 2.

Test yourself

  1. Distinguish concurrency from parallelism. Concurrency means several tasks are in progress

and may share one processor by taking turns. Parallelism means several are executing at the same instant, which requires more than one core.

  1. Distinguish data parallelism from task parallelism, with an example of each. Data

parallelism divides the data and gives every thread the same operation, for example eight threads each summing one eighth of an array. Task parallelism gives each thread a different job, for example one drawing and one saving. Only the first scales with the number of cores.

  1. Name the five challenges of multicore programming. Identifying tasks, balance, data

splitting, data dependency, and testing and debugging.

  1. State Amdahl's law. The speedup from N processing cores is at most 1 / (S + (1 - S) / N),
munotes.in111

Multicore Programming, and the Limit on It

where S is the fraction of the program that must be done serially.

  1. A program is 25 per cent serial. What is the most it can ever be sped up? Four times,

because the limit of the law as N grows is 1 / S = 1 / 0.25 = 4. 6. A 100 second job spends 60 seconds on unsplittable reading. Is threading the arithmetic worth it on four cores? Barely. The speedup is at most 10 / 7, which is about 1.43, so 100 seconds becomes about 70. Making the reading faster is the better investment.

Contents This chapter on its own page

munotes.in112

Chapter Twenty-Nine

The Three Multithreading Models

Syllabus topic Module 1, "Processes - Multithreading Models"

In one line

A thread the programmer makes has to be mapped onto a thread the kernel schedules, and there are three ways to do it.

Why there are two kinds of thread at all

This is the idea the whole topic rests on, and it has to come first.

A user thread is managed by a library in the program: making one, switching between them and scheduling them are all done by library code, in user mode, with no system call. The kernel does not know they exist.

A kernel thread is managed by the operating system: the kernel knows about it, schedules it, and can put it on a processor.

Only a kernel thread can be given a processor. So a user thread is only ever running because some kernel thread is running it. The three models are the three answers to the question of how many user threads are mapped onto how many kernel threads.

Many to one

Many user threads are mapped onto one kernel thread. All the thread management is in the library and is very fast, because a switch between two user threads is a jump, not a system call.

Two consequences follow, and both are examined.

  1. If one thread makes a blocking system call, the whole process blocks. The kernel sees one

thread, and that thread is blocked, so no other user thread can run either, even though they are ready.

  1. The process can never use more than one processor. There is one kernel thread and one

kernel thread can be on one core.

Used by: the older Java green threads, and some small libraries. Effectively obsolete on a machine with several cores.

One to one

Each user thread is mapped onto its own kernel thread. This is what pthread_create does on Linux and on Windows.

Gaineda blocking call blocks only that thread
Gainedthreads run on as many cores as the machine has
Lostmaking a user thread now means making a kernel thread, which costs a system call and kernel memory
Limitthe number of threads a process may have is limited by the system

This is the model that won, because both of its gains matter on real machines and its cost has been driven down. clone on Linux, which Chapter eleven's trace showed, is what makes a kernel thread, and it is the same call that makes a process.

Many to many

Many user threads are multiplexed onto a smaller number of kernel threads. The library keeps, say, eight kernel threads on an eight core machine and runs a hundred user threads on them.

Gainedas many user threads as the program likes, with the cheapness of the library
Gainedreal parallelism, up to the number of kernel threads
Lostthe library and the kernel are both scheduling, and they cannot see each other's decisions
munotes.in113

The Three Multithreading Models

The two level model, sometimes listed as a fourth, is many to many with the addition that a particular user thread may be bound to a kernel thread of its own when it needs to be.

The difficulty of many to many is the reason one to one won. When a user thread blocks in the kernel, the library would like to run another user thread on that kernel thread, and it cannot know that the block happened. Solving it needs the kernel to tell the library, a mechanism named scheduler activations, and it was hard enough that most systems gave up and used one to one instead.

Which model this machine uses

$ getconf GNU_LIBPTHREAD_VERSION
NPTL 2.39
$ grep Threads /proc/self/status
Threads:	1

NPTL is the Native POSIX Thread Library, and it is one to one. The proof is in Chapter twenty seven's count-threads, where a process with four created threads reported five to the kernel: the kernel knew about every one of them, which a many to one library would have hidden.

Worked example: one blocking read, three models

A program has ten threads. One of them reads from a file and waits. The machine has four cores.

Many to one. The kernel has one thread and it is blocked. All ten user threads stop, including the nine with work to do. Nothing runs. The machine's other three cores were never usable anyway.

One to one. The kernel has ten threads and one is blocked. The other nine are ready, and up to four of them run at once. The blocked one wakes when the read completes.

Many to many with four kernel threads. One of the four kernel threads is blocked. The library has three left and nine user threads that want to run, so three run at once and the fourth core is idle until the read finishes. Better than many to one and worse than one to one, for this case.

Change the case and the ranking changes: with ten thousand user threads, one to one means ten thousand kernel threads and the kernel's own tables become the problem, while many to many still keeps four. That is why the model is a trade and not a ranking.

Distinctions that carry marks

Many to oneOne to oneMany to many
Kernel threadsoneone per user threadfewer than the user threads
Kernel knows about user threadsnoyesonly the kernel threads
A blocking callblocks the whole processblocks that thread onlyblocks one kernel thread
Uses several coresnoyesup to its kernel thread count
Switch costa jump, no system calla system calla jump inside the library
Thread creation costvery lowhigherlow
Exampleold green threadsLinux NPTL, WindowsSolaris before 9, some runtimes
munotes.in114

The Three Multithreading Models

User threadKernel thread
Managed bya library, in user modethe operating system
Can be given a processornot directlyyes
Creationno system calla system call
The kernel schedulesnoyes

What it does not mean

A user thread is not a fake thread. It is a real flow of execution with its own stack. What it lacks is the kernel's knowledge of it.

One to one is not always best. It is best for the usual number of threads on the usual machine. For very large numbers of threads a library on top is still used, which is what the "green thread" runtimes of several modern languages are.

Many to many is not a compromise nobody uses. It is what a language runtime does when it offers cheap threads of its own on top of the operating system's.

Quick revision

  • A user thread is managed by a library in user mode; a kernel thread is managed and

scheduled by the operating system. Only a kernel thread can be given a processor.

  • Many to one: cheap, and one blocking call blocks the whole process; can never use more than

one core.

  • One to one: a kernel thread each, so a block affects one thread and the program uses every

core; costs a system call per thread and the count is limited.

  • Many to many: many user threads on fewer kernel threads; cheap and parallel, but two

schedulers that cannot see each other.

  • The two level model is many to many plus the ability to bind one user thread to a kernel

thread.

  • Linux uses one to one, through NPTL, and clone is the call.

Test yourself

  1. Distinguish a user thread from a kernel thread. A user thread is created and scheduled by

a library in user mode and the kernel does not know it exists. A kernel thread is known to and scheduled by the operating system, and only a kernel thread can be put on a processor.

  1. Name the three models. Many to one, one to one, and many to many.
  2. Why does a blocking system call stop everything in the many to one model? Because the

kernel sees one thread for the whole process. When it blocks, there is no other kernel thread for the library's ready user threads to run on.

  1. Give one advantage and one disadvantage of one to one. A blocking call blocks only the
munotes.in115

The Three Multithreading Models

thread that made it and the program can use every core; but each thread costs a system call and kernel memory, and the number of threads is limited.

  1. What makes many to many hard to implement well? The library and the kernel both schedule

and neither can see the other's decisions, so the library cannot tell when one of its kernel threads has blocked.

  1. Which model does Linux use, and what is the evidence? One to one. A process with four

created threads reports five threads to the kernel, so the kernel knows about each one.

Contents This chapter on its own page

munotes.in116

Chapter Thirty

Thread Pools, and Handing Work Out

Syllabus topic Computer Science Practical 3, Module 1, "Introduce concepts of thread pooling and task delegation."

In one line

A thread pool makes a fixed number of threads once, and hands them work as it arrives, instead of making a new thread for every job.

Why a thread per job stops working

The obvious design for a server is: a request arrives, make a thread, the thread answers it, the thread ends. It is simple and it fails in three ways.

  1. Creation costs. Chapter twenty nine said that on a one to one system making a user thread

means making a kernel thread, which is a system call and kernel memory. For a job that takes a microsecond, the thread costs more than the work.

  1. There is no limit. Ten thousand requests arrive at once and the program tries to make ten

thousand threads. Each has a stack, by default eight megabytes of address space, and the machine runs out of something.

  1. More threads than cores is not faster. Chapter twenty eight said the machine has two

cores. Two hundred threads on two cores do the same total work as two threads, plus two hundred times the switching.

A pool fixes all three with one idea: separate the number of workers from the number of jobs. The workers are made once, their number is chosen to fit the machine, and the jobs queue up.

What a pool is made of

PartWhat it is
The workersa fixed number of threads, each in a loop: take a job, do it, take the next
The work queuea list of jobs waiting, with a lock and a way to wait when it is empty
The submit operationput a job on the queue and wake a worker
The shutdowntell the workers to stop when the queue is empty, and join them

The queue is shared by every worker and by whoever submits, so it needs a lock. A pool is therefore not a way of avoiding Chapters thirty seven to forty two; it is one of their customers. The version below uses a mutex and a condition variable, which are Chapters thirty seven and forty two, and is written so it can be read again after them.

Task delegation, which is the other half of the practical's label

Delegation is the idea that the code which has work to do does not decide who does it or when. It describes the job, hands it over, and gets on with something else.

That separation buys three things: the submitter never blocks on the work, the number of workers can be changed without touching the submitter, and the same pool serves every part of the program.

A pool, in code that runs

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <stdlib.h>
#include <pthread.h>

#define WORKERS 4
#define JOBS    20

struct pool {
    int             job[JOBS];            /* the queue: a job is a number to square */
    int             head;                 /* next job to take */
    int             tail;                 /* next free slot */
    int             closed;
    long            total;                /* where the answers are added up */
    int             done_by[WORKERS];     /* how many jobs each worker did */
    pthread_mutex_t lock;
    pthread_cond_t  work_arrived;
};

struct hand {
    struct pool *p;
    int          which;
};

static void *worker(void *argument)
{
    struct hand *h = argument;
    struct pool *p = h->p;

    for (;;) {
        pthread_mutex_lock(&p->lock);
        while (p->head == p->tail && !p->closed) {
            pthread_cond_wait(&p->work_arrived, &p->lock);
        }
        if (p->head == p->tail && p->closed) {
            pthread_mutex_unlock(&p->lock);
            return NULL;                  /* the queue is empty and closed: go home */
        }
        int n = p->job[p->head++];
        pthread_mutex_unlock(&p->lock);

        long answer = (long)n * n;        /* the work itself, outside the lock */

        pthread_mutex_lock(&p->lock);
        p->total += answer;
        p->done_by[h->which]++;
        pthread_mutex_unlock(&p->lock);
    }
}

int main(void)
{
    struct pool p = {0};
    struct hand hands[WORKERS];
    pthread_t id[WORKERS];

    pthread_mutex_init(&p.lock, NULL);
    pthread_cond_init(&p.work_arrived, NULL);

    for (int i = 0; i < WORKERS; i++) {
        hands[i].p = &p;
        hands[i].which = i;
        pthread_create(&id[i], NULL, worker, &hands[i]);
    }
    for (int n = 1; n <= JOBS; n++) {     /* submit, and do not wait */
        pthread_mutex_lock(&p.lock);
        p.job[p.tail++] = n;
        pthread_cond_signal(&p.work_arrived);
        pthread_mutex_unlock(&p.lock);
    }
    pthread_mutex_lock(&p.lock);          /* shut down */
    p.closed = 1;
    pthread_cond_broadcast(&p.work_arrived);
    pthread_mutex_unlock(&p.lock);

    for (int i = 0; i < WORKERS; i++) {
        pthread_join(id[i], NULL);
    }
    int counted = 0;
    for (int i = 0; i < WORKERS; i++) {
        counted += p.done_by[i];
    }
    printf("%d workers did %d jobs between them\n", WORKERS, counted);
    printf("the total of the squares from 1 to %d is %ld\n", JOBS, p.total);
    return 0;
}
munotes.in117

Thread Pools, and Handing Work Out

$ gcc -std=c17 -Wall -Wextra -pthread -o pool pool.c
$ ./pool
4 workers did 20 jobs between them
the total of the squares from 1 to 20 is 2870

Twenty jobs, four threads, and every job done exactly once. The four workers did not do five each, and they were never meant to. Which worker takes which job depends on which one reaches the queue first, and that is the point: the work is balanced by the queue rather than by the programmer.

$ for i in $(seq 5); do ./pool | head -1; done | sort -u
4 workers did 20 jobs between them

Five runs and the same answer every time: every job done once, whatever order the workers took them in. That is what a correct pool guarantees and what the lock is for.

A thread per job, for comparison

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>

#define JOBS 20

struct job {
    int  number;
    long answer;
};

static void *one(void *argument)
{
    struct job *j = argument;

    j->answer = (long)j->number * j->number;
    return NULL;
}

int main(void)
{
    pthread_t id[JOBS];
    struct job jobs[JOBS];
    long total = 0;

    for (int i = 0; i < JOBS; i++) {
        jobs[i].number = i + 1;
        pthread_create(&id[i], NULL, one, &jobs[i]);
    }
    for (int i = 0; i < JOBS; i++) {
        pthread_join(id[i], NULL);
        total += jobs[i].answer;
    }
    printf("%d threads made and joined, total %ld\n", JOBS, total);
    return 0;
}
munotes.in118

Thread Pools, and Handing Work Out

$ gcc -std=c17 -Wall -Wextra -pthread -o perjob perjob.c
$ ./perjob
20 threads made and joined, total 2870
$ /usr/bin/time -f "%e seconds, %c involuntary switches" ./pool > /dev/null
0.03 seconds, 43 involuntary switches
$ /usr/bin/time -f "%e seconds, %c involuntary switches" ./perjob > /dev/null
0.02 seconds, 3 involuntary switches

Both print 2870, which is the same right answer by two routes, and then the measurement says something the chapter was not written to say. At twenty jobs the pool is not faster, and it was switched far more often: its four workers sleep on the queue and are woken, twenty times between them, while the thread per job version simply runs each thread once and joins it.

That is the honest result and it is worth more than the expected one. A pool costs a queue operation, a lock and a wake up per job; a thread costs a creation. The pool wins when a creation costs more than a wake up, which means when the jobs are many and short, and twenty squarings is neither. The next chapter measures the same kind of question and refuses to rank two figures that are within noise of each other.

How many workers

The usual answers, and a question may ask for the reasoning rather than the number.

The work isWorkers
processor boundabout the number of cores: more only adds switching
input and output boundmore than the number of cores, because most are waiting
a mixturemeasured, not guessed

A pool sized at the number of cores is wrong for a program that waits. Four cores and four workers, each waiting on a disk, leaves the machine idle. This is Chapter seventeen's mixture, seen from the programmer's side.

Distinctions that carry marks

A thread per jobA thread pool
Threads madeone per joba fixed number, once
Cost per joba thread creationa queue operation
Limit on loadnone, so the machine can be exhaustedthe queue grows, the thread count does not
Switchinggrows with the jobsfixed
Needs a lockno, if each job is separateyes, for the queue
Right whenjobs are few and longjobs are many and short

What it does not mean

A pool is not faster per job. It removes the creation cost and the unbounded thread count. A single long job is no faster in a pool.

munotes.in119

Thread Pools, and Handing Work Out

A pool does not remove synchronisation. It concentrates it in one place, the queue.

More workers is not more throughput. Past the number the machine can run, extra workers only add switching, and Chapter twenty eight's law still caps the whole thing.

Quick revision

  • A thread pool makes a fixed number of worker threads once and feeds them from a

work queue.

  • It fixes three faults of a thread per job: the creation cost, the unbounded thread count, and

the switching from having more threads than cores.

  • Parts: workers, a work queue with a lock and a condition variable, a submit operation, a

shutdown.

  • Task delegation means the code with work to do describes the job and hands it over, rather

than deciding who does it and when.

  • Workers are not given equal shares; the queue balances the work by itself.
  • Size the pool by the kind of work: about the core count for processor bound work, more for

work that waits.

  • The queue is shared, so a pool is a customer of the synchronisation chapters, not an escape

from them.

Test yourself

  1. What is a thread pool? A fixed set of threads created once, each taking jobs from a shared

work queue, so that no thread is created per job.

  1. Give three problems with making one thread per job. Creation cost dominates short jobs;

nothing bounds the number of threads, so a burst of work can exhaust the machine; and more threads than cores adds switching without adding throughput.

  1. What is task delegation? Separating the code that has work from the code that does it: the

job is described, handed over, and carried out by somebody else at some other time.

  1. How many workers should a pool have? About the number of cores for processor bound work,

more than that for work that spends its time waiting, and for a mixture it should be measured.

  1. Do the workers each do an equal share? No, and they are not meant to. Whichever worker

reaches the queue first takes the next job, which balances the work without the programmer dividing it.

  1. What must a pool's queue have that a thread per job does not need? A lock, and a way for a

worker to wait when the queue is empty: it is shared by every worker and by the submitter.

Contents This chapter on its own page

munotes.in120

Chapter Thirty-One

Measuring It: Sequential Against Threaded

Syllabus topic Computer Science Practical 3, Module 1, "Measure execution time for sequential vs threaded execution."

In one line

Threading the same work over two cores makes it about twice as fast, and measuring it is the only way to know whether it did.

What has to be measured, and how

Three different times can be measured and a question may name any of them.

TimeWhat it isRead with
Real, or wall clock, or elapsedhow long you waitedclock_gettime(CLOCK_MONOTONIC)
Userprocessor time in the program's own code, added over all threadsgetrusage, or %U
Systemprocessor time inside the kernel on the program's behalfgetrusage, or %S

Real time is the one that falls when a program is threaded. User time does not. Two threads each computing for one second use two seconds of user time and one second of real time. That single sentence is the whole measurement, and a student who reports user time as the speedup gets it backwards.

Never measure with clock(). It measures processor time, added over threads, so a perfectly threaded program appears to get slower.

CLOCK_MONOTONIC rather than CLOCK_REALTIME: the monotonic clock cannot jump backwards when the machine's time is corrected, and a negative measurement is a real bug in real programs.

The measurement, made and reported honestly

The work is deliberately pure arithmetic with no memory traffic to speak of, because anything else measures the memory system rather than the threads.

#define _GNU_SOURCE
#include <stdio.h>
#include <pthread.h>
#include <time.h>
#include <sys/resource.h>

#define SLICES 2
#define ROUNDS 5
#define SIZE   30000000L

struct part {
    long from;
    long to;
    long sum;
};

static void *add_up(void *argument)
{
    struct part *p = argument;
    long total = 0;

    for (long i = p->from; i < p->to; i++) {
        total += i % 7;
    }
    p->sum = total;
    return NULL;
}

static double now(void)
{
    struct timespec t;

    clock_gettime(CLOCK_MONOTONIC, &t);
    return (double)t.tv_sec + (double)t.tv_nsec / 1e9;
}

static double user_seconds(void)
{
    struct rusage r;

    getrusage(RUSAGE_SELF, &r);
    return (double)r.ru_utime.tv_sec + (double)r.ru_utime.tv_usec / 1e6;
}

int main(void)
{
    double best_one = 1e9;
    double best_many = 1e9;
    double used_by_best_many = 0;
    long answer_one = 0;
    long answer_many = 0;

    for (int round = 0; round < ROUNDS; round++) {
        struct part whole = {0, SIZE, 0};
        double t0 = now();
        add_up(&whole);
        double t1 = now();
        if (t1 - t0 < best_one) {
            best_one = t1 - t0;
        }
        answer_one = whole.sum;

        struct part slice[SLICES];
        pthread_t id[SLICES];
        double before = user_seconds();
        t0 = now();
        for (int i = 0; i < SLICES; i++) {
            slice[i].from = SIZE / SLICES * i;
            slice[i].to = SIZE / SLICES * (i + 1);
            slice[i].sum = 0;
            pthread_create(&id[i], NULL, add_up, &slice[i]);
        }
        long total = 0;
        for (int i = 0; i < SLICES; i++) {
            pthread_join(id[i], NULL);
            total += slice[i].sum;
        }
        t1 = now();
        if (t1 - t0 < best_many) {
            best_many = t1 - t0;
            used_by_best_many = user_seconds() - before;
        }
        answer_many = total;
    }

    /* Every line below prints the same WORDS on every run and only its digits
    change, so a recorded run can be compared with a later one. The reader is
        given the two numbers a verdict has to be read against: what the machine
        actually gave, and what the core count allows. */
        printf("the two answers agree        : %s\n",
        answer_one == answer_many ? "yes" : "NO");
    printf("speedup from %d threads      : %.2f\n", SLICES, best_one / best_many);
    printf("processors the machine gave  : %.2f\n", used_by_best_many / best_many);
    printf("the ceiling the cores allow  : %d\n", SLICES);
    printf("rounds run, best of each     : %d\n", ROUNDS);
    return 0;
}
munotes.in121

Measuring It: Sequential Against Threaded

$ gcc -std=c17 -Wall -Wextra -O2 -pthread -o measure measure.c
$ ./measure
the two answers agree        : yes
speedup from 2 threads      : 1.95
processors the machine gave  : 1.90
the ceiling the cores allow  : 2
rounds run, best of each     : 5

Five lines, and every one of them is part of the answer.

  1. The two answers agree. A faster wrong answer is not a result, so the program checks before

it reports anything. Chapter thirty four is what happens when it does not.

  1. The speedup is 1.95 and the ceiling is 2. Two threads on two cores came within three per

cent of the best that two cores allow, and the missing part is the cost of creating two threads, joining them and adding their two answers.

  1. The machine gave 1.90 processors, which is the number that makes the speedup believable.

It is the processor seconds the threaded part was charged divided by the seconds it took, so 1.90 means both cores were genuinely working for almost all of it.

  1. The ceiling is printed beside the speedup on purpose. A speedup means nothing without the

number it is being measured against.

  1. The best of five rounds, not the average. A machine shared with other work produces

occasional slow rounds and never occasional fast ones, so the best round is the closest estimate of what the program can do and the average measures the neighbours.

The third line is the one to copy into your own work. A speedup of 1.95 with 1.90 processors given is a real result. A speedup of 1.02 with 0.13 processors given is not a result about threads at all: it says the machine was busy with something else, and no conclusion about the program can be drawn from it. Always print both.

Where the time went, and a trap in reading it

$ /usr/bin/time -f "real %e, user %U, system %S" ./measure > /dev/null
real 0.18, user 0.22, system 0.00
munotes.in122

Measuring It: Sequential Against Threaded

That line is about the whole program and it is not the speedup. The program runs the single threaded version five times and the threaded version five times, so its overall processor time is roughly one processor's worth however well the threads did. Read it for one thing only: user time close to or above real time is the sign that more than one core was working at some point. For the number that matters, the program's own third line divides the processor time of the threaded part alone by its own elapsed time.

The general rule, which is worth more than any of these numbers. Before believing a parallel measurement, print the processors the machine actually gave. If it is close to the core count, the speedup is a fact about your program. If it is well below, the speedup is a fact about the machine's other work, and the honest report is that the measurement could not be taken. This book was written on a machine that did both at different times, which is how the rule was learnt.

More threads than cores

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>
#include <time.h>

#define SIZE 20000000L

static long slice_size;

static void *add_up(void *argument)
{
    long which = *(long *)argument;
    long total = 0;

    for (long i = which * slice_size; i < (which + 1) * slice_size; i++) {
        total += i % 7;
    }
    return (void *)total;
}

static double now(void)
{
    struct timespec t;

    clock_gettime(CLOCK_MONOTONIC, &t);
    return (double)t.tv_sec + (double)t.tv_nsec / 1e9;
}

static double run(int threads)
{
    pthread_t id[64];
    long which[64];
    double best = 1e9;

    slice_size = SIZE / threads;
    for (int round = 0; round < 3; round++) {
        double t0 = now();
        for (int i = 0; i < threads; i++) {
            which[i] = i;
            pthread_create(&id[i], NULL, add_up, &which[i]);
        }
        for (int i = 0; i < threads; i++) {
            pthread_join(id[i], NULL);
        }
        double t = now() - t0;
        if (t < best) {
            best = t;
        }
    }
    return best;
}

int main(void)
{
    double two = run(2);
    double many = run(32);

    /* A tenth is within this machine's noise, so the verdict is set at half:
    thirty two threads would have to be half as fast again to mean anything. */
        printf("thirty two threads against two: %s\n",
        many > two * 1.5 ? "MUCH slower" :
    many < two * 0.67 ? "MUCH faster" : "no real difference either way");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -O2 -pthread -o toomany toomany.c
$ ./toomany
thirty two threads against two: no real difference either way
munotes.in123

Measuring It: Sequential Against Threaded

Thirty two threads on two cores is not faster, and on this machine it is not measurably slower either. The work is divided sixteen ways more finely and the same two cores do it, so the total is the same and the extra cost is the thirty extra creations and the extra switching, which on this arithmetic is too small to see. The honest report is the one printed: too close to rank. A program that claimed a difference here would be reading noise.

Worked example: what to write in a practical journal

The practical asks for this measurement, and a journal entry that gets full marks says five things.

  1. The machine. Two cores, reported by nproc. Without this the numbers mean nothing.
  2. The work. Adding thirty million remainders, chosen because it touches almost no memory.
  3. The method. Best of five rounds, CLOCK_MONOTONIC, real time.
  4. The result. One thread against two threads: a factor of about 1.9.
  5. The interpretation. Below the ceiling of 2 that two cores allow, and the difference is the

cost of creating, joining and combining.

A journal that reports only step 4 is reporting a number nobody can check.

Distinctions that carry marks

Real timeUser time
Measureshow long you waitedprocessor seconds spent in the program
Added over threadsnoyes
Falls when threadedyesno
Can exceed the othernoyes, on several cores
Read withclock_gettime(CLOCK_MONOTONIC)getrusage, %U
The best roundThe average round
Measureswhat the program can dowhat the machine did while other things ran
Affected by a neighbour's workhardlygreatly
Use whenmeasuring a programmeasuring a loaded system

What it does not mean

Two cores do not make everything twice as fast. Only the part that was split, and only if it was not waiting for memory or a disk.

A factor of 1.9 is not a disappointment. It is 95 per cent of the theoretical ceiling.

clock() is not a stopwatch. It returns processor time, so a threaded program appears slower.

A single measurement is not a result. Repeat, take the best, and say how many rounds.

Quick revision

  • Measure real time with clock_gettime(CLOCK_MONOTONIC). Never with clock(), which

returns processor time added over threads.

  • Real time falls when a program is threaded; user time does not. User time larger than real

time is the proof that more than one core was used.

  • Check the answers agree before reporting a speedup.
  • Take the best of several rounds, not the average: slow rounds happen and fast ones do not.
  • Two threads on two cores measured a factor of about 1.9, below the ceiling of 2, and the

difference is creation, joining and combining.

munotes.in124

Measuring It: Sequential Against Threaded

  • Thirty two threads on two cores is not faster: the same cores do the same work.
  • Refuse to rank two measurements within about ten per cent of each other.

Test yourself

  1. Which time falls when a program is threaded, and which does not? Real, elapsed time falls.

User processor time does not, because it is added over every thread.

  1. Why must clock() not be used? It measures processor time, so a program that uses two

cores appears to take twice as long.

  1. time reports real 1.1 seconds and user 2.0 seconds. What does that tell you? That more

than one core was working: the program spent about two processor seconds within about one second of clock time.

  1. Why take the best of several rounds rather than the average? Other work on the machine

makes rounds slower and never faster, so the best round is the closest estimate of the program itself.

  1. Two threads on two cores gave a factor of 1.9. Where did the other tenth go? Into creating

and joining the threads and combining their two answers, which is work the single threaded version never did.

  1. Thirty two threads on two cores are no faster than two. Why? The cores do the same total

work whatever the number of threads; dividing it more finely only adds creations and switches.

  1. What five things should a journal entry for this practical state? The machine and its core

count, the work being timed, the method and the clock used, the result, and the interpretation against the ceiling the core count allows.

Contents This chapter on its own page

munotes.in125

Chapter Thirty-Two

The Shape of Every Concurrent Program

Syllabus topic Module 1, "Process Synchronization - General structure of a typical process"

In one line

A process that shares anything with another has four parts: the entry section, the critical section, the exit section, and the remainder.

The four sections

MU's own label names the general structure, and this is it, written the way her text book writes it.

do {

entry section

critical section

exit section

remainder section

} while (true);

SectionWhat is in itWho writes it
Entryask permission to go in, and wait if permission is refusedthe synchronisation mechanism
Criticalthe code that touches the shared thingthe programmer, and it is the program's real work
Exitsay you have finished, so somebody else may go inthe synchronisation mechanism
Remaindereverything else this process does, which touches nothing sharedthe programmer

The entry and exit sections are the only parts any of the next ten chapters change. A mutex lock is an entry section and an unlock is an exit section. A semaphore's wait is an entry section and its signal is an exit section. A monitor writes both of them for you. Recognising that makes the whole row one idea with several implementations instead of ten unrelated tricks.

What a critical section is, precisely

A critical section is a piece of code in which a process reads or writes something that other processes also read or write, and which must not be executed by two processes at the same time.

Two parts of that definition are examined and both are commonly got wrong.

  1. It is a piece of CODE, not the data. The shared variable is not the critical section; the

lines that touch it are. The same variable may be touched by three different critical sections in three different functions, and all three must exclude each other.

  1. Not being executed at the same time is a requirement, not a description. Nothing in the

language or the machine stops it. Chapter thirty four shows the machine cheerfully doing it.

Read only sharing needs no critical section at all. Ten threads reading a variable that nobody writes cannot interfere. The moment one of them writes, every reader needs protecting too.

The rule for finding one

A question may give a program and ask which part is the critical section. The test is three questions.

  1. Does this code touch something more than one process can reach: a global, a heap object shared

by pointer, a file, a device, a shared memory segment?

  1. Does it write to that thing, or read it while somebody else may write?
  2. Would two processes doing this at the same time give a result that neither ordering alone

would give?

If all three are yes, it is a critical section and it needs an entry and an exit.

munotes.in126

The Shape of Every Concurrent Program

Why the shape is written as a loop

The do while (true) in the shape is not decoration. It says that a process comes back, over and over. That matters for two of the three requirements of the next chapter:

  • a process that has just left must not be able to walk straight back in and starve the one

waiting;

  • a mechanism that works once and deadlocks on the second pass is worthless.

So when reading a proposed solution, always ask what happens on the second time round. Most wrong solutions are correct on the first pass.

The shape, filled in three ways

The same counter, protected three ways, so the four sections can be pointed at. Each of these is a later chapter and none of them is explained here.

Entry sectionExit sectionChapter
A mutexpthread_mutex_lock(&m)pthread_mutex_unlock(&m)thirty seven
A semaphoresem_wait(&s)sem_post(&s)thirty eight
A monitorthe language does it on entrythe language does it on exitforty two
#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>

static long            counter;           /* the shared thing */
static pthread_mutex_t gate = PTHREAD_MUTEX_INITIALIZER;

static void *worker(void *unused)
{
    (void)unused;
    for (int i = 0; i < 100000; i++) {
        pthread_mutex_lock(&gate);        /* entry section */
        counter++;                        /* critical section */
        pthread_mutex_unlock(&gate);      /* exit section */
        /* remainder section: nothing here */
    }
    return NULL;
}

int main(void)
{
    pthread_t a, b;

    pthread_create(&a, NULL, worker, NULL);
    pthread_create(&b, NULL, worker, NULL);
    pthread_join(a, NULL);
    pthread_join(b, NULL);
    printf("two threads counted to 100000 each, and the counter is %ld\n", counter);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o shape shape.c
$ ./shape
two threads counted to 100000 each, and the counter is 200000
$ for i in $(seq 5); do ./shape | awk '{print $NF}'; done | sort -u
200000

Five runs and the answer is 200000 every time. The next chapter is this program with the entry and exit sections deleted, and its answer is different every run and never right.

Worked example: finding the critical section in somebody else's code

A booking program for a college hall has this in it, run by several staff at once.

if (seats_free > 0) {          /* line 1 */
    seats_free = seats_free - 1;  /* line 2 */
    print_ticket();               /* line 3 */
}

Apply the three questions. seats_free is shared: yes. It is written: yes. Two staff at the same time give a result neither ordering gives: yes, and here is how.

  • One seat is left. Both run line 1 at almost the same moment. Both see 1, which is greater than

0.

  • Both run line 2. Each reads 1, subtracts, writes 0. The variable is 0 and two tickets are
munotes.in127

The Shape of Every Concurrent Program

printed.

  • The hall is oversold by one, and seats_free is 0 rather than the -1 that would at least have

shown the error.

The critical section is lines 1 to 3 together, not line 2 alone. Protecting only the subtraction would still let both staff pass the test. That is the commonest mistake in answering this kind of question: the critical section begins at the test that the decision depends on.

Whether print_ticket belongs inside is a real design question. Inside, and the hall cannot be oversold but staff wait for the printer. Outside, and a failed print leaves the seat sold to nobody. The usual answer is to decrement inside and print outside, and to put the seat back if the printing fails.

Distinctions that carry marks

Critical sectionShared variable
Isa piece of codea piece of data
There may beseveral, in different functions, for one variableone
Protected bythe entry and exit sections around itnothing by itself
Entry sectionRemainder section
Purposeget permission, and wait if refusedthe rest of the process's work
Touches shared datano: it manages permissionno
Written bythe synchronisation mechanismthe programmer

What it does not mean

A critical section is not slow code. It may be one instruction. What makes it critical is that two processes inside it at once is wrong.

The entry section is not a test. if (free) enter(); is not an entry section: the test and the entry must be one indivisible step, which is the whole difficulty of the next three chapters.

Not all shared data needs a critical section. Data that nobody writes does not.

Making the critical section as small as possible is not always right. It must be large enough to contain every step that depends on the others, which the booking example shows: too small is wrong, not merely slow.

Quick revision

  • The general structure: entry section, critical section, exit section,

remainder section, in a loop.

  • A critical section is code that touches shared data and must not be run by two processes at

once.

  • Every mechanism in this row writes only the entry and exit sections: a mutex lock and

unlock, a semaphore wait and signal, a monitor's automatic entry and exit.

  • Find a critical section with three questions: is the data shared, is it written, and would two

processes at once give a result neither order gives?

  • The critical section starts at the test the decision rests on, not at the write.
  • Read only sharing needs no protection.
  • Always ask what a proposed solution does on the second time round.

Test yourself

  1. Write the general structure of a process that shares data. A loop containing an entry
munotes.in128

The Shape of Every Concurrent Program

section, the critical section, an exit section and the remainder section.

  1. Define a critical section. A piece of code that reads or writes data shared with other

processes and that must not be executed by two processes at the same time.

  1. Which sections does a synchronisation mechanism provide? Only the entry and the exit

sections.

  1. Is a shared variable a critical section? No. The critical section is the code that touches

it, and there may be several such pieces of code for one variable. 5. A program tests seats_free > 0 and then decrements it. Where does the critical section begin, and why? At the test. Protecting only the decrement lets two processes both pass the test with one seat left and sell it twice.

  1. When does shared data need no protection? When nobody writes it.
  2. Why is the structure written as a loop? Because a process returns to its critical section

repeatedly, and a mechanism must still be correct and must not starve anybody on the second and every later pass.

Contents This chapter on its own page

munotes.in129

Chapter Thirty-Three

A Race Condition, Made to Happen

Syllabus topic Module 1, "Process Synchronization - race condition"

In one line

A race condition is when the answer depends on which of two processes happens to get there first.

The form to write: a race condition is a situation in which several processes access and manipulate the same data concurrently, and the outcome depends on the particular order in which the accesses take place.

Why one line of C is three instructions

counter++ looks indivisible. It is not. The processor has no instruction that adds one to a number in memory; it must fetch it into a register, add, and store it back.

StepWhat happens
1. Loadread counter from memory into a register
2. Addadd one to the register
3. Storewrite the register back to counter

A context switch can happen between any two of those, and on a machine with two cores the two threads need not even take turns: they can be in step 1 at the same instant. Then both read the same value, both add one to it, both write the same value back, and one of the two increments has vanished.

The word for this is a lost update, and it is the commonest race there is.

Asking the machine, and being told something inconvenient

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>

#define EACH 1000000

static long counter;                      /* shared, and unprotected */

static void *bump(void *unused)
{
    (void)unused;
    for (long i = 0; i < EACH; i++) {
        counter++;                        /* three instructions, not one */
    }
    return NULL;
}

int main(void)
{
    pthread_t a, b;

    pthread_create(&a, NULL, bump, NULL);
    pthread_create(&b, NULL, bump, NULL);
    pthread_join(a, NULL);
    pthread_join(b, NULL);

    long should_be = 2 * (long)EACH;
    printf("expected %ld, got %ld, lost %ld\n",
        should_be, counter, should_be - counter);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o race race.c
$ bad=$(for i in $(seq 20); do ./race; done | grep -vc 'lost 0$')
$ echo "of twenty runs, $bad gave the wrong answer and $((20 - bad)) gave the right one"
of twenty runs, 3 gave the wrong answer and 17 gave the right one

Most runs are right. Two million increments, no protection at all, and on most runs not one is lost; on the others hundreds of thousands vanish. Which runs are which is not the program's choice and not yours.

The lab machine is a container that receives a fraction of one core, which Chapter thirty one measured, so the two threads rarely run at the same instant and each usually gets a long stretch without interruption. When one of them is interrupted in the wrong place, a great deal is lost at once.

The program is not correct. On a good run it is lucky, and it is lucky in exactly the way that puts a program into production and keeps it there until the day the machine changes.

munotes.in130

A Race Condition, Made to Happen

Making the machine show its hand

The three steps are the same three steps. Written out, with the thread giving up the processor between reading and writing, the interleaving the machine is allowed to produce at any moment is produced every time.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>
#include <sched.h>

#define EACH 20000

static long counter;

static void *bump(void *unused)
{
    (void)unused;
    for (long i = 0; i < EACH; i++) {
        long seen = counter;              /* 1. load  */
        sched_yield();                    /*    the switch that is always allowed */
        counter = seen + 1;               /* 2. add and 3. store */
    }
    return NULL;
}

int main(void)
{
    pthread_t a, b;

    pthread_create(&a, NULL, bump, NULL);
    pthread_create(&b, NULL, bump, NULL);
    pthread_join(a, NULL);
    pthread_join(b, NULL);

    long should_be = 2 * (long)EACH;
    printf("expected %ld, got %ld, lost %ld\n",
        should_be, counter, should_be - counter);
    printf("the answer is %s\n", counter == should_be ? "RIGHT" : "WRONG");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o race-shown race-shown.c
$ ./race-shown
expected 40000, got 20000, lost 20000
the answer is WRONG
$ for i in $(seq 5); do ./race-shown | tail -1; done | sort | uniq -c
      5 the answer is WRONG

Half the increments vanished, five runs out of five.

sched_yield changed nothing about what the program computes. It asks the scheduler to run somebody else now, which the scheduler was free to do at that point anyway, and at a thousand other points besides. The two programs are the same program: one hides the bug and one shows it.

So the lesson is the opposite of a demonstration. A race condition is not something you can decide is absent by running the program. Chapter thirty two's shape.c, with a lock, gave 200000 five times out of five and is correct; race.c gave 2000000 five times out of five and is wrong. The two transcripts look the same. Only the code tells you which is which.

And the part that makes it dangerous

Run the second program with far fewer increments.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>
#include <sched.h>

#define EACH 50

static long counter;

static void *bump(void *unused)
{
    (void)unused;
    for (long i = 0; i < EACH; i++) {
        long seen = counter;
        sched_yield();
        counter = seen + 1;
    }
    return NULL;
}

int main(void)
{
    pthread_t a, b;

    pthread_create(&a, NULL, bump, NULL);
    pthread_create(&b, NULL, bump, NULL);
    pthread_join(a, NULL);
    pthread_join(b, NULL);
    printf("%s\n", counter == 2 * (long)EACH ? "right" : "wrong");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o race-small race-small.c
$ bad=$(for i in $(seq 20); do ./race-small; done | grep -c wrong)
$ echo "fifty increments each: $bad of twenty runs wrong, $((20 - bad)) right"
fifty increments each: 11 of twenty runs wrong, 9 right
munotes.in131

A Race Condition, Made to Happen

Fifty increments each is enough to go wrong, and not enough to go wrong every time. The size of the loop is not what decides it. Whether the two threads overlap is, and that is the scheduler's business and not the programmer's.

That is the whole danger of a race condition, and it is what an examination answer should say. The bug is not in the program's logic, it is in the program's timing, so:

  • it passes testing, as race.c did five times over;
  • it works on the developer's machine and fails on the customer's;
  • it appears under load, which is when it costs most;
  • it can disappear when you add a print statement to look for it, because the print changes the

timing;

  • and it can sit for years in code everybody believes is correct.

The interleaving, written out

Two threads, one counter, starting at 5. The correct answer is 7.

TimeThread AThread BRegister ARegister Bcounter
1load55
2load555
3add655
4add665
5store666
6store666

The counter ends at 6 and should be 7. Neither thread did anything wrong. A question that asks you to demonstrate a race condition wants exactly this table, with the shared value in the last column.

Note that the ordering 1, 3, 5, 2, 4, 6, where A finishes before B starts, gives 7. Both orderings are allowed and only one is right, which is the definition.

A race is not only a counter

Shared thingThe race
A countera lost update, as above
A linked listtwo inserts at the head, and one is lost or the list is broken in half
A filetwo appends, and one overwrites the other
A bank balancetwo withdrawals both check the balance, both succeed, the account goes negative
A seat countChapter thirty two's booking, sold twice
A file's existenceone process checks a name is free and another creates it before the first does

The last one has a name, time of check to time of use, and it is a security bug rather than a counting bug: a program that checks it may write a file and then writes it can be tricked by anybody who changes the file in between.

Distinctions that carry marks

A race conditionAn ordinary bug
Inthe timingthe logic
Reproduciblesometimes, and not on demandyes
Found by testingrarelyusually
Changes when you add a printyesno
Fixed bymutual exclusioncorrecting the code
munotes.in132

A Race Condition, Made to Happen

The three instructionsWhat a lock does
Load, add, storecan be split by a switch or by another coreare made indivisible as a group
Costnonea lock and an unlock

What it does not mean

A race condition is not a compiler bug. The compiler did what the language allows.

It is not solved by making the variable volatile. volatile stops the compiler caching the value in a register across statements. It does nothing about two threads doing load, add and store at the same time, and a student who offers it as the answer has not understood the problem.

It is not solved by making the loop shorter, or by hoping. The small version above was right twenty times and is exactly as wrong as the large one.

It is not limited to threads. Two processes sharing memory (Chapter twenty two) race in the same way, and so do two programs appending to one file.

Quick revision

  • A race condition is when several processes touch the same data at once and the outcome

depends on the order.

  • counter++ is load, add, store. A switch or a second core between any two of them loses an

update.

  • One million increments each, two threads, no protection: the answer was wrong on every run and

nearly a million increments disappeared.

  • One thousand each: right twenty times out of twenty, because the threads never overlapped.

A race that does not appear is still there.

  • It is a bug in timing, so it survives testing, appears under load, and vanishes when you

add a print to look for it.

  • volatile does not fix it. Mutual exclusion does.
  • The examinable demonstration is the interleaving table with the shared value in the last

column.

Test yourself

  1. Define a race condition. A situation in which several processes access and change shared

data concurrently and the result depends on the order in which the accesses happen.

  1. Why is counter++ unsafe for two threads? It is three machine steps, load, add and store,

and two threads can both load the same value before either stores, so one increment is lost.

  1. Draw the interleaving that loses an increment. Both threads load the same value, both add

one to their own register, both store the same result. Two increments, one effect.

  1. A program with a race was tested a hundred times and passed. Is it correct? No. A race is

a timing bug; passing a test means the bad interleaving did not happen, not that it cannot.

  1. Does volatile fix a race condition? No. It affects how the compiler keeps the value, not
munotes.in133

A Race Condition, Made to Happen

whether two threads can be inside load, add and store at the same time.

  1. Why does adding a print statement often make a race disappear? Printing takes time and

changes the timing, so the window in which the two threads overlap moves or closes.

  1. Name a race that is a security problem rather than a counting problem. Time of check to

time of use: a program checks that it may write a file and the file is changed between the check and the write.

Contents This chapter on its own page

munotes.in134

Chapter Thirty-Four

The Critical Section Problem, and the Three Conditions

Syllabus topic Module 1, "Process Synchronization - The Critical-Section Problem"

In one line

The critical section problem is to write the entry and exit sections so that only one process is ever inside, nobody is kept out for no reason, and nobody waits for ever.

The problem, stated

Given n processes, each with a critical section as in Chapter thirty two, design a protocol the processes can use to co-operate. Each may share some variables for that purpose, and nothing else may be assumed: in particular nothing about their relative speeds, and nothing about the number of processors.

That last sentence is the difficulty. A solution that works because one process happens to be faster is not a solution.

The three conditions

1. Mutual exclusion

If a process is executing in its critical section, no other process may be executing in its critical section.

This is the obvious one and the only one most students remember. It is also the easiest to achieve on its own: a solution that never lets anybody in at all has perfect mutual exclusion.

2. Progress

If no process is in its critical section and some processes want to enter, then only those processes that want to enter may take part in deciding who goes next, and that decision cannot be postponed indefinitely.

Read the second half twice. It rules out two different faults:

  • a process that does not want to enter must not be able to block one that does;
  • the decision must actually be made, not deferred for ever.

The first is what strict alternation breaks, and it is the standard example below.

3. Bounded waiting

There is a limit on the number of times other processes are allowed to enter their critical sections after a process has asked to enter and before that request is granted.

Bounded waiting is not "no starvation", although it implies it. It is stronger: it demands a number. A promise that a process will get in eventually is not bounded waiting. A promise that at most n minus 1 others will enter before it is.

A solution that breaks progress: strict alternation

The simplest idea anybody has: a shared variable says whose turn it is.

/* process 0 */                      /* process 1 */
while (turn != 0) { }                while (turn != 1) { }
    critical section                     critical section
turn = 1;                            turn = 0;

Mutual exclusion: satisfied. turn holds one value, so only one loop can be past its test.

Progress: broken, and here is the case. Process 0 enters, leaves, sets turn to 1, and then goes off to do half an hour of unrelated work. Process 1 enters, leaves, sets turn to 0. Now process 1 wants to enter again. Nobody is inside. But turn is 0, so process 1 must wait for process 0, which does not want to enter at all. A process that is not interested is blocking one that is.

munotes.in135

The Critical Section Problem, and the Three Conditions

That is precisely what the second condition's wording forbids: only processes that want to enter may take part in the decision.

Bounded waiting: satisfied, trivially, because the turns alternate. A solution can satisfy two conditions and be useless.

A solution that breaks mutual exclusion: two flags

The next idea: each process raises a flag to say it wants in, and waits while the other's flag is up.

/* process i */
flag[i] = true;
while (flag[j]) { }
    critical section
flag[i] = false;

Progress: better. A process that does not want in has its flag down and blocks nobody.

Mutual exclusion: still satisfied.

But it deadlocks, which breaks progress in the other way the condition forbids: the decision is postponed indefinitely. Both processes set their flag, then both test the other's, then both wait for ever. Neither is in the critical section and neither can get in.

Swapping the two lines, testing before setting, breaks mutual exclusion instead: both test, both see a clear flag, both set, both enter. The order of those two lines is the whole problem, and Chapter thirty five is the solution that needs one more variable to settle it.

Seeing progress broken

Strict alternation, with one process that stops being interested.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>
#include <unistd.h>

static volatile int turn;                 /* whose turn it is */
static volatile int entries[2];           /* how many times each got in */
static volatile int stop;

static void *loop(void *argument)
{
    int me = *(int *)argument;
    int other = 1 - me;

    for (int round = 0; round < 3 && !stop; round++) {
        while (turn != me && !stop) {
            /* wait for my turn */
        }
        if (stop) {
            break;
        }
        entries[me]++;
        turn = other;
        if (me == 0) {
            /* process 0 now loses interest for good */
            break;
        }
    }
    return NULL;
}

int main(void)
{
    pthread_t id[2];
    int name[2] = {0, 1};

    pthread_create(&id[0], NULL, loop, &name[0]);
    pthread_create(&id[1], NULL, loop, &name[1]);
    sleep(1);
    stop = 1;                             /* give up waiting and report */
    pthread_join(id[0], NULL);
    pthread_join(id[1], NULL);

    printf("process 0 entered %d time(s), process 1 entered %d time(s)\n",
        entries[0], entries[1]);
    printf("process 1 wanted three turns and got %d, because it had to wait for a\n"
        "process that had stopped wanting to enter\n", entries[1]);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o alternate alternate.c
$ ./alternate
process 0 entered 1 time(s), process 1 entered 1 time(s)
process 1 wanted three turns and got 1, because it had to wait for a
process that had stopped wanting to enter
munotes.in136

The Critical Section Problem, and the Three Conditions

Process 1 asked for three turns and got one. Nobody was in the critical section for the whole of that second. That is a progress failure, and it is not a deadlock: process 1 is not waiting for a resource somebody holds, it is waiting for permission from somebody who has gone home.

What a correct solution may not assume

A question sometimes asks what assumptions are allowed, and the list is short.

May assumeMay not assume
a load and a store of one word are atomicthat two of them together are
the processes will leave their critical sectionsanything about their relative speeds
anything about the number of processors
that a process outside its critical section will ever run again

The last row is the one that kills strict alternation, and it is also why the classical software solutions are hard: a correct one must work when one process is stopped for an hour at any point outside its critical section.

Two more ways to get it wrong

Disabling interrupts. On a single processor, turning interrupts off around the critical section does give mutual exclusion. It is wrong anyway: it is privileged, so a user program cannot do it; it stops the timer, so the whole machine is held; and on a machine with several processors it does nothing at all, because the other processor is not interrupted.

Busy waiting for a long critical section. A while loop that spins is called a spinlock and it costs a whole processor while it waits. For a critical section of three instructions that is cheaper than a context switch; for one that reads a file it wastes seconds. Chapter thirty seven measures the trade.

Distinctions that carry marks

ConditionBroken byThe symptom
Mutual exclusiontesting before setting your flagtwo processes inside at once
Progressstrict alternation, or both flags set then both testednobody inside and nobody able to get in
Bounded waitinga rule that lets the same process keep winningone process waits while others go in repeatedly
DeadlockStarvation
Who is stuckeverybody involvedone unlucky process
Others make progressnoyes
Breaksprogressbounded waiting

What it does not mean

Mutual exclusion alone is not a solution. A lock nobody can ever take has it.

Bounded waiting is not fairness in the ordinary sense. It does not require first come first served, only a limit on how many others may overtake you.

"Progress" is not about speed. It is about not being blocked by processes that do not want in, and about the decision being made at all.

munotes.in137

The Critical Section Problem, and the Three Conditions

Quick revision

  • The critical section problem: write the entry and exit sections for n processes, assuming

nothing about their speeds or the number of processors.

  • Mutual exclusion: at most one process inside at a time.
  • Progress: if nobody is inside, only the processes that want in decide who goes, and the

decision cannot be postponed for ever.

  • Bounded waiting: a limit on how many others may enter after you have asked and before

you are let in.

  • Strict alternation gives mutual exclusion and bounded waiting and breaks progress: a

process that has lost interest still blocks the other.

  • Two flags, set then test deadlocks. Test then set lets both in. The order of those two

lines is the whole problem.

  • Disabling interrupts is privileged, holds the whole machine, and does nothing on a

multiprocessor.

  • A correct solution must work when one process is stopped for an hour outside its critical

section.

Test yourself

  1. State the three requirements for a solution to the critical section problem. Mutual

exclusion: only one process in its critical section at a time. Progress: if no process is inside and some wish to enter, only those wishing to enter may decide who goes next, and the decision may not be postponed indefinitely. Bounded waiting: a limit on how many other processes may enter after a process has requested entry and before that request is granted.

  1. Which condition does strict alternation break, and how? Progress. A process that no longer

wants to enter still holds the turn, so the other process waits although nobody is inside.

  1. Two processes each set their own flag and then test the other's. What goes wrong? Both

set, both test, both wait for ever. The decision is postponed indefinitely, which is a progress failure.

  1. They test first and then set. Now what goes wrong? Both see a clear flag, both set theirs,

both enter. Mutual exclusion is broken.

  1. Why is "you will get in eventually" not bounded waiting? Bounded waiting requires a

number: a limit on how many others may go in first. An eventual guarantee with no bound allows an unbounded wait.

  1. Give three reasons disabling interrupts is not an acceptable solution. It is a privileged

operation a user program cannot perform; it stops the timer and holds the whole machine; and on a machine with more than one processor it does not exclude the other processor at all.

  1. What may a solution assume about the processes? Only that a single load or store of a word

is atomic and that a process will eventually leave its critical section. Nothing about speeds, the number of processors, or whether a process outside its critical section will run again.

Contents This chapter on its own page

munotes.in138

Chapter Thirty-Five

Peterson's Solution

Syllabus topic Module 1, "Process Synchronization - The Critical-Section Problem"

In one line

Peterson's solution combines the two broken ideas of the last chapter: each process says it wants in, and then politely gives the other one the turn.

The algorithm

Two processes, numbered 0 and 1. Two shared variables.

VariableMeaning
flag[i]process i wants to enter
turnwhose turn it is to go if both want in
/* process i, with j the other one */
flag[i] = true;
turn = j;                        /* offer the turn to the other */
while (flag[j] && turn == j) { } /* wait only if the other wants in AND it is their turn */
    critical section
flag[i] = false;

The whole trick is turn = j. A process does not claim the turn, it gives it away. If both processes do this at nearly the same moment, the second write wins, so exactly one of them is left looking at a turn that is not its own, and that one waits. There is no case where both wait and no case where both proceed.

The proof of each condition

Mutual exclusion. Suppose both are inside. Then both passed the while, so for process 0 either flag[1] was false or turn was 0, and for process 1 either flag[0] was false or turn was 1. But both had set their own flag true before the test, so neither flag was false. So turn was 0 and turn was 1 at the same moment, which is impossible for one variable. Contradiction, so they cannot both be inside.

Progress. Process i can be held up only by flag[j] && turn == j. If process j does not want in, its flag is false and the test fails at once. If process j does want in, then turn is 0 or 1, and whichever it is, that process proceeds. So somebody always gets in.

Bounded waiting. When process j leaves, it sets flag[j] false, which releases process i immediately. If process j wants in again, it must first set turn = i, which again releases process i. So process i waits for at most one turn of process j, which is the bound.

Notice that the bound is a number, which is what Chapter thirty four said bounded waiting demands.

It runs, and it works

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>

#define EACH 200000

static volatile int  flag[2];
static volatile int  turn;
static long          counter;             /* the shared thing being protected */

static void *worker(void *argument)
{
    int me = *(int *)argument;
    int other = 1 - me;

    for (int n = 0; n < EACH; n++) {
        flag[me] = 1;                     /* entry section */
        turn = other;
        while (flag[other] && turn == other) {
            /* wait */
        }
        counter++;                        /* critical section */
        flag[me] = 0;                     /* exit section */
    }
    return NULL;
}

int main(void)
{
    pthread_t id[2];
    int name[2] = {0, 1};

    pthread_create(&id[0], NULL, worker, &name[0]);
    pthread_create(&id[1], NULL, worker, &name[1]);
    pthread_join(id[0], NULL);
    pthread_join(id[1], NULL);

    long should_be = 2 * (long)EACH;
    printf("expected %ld, got %ld, %s\n", should_be, counter,
        counter == should_be ? "correct" : "WRONG");
    return 0;
}
munotes.in139

Peterson's Solution

$ gcc -std=c17 -Wall -Wextra -pthread -o peterson peterson.c
$ bad=$(for i in $(seq 10); do ./peterson; done | grep -c WRONG)
$ echo "ten runs of Peterson's solution, $bad of them wrong"
ten runs of Peterson's solution, 0 of them wrong

Ten runs and four hundred thousand increments each time, with no lock, no system call and no help from the operating system at all. Compare Chapter thirty three, where the same counter with no protection lost hundreds of thousands.

Why nobody uses it

Three reasons, and the third is the one this machine can show.

1. It is for two processes

The n process version is Lamport's bakery algorithm: every process takes a numbered ticket, as at a bakery counter, and the lowest number goes first, with the process id breaking ties because two processes can take the same number. It works, it needs an array of two variables per process, and it is far more elaborate than anything a real system would run.

2. It busy waits

The while loop burns a processor for as long as the other process is inside. On one processor that is worse than useless: the waiting process holds the processor that the process inside needs in order to finish. Chapter thirty seven measures what busy waiting costs and when it is worth it.

3. It assumes the machine does what the program says, in the order the program says

It does not. A modern processor and a modern compiler may both reorder a store and a following load of a different address, because on one processor nobody can tell the difference. Peterson's solution depends on exactly that ordering: flag[me] = 1 must become visible to the other processor before flag[other] is read. If the read happens first, both processes can pass the test.

This is called a weak memory model, and the machine this book runs on has one: it is an ARM processor, where stores become visible to other cores in an order the hardware is free to choose.

The fix is a memory barrier, an instruction that forbids the reordering across it. In C it is written atomic_thread_fence(memory_order_seq_cst), and the listing above compiles to correct code on this machine only because volatile and the compiler's own choices happen to leave the order alone.

munotes.in140

Peterson's Solution

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <stdatomic.h>
#include <pthread.h>

#define EACH 200000

static volatile int flag[2];
static volatile int turn;
static long         counter;

static void *worker(void *argument)
{
    int me = *(int *)argument;
    int other = 1 - me;

    for (int n = 0; n < EACH; n++) {
        flag[me] = 1;
        turn = other;
        atomic_thread_fence(memory_order_seq_cst);   /* the missing instruction */
        while (flag[other] && turn == other) {
            atomic_thread_fence(memory_order_seq_cst);
        }
        counter++;
        atomic_thread_fence(memory_order_seq_cst);
        flag[me] = 0;
    }
    return NULL;
}

int main(void)
{
    pthread_t id[2];
    int name[2] = {0, 1};

    pthread_create(&id[0], NULL, worker, &name[0]);
    pthread_create(&id[1], NULL, worker, &name[1]);
    pthread_join(id[0], NULL);
    pthread_join(id[1], NULL);
    printf("with barriers: expected %ld, got %ld, %s\n", 2 * (long)EACH, counter,
        counter == 2 * (long)EACH ? "correct" : "WRONG");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o barrier barrier.c
$ ./barrier
with barriers: expected 400000, got 400000, correct

Both versions gave the right answer here, and only one of them is guaranteed to. That is the honest statement, and it is the same lesson as Chapter thirty three: a run that works does not make a concurrent program correct. What makes the second version correct is the barrier, and what made the first one work is luck about how this compiler and this processor happened to behave.

Write this in an examination as: Peterson's solution is correct on a machine with a strong memory model, and on a modern processor it requires memory barriers, which is one of the reasons real systems use hardware instructions or operating system locks instead.

Distinctions that carry marks

Strict alternationTwo flagsPeterson
Mutual exclusionyesyesyes
Progressnono, it deadlocksyes
Bounded waitingyesnot reachedyes, at most one turn
Variablesturnflag[2]flag[2] and turn
Why it failsan uninterested process holds the turnboth set, both waitnothing, in theory
PetersonLamport's bakery
Processestwon
Ideagive the turn awaytake a numbered ticket, lowest first
Tiesimpossiblebroken by process id

What it does not mean

Peterson's solution is not obsolete as an idea. It is the proof that mutual exclusion is achievable in software alone, which is why it is taught. What is obsolete is using it.

The while loop is not a mistake. It is the entry section of Chapter thirty two, implemented by busy waiting. Busy waiting is a cost, not an error.

volatile is not a memory barrier. It stops the compiler caching a value in a register. It says nothing to the processor about the order in which its stores become visible to other cores, which is why atomic_thread_fence exists.

Quick revision

  • Peterson's solution: flag[i] = true; turn = j; while (flag[j] && turn == j) ; then the
munotes.in141

Peterson's Solution

critical section, then flag[i] = false.

  • The trick is that a process gives the turn away rather than claiming it, so one of two

simultaneous writers is always left waiting.

  • Mutual exclusion: both inside would need turn to hold two values at once.
  • Progress: a process whose flag is false blocks nobody.
  • Bounded waiting: at most one turn of the other process, which is a number.
  • It is for two processes; the n process version is Lamport's bakery algorithm.
  • It busy waits, which wastes a processor and is actively harmful on a single processor

machine.

  • On a weak memory model, which every modern processor has, it needs memory barriers.

volatile is not a barrier.

Test yourself

  1. Write Peterson's solution. Set your own flag, set turn to the other process, wait while

the other's flag is up and the turn is theirs, run the critical section, clear your flag.

  1. Why does the algorithm set turn to the other process rather than to itself? So that when

both processes are entering, the later write decides, leaving exactly one of them looking at a turn that is not its own. If each claimed the turn, both could see their own value.

  1. Prove mutual exclusion. If both were inside, both must have found either the other's flag

false or the turn its own. Both flags were true, so turn would have to equal 0 and 1 at once, which is impossible.

  1. What is the bound in bounded waiting here? One. A waiting process is released either when

the other clears its flag or when the other sets the turn back on its next attempt.

  1. Give three reasons a real system does not use it. It handles only two processes; it busy

waits, which wastes a processor and is harmful on a single processor machine; and on a processor with a weak memory model it is incorrect without memory barriers.

  1. What is the n process version called and how does it work? Lamport's bakery algorithm.

Each process takes a numbered ticket and the lowest number enters first, with the process id breaking ties because two processes may take the same number.

  1. Does volatile make Peterson's solution correct? No. It prevents the compiler from

caching the values, but not the processor from making stores visible to other cores out of order. A memory barrier does that.

Contents This chapter on its own page

munotes.in142

Chapter Thirty-Six

Hardware Help: Test and Set, and Compare and Swap

Syllabus topic Module 1, "Process Synchronization - Mutex Locks"

In one line

The processor provides one instruction that reads a word and writes it in the same indivisible step, and every lock in every operating system is built on it.

Why an instruction is needed at all

Chapter thirty five achieved mutual exclusion in software alone, and it took two shared variables, a proof and a memory barrier. The reason it was hard is one sentence: there is no way to test a variable and set it without a gap in between.

Everything in Chapter thirty four's list of failures is that gap. Test then set: both test in the gap. Set then test: both set, then both test. Peterson closes the gap with a third variable and an argument.

A single instruction that does both has no gap, by definition, because the hardware refuses to let anything else touch that word while it runs. Once you have one, a lock is four lines.

Test and set

The instruction, written as the function it behaves like.

boolean test_and_set(boolean *target)
{
    boolean was = *target;
    *target = true;
    return was;
}

The whole of that happens atomically: no other processor and no interrupt can see the word between the read and the write.

A lock follows at once.

/* lock is false when free */
while (test_and_set(&lock)) { }      /* entry section */
    critical section
lock = false;                        /* exit section */

If the lock was free, test_and_set returns false and sets it true: you are in, and it is now held. If it was held, it returns true and sets it true again, which changes nothing: you go round the loop.

This satisfies mutual exclusion and progress, and NOT bounded waiting. Nothing chooses among the waiting processes, so one of them can be unlucky every single time while others come and go. A version with bounded waiting needs an array of waiting flags and passes the lock along a circle of processes, which is exactly Silberschatz's bounded waiting mutual exclusion algorithm.

Compare and swap

The more useful instruction, and the one modern processors actually provide.

int compare_and_swap(int *value, int expected, int new_value)
{
    int was = *value;
    if (*value == expected) {
        *value = new_value;
    }
    return was;
}

Again all of it is atomic. It is strictly more powerful than test and set, because it can change a word to anything conditional on it having an expected value, so it can be used for counters, list heads and anything else, not only for a boolean lock.

while (compare_and_swap(&lock, 0, 1) != 0) { }
    critical section
lock = 0;

On this machine the instructions are ARM's LDXR and STXR, a load exclusive and a store exclusive, which between them provide compare and swap: the store fails if anything touched the word since the load, and the loop retries. On an Intel processor it is XCHG and CMPXCHG, and on both the C compiler hides the difference behind one name.

munotes.in143

Hardware Help: Test and Set, and Compare and Swap

Built and run

The C language provides these directly, so the counter of Chapter thirty three can be fixed without a lock at all.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <stdatomic.h>
#include <pthread.h>

#define EACH 200000

static atomic_flag lock = ATOMIC_FLAG_INIT;   /* test and set, in the language */
static long        by_lock;

static atomic_long by_atomic_add;             /* no lock at all */

static void *worker(void *unused)
{
    (void)unused;
    for (int n = 0; n < EACH; n++) {
        while (atomic_flag_test_and_set(&lock)) {
            /* busy wait: the entry section */
        }
        by_lock++;                            /* critical section */
        atomic_flag_clear(&lock);             /* exit section */

        atomic_fetch_add(&by_atomic_add, 1);  /* the same job, with no lock */
    }
    return NULL;
}

int main(void)
{
    pthread_t a, b;

    pthread_create(&a, NULL, worker, NULL);
    pthread_create(&b, NULL, worker, NULL);
    pthread_join(a, NULL);
    pthread_join(b, NULL);

    long want = 2 * (long)EACH;
    printf("with a test and set lock : %ld, %s\n", by_lock,
        by_lock == want ? "correct" : "WRONG");
    printf("with an atomic add      : %ld, %s\n", atomic_load(&by_atomic_add),
        atomic_load(&by_atomic_add) == want ? "correct" : "WRONG");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o atomics atomics.c
$ bad=$(for i in $(seq 10); do ./atomics; done | grep -c WRONG)
$ echo "ten runs, twenty results, $bad of them wrong"
ten runs, twenty results, 0 of them wrong

Two different correct answers to the same problem. The first builds a lock and protects an ordinary variable, which is what the rest of this row is about. The second has no lock at all: the addition itself is one atomic instruction, so there is no critical section to protect.

Which to use is a real decision. An atomic add is far faster and works only when the whole critical section is one such operation. A lock costs more and protects anything, including three statements that must happen together.

Compare and swap, doing something a lock cannot

A lock free counter that only ever increases, written with compare and swap, to show the shape of the loop.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <stdatomic.h>
#include <pthread.h>

#define EACH 100000

static atomic_long biggest;

static void *worker(void *argument)
{
    long me = *(long *)argument;

    for (long n = 1; n <= EACH; n++) {
        long candidate = n * me;
        long seen = atomic_load(&biggest);
        /* keep trying until either we won, or somebody put a bigger one there */
        while (candidate > seen &&
            !atomic_compare_exchange_weak(&biggest, &seen, candidate)) {
            /* seen has been reloaded by the call: go round again */
        }
    }
    return NULL;
}

int main(void)
{
    pthread_t id[2];
    long name[2] = {1, 2};

    pthread_create(&id[0], NULL, worker, &name[0]);
    pthread_create(&id[1], NULL, worker, &name[1]);
    pthread_join(id[0], NULL);
    pthread_join(id[1], NULL);
    printf("the largest value any thread offered was %ld\n", atomic_load(&biggest));
    printf("it should be %ld, and it is %s\n", 2L * EACH,
        atomic_load(&biggest) == 2L * EACH ? "correct" : "WRONG");
    return 0;
}
munotes.in144

Hardware Help: Test and Set, and Compare and Swap

$ gcc -std=c17 -Wall -Wextra -pthread -o cas cas.c
$ ./cas
the largest value any thread offered was 200000
it should be 200000, and it is correct

That loop is the shape of every lock free algorithm: read the current value, work out the new one, try to swap it in, and go round again if somebody changed it while you were thinking. Nothing is ever locked, and no thread can be blocked by another thread being descheduled, which is the property lock free algorithms are written for.

It is also much harder to get right than a lock, which is why a first course teaches the lock.

Distinctions that carry marks

Test and setCompare and swap
Reads and writes atomicallyyesyes
Writesalways trueonly if the value was as expected
Returnsthe old valuethe old value
Can build a lockyesyes
Can update a counter without a locknoyes
On this machineLDXR and STXRthe same pair
A software solution, PetersonA hardware instruction
Needstwo variables, a proof, a barrierone instruction
Processestwo, or the bakery algorithm for nany number
Memory modelmust be reasoned aboutthe instruction is defined to be atomic
Used by real systemsnoyes, underneath every lock
A lockAn atomic operation
Protectsany amount of codeone operation
Costhigherlower
Can a thread be stuck holding ityesno
Use whenseveral statements must happen togetherthe whole change is one step

What it does not mean

Atomic does not mean fast. An atomic operation has to co-ordinate between processor caches, and on a machine with many cores a heavily contended atomic variable is slow. It is still far faster than a system call.

Test and set does not give bounded waiting. It gives mutual exclusion and progress. Nothing decides who goes next, so a process can be unlucky indefinitely.

These instructions are not privileged. An ordinary user program uses them, which is exactly why locks can be built in user space with no system call in the common case. Chapter thirty seven is that design.

Atomic operations do not remove the need for barriers. They include the necessary ordering themselves, which is why the atomic add above needed no atomic_thread_fence while Chapter thirty five's plain variables did.

munotes.in145

Hardware Help: Test and Set, and Compare and Swap

Quick revision

  • The hardware provides one instruction that reads and writes a word atomically, closing the

gap between a test and a set.

  • test and set returns the old value and always writes true. A lock:

while (test_and_set(&lock)) ; then the critical section, then lock = false.

  • It gives mutual exclusion and progress but not bounded waiting: nothing chooses among the

waiters.

  • compare and swap writes only if the word held the expected value, and is strictly more

powerful: it can update counters and pointers with no lock.

  • The lock free shape: read, compute, try to swap, and go round again if somebody changed it.
  • On this ARM machine the instructions are load exclusive and store exclusive; on Intel, exchange

and compare exchange.

  • Atomic operations carry their own ordering, so they need no separate memory barrier.
  • An atomic operation protects one step; a lock protects any amount of code.

Test yourself

  1. What does test_and_set do? It reads a boolean, sets it to true, and returns what it

read, all in one indivisible step.

  1. Write a lock using it. while (test_and_set(&lock)) ; then the critical section, then

lock = false.

  1. Which of the three conditions does that lock fail, and why? Bounded waiting. Nothing

chooses among the processes spinning on the lock, so one of them may lose the race any number of times.

  1. How is compare and swap more powerful than test and set? It writes only if the word holds

an expected value, so it can be used to change a counter or a pointer to a computed new value, not only to set a flag.

  1. Why is one instruction enough when two statements were not? Because the hardware

guarantees that nothing else touches the word between the read and the write, so there is no gap for another process to act in.

  1. Write the shape of a lock free update. Read the current value; compute the new one; try to

compare and swap it in; if the swap fails because somebody else changed the value, reload and try again.

  1. When is an atomic add the right answer and when is a lock? An atomic add when the whole

critical section is that one operation. A lock when several statements must happen together as a unit.

Contents This chapter on its own page

munotes.in146

Chapter Thirty-Seven

The Mutex Lock

Syllabus topic Module 1, "Process Synchronization - Mutex Locks"

In one line

A mutex is a lock with two operations, acquire and release, and only the thread that holds it may be inside the critical section.

The name, and what it promises

Mutex is short for mutual exclusion, which is the first of Chapter thirty four's three conditions and the only one the lock itself promises. It has exactly two states.

StateMeaning
available, or unlockedanybody may acquire it
held, or lockedone thread has it; everybody else waits

Two operations, and a rule.

OperationWhat it does
acquire(), or lockwait until the mutex is available, then take it
release(), or unlockmake it available again

The rule: the thread that acquired it is the thread that must release it. A mutex has an owner. That distinguishes it from the semaphore of the next chapter, where any process may signal, and it is the single most examinable difference between the two.

Built out of the last chapter

acquire is exactly the loop of Chapter thirty six, and that is the whole implementation.

acquire()
{
    while (test_and_set(&available_is_false)) { }
}

release()
{
    available = true;
}

A lock is therefore not a new idea. It is a name for the pattern, given so that a programmer writes lock and unlock instead of reasoning about a spin loop every time.

The counter, fixed

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>

#define EACH 500000

static long            counter;
static pthread_mutex_t gate = PTHREAD_MUTEX_INITIALIZER;

static void *bump(void *unused)
{
    (void)unused;
    for (long i = 0; i < EACH; i++) {
        pthread_mutex_lock(&gate);
        counter++;
        pthread_mutex_unlock(&gate);
    }
    return NULL;
}

int main(void)
{
    pthread_t a, b;

    pthread_create(&a, NULL, bump, NULL);
    pthread_create(&b, NULL, bump, NULL);
    pthread_join(a, NULL);
    pthread_join(b, NULL);
    printf("expected %ld, got %ld, %s\n", 2 * (long)EACH, counter,
        counter == 2 * (long)EACH ? "correct" : "WRONG");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o mutex mutex.c
$ bad=$(for i in $(seq 10); do ./mutex; done | grep -c WRONG)
$ echo "ten runs with a mutex, $bad wrong"
ten runs with a mutex, 0 wrong

One million increments, ten runs, no losses. Compare Chapter thirty three, the same program with those two lines deleted.

The spinlock, and what it costs

A spinlock is a mutex whose acquire busy waits: it goes round a loop until the lock is free, holding the processor the whole time.

Gainedno context switch at all when the wait is short
Losta whole processor for as long as the wait lasts
On one processoractively harmful: the spinner holds the processor the holder needs to finish
Right whenthe critical section is a few instructions and there is more than one processor
munotes.in147

The Mutex Lock

So the decision is a comparison of two times: the length of the critical section against the cost of a context switch. Spin if the wait will be shorter than the switch; sleep if it will be longer. A real mutex does both: it spins for a short while and then asks the kernel to put the thread to sleep, which is called an adaptive mutex, and it is what the library on this machine does.

The difference can be measured, and the measurement to take is not the clock but the count of voluntary context switches: the number of times a thread gave the processor up on purpose. A sleeping mutex gives it up; a spinlock never does.

#define _GNU_SOURCE
#include <stdio.h>
#include <stdatomic.h>
#include <pthread.h>
#include <time.h>
#include <sys/resource.h>

static atomic_flag     spin = ATOMIC_FLAG_INIT;
static pthread_mutex_t sleeper = PTHREAD_MUTEX_INITIALIZER;
static int             use_spin;

static void pause_for(long milliseconds)
{
    struct timespec t = {milliseconds / 1000, (milliseconds % 1000) * 1000000};

    nanosleep(&t, NULL);
}

/* The holder takes the lock, keeps it for half a second, and gives it back. */
static void *holder(void *unused)
{
    (void)unused;
    if (use_spin) {
        while (atomic_flag_test_and_set(&spin)) {
        }
        pause_for(500);
        atomic_flag_clear(&spin);
    } else {
        pthread_mutex_lock(&sleeper);
        pause_for(500);
        pthread_mutex_unlock(&sleeper);
    }
    return NULL;
}

/* The waiter arrives late, so it is certain to find the lock held. */
static void *waiter(void *unused)
{
    struct rusage before, after;

    (void)unused;
    pause_for(100);
    getrusage(RUSAGE_THREAD, &before);
    if (use_spin) {
        while (atomic_flag_test_and_set(&spin)) {
            /* busy wait for about four tenths of a second */
        }
        atomic_flag_clear(&spin);
    } else {
        pthread_mutex_lock(&sleeper);
        pthread_mutex_unlock(&sleeper);
    }
    getrusage(RUSAGE_THREAD, &after);
    printf("%s: the waiter gave the processor up %ld time(s) on purpose\n",
        use_spin ? "spinning" : "sleeping",
        after.ru_nvcsw - before.ru_nvcsw);
    return NULL;
}

int main(int argc, char **argv)
{
    pthread_t h, w;

    use_spin = (argc > 1 && argv[1][0] == 's' && argv[1][1] == 'p');
    pthread_create(&h, NULL, holder, NULL);
    pthread_create(&w, NULL, waiter, NULL);
    pthread_join(h, NULL);
    pthread_join(w, NULL);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o spin-vs-sleep spin-vs-sleep.c
$ ./spin-vs-sleep spin
spinning: the waiter gave the processor up 0 time(s) on purpose
$ ./spin-vs-sleep sleep
sleeping: the waiter gave the processor up 1 time(s) on purpose

Zero against one, and that one is the whole difference. Both waiters waited about four tenths of a second for the same lock. The spinning one held a processor for all of it and never gave it up. The sleeping one gave it up once, immediately, and was woken when the lock came free.

Four tenths of a second of a processor, thrown away, to avoid one context switch that costs a few microseconds. That is why a lock held for any length of time must sleep, and it is also why a lock held for three instructions should not: there, one switch each way costs more than the wait.

munotes.in148

The Mutex Lock

Three ways to get a mutex wrong

Forgetting to unlock on an early return. A function that locks and then returns from the middle of an if leaves the mutex held for ever. Every later thread waits and the program stops. This is the single commonest lock bug there is.

Locking twice in the same thread. An ordinary mutex is not recursive: a thread that already holds it and locks it again waits for itself, for ever. A recursive mutex counts the acquisitions and is released the same number of times, and it exists because this happens.

Locking two mutexes in different orders in different threads. Thread A takes lock 1 then lock 2; thread B takes lock 2 then lock 1; both wait. That is a deadlock and it is Chapter fifty six.

Distinctions that carry marks

A mutexA binary semaphore
Has an owneryes: only the locker may unlockno: anybody may signal
Used forprotecting a critical sectionprotecting, or signalling between threads
Countsno, two states onlyyes, from 0 upwards
Initial valueavailablewhatever you choose
A spinlockA sleeping, or blocking, mutex
While waitingholds the processoris in the waiting state
Context switchesnonetwo, into and out of sleep
Right whenthe critical section is very short and there are several processorsthe wait may be long
On one processorharmfulcorrect

What it does not mean

A mutex does not give bounded waiting by itself. The underlying test and set does not choose among the waiters; a real implementation usually keeps a queue, and then it does.

Locking does not protect data. It protects code. Two functions that touch the same variable must lock the same mutex. A mutex per function protects nothing.

A mutex is not free even when it is uncontended. It is an atomic operation and a memory barrier, which on a modern processor costs tens of cycles. It is still hundreds of times cheaper than a system call, which is why the library avoids the kernel when the lock is free.

Unlocking a mutex you do not hold is not harmless. It is undefined behaviour, and in practice it lets a second thread into the critical section.

Quick revision

  • A mutex, for mutual exclusion, has two states and two operations: acquire and

release.

  • The thread that acquires it must release it: a mutex has an owner. A semaphore does

not.

  • It is built from the test and set of Chapter thirty six; acquire is that spin loop.
  • A spinlock busy waits: no context switch, but it holds a processor, and on a single
munotes.in149

The Mutex Lock

processor machine it is harmful because the spinner holds the processor the holder needs.

  • Spin if the wait is shorter than a context switch; sleep otherwise. An adaptive mutex does

both.

  • Measured here: the spinning version made almost no voluntary switches, the sleeping one made

fifty five on purpose and finished sooner on a machine short of processor time.

  • Three bugs: forgetting to unlock on an early return, locking the same mutex twice in one

thread, and locking two mutexes in different orders.

  • Two functions touching one variable must lock the same mutex.

Test yourself

  1. What are the two operations on a mutex and what does each do? acquire waits until the

mutex is available and then takes it; release makes it available again.

  1. What is the ownership rule, and why does it matter? Only the thread that acquired the

mutex may release it. It is what distinguishes a mutex from a semaphore, where any process may signal.

  1. What is a spinlock and when is it the right choice? A mutex whose acquire busy waits. It

is right when the critical section is a few instructions and the machine has more than one processor, so that the wait is shorter than a context switch.

  1. Why is a spinlock harmful on a single processor machine? The spinning thread holds the

only processor, which is the processor the lock holder needs in order to finish and release the lock.

  1. What is an adaptive mutex? One that spins briefly and then sleeps, so that short waits

cost no switch and long waits cost no processor. 6. A thread locks a mutex and returns early from the function without unlocking. What happens? The mutex stays held for ever and every other thread that wants it waits for ever.

  1. Two functions each touch the same global. Is a mutex in each of them enough? No. They must

lock the same mutex; two different mutexes exclude nobody from anything.

Contents This chapter on its own page

munotes.in150

Chapter Thirty-Eight

The Semaphore

Syllabus topic Module 1, "Process Synchronization - Semaphores"

In one line

A semaphore is an integer that may only be increased and decreased, and a process that tries to decrease it below zero waits.

Where it comes from

Edsger Dijkstra introduced the semaphore in Cooperating Sequential Processes, written in 1965. He called the two operations the P-operation and the V-operation and did not say in that paper what the letters stand for. English text books write them wait and signal, and an examination question may use either pair of names.

His own definition of the P-operation, in the paper that introduced it:

Its function is to decrease the value of its argument semaphore by 1 as soon as the resulting value

would be non-negative.

Read that carefully. As soon as the resulting value would be non-negative: it is not "if", it is "as soon as". The operation does not fail and it does not return an error. It waits until subtracting one leaves a value of zero or more, and then it subtracts. And the completion, meaning the decision that this is the moment and the decrease itself, is one indivisible action.

He is equally clear about what is not promised. When several processes are waiting on a semaphore and it becomes available, which one proceeds is, in his words, "left unspecified, i,e, at least outside our control". The typing slip is his and is reproduced as printed.

The V operation is the other half: it increases the value by one, and if anybody was waiting, one of them may now proceed.

The two operations, written out

wait(S)                          signal(S)
{                                {
    while (S <= 0) { }               S = S + 1;
    S = S - 1;                   }
}

Both of those must be atomic as a whole, and the loop is written that way only to show the intention. A real implementation uses the hardware of Chapter thirty six for the test and the decrement together, and puts the waiting process in a queue rather than spinning, which is what gives bounded waiting.

The busy waiting version has a name, a spinlock semaphore, and the sleeping version is what every real system provides: wait puts the process in the waiting state of Chapter fifteen and signal moves one waiting process to ready.

Two kinds, and the difference is the initial value

Binary semaphoreCounting semaphore
Values0 and 1 only0 upwards, with no limit
Initialised to1, usuallythe number of instances of the resource
Behaves likea mutex, without an ownera count of what is left
Used formutual exclusionresource counting, and signalling

A counting semaphore initialised to n lets n processes through before the n plus first waits. That single sentence is what a semaphore is for, and it is what a mutex cannot do.

munotes.in151

The Semaphore

The three things a semaphore does that a mutex cannot

1. Count a resource

A printer room with three printers. The semaphore starts at 3. Each job waits, which takes one, and signals when finished, which gives one back. The fourth job waits until one is free, and nobody counts anything by hand.

2. Signal from one process to another

Suppose statement A in process 1 must happen before statement B in process 2. A semaphore starting at 0 does it: process 2 waits on it, which blocks immediately because it is 0, and process 1 signals it after doing A.

A mutex cannot do this at all, because of the ownership rule of Chapter thirty seven: process 1 would have to release a mutex it never acquired. That is the practical meaning of "a semaphore has no owner", and it is the answer to the commonest examination question on the pair.

3. Let several readers in at once

The readers and writers problem of Chapter forty is built on a counting semaphore, and no arrangement of mutexes gives it as directly.

All three, run

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <semaphore.h>
#include <pthread.h>
#include <time.h>

static sem_t printers;                    /* counting: three at a time */
static sem_t a_is_done;                   /* signalling: starts at zero */
static sem_t gate;                        /* binary: mutual exclusion */

static int  at_once;
static int  most_at_once;
static long counter;

static void *job(void *unused)
{
    (void)unused;
    sem_wait(&printers);                  /* take a printer, or wait */

    sem_wait(&gate);                      /* the binary semaphore, as a mutex */
    at_once++;
    if (at_once > most_at_once) {
        most_at_once = at_once;
    }
    sem_post(&gate);

    nanosleep(&(struct timespec){0, 20000000}, NULL);   /* printing */

    sem_wait(&gate);
    at_once--;
    counter++;
    sem_post(&gate);

    sem_post(&printers);                  /* give the printer back */
    return NULL;
}

static void *waits_for_a(void *unused)
{
    (void)unused;
    sem_wait(&a_is_done);                 /* blocks: the semaphore starts at 0 */
    printf("B ran, and it ran after A\n");
    return NULL;
}

int main(void)
{
    pthread_t id[12], b;

    sem_init(&printers, 0, 3);
    sem_init(&a_is_done, 0, 0);
    sem_init(&gate, 0, 1);

    pthread_create(&b, NULL, waits_for_a, NULL);
    nanosleep(&(struct timespec){0, 50000000}, NULL);   /* B is waiting by now */
    printf("A ran first\n");
    sem_post(&a_is_done);                 /* release B */
    pthread_join(b, NULL);

    for (int i = 0; i < 12; i++) {
        pthread_create(&id[i], NULL, job, NULL);
    }
    for (int i = 0; i < 12; i++) {
        pthread_join(id[i], NULL);
    }
    printf("12 jobs finished, %ld counted, and never more than %d printed at once\n",
        counter, most_at_once);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o semaphores semaphores.c
$ ./semaphores
A ran first
B ran, and it ran after A
12 jobs finished, 12 counted, and never more than 3 printed at once
munotes.in152

The Semaphore

Three results in one program.

  1. The order was forced. B was created first and still ran second, because it waited on a

semaphore that started at 0 and only A's signal released it. No mutex can do that.

  1. Twelve jobs, three at a time. most_at_once never exceeded 3, which the counting

semaphore guaranteed with no counting in the program.

  1. The count is exactly 12. The binary semaphore protected at_once and counter as a mutex

would.

The two ways to get a semaphore wrong

Signalling instead of waiting, or waiting twice. Write wait where signal belongs and two processes enter the critical section together. Write wait twice and the process blocks for ever on its own second call. Neither mistake is possible with a mutex used through a lock and unlock pair, which is one reason the monitor of Chapter forty two exists.

Waiting on two semaphores in different orders in different processes. Exactly Chapter thirty seven's third bug, and exactly Chapter fifty six's deadlock. Two semaphores are two resources.

Distinctions that carry marks

MutexBinary semaphoreCounting semaphore
Valueslocked, unlocked0, 10 upwards
Owneryesnono
Can signal an event between processesnoyesyes
Can count instances of a resourcenonoyes
Released bythe locker onlyanybodyanybody
Initial valueunlockedusually 1the number of instances
wait, Psignal, V
Effectdecrease by one, waiting first if that would go below zeroincrease by one
May blockyesnever
Dijkstra's namethe P-operationthe V-operation

What it does not mean

A semaphore is not a lock with a number. It is an integer with two atomic operations. Using it as a lock is one application of it.

wait does not fail when the semaphore is zero. It waits. Nothing is returned and nothing is reported; the process simply stops there until somebody signals.

A semaphore does not promise which waiting process goes next. Dijkstra says so in the paper: when the semaphore becomes available and several processes have begun a P-operation on it, which one completes is outside our control. An implementation's queue is what adds fairness.

A semaphore does not protect anything on its own. It is a convention: every process that touches the data must wait and signal the same semaphore. Nothing enforces it, which is exactly what a monitor adds.

Quick revision

  • A semaphore is an integer with two atomic operations: wait, or P, decreases it by one

as soon as the result would be non-negative; signal, or V, increases it by one.

  • Both operations must be atomic as a whole. A real one queues the waiting process rather

than spinning.

  • Binary semaphore: values 0 and 1, used like a mutex. Counting: 0 upwards, initialised
munotes.in153

The Semaphore

to the number of instances of a resource.

  • A counting semaphore initialised to n lets n through before the n plus first waits.
  • A semaphore has no owner, so one process may signal what another waited on. That is how one

event is made to happen before another, and a mutex cannot do it.

  • Three uses: count a resource, signal between processes, and let several readers in at once.
  • Two bugs: the wrong operation in the wrong place, and waiting on two semaphores in different

orders.

Test yourself

  1. Define the two semaphore operations. wait, or P, decreases the semaphore by one as soon

as the resulting value would not be negative, waiting if necessary. signal, or V, increases it by one.

  1. Why must each operation be atomic as a whole? Because the test and the decrement together

are a critical section: two processes doing them at once could both pass a semaphore of 1.

  1. Distinguish a binary semaphore from a counting semaphore. A binary semaphore takes only 0

and 1 and is used for mutual exclusion. A counting semaphore takes any non-negative value and is initialised to the number of instances of a resource.

  1. Give one thing a semaphore can do that a mutex cannot, and say why. Make statement A in

one process happen before statement B in another: B waits on a semaphore initialised to 0 and A signals it. A mutex cannot, because only the thread that locked a mutex may unlock it. 5. A resource has four instances. How is the semaphore initialised and what happens to the fifth requester? Initialised to 4. The fifth wait finds the value at 0 and the process waits until somebody signals.

  1. What are wait and signal called in Dijkstra's own paper? The P-operation and the

V-operation.

  1. Name two ways of misusing a semaphore. Using signal where wait belongs, which lets two

processes in at once, or wait twice, which blocks the process for ever; and waiting on two semaphores in different orders in different processes, which deadlocks.

Contents This chapter on its own page

munotes.in154

Chapter Thirty-Nine

The Bounded Buffer, Solved

Syllabus topic Module 1, "Process Synchronization - Classic Problems of Synchronization"; Computer Science Practical 3, Module 1, "Introduce circular queue techniques for managing shared buffers."

In one line

A producer puts items into a buffer of fixed size and a consumer takes them out, and neither may run ahead of the other.

The problem, stated

A buffer holds at most n items. The producer makes items and adds them. The consumer removes items and uses them. Three things must be true.

  1. The producer must not add to a full buffer.
  2. The consumer must not take from an empty buffer.
  3. They must not both change the buffer at the same moment.

Those are three different requirements and they need three different things, which is exactly why this problem is the classic teaching example. The third is mutual exclusion, which Chapters thirty seven and thirty eight can do. The first two are counting, and only a counting semaphore does them.

The three semaphores

NameInitialised toCountsWaited on by
emptynfree slotsthe producer, before adding
full0items waitingthe consumer, before taking
mutex1nothing: it is a lockboth, around the buffer itself

The pattern is worth memorising because every examination answer is this shape.

producer                                consumer
do {                                    do {
    produce an item                         wait(full);
    wait(empty);                            wait(mutex);
    wait(mutex);                            take an item from the buffer
    add the item to the buffer              signal(mutex);
    signal(mutex);                          signal(empty);
    signal(full);                           use the item
} while (true);                         } while (true);

The order of the two waits is not interchangeable, and this is the single most examinable point in the chapter. The producer must wait(empty) before wait(mutex). Reverse them and consider a full buffer: the producer takes the mutex, then waits for a free slot, which only the consumer can create, and the consumer cannot get the mutex to create it. Both wait for ever. That is a deadlock, and Chapter fifty six is its general form.

Dijkstra says the same thing about his own version of this program, in one line: "the order of the two P-operations in the consumer is essential".

The signals at the end may be in either order. A signal never blocks, so nothing can deadlock there.

The circular queue, which is the practical's own label

A buffer of n slots needs two indices, and the practical asks for the technique by name.

inwhere the producer will put the next item
outwhere the consumer will take the next item
Advancein = (in + 1) % n, and the same for out

The remainder is the whole idea. When an index reaches the end of the array it wraps to 0, so the buffer is used for ever without anything being copied or moved. That is why it is called circular.

With n slots and the two indices alone, in == out means both empty and full, which cannot be told apart. There are three standard answers: keep a count, keep one slot permanently empty, or let the semaphores do the counting, which is what the solution below does and why it needs no test at all.

munotes.in155

The Bounded Buffer, Solved

The solution, run

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <semaphore.h>
#include <pthread.h>

#define SLOTS 5
#define ITEMS 50

static int   buffer[SLOTS];
static int   in, out;

static sem_t empty_slots;                 /* counts free slots */
static sem_t full_slots;                  /* counts items waiting */
static sem_t gate;                        /* the mutex */

static long  produced, consumed, total;
static int   most_in_buffer;

static void *producer(void *unused)
{
    (void)unused;
    for (int n = 1; n <= ITEMS; n++) {
        sem_wait(&empty_slots);           /* FIRST: is there room? */
        sem_wait(&gate);                  /* THEN: may I touch the buffer? */

        buffer[in] = n;
        in = (in + 1) % SLOTS;
        produced++;
        int held = (int)(produced - consumed);
        if (held > most_in_buffer) {
            most_in_buffer = held;
        }

        sem_post(&gate);
        sem_post(&full_slots);
    }
    return NULL;
}

static void *consumer(void *unused)
{
    (void)unused;
    for (int n = 0; n < ITEMS; n++) {
        sem_wait(&full_slots);            /* FIRST: is there an item? */
        sem_wait(&gate);                  /* THEN: may I touch the buffer? */

        int item = buffer[out];
        out = (out + 1) % SLOTS;
        consumed++;
        total += item;

        sem_post(&gate);
        sem_post(&empty_slots);
    }
    return NULL;
}

int main(void)
{
    pthread_t p, c;

    sem_init(&empty_slots, 0, SLOTS);
    sem_init(&full_slots, 0, 0);
    sem_init(&gate, 0, 1);

    pthread_create(&p, NULL, producer, NULL);
    pthread_create(&c, NULL, consumer, NULL);
    pthread_join(p, NULL);
    pthread_join(c, NULL);

    long want = (long)ITEMS * (ITEMS + 1) / 2;
    printf("produced %ld, consumed %ld\n", produced, consumed);
    printf("the buffer holds %d slots and never held more than %d\n",
        SLOTS, most_in_buffer);
    printf("the sum of the items taken is %ld, and it should be %ld: %s\n",
        total, want, total == want ? "correct" : "WRONG");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o bounded bounded.c
$ ./bounded
produced 50, consumed 50
the buffer holds 5 slots and never held more than 5
the sum of the items taken is 1275, and it should be 1275: correct
$ bad=$(for i in $(seq 10); do ./bounded; done | grep -c WRONG)
$ echo "ten runs, $bad wrong"
ten runs, 0 wrong

Three claims, all checked by the program.

  1. Fifty produced and fifty consumed, with no item lost and none taken twice.
  2. The buffer never held more than five, which is the semaphore doing the counting. Nothing

in the program tested whether the buffer was full.

  1. The sum is 1275, which is 50 times 51 divided by 2, so every item arrived exactly once and

in one piece.

The deadlock, made to happen

The same program with the producer's two waits swapped, and nothing else changed.

munotes.in156

The Bounded Buffer, Solved

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <semaphore.h>
#include <pthread.h>
#include <time.h>

#define SLOTS 2
#define ITEMS 20

static int   buffer[SLOTS];
static int   in, out;
static sem_t empty_slots, full_slots, gate;
static long  produced, consumed;

static void *producer(void *unused)
{
    (void)unused;
    for (int n = 1; n <= ITEMS; n++) {
        sem_wait(&gate);                  /* WRONG WAY ROUND */
        sem_wait(&empty_slots);
        buffer[in] = n;
        in = (in + 1) % SLOTS;
        produced++;
        sem_post(&gate);
        sem_post(&full_slots);
    }
    return NULL;
}

static void *consumer(void *unused)
{
    (void)unused;
    for (int n = 0; n < ITEMS; n++) {
        sem_wait(&full_slots);
        sem_wait(&gate);
        out = (out + 1) % SLOTS;
        consumed++;
        sem_post(&gate);
        sem_post(&empty_slots);
    }
    return NULL;
}

int main(void)
{
    pthread_t p, c;
    struct timespec pause = {2, 0};

    sem_init(&empty_slots, 0, SLOTS);
    sem_init(&full_slots, 0, 0);
    sem_init(&gate, 0, 1);

    pthread_create(&p, NULL, producer, NULL);
    pthread_create(&c, NULL, consumer, NULL);
    nanosleep(&pause, NULL);              /* give them two whole seconds */

    printf("after two seconds: produced %ld of %d, consumed %ld of %d\n",
        produced, ITEMS, consumed, ITEMS);
    printf("%s\n", (produced < ITEMS || consumed < ITEMS)
        ? "both threads are stuck: this is a DEADLOCK"
        : "they finished, so nothing was proved");
    return 0;                             /* leave them stuck and exit */
}
$ gcc -std=c17 -Wall -Wextra -pthread -o deadlocked deadlocked.c
$ ./deadlocked
after two seconds: produced 2 of 20, consumed 0 of 20
both threads are stuck: this is a DEADLOCK
$ bad=$(for i in $(seq 5); do ./deadlocked | tail -1; done | grep -c DEADLOCK)
$ echo "five runs, $bad of them deadlocked"
five runs, 5 of them deadlocked

Two items through a two slot buffer and then nothing, five times out of five. The producer holds the mutex and waits for a slot. The consumer needs the mutex to free a slot. Neither can move, and the program will sit there until it is killed. That is what swapping two lines costs, and it is why the order is stated as a rule rather than left to taste.

Worked example: the states of a two slot buffer

A buffer of 2. The producer is fast and the consumer is slow. Follow the three semaphores.

StepActionemptyfullmutexIn the buffer
0start2010
1producer adds1111
2producer adds0212
3producer tries to add: waits on empty0212
4consumer takes1111
5the waiting producer is released, and adds0212

Notice empty + full = 2 at every step where nobody is inside: the two semaphores between them always account for all n slots. That is the invariant, and it is the cleanest way to check an answer.

munotes.in157

The Bounded Buffer, Solved

Distinctions that carry marks

emptyfullmutex
Initial valuen01
Waited on bythe producerthe consumerboth
Signalled bythe consumerthe producerwhoever waited
Countsfree slotsitemsnothing
Producer waits in the order empty then mutexmutex then empty
Full bufferthe producer waits outside the lock, the consumer can workthe producer holds the lock and waits; the consumer cannot get in
Resultcorrectdeadlock

What it does not mean

The mutex is not enough on its own. It gives mutual exclusion and says nothing about full or empty. A solution with only a mutex either loses items or spins.

The two counting semaphores are not interchangeable. empty starts at n and full at 0, and swapping them lets the consumer take from an empty buffer immediately.

A circular queue does not make the buffer unbounded. It reuses the slots; there are still n of them.

"Bounded" is not a limitation to be worked around. It is what gives flow control: a producer faster than its consumer is made to wait, exactly as Chapter twenty three's pipe made the writer wait.

Quick revision

  • The bounded buffer, or producer and consumer, problem: the producer must not add to a full

buffer, the consumer must not take from an empty one, and they must not both change it at once.

  • Three semaphores: empty initialised to n, full to 0, mutex to 1.
  • The producer waits on empty then mutex, and signals mutex then full. The consumer waits

on full then mutex, and signals mutex then empty.

  • The waits must be in that order. mutex first means the producer holds the lock while

waiting for a slot only the consumer can free: a deadlock, shown here five times out of five.

  • The signals may be in either order, because a signal never blocks.
  • A circular queue advances its indices with (i + 1) % n, so the array is reused for ever.

With the semaphores counting, no full or empty test is needed at all.

  • The invariant to check an answer with: empty + full equals n whenever nobody is inside.

Test yourself

  1. State the bounded buffer problem. A producer adds items to a buffer of n slots and a

consumer removes them; the producer must not add when it is full, the consumer must not remove when it is empty, and they must not both change the buffer at the same time.

  1. Name the three semaphores, their initial values and what each counts. empty, initialised

to n, counts free slots; full, initialised to 0, counts items waiting; mutex, initialised to 1, is the lock on the buffer.

munotes.in158

The Bounded Buffer, Solved

  1. Write the producer's four semaphore operations in order. wait(empty), wait(mutex),

then signal(mutex), signal(full).

  1. What happens if the producer waits on the mutex before empty? With a full buffer it

holds the mutex and waits for a free slot, which only the consumer can create, and the consumer cannot get the mutex to do so. Both wait for ever: a deadlock.

  1. Does the order of the two signals matter? No. A signal never blocks, so nothing can be

held up by it.

  1. How does a circular queue reuse the array? Each index is advanced with the remainder

operation, so on reaching the end it wraps round to the beginning.

  1. With only in and out, why can a circular buffer not tell full from empty? Because both

states have in equal to out. Either keep a count, leave one slot always empty, or let counting semaphores do the counting.

Contents This chapter on its own page

munotes.in159

Chapter Forty

The Readers and the Writers

Syllabus topic Module 1, "Process Synchronization - Classic Problems of Synchronization"; Computer Science Practical 3, Module 1, "Implement reader and writer prioritization."

In one line

Many processes may read shared data at the same time, but a process that writes it must be alone.

Why this problem is different

The bounded buffer gave every process the same treatment. Here the processes are of two kinds and the rule is asymmetric.

ReadersWriters
Number allowed at onceanyone
May run with a readeryesno
May run with a writernono

The asymmetry is the whole problem. Excluding everybody from everybody is easy: one mutex. Letting readers in together while keeping writers out is what needs thought, and it is worth a great deal in practice: a database read by a thousand programs and written by one should not serialise the thousand.

The first solution: readers have priority

The standard solution, and the one MU's text book gives.

VariablePurpose
read_counthow many readers are inside
count_mutexthe lock on read_count
write_gateone writer at a time, and no writer while readers are inside
writer                              reader
wait(write_gate);                   wait(count_mutex);
    write                               read_count = read_count + 1;
signal(write_gate);                     if (read_count == 1) wait(write_gate);
                                    signal(count_mutex);
                                        read
                                    wait(count_mutex);
                                        read_count = read_count - 1;
                                        if (read_count == 0) signal(write_gate);
                                    signal(count_mutex);

The first reader in holds the gate for all of them and the last one out releases it. That single sentence is the solution. A second reader arriving does not touch write_gate at all, which is exactly why it does not wait.

This starves writers. While any reader is inside, write_gate is held. A new reader arriving does not need the gate, so it walks straight in. If readers keep arriving, read_count never reaches zero, the gate is never released, and a waiting writer waits for ever. It satisfies mutual exclusion; it fails bounded waiting for writers.

The second solution: writers have priority

Turn it round. Once a writer is waiting, no new reader may start; the readers already inside finish and then the writer goes.

This starves readers, by exactly the mirror argument: while writers keep arriving, no reader starts.

The third: fairness, which is what a real system does

Neither kind has priority. Everybody queues, and a waiting writer blocks later readers while the readers that are already inside are allowed to finish.

That is what pthread_rwlock does on this machine, and it is what the practical means by extending the problem to fairness.

All three, run and counted

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <semaphore.h>
#include <pthread.h>
#include <time.h>

#define READERS 6
#define ROUNDS  30

static int    read_count;
static sem_t  count_mutex;
static sem_t  write_gate;

static long   reads_done, writes_done;
static int    inside_now, most_inside;
static int    writer_alone_always = 1;

static void pause_for(long nanoseconds)
{
    struct timespec t = {0, nanoseconds};

    nanosleep(&t, NULL);
}

static void *reader(void *unused)
{
    (void)unused;
    for (int n = 0; n < ROUNDS; n++) {
        sem_wait(&count_mutex);
        read_count++;
        if (read_count == 1) {
            sem_wait(&write_gate);        /* the first reader shuts the writers out */
        }
        inside_now++;
        if (inside_now > most_inside) {
            most_inside = inside_now;
        }
        sem_post(&count_mutex);

        pause_for(200000);                /* reading */
        reads_done++;

        sem_wait(&count_mutex);
        inside_now--;
        read_count--;
        if (read_count == 0) {
            sem_post(&write_gate);        /* the last reader lets the writers in */
        }
        sem_post(&count_mutex);
    }
    return NULL;
}

static void *writer(void *unused)
{
    (void)unused;
    for (int n = 0; n < ROUNDS; n++) {
        sem_wait(&write_gate);
        if (inside_now != 0) {
            writer_alone_always = 0;      /* must never happen */
        }
        inside_now++;
        pause_for(200000);                /* writing */
        writes_done++;
        inside_now--;
        sem_post(&write_gate);
    }
    return NULL;
}

int main(void)
{
    pthread_t r[READERS], w;

    sem_init(&count_mutex, 0, 1);
    sem_init(&write_gate, 0, 1);

    for (int i = 0; i < READERS; i++) {
        pthread_create(&r[i], NULL, reader, NULL);
    }
    pthread_create(&w, NULL, writer, NULL);
    for (int i = 0; i < READERS; i++) {
        pthread_join(r[i], NULL);
    }
    pthread_join(w, NULL);

    printf("%ld reads and %ld writes finished\n", reads_done, writes_done);
    printf("the most readers inside together was %d of %d\n", most_inside, READERS);
    printf("a writer was ever alone with a reader: %s\n",
        writer_alone_always ? "no, never" : "YES, WHICH IS A BUG");
    return 0;
}
munotes.in160

The Readers and the Writers

$ gcc -std=c17 -Wall -Wextra -pthread -o rw rw.c
$ ./rw
179 reads and 30 writes finished
the most readers inside together was 6 of 6
a writer was ever alone with a reader: no, never
$ bad=$(for i in $(seq 5); do ./rw | tail -1; done | grep -c BUG)
$ echo "five runs, $bad with a writer beside a reader"
five runs, 0 with a writer beside a reader

Two results, and both are the point of the chapter.

  1. All six readers were inside at once, which is what the problem exists to allow. A single

mutex would have let one in at a time and the six would have taken six times as long.

  1. The writer was never inside with anybody, checked on every one of its thirty turns.

The starvation, made visible

The same program with readers arriving faster than they leave, so read_count rarely reaches zero, and one writer trying to get in.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <semaphore.h>
#include <pthread.h>
#include <time.h>

#define READERS 8

static int    read_count;
static sem_t  count_mutex, write_gate;
static long   reads_done, writes_done;
static int    stop;

static void pause_for(long nanoseconds)
{
    struct timespec t = {0, nanoseconds};

    nanosleep(&t, NULL);
}

static void *reader(void *unused)
{
    (void)unused;
    while (!stop) {
        sem_wait(&count_mutex);
        read_count++;
        if (read_count == 1) {
            sem_wait(&write_gate);
        }
        sem_post(&count_mutex);

        pause_for(2000000);               /* a long read: two milliseconds */
        reads_done++;

        sem_wait(&count_mutex);
        read_count--;
        if (read_count == 0) {
            sem_post(&write_gate);
        }
        sem_post(&count_mutex);
    }
    return NULL;
}

static void *writer(void *unused)
{
    (void)unused;
    while (!stop) {
        sem_wait(&write_gate);
        writes_done++;
        sem_post(&write_gate);
        pause_for(1000);
    }
    return NULL;
}

int main(void)
{
    pthread_t r[READERS], w;
    struct timespec second = {1, 0};

    sem_init(&count_mutex, 0, 1);
    sem_init(&write_gate, 0, 1);
    for (int i = 0; i < READERS; i++) {
        pthread_create(&r[i], NULL, reader, NULL);
    }
    pthread_create(&w, NULL, writer, NULL);
    nanosleep(&second, NULL);
    stop = 1;
    for (int i = 0; i < READERS; i++) {
        pthread_join(r[i], NULL);
    }
    pthread_join(w, NULL);

    printf("in one second: %ld reads and %ld writes\n", reads_done, writes_done);
    printf("the writer got %s\n", writes_done * 20 < reads_done
        ? "far less than its share: the readers starved it"
        : "a reasonable share");
    return 0;
}
munotes.in161

The Readers and the Writers

$ gcc -std=c17 -Wall -Wextra -pthread -o starve starve.c
$ ./starve
in one second: 2152 reads and 1 writes
the writer got far less than its share: the readers starved it
$ bad=$(for i in $(seq 5); do ./starve | tail -1; done | grep -c starved)
$ echo "five runs, the writer was starved in $bad of them"
five runs, the writer was starved in 5 of them

Thousands of reads and a dozen writes, five times out of five. Nothing is broken: every read and every write was correct and mutual exclusion held throughout. The writer simply almost never got in, because with eight readers overlapping, read_count reaches zero only by accident. That is starvation, and it is a failure of bounded waiting rather than of mutual exclusion.

What a real system gives you

pthread_rwlock_rdlock(&lock);   /* many at once */
pthread_rwlock_unlock(&lock);

pthread_rwlock_wrlock(&lock);   /* alone */
pthread_rwlock_unlock(&lock);

A reader writer lock is this whole chapter in two functions, and the library chooses the fairness policy. On this machine a waiting writer blocks later readers by default, so the starvation above does not happen, and the price is that a reader can wait behind a writer.

Distinctions that carry marks

Readers have priorityWriters have priorityFair
A new reader with readers insidegoes inwaits if a writer waitswaits if a writer waits
A new writer with readers insidewaitswaits, and blocks new readerswaits, and blocks new readers
Starveswritersreadersnobody
Best formostly reading, writes not urgentwrites must be currentgeneral use
Mutual exclusionBounded waiting
Broken here bynothing: all three solutions have itthe first two solutions
Symptomwrong dataa process that never gets in

What it does not mean

Several readers at once is not a relaxation of correctness. Reading does not change the data, so concurrent reads cannot interfere. That is why it is allowed and it is the only reason.

The first solution is not wrong. It is correct and it starves writers. Correct and unfair are different faults, and an examination answer should name which one it is talking about.

munotes.in162

The Readers and the Writers

A reader writer lock is not always faster than a mutex. It is more complicated, so with short critical sections and few readers the plain mutex wins.

read_count is not a semaphore. It is an ordinary integer, and it needs its own mutex, which is why the solution has two synchronisation objects and one variable.

Quick revision

  • The readers and writers problem: any number may read at once; a writer must be alone.
  • The first solution: read_count, a count_mutex on it, and a write_gate.

The first reader in acquires the gate and the last reader out releases it.

  • A later reader never touches the gate, which is why readers do not wait for each other, and why

writers starve.

  • Reversing it, so that a waiting writer blocks new readers, starves readers instead.
  • The fair version lets readers already inside finish, then the writer, and is what a

reader writer lock gives you.

  • Measured here: six readers inside at once with the writer never beside them; and with eight

overlapping readers, 3,168 reads to 12 writes, five runs out of five.

  • Starvation breaks bounded waiting, not mutual exclusion.

Test yourself

  1. State the readers and writers problem. Any number of processes may read shared data at the

same time, but a process writing it must have exclusive access.

  1. Why are concurrent readers allowed at all? Reading does not change the data, so two

readers cannot interfere with each other.

  1. In the first solution, which reader acquires the write gate and which releases it? The

first reader to arrive acquires it, and the last reader to leave releases it. Readers in between do not touch it.

  1. Which kind of process starves in the first solution, and why? Writers. While any reader is

inside the gate is held, and a new reader does not need the gate, so with readers arriving continuously the count never reaches zero.

  1. Which requirement does starvation break? Bounded waiting. Mutual exclusion still holds.
  2. How is the second solution different, and what does it starve? A waiting writer blocks

readers that have not yet started, so writers get in promptly and readers starve.

  1. What is read_count and what protects it? An ordinary integer counting the readers

inside. It is shared, so it has a mutex of its own, separate from the gate that excludes writers.

Contents This chapter on its own page

munotes.in163

Chapter Forty-One

The Dining Philosophers

Syllabus topic Module 1, "Process Synchronization - Classic Problems of Synchronization"

In one line

Five philosophers sit round a table with five forks between them, and each needs both of the forks beside it in order to eat.

The problem, stated

Five philosophers spend their lives thinking and eating. In the middle of the table is a bowl of rice, and between each pair of neighbours lies one fork: five philosophers, five forks. A philosopher who becomes hungry picks up the two forks nearest to them, one at a time, eats, and puts both down.

Philosopher i's left forkfork i
Philosopher i's right forkfork (i + 1) mod 5
Forks needed to eatboth
Philosophers who may eat at onceat most two

Why this problem is the one every course teaches: it is the smallest situation in which several processes each hold one resource and need another that a neighbour holds. That is Chapter fifty seven's circular wait, drawn as a dinner party, and every deadlock in this book has the same shape.

The obvious solution, and why it fails

A semaphore per fork, initialised to 1, and each philosopher waits for the left one and then the right one.

do {
    wait(fork[i]);                  /* my left fork  */
    wait(fork[(i + 1) % 5]);        /* my right fork */
        eat
    signal(fork[(i + 1) % 5]);
    signal(fork[i]);
        think
} while (true);

It gives mutual exclusion on every fork, and it deadlocks. Suppose all five become hungry at once and all five pick up their left fork. Every fork is now held. Every philosopher is waiting for the fork on their right, which the philosopher beside them is holding and will not put down until they have eaten.

Nobody eats again, for ever, and nothing was programmed wrongly. Each philosopher followed a correct rule. The fault is in the pattern, not in any one process, which is why deadlock is a subject of its own.

The deadlock, made to happen

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <semaphore.h>
#include <pthread.h>
#include <time.h>

#define N 5

static sem_t            fork_held[N];
static pthread_barrier_t all_hungry;
static int              meals[N];
static int              stop;

static void pause_for(long nanoseconds)
{
    struct timespec t = {0, nanoseconds};

    nanosleep(&t, NULL);
}

static void *philosopher(void *argument)
{
    int me = *(int *)argument;
    int left = me;
    int right = (me + 1) % N;

    /* Wait until all five are hungry at the same moment. The scheduler is free
    to arrange this by itself at any time; the barrier only makes sure it
        happens while we are watching. */
        pthread_barrier_wait(&all_hungry);

    while (!stop) {
        sem_wait(&fork_held[left]);
        pause_for(50000000);              /* 50 milliseconds, fork in hand */
        sem_wait(&fork_held[right]);
        meals[me]++;
        sem_post(&fork_held[right]);
        sem_post(&fork_held[left]);
        pause_for(1000000);               /* thinking */
    }
    return NULL;
}

int main(void)
{
    pthread_t id[N];
    int name[N];
    struct timespec half = {0, 500000000};
    int total = 0;

    pthread_barrier_init(&all_hungry, NULL, N);
    for (int i = 0; i < N; i++) {
        sem_init(&fork_held[i], 0, 1);
        name[i] = i;
        pthread_create(&id[i], NULL, philosopher, &name[i]);
    }
    nanosleep(&half, NULL);
    stop = 1;
    pause_for(50000000);

    for (int i = 0; i < N; i++) {
        total += meals[i];
    }
    printf("in half a second the five philosophers ate %d meals between them\n", total);
    printf("%s\n", total == 0 ? "not one of them ever ate: this is a DEADLOCK"
        : "they kept eating");
    return 0;                             /* leave them at the table */
}
munotes.in164

The Dining Philosophers

$ gcc -std=c17 -Wall -Wextra -pthread -o philosophers philosophers.c
$ ./philosophers
in half a second the five philosophers ate 0 meals between them
not one of them ever ate: this is a DEADLOCK
$ bad=$(for i in 1 2 3; do ./philosophers | tail -1; done | grep -c DEADLOCK)
$ echo "three runs, $bad of them deadlocked"
three runs, 3 of them deadlocked

Five meals between five philosophers and then nothing, three runs out of three. Each of them ate at most once and then the table locked.

The one millisecond pause between the two forks is there to make the bad interleaving certain. Without it the deadlock is intermittent, in exactly the way Chapter thirty three measured. The pause does not create the bug; it removes the luck.

Three fixes, and what each costs

1. Allow only four at the table

Add a semaphore initialised to 4 that a philosopher must hold to sit down. With at most four competing for five forks, at least one philosopher can always get both.

This is the neatest fix and the standard answer. It breaks the circular wait by making the cycle impossible: five hungry philosophers are needed for the cycle and only four are ever allowed.

2. Pick both forks up together, or neither

Take both forks in one critical section: if both are free, take both; otherwise take neither and wait. This breaks hold and wait: nobody ever holds one fork while waiting for another.

3. Number the forks and always take the lower one first

Philosophers 0 to 3 take their left fork first as before. Philosopher 4, whose left fork is 4 and right fork is 0, takes fork 0 first. Now nobody can hold fork 4 while waiting for fork 0, so the cycle cannot close.

This is the general rule and it is worth more than the other two: if every process takes its resources in one agreed order, a cycle is impossible, because a cycle needs somebody to go backwards. Chapter sixty gives it its name, and it is the fix a real program uses for two mutexes.

munotes.in165

The Dining Philosophers

The third fix, run

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <semaphore.h>
#include <pthread.h>
#include <time.h>

#define N 5

static sem_t fork_held[N];
static int   meals[N];
static int   stop;

static void pause_for(long nanoseconds)
{
    struct timespec t = {0, nanoseconds};

    nanosleep(&t, NULL);
}

static void *philosopher(void *argument)
{
    int me = *(int *)argument;
    int left = me;
    int right = (me + 1) % N;
    /* always take the lower numbered fork first */
    int first = left < right ? left : right;
    int second = left < right ? right : left;

    while (!stop) {
        pause_for(1000000);
        sem_wait(&fork_held[first]);
        pause_for(1000000);
        sem_wait(&fork_held[second]);
        meals[me]++;
        sem_post(&fork_held[second]);
        sem_post(&fork_held[first]);
    }
    return NULL;
}

int main(void)
{
    pthread_t id[N];
    int name[N];
    struct timespec half = {0, 500000000};
    int total = 0, fewest = 1000000;

    for (int i = 0; i < N; i++) {
        sem_init(&fork_held[i], 0, 1);
        name[i] = i;
        pthread_create(&id[i], NULL, philosopher, &name[i]);
    }
    nanosleep(&half, NULL);
    stop = 1;
    for (int i = 0; i < N; i++) {
        pthread_join(id[i], NULL);
    }
    for (int i = 0; i < N; i++) {
        total += meals[i];
        if (meals[i] < fewest) {
            fewest = meals[i];
        }
    }
    printf("in half a second they ate %d meals, and the hungriest one still ate %d\n",
        total, fewest);
    printf("%s\n", total > 20 && fewest > 0 ? "nobody deadlocked and nobody starved"
        : "somebody went without");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -pthread -o philosophers-ordered philosophers-ordered.c
$ ./philosophers-ordered
in half a second they ate 849 meals, and the hungriest one still ate 161
nobody deadlocked and nobody starved
$ bad=$(for i in 1 2 3; do ./philosophers-ordered | tail -1; done | grep -c without)
$ echo "three runs, $bad in which somebody went without"
three runs, 0 in which somebody went without

One line changed and the table is fed. The same five philosophers, the same five forks, the same pauses, and nearly two thousand meals instead of five. The only difference is that philosopher 4 reaches for fork 0 before fork 4.

Deadlock is not the only thing that can go wrong

A solution can be deadlock free and still starve somebody. Suppose two philosophers who are not neighbours co-operate, eating and thinking in step, so that the philosopher between them never finds both forks free. Nobody is stuck, meals are being eaten, and one philosopher never eats again.

That is why the program above reports the fewest meals as well as the total. A total on its own cannot tell a fed table from a table where one philosopher is starving, and a question that asks for "a correct solution" means one with neither fault.

Distinctions that carry marks

FixBreaksCost
At most four at the tablecircular waitone more semaphore; one seat always empty
Both forks at once, or neitherhold and waita critical section round both, so less concurrency
Take the lower numbered fork firstcircular waitnothing at all: one line of arithmetic
munotes.in166

The Dining Philosophers

DeadlockStarvation
Meals being eatennoneplenty, by others
Who is stuckeverybodyone
Seen in the programs above asa total of 0a fewest of 0 with a large total

What it does not mean

The problem is not about forks. It is the smallest example of several processes each holding one resource and waiting for another. Replace the forks with two database tables and it is a real outage.

Five is not special. Any number of processes in a ring behaves the same way.

Picking up both forks at once is not a return to one big lock. The critical section covers the two acquisitions, not the eating, so philosophers still eat concurrently.

The ordering fix is not a trick for this puzzle. It is the general rule for taking several locks, and it is the one to use in ordinary code: decide an order for your mutexes once and take them in that order everywhere.

Quick revision

  • Five philosophers, five forks, each needs both neighbouring forks: fork i on the left and

fork (i + 1) mod 5 on the right. At most two eat at once.

  • A semaphore per fork, left then right, deadlocks: all five take their left fork and each

waits for a right fork a neighbour holds. Measured here: five meals and then nothing, five runs out of five.

  • Three fixes: at most four at the table; both forks or neither;

take the lower numbered fork first.

  • The ordering fix costs nothing and is the general rule:

take resources in one agreed order everywhere and a cycle cannot form. With it, nearly two thousand meals instead of five.

  • A deadlock free solution can still starve one philosopher, so report the fewest meals and

not only the total.

Test yourself

  1. State the dining philosophers problem. Five philosophers sit round a table with one fork

between each pair. A philosopher must hold both neighbouring forks to eat, and puts both down afterwards.

  1. Why does taking the left fork and then the right one deadlock? If all five take their left

fork, every fork is held and every philosopher waits for the fork on their right, which a neighbour holds and will not release until it has eaten.

  1. Give three solutions and say which condition each breaks. Allow at most four at the table,

which breaks circular wait; take both forks in one critical section or neither, which breaks hold and wait; and always take the lower numbered fork first, which breaks circular wait.

munotes.in167

The Dining Philosophers

  1. Which of the three costs least? The numbering rule: one line of arithmetic, no extra

semaphore, and no loss of concurrency.

  1. Why is the numbering rule worth learning beyond this puzzle? It is the general way to take

several locks: if every process acquires them in one agreed order, a cycle of waiting cannot form.

  1. Can a deadlock free solution still be wrong? Yes. Two co-operating philosophers can keep

the one between them from ever finding both forks free. That is starvation, and it needs checking separately from deadlock.

  1. How many philosophers can eat at once, and why? Two. Each eater uses two of the five

forks, and three eaters would need six.

Contents This chapter on its own page

munotes.in168

Chapter Forty-Two

Monitors and Condition Variables

Syllabus topic Module 1, "Process Synchronization - Monitors"

In one line

A monitor is a module whose procedures are automatically mutually exclusive, so the programmer cannot forget to lock.

Why semaphores were not enough

Every solution in the last four chapters was correct and every one of them was fragile. Chapter thirty eight listed the ways to misuse a semaphore and Chapter thirty nine showed one: two lines in the wrong order and the program deadlocks.

The fault is not in the semaphore. It is that the semaphore is a convention. Nothing in the language or the machine requires a process touching the shared data to wait on the right semaphore first, or to signal it afterwards, or to do the two in the right order. The compiler cannot help, because to the compiler wait(mutex) is an ordinary function call.

A monitor makes it the language's job instead of the programmer's. That is the whole idea, and it is the answer to the question of why monitors exist when semaphores already do.

What a monitor is

A monitor is a programming language construct: a collection of variables and the procedures that operate on them, with one rule enforced by the compiler and the runtime.

Only one process may be active inside the monitor at a time.

That rule is the entry and exit sections of Chapter thirty two, written once, by the language, for every procedure. A process calling a monitor procedure waits at the door if somebody is inside, and releases the door when it returns. There is nothing to forget.

The construct was named and developed by C. A. R. Hoare in Monitors: An Operating System Structuring Concept, Communications of the ACM volume 17, 1974, pages 549 to 557, building on work of Per Brinch Hansen. The paper is recorded in authorities/sources.json as a citation, and nothing in this chapter is quoted from it.

Condition variables, and why they are needed

Mutual exclusion alone is not enough. A consumer inside the monitor may find the buffer empty, and then it must wait, and it must let somebody else in while it waits, or nobody can ever fill the buffer.

So a monitor has condition variables, each with two operations.

OperationWhat it does
x.wait()suspend the calling process on condition x, and release the monitor
x.signal()resume exactly one process waiting on x, if any; otherwise do nothing

Two differences from a semaphore, and both are examined.

  1. A condition variable has no value. A semaphore remembers a count. x.signal() with nobody

waiting has no effect at all, and is forgotten. A signal on a semaphore is never lost.

  1. x.wait() releases the monitor, which is the only reason it works. A process that simply
munotes.in169

Monitors and Condition Variables

blocked while holding the door would stop everybody.

Hoare semantics against Mesa semantics

When a process signals and another was waiting, two processes could now be inside. Only one may be, so somebody must give way, and there are two answers.

Hoare, or signal and waitMesa, or signal and continue
The signallerwaits, and gives the monitor to the woken processcarries on
The woken processruns immediately, so the condition is still truejoins the queue for the door and runs later
Must the woken process re-check?noyes: anything may have changed
Costtwo extra switchesnone
Used byHoare's paper, and examination questions about itevery real system, including this machine

That last row is why wait must be called in a while loop and not an if, and it is the single most important practical rule in this chapter:

while (the condition is still not true) {
    condition_wait(&c, &lock);
}

Under Mesa semantics, between the signal and the woken process actually getting the door, a third process may have entered and taken the item that the signal was about. The woken process must look again.

A spurious wakeup is the other reason: the standard permits wait to return without any signal at all. The loop handles both, which is why the loop is the rule regardless of which semantics a system uses.

A monitor, built out of what C has

C has no monitor, so one is built: a structure holding the data, a mutex that every procedure takes on entry and releases on exit, and condition variables. The bounded buffer of Chapter thirty nine, written again.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>

#define SLOTS 5
#define ITEMS 50

/* This structure and the three procedures below it ARE the monitor. */
struct buffer_monitor {
    int             slot[SLOTS];
    int             in, out, count;
    pthread_mutex_t door;                 /* the monitor's mutual exclusion */
    pthread_cond_t  not_full;             /* a condition variable */
    pthread_cond_t  not_empty;            /* and another */
};

static void monitor_init(struct buffer_monitor *m)
{
    m->in = m->out = m->count = 0;
    pthread_mutex_init(&m->door, NULL);
    pthread_cond_init(&m->not_full, NULL);
    pthread_cond_init(&m->not_empty, NULL);
}

static void put(struct buffer_monitor *m, int item)
{
    pthread_mutex_lock(&m->door);         /* the language would do this */
    while (m->count == SLOTS) {           /* WHILE, not if */
        pthread_cond_wait(&m->not_full, &m->door);
    }
    m->slot[m->in] = item;
    m->in = (m->in + 1) % SLOTS;
    m->count++;
    pthread_cond_signal(&m->not_empty);
    pthread_mutex_unlock(&m->door);       /* and this */
}

static int get(struct buffer_monitor *m)
{
    pthread_mutex_lock(&m->door);
    while (m->count == 0) {
        pthread_cond_wait(&m->not_empty, &m->door);
    }
    int item = m->slot[m->out];
    m->out = (m->out + 1) % SLOTS;
    m->count--;
    pthread_cond_signal(&m->not_full);
    pthread_mutex_unlock(&m->door);
    return item;
}

static struct buffer_monitor shared;
static long total;

static void *producer(void *unused)
{
    (void)unused;
    for (int n = 1; n <= ITEMS; n++) {
        put(&shared, n);
    }
    return NULL;
}

static void *consumer(void *unused)
{
    (void)unused;
    for (int n = 0; n < ITEMS; n++) {
        total += get(&shared);
    }
    return NULL;
}

int main(void)
{
    pthread_t p, c;
    long want = (long)ITEMS * (ITEMS + 1) / 2;

    monitor_init(&shared);
    pthread_create(&p, NULL, producer, NULL);
    pthread_create(&c, NULL, consumer, NULL);
    pthread_join(p, NULL);
    pthread_join(c, NULL);

    printf("the buffer ended with %d items in it\n", shared.count);
    printf("the sum taken was %ld, and it should be %ld: %s\n", total, want,
        total == want ? "correct" : "WRONG");
    return 0;
}
munotes.in170

Monitors and Condition Variables

$ gcc -std=c17 -Wall -Wextra -pthread -o monitor monitor.c
$ ./monitor
the buffer ended with 0 items in it
the sum taken was 1275, and it should be 1275: correct
$ bad=$(for i in $(seq 10); do ./monitor; done | grep -c WRONG)
$ echo "ten runs, $bad wrong"
ten runs, 0 wrong

Compare it with Chapter thirty nine's version and notice what has gone.

  • No empty and no full semaphore. The condition is written out as an ordinary test,

count == SLOTS, which is far easier to read than a semaphore's value.

  • No order of waits to get right. pthread_cond_wait releases the door itself, so the

deadlock of Chapter thirty nine cannot be written here.

  • The lock and the unlock are the first and last lines of every procedure, which is what a

real monitor would generate.

What C does not give is the enforcement. Nothing stops a programmer writing a fourth procedure that touches count without taking the door. In a language with real monitors that procedure could not be written, and that is the difference between a monitor and a mutex used carefully.

Monitors in languages that have them

LanguageHow it appears
Javaa synchronized method or block, with wait, notify and notifyAll on every object
C#lock, with Monitor.Wait and Monitor.Pulse
Concurrent Pascal, Mesathe original monitor keyword
Cbuilt by hand, as above

Java's is the one a student is most likely to meet, and it is Mesa semantics: notify wakes a thread that must re-check, so the loop rule applies there too.

Distinctions that carry marks

SemaphoreMonitor
Isan integer with two operationsa language construct: data plus procedures
Mutual exclusionthe programmer's job, by conventionthe compiler's job, automatically
Wrong order possibleyes, and it deadlocksno
Remembers a missed signalyes, in its countno: a signal with nobody waiting is lost
Waitingwait on the semaphorex.wait() on a condition variable, which releases the monitor
x.signal() on a condition variablesignal(S) on a semaphore
Nobody waitingnothing happens, for everthe count goes up, and the next waiter passes
Effect on the signallerHoare: it waits. Mesa: it continuesnone
munotes.in171

Monitors and Condition Variables

What it does not mean

A monitor is not a mutex with a nicer name. The difference is that the mutual exclusion is enforced by the language over every procedure, so it cannot be omitted.

A condition variable is not a semaphore. It has no value and a signal to nobody is lost. Using one where a counting semaphore belongs loses events.

if is not good enough before wait. Under Mesa semantics, and because of spurious wakeups, the condition must be re-checked in a while loop.

A monitor does not prevent deadlock. Two monitors called in different orders by two processes deadlock exactly as two mutexes do. What it prevents is forgetting the lock and getting the order of waits wrong inside one monitor.

Quick revision

  • A monitor is a language construct: shared variables plus the procedures on them, with

only one process active inside at a time, enforced by the compiler.

  • It exists because a semaphore is only a convention that the programmer can get wrong.
  • Condition variables carry x.wait(), which suspends the caller

and releases the monitor, and x.signal(), which wakes one waiter.

  • A condition variable has no value: a signal with nobody waiting is lost. A semaphore's is

not.

  • Hoare semantics: the signaller gives way, so the woken process finds the condition true.

Mesa semantics: the signaller continues, so the woken process must re-check.

  • Therefore wait goes in a while loop, never an if. Spurious wakeups are the second

reason.

  • Java's synchronized with wait and notify is a monitor, with Mesa semantics.
  • A monitor does not prevent deadlock between two monitors.

Test yourself

  1. What is a monitor? A language construct collecting shared variables and the procedures

that operate on them, in which only one process may be active at a time, with the mutual exclusion enforced by the compiler rather than by the programmer.

  1. Why were monitors introduced when semaphores already existed? Because a semaphore is a

convention: nothing forces a programmer to wait and signal correctly, and one mistake gives a race or a deadlock. A monitor makes the exclusion automatic.

  1. What are the two operations on a condition variable and what does each do? x.wait()

suspends the calling process on x and releases the monitor; x.signal() resumes one process waiting on x, and does nothing at all if none is.

  1. Give two differences between a condition variable and a semaphore. A condition variable

has no value, so a signal with nobody waiting is lost, while a semaphore's count remembers it; and wait on a condition variable releases the monitor, which a semaphore knows nothing about.

munotes.in172

Monitors and Condition Variables

  1. Distinguish Hoare from Mesa semantics. Under Hoare, the signalling process gives the

monitor to the woken one, which therefore finds the condition true. Under Mesa the signaller continues and the woken process joins the queue, so it must re-check the condition.

  1. Why must wait be called inside a while loop? Under Mesa semantics another process may

change the condition between the signal and the woken process resuming, and the standard also permits a wakeup with no signal at all.

  1. Does a monitor prevent deadlock? No. Two monitors entered in different orders by two

processes deadlock just as two mutexes do. It prevents forgetting the lock and misordering the waits within one monitor.

Contents This chapter on its own page

munotes.in173

Chapter Forty-Three

Why Scheduling Exists: The Burst Cycle

Syllabus topic Module 1, "CPU Scheduling - Basic Concepts"

In one line

Every process alternates between computing and waiting, and the scheduler's whole job is to give the processor to somebody else during the waiting.

The cycle

Process execution is a cycle of CPU execution and input and output wait. A process starts with a CPU burst, then an input and output burst, then another CPU burst, and so on, and it ends with a CPU burst that finishes with a request to terminate.

BurstWhat the process is doingWhich state, from Chapter fifteen
CPU burstcomputingrunning, or ready
Input and output burstwaiting for a devicewaiting

That alternation is the entire reason scheduling is worth doing. If a program computed from start to finish and never waited, the best possible schedule would be to run each one to completion, and there would be nothing to decide. Because it waits, the processor is free for somebody else, and a machine that does not use those gaps wastes most of itself.

The shape of the histogram

Measure every CPU burst on a real machine and draw how many bursts there were of each length. The shape is the same on every system ever measured: an enormous number of very short bursts and a small number of long ones. The curve falls away steeply and has a long tail.

Burst lengthHow many
very shortvery many
mediumfew
longvery few

Three consequences follow, and they are the justification for the next twelve chapters.

  1. Most bursts are short, so the time spent choosing must be much shorter still. A scheduler

that took a millisecond to pick would double the cost of a one millisecond burst.

  1. A process's next burst can be guessed from its past ones, which is what makes shortest job

first implementable at all (Chapter forty eight).

  1. A few processes have long bursts, and a scheduler that ignored them would let them

monopolise the machine, which is what round robin's quantum prevents (Chapter fifty two).

Two kinds of process, from the same histogram

Chapter seventeen's vocabulary, now with the reason behind it.

Input and output boundProcessor bound
Its CPU bursts aremany and very shortfew and long
It spends its timewaiting for devicescomputing
Wants from the schedulerto be served quickly when it wakesa long uninterrupted run
Examplean editor, a shell, a web servera compiler, a video encoder

The two want opposite things, which is why no single algorithm is best and why Chapter fifty three's multilevel queue exists: give them different queues and treat them differently.

When a scheduling decision is made

A scheduler is asked to choose at exactly four moments, and the list is a standard question.

munotes.in174

Why Scheduling Exists: The Burst Cycle

MomentWhat happened
1a process switches from running to waiting: it asked for input or output, or for a child to finish
2a process switches from running to ready: an interrupt, usually the timer
3a process switches from waiting to ready: its input or output finished
4a process terminates

Moments 1 and 4 leave the processor with nothing to run, so a choice must be made. Moments 2 and 3 leave the running process able to continue, so a choice is optional. That distinction is the definition of the next chapter's word:

  • a scheduler that chooses only at 1 and 4 is non-preemptive, or co-operative;
  • a scheduler that also chooses at 2 and 3 is preemptive.

The vocabulary to have straight

WordMeaning
CPU burstone stretch of computing between two waits
Short term scheduler, or CPU schedulerthe code that chooses the next ready process
Dispatcherthe code that actually gives it the processor (Chapter forty four)
Dispatch latencythe time the dispatcher takes to stop one process and start another
Preemptiontaking the processor away from a process that could still use it

The scheduler decides and the dispatcher acts. They are named separately because they are measured separately: the scheduler's cost is the choosing, the dispatcher's is the switching.

Worked example: the same work, scheduled two ways

Two processes on one processor. Each alternates: 1 unit of computing, then 3 units of waiting for a disk, five times over.

Run them one after the other. P1 takes 1 + 3 + 1 + 3 + 1 + 3 + 1 + 3 + 1 = 17 units, and then P2 takes another 17. Total 34 units. The processor computed for 10 of those and was idle for 24.

Interleave them. While P1 waits for its disk, P2 computes, and the other way round. Both finish in about 18 units. Total 18 units for the same work, and the processor computed for 10 of the 18.

processor utilisation, one after the other = 10 / 34

processor utilisation, interleaved = 10 / 18

Neither fraction has an exact decimal. The first is about 29 per cent and the second about 56 per cent: the machine did nearly twice as much with the same processor.

Nothing about either program changed. The only difference is that somebody used the gaps, and that is the value the next twelve chapters are about.

What it does not mean

A CPU burst is not a time slice. A burst is how long a process wants the processor before it waits for something. A slice, or quantum, is how long the scheduler allows it. They are unrelated numbers and a burst may take many slices.

munotes.in175

Why Scheduling Exists: The Burst Cycle

The histogram is not a rule about one program. It describes the mixture on a machine. One program may have a single enormous burst.

Scheduling does not make a single program faster. It makes the machine busier. A program alone on an idle machine is not helped by any algorithm in this row.

Quick revision

  • Process execution alternates between a CPU burst and an input and output burst, and it

ends with a CPU burst.

  • The histogram of burst lengths always has very many short bursts and a few long ones.

Therefore: the scheduler must be fast, the next burst can be predicted from the past, and the long ones must be preventable from monopolising the machine.

  • Input and output bound processes have many short bursts; processor bound processes have

few long ones. They want opposite treatment.

  • Four scheduling moments: running to waiting, running to ready, waiting to ready, and

terminating.

  • Choosing only at moments 1 and 4 is non-preemptive; also choosing at 2 and 3 is

preemptive.

  • The scheduler decides, the dispatcher acts, and dispatch latency is what the

dispatcher costs.

Test yourself

  1. Describe the CPU and input and output burst cycle. A process alternates between stretches

of computing, the CPU bursts, and stretches of waiting for devices, the input and output bursts, and ends with a CPU burst.

  1. What shape does the histogram of CPU burst lengths have, and what follows from it? A very

large number of short bursts and a small number of long ones. It follows that the scheduler must be very fast, that a process's next burst can be estimated from its earlier ones, and that long bursts must be preemptable.

  1. Name the four moments at which a scheduling decision may be made. When a process moves

from running to waiting, from running to ready, from waiting to ready, and when it terminates.

  1. Which of those four force a decision, and what does that define? Running to waiting and

terminating leave the processor idle, so a choice must be made; a scheduler that chooses only then is non-preemptive. Choosing at the other two as well makes it preemptive.

  1. Distinguish the scheduler from the dispatcher. The scheduler chooses which ready process

runs next; the dispatcher performs the switch that gives it the processor. The dispatcher's cost is the dispatch latency. 6. Two processes each compute for 1 unit and wait 3 units, five times. Compare running them in turn with interleaving them. In turn takes 34 units with utilisation 10 / 34; interleaved takes about 18 units with utilisation 10 / 18. The same work, with the gaps used.

munotes.in176

Why Scheduling Exists: The Burst Cycle

  1. Is a CPU burst the same as a time slice? No. A burst is how long the process wants the

processor; a slice is how long the scheduler allows before preempting it.

Contents This chapter on its own page

munotes.in177

Chapter Forty-Four

The Dispatcher, Preemption, and What a Switch Costs

Syllabus topic Module 1, "CPU Scheduling - Basic Concepts"; Computer Science Practical 3, Module 1, "Track context switches and improve queue management."

In one line

The dispatcher is the code that actually hands the processor over, and every hand over costs time in which no useful work is done.

What the dispatcher does

The scheduler chooses. The dispatcher acts, and it does exactly three things.

  1. Switch context: save the outgoing process's state into its control block and load the

incoming one's (Chapter sixteen).

  1. Switch to user mode, because the dispatcher itself runs in kernel mode.
  2. Jump to the saved program counter of the incoming process, so it resumes where it was

interrupted.

Dispatch latency is how long those three steps take. It is pure overhead and it happens on every single switch, so it must be as small as the hardware allows.

Preemptive against non-preemptive

Chapter forty three named the four scheduling moments. The definition follows from which of them the scheduler uses.

Non-preemptive, co-operativePreemptive
Chooses atonly moments 1 and 4: a process blocks or endsall four moments
A running processkeeps the processor until it gives it upmay lose it at any moment
Needsnothing speciala timer (Chapter four)
A long processholds the machineis interrupted
Response to a keystrokeafter the current process blockswithin one quantum
Shared datasafer: no switch in the middle of an updateneeds the whole of the synchronisation row

The last row is the cost of preemption that students forget. A preemptive kernel can be interrupted in the middle of updating its own data structures, so the kernel itself needs locks. That is why early Unix kernels were non-preemptive inside the kernel even while preempting user processes: the kernel ran a system call to completion, which made kernel data safe for free.

Preemption also affects processes that co-operate. If P1 is preempted while updating shared data and P2 reads it, P2 sees a half finished update: exactly Chapter thirty three.

Counting the switches, which the practical asks for

The kernel counts every switch for every process and publishes the two kinds separately.

CountMeaningCaused by
voluntarythe process gave the processor upit blocked: a read, a lock, a sleep
involuntary, or nonvoluntaryit was taken awaythe timer, or a more urgent process became ready

A program can ask for its own counts.

#define _GNU_SOURCE
#include <stdio.h>
#include <time.h>
#include <sys/resource.h>

static void pause_for(long nanoseconds)
{
    struct timespec t = {0, nanoseconds};

    nanosleep(&t, NULL);
}

int main(void)
{
    struct rusage before, after;

    getrusage(RUSAGE_SELF, &before);
    for (int i = 0; i < 100; i++) {
        pause_for(1000000);              /* block for a millisecond, a hundred times */
    }
    getrusage(RUSAGE_SELF, &after);

    long gave_up = after.ru_nvcsw - before.ru_nvcsw;
    long taken = after.ru_nivcsw - before.ru_nivcsw;

    printf("a hundred sleeps: gave the processor up %ld times, had it taken %ld times\n",
        gave_up, taken);
    printf("%s\n", gave_up >= 100 ? "at least one voluntary switch per sleep, as expected"
        : "FEWER than one per sleep, which cannot be right");

    getrusage(RUSAGE_SELF, &before);
    volatile long spin = 0;
    for (long i = 0; i < 100000000L; i++) {
        spin += i & 1;                   /* compute, and never block */
    }
    getrusage(RUSAGE_SELF, &after);
    printf("then a long computation: gave up %ld, taken %ld\n",
        after.ru_nvcsw - before.ru_nvcsw, after.ru_nivcsw - before.ru_nivcsw);
    printf("%s\n", after.ru_nivcsw - before.ru_nivcsw > 0
        ? "it was preempted, which is the timer doing its job"
        : "it was never preempted");
    return 0;
}
munotes.in178

The Dispatcher, Preemption, and What a Switch Costs

$ gcc -std=c17 -Wall -Wextra -o switches switches.c
$ ./switches
a hundred sleeps: gave the processor up 100 times, had it taken 0 times
at least one voluntary switch per sleep, as expected
then a long computation: gave up 0, taken 44
it was preempted, which is the timer doing its job

Two behaviours, cleanly separated. A hundred sleeps produced a hundred voluntary switches and no preemption: the process was never taken off the processor because it kept giving it up. The computation produced no voluntary switches and dozens of involuntary ones: it never gave the processor up and the timer took it away, over and over.

That is the practical's answer in two numbers, and it is the diagnosis Chapter seventeen promised: a large voluntary count means input and output bound, a large involuntary count means processor bound.

What a switch costs, and why the quantum follows from it

Suppose a switch costs s and the quantum is q. Of every q + s units the machine spends s doing nothing useful.

overhead = s / (q + s)

At s = 5 microseconds:

QuantumOverhead, as a fractionRoughly
100 microseconds5 / 1054.8 per cent
1 millisecond5 / 1005half a per cent
10 milliseconds5 / 10005a twentieth of a per cent
100 milliseconds5 / 100005a five hundredth of a per cent

A long quantum wastes almost nothing and responds slowly; a short quantum responds quickly and wastes a great deal. That is the whole of Chapter fifty two's choice, and it is arithmetic.

The real cost of a switch is usually not the registers. It is the memory management information of Chapter sixteen step 6, and the caches: the incoming process's data is not in the cache, so its first thousand memory references are slow. That hidden cost is called the cache footprint and it is why switching between two threads of one process is cheaper than between two processes.

Distinctions that carry marks

SchedulerDispatcher
Doeschooses which ready process runs nextperforms the switch
Cost is calledscheduling overheaddispatch latency
Runsat the four scheduling momentson every switch
munotes.in179

The Dispatcher, Preemption, and What a Switch Costs

Voluntary switchInvoluntary switch
Causethe process blockedthe timer or a more urgent process
Means the process isinput and output boundprocessor bound
Chapter fifteen transitionrunning to waitingrunning to ready

What it does not mean

Dispatch latency is not the quantum. The quantum is how long a process runs; the latency is how long the hand over takes.

A preemptive scheduler does not preempt constantly. Most timer interrupts return to the same process.

Zero involuntary switches is not good news. It means either the process always blocks first, or nothing else wanted the processor.

Quick revision

  • The dispatcher switches context, switches to user mode, and jumps to the saved program

counter. Dispatch latency is its cost, and it is pure overhead.

  • Non-preemptive: the scheduler chooses only when a process blocks or ends. Preemptive:

at all four moments, which needs a timer.

  • Preemption's hidden cost is that the kernel's own data and co-operating processes' data can be

interrupted mid update, so both need locks.

  • Voluntary switches mean the process blocked; involuntary mean it was preempted.

Measured here: a hundred sleeps gave 100 voluntary and 0 involuntary; a long computation gave 0 voluntary and dozens of involuntary.

  • Overhead is s / (q + s). At a five microsecond switch, a 100 microsecond quantum wastes about

4.8 per cent and a 10 millisecond quantum about 0.05 per cent.

  • The real cost of a switch is the memory mapping and the cold caches, not the registers.

Test yourself

  1. What three things does the dispatcher do? Switch context, switch the processor to user

mode, and jump to the location saved in the incoming process's program counter.

  1. Define dispatch latency. The time the dispatcher takes to stop one process running and

start another.

  1. Distinguish preemptive from non-preemptive scheduling. Non-preemptive schedules only when

a process blocks or terminates, so a running process keeps the processor until it gives it up. Preemptive also schedules when an interrupt arrives or a process becomes ready, so a running process can lose the processor at any moment.

  1. Give one cost of preemption besides the switch itself. The kernel and any co-operating

processes may be interrupted in the middle of updating shared data, so both need locks; without preemption a system call ran to completion and kernel data was safe for nothing.

  1. A process shows 40,000 voluntary and 3 involuntary switches. What kind of process is it?

Input and output bound: it gives the processor up itself, over and over, rather than being preempted. 6. A switch costs 5 microseconds. Compare the overhead at a 100 microsecond and a 10 millisecond quantum. the fraction is 5 / 105 at the short quantum, about 4.8 per cent, and 5 / 10005 at the long one, which is a twentieth of a per cent.

munotes.in180

The Dispatcher, Preemption, and What a Switch Costs

  1. Why is switching between threads of one process cheaper than between processes? The

address space does not change, so the memory management hardware and the caches are not invalidated.

Contents This chapter on its own page

munotes.in181

Chapter Forty-Five

The Five Criteria, and the Arithmetic of Each

Syllabus topic Module 1, "CPU Scheduling - Scheduling Criteria"

In one line

A schedule is judged by five measures: three you want as large as possible and two you want as small as possible.

The five

CriterionDefinitionWant it
CPU utilisationthe fraction of the time the processor is busyas high as possible
Throughputthe number of processes completed per unit of timeas high as possible
Turnaround timecompletion time minus arrival timeas low as possible
Waiting timeturnaround time minus burst timeas low as possible
Response timethe first moment on the processor, minus arrival timeas low as possible

Turnaround, waiting and response are the three that are examined in sums, and the three formulas are the whole of the arithmetic:

turnaround = completion - arrival

waiting = turnaround - burst

response = first on the processor - arrival

The second formula is the one to hold on to. Waiting time is turnaround minus burst, so it is turnaround with the process's own work taken out: the time it spent doing nothing. It is never negative, and a negative answer means a mistake in the completion times.

What each one is for

  • Utilisation matters to whoever paid for the machine. A server at 15 per cent has been over

bought.

  • Throughput matters to a batch system: how many jobs got through today.
  • Turnaround matters to whoever submitted one job: how long from handing it in to getting the

answer.

  • Waiting matters because it is the only one the scheduler can actually change. It cannot

make a process's burst shorter or a disk faster; it can only change how long the process spends in the ready queue.

  • Response matters to whoever is sitting at the screen. A word processor that finishes a

keystroke in half a second feels broken however good its turnaround is.

Response time and turnaround time can pull in opposite directions, and a question may ask for a case. Round robin with a very short quantum gives excellent response and poor turnaround, because every process is started early and finished late. That is Chapter fifty two.

Averages, and the honest way to write one

An examination wants the average, over all the processes, of turnaround and of waiting. Write it as the division, not only the answer.

A division must be shown, and the number of decimals must be honest. 17 divided by 3 is not 5.6 and it is not 5.66; if two places are printed it is 5.67. The check-sums.py gate in this book proves both sides of every such line: that the total is the sum of the column above it, and that the quotient is right to the number of places printed.

munotes.in182

The Five Criteria, and the Arithmetic of Each

One problem, all five, worked

Schedule: first come first served

ProcessArrivalBurst
P107
P224
P341
P454
SliceProcessFromTo
1P107
2P2711
3P31112
4P41216
ProcessCompletionTurnaroundWaiting
P1770
P21195
P31287
P416117

Average turnaround time: 35 / 4 = 8.75

Average waiting time: 19 / 4 = 4.75

Context switches: 3

Now the other three criteria on the same schedule.

Response time. Every process here runs to completion the first time it is chosen, so its first moment on the processor is the same as the start of its only slice.

ProcessFirst on the processorArrivalResponse
P1000
P2725
P31147
P41257

average response = 19 / 4 = 4.75

Response equals waiting here, for every process, and that is not a coincidence. Under a non-preemptive algorithm a process runs once, without interruption, so the only time it waits is before it starts. Under round robin the two differ, because a process starts early and then waits again between its slices. If a question gives you a non-preemptive schedule and asks for both, they are the same number.

Throughput. Four processes finished in the sixteen units from 0 to 16.

throughput = 4 / 16 = 0.25

Or one process every four units.

Utilisation. The Gantt chart has no idle stretch: the processor worked from 0 to 16 without a gap.

utilisation = 16 / 16 = 1

Utilisation is 1 here only because P1 arrived at time 0 and there was always something ready afterwards. Change P1's arrival to 3 and the processor is idle for the first three units, so utilisation is 13 / 16.

Which criterion to optimise

No algorithm optimises all five, and a question that asks which algorithm is best is asking about the workload.

The machine isOptimiseBecause
a batch systemthroughput and turnaroundnobody is waiting at a screen
an interactive desktopresponse timea person notices a tenth of a second
a real time controllerthe worst case, not the averagea late answer is a wrong answer
a shared serverfairness, and the variance of response timepredictable is worth more than fast on average

The last row names a sixth measure an examination sometimes wants: the variance. A system whose response time is always 1 second is better to use than one that averages half a second and occasionally takes five, and an average alone cannot tell them apart.

What it does not mean

Waiting time is not the time spent waiting for input or output. It is time in the ready queue, wanting the processor and not having it. A process blocked on a disk is not counted as waiting in this arithmetic, which is why waiting time can be computed from the Gantt chart alone.

munotes.in183

The Five Criteria, and the Arithmetic of Each

Turnaround time is not the burst time. It includes every moment from arrival to completion, waiting and computing together.

High utilisation is not the goal by itself. A machine at 100 per cent utilisation with a twenty second response time is a bad machine.

Quick revision

  • Five criteria: CPU utilisation and throughput, to be maximised; turnaround time,

waiting time and response time, to be minimised.

  • turnaround = completion - arrival. waiting = turnaround - burst. response = first on the

processor - arrival.

  • Waiting time can never be negative; a negative answer means the completion times are wrong.
  • Under a non-preemptive algorithm, response time equals waiting time for every process.
  • Show the average as a division, and round honestly: 17 / 3 to two places is 5.67.
  • Throughput is processes finished divided by the elapsed time; utilisation is busy time divided

by elapsed time.

  • No algorithm is best at all five. A batch system wants throughput, a desktop wants response, a

real time system wants the worst case, and a shared server wants low variance.

Test yourself

  1. Name the five scheduling criteria and say which are maximised. CPU utilisation and

throughput are maximised; turnaround time, waiting time and response time are minimised.

  1. Give the three formulas. Turnaround is completion minus arrival; waiting is turnaround

minus burst; response is the first moment on the processor minus arrival.

  1. Why can waiting time never be negative? It is the part of the turnaround that the process

did not spend computing, and a process cannot finish before it has run for its whole burst.

  1. When are response time and waiting time the same for every process? Under a non-preemptive

algorithm, because each process runs once without interruption, so the only waiting is before it starts.

  1. Which criterion can the scheduler actually change, and why only that one? Waiting time. It

cannot shorten a burst or speed up a device; it can only change how long a process sits in the ready queue.

  1. Four processes finish by time 16 with no idle processor. Give throughput and utilisation.

Throughput is 4 / 16 = 0.25 processes per unit; utilisation is 16 / 16 = 1.

  1. Why is a low variance of response time worth having? A system that always answers in one

second is more usable than one that averages half a second and sometimes takes five, and an average cannot tell them apart.

Contents This chapter on its own page

munotes.in184

Chapter Forty-Six

Reading and Drawing a Gantt Chart

Syllabus topic Computer Science Practical 3, Module 1, "Analyze waiting time, turnaround time, and Gantt chart generation."

In one line

A Gantt chart is a picture of who had the processor when, and every answer in the next eight chapters is read off one.

What it is

A horizontal line for time, divided into blocks. Each block is labelled with the process that held the processor during it, and the boundaries between blocks are marked with the time.

On paper it looks like this, for three processes that arrive together with bursts of 24, 3 and 3:

| P1 | P2 | P3 |

0 24 27 30

Everything in this row of the syllabus comes out of that picture, so drawing it correctly is the first step of every answer and the step where marks are lost.

How this book sets one

A drawing does not survive being read on a phone, and a chart of twenty slices does not fit across a page. So in this book a Gantt chart is a table with one row per slice: the slice number, the process, and the two times.

SliceProcessFromTo
1P1024
2P22427
3P32730

That table holds exactly the same information as the drawing and has two advantages: it reads at any width, and it can be checked. Every one of these tables in this book is re-derived from the problem beside it by check-sums.py.

Draw the picture in an examination. A table is right and a chart is what the paper asks for. Learn to produce both from the same reasoning.

Reading the three times off it

Once the chart is drawn, no thinking is left: the three answers are lookups and two subtractions.

WantRead from the chart
Completion time of Pthe To of P's last slice
Turnaround of Pcompletion minus P's arrival
Waiting of Pturnaround minus P's burst
Response of Pthe From of P's first slice, minus P's arrival

Waiting time can also be read directly, as the total of the gaps in which the process was ready and not running, and it is worth doing both ways once: the two must agree, and if they do not, the chart is wrong.

Two mistakes that lose marks

1. Forgetting the idle stretches. If no process has arrived yet, the processor is idle and the chart must show it. A chart that runs P2 at time 0 when P2 arrives at time 3 is wrong, and every figure derived from it is wrong.

SliceProcessFromTo
1idle03
2P138

2. Treating the chart as continuous when the algorithm is preemptive. Under round robin or shortest remaining time first a process appears several times, and its completion time is the end of its last appearance, not its first. Adding up the wrong end is the commonest arithmetic mistake in this whole row.

munotes.in185

Reading and Drawing a Gantt Chart

SliceProcessFromTo
1P104
2P247
3P1727

P1's completion time is 27, not 4.

Worked example: one chart, every answer

Schedule: round robin, quantum 4

ProcessArrivalBurst
P1024
P203
P303
SliceProcessFromTo
1P104
2P247
3P3710
4P11030
ProcessCompletionTurnaroundWaiting
P130306
P2774
P310107

Average turnaround time: 47 / 3, about 15.67

Average waiting time: 17 / 3, about 5.67

Context switches: 3

Now read the fourth criterion off the same chart. Response time is the From of each process's first slice minus its arrival.

ProcessFirst slice beginsArrivalResponse
P1000
P2404
P3707

Compare P1's two numbers: waiting 6 and response 0. P1 was on the processor from the very first instant, so its response time is nothing, and it still waited six units, in the gap from 4 to 10 while the other two had their turn. Under a preemptive algorithm the two are different numbers and Chapter forty five's rule about them being equal does not apply. A question that asks for both is checking exactly this.

P1's waiting time can be checked the other way as promised: P1 was ready and not running from 4 to 10, which is 6 units, and nowhere else.

waiting of P1 = 10 - 4 = 6

turnaround - burst = 30 - 24 = 6

The two agree, so the chart is right.

Drawing one, step by step

The method, which works for every algorithm in this row.

  1. List the arrival times and start the clock at the earliest one, or at 0 with an idle block

if the earliest is later.

  1. At each decision point, work out which processes have arrived and not finished. That set

is the ready queue.

  1. Apply the algorithm's rule to choose from that set.
  2. Draw the block, ending it where the algorithm says: at the end of the burst for a

non-preemptive algorithm, at the end of the quantum for round robin, or at the next arrival for a preemptive algorithm that might change its mind.

  1. Advance the clock to the end of the block and go back to step 2.
  2. When every process has finished, write the times under the boundaries.

Step 4 is where the algorithms differ and it is the only place they differ. Everything else on the list is the same for all eight of them.

munotes.in186

Reading and Drawing a Gantt Chart

What it does not mean

A Gantt chart is not a state diagram. It shows only who was running. A process that is ready or waiting does not appear, which is why waiting time has to be worked out rather than read directly.

A block is not a process. A preemptive algorithm gives one process several blocks.

The numbers under the chart are not durations. They are instants: the times at which one block ends and the next begins.

Quick revision

  • A Gantt chart shows which process held the processor during each interval. In this book it

is a table of Slice, Process, From, To; in an examination, draw the bar.

  • Completion is the To of the last slice. Turnaround is completion minus arrival. Waiting is

turnaround minus burst. Response is the From of the first slice minus arrival.

  • Show idle blocks. Under a preemptive algorithm a process has several blocks and its

completion is the end of the last one.

  • Under a preemptive algorithm response and waiting are different; under a non-preemptive one

they are equal.

  • Check a chart by computing one waiting time both ways: turnaround minus burst, and the total of

the gaps.

  • The six step method is the same for every algorithm except step 4, where the block ends.

Test yourself

  1. What does a Gantt chart show? Which process was running during each interval of time, with

the instants at which the intervals begin and end.

  1. How do you read a process's completion time off the chart? It is the end of that process's

last block, not its first.

  1. How do you get waiting time from the chart? Take the completion time, subtract the arrival

time to get turnaround, then subtract the burst. Or add up the intervals in which the process was ready and not running; the two must agree.

  1. A process arrives at 3 and the chart starts at 0. What must the chart show? An idle block

from 0 to 3. Omitting it makes every figure wrong.

  1. A process appears in three blocks. Which one gives its completion time? The last.

6. P1 has response 0 and waiting 6. Which kind of algorithm is this, and why are they different? A preemptive one. P1 started immediately, so its response is 0, but it was taken off the processor and waited six units in the middle.

  1. At which step of the drawing method do the eight algorithms differ? Only at the step where

the block ends: at the end of the burst, at the end of the quantum, or at the next arrival.

Contents This chapter on its own page

munotes.in187

Chapter Forty-Seven

First Come First Served

Syllabus topic Module 1, "CPU Scheduling - Scheduling Algorithms (FCFS ...)"

In one line

The processor goes to whichever ready process asked for it first, and keeps it until that process gives it up.

The rule

The ready queue is a first in, first out queue. A process joining goes to the tail; the scheduler takes the head. It is non-preemptive: once a process has the processor it keeps it until it blocks or finishes.

Choosesthe process that has been in the ready queue longest
Preemptiveno
Needs to knownothing: no burst lengths, no priorities
Data structureone queue
Starvationimpossible: every process reaches the head

It is the only algorithm in this row that needs no information about the processes at all, which is why it is the baseline the others are compared with.

Worked in full

Schedule: first come first served

ProcessArrivalBurst
P1024
P213
P323
SliceProcessFromTo
1P1024
2P22427
3P32730
ProcessCompletionTurnaroundWaiting
P124240
P2272623
P3302825

Average turnaround time: 78 / 3 = 26

Average waiting time: 48 / 3 = 16

Context switches: 2

Two short processes waited more than twenty units each for three units of work.

The convoy effect

Now change nothing except the order in which they arrive. The same three processes with the same bursts, the long one arriving last.

Schedule: first come first served

ProcessArrivalBurst
P203
P313
P1224
SliceProcessFromTo
1P203
2P336
3P1630
ProcessCompletionTurnaroundWaiting
P2330
P3652
P130284

Average turnaround time: 36 / 3 = 12

Average waiting time: 6 / 3 = 2

Context switches: 2

Average waiting fell from 16 to 2, eight times better, and not one burst changed. The work took thirty units either way; the processor was busy for all thirty either way. All that changed is who waited.

That is the convoy effect, and the name is the explanation: one long process at the front of the queue holds up a convoy of short ones behind it, exactly as one slow lorry holds up the cars on a single track road. The words to use in an answer are that the average waiting time under first come first served varies greatly with the arrival order, and is very poor when a long process arrives first.

The other half of the convoy effect

The version above is about arrival order. The classic statement is about the mixture of Chapter forty three.

One processor bound process and many input and output bound ones. The processor bound process gets the processor and holds it for a long burst. The input and output bound processes finish their tiny bursts, queue for the disk, and are soon all waiting on devices while the processor bound one computes. Then it releases the processor and goes to the disk, and all the little processes now rush through the processor and pile up behind it at the disk. The devices are idle while the processor is busy and the processor is idle while the devices are busy, and the machine achieves far less than it could.

munotes.in188

First Come First Served

What it is good for

It is not useless, and a question may ask when it is right.

Right forBecause
a batch system where turnaround per job does not matterthroughput is the same whatever the order
the lowest level queue of a feedback scheduler (Chapter fifty four)processes there are long and nobody is waiting on them
any case where predictability matters more than speedit is completely predictable
a system with no way to estimate burst lengthsit needs none

Distinctions that carry marks

First come first servedShortest job first
Choosesthe oldest requestthe smallest burst
Needs to know the burstnoyes
Average waiting timepoor, and depends on the orderthe best possible
Starvationimpossiblepossible
Convoy effectyesno

What it does not mean

First come first served is not fair in the useful sense. Everybody is served in order, and a short job behind a long one waits far longer than its own work takes. Fair ordering and fair waiting are different things.

It does not waste the processor. Utilisation is as good as any algorithm's; it is the waiting times that are bad.

The convoy effect is not caused by having too few processors. It is caused by not being allowed to interrupt.

Quick revision

  • First come first served gives the processor to the oldest request and is

non-preemptive.

  • It needs no information about the processes and cannot starve anybody.
  • The convoy effect: one long process at the head of the queue delays every short process

behind it. Measured here, the same three processes gave an average waiting time of 16 in one order and 2 in another.

  • In its classic form, one processor bound process and many input and output bound ones

alternately idle the devices and the processor.

  • It is right where predictability matters, where nothing is known about burst lengths, and at

the bottom of a multilevel feedback queue.

Test yourself

  1. State the first come first served rule. The processor is given to the process that

requested it first, and is not taken away until that process blocks or finishes.

munotes.in189

First Come First Served

  1. What does it need to know about the processes? Nothing at all.
  2. Can it starve a process? No. Every process reaches the head of a first in, first out

queue.

  1. Explain the convoy effect. A long process at the front of the ready queue makes every

short process behind it wait for the whole of its burst, so the average waiting time is very poor and depends heavily on the arrival order. 5. Three processes with bursts 24, 3 and 3 arrive in that order, then in the reverse order. Compare the average waiting times. 48 / 3 = 16 in the first order and 6 / 3 = 2 in the second: eight times better, with no burst changed. 6. Describe the convoy effect with one processor bound process and several input and output bound ones. The processor bound process holds the processor while the others finish their short bursts and queue at the devices; when it releases the processor they rush through and pile up at the device again, so the processor and the devices take turns at being idle.

  1. Name a place where first come first served is the right choice. The lowest priority queue

of a multilevel feedback scheduler, or any batch system where per job turnaround does not matter and nothing is known about burst lengths.

Contents This chapter on its own page

munotes.in190

Chapter Forty-Eight

Shortest Job First

Syllabus topic Module 1, "CPU Scheduling - Scheduling Algorithms (... SJF ...)"

In one line

Of the processes waiting, run the one whose next burst is shortest.

The rule

Choosesthe ready process with the smallest next CPU burst
Preemptiveno, in this form: see Chapter forty nine
Tiesbroken by first come first served
Needs to knowthe length of each process's next burst
Starvationpossible: a long process can be overtaken for ever

The proper name is shortest NEXT CPU BURST first, and the short name misleads: it is not about the total length of the job but about its next burst. A process with a long total and a short next burst goes first.

Worked in full

Schedule: shortest job first

ProcessArrivalBurst
P106
P208
P307
P403
SliceProcessFromTo
1P403
2P139
3P3916
4P21624
ProcessCompletionTurnaroundWaiting
P1993
P2242416
P316169
P4330

Average turnaround time: 52 / 4 = 13

Average waiting time: 28 / 4 = 7

Context switches: 3

Compare first come first served on the same four processes, which would run them in the order P1, P2, P3, P4 and give an average waiting time of (0 + 6 + 14 + 21) / 4 = 41 / 4 = 10.25. Shortest job first gives 7.

Why it is optimal, in one paragraph

Shortest job first gives the minimum possible average waiting time for a given set of processes, and the reason is worth understanding rather than memorising.

Every process that is not running is waiting, so each unit of a running process's burst adds one unit of waiting to every process still queued behind it. Running a short process first therefore adds its small burst to many waiting totals, while running a long one first adds its large burst to the same number. Putting the shortest first at every step minimises the sum, and any swap of two neighbours in which the longer comes first makes the total worse.

"Optimal" means optimal for average waiting time and nothing else. It is not optimal for response time, not for the worst case, and not for fairness.

Why it cannot be implemented

The next burst length is not known. The scheduler is asked to choose now, and nothing in the process tells it how long it will compute before its next read. For a batch system a user's estimate can be asked for, and users understate. For short term scheduling there is no estimate at all.

So the length is predicted from the process's own history, by an exponential average:

munotes.in191

Shortest Job First

next prediction = alpha last actual burst + (1 - alpha) last prediction

SymbolMeaning
the last actual bursthow long the burst that just ended really was
the last predictionwhat was predicted for it
alphabetween 0 and 1, how much to trust the most recent burst

The two extremes explain the formula. alpha = 0: the prediction never changes, so only the initial guess counts and history is ignored. alpha = 1: the prediction is the last burst, so only the most recent burst counts and the older history is ignored. alpha = 0.5, the usual choice, weights each older burst half as much as the one after it.

The prediction worked as a sum

A process's real bursts are 6, 4, 6, 4, 13, 13, 13. Start with a prediction of 10 and alpha = 0.5.

Burst numberActualPrediction for itPrediction for the next
16108
2486
3666
4465
51359
613911
7131112

Check two rows:

after burst 1: 0.5 6 + 0.5 10 = 3 + 5 = 8

after burst 5: 0.5 13 + 0.5 5 = 6.5 + 2.5 = 9

Read the last three rows. The process changed behaviour at burst 5, from about 5 units to 13, and the prediction chased it: 5, then 9, then 11, then 12. It is always behind, and it gets closer each time. That is what an exponential average does and what it cannot do: it follows a change smoothly and never predicts one.

Starvation

A long process is passed over every time a shorter one arrives. If short processes keep arriving, it waits for ever.

There is no fix inside the algorithm. The fix is ageing, which is Chapter fifty's subject: raise a waiting process's standing as it waits.

Distinctions that carry marks

First come first servedShortest job first
Average waiting timepoorprovably the minimum
Information needednonethe next burst length
Implementable for short term schedulingyesno, only by prediction
Starvationimpossiblepossible
alpha near 0alpha near 1
Truststhe initial guess and the distant pastthe most recent burst
Reacts to a changeslowlyimmediately
Affected by one unusual bursthardlygreatly

What it does not mean

It is not shortest total job first. It is shortest next burst first, and a process with a long total may still go first.

It is not usable as written. Every real use is the prediction, or a batch system with user estimates.

Optimal is not best. It minimises average waiting time and it can starve a process, which no usable scheduler may do.

munotes.in192

Shortest Job First

Quick revision

  • Shortest job first, properly shortest next CPU burst first, runs the ready process with

the smallest next burst. Non-preemptive in this form, ties by arrival.

  • It gives the provably minimum average waiting time for a set of processes, because each

unit of a running burst is added to the waiting time of everybody queued behind it.

  • It cannot be implemented for short term scheduling: the next burst is unknown.
  • The next burst is predicted by an exponential average: next = alpha times the last actual

burst plus (1 minus alpha) times the last prediction.

  • alpha = 0 ignores history; alpha = 1 ignores everything but the last burst;

alpha = 0.5 halves the weight of each older burst.

  • A prediction follows a change in behaviour and always lags behind it.
  • It can starve a long process. The remedy is ageing.

Test yourself

  1. State the rule, with its proper name. Shortest next CPU burst first: run the ready process

whose next CPU burst is shortest, with ties broken by arrival order.

  1. Why is it optimal, and optimal for what? For average waiting time. Every unit of the

running process's burst is added to the waiting time of each process queued behind it, so putting the shortest first at every step minimises the total.

  1. Why can it not be implemented for short term scheduling? The length of a process's next

burst is not known when the choice has to be made.

  1. Write the exponential average and say what alpha does. The next prediction is alpha times

the last actual burst plus one minus alpha times the last prediction. A large alpha follows recent behaviour closely; a small alpha smooths it and reacts slowly.

  1. Predictions were 10 then 8 after a burst of 6, with alpha 0.5. Check the arithmetic. 0.5

times 6 plus 0.5 times 10 equals 3 plus 5 equals 8.

  1. A process's bursts jump from about 5 to 13. What does the prediction do? It rises towards

13 over several bursts, always behind: 9, then 11, then 12. An exponential average follows a change and cannot anticipate one.

  1. What can go wrong with shortest job first and what is the remedy? A long process can be

overtaken indefinitely by shorter arrivals, which is starvation. The remedy is ageing.

Contents This chapter on its own page

munotes.in193

Chapter Forty-Nine

Shortest Remaining Time First

Syllabus topic Module 1, "CPU Scheduling - Scheduling Algorithms (... SRTF ...)"

In one line

Whenever a process arrives whose burst is shorter than what the running process has left, the running process is thrown off and the new one runs.

The rule

Choosesthe ready process with the smallest remaining time
Preemptiveyes
Decision pointsevery arrival, and every completion
Tiesthe running process keeps the processor, so that no switch happens for nothing
Also calledpreemptive shortest job first
Starvationpossible, and worse than in Chapter forty eight

The word is REMAINING, and that is the whole difference from Chapter forty eight. A process that has already run for 6 of its 8 units has 2 left, and it is compared on 2, not on 8. Comparing on the original burst is the commonest error in a hand worked answer.

Worked in full

Schedule: shortest remaining time first

ProcessArrivalBurst
P108
P214
P329
P435
SliceProcessFromTo
1P101
2P215
3P4510
4P11017
5P31726
ProcessCompletionTurnaroundWaiting
P117179
P2540
P3262415
P41072

Average turnaround time: 52 / 4 = 13

Average waiting time: 26 / 4 = 6.50

Context switches: 4

Non-preemptive shortest job first on the same problem gives an average waiting time of 7.75, so preemption bought 1.25 units per process.

The decision at every instant, written out

This is the part to be able to reproduce, because it is what the marks are for.

TimeWhat happenedReady, with time remainingChosenWhy
0P1 arrivesP1 has 8P1it is the only one
1P2 arrives with 4P1 has 7, P2 has 4P24 is less than 7, so P1 is preempted
2P3 arrives with 9P2 has 3, P1 has 7, P3 has 9P23 is still the smallest
3P4 arrives with 5P2 has 2, P1 has 7, P3 has 9, P4 has 5P22 is still the smallest
5P2 finishesP1 has 7, P3 has 9, P4 has 5P45 is the smallest
10P4 finishesP1 has 7, P3 has 9P17 is less than 9
17P1 finishesP3 has 9P3it is the only one left

Only two things ever cause a decision: an arrival, or a completion. Nothing happens between them, so a chart of twenty units may have only five decision points, and looking only at those is what makes the method quick.

At time 2 and time 3 the arriving process did not preempt, because the running process had less left than the newcomer's whole burst. A student who preempts at every arrival gets a different and wrong chart.

munotes.in194

Shortest Remaining Time First

Where the extra context switches come from

Compare the two shortest job algorithms on this problem.

Shortest job firstShortest remaining time first
Average waiting7.756.50
Average turnaround14.2513
Context switches34

Preemption bought better waiting times and cost one more switch. On a real machine that trade has to be priced with Chapter forty four's arithmetic: four switches at five microseconds is twenty microseconds, against a saving of 1.25 time units per process, and whether that is worth it depends entirely on what a time unit is.

Starvation, and why it is worse here

Chapter forty eight could starve a long process only when shorter ones kept arriving while the processor was free. Here a long process can be thrown off while it is running, over and over, and it makes no progress at all between preemptions.

P3 in the worked example waited fifteen units to do nine units of work, and it was the longest process. With a steady stream of short arrivals it would never have run.

Distinctions that carry marks

Shortest job firstShortest remaining time first
Preemptivenoyes
Comparesthe whole next burstthe remaining time
Decides atcompletions onlyarrivals and completions
Average waitinggoodbetter, and optimal among preemptive algorithms
Switchesfewermore
Starvationpossiblepossible, and worse

What it does not mean

It is not a different algorithm from shortest job first. It is the same rule applied at more moments, which is why the pair is often written as the non-preemptive and preemptive forms of one algorithm.

Preemption does not happen at every arrival. Only when the arriving process's burst is shorter than what the running one has left.

The remaining time is not recomputed by the process. The scheduler knows it: the burst it was predicted to need, minus what it has had.

Quick revision

  • Shortest remaining time first is preemptive shortest job first: at every arrival and every

completion, run the ready process with the least remaining time.

  • Compare on remaining time, not the original burst.
  • A decision is needed only at an arrival or a completion.
  • An arriving process preempts only if its burst is less than the running process's remaining

time; a tie leaves the running process alone.

  • On the worked problem it gave an average waiting time of 6.50 against 7.75 for the

non-preemptive form, at the cost of one more context switch.

  • Starvation is worse than in the non-preemptive form, because a long process can be thrown off

while running.

Test yourself

  1. State the rule. At every arrival and every completion, give the processor to the ready
munotes.in195

Shortest Remaining Time First

process with the smallest remaining execution time, preempting the running process if necessary.

  1. What is compared, and what is the common mistake? The remaining time. The mistake is to

compare the original burst lengths and so to preempt or not preempt wrongly.

  1. When does an arriving process preempt the running one? Only when its burst is shorter than

the time the running process has left.

  1. How many decision points does a schedule have? One per arrival and one per completion, and

no others. 5. P1 arrives at 0 with 8, P2 at 1 with 4, P3 at 2 with 9, P4 at 3 with 5. Give the order of the slices. P1 from 0 to 1, P2 from 1 to 5, P4 from 5 to 10, P1 from 10 to 17, P3 from 17 to 26.

  1. Why is starvation worse here than under the non-preemptive form? A long process can be

taken off the processor part way through, over and over, so it makes no progress at all rather than merely waiting to start.

  1. What does preemption cost? More context switches: four rather than three on the worked

problem, each one pure overhead.

Contents This chapter on its own page

munotes.in196

Chapter Fifty

Priority Scheduling, Starvation and Ageing

Syllabus topic Module 1, "CPU Scheduling - Scheduling Algorithms (... Priority ...)"

In one line

Every process carries a number saying how important it is, and the most important ready process runs.

The convention, stated

In this book, and in MU's text book, a SMALLER number means a HIGHER priority. So priority 1 beats priority 5.

There is no universal rule. Linux's nice value runs the other way, where a larger number means a lower priority, and some textbooks number upwards. An examination answer should say which convention it is using in one line, and then it cannot be marked wrong for the other.

The rule

Choosesthe ready process with the highest priority, which is the smallest number here
Tiesbroken by first come first served
Preemptive forman arriving process of higher priority throws the running one off
Non-preemptive formthe arriving process waits at the head of the queue
Starvationpossible, and it is the defining problem of this algorithm

Shortest job first is a special case of priority scheduling, where the priority is the inverse of the predicted next burst. That is worth knowing because it explains why the two share the same fault.

Where a priority comes from

KindSet fromExample
Internalsomething the system can measurememory used, open files, the burst ratio, time limits
Externalsomething outside the systemwho is paying, which department, how urgent the work is

Non-preemptive, worked in full

Schedule: priority

ProcessArrivalBurstPriority
P10103
P2011
P3024
P4015
P5052
SliceProcessFromTo
1P201
2P516
3P1616
4P31618
5P41819
ProcessCompletionTurnaroundWaiting
P116166
P2110
P3181816
P4191918
P5661

Average turnaround time: 60 / 5 = 12

Average waiting time: 41 / 5 = 8.20

Context switches: 4

The five processes ran in priority order, 1, 2, 3, 4, 5, and P4 with the worst priority waited eighteen units for one unit of work.

Preemptive, worked in full

The same five processes and priorities, now arriving one unit apart so that preemption has something to do.

Schedule: priority (preemptive)

ProcessArrivalBurstPriority
P10103
P2111
P3224
P4315
P5452
SliceProcessFromTo
1P101
2P212
3P124
4P549
5P1916
6P31618
7P41819
ProcessCompletionTurnaroundWaiting
P116166
P2210
P3181614
P4191615
P5950
munotes.in197

Priority Scheduling, Starvation and Ageing

Average turnaround time: 54 / 5 = 10.80

Average waiting time: 35 / 5 = 7

Context switches: 6

P1 appears three times. It started, was thrown off at 1 by P2 of priority 1, resumed, was thrown off again at 4 by P5 of priority 2, and finally finished at 16. Its completion time is the end of its last slice, which is Chapter forty six's second warning, and its two extra appearances are two extra context switches.

Starvation, and the fix

Indefinite blocking, or starvation, is the problem of priority scheduling. A low priority process is ready, and a higher priority one is always ready too, so the low priority process never runs. On a loaded system it may never run at all.

The story every textbook tells is worth repeating because it fixes the idea: when an IBM 7094 at MIT was shut down in 1973 a low priority process was found that had been submitted in 1967 and had not yet run.

Ageing

Ageing is the fix: increase the priority of a process the longer it waits. However bad its priority starts, it improves while it waits, and eventually it is the best in the queue.

Worked as a sum. Priorities run from 127, the worst, to 0, the best, and a waiting process's priority number is reduced by 1 for every 15 minutes it waits.

minutes to reach priority 0 from 127 = 127 * 15 = 1905

hours = 1905 / 60 = 31.75

So the worst possible process is guaranteed to run within about thirty two hours, whatever else arrives. That turns "it might never run" into a number, which is Chapter thirty four's bounded waiting.

Ageing costs something. A process whose priority rises can overtake a process that genuinely is more important, so a real system usually resets a process's priority after it has run.

Distinctions that carry marks

Non-preemptive priorityPreemptive priority
A higher priority arrivalwaits until the running process blocks or endsthrows the running process off
Context switchesfewermore
Response for an urgent processup to a whole burstimmediate
Used bybatch systemsevery interactive and real time system
StarvationDeadlock
The process isready, and never chosenblocked, waiting for something held
Others progressyesno
Fixed byageingChapters fifty six to sixty five

What it does not mean

A priority is not a guarantee of speed. It decides order, not duration.

Preemptive priority is not unfair to the preempted process. It loses no work: its state is saved and it resumes exactly where it was.

Ageing is not the same as round robin. Round robin gives everybody a turn regardless of priority. Ageing keeps the priority order and makes the order change over time.

munotes.in198

Priority Scheduling, Starvation and Ageing

Quick revision

  • Priority scheduling runs the ready process with the highest priority.

In this book a smaller number is a higher priority, and an answer must say which convention it uses.

  • Both forms exist. The preemptive form throws the running process off when a better one

arrives; the non-preemptive form makes it wait.

  • Priorities are internal, from something the system can measure, or external, from

outside it.

  • Shortest job first is priority scheduling with the priority set to the inverse of the next

burst.

  • The problem is indefinite blocking, or starvation: a low priority process may never

run.

  • The fix is ageing: raise a process's priority the longer it waits. From 127 to 0 at one

step per fifteen minutes is 1,905 minutes, about 31.75 hours, which is a bound.

  • Under the preemptive form a process appears several times on the chart and its completion is

the end of the last one.

Test yourself

  1. State the rule and the convention you are using. The processor goes to the ready process

with the highest priority; in this answer a smaller number means a higher priority.

  1. Why must the convention be stated? Because books and systems differ: Linux's nice value

treats a larger number as a lower priority, and an answer is unreadable without saying which way round it is.

  1. Distinguish internal from external priorities. Internal priorities are computed from

quantities the system can measure, such as memory used or time limits. External ones come from outside, such as who is paying or how urgent the work is.

  1. How is shortest job first a priority algorithm? Its priority is the inverse of the

predicted next CPU burst: the shorter the burst, the higher the priority.

  1. What is the major problem with priority scheduling, and what is the solution? Indefinite

blocking, or starvation, of low priority processes. The solution is ageing: increasing a process's priority the longer it has waited. 6. Priorities run 127 to 0 and a waiting process gains one step every 15 minutes. What is the guaranteed bound? 127 times 15 equals 1,905 minutes, which is 1905 / 60 = 31.75 hours. 7. Under preemptive priority a process appears three times in the chart. Which slice gives its completion time, and what do the extra appearances cost? The last slice. Each extra appearance is an extra context switch, which is pure overhead.

Contents This chapter on its own page

munotes.in199

Chapter Fifty-One

Round Robin, and Choosing the Quantum

Syllabus topic Module 1, "CPU Scheduling - Scheduling Algorithms (... RR ...)"

In one line

Everybody gets a turn of the same fixed length, over and over, until they are finished.

The rule

Choosesthe head of a first in, first out ready queue
Runs it forone time quantum, or time slice, or until it blocks or finishes, whichever is sooner
Thenthe process goes to the tail of the queue
Preemptiveyes, by the timer of Chapter four
Starvationimpossible
Designed fortime sharing: response time, not turnaround

With n processes and a quantum of q, no process waits more than (n - 1) times q before its next turn. That is bounded waiting with an actual number, which no other algorithm in this row gives, and it is the reason round robin is the basis of every interactive system.

The convention when a process arrives at the same instant as another is preempted: this book puts the arriving process in the queue first and the preempted one behind it. Books differ, the answer differs, and an examination answer should state which it used.

Worked in full

Schedule: round robin, quantum 4

ProcessArrivalBurst
P1024
P203
P303
SliceProcessFromTo
1P104
2P247
3P3710
4P11030
ProcessCompletionTurnaroundWaiting
P130306
P2774
P310107

Average turnaround time: 47 / 3, about 15.67

Average waiting time: 17 / 3, about 5.67

Context switches: 3

Notice slice 2: P2 ran for 3 units, not 4. A process that finishes inside its quantum gives the processor back at once, and the remaining quantum is not wasted.

Compare Chapter forty seven: first come first served on the same three processes gave an average waiting time of 17 as well when they arrived together in this order, and an average response time of 17 against round robin's 11 / 3, about 3.67. That is the trade round robin exists to make.

The quantum, at four values

The same three processes, four quanta, nothing else changed.

QuantumAverage waitingAverage turnaroundContext switches
219 / 3, about 6.3349 / 3, about 16.336
417 / 3, about 5.6747 / 3, about 15.673
825 / 3, about 8.3355 / 3, about 18.333
2049 / 3, about 16.3379 / 3, about 26.333

Read the last row. At a quantum of 20, no process is ever preempted, because P1's first slice of 20 is longer than anything the others need and they simply queue behind it. The averages are first come first served's averages, and the switches are first come first served's switches.

munotes.in200

Round Robin, and Choosing the Quantum

That is the first of the two rules about the quantum, and it is a guaranteed question:

If the quantum is larger than the longest CPU burst, round robin degenerates into first come first served.

And the second rule is Chapter forty four's arithmetic:

If the quantum is very small, the overhead of switching becomes a large fraction of the machine.

So the quantum is chosen between the two, and the usual statement of the rule is that 80 per cent of the CPU bursts should be shorter than the quantum. In practice that is between ten and one hundred milliseconds.

Turnaround does not improve smoothly

A student who expects turnaround to fall steadily as the quantum grows is surprised by a real table, and a question may exploit it. Consider four processes of 6 units each, arriving together.

QuantumWhat happensAverage turnaround
6each runs once, to completion15
1all four interleave and all finish near the end22.5

Turnaround time improves when most processes finish inside one quantum, and it gets worse when they do not. It is not monotonic in the quantum, and the rule to quote is that the quantum should be large compared with the context switch time and large enough that most bursts complete within it.

Distinctions that carry marks

First come first servedRound robin
Preemptivenoyes
Needs a timernoyes
Response timepoorgood, and bounded by (n - 1) times q
Turnaround timecan be betterusually worse
Starvationimpossibleimpossible
Quantum larger than every burstthis is what it becomesbecomes this
A small quantumA large quantum
Responsevery goodpoor
Switching overheadlargesmall
Turnaroundworsebetter, up to a point
At the extremethe machine does nothing but switchit is first come first served

What it does not mean

Round robin does not give everybody an equal share of the processor. It gives everybody an equal turn. A process with a long burst gets many turns and therefore more processor time in total, which is correct.

The quantum is not the burst. A process may need twenty quanta.

Round robin is not fair to processor bound processes. They are preempted constantly and pay the switching cost, which is what Chapter fifty four's feedback queue fixes by demoting them to a queue with a longer quantum.

A process that blocks does not use its whole quantum. It goes to the waiting queue and, when it is ready again, joins the tail with a fresh quantum.

Quick revision

  • Round robin runs the head of a first in, first out queue for one quantum, then moves it
munotes.in201

Round Robin, and Choosing the Quantum

to the tail. Preemptive, and it cannot starve anybody.

  • No process waits more than (n - 1) times q for its next turn, which is bounded waiting with

a number.

  • A process that finishes inside its quantum returns the processor at once.
  • A quantum larger than the longest burst makes it first come first served. Measured here: at

a quantum of 20 the averages and the switch count were identical to that algorithm's.

  • A very small quantum spends the machine on switching: Chapter forty four's s / (q + s).
  • The rule of thumb: the quantum should be large compared with the switch time, and about

80 per cent of bursts should finish inside it. Ten to a hundred milliseconds in practice.

  • Turnaround is not monotonic in the quantum: it is best when most processes finish within

one quantum.

Test yourself

  1. State the round robin rule. The ready queue is treated as a first in, first out queue; the

process at the head runs for at most one time quantum, then goes to the tail.

  1. What is the bound on how long a process waits for its next turn? With n processes and a

quantum q, no more than (n - 1) times q.

  1. What happens when the quantum is larger than every CPU burst? No process is ever

preempted, so round robin becomes first come first served, with the same averages and the same number of switches.

  1. What happens when the quantum is very small? The machine spends a large fraction of its

time switching context: s / (q + s) of it, for a switch costing s.

  1. State the rule of thumb for choosing a quantum. It should be large compared with the

context switch time, and large enough that about 80 per cent of CPU bursts complete within it, which in practice is ten to a hundred milliseconds.

  1. Does round robin give every process an equal share of the processor? No, an equal

turn. A process with more work takes more turns and so gets more processor time.

  1. Is average turnaround time always better with a larger quantum? No. It is best when most

processes finish inside a single quantum, and it can get worse either side of that.

Contents This chapter on its own page

munotes.in202

Chapter Fifty-Two

Multilevel Queue Scheduling

Syllabus topic Module 1, "CPU Scheduling - Scheduling Algorithms (... Multilevel Queue Scheduling ...)"

In one line

Divide the processes into classes, give each class a queue of its own with its own algorithm, and always serve the most important queue that is not empty.

Why one queue is not enough

Chapter forty three said that input and output bound and processor bound processes want opposite treatment. One queue and one algorithm must treat them the same, and whichever algorithm is chosen is wrong for one of them.

A multilevel queue keeps them apart and treats each properly. That is the whole idea, and it is the answer to the question of why a system would have more than one ready queue.

The rule

Processes are divided intoclasses, by a property that does not change
Each class hasits own queue and its own scheduling algorithm
Between queuesa fixed priority: the highest non-empty queue is served
A processstays in its queue for life
Preemptive between queuesyes, usually: a process arriving in a higher queue throws the running one off

A process never moves between queues, and that is what distinguishes this from the next chapter. If the classification was wrong, it stays wrong. That single sentence answers the commonest examination question on the pair.

A typical arrangement

The classic five queue example, highest priority first.

QueueHoldsAlgorithmWhy
0system processesround robin, short quantummust respond, and are trusted
1interactive processesround robin, short quantuma person is waiting
2interactive editinground robina person is waiting, less urgently
3batch processesfirst come first servednobody is waiting
4student processesfirst come first servedthey can wait

Each queue's algorithm is chosen for what is in it: round robin where response matters, first come first served where it does not and the switching would be waste.

Two ways to share between queues

The choice is examined and both answers are legitimate.

Fixed priority, or absolute priority. The highest non-empty queue is served, always. Simple, and it starves the lower queues: while any interactive process is ready, a batch process never runs.

Time slice between queues. Each queue is given a percentage of the processor: for example 80 per cent to the foreground queue and 20 per cent to the background one. Nothing starves, and the guarantee is weaker: a foreground process may wait while the background queue takes its share.

Worked in full

Four processes. P1 and P3 are interactive and belong in queue 0, run round robin with a quantum of

  1. P2 and P4 are batch and belong in queue 1, run first come first served. Queue 0 has absolute

priority.

Schedule: multilevel queue, queue 0 round robin quantum 3, queue 1 first come first served

munotes.in203

Multilevel Queue Scheduling

ProcessArrivalBurstQueue
P1060
P2081
P3140
P4251
SliceProcessFromTo
1P103
2P336
3P169
4P3910
5P21018
6P41823
ProcessCompletionTurnaroundWaiting
P1993
P2181810
P31095
P4232116

Average turnaround time: 57 / 4 = 14.25

Average waiting time: 34 / 4 = 8.50

Context switches: 5

Read the chart in two halves.

  1. From 0 to 10, only queue 0 runs. P1 and P3 take turns of three units. P2 and P4 are ready

the whole time and are never chosen, because queue 1 is only served when queue 0 is empty.

  1. From 10 onwards, queue 0 is empty, so queue 1 runs, first come first served: P2 then P4.

P4 arrived at 2 and started at 18. Sixteen units of waiting for five units of work, because of a classification it cannot change.

Starvation, and what fixes it

Absolute priority between queues starves the lower queues, and it is not a subtle failure: if interactive processes keep arriving, a batch process never runs at all. Three answers, and a question may want any of them.

AnswerWhat it doesCost
Time slice between queueseach queue gets a percentagea high priority process may wait
Ageinga waiting process's priority improvesneeds the priority to be able to change
Let processes move between queuesthis is the next chaptermore bookkeeping

Ageing does not fit comfortably here, because a process's queue is supposed to be fixed. That tension is exactly why the multilevel feedback queue was invented.

Distinctions that carry marks

Priority schedulingMultilevel queue
Structureone queue, ordered by priorityseveral queues, each with its own algorithm
Within a priority levelfirst come first servedwhatever that queue's algorithm is
Chosen forurgencythe kind of process
Multilevel queueMultilevel feedback queue
A process moves between queuesneveryes, on its behaviour
Classificationpermanent, and set when the process startsdiscovered while it runs
If the classification is wrongit stays wrongit corrects itself

What it does not mean

The queues are not priority levels of one queue. Each has its own algorithm, which a priority level does not.

A multilevel queue does not measure processes. It classifies them once, by what they are, not by what they do. Measuring is the next chapter.

Absolute priority is not always wrong. For a real time queue above everything else it is exactly right: a brake controller must not wait behind an editor.

munotes.in204

Multilevel Queue Scheduling

Quick revision

  • A multilevel queue divides processes into classes, gives each class its own queue

and its own algorithm, and serves the highest non empty queue.

  • A process never leaves its queue.
  • The classic arrangement: system, interactive, interactive editing, batch, student, with round

robin above and first come first served below.

  • Sharing between queues is either absolute priority, which starves the lower queues, or a

time slice per queue, for example 80 per cent to the foreground and 20 to the background.

  • Measured here: with absolute priority, the batch process P4 waited sixteen units for five units

of work and the two interactive processes finished first.

  • The fix for the starvation is a time slice per queue, ageing, or letting processes move, which

is the next chapter.

Test yourself

  1. What is a multilevel queue? Several ready queues, one per class of process, each with its

own scheduling algorithm, with the scheduler serving the highest priority queue that is not empty.

  1. What does a process's queue depend on, and can it change? The class it belongs to, decided

when it enters the system, and it cannot change.

  1. Why does each queue have its own algorithm? Because the classes want different things:

round robin for processes a person is waiting on, first come first served for batch work where switching is only overhead.

  1. Give the two ways of sharing the processor between queues, with a fault of each. Absolute

priority, which starves the lower queues; and a time slice per queue, under which a high priority process may wait while a lower queue takes its share.

  1. Give the classic five queue arrangement. System processes, interactive processes,

interactive editing, batch processes, student processes, in that order of priority.

  1. What is the essential difference from a multilevel feedback queue? A process never moves

between queues here, so a wrong classification stays wrong; in a feedback queue it moves according to its behaviour.

Contents This chapter on its own page

munotes.in205

Chapter Fifty-Three

Multilevel Feedback Queue Scheduling

Syllabus topic Module 1, "CPU Scheduling - Scheduling Algorithms (... Multilevel Feedback Queue Scheduling ...)"

In one line

Several queues as in the last chapter, but a process that uses up its whole quantum is moved down to a queue with a longer one, so the scheduler discovers what kind of process it is instead of being told.

Why moving matters

The last chapter's fault was that the classification was permanent and had to be got right in advance. Nobody can get it right in advance: a program may be interactive for an hour and then spend ten minutes computing.

A feedback queue measures rather than classifies. A process that gives the processor up before its quantum expires is behaving like an interactive process and stays high. A process that uses its whole quantum is behaving like a processor bound one and is moved down. Nobody declares anything and the scheduler is never wrong for long.

This is also shortest job first without knowing the burst lengths, which is the deepest thing to say about it: a process with short bursts naturally ends up in the high queue with the short quantum, so short bursts are served first, which is what Chapter forty eight wanted and could not implement.

The rule

Queuesseveral, numbered 0 at the top
Each queue hasits own quantum, getting longer further down, and its own algorithm
A new process entersthe highest queue
Uses its whole quantumit is demoted one queue
Gives the processor up earlyit stays where it is, or is promoted
The bottom queueusually first come first served, with no demotion below it
Scheduling between queuesthe highest non empty queue is served

The parameters a designer must choose

An examination question often asks what defines a multilevel feedback queue, and the answer is this list of six.

  1. The number of queues.
  2. The scheduling algorithm for each queue.
  3. The method used to decide when to promote a process.
  4. The method used to decide when to demote a process.
  5. The method used to decide which queue a process enters when it needs service.
  6. Whether and how the processor is shared between the queues.

It is the most general of the algorithms because it has the most parameters, and it is the hardest to tune for the same reason.

Worked in full

Three queues. Queue 0 has a quantum of 4, queue 1 a quantum of 8, and queue 2 is first come first served. Every process enters queue 0 and drops one queue each time it uses its whole quantum.

Schedule: multilevel feedback queue, quanta 4 8, then first come first served

ProcessArrivalBurst
P1017
P205
P333
SliceProcessFromTo
1P104
2P248
3P3811
4P11119
5P21920
6P12025
munotes.in206

Multilevel Feedback Queue Scheduling

ProcessCompletionTurnaroundWaiting
P125258
P2202015
P31185

Average turnaround time: 53 / 3, about 17.67

Average waiting time: 28 / 3, about 9.33

Context switches: 5

The demotions, traced

This is the table to be able to produce, because it is where the marks are.

TimeProcessQueue it ran inQuantumUsedWhat happened next
0 to 4P104all 4demoted to queue 1, with 13 left
4 to 8P204all 4demoted to queue 1, with 1 left
8 to 11P3043 of 4finished inside its quantum, so it never dropped
11 to 19P118all 8demoted to queue 2, with 5 left
19 to 20P2181 of 8finished
20 to 25P12none5finished at the bottom

Read P3's row. It arrived at 3, ran at 8, needed only 3 units and finished inside its quantum, so it was never demoted and never waited behind the long process's later slices. A short process is served quickly without anybody declaring it short, which is the whole purpose of the algorithm.

And read P1's three rows. It was demoted twice, so it ran in all three queues, and in the bottom one it got a long uninterrupted stretch. A processor bound process ends up where switching costs least, which is Chapter fifty one's complaint about round robin, fixed.

Starvation, and the promotion that prevents it

A long process sinks to the bottom queue. If short processes keep arriving, the top queue is never empty and the bottom one is never served. Demotion alone starves long processes.

Two standard remedies:

RemedyWhat it does
Promotion by ageinga process that has waited too long in a low queue is moved up
Periodic boostevery so often, every process is put back into the top queue

The second is what several real systems do, because it is simple and it cannot be got wrong: no process can be forgotten for longer than the boost interval, which is bounded waiting with a number again.

There is a cheat to know about: a process that deliberately gives the processor up just before its quantum expires stays in the top queue for ever and gets more than its share. Real schedulers count the processor time a process has used rather than only whether it finished its quantum, exactly to close that hole.

Distinctions that carry marks

Multilevel queueMultilevel feedback queue
Processes moveneveryes, by demotion and promotion
Classificationgiven in advancemeasured while running
Quantum per queuemay be the samelonger further down
Approximatesnothingshortest job first, without knowing the bursts
Starvationof the lower queuesof long processes, fixed by promotion or a boost
munotes.in207

Multilevel Feedback Queue Scheduling

DemotionPromotion
Happens whena process uses its whole quantumit has waited too long, or on a periodic boost
Means the process isprocessor boundin danger of starving
Effectlonger quantum, lower priorityshorter quantum, higher priority

What it does not mean

The quantum does not get shorter further down. It gets longer, so that a process which needs a lot of processor gets it in fewer, larger pieces and pays less switching.

A demotion is not a punishment. It is the scheduler learning what the process is, and the bottom queue is the right place for a long computation.

It is not the same as ageing in Chapter fifty. Ageing changes a priority number; here the process changes queue, which also changes its quantum and its algorithm.

Quick revision

  • A multilevel feedback queue has several queues with increasing quanta, a new process

starting at the top, and demotion when a process uses its whole quantum.

  • It measures behaviour instead of accepting a classification, and it

approximates shortest job first without knowing any burst length.

  • Six parameters define one: the number of queues, each queue's algorithm, when to promote, when

to demote, which queue a process enters, and how the processor is shared between the queues.

  • A process that finishes inside its quantum is never demoted, so short processes stay fast.
  • A process that is demoted twice runs in the bottom queue in long uninterrupted stretches, where

switching costs least.

  • Demotion alone starves long processes. The remedies are promotion by ageing or a

periodic boost of everybody to the top queue.

  • A process can cheat by yielding just before its quantum expires, so real schedulers count

processor time used.

Test yourself

  1. How does a multilevel feedback queue differ from a multilevel queue? Processes move

between the queues according to how they behave, instead of staying in a queue fixed when they started.

  1. When is a process demoted, and what does that say about it? When it uses up its whole

quantum, which means it is processor bound rather than interactive.

  1. Why do the quanta get longer further down? So that a process needing a lot of processor

time gets it in fewer, larger pieces and pays less context switching.

  1. In what sense does it approximate shortest job first? Processes with short bursts finish

inside the short quantum of the top queue and stay there, so short bursts are served first without any burst length being known.

munotes.in208

Multilevel Feedback Queue Scheduling

  1. Name the six parameters that define one. The number of queues; the algorithm for each

queue; the method for promoting a process; the method for demoting one; the method that decides which queue a process enters; and how the processor is divided between the queues.

  1. What starves, and what are the two remedies? Long processes, which sink to the bottom

queue and are never reached while higher queues have work. The remedies are promotion by ageing and a periodic boost of every process to the top queue.

  1. How can a process cheat this scheduler, and how is that closed? By giving the processor up

just before its quantum expires, so it is never demoted. Real schedulers count the processor time a process has actually used rather than only whether it finished its quantum.

Contents This chapter on its own page

munotes.in209

Chapter Fifty-Four

All Seven Algorithms on One Problem

Syllabus topic Module 1, "CPU Scheduling - Scheduling Criteria; Scheduling Algorithms"

In one line

The same four processes take seven different amounts of waiting depending on nothing but the rule used to choose between them.

The problem

Four processes, with arrivals, bursts and priorities. A smaller priority number is a higher priority.

ProcessArrivalBurstPriority
P1072
P2241
P3413
P4542

The total work is 16 units and nothing waits for a device, so every algorithm finishes at time 16 and every one keeps the processor busy from 0 to 16. Utilisation and throughput are therefore identical for all seven, which is worth saying before the table: the only things that change are the waiting, the turnaround, the response and the number of switches.

Each schedule below is shown as a chart on one line: the process, and the interval it held the processor.

The seven, worked

First come first served

Schedule: first come first served

ProcessArrivalBurstPriority
P1072
P2241
P3413
P4542

P1[0-7] P2[7-11] P3[11-12] P4[12-16]

Average waiting time: 19 / 4 = 4.75

Average turnaround time: 35 / 4 = 8.75

Average response time: 19 / 4 = 4.75

Context switches: 3

Shortest job first

Schedule: shortest job first

ProcessArrivalBurstPriority
P1072
P2241
P3413
P4542

P1[0-7] P3[7-8] P2[8-12] P4[12-16]

Average waiting time: 16 / 4 = 4

Average turnaround time: 32 / 4 = 8

Average response time: 16 / 4 = 4

Context switches: 3

Shortest remaining time first

Schedule: shortest remaining time first

ProcessArrivalBurstPriority
P1072
P2241
P3413
P4542

P1[0-2] P2[2-4] P3[4-5] P2[5-7] P4[7-11] P1[11-16]

Average waiting time: 12 / 4 = 3

Average turnaround time: 28 / 4 = 7

Average response time: 2 / 4 = 0.50

Context switches: 5

Priority

Schedule: priority

ProcessArrivalBurstPriority
P1072
P2241
P3413
P4542

P1[0-7] P2[7-11] P4[11-15] P3[15-16]

Average waiting time: 22 / 4 = 5.50

Average turnaround time: 38 / 4 = 9.50

Average response time: 22 / 4 = 5.50

Context switches: 3

Priority (preemptive)

Schedule: priority (preemptive)

ProcessArrivalBurstPriority
P1072
P2241
P3413
P4542

P1[0-2] P2[2-6] P1[6-11] P4[11-15] P3[15-16]

Average waiting time: 21 / 4 = 5.25

Average turnaround time: 37 / 4 = 9.25

Average response time: 17 / 4 = 4.25

Context switches: 4

Round robin, quantum 2

Schedule: round robin, quantum 2

ProcessArrivalBurstPriority
P1072
P2241
P3413
P4542
munotes.in210

All Seven Algorithms on One Problem

P1[0-2] P2[2-4] P1[4-6] P3[6-7] P2[7-9] P4[9-11] P1[11-13] P4[13-15] P1[15-16]

Average waiting time: 20 / 4 = 5

Average turnaround time: 36 / 4 = 9

Average response time: 6 / 4 = 1.50

Context switches: 8

Round robin, quantum 4

Schedule: round robin, quantum 4

ProcessArrivalBurstPriority
P1072
P2241
P3413
P4542

P1[0-4] P2[4-8] P3[8-9] P1[9-12] P4[12-16]

Average waiting time: 18 / 4 = 4.50

Average turnaround time: 34 / 4 = 8.50

Average response time: 13 / 4 = 3.25

Context switches: 4

The comparison

AlgorithmWaitingTurnaroundResponseSwitches
First come first served4.758.754.753
Shortest job first4843
Shortest remaining time first370.505
Priority5.509.505.503
Priority (preemptive)5.259.254.254
Round robin, quantum 2591.508
Round robin, quantum 44.508.503.254

Six things to read off that table, and each is a possible question.

  1. Shortest remaining time first wins on waiting time, at 3 units, and it is not close. That

is Chapter forty eight's optimality, applied at every instant rather than only at completions.

  1. It also wins on response time, at half a unit, because it starts a short newcomer

immediately.

  1. Round robin at quantum 2 is second on response at 1.50, and it pays for it with eight

context switches, more than twice anybody else's. That is the trade of Chapter fifty one in one line of a table.

  1. Round robin at quantum 4 is better than at quantum 2 on every measure except response. A

larger quantum means fewer switches and less waiting, and worse response. There is no quantum that is best at everything.

  1. Priority is the worst on waiting and turnaround here, at 5.50 and 9.50, because the

priorities were not chosen to match the burst lengths. Priority scheduling optimises what the priorities say, and if they say the wrong thing it optimises the wrong thing.

  1. Preemptive priority beats non preemptive priority on all three times and costs one more

switch. That is the general shape of preemption: better times, more switches.

The ranking, and what it is worth

CriterionBest hereWorst here
Waiting timeshortest remaining time first, 3priority, 5.50
Turnaround timeshortest remaining time first, 7priority, 9.50
Response timeshortest remaining time first, 0.50priority, 5.50
Context switchesfirst come first served, shortest job first and priority, 3round robin at quantum 2, 8

And now the sentence that matters more than the table. Shortest remaining time first won everything on this problem and it is not the algorithm any general purpose system uses, for three reasons already given: it needs burst lengths nobody knows (Chapter forty eight), it can starve a long process while it is running (Chapter forty nine), and this problem has no input and output in it at all.

munotes.in211

All Seven Algorithms on One Problem

A comparison on one problem ranks the algorithms on that problem. What makes a scheduler good is how it behaves over every workload, including the ones that arrive tomorrow, and that is why the algorithm real systems use is the one that measures rather than assumes: the multilevel feedback queue of Chapter fifty three, which approximates the winner of this table without needing anything it cannot know.

What it does not mean

These averages are not properties of the algorithms. They are properties of this problem under those algorithms. Change one arrival time and the ranking can change.

Fewer context switches is not better by itself. First come first served has three and the worst response time of the non priority algorithms.

Utilisation is not a way to tell these apart. With no input and output in the problem, every algorithm keeps the processor busy for all sixteen units.

Quick revision

  • The same four processes give seven different sets of averages. The

total work, the finishing time, utilisation and throughput are identical for all seven.

  • Shortest remaining time first gave the lowest waiting, turnaround and response on this

problem.

  • Round robin at quantum 2 gave the second best response and eight switches, against

three for the non preemptive algorithms.

  • A larger quantum means fewer switches and lower waiting, and worse response.
  • Priority was worst here because the priorities did not match the burst lengths: it

optimises what the priorities say.

  • Preemption improved all three times and cost one more switch, which is its general shape.
  • A comparison on one problem is a ranking on that problem. The winner here is unusable in

practice, and the multilevel feedback queue is what a real system runs.

Test yourself

  1. Why are utilisation and throughput the same for all seven algorithms here? No process

waits for a device, so the processor is busy from 0 to 16 whatever the order, and all four finish by 16.

  1. Which algorithm gave the lowest average waiting time, and why? Shortest remaining time

first, because at every instant it runs the process with the least work left, which is the optimal choice for average waiting time.

  1. Which gave the best response time, and which was second? Shortest remaining time first at

0.50, then round robin at quantum 2 at 1.50.

  1. What did round robin at quantum 2 pay for its response time? Eight context switches, more
munotes.in212

All Seven Algorithms on One Problem

than twice as many as any non preemptive algorithm.

  1. Why did priority scheduling do worst here? Because the priorities were not related to the

burst lengths. Priority scheduling optimises the order the priorities specify, and here they specified an order that was bad for waiting time.

  1. Compare preemptive and non preemptive priority on this problem. The preemptive form was

better on waiting, turnaround and response, and cost one more context switch.

  1. If shortest remaining time first wins, why does no general purpose system use it? It needs

burst lengths that cannot be known, it can starve a long process, and a comparison on one problem with no input and output is not a comparison of workloads. A multilevel feedback queue approximates it using only what can be measured.

Contents This chapter on its own page

munotes.in213

Chapter Fifty-Five

Thread Scheduling, and What Linux Actually Does

Syllabus topic Module 1, "CPU Scheduling - Thread Scheduling"

In one line

It is threads, not processes, that the kernel schedules, and a library may schedule its own user threads on top of that.

Contention scope

Chapter twenty nine said that a user thread must be mapped onto a kernel thread before it can run. That gives two different competitions, and the word for which one a thread is in is its contention scope.

Process contention scope, PCSSystem contention scope, SCS
The thread competes withthe other threads of its own processevery thread on the machine
Scheduled bythe thread library, in user modethe kernel
Used bymany to one and many to many systemsone to one systems
Priority set bythe programmer, and the library obeys itthe kernel, from its own policy

On a one to one system, which Chapter twenty nine showed this machine to be, every thread has system contention scope and the library does no scheduling at all. So on Linux, on Windows and on macOS the whole of the process contention scope column is theory, and an examination answer should say so rather than pretending both are in use.

POSIX lets a program ask for either, with pthread_attr_setscope and the values PTHREAD_SCOPE_PROCESS and PTHREAD_SCOPE_SYSTEM. A system that supports only one of them refuses the other, and Linux supports only PTHREAD_SCOPE_SYSTEM.

The policies a real kernel offers

Linux schedules threads under one of several policies, and a thread's policy decides which algorithm from this module applies to it.

PolicyWhat it isWhich algorithm
SCHED_OTHER, also SCHED_NORMALthe ordinary policy every process starts witha fair share scheduler, described below
SCHED_FIFOreal time, first in first outpriority scheduling, non preemptive within a priority
SCHED_RRreal time, round robinround robin within each priority level
SCHED_BATCHfor work nobody is waiting onas SCHED_OTHER, but never treated as interactive
SCHED_IDLErun only when nothing else wants the processorthe lowest possible

The two real time policies are absolute priorities above everything else. A SCHED_FIFO thread at priority 50 runs in preference to every ordinary process on the machine, and if it loops for ever, nothing else runs at all. That is why creating one requires privilege, and it is Chapter fifty's starvation, permitted on purpose because a brake controller must not wait behind an editor.

The nice value, and what it really controls

An ordinary process carries a nice value from -20 to 19.

A HIGHER nice value means a LOWER priority. The name is the explanation: a process that is nice to others takes less. This is the opposite convention from Chapter fifty's, which is exactly why that chapter insisted an answer must state which convention it uses.

munotes.in214

Thread Scheduling, and What Linux Actually Does

niceMeaning
-20the greediest, and it needs privilege to ask for
0the default
19the least demanding

Nice does not set a fixed share. On the fair share scheduler it weights a share: a difference of one nice level changes a process's share of the processor by about ten per cent, so a difference of ten levels is about a factor of ten.

Read off the machine

$ chrt -p $$
pid 8's current scheduling policy: SCHED_OTHER
pid 8's current scheduling priority: 0
$ nice
0
$ chrt -m | head -6
SCHED_OTHER min/max priority	: 0/0
SCHED_FIFO min/max priority	: 1/99
SCHED_RR min/max priority	: 1/99
SCHED_BATCH min/max priority	: 0/0
SCHED_IDLE min/max priority	: 0/0
SCHED_DEADLINE min/max priority	: 0/0
$ nice -n 10 chrt -p $$ | tail -1
pid 8's current scheduling priority: 0
$ nice -n 10 nice
10

Read it against the tables above.

  • The shell runs SCHED_OTHER with priority 0, which is the only priority that policy has:

an ordinary process has no real time priority at all, and its treatment comes from its nice value instead.

  • The two real time policies have priorities 1 to 99. Those are genuine fixed priorities and they

sit above every ordinary process.

  • SCHED_DEADLINE appears, a policy where a thread states how much processor it needs and by

when. It is outside MU's syllabus and worth knowing exists.

  • nice -n 10 ran a command with a nice value of 10, and the value is inherited by what it

starts.

The kernel publishes what it counts for each thread, which is the fair share scheduler's own bookkeeping.

$ grep -E '^(nr_switches|policy|prio)' /proc/self/sched
nr_switches                                  :                    6
policy                                       :                    0
prio                                         :                  120
$ awk '/se.sum_exec_runtime/ {print "processor milliseconds used:", $3}' /proc/self/sched
processor milliseconds used: 2.947250

prio is 120, which is 120 minus the nice value of 0 in the kernel's own internal numbering where 100 is the best ordinary priority and 139 the worst. The two numbering systems, nice from -20 to 19 and the internal 100 to 139, describe the same thing: nice plus 120 is the internal number.

What the fair share scheduler does, in one section

MU's syllabus stops at the seven algorithms, and a question may ask what a real system uses. The answer, from the kernel's own documentation, is none of them exactly.

Instead of a queue in an order, the scheduler keeps for each thread the amount of processor time it has already had, weighted by its nice value, and always runs the thread that has had the least. There is no quantum in the round robin sense: a thread runs until somebody else's weighted time falls below its own.

munotes.in215

Thread Scheduling, and What Linux Actually Does

That is shortest remaining time first turned inside out. Chapter forty eight could not implement shortest job first because it needed the future. This needs only the past, which is known exactly, and it gets much of the same effect: a thread that has used little processor time, which is what an interactive thread looks like, is always chosen first.

The name to know is that this family of schedulers is described in authorities/kernel/scheduler-sched-design-CFS.txt and its successor in authorities/kernel/scheduler-sched-eevdf.txt, both fetched from the kernel's own documentation. Neither is examinable on this paper; what is examinable is being able to say that a modern system uses a fair share scheduler based on processor time already consumed, rather than any of the seven textbook algorithms.

Distinctions that carry marks

Process contention scopeSystem contention scope
Competes withthe threads of its own processevery thread on the machine
Scheduled bythe librarythe kernel
Exists on Linuxnoyes, for every thread
SCHED_OTHERSCHED_FIFO and SCHED_RR
Priority range0 only1 to 99
Ordered byfair share, weighted by niceabsolute priority
Can starve everything elsenoyes, deliberately
Needs privilegenoyes
SCHED_RR differs from SCHED_FIFO byhaving a quantum within each priority
A priority number in Chapter fiftyA nice value
Smaller meanshigher prioritylower priority
Rangewhatever the book says-20 to 19

What it does not mean

Thread scheduling is not a different set of algorithms. It is the same algorithms applied to threads, plus the question of who applies them.

A nice value of -20 does not make a process real time. It weights its share among the ordinary processes. Real time means one of the two real time policies.

SCHED_FIFO is not first come first served for the machine. It is first come first served within one priority level, and the levels are absolute.

Quick revision

  • Contention scope: process contention scope means a thread competes only with its own

process's threads, scheduled by the library; system contention scope means it competes with every thread, scheduled by the kernel.

  • On a one to one system, which Linux is, every thread has system contention scope and

the library schedules nothing.

  • Linux policies: SCHED_OTHER for ordinary work, SCHED_FIFO and SCHED_RR for real time with

absolute priorities 1 to 99, SCHED_BATCH, SCHED_IDLE.

  • A higher nice value means a lower priority, from -20 to 19, which is the opposite of

Chapter fifty's convention. One nice level is about ten per cent of a process's share.

  • chrt -p shows a process's policy and priority; chrt -m shows the ranges; nice shows the

value.

  • An ordinary process has no real time priority: SCHED_OTHER's only priority is 0.
  • A real system uses a fair share scheduler that runs whichever thread has had the least
munotes.in216

Thread Scheduling, and What Linux Actually Does

weighted processor time so far, which achieves much of what shortest job first wanted using only the past.

Test yourself

  1. What is contention scope? Whether a thread competes for the processor only with the other

threads of its own process, which is process contention scope, or with every thread on the machine, which is system contention scope.

  1. Which does Linux use, and what follows? System contention scope for every thread, because

it maps one user thread to one kernel thread. The thread library therefore does no scheduling at all.

  1. Name the Linux scheduling policies and which are real time. SCHED_OTHER, SCHED_BATCH

and SCHED_IDLE are ordinary; SCHED_FIFO and SCHED_RR are real time, with priorities 1 to 99 above every ordinary process.

  1. What is the difference between SCHED_FIFO and SCHED_RR? Within one priority level,

SCHED_FIFO runs a thread until it blocks or yields, and SCHED_RR gives each thread a quantum in turn.

  1. What is a nice value and which direction does it run? A weighting on an ordinary process's

share of the processor, from -20 to 19, in which a higher number means a lower priority.

  1. Why must a real time policy require privilege? A real time thread has absolute priority

over every ordinary process, so one that loops for ever stops the machine doing anything else.

  1. What does a modern general purpose scheduler actually do? It keeps for each thread the

processor time it has already used, weighted by its nice value, and runs whichever has used the least. That needs only the past, and it gets much of the benefit shortest job first wanted from the future.

Contents This chapter on its own page

munotes.in217

Module II

Deadlocks, Memory Management, Virtual Memory and Mass Storage, and the File System

munotes.in

Chapter Fifty-Six

What a Deadlock Is, and the System Model

Syllabus topic Module 2, "Deadlocks - System Model"

In one line

A deadlock is a set of processes in which every one of them is waiting for something that another one in the set is holding, so none of them can ever move again.

The form to write: a set of processes is deadlocked when every process in the set is waiting for an event that can be caused only by another process in the set.

Why that definition is worded so carefully

Two clauses do all the work.

  • Every process in the set. If one of them can move, it will eventually release what it holds

and the others follow. A deadlock is not one stuck process; it is a closed set.

  • Can be caused only by another process in the set. The event will never happen, because the

only processes that could cause it are themselves waiting. Nothing outside the set can help.

So a deadlock is permanent and it is self inflicted. No timer expires, no device interrupts, no operator intervention arrives. Left alone, those processes are still there when the machine is switched off. That is what makes it different from every other kind of waiting in this book.

The system model

MU's label is "System Model", and it is the vocabulary that makes the rest of the module precise.

A system has a finite number of resources to be shared among competing processes. The resources are of types, and a type has some number of identical instances.

WordMeaningExample
Resource typea kind of thing a process can needprocessor, memory, printer, a lock, a file
Instanceone interchangeable unit of a typeone of the three printers
Requesta process asks for instances of a typeit may have to wait
Usethe process operates on them
Releasethe process gives them back

If instances of a type are truly interchangeable, any instance satisfies a request for that type. That sentence is the test for whether two things are one type or two. Three identical printers are one type with three instances. A colour printer and a black and white one are two types, because a request for colour is not satisfied by the other.

Request, use, release is the three step cycle of every resource in the system, and every deadlock happens between the first two steps.

A request is made with a system call: open for a file, wait for a semaphore, pthread_mutex_lock for a lock, and the allocating calls for memory. The release is the matching call. A resource a process obtains without the kernel's knowledge, such as a lock held entirely in shared memory, is one the kernel cannot help with, which is why a deadlock between two threads is usually invisible to the operating system.

munotes.in218

What a Deadlock Is, and the System Model

A deadlock, made to happen

Two mutexes and two threads that take them in opposite orders. This is the shortest deadlock that can be written and the commonest one in real programs.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <pthread.h>
#include <time.h>

static pthread_mutex_t first = PTHREAD_MUTEX_INITIALIZER;
static pthread_mutex_t second = PTHREAD_MUTEX_INITIALIZER;
static int             got_both[2];

static void pause_for(long milliseconds)
{
    struct timespec t = {milliseconds / 1000, (milliseconds % 1000) * 1000000};

    nanosleep(&t, NULL);
}

static void *takes_first_then_second(void *unused)
{
    (void)unused;
    pthread_mutex_lock(&first);
    pause_for(100);                       /* hold one, and pause: see the note */
    pthread_mutex_lock(&second);
    got_both[0] = 1;
    pthread_mutex_unlock(&second);
    pthread_mutex_unlock(&first);
    return NULL;
}

static void *takes_second_then_first(void *unused)
{
    (void)unused;
    pthread_mutex_lock(&second);          /* the OTHER order */
    pause_for(100);
    pthread_mutex_lock(&first);
    got_both[1] = 1;
    pthread_mutex_unlock(&first);
    pthread_mutex_unlock(&second);
    return NULL;
}

int main(void)
{
    pthread_t a, b;

    pthread_create(&a, NULL, takes_first_then_second, NULL);
    pthread_create(&b, NULL, takes_second_then_first, NULL);
    pause_for(1000);                      /* a whole second is plenty */

    printf("thread A got both locks: %s\n", got_both[0] ? "yes" : "no");
    printf("thread B got both locks: %s\n", got_both[1] ? "yes" : "no");
    printf("%s\n", (!got_both[0] && !got_both[1])
        ? "neither of them ever will: this is a DEADLOCK"
        : "they got through");
    return 0;                             /* leave them holding their locks */
}
$ gcc -std=c17 -Wall -Wextra -pthread -o deadlock2 deadlock2.c
$ ./deadlock2
thread A got both locks: no
thread B got both locks: no
neither of them ever will: this is a DEADLOCK
$ bad=$(for i in 1 2 3; do ./deadlock2 | tail -1; done | grep -c DEADLOCK)
$ echo "three runs, $bad of them deadlocked"
three runs, 3 of them deadlocked

Read the program against the definition. Thread A holds first and waits for second. Thread B holds second and waits for first. Each is waiting for an event, the release of a mutex, that only the other one in the set can cause, and the other one is waiting too. Every clause of the definition is satisfied, and the program will sit there for ever.

The hundred millisecond pause is there for the same reason as Chapter forty one's: it makes the bad interleaving certain rather than occasional. Take it out and the deadlock happens sometimes, which in a real program means it happens in production and not in testing.

Notice what the operating system did about it: nothing. It does not know the two threads are stuck. It sees two threads waiting on two locks, which is an entirely normal thing for threads to do. Chapters sixty four and sixty five are about a system that does look.

Four kinds of waiting, told apart

This table is the reason the definitions of this module matter, and a question will ask for the differences.

munotes.in219

What a Deadlock Is, and the System Model

What the process is doingWill it end by itself?
Blocked on input or outputwaiting for a device (Chapter fifteen)yes, when the device finishes
Starvedready, and never chosen (Chapter fifty)perhaps never, but it could be chosen at any moment
In a livelockrunning, and making no progress: two processes politely stepping aside for each other for everno, and it is using the processor while not progressing
Deadlockedwaiting for something held by another waiting processnever

Starvation and deadlock are the pair most often confused. A starved process could run: give it the processor and it proceeds. A deadlocked process cannot: give it the processor and it goes straight back to waiting for something that will never arrive.

A livelock is the one students have not usually heard of and examiners like. The processes are not blocked at all: they are running, and their state keeps changing, and no work gets done. Two people stepping aside in a corridor, each way, for ever.

Worked example: a bank transfer

The commonest real deadlock in commercial code, and the one worth recognising.

A transfer locks the account it takes money from and then the account it puts money into.

  1. Anita transfers to Bharat. Her thread locks Anita's account, then asks for Bharat's.
  2. At the same moment Bharat transfers to Anita. His thread locks Bharat's account, then asks for

Anita's.

  1. Neither can proceed. Two customers, two accounts, no money moved, and the bank's software

stops.

It is the same program as the one above with the locks renamed, and the fix is the same as Chapter forty one's third fix: lock the accounts in a fixed order, for example by account number, whichever way the money is going. That is Chapter sixty of this module, and it costs nothing.

What it does not mean

A deadlock is not a crash. Nothing has failed and nothing will be reported. The processes are perfectly healthy and waiting politely.

A deadlock is not a performance problem. It does not get better under a lighter load.

Two processes waiting on one resource are not deadlocked. One of them has it and will release it. A deadlock needs a cycle of waiting.

Deadlock is not confined to locks. Memory, files, devices, database rows and network connections all deadlock the same way. The resource being a lock is the commonest case and not the definition.

Quick revision

  • A set of processes is deadlocked when every process in the set waits for an event that only

another process in the set can cause.

  • Both clauses matter: every process, and the event can come only from inside the set. So
munotes.in220

What a Deadlock Is, and the System Model

a deadlock is permanent and nothing outside can free it.

  • System model: resources come in types, each with interchangeable instances. The

cycle is request, use, release, and a deadlock happens between request and use.

  • Two things are one type if any instance satisfies a request; a colour printer and a mono

printer are two types.

  • Two mutexes taken in opposite orders by two threads is the shortest deadlock there is, and the

operating system does not notice it.

  • Blocked ends by itself; starved could be chosen at any moment; a livelock is

running and making no progress; deadlocked never ends.

Test yourself

  1. Define a deadlock. A set of processes is deadlocked when every process in the set is

waiting for an event that can be caused only by another process in the same set.

  1. Why does the definition say "every process in the set"? Because if one of them could move

it would eventually release what it holds and the others would follow. A deadlock is a closed set, not one stuck process.

  1. State the three steps of using a resource. Request, use, release.
  2. When are two resources of the same type? When the instances are interchangeable, so that a

request for the type is satisfied by any of them. A colour printer and a black and white printer are different types.

  1. Distinguish deadlock from starvation. A starved process is ready and could run the moment

it is chosen. A deadlocked process cannot run even if it is given the processor, because what it waits for will never arrive.

  1. What is a livelock? Processes that are running and changing state but making no progress,

for example two that keep standing aside for each other.

  1. Write the shortest deadlock you can. Two threads and two mutexes: one takes A then B, the

other takes B then A. With any pause between the two acquisitions it deadlocks every time.

Contents This chapter on its own page

munotes.in221

Chapter Fifty-Seven

The Four Conditions

Syllabus topic Module 2, "Deadlocks - Deadlock Characterization"

In one line

A deadlock can happen only if all four of these hold at the same time: mutual exclusion, hold and wait, no preemption, and circular wait.

The four, and what each one says

ConditionIt saysTake it away and
Mutual exclusionat least one resource is held in a non sharable mode: only one process at a timea process never has to wait for the resource at all
Hold and waita process holds at least one resource and is waiting for another one held by somebody elsenobody ever waits while holding, so nothing accumulates
No preemptiona resource can be released only voluntarily, by the process holding itthe resource can be taken back, so a wait can be broken
Circular waitthere is a set P0, P1, ... Pn in which P0 waits for something P1 holds, P1 for something P2 holds, and Pn for something P0 holdsthe chain of waiting has an end, and whoever is at the end proceeds

All four must hold together. They are necessary conditions, not sufficient ones: the next chapter shows a case where all four hold and there is no deadlock. An answer that calls them sufficient is wrong.

And the four are not independent. Circular wait implies hold and wait, because a process in a cycle is holding something and waiting for something. That is worth saying in an answer, because it explains why Chapter sixty's prevention methods attack them one at a time and why breaking circular wait is the one that costs least.

The four are named after the paper that set them out: Coffman, Elphick and Shoshani, System Deadlocks, ACM Computing Surveys volume 3, 1971, pages 67 to 78, recorded in authorities/sources.json. They are often called the Coffman conditions.

Each one removed, on the lab machine

The deadlock of Chapter fifty six is the test bed: two threads, two mutexes, opposite orders. Four variants, each with exactly one condition taken away, and none of them deadlocks.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <pthread.h>
#include <time.h>

static pthread_mutex_t first = PTHREAD_MUTEX_INITIALIZER;
static pthread_mutex_t second = PTHREAD_MUTEX_INITIALIZER;
static pthread_mutex_t both_at_once = PTHREAD_MUTEX_INITIALIZER;
static pthread_rwlock_t shared = PTHREAD_RWLOCK_INITIALIZER;
static int  done[2];
static char which[32];

static void pause_for(long milliseconds)
{
    struct timespec t = {milliseconds / 1000, (milliseconds % 1000) * 1000000};

    nanosleep(&t, NULL);
}

static void *work(void *argument)
{
    int me = *(int *)argument;

    if (strcmp(which, "deadlock") == 0) {
        /* all four conditions hold */
        pthread_mutex_t *a = me == 0 ? &first : &second;
        pthread_mutex_t *b = me == 0 ? &second : &first;
        pthread_mutex_lock(a);
        pause_for(100);
        pthread_mutex_lock(b);
        done[me] = 1;
        pthread_mutex_unlock(b);
        pthread_mutex_unlock(a);
    } else if (strcmp(which, "no-mutual-exclusion") == 0) {
        /* the resource is sharable: both may hold it as readers */
        pthread_rwlock_rdlock(&shared);
        pause_for(100);
        pthread_rwlock_rdlock(&shared);
        done[me] = 1;
        pthread_rwlock_unlock(&shared);
        pthread_rwlock_unlock(&shared);
    } else if (strcmp(which, "no-hold-and-wait") == 0) {
        /* take both together under one gate, or neither */
        pthread_mutex_lock(&both_at_once);
        pthread_mutex_lock(&first);
        pthread_mutex_lock(&second);
        pthread_mutex_unlock(&both_at_once);
        pause_for(100);
        done[me] = 1;
        pthread_mutex_unlock(&second);
        pthread_mutex_unlock(&first);
    } else if (strcmp(which, "preemption") == 0) {
        /* the second lock may be given up: try, and back out if it is held */
        pthread_mutex_t *a = me == 0 ? &first : &second;
        pthread_mutex_t *b = me == 0 ? &second : &first;
        for (;;) {
            pthread_mutex_lock(a);
            pause_for(100);
            if (pthread_mutex_trylock(b) == 0) {
                break;
            }
            pthread_mutex_unlock(a);       /* release what we hold and retry */
            pause_for(10);
        }
        done[me] = 1;
        pthread_mutex_unlock(b);
        pthread_mutex_unlock(a);
    } else {                              /* no-circular-wait */
        /* both take them in the SAME order */
        pthread_mutex_lock(&first);
        pause_for(100);
        pthread_mutex_lock(&second);
        done[me] = 1;
        pthread_mutex_unlock(&second);
        pthread_mutex_unlock(&first);
    }
    return NULL;
}

int main(int argc, char **argv)
{
    pthread_t id[2];
    int name[2] = {0, 1};

    snprintf(which, sizeof which, "%s", argc > 1 ? argv[1] : "deadlock");
    pthread_create(&id[0], NULL, work, &name[0]);
    pthread_create(&id[1], NULL, work, &name[1]);
    pause_for(1500);
    printf("%-20s both threads finished: %s\n", which,
        (done[0] && done[1]) ? "yes" : "NO, they are deadlocked");
    exit(0);                              /* leave any stuck threads where they are */
}
munotes.in222

The Four Conditions

$ gcc -std=c17 -Wall -Wextra -pthread -o four four.c
$ ./four deadlock
deadlock             both threads finished: NO, they are deadlocked
$ ./four no-mutual-exclusion
no-mutual-exclusion  both threads finished: yes
$ ./four no-hold-and-wait
no-hold-and-wait     both threads finished: yes
$ ./four preemption
preemption           both threads finished: yes
$ ./four no-circular-wait
no-circular-wait     both threads finished: yes

One line deadlocks and four do not, and the only difference in each case is one condition. That is what "necessary" means, demonstrated rather than asserted: remove any one of the four and the deadlock cannot form, whatever the timing.

Read what each variant actually did.

  • No mutual exclusion: the resource became a reader writer lock and both threads held it as

readers. A sharable resource is never a party to a deadlock.

  • No hold and wait: both locks are taken under a single outer gate, so a thread never holds

one while waiting for the other. This is Chapter forty one's second fix.

  • No preemption: pthread_mutex_trylock fails instead of waiting, and the thread

releases what it holds and tries again. That is preemption of its own resources, done voluntarily.

  • No circular wait: both threads take first then second. There is no cycle to close. This

is Chapter forty one's third fix, and it is one line.

The third variant is worth a warning. Releasing and retrying can livelock: two threads can keep grabbing, failing, releasing and grabbing again for ever. The ten millisecond pause before retrying is what makes that unlikely, and a real implementation waits a random time, which is why network protocols do the same.

munotes.in223

The Four Conditions

What "necessary but not sufficient" means here

A question sometimes asks whether the four conditions guarantee a deadlock. They do not.

  • If a deadlock exists, all four hold. Necessary.
  • All four holding does not mean a deadlock exists. Not sufficient.

The next chapter gives the case: with several instances of a resource type, a cycle in the resource allocation graph can exist while every process still finishes.

So the four conditions are what you check to know that a deadlock is possible, which is exactly what Chapter sixty's prevention and Chapter sixty one's avoidance need.

Worked example: naming the four in a real system

A print server has one printer and one scanner. Job A has the printer and wants the scanner to copy a page. Job B has the scanner and wants the printer.

ConditionWhere it is in this story
Mutual exclusionneither the printer nor the scanner can be used by two jobs at once
Hold and waitA holds the printer and waits for the scanner
No preemptionthe printer cannot be taken from A halfway through a page
Circular waitA waits for B's scanner, B waits for A's printer

And the cheapest fix, as always: make every job take the printer before the scanner. Circular wait cannot form, nothing else changes, and no job loses anything.

What it does not mean

Mutual exclusion is not a fault. It is why the resource is worth protecting. Removing it is only possible where the resource is genuinely sharable, which a printer is not.

No preemption does not mean the processor cannot be preempted. It is about the resource in question. A processor is preemptible and so never deadlocks; a printer halfway through a page is not.

Circular wait is not the same as a cycle in a graph. It is a cycle of waiting, and the next chapter shows that a graph cycle is not always one.

Breaking a condition is not free. Each of the four costs something, and Chapter sixty prices them.

Quick revision

  • Four necessary conditions, all of which must hold at once: mutual exclusion,

hold and wait, no preemption, circular wait.

  • They are necessary and not sufficient: all four can hold with no deadlock, when a resource

type has several instances.

  • Circular wait implies hold and wait, so the four are not independent.
  • Named after Coffman, Elphick and Shoshani, System Deadlocks, ACM Computing Surveys 3, 1971.
  • Measured here: the same program deadlocks with all four, and finishes with any one of them
munotes.in224

The Four Conditions

removed.

  • The cheapest condition to break is circular wait: order the resources and take them in that

order.

  • Breaking no preemption by releasing and retrying risks a livelock, so the retry waits a

little.

Test yourself

  1. Name the four necessary conditions for deadlock. Mutual exclusion, hold and wait, no

preemption, and circular wait.

  1. Are they sufficient? No, only necessary. With several instances of a resource type all

four can hold while every process still finishes.

  1. Which two are not independent, and why? Circular wait implies hold and wait: a process in

the cycle is holding one resource and waiting for another.

  1. How would you remove mutual exclusion, and when is that possible? By making the resource

sharable, for example allowing many readers at once. It is possible only for resources that genuinely can be shared: a read only file can, a printer cannot.

  1. Give one way to break hold and wait. Require a process to request all the resources it

needs at once, so that it never holds one while waiting for another.

  1. How is no preemption broken, and what is the risk? By having a process release what it

holds when it cannot get the next thing, and retry. The risk is a livelock, in which the processes keep grabbing and releasing for ever; a random wait before retrying makes that unlikely.

  1. What is the cheapest condition to break and how? Circular wait: number the resource types

and require every process to request them in increasing order, so a cycle cannot close.

Contents This chapter on its own page

munotes.in225

Chapter Fifty-Eight

The Resource Allocation Graph

Syllabus topic Module 2, "Deadlocks - Deadlock Characterization"

In one line

Draw the processes and the resources as dots and the requests and allocations as arrows, and a deadlock shows up as a cycle.

How to draw one

A resource allocation graph has two kinds of vertex and two kinds of edge.

Drawn asMeans
A circle, labelled P1, P2a process
A rectangle, labelled R1, R2, with one dot inside per instancea resource type
An arrow from a process to a resource typea request edge: P is asking and waiting
An arrow from a dot in a resource to a processan assignment edge: that instance is held by P

A request edge points from the process to the resource; an assignment edge points from the resource to the process. Getting the direction wrong is the commonest way to lose the marks on this question, and the way to remember it is that the arrow points at what you want or at who has it.

When a request is granted, the request edge is turned round and becomes an assignment edge. When the resource is released, the edge is removed. That is the whole of how the graph changes.

How this book writes one

A drawing does not survive being read on a phone, so a graph here is written as its list of edges, which is exactly what a drawing is.

FromToKind
P1R1request: P1 is waiting for R1
R1P2assignment: P2 holds R1
P2R2request: P2 is waiting for R2
R2P1assignment: P1 holds R2

Read it as a sentence: P1 wants R1, which P2 has; P2 wants R2, which P1 has. The arrows lead P1, R1, P2, R2, back to P1, and that is a cycle.

The rule, in two halves

This is the examinable sentence and it has two halves that must both be given.

If the graph contains no cycle, then no process is deadlocked.

If the graph contains a cycle, then a deadlock may exist. If every resource type in the cycle has exactly one instance, the cycle means a deadlock has occurred. If any resource type in the cycle has several instances, the cycle is not sufficient.

The first half is the useful one for a checker: no cycle, no deadlock, with no exceptions. That is why the detection algorithm of Chapter sixty four for single instance types is simply a cycle search.

A cycle that IS a deadlock

The four edges above, with one instance each of R1 and R2.

FromToKind
P1R1request
R1P2assignment
P2R2request
R2P1assignment

The cycle: P1, R1, P2, R2, P1.

Every resource type has one instance, so the only process that could release R1 is P2, and P2 is waiting for R2, which only P1 can release, and P1 is waiting for R1. Deadlocked, and it is Chapter fifty six's two mutexes drawn as a graph.

munotes.in226

The Resource Allocation Graph

A cycle that is NOT a deadlock

Now R2 has two instances, and one of them is held by a process that is outside the cycle's waiting.

FromToKind
P1R1request
R1P2assignment
P2R3request
R3P3assignment
P3R2request
R2P1assignment: the first instance
R2P2assignment: the second instance

There are cycles here, and sim/deadlock.py finds two of them: P1, R1, P2, R3, P3, R2, P1, and P2, R3, P3, R2, P2.

And nothing is deadlocked. Follow it: P3 is waiting for an instance of R2, and R2 has two, one held by P1 and one held by P2. P2 is not waiting for R2; it is waiting for R3. So the moment P2 or P1 finishes its current work and releases its instance of R2, P3 gets one and proceeds.

The lesson in one line: a cycle is a deadlock only when the processes in it are the only ones that could release what is wanted. With several instances, somebody outside the cycle of waiting may be holding one.

Drawing the graph as a state, step by step

A graph is a snapshot. A question that gives a sequence of requests wants a graph after each step.

Three processes, R1 with one instance, R2 with two.

StepEventEdges after itCycle?
1P1 requests and gets R1R1 to P1no
2P2 requests and gets an R2R1 to P1, R2 to P2no
3P1 requests R2 and gets the second oneR1 to P1, R2 to P2, R2 to P1no
4P2 requests R1 and waitsabove, plus P2 to R1no: P1 holds R1 and is not waiting
5P1 requests another R2 and waitsabove, plus P1 to R2yes: P1 to R2 to P2 to R1 to P1

At step 5 both instances of R2 are held, by P1 and P2, and P2 is waiting for R1 which P1 holds. The cycle closes and nothing can move: deadlocked, even though R2 has two instances, because both of them are held by processes in the cycle.

Compare that with the previous section, where R2's second instance was held by a process that was not waiting on anything in the cycle. That is the whole difference, and it is what a question about "a cycle with multiple instances" is testing.

Distinctions that carry marks

Request edgeAssignment edge
Points froma processa resource instance
Points toa resource typea process
Meansthe process is waitingthe process is holding
Becomesan assignment edge when grantedremoved when released
munotes.in227

The Resource Allocation Graph

No cycleA cycle, one instance per typeA cycle, several instances
Deadlockneveralwaysmay or may not
What to donothingit is a deadlocklook at who holds the other instances
Resource allocation graphWait for graph (Chapter sixty four)
Verticesprocesses and resourcesprocesses only
An edge meansa request or an allocationP is waiting for Q
Built bydrawing the statecollapsing the resource vertices out
Used forcharacterising a deadlockdetecting one, single instances only

What it does not mean

The graph is not a history. It is the state at one instant. A cycle that existed a moment ago and has gone is not a deadlock.

A rectangle is not one resource. It is a resource type, and the dots inside it are the instances. A question that draws one dot has told you there is one instance.

A cycle is not the definition of deadlock. The definition is Chapter fifty six's. The cycle is a way of seeing one, and only an exact way when every type has a single instance.

No cycle does not mean no problem. The processes may still be starving, or livelocked, or simply slow.

Quick revision

  • A resource allocation graph has processes as circles, resource types as

rectangles with one dot per instance, request edges from process to resource, and assignment edges from a resource instance to a process.

  • A granted request turns the request edge round into an assignment edge.
  • No cycle means no deadlock, always. A cycle means a deadlock

if every type in it has one instance, and may or may not otherwise.

  • With several instances, the question is whether every instance is held by a process that is

itself in the cycle of waiting. If one is held by a process that can finish, there is no deadlock.

  • The graph is a snapshot of one instant, not a history.
  • A wait for graph is the same information with the resources collapsed out: processes only,

an edge meaning one waits for another.

Test yourself

  1. What do the two kinds of vertex and the two kinds of edge mean? Circles are processes and

rectangles are resource types with one dot per instance. An arrow from a process to a resource is a request; an arrow from a resource instance to a process is an assignment.

  1. Which way round do the edges point? A request edge points from the process to the resource

it wants; an assignment edge points from the resource instance to the process holding it.

munotes.in228

The Resource Allocation Graph

  1. State the rule connecting cycles and deadlock. If there is no cycle there is no deadlock.

If there is a cycle and every resource type in it has one instance, there is a deadlock. If a type in the cycle has several instances, a deadlock may or may not exist.

  1. Give a cycle that is not a deadlock. A cycle in which a resource type has two instances

and one of them is held by a process that is not waiting for anything in the cycle: that process will finish and release its instance, and the waiting process proceeds.

  1. Two instances of R2 exist and a cycle includes R2. When is it a deadlock? When both

instances are held by processes that are themselves in the cycle, so neither can finish.

  1. What happens to a request edge when the request is granted? It is reversed and becomes an

assignment edge.

  1. Distinguish a resource allocation graph from a wait for graph. The first has processes and

resources as vertices; the second has only processes, with an edge from P to Q meaning P is waiting for a resource Q holds, and it is used for detection when every type has one instance.

Contents This chapter on its own page

munotes.in229

Chapter Fifty-Nine

The Four Ways to Handle a Deadlock

Syllabus topic Module 2, "Deadlocks - Methods for Handling Deadlocks"

In one line

A system can stop deadlocks happening, or let them happen and detect them, or pretend they cannot happen and leave it to the programmer.

The four

MethodWhat it doesWhen the check happensChapters
Preventionmake one of the four conditions impossible, by designnever: the design forbids itsixty
Avoidanceallow the conditions, but refuse any request that could lead to a deadlockbefore granting every requestsixty one, sixty two, sixty three
Detection and recoverylet deadlocks happen, look for them, and break themperiodically, or when the machine looks stucksixty four, sixty five
Ignore itassume deadlocks do not occur, and if the machine stops, restart itneverthis chapter

Prevention and avoidance are both "deadlock will never happen", and they are not the same thing. That distinction is the commonest examination question in this row, and the difference is when the restriction bites:

  • Prevention restricts how a request may be made: you must ask for everything at once, or

in a fixed order, or you must release what you hold. The rule applies to every request, always, whatever the state.

  • Avoidance lets a process ask for anything at any time and

decides each request on the state of the system: it needs to know in advance what each process might eventually want, and it refuses a request that would take the system somewhere unsafe.

So prevention needs no information about the future and costs utilisation all the time. Avoidance needs a declaration of the maximum each process will ever need and costs utilisation only when the state is tight.

The one most systems use

Linux, Windows and macOS do none of the first three for ordinary resources. They ignore the problem.

That sounds like negligence and it is a considered decision, for three reasons a question may ask for.

  1. Deadlocks are rare. They need four conditions to coincide, and on a machine where most

resources are preemptible or sharable they mostly cannot.

  1. The cost of the alternatives is paid all the time. Avoidance requires every process to

declare its maximum needs in advance, which no ordinary program can do: a text editor does not know how many files the user will open. Prevention wastes resources on every request.

  1. The consequence is tolerable. A deadlocked pair of processes is noticed by a person, who

kills one of them. The machine does not fall over; two processes stop.

Where the consequence is not tolerable, the other methods are used. A database engine detects deadlocks between transactions and aborts one, automatically, several times a second on a busy server. A real time controller prevents them by design, because there is nobody to notice.

munotes.in230

The Four Ways to Handle a Deadlock

And the kernel does not ignore them for its own locks. Kernel code is written to a fixed lock ordering, which is prevention, and Linux ships a checker that watches the order at run time and reports a violation before it ever deadlocks.

The cost of each, side by side

This is the table that answers a request to compare the methods.

PreventionAvoidanceDetectionIgnoring
Deadlock can occurnonoyes, and is then brokenyes, and stays
Needs to know the futurenoyes: each process's maximumnono
Cost when nothing is wronglow utilisation, alwaysa check on every requestnothingnothing
Cost when something is wrongnothing to donothing to dothe detection run, plus a victim's lost worka person's time
Resource utilisationworstpoorbest of the threebest
Used byreal time systems, kernel lock orderingrarely, and in some schedulersdatabase enginesgeneral purpose operating systems

Read the utilisation row. Every method that guarantees no deadlock pays for the guarantee by leaving resources idle that could safely have been used. That is the trade, and it is why the method a system chooses follows from what a deadlock would cost it.

Worked example: choosing a method four times

A bank's transaction engine. Two transfers can deadlock on two accounts and it happens constantly. Detection and recovery: detect the cycle, abort the younger transaction, and let the client retry. The work lost is one transaction and the client never knows.

An aircraft's flight control computer. A deadlock is a crash and there is nobody to notice. Prevention: a fixed order for every resource, checked when the software is written and proved before it flies.

A student's laptop. A deadlock costs the student one killed program. Ignore it, and the operating system spends nothing.

A print room with three printers, a scanner and a plotter, used by long unattended jobs. Jobs can declare what they need when they are submitted, and a stuck job wastes an hour of machine time. Avoidance: ask each job for its maximum needs and refuse to start one that could deadlock.

Four systems, four answers, and the question each time is the same: what does a deadlock cost here, and what is the guarantee worth?

What it does not mean

Ignoring the problem is not the same as not knowing about it. It is a decision that the cost of the cure exceeds the cost of the disease, and it is revisited for every kind of resource: the same kernel that ignores deadlocks between user processes prevents them among its own locks.

Prevention and avoidance are not two names for one thing. Prevention constrains how requests may be made. Avoidance examines the state before granting each one.

munotes.in231

The Four Ways to Handle a Deadlock

Detection does not fix anything by itself. It finds the deadlock; recovery is a separate decision, and Chapter sixty five shows that it always costs somebody their work.

No method is free. The three that guarantee anything all lower utilisation, and a question that asks for the "best" method is asking what the deadlock would cost.

Quick revision

  • Four methods: prevention, avoidance, detection and recovery, and ignoring it.
  • Prevention makes one of the four conditions impossible by restricting

how a request may be made, and needs nothing known in advance.

  • Avoidance allows the conditions and decides each request on the current state, and

needs each process's maximum future need declared.

  • Detection lets deadlocks happen, looks for them periodically, and recovers by killing a

process or taking a resource back.

  • Most general purpose operating systems ignore the problem, because deadlocks are rare, the

alternatives cost utilisation all the time, and a person can kill a stuck process.

  • The same kernel prevents deadlocks among its own locks, by a fixed lock ordering.
  • Every guarantee is paid for in utilisation: resources left idle that could safely have been

used.

Test yourself

  1. Name the four methods of handling deadlocks. Prevention, avoidance, detection with

recovery, and ignoring the problem.

  1. Distinguish prevention from avoidance. Prevention restricts how requests may be made so

that one of the four conditions can never hold. Avoidance permits any request and examines the state of the system before granting it, refusing any that could lead to an unsafe state.

  1. What extra information does avoidance need? The maximum number of instances of each

resource type that each process will ever need.

  1. Why do most operating systems ignore deadlocks? They are rare, the alternatives cost

resource utilisation on every request, and the consequence is that a person kills one stuck process rather than the machine failing.

  1. Give a system that must use prevention, and say why. A real time controller such as a

flight computer: a deadlock is a failure, there is nobody to intervene, and the guarantee has to be provable before the system runs.

  1. Which method does a database engine use, and what does recovery cost there? Detection and

recovery. It aborts one of the transactions in the cycle, so one transaction's work is lost and the client retries.

  1. What does every guarantee cost? Resource utilisation: something that could safely have

been granted is refused or delayed, on every request.

Contents This chapter on its own page

munotes.in232

Chapter Sixty

Deadlock Prevention

Syllabus topic Module 2, "Deadlocks - Deadlock Prevention"

In one line

Make one of the four conditions impossible and a deadlock can never form, and every way of doing it wastes something.

The method

Chapter fifty seven proved that all four conditions are necessary. So denying any one of them prevents deadlock absolutely, with no checking and no knowledge of the future. The question is only which one to deny and what that costs.

Attempt one: deny mutual exclusion

Make the resource sharable, so no process ever waits for it.

This works only where the resource genuinely can be shared, and for most resources it cannot. A read only file can be opened by everybody at once. A printer cannot print two documents at once, a mutex that let two processes in would not be a mutex, and a record being updated cannot be updated by two writers.

Where it worksread only data, and anything a reader writer lock covers
Costnothing, where it applies
Why it is not the answerit applies to almost nothing that matters

So mutual exclusion is normally not denied, and an answer should say so rather than listing it as a real option.

Attempt two: deny hold and wait

Guarantee that a process never holds one resource while waiting for another. Two ways.

Ask for everything at once. A process requests all the resources it will need before it starts, and is given all of them or none.

Release everything first. A process holding resources must release them all before requesting anything new, and then ask for the whole set again.

Gainedhold and wait is impossible, so no deadlock
Cost, onevery low utilisation: a job that needs the printer only at the end holds it from the start
Cost, twostarvation is possible: a process needing several popular resources may never find them all free at once
Cost, threea process must know its whole need in advance, which many cannot

The utilisation cost is the one to name. A job that copies a tape to a disk and then prints the result needs the tape drive, the disk and the printer. Under the first rule it holds all three for the whole job, although the printer is idle for almost all of it. Under the second rule it must release the disk to ask for the printer, and then it may lose the disk to somebody else.

Attempt three: deny no preemption

Allow a resource to be taken away from a process that is holding it.

The standard protocol: if a process holding some resources requests another that cannot be granted at once, all the resources it currently holds are preempted. They are added to the list of things it is waiting for, and it restarts only when it can have everything.

munotes.in233

Deadlock Prevention

Works well forresources whose state is easy to save and restore: the processor, memory, a register set
Does not work fora printer halfway through a page, a tape drive halfway through a reel, a mutex protecting a half finished update
Costwork is thrown away and repeated, and a process can be preempted repeatedly and starve

Whether a resource is preemptible is a property of the resource, not a choice. Memory is preemptible because its contents can be written to disk and read back. A printer is not, because ink on paper cannot be un-printed. That single distinction decides where this method can be used at all.

Chapter fifty seven's third variant showed the programmer's version of this: use trylock, and if it fails, release what you hold and start again. It works, and it risks the livelock of two processes releasing and grabbing for ever, which is why a real retry waits a random time.

Attempt four: deny circular wait

This is the one that is actually used, and it costs almost nothing.

Give every resource type a number, once, for the whole system, and require every process to request resources in increasing order of number. A process holding a resource numbered n may request only resources numbered above n; to get something lower it must first release everything from n upwards.

Why that makes a cycle impossible

Suppose a cycle existed: P0 waits for something P1 holds, P1 for something P2 holds, and so on back to P0. Take the numbers of the resources involved. Every process is waiting for a higher numbered resource than the one it holds, so going round the cycle the numbers must increase at every step. After a full circuit you arrive back where you started with a larger number than you began with, which is impossible. So the cycle cannot exist.

That argument is worth being able to write: it is a complete proof in four lines, and a question that asks why the ordering works wants exactly it.

What it costs

Utilisationalmost nothing: a process asks when it needs something, not in advance
Knowledge needednone about the future, only the fixed order
Costthe programmer must know the order and obey it, everywhere, in every piece of code
Where it failswhen a process genuinely needs resources in the other order and must release and re-acquire

The ordering is a discipline, not a mechanism, and that is its only weakness: nothing in the hardware enforces it. One function in one library that takes two locks the other way round is enough to bring the whole guarantee down, which is why the Linux kernel ships a run time checker that watches the order in which locks are taken and reports a violation before it ever deadlocks.

munotes.in234

Deadlock Prevention

Numbering, worked

A system with a tape drive, a disk drive and a printer, numbered in the order a job naturally uses them.

ResourceNumber
tape drive1
disk drive5
printer12

A job that copies tape to disk and prints requests 1, then 5, then 12: increasing, so it is legal.

A job that prints a header first and then reads the tape wants 12 then 1, which is not legal. It must either ask for 1 and 12 in that order at the start, or release the printer before asking for the tape.

Choosing the numbers is the real design work. They should follow the order in which resources are normally used, so that most processes ask in increasing order without having to think about it.

The four, priced

Condition deniedHowUtilisation costWorks for
Mutual exclusionmake the resource sharablenonealmost nothing: only sharable resources
Hold and waitrequest everything at once, or release all before askingvery highanything, if the total need is known
No preemptiontake resources back from a waiting holderhigh: work is repeatedonly resources whose state can be saved
Circular waitnumber the resources and request in orderalmost noneeverything

What it does not mean

Prevention does not need to know anything about the future. That is avoidance, in the next chapter. All four methods here are rules about the form of a request.

Denying mutual exclusion is not a way of ignoring locking. It applies only where the resource really is sharable.

Preemption of a resource is not preemption of the processor. The processor is preempted constantly and is never part of a deadlock for exactly that reason.

The ordering rule is not a trick for the dining philosophers. It is the general answer for taking several locks, and it is what to do in your own code: pick an order for your mutexes once, and take them in that order everywhere.

Quick revision

  • Prevention denies one of the four necessary conditions, so a deadlock cannot form. It needs

no knowledge of the future.

  • Mutual exclusion: deny it by making the resource sharable. Possible for read only data and

almost nothing else.

  • Hold and wait: request everything at once, or release everything before asking again.

Very low utilisation, and starvation is possible.

  • No preemption: take the resources back from a process that must wait. Only for resources

whose state can be saved; work is repeated, and a process may be preempted repeatedly.

munotes.in235

Deadlock Prevention

  • Circular wait: number the resource types and require requests in increasing order.

Almost free, works everywhere, and is what real systems use.

  • The proof: round a cycle the numbers must increase at every step, so a full circuit returns to

the start with a larger number, which is impossible.

  • The ordering is a discipline nothing enforces, which is why the Linux kernel checks it at

run time.

Test yourself

  1. What does deadlock prevention do, and what does it need to know? It denies one of the four

necessary conditions so that a deadlock cannot form. It needs nothing known in advance: the restriction is on the form of a request.

  1. Why is denying mutual exclusion rarely useful? Because most resources genuinely cannot be

shared: a printer, a tape drive or a lock protecting an update would stop working if two processes held it at once.

  1. Give the two ways to deny hold and wait and the cost of each. Request all the resources at

once before starting, which holds resources that are idle for most of the job; or release everything before requesting again, which risks losing what you had. Both give low utilisation and allow starvation.

  1. Which resources can no preemption be denied for? Those whose state can be saved and

restored, such as memory or the processor. Not a printer halfway through a page or a lock guarding a half finished update.

  1. State the circular wait prevention rule. Give every resource type a number and require

every process to request resources in increasing order of number, releasing anything at or above a number before requesting below it.

  1. Prove that the ordering rule prevents a cycle. In a cycle every process waits for a

resource numbered higher than the one it holds, so the numbers increase at every step. A full circuit would return to the starting resource with a higher number than itself, which is impossible.

  1. Which condition is cheapest to deny and what is its one weakness? Circular wait. Its

weakness is that the ordering is a discipline nothing enforces, so one piece of code that takes two locks the other way round destroys the guarantee.

Contents This chapter on its own page

munotes.in236

Chapter Sixty-One

Safe States, and Avoidance

Syllabus topic Module 2, "Deadlocks - Deadlock Avoidance"

In one line

A state is safe if the processes could all be finished in some order from where they stand, and avoidance means never leaving a safe state.

Why avoidance exists at all

Chapter sixty's prevention worked and cost utilisation on every single request. Avoidance buys the utilisation back by being cleverer, and it pays for that with one requirement.

Avoidance requires every process to declare, in advance, the maximum number of instances of each resource type it will ever need at once. With that declaration the system can look at any request and work out whether granting it could possibly lead to trouble.

That requirement is also why avoidance is almost never used on a general purpose machine: a text editor cannot say how many files the user will open. It is usable where jobs are submitted with a description of what they need, which is what a batch system or a print room has.

The four matrices

Every avoidance question is worked in these, and the third is computed, never given.

NameMeaning
Allocationhow many instances of each type each process holds now
Maximum, or Maxthe most it will ever need at once, declared in advance
NeedMaximum minus Allocation: what it might still ask for
Availablehow many instances of each type are unheld

Need is derived, and a question that gives you Need has done half the work for you. If it gives Max and Allocation, compute Need first and write it down. Nearly every mistake in this row is a subtraction in that step.

A negative Need is impossible: a process cannot hold more than its declared maximum. sim/deadlock.py refuses such a state rather than quietly negating it, because a question sometimes contains one as a trap.

What a safe state is

A state is safe if there exists a sequence of all the processes such that, taking them in that order, each one can be given everything it still needs from what is then available, finish, and return everything it holds.

Such a sequence is a safe sequence. If no safe sequence exists, the state is unsafe.

Read the definition again and notice what it does not say. It does not say the processes will run in that order. It does not say any particular process will ask for anything. It says only that a way out exists, and that is the whole of the guarantee: as long as one way out exists, the system can always be steered to completion however the processes behave.

Safe, unsafe and deadlocked

This is the most examinable idea in the chapter and the one to be precise about.

munotes.in237

Safe States, and Avoidance

SafeUnsafeDeadlocked
A way to finish everybody existsyesno guaranteeno
Deadlock nownonoyes
Deadlock possible laternoyesit has happened
Who decidesthe statethe statethe state

An unsafe state is not a deadlocked state. It is a state from which a deadlock can be reached if the processes make the wrong requests. They may not: a process might finish without asking for its maximum at all, and the system recovers. Unsafe means the operating system has lost control of the outcome, and that is precisely what avoidance refuses to allow.

The relationship, in one line to memorise: every deadlocked state is unsafe, and an unsafe state may or may not become deadlocked.

Why an unsafe state is refused anyway

A question sometimes asks why the system should refuse a request that might well have been harmless. The answer is that the operating system cannot control what the processes ask for next. Its only lever is which requests it grants. Once the state is unsafe, there is a sequence of perfectly legal requests that deadlocks the system, and nothing the operating system may then do will prevent it.

So avoidance is deliberately conservative: it refuses requests that would have been safe in hindsight. That lost utilisation is what the guarantee costs, and it is the answer to a question about the disadvantage of the banker's algorithm.

Worked example: twelve tape drives

The classic single resource example, and it shows safety and unsafety in six lines.

Twelve tape drives, three processes.

ProcessHolds nowMaximum needStill needs
P05105
P1242
P2297

Available: 12 minus 9 held, which is 3.

Is it safe? Try to find a sequence.

StepAvailableWho can finishAfter it returns everything
13P1, which needs 23 + 2 = 5
25P0, which needs 55 + 5 = 10
310P2, which needs 710 + 2 = 12

The sequence P1, P0, P2 works, so the state is safe.

Now change one thing. Suppose P2 asks for one more drive and is given it.

ProcessHolds nowMaximum needStill needs
P05105
P1242
P2396

Available: 12 minus 10, which is 2.

Try again. P1 needs 2 and can finish, leaving 2 + 2 = 4. Now P0 needs 5 and cannot, and P2 needs 6 and cannot. No sequence completes, so the state is unsafe, and the request for that one drive must be refused.

P2 asked for one drive that was sitting there free and was refused. Nothing was wrong at that instant and no deadlock existed. The refusal is the operating system keeping the one thing it controls, which is its ability to guarantee a way out.

munotes.in238

Safe States, and Avoidance

The claim graph, for single instances

For resource types with one instance each, the resource allocation graph of Chapter fifty eight can do the avoidance job with one addition.

EdgeDrawn asMeans
Assignmenta solid arrow from resource to processP holds it
Requesta solid arrow from process to resourceP is waiting for it
Claima dashed arrow from process to resourceP may request it at some time in the future

The rules:

  1. A process must declare all its claim edges before it starts.
  2. When a process actually requests the resource, the claim edge becomes a request edge.
  3. When it is granted, the request edge becomes an assignment edge.
  4. When it is released, the assignment edge becomes a claim edge again, because it may ask

once more.

A request is granted only if turning the request edge into an assignment edge leaves no cycle, counting the dashed claim edges as if they were real. A cycle in that graph means an unsafe state; with one instance per type, an unsafe state means a deadlock is reachable.

The cost of this method is a cycle detection run on every request, which is the same order of work as the banker's algorithm, and it works only for single instance types. For several instances the banker's algorithm of the next chapter is used instead.

What it does not mean

Safe does not mean no process will ever wait. Processes wait constantly in a safe state. It means nobody waits for ever.

Unsafe does not mean deadlocked. It means the operating system can no longer guarantee an outcome. The processes may still all finish.

A safe sequence is not a schedule. Nothing will be run in that order. It is a proof that an order exists.

Avoidance is not detection. Detection looks at what has happened; avoidance decides before anything happens, on every request.

Quick revision

  • Avoidance grants a request only if the resulting state is safe, and needs every

process's maximum need declared in advance.

  • Four matrices: Allocation, Maximum, Need = Maximum minus Allocation, and

Available. Need is always computed.

  • A state is safe if a safe sequence exists: an order in which each process can be given

all it still needs, finish, and return everything.

  • Unsafe is not deadlocked. Every deadlocked state is unsafe; an unsafe state may never

deadlock. Unsafe means the operating system has lost control of the outcome.

  • Avoidance is deliberately conservative, and the utilisation it gives up is what the
munotes.in239

Safe States, and Avoidance

guarantee costs.

  • Twelve tape drives with P0 holding 5 of 10, P1 holding 2 of 4 and P2 holding 2 of 9 is safe, by

P1, P0, P2. Grant P2 one more and it becomes unsafe.

  • For single instance types, a claim graph with dashed future request edges does the same

job: grant only if no cycle results.

Test yourself

  1. Define a safe state. A state is safe if there is a sequence of all the processes such that

each one in turn can be granted everything it still needs from what is then available, complete, and release everything it holds.

  1. What extra information does avoidance need? The maximum number of instances of each

resource type that every process will ever need at one time.

  1. How is Need obtained? Need equals Maximum minus Allocation, computed per process and per

resource type.

  1. Is an unsafe state a deadlocked state? No. It is a state from which a deadlock is

reachable if the processes make particular legal requests. Every deadlocked state is unsafe, but not the reverse.

  1. Why refuse a request when the resources are free and no deadlock exists? Because the

operating system cannot control what the processes request next. Once the state is unsafe there is a legal sequence of requests that deadlocks it, and nothing can then prevent that.

  1. Twelve drives; P0 holds 5 of a maximum 10, P1 holds 2 of 4, P2 holds 2 of 9. Is it safe?

Yes. Three are available; P1 needs 2 and finishes leaving 5; P0 needs 5 and finishes leaving 10; P2 needs 7 and finishes. The safe sequence is P1, P0, P2.

  1. What is a claim edge and when is a request granted in a claim graph? A dashed edge from a

process to a resource it may request in future. A request is granted only if converting its request edge to an assignment edge leaves the graph without a cycle, counting claim edges as real.

Contents This chapter on its own page

munotes.in240

Chapter Sixty-Two

The Banker's Algorithm

Syllabus topic Module 2, "Deadlocks - Deadlock Avoidance"

In one line

The banker's algorithm decides whether a state is safe by repeatedly finding a process that can finish with what is available, pretending it has finished, and seeing whether everybody can be got through that way.

Where the name comes from

A banker with a fixed amount of cash lends to customers who have each declared a credit limit. A customer may draw up to that limit, and will eventually repay everything. The banker must never lend so much that some customer could ask for the rest of their limit and be refused, leaving everybody waiting.

The algorithm is Dijkstra's, from 1965. His own note on it is in authorities/papers/ as a scan of his handwriting, in Dutch, with no text layer, so authorities/sources.json records it as a source this book may name and may not quote.

The banker analogy is worth carrying because it explains the one thing students find strange: the banker refuses to lend money it has, to a customer within their limit, because of what a different customer might ask for later. That is exactly what Chapter sixty one's conservatism is.

The safety algorithm, step by step

This is the part to be able to write out.

Let there be n processes and m resource types.

  1. Let Work be a copy of Available, and Finish be an array of n falses.
  2. Find an i such that Finish[i] is false and Need[i] is less than or equal to Work, in

every resource type. If there is no such i, go to step 4.

  1. Set Work = Work + Allocation[i], set Finish[i] = true, record i in the sequence, and

go back to step 2.

  1. If Finish[i] is true for every i, the state is safe and the recorded order is a safe

sequence. Otherwise it is unsafe.

Step 3 is the step that is misunderstood. Adding Allocation[i] to Work is the algorithm pretending that process i has been given everything it needs, has finished, and has returned everything it holds. No process actually runs. Nothing is granted. The whole algorithm is a thought experiment on paper.

Its cost is m times n squared operations, because in the worst case each of the n passes searches all n processes over m resource types. That figure is sometimes asked for.

A safe sequence is not unique. Step 2 says "find an i", not "find the smallest i". The convention in this book, and the one a teacher expects, is to take the lowest numbered process that qualifies at each step, so that the answer is reproducible. sim/deadlock.py follows it, and it can also count how many safe sequences exist altogether.

munotes.in241

The Banker's Algorithm

The standard problem, worked in full

Five processes, three resource types A, B and C, with 10, 5 and 7 instances in all.

Banker: available 3 3 2

ProcessAllocationMaximumNeed
P00 1 07 5 37 4 3
P12 0 03 2 21 2 2
P23 0 29 0 26 0 0
P32 1 12 2 20 1 1
P40 0 24 3 34 3 1

Resource types: A, B, C

StepProcessNeedWork beforeWork after
1P11 2 23 3 25 3 2
2P30 1 15 3 27 4 3
3P07 4 37 4 37 5 3
4P26 0 07 5 310 5 5
5P44 3 110 5 510 5 7

The state: safe

Safe sequence: P1, P3, P0, P2, P4

(16 safe sequences exist in all)

Read the work down the last two columns. Work starts at Available and grows by exactly the Allocation of each process as that process is pretended to finish. By step 5 it has reached 10 5 7, which is every instance of every type, and that is the proof that everybody got through.

Check the arithmetic on any row: at step 2, P3's Need is 0 1 1 and Work is 5 3 2, which covers it, so P3 qualifies; adding P3's Allocation of 2 1 1 gives 7 4 3.

And check step 3, which is the interesting one. P0's Need is 7 4 3 and Work is exactly 7 4 3. Equal counts as satisfied: Need must be less than or equal to Work. A student who requires strict inequality gets stuck at this step and concludes the state is unsafe.

Why the total matters

The three totals, 10, 5 and 7, never appear in the algorithm and they are worth knowing anyway: they are the check on the question. For each resource type,

total = Available + the sum of that column of Allocation

total of A = 3 + (0 + 2 + 3 + 2 + 0) = 3 + 7 = 10

total of B = 3 + (1 + 0 + 0 + 1 + 0) = 3 + 2 = 5

total of C = 2 + (0 + 0 + 2 + 1 + 2) = 2 + 5 = 7

If those do not come out to the totals the question gives, the question has been copied down wrongly, and that is worth thirty seconds before starting.

munotes.in242

The Banker's Algorithm

How many safe sequences are there

The algorithm finds one. There are sixteen for this state, and a question that asks for "a" safe sequence will accept any of them.

So an answer should say a safe sequence, not the safe sequence, and a student whose answer differs from the printed one is not necessarily wrong. What makes an answer right is that every step of it is justified: at each step the chosen process's Need really is within the Work at that moment.

An unsafe state, for comparison

Take the same five processes and give the system one instance of A spare instead of three, so Available is 1 3 2.

StepProcessNeedWork beforeWork after
1P11 2 21 3 23 3 2
2P30 1 13 3 25 4 3
3P44 3 15 4 35 4 5

And there it stops. P0 needs 7 4 3 and Work is 5 4 5: the A column is short. P2 needs 6 0 0 and A is short again. Finish is false for P0 and P2, so the state is unsafe.

Notice that P4 did get through here although it did not in the safe version's ordering, and the state is still unsafe. The order the algorithm happens to try is not what makes a state safe: what makes it safe is whether any order finishes everybody.

What it does not mean

The banker's algorithm does not detect deadlock. It is asked before granting a request, and its answer is about the future. Chapter sixty four detects one that has already happened.

It does not run the processes. Nothing is scheduled and nothing is granted by it. It answers one question: is this state safe?

A safe sequence is not a plan. The processes will run in whatever order the scheduler chooses. The sequence is only evidence that an order exists.

Need is not a request. Need is what a process might still ask for in total. What it is asking for now is the next chapter.

Quick revision

  • The banker's algorithm decides whether a state is safe, using Allocation,

Maximum, Need = Maximum minus Allocation and Available.

  • The safety algorithm: copy Available into Work; repeatedly find an unfinished process whose

Need is less than or equal to Work, add its Allocation to Work and mark it finished; if all finish, the state is safe.

  • Adding Allocation is pretending the process finished and returned everything. Nothing is

really granted.

  • Need less than or equal to Work: equality counts, and P0's step in the worked example

depends on it.

  • A safe sequence is not unique: this state has sixteen. Say "a safe sequence".
  • Cost: m times n squared operations.
  • Check the question first: Available plus each Allocation column must equal the total of
munotes.in243

The Banker's Algorithm

that resource type.

  • Work ends at the total of every resource type when the state is safe.

Test yourself

  1. What question does the banker's algorithm answer? Whether the current state is safe:

whether some order exists in which every process can be given all it still needs, finish, and return everything.

  1. Write the safety algorithm. Copy Available into Work and set every Finish to false.

Repeatedly find a process not yet finished whose Need is less than or equal to Work; add its Allocation to Work, mark it finished, and record it. If every process is finished the state is safe and the recorded order is a safe sequence; otherwise it is unsafe.

  1. What does adding Allocation to Work represent? Pretending that the chosen process has been

given everything it needs, has completed, and has released all the resources it held.

  1. Does Need have to be strictly less than Work? No. Need less than or equal to Work is

enough, and the worked example depends on it: P0's Need of 7 4 3 is satisfied by a Work of exactly 7 4 3.

  1. Is the safe sequence unique? No. The worked state has sixteen safe sequences, and any of

them is a correct answer provided each step is justified.

  1. How would you check that a question has been copied correctly? For each resource type,

Available plus the sum of that column of Allocation must equal the total number of instances of that type.

  1. What is the running cost of the algorithm? Of the order of m times n squared, for n

processes and m resource types.

Contents This chapter on its own page

munotes.in244

Chapter Sixty-Three

The Resource Request Algorithm

Syllabus topic Module 2, "Deadlocks - Deadlock Avoidance"

In one line

When a process asks for resources, check that it is within its declared maximum, check that the resources exist, pretend to grant it, and refuse unless the pretended state is still safe.

The algorithm

Let Request[i] be what process i is asking for now.

  1. If Request[i] is greater than Need[i] in any resource type, the process has asked for more

than it declared it would ever need. This is an error, not a wait: the process has broken its own promise and is killed or reported.

  1. If Request[i] is greater than Available in any resource type, the resources do not exist

to give. The process waits.

  1. Otherwise, pretend to grant it:
  • Available = Available minus Request[i]
  • Allocation[i] = Allocation[i] plus Request[i]
  • Need[i] = Need[i] minus Request[i]
  1. Run the safety algorithm of Chapter sixty two on that pretended state. If it is safe,

the change is made for real and the process gets its resources. If it is unsafe, the pretended change is undone and the process waits.

The three outcomes are different and a question wants them told apart: an error, a wait because nothing is available, and a wait although everything asked for is available. Only the third is the banker doing anything interesting.

Notice step 3's third line. Granting a request reduces Need as well as raising Allocation, because the process has now got part of what it might have asked for. Forgetting that line is a common slip and it makes every later safety check wrong.

Three requests on one state

The state is Chapter sixty two's: five processes, A B C with 10 5 7 instances, Available 3 3 2.

ProcessAllocationMaximumNeed
P00 1 07 5 37 4 3
P12 0 03 2 21 2 2
P23 0 29 0 26 0 0
P32 1 12 2 20 1 1
P40 0 24 3 34 3 1

Request one: P1 asks for 1 0 2, and is granted

Test 1. Request 1 0 2 against P1's Need of 1 2 2: 1 is at most 1, 0 is at most 2, 2 is at most

  1. Within its maximum, so not an error.

Test 2. Request 1 0 2 against Available 3 3 2: all three fit. So the resources exist.

Pretend to grant it. Available becomes 3 3 2 minus 1 0 2, which is 2 3 0. P1's Allocation becomes 2 0 0 plus 1 0 2, which is 3 0 2. P1's Need becomes 1 2 2 minus 1 0 2, which is 0 2 0.

munotes.in245

The Resource Request Algorithm

Test 3, the safety algorithm on the pretended state:

StepProcessNeedWork beforeWork after
1P10 2 02 3 05 3 2
2P30 1 15 3 27 4 3
3P07 4 37 4 37 5 3
4P26 0 07 5 310 5 5
5P44 3 110 5 510 5 7

Everybody finishes, so the state is safe and the request is granted. A safe sequence is P1, P3, P0, P2, P4.

Request two: P4 asks for 3 3 0, and waits because nothing is available

This is asked in the state after request one, where Available is 2 3 0.

Test 1. Request 3 3 0 against P4's Need of 4 3 1: within its maximum.

Test 2. Request 3 3 0 against Available 2 3 0: 3 is more than 2 in resource A. So the resources are not there and P4 waits.

The safety algorithm is never reached. There is nothing to test: you cannot give away what you do not have.

Request three: P0 asks for 0 2 0, and waits although the resources are free

Also in the state after request one, where Available is 2 3 0.

Test 1. Request 0 2 0 against P0's Need of 7 4 3: within its maximum.

Test 2. Request 0 2 0 against Available 2 3 0: 0 is at most 2, 2 is at most 3, 0 is at most 0. Everything asked for is available.

Pretend to grant it. Available becomes 2 1 0. P0's Allocation becomes 0 3 0, and P0's Need becomes 7 2 3.

Test 3, the safety algorithm:

ProcessNeedIs Need within Work of 2 1 0?
P07 2 3no: A and C are short
P10 2 0no: B is short, 2 is more than 1
P26 0 0no: A is short
P30 1 1no: C is short, 1 is more than 0
P44 3 1no: all three short

No process can finish, so the state is unsafe, the pretended grant is undone, and P0 waits.

Read that again: two units of B were sitting there unheld and P0 was refused them. Nothing was wrong, nobody was deadlocked, and the request was legal. It was refused because after granting it there would have been no order at all in which the five processes could be finished, and from that point a perfectly legal sequence of later requests would have deadlocked the system.

That is the banker's algorithm doing the only thing it does, and it is the answer to "what is the disadvantage of deadlock avoidance": resources go unused and processes wait when they need not have, because the guarantee is worth more than the utilisation.

munotes.in246

The Resource Request Algorithm

And a fourth: P0 asks for 8 0 0

Test 1. Request 8 0 0 against P0's Need of 7 4 3: 8 is more than 7. P0 has asked for more than it declared it would ever need.

This is an error and not a wait. The process broke the declaration the whole algorithm rests on, so nothing can be guaranteed about it and the system raises an error. A student who answers "it waits" has missed the point of test 1.

The three outcomes, side by side

OutcomeWhich test failedWhat happens to the process
Errortest 1: Request exceeds Needit is reported or killed: it broke its declaration
Waittest 2: Request exceeds Availableit waits until the resources exist
Waittest 3: the resulting state is unsafeit waits although the resources are free
GrantednoneAllocation, Need and Available are updated for real

What it does not mean

A refusal is not a deadlock. The process waits, and it will be granted the request later, when somebody releases something and the state can take it.

The pretended grant is not a grant. If the safety test fails, every one of the three lines in step 3 is undone. A question that asks for the state after a refused request wants the original state.

Test 2 failing does not mean the system is short of resources overall. It means they are held right now.

The algorithm does not choose which process to serve. It answers yes or no to one request. Which waiting process is tried next is a scheduling decision outside it.

Quick revision

  • Request[i] is what process i wants now. Three tests, in order.
  • Test 1: Request greater than Need is an error: the process exceeded its declared

maximum.

  • Test 2: Request greater than Available means it waits: the resources are not there.
  • Test 3: pretend to grant it, with Available minus Request, Allocation plus Request

and Need minus Request, then run the safety algorithm. Unsafe means undo it and wait.

  • Granting reduces Need as well as raising Allocation. Forgetting that makes every later

check wrong.

  • Worked on the standard state: P1 asking 1 0 2 is granted; P4 asking 3 3 0 waits because A is

short; P0 asking 0 2 0 waits although both units of B are free, because the resulting state is unsafe; P0 asking 8 0 0 is an error.

munotes.in247

The Resource Request Algorithm

  • The third case is the whole point of avoidance, and the lost utilisation is what the guarantee

costs.

Test yourself

  1. State the three tests of the resource request algorithm. Request greater than Need is an

error; Request greater than Available means the process waits; otherwise pretend to grant it and run the safety algorithm, granting only if the resulting state is safe.

  1. What three quantities change when a request is pretended? Available falls by the request,

the process's Allocation rises by it, and the process's Need falls by it.

  1. What happens if a process asks for more than its declared maximum? It is an error, not a

wait: the process has broken the declaration the algorithm depends on and is reported or killed.

  1. P0 asks for 0 2 0, two units of B are free, and the request is refused. Why? Because after

granting it no process's Need would be within the remaining Available, so no order could finish them all: the state would be unsafe, and a legal sequence of later requests could then deadlock the system.

  1. What is the state after a request is refused by the safety test? Exactly what it was

before. The pretended grant is undone in all three quantities.

  1. When is the safety algorithm not run at all? When test 1 or test 2 fails: an error needs

no safety check, and resources that are not available cannot be granted.

  1. What is the disadvantage of avoidance that this chapter demonstrates? Resources sit unused

and processes wait when they need not have. The guarantee is paid for in utilisation.

Contents This chapter on its own page

munotes.in248

Chapter Sixty-Four

Deadlock Detection

Syllabus topic Module 2, "Deadlocks - Deadlock Detection"

In one line

Let deadlocks happen, and run an algorithm every so often that looks at the current state and reports which processes, if any, are stuck.

Why a system would choose this

Prevention and avoidance both cost utilisation all the time, whether or not anything ever goes wrong. Detection costs nothing until it is run, and what it costs then is one algorithm and, if it finds something, somebody's work.

That trade is why database engines detect rather than prevent: transactions deadlock often enough to matter and cheaply enough to abort, and no transaction can declare its maximum needs in advance.

Two cases, two algorithms

Which algorithm applies depends on whether every resource type has one instance. A question that gives you instance counts is telling you which method it wants.

Every type hasUseCost
one instancethe wait for graph: look for a cycleof the order of n squared
several instancesthe detection algorithm: like the banker's, with Request instead of Needof the order of m times n squared

Case one: the wait for graph

The textbooks hyphenate it, as the wait-for graph, and it is worth recognising in that spelling because that is how a question will print it. Take the resource allocation graph of Chapter fifty eight and collapse the resource vertices out of it. What is left is processes only, with an edge from P to Q meaning P is waiting for a resource that Q holds.

To build it: for every pair of edges P to R and R to Q in the resource allocation graph, draw an edge P to Q, and then remove the resource vertices.

Example. The resource allocation graph has P1 waiting for R1 which P2 holds, P2 waiting for R2 which P3 holds, and P3 waiting for R3 which P1 holds. Collapsed, the wait for graph is:

FromToMeans
P1P2P1 waits for something P2 holds
P2P3P2 waits for something P3 holds
P3P1P3 waits for something P1 holds

The cycle is P1, P2, P3, P1, found by sim/deadlock.py rather than by eye, and with one instance of every type a cycle in the wait for graph is a deadlock.

A cycle in the WAIT FOR graph is sufficient for deadlock; a cycle in the RESOURCE ALLOCATION graph is not. That is because the wait for graph can only be built at all when every type has one instance, so the ambiguity Chapter fifty eight described cannot arise. Getting those two sentences the right way round is worth marks.

To use this, the system must maintain the wait for graph and search it for a cycle periodically. The search is of the order of n squared in the number of processes.

munotes.in249

Deadlock Detection

Case two: the detection algorithm

For several instances, the same shape as Chapter sixty two's safety algorithm with one difference that is the whole point of this chapter.

MatrixThe banker usesDetection uses
AllocationAllocationAllocation
The demandNeed: what the process might ask for in totalRequest: what it is asking for right now
AvailableAvailableAvailable

Need is the future and Request is the present, and swapping them is the commonest error in this row. The banker asks whether the processes could all finish if they each asked for everything they are entitled to. Detection asks whether they can finish given only what they are actually waiting for. Detection is therefore optimistic where the banker is pessimistic, which is right: the banker is preventing something, and detection is reporting something that has already happened or has not.

The algorithm:

  1. Let Work be a copy of Available. For each i, set

Finish[i] = false if Allocation[i] is not all zero, and true otherwise.

  1. Find an i with Finish[i] false and Request[i] less than or equal to Work. If there is

none, go to 4.

  1. Work = Work + Allocation[i], Finish[i] = true, and go back to 2.
  2. If Finish[i] is false for some i, the system is deadlocked, and those processes are

the deadlocked ones.

Step 1's second half is worth understanding rather than memorising: a process holding nothing cannot be part of a deadlock, because nobody can be waiting for anything it has. So it is marked finished at the start and takes no further part.

Worked: no deadlock

Five processes, three types A B C with 7 2 6 instances, Available 0 0 0.

ProcessAllocationRequest
P00 1 00 0 0
P12 0 02 0 2
P23 0 30 0 0
P32 1 11 0 0
P40 0 20 0 2

Work starts at 0 0 0. Every process holds something, so no Finish starts true.

StepProcessRequestWork beforeWork after
1P00 0 00 0 00 1 0
2P20 0 00 1 03 1 3
3P12 0 23 1 35 1 3
4P31 0 05 1 37 2 4
5P40 0 27 2 47 2 6

Every process finishes, so there is no deadlock, and the sequence found is P0, P2, P1, P3, P4.

Notice how it starts. Nothing at all is available, and yet the algorithm gets going, because P0 and P2 are requesting nothing. A process that is not waiting for anything will finish, and what it releases is what frees everybody else. That is the engine of the whole algorithm.

munotes.in250

Deadlock Detection

Worked: a deadlock, after one more request

Change one number: P2 now requests one instance of C.

ProcessAllocationRequest
P00 1 00 0 0
P12 0 02 0 2
P23 0 30 0 1
P32 1 11 0 0
P40 0 20 0 2
StepProcessRequestWork beforeWork after
1P00 0 00 0 00 1 0

And it stops. Work is 0 1 0, and every remaining request needs A or C, of which there are none.

Finish is false for P1, P2, P3 and P4, so those four processes are deadlocked. P0 is not: it was requesting nothing and finished.

One instance of one resource type turned a healthy system into a four process deadlock, and nothing else changed. That is worth a sentence in any answer about why deadlock is hard to test for: the state one request before a deadlock looks entirely healthy.

When to run it, which is a real decision

Detection is only half the method. How often to run it is the other half, and a question asks for the trade.

Run itCostWhat you get
On every request that must waitvery high: it is the most expensive thing the kernel doesyou know exactly which request closed the cycle, so the victim is obvious
Periodically, say once an hourlowseveral processes may be in the cycle by then, so choosing a victim is harder
When processor utilisation drops below some level, say 40 per centlow, and self triggeringdeadlocked processes do not compute, so a deadlock shows up as an idle machine

The third is the clever one and it is worth naming: a deadlock makes the machine idle, because the stuck processes are waiting rather than computing. So low utilisation with a long ready queue is itself the signal.

The cost of running it rarely is not only a harder choice of victim. A cycle can grow: once two processes are stuck holding resources, a third that wants one of those resources joins the deadlock, then a fourth. Running detection an hour later may find twenty processes where there were two.

Distinctions that carry marks

The banker's safety algorithmThe detection algorithm
UsesNeedRequest
Askscould they all finish if each asked for its maximumcan they all finish given what they are asking now
Runbefore granting a requestafter the fact, periodically
Answersafe or unsafedeadlocked or not, and who
Processes holding nothingtake part normallymarked finished at the start
munotes.in251

Deadlock Detection

Resource allocation graph cycleWait for graph cycle
Verticesprocesses and resourcesprocesses only
A cycle meansdeadlock only if one instance per typedeadlock, and the graph exists only in that case
Cost to searchlarger graphof the order of n squared

What it does not mean

Detection does not prevent anything. The deadlock has already happened when the algorithm finds it.

Request is not Need. Request is what a process is waiting for at this instant; Need is what it might ask for in total. Using Need here would report deadlocks that do not exist.

A process holding nothing is not idle. It may be computing happily. It is excluded because nobody can be waiting for it.

Finding a deadlock is not fixing it. The next chapter is the fixing, and it always costs somebody their work.

Quick revision

  • Detection lets deadlocks happen and looks for them. It costs nothing until it runs.
  • One instance per type: build the wait for graph by collapsing the resources out, and

look for a cycle. A cycle there is a deadlock. Cost of the order of n squared.

  • Several instances: the detection algorithm, which is the safety algorithm with

Request in place of Need. Cost of the order of m times n squared.

  • Need is the future, Request is the present. Swapping them is the error this row is full of.
  • A process holding nothing is marked finished at the start: nobody can be waiting for it.
  • Worked here: a five process state with nothing available is not deadlocked, by P0, P2, P1,

P3, P4; add one request for one instance of C and four of the five are deadlocked.

  • When to run it: on every waiting request, which is too expensive; periodically; or when

processor utilisation drops, because deadlocked processes do not compute.

  • Running it rarely makes the cycle bigger and the victim harder to choose.

Test yourself

  1. What are the two detection methods and when is each used? The wait for graph with a cycle

search, when every resource type has one instance; and the detection algorithm, when types have several instances.

  1. How is a wait for graph built? By removing the resource vertices from the resource

allocation graph: for every P to R and R to Q, draw P to Q.

  1. Is a cycle in a wait for graph sufficient for deadlock? Yes. The graph can only be built

when every type has one instance, so the ambiguity of the resource allocation graph does not arise. 4. What is the one difference between the detection algorithm and the banker's safety algorithm? It uses Request, what each process is waiting for now, instead of Need, what it might eventually ask for.

munotes.in252

Deadlock Detection

  1. Why is a process holding no resources marked finished at the start? Because no other

process can be waiting for anything it holds, so it cannot be part of a deadlock.

  1. Give three policies for when to run detection, with a cost of each. On every request that

must wait, which is very expensive; periodically, which lets the cycle grow before it is found; or when processor utilisation falls below a threshold, which is cheap and self triggering because deadlocked processes do not compute.

  1. Why does running detection rarely make recovery harder? A cycle grows: processes that want

the resources held by the stuck ones join the deadlock, so an hour later there may be twenty processes in it rather than two, and choosing a victim is harder.

Contents This chapter on its own page

munotes.in253

Chapter Sixty-Five

Recovery from Deadlock

Syllabus topic Module 2, "Deadlocks - Recovery from Deadlock"

In one line

A deadlock is broken by ending a process or by taking a resource away from one, and either way somebody loses work.

The three ways out

WayWhat happensWho decides
Tell somebodyreport it and let a person deal with itthe operator, manually
Process terminationend one or more of the deadlocked processesthe system
Resource preemptiontake resources from a process and give them to anotherthe system

The first is what a general purpose operating system does, because Chapter fifty nine said it ignores deadlocks: the person notices the machine is stuck and kills something. The other two are what a system that detects automatically must then do.

There is no fourth way, and in particular there is no way that costs nothing. Every recovery destroys work that had already been done. That is the price of having let the deadlock happen, and a question that asks what is wrong with detection wants exactly this.

Process termination

Two policies, and the trade between them is examinable.

Abort all deadlocked processesAbort one at a time
Effectthe deadlock is certainly brokenit may be broken; check again after each one
Work lostall of it, from every process in the cycleonly the victim's, and only as much as needed
Cost of the methodone actiona detection run after every abort
Used whenthe cycle is large or time is criticalwork is expensive to repeat

The second needs the detection algorithm run again after each abort, because killing one process may free enough resources to release the rest, or may not. That repeated running is its cost.

Which victim

The system must choose, and a question asks for the factors. Six, and they are a list to be able to reproduce.

  1. The priority of the process.
  2. How long it has computed, and how much longer it needs to finish.
  3. How many and what kind of resources it holds: killing a process that holds the thing

everybody wants breaks more cycles.

  1. How many more resources it will need to complete.
  2. How many processes will need to be terminated if this one is chosen.
  3. Whether it is interactive or batch. A batch job can be resubmitted without a person

noticing; an interactive one has a user in front of it.

The principle underneath the list: choose the victim whose loss is cheapest, not the one that is easiest to kill. A process three hours into a four hour computation is a terrible victim however low its priority.

And the part students forget

Killing a process may leave the system inconsistent, and that is worse than the deadlock. If the victim was halfway through updating a file, or held a lock protecting a data structure it had half rewritten, the file or the structure is now broken and nothing will notice until much later.

munotes.in254

Recovery from Deadlock

That is why a database engine can recover safely and a general program often cannot: a database wraps the work in a transaction it can roll back exactly. A process holding a mutex over a half finished update has no such record, and killing it is a decision to corrupt something.

Resource preemption

Take resources from processes, one at a time, and give them to others until the cycle is broken. Three questions must be answered, and they are the standard three.

1. Selecting a victim

Which resources from which processes? The same cost calculation as above: the number of resources a deadlocked process holds and the amount of time it has consumed.

2. Rollback

A process that has had a resource taken away cannot continue. It was using that resource; without it, it is not in a valid state. So it must be rolled back to some safe state and restarted from there.

Rollback toMeans
Total rollbackabort the process and start it again from the beginning
Partial rollbackreturn it to a checkpoint, a recorded earlier state, and restart from there

Partial rollback is cheaper and needs the system to have been saving checkpoints all along, which costs time and space on every process whether or not a deadlock ever happens. Total rollback needs nothing kept and throws everything away. That is the same shape of trade as everything else in this module.

3. Starvation

The same process can be chosen as the victim every time. If the cost calculation says it is the cheapest victim now, it will say so again in an hour, and that process is rolled back for ever and never completes. That is starvation, caused by the recovery rather than by the deadlock.

The fix is to include the number of rollbacks in the cost. A process that has been rolled back five times is expensive to roll back again, so eventually it stops being the cheapest victim and gets through. Counting the rollbacks is the same idea as Chapter fifty's ageing, applied to victim selection.

Worked example: a print room deadlock, resolved four ways

Four jobs are deadlocked over a printer, a plotter and a scanner. Job A is three hours into a four hour render. Job B started a minute ago. Job C is an interactive session with a person waiting. Job D is a batch job submitted overnight.

ChoiceWhat it costsVerdict
Abort all fourthree hours of A, plus everything elsebreaks the deadlock certainly, and wastes the most
Abort A, the lowest prioritythree hourswrong: cheapest to kill, dearest to lose
Abort D, the overnight batch joba resubmission nobody seesbest: it can be run again tonight
Preempt the plotter from A, with a checkpoint at the last hourone hour of Agood, and only if checkpoints were being saved
munotes.in255

Recovery from Deadlock

The right answer changes with one fact. If no checkpoints were being kept, the last row costs three hours rather than one, and aborting D wins again. What makes recovery a design question rather than an algorithm is that the costs are outside the operating system.

Distinctions that carry marks

Process terminationResource preemption
The processendscontinues, after a rollback
Work lostall of the victim'sback to a checkpoint, or all of it
Needsnothing kept in advancecheckpoints, to be cheap
Riskan inconsistent file or data structurestarvation of the repeated victim
Total rollbackPartial rollback
Restart fromthe beginningthe last checkpoint
Needsnothingcheckpoints saved throughout
Costseverything the process had donethe work since the checkpoint
Abort allAbort one at a time
Deadlock brokencertainlyperhaps: check again
Work losteverybody'sthe minimum
Extra costnonea detection run after each abort

What it does not mean

Recovery is not a way of avoiding the cost of a deadlock. It is a way of paying it in the cheapest currency available.

Killing a process is not always safe. If it held a lock over a half finished update, the data is now wrong, and that can be worse than leaving the deadlock in place.

Preemption is not possible for every resource. Chapter sixty's rule stands: the resource must be one whose state can be saved and restored.

Choosing the lowest priority process is not choosing the cheapest victim. Time already spent, resources held and whether a person is waiting all matter more.

Quick revision

  • Three ways out: tell somebody, process termination, resource preemption. None is

free.

  • Termination: abort all the deadlocked processes, which certainly works and wastes the most;

or abort one at a time, running detection again after each, which wastes least and costs the repeated detection.

  • Choosing a victim: priority, time already spent and time still needed,

resources held, resources still needed, how many processes must die, and interactive or batch.

  • Killing a process that held a lock over a half finished update leaves the data

inconsistent, which can be worse than the deadlock. A database can recover because its work is in a transaction.

  • Preemption raises three questions: which victim, how far to roll back (total to the
munotes.in256

Recovery from Deadlock

beginning, or partial to a checkpoint), and how to avoid starvation.

  • The starvation fix: count the rollbacks in the cost, so a repeated victim stops being the

cheapest one.

Test yourself

  1. Name the two automatic ways of recovering from a deadlock. Process termination and

resource preemption.

  1. Compare aborting all the deadlocked processes with aborting one at a time. Aborting all

certainly breaks the deadlock and destroys the most work. Aborting one at a time destroys the least, and costs a detection run after every abort to see whether the deadlock is gone.

  1. List the factors in choosing a victim. Its priority; how long it has computed and how much

longer it needs; how many and what kind of resources it holds; how many more it will need; how many processes must be terminated if it is chosen; and whether it is interactive or batch.

  1. Why can killing a process be worse than the deadlock? Because it may have held a lock over

a half finished update, so a file or a data structure is left inconsistent and nothing notices until much later.

  1. What are the three questions resource preemption raises? Which victim to take resources

from; how far to roll the victim back; and how to stop the same process being chosen every time.

  1. Distinguish total from partial rollback. Total rollback aborts the process and restarts it

from the beginning and needs nothing saved. Partial rollback returns it to a checkpoint and needs checkpoints to have been recorded all along.

  1. How is starvation in recovery prevented? By including the number of times a process has

already been rolled back in the cost of choosing it, so a repeated victim eventually stops being the cheapest.

Contents This chapter on its own page

munotes.in257

Chapter Sixty-Six

Why Memory Needs Managing

Syllabus topic Module 2, "Memory Management - Main memory background"

In one line

Main memory is the only storage the processor can reach directly, there is never enough of it, and two programs in it must not be able to touch each other.

What the hardware gives, and what it does not

The processor can do exactly two things with memory: read a word at an address, and write a word at an address. That is the whole interface.

There is no instruction that asks whose memory it is. A load from address 5000 loads address

  1. The hardware has no idea which process is running and no opinion about whether it should be

allowed. So protection cannot come from the instruction: it must come from something that sits between the address the program uses and the address the memory sees.

That something is the memory management unit of Chapter sixty eight, and everything in this row of the syllabus is about what it does.

The three things memory management must achieve

GoalMeans
Protectionone process must not read or write another's memory, or the kernel's
Relocationa program must run wherever it is put, because where it is put depends on what else is in memory
Sharingtwo processes must be able to share memory deliberately, which is Chapter twenty two of Module 1

Protection and sharing sound like opposites and they are the same mechanism used two ways: whatever can stop an address reaching another process's memory can also let it reach memory that has been shared on purpose.

What a memory access costs

The numbers matter because they explain the rest of the row.

Roughly
One processor instructionwell under a nanosecond
One access to main memorytens of nanoseconds, so tens of instructions' worth
One access to a diskmilliseconds, so millions of instructions' worth

Main memory is slow compared with the processor and fast compared with the disk, and the whole of memory management lives in that gap. A cache exists because memory is too slow for the processor; virtual memory exists because the disk is too slow to use directly and too large to ignore.

The cache is hardware and the operating system does not manage it. That is worth saying because a question sometimes asks: the operating system manages main memory and the disk; the processor manages the cache.

The simplest possible protection: base and limit

Two registers, and they are the smallest thing that works.

RegisterHolds
base, or relocation registerthe smallest physical address the process may use
limitthe size of the range it may use

On every single memory access the hardware checks:

address < limit

and if it passes, the memory actually used is:

physical address = base + address

munotes.in258

Why Memory Needs Managing

If the check fails, the hardware raises a trap, the kernel's handler runs, and the process is killed. That is Chapter two of Module 1's segmentation fault, from the memory side.

The limit register holds the SIZE, not the highest address. A process with base 300040 and limit 120900 may use addresses 0 to 120899, which land at 300040 to 420939. A student who reads the limit as the top address gets every translation wrong by the base.

Worked, three addresses

Base 300040, limit 120900.

The program's addressTest: is it below 120900?Physical address
0yes300040 + 0 = 300040
5000yes300040 + 5000 = 305040
120899yes, just300040 + 120899 = 420939
120900no: equal is not belowtrap
500000notrap

Notice the fourth row. The limit test is strictly less than, so an address equal to the limit is already outside. Off by one here is a real bug in real hardware manuals and it is worth being exact about.

Who is allowed to load those registers

Loading the base and limit registers is a privileged instruction, so only the kernel does it, and it does it during the context switch of Chapter sixteen. If a process could load its own base register it could point at anybody's memory, and the protection would be a suggestion.

That is the same argument as the mode bit in Chapter three, and it is the general shape of every protection mechanism in this book: the hardware checks, and only the kernel may change what it checks against.

Where a program's memory comes from

A process's memory is not one lump. Chapter fourteen of Module 1 gave the four sections; here is where they come from and who decides how big they are.

SectionSize decidedGrows while running
Textby the compiler, fixed in the fileno
Databy the compiler, fixed in the fileno
Heapby the program, as it asksyes, upwards
Stackby how deep the calls goyes, downwards

So the total memory a process needs is not known when it starts. That single fact is why contiguous allocation in Chapter seventy is awkward and why paging in Chapter seventy four wins.

Seeing it on the lab machine

The kernel publishes how much memory a process is using, in two different senses.

$ set +m
$ sleep 30 & job=$!
$ sleep 0.2
$ grep -E '^(VmSize|VmRSS|VmData|VmStk|VmExe)' /proc/$job/status
VmSize:	    2868 kB
VmRSS:	    1828 kB
VmData:	     244 kB
VmStk:	     132 kB
VmExe:	      20 kB
$ kill $job 2>/dev/null
$ wait 2>/dev/null

VmSize and VmRSS are the two numbers to tell apart, and a question asks for the difference.

munotes.in259

Why Memory Needs Managing

  • VmSize is the size of the process's address space: everything it could address.
  • VmRSS, the resident set size, is how much of that is actually in physical memory

right now.

VmSize is larger, and on this machine much larger. The gap is the whole subject of virtual memory, Chapters eighty to ninety one: a process may have an address space bigger than the memory it is given.

VmExe is the text, VmData the data and heap, and VmStk the stack: the sections of the table above, with numbers.

What it does not mean

Memory management is not about making memory bigger. It is about handing out what there is, protecting it, and using the disk to pretend there is more.

The limit register is not the end address. It is the size.

Protection is not a check the operating system performs. The hardware performs it, on every access, because the operating system is not running while the process is. The operating system only sets up what the hardware checks against.

VmSize is not how much memory a process is using. It is how much it could address. VmRSS is what it is using.

Quick revision

  • Main memory is the only storage the processor addresses directly, and the hardware can only

read and write a word: it has no notion of whose memory it is.

  • Memory management must give protection, relocation and deliberate sharing.
  • Costs: an instruction well under a nanosecond, a memory access tens of nanoseconds, a disk

access milliseconds. The cache is the processor's answer to the first gap, virtual memory the operating system's answer to the second.

  • The simplest protection is a base register and a limit register: the check is

address < limit and the translation is base + address.

  • The limit holds the size, not the top address, and the test is strictly less than.
  • Loading those registers is privileged, done by the kernel during a context switch.
  • A process's total need is not known when it starts, because the heap and the stack grow.
  • VmSize is the address space; VmRSS is how much of it is in physical memory.

Test yourself

  1. Why can protection not come from the instruction set? Because a load or a store just uses

the address it is given; the hardware has no idea which process is running. Protection has to sit between the address the program uses and the address the memory sees.

  1. What three things must memory management achieve? Protection of one process from another,

relocation so a program runs wherever it is placed, and deliberate sharing where it is wanted.

  1. Give the base and limit check and translation. The access is legal if the address is
munotes.in260

Why Memory Needs Managing

strictly less than the limit, and the physical address is then the base plus the address. 4. A process has base 300040 and limit 120900. Where does address 5000 land, and what about 120900? 5000 lands at 305040. 120900 is not strictly less than the limit, so it traps.

  1. Why must loading the base register be privileged? Otherwise a process could point its base

at another process's memory and the protection would mean nothing.

  1. Distinguish VmSize from VmRSS. VmSize is the size of the process's address space,

everything it could address. VmRSS is how much of that is in physical memory at the moment.

  1. Who manages the cache? The processor, in hardware. The operating system manages main

memory and the disk.

Contents This chapter on its own page

munotes.in261

Chapter Sixty-Seven

Binding an Address: Compile, Load, Run

Syllabus topic Module 2, "Memory Management - Logical address space, Physical address space"

In one line

An address in a program can be turned into a real memory address at compile time, at load time, or at the moment the instruction runs, and only the last one gives the operating system any freedom.

The two address spaces

NameAlso calledWhat it isWho sees it
Logical addressvirtual addressthe address the program usesthe program, and the programmer
Physical addressthe address the memory unit seesthe memory, and the hardware

The set of all logical addresses a process can generate is its logical address space; the set of physical addresses that correspond to them is its physical address space.

The two are the same number only under compile time or load time binding. Under execution time binding they differ, and everything in the rest of this module depends on their differing.

"Virtual address" and "logical address" mean the same thing in this paper. Some books reserve "virtual" for systems with virtual memory; MU's text book uses the two interchangeably, and so does this chapter.

The three binding times

Bound atThe compiler and loader produceCan the program be moved after that?Used by
Compile timeabsolute addresses, fixed in the fileno: to move it you must recompileMS-DOS .COM programs, small embedded controllers
Load timerelocatable code, which the loader adjusts as it loadsno, not after loading: to move it you must reloadolder systems without hardware support
Execution timelogical addresses, translated on every accessyes, at any momentevery modern system

Read the third column. It is the only column that matters, and it explains the whole table: binding early is simpler and cheaper and takes away the operating system's freedom to move a process. Binding at execution time costs a translation on every single memory access, and buys everything: swapping, paging, virtual memory and sharing.

Execution time binding requires hardware, the memory management unit of the next chapter. It cannot be done in software, because it happens on every access and software cannot get between an instruction and the memory.

Why the operating system needs to move a process

A question asks why late binding is worth its cost, and there are four answers.

  1. The operating system does not know where there will be room. It decides where to put a

process when it starts it, and what is free then depends on what else is running.

  1. A process may be swapped out and brought back somewhere else (Chapter sixty nine).
  2. A process may grow. Its heap and stack grow while it runs, so its memory may have to be

moved to somewhere with more room around it.

  1. Its pieces need not be together at all. Paging in Chapter seventy four puts a process's
munotes.in262

Binding an Address: Compile, Load, Run

pages wherever there are free frames, which needs a translation per page and is impossible without late binding.

Where in the toolchain each binding happens

A chapter on binding is clearer with the stages named, and a question sometimes asks for them.

StageWhat it doesWhich binding could happen here
Compilersource to an object file, addresses relative to that filecompile time, if it is told the final address
Linker, or link editorseveral object files plus libraries into one load moduleit resolves references between files
Loaderthe load module into memory, ready to runload time, by adjusting every address
Runthe program executesexecution time, by the memory management unit

A dynamically linked library is bound later still. The call to a library routine is left as a stub, and the first time it is called the library is found, loaded if it is not already there, and the stub replaced. That is dynamic linking, and it is why one copy of the C library serves every program on the machine. Chapter ten of Module 1's trace showed it happening: openat on libc.so.6 before the program's own first instruction.

Both spaces, on the lab machine

Every address a program holds is a logical one, and the machine can be asked to prove that the physical ones are none of its business.

$ grep -c . /proc/self/maps
35
$ awk 'NR==1 {print "the lowest mapped logical address of awk starts with", substr($1,1,2)}' /proc/self/maps
the lowest mapped logical address of awk starts with 00
$ test -r /proc/self/pagemap && echo "pagemap exists, and it is the physical mapping" || echo "pagemap is not readable"
pagemap exists, and it is the physical mapping
$ head -c 1 /proc/self/pagemap > /dev/null 2>&1 && echo "and it can be read" || echo "and an ordinary user may NOT read it"
and an ordinary user may NOT read it

That last line is the chapter's point, proved. /proc/self/pagemap is the file that would tell a program which physical frame each of its pages is in, and an ordinary user is not allowed to read it. A process cannot find out its own physical addresses, because knowing them would be the first step to reaching somebody else's.

So the logical address space is not merely a convenience for the operating system. It is the boundary of what a process is allowed to know, and that is a stronger statement than the tables above.

Worked example: the same program at three bindings

A program 100 kilobytes long, whose first instruction refers to a variable at offset 500.

munotes.in263

Binding an Address: Compile, Load, Run

Compile time binding, told it will load at 0. The instruction contains the address 500. The program must be loaded at 0 and nowhere else. If address 0 is taken, the program cannot run until it is free.

Load time binding, loaded at 14000. The file contains offset 500 plus a note that it is relocatable. As the loader copies the program in, it changes that 500 to 14500. The program now runs only at 14000; to move it, load it again.

Execution time binding, base register 14000. The instruction still contains 500. When it executes, the hardware adds the base and uses 14500. Change the base register to 90000 and the same program, unmoved and unaltered, now uses 90500. That is what the operating system buys.

Notice that the second and third produce the same physical address. The difference is not the answer; it is when it was worked out, and therefore whether it can be worked out again differently.

Distinctions that carry marks

Logical addressPhysical address
Generated bythe processor, running the programthe memory management unit
Seen bythe programthe memory hardware
Also calledvirtual addressreal address
Can the program learn ityes: it is what a pointer holdsno: /proc/self/pagemap is not readable by an ordinary user
Compile timeLoad timeExecution time
Addresses in the file areabsoluterelocatablelogical
Moving the program needsrecompilingreloadingnothing
Hardware needednonenonea memory management unit
Costnonea pass over the addresses at loada translation on every access
Static linkingDynamic linking
The library iscopied into the programfound and loaded at first use
Copies in memoryone per programone for the whole machine
Bound atlink timerun time, at the first call

What it does not mean

A logical address is not fake. It is the only address the program has, and it is what every pointer in every program contains.

Execution time binding is not slow. The translation is done by hardware in parallel with the access, and Chapter seventy six's translation cache is what keeps it that way.

Relocatable code is not position independent code. Relocatable code is adjusted once by the loader. Position independent code needs no adjustment at all because every reference is relative, and it is what a shared library must be.

Load time binding is not execution time binding done early. It cannot be undone: after loading, the addresses in memory are absolute.

Quick revision

  • A logical, or virtual, address is what the program uses; a physical address is what

the memory unit sees. They are equal only under compile or load time binding.

  • Three binding times: compile time (absolute, cannot move), load time (relocatable,
munotes.in264

Binding an Address: Compile, Load, Run

adjusted once by the loader, cannot move after), execution time (translated on every access, can move at any moment).

  • Execution time binding needs hardware: the memory management unit.
  • It is worth its cost because the operating system does not know in advance where there is room,

a process may be swapped out and back, it may grow, and its pieces need not be contiguous.

  • Toolchain: compiler, linker, loader, run. Dynamic linking binds a library call at its

first use, so one copy serves the machine.

  • A process cannot learn its physical addresses: /proc/self/pagemap exists and an ordinary

user may not read it. The logical address space is the boundary of what a process may know.

Test yourself

  1. Distinguish a logical from a physical address. A logical, or virtual, address is generated

by the program; a physical address is what the memory unit is given. Under execution time binding they differ, and the memory management unit converts one to the other.

  1. Name the three binding times and what each prevents. Compile time gives absolute addresses

and the program cannot be moved without recompiling. Load time gives relocatable code that the loader adjusts, and it cannot be moved after loading. Execution time translates on every access and the program can be moved at any moment.

  1. Which binding needs hardware, and what hardware? Execution time binding, which needs a

memory management unit, because the translation happens on every memory access and software cannot get between an instruction and the memory.

  1. Give two reasons the operating system needs to be able to move a process. It does not know

in advance where there will be room, and a process may be swapped out and brought back somewhere else. A process may also grow, and under paging its pieces need not be contiguous at all.

  1. What is dynamic linking and what does it save? Binding a library call at its first use

rather than copying the library into the program. One copy of the library serves every program on the machine.

  1. Can a program find out its own physical addresses? No. The file that would tell it,

/proc/self/pagemap, is not readable by an ordinary user, because knowing a physical address is the first step to reaching somebody else's memory.

  1. Distinguish relocatable code from position independent code. Relocatable code is adjusted

once by the loader. Position independent code needs no adjustment because every reference is relative, which is what a shared library must be.

Contents This chapter on its own page

munotes.in265

Chapter Sixty-Eight

The Memory Management Unit

Syllabus topic Module 2, "Memory Management - MMU"

In one line

The memory management unit is the hardware that turns every logical address the processor produces into a physical address, on every access, before the memory ever sees it.

Where it sits

The processor produces an address. The memory takes an address. The memory management unit is between them, and nothing gets past it.

In orderWhat happens
1the instruction produces a logical address
2the memory management unit checks it and translates it
3the physical address goes on the address bus
4the memory answers

It is hardware, and it is not part of the operating system. The operating system programs it, by loading its registers during a context switch, and then it works unaided on every one of the millions of accesses a process makes per second. A question that says the operating system translates each address is wrong: it could not possibly, because it is not running while the process is.

And it is not optional. Remove it and there is no protection, no relocation and no virtual memory, because all three are the same mechanism.

Its simplest form: the relocation register

The base and limit of Chapter sixty six is a memory management unit, and the simplest one there is.

if address < limit then physical = base + address else trap

In that form the base register is called the relocation register, because loading a different value into it relocates the whole process without touching a byte of it.

Worked, and the arithmetic written out

Base 14000, limit 100000. A process refers to addresses 0, 346, and 100000.

LogicalTestPhysical
00 is below 10000014000 + 0 = 14000
346346 is below 10000014000 + 346 = 14346
99999just below14000 + 99999 = 113999
100000not belowtrap

Now change the relocation register to 90000 and run the same process.

LogicalPhysical
090000 + 0 = 90000
34690000 + 346 = 90346

Nothing in the program changed. That is relocation, and it is one register.

What it does on every access, in full

A real unit does more than add. Whatever the scheme, it performs some subset of these four things, and a question about what the unit does wants them.

  1. Translate: turn the logical address into a physical one, by adding a base, or by looking

up a segment table, or by looking up a page table.

  1. Check the bounds: refuse an address outside what the process owns, and trap.
  2. Check the permissions: refuse a write to a read only region, or an execute of data, and

trap.

  1. Cache the translation: remember recent lookups so that most accesses need no table read at
munotes.in266

The Memory Management Unit

all. That is Chapter seventy six's translation lookaside buffer.

Steps 2 and 3 are the reason the unit is a protection device and not merely an adder, and they are why a program that reads address 1 dies (Chapter two) and a program that writes to its own code dies (Chapter fourteen of Module 1 showed the text section mapped r-xp, read and execute and not write).

The three schemes it can implement

Everything in the rest of this row is one of these, and they differ only in what the unit looks the address up in.

SchemeThe unit holdsThe translation isChapter
Base and limittwo registersadd the baseeleven, this one
Segmentationa pointer to a segment tablelook the segment up, check its limit, add its baseeighteen
Paginga pointer to a page tablesplit the address, look the page up, join the frame to the offsetnineteen, twenty

So this chapter is the shape of the next eight. They are all the same three steps, with a different table and a different arithmetic in the middle.

What it costs, and why that matters

A table lookup is a memory access. So a naive paged unit makes two memory accesses for every one the program asked for: one to read the table, one to get the data.

That would halve the speed of the machine, and it is the reason the fourth thing in the list above exists. Chapter seventy six works the arithmetic and shows how a cache of translations brings the cost back to a few per cent.

The bounds and permission checks cost nothing: they happen in parallel with the translation, in hardware, and a legal access is not slowed at all by being checked.

Seeing the permission check work

The text section of a program is mapped read and execute but not write. Writing to it is not a mistake the compiler can catch, and the unit catches it on the instruction.

#include <stdio.h>
#include <string.h>

int main(void)
{
    /* a pointer to this function's own first instruction */
    void *code = (void *)main;

    printf("about to write one byte over my own code\n");
    memset(code, 0, 1);
    printf("this line is never reached\n");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o writecode writecode.c
$ sh -c ./writecode
about to write one byte over my own code
Segmentation fault (core dumped)
$ echo $?
139

The program was allowed to compute that address, hold it in a pointer, and pass it to a function. Nothing stopped it until the instruction that actually wrote, and what stopped it then was the memory management unit noticing that the region is not writable. No software check happened: the kernel only found out because the hardware trapped.

munotes.in267

The Memory Management Unit

That is the whole argument for doing protection in hardware rather than in the operating system: the check has to happen on every access, and only the hardware is present on every access.

Distinctions that carry marks

The memory management unitThe operating system
Ishardwaresoftware
Acts onevery memory accessa context switch, a trap, a system call
Translates addressesyesno: it programs the unit that does
Runs while a process runsyesno
Bounds checkPermission check
Refusesan address outside the process's own memorya write to read only memory, or executing data
Costsnothing: parallel with the translationnothing
Failure givesa trap, then the kernel kills the processthe same
Relocation registerPage table
Sizeone registerone entry per page, thousands of them
Translationaddlook up, then join
Process must be contiguousyesno

What it does not mean

The unit is not a cache. It has one inside it, the translation lookaside buffer, but its job is translation and protection.

It does not know about processes. It knows what is in its registers, which the kernel changes at every context switch. That change is the expensive part of Chapter sixteen's switch.

A trap is not the unit killing the process. The unit refuses and raises a trap; the kernel decides what to do, and what it usually does is send a signal.

Translation is not slow. A single translation is done in hardware in parallel with the access. What is slow is reading a table from memory to do it, which is why the cache exists.

Quick revision

  • The memory management unit is hardware between the processor and the memory. It

translates every logical address to a physical one, on every access.

  • The operating system does not translate addresses: it programs the unit, at a context

switch. It is not running while the process is.

  • Four jobs: translate, check the bounds, check the permissions, and

cache the translation.

  • Its simplest form is a relocation register plus a limit:

if address < limit then base + address else trap. Changing the relocation register moves the whole process.

  • Three schemes, all the same three steps with a different table: base and limit,

segmentation, paging.

  • A table lookup is itself a memory access, so a naive paged unit doubles the cost of every

access. That is what Chapter seventy six's cache is for.

  • The bounds and permission checks are free: they happen in parallel with the translation.
  • Writing to your own code compiles, runs, and dies on the instruction: the hardware refused, not
munotes.in268

The Memory Management Unit

the operating system.

Test yourself

  1. What does the memory management unit do? It converts every logical address the processor

generates into a physical address, checking the bounds and the permissions as it does so, and traps if either fails.

  1. Is it hardware or software, and why does that matter? Hardware. The check must happen on

every memory access, and the operating system is not running while the process is, so only hardware can be present.

  1. Give the translation and check for a relocation register. If the logical address is

strictly less than the limit, the physical address is the base plus the address; otherwise the hardware traps.

  1. How is a process relocated under that scheme? By loading a different value into the

relocation register. Nothing in the program changes.

  1. Name the four things a memory management unit does on an access. Translate the address,

check it is within the process's bounds, check the permissions for that region, and cache the translation for next time.

  1. Why would a naive paged unit halve the speed of the machine? Because reading the page

table is itself a memory access, so every access the program makes costs two.

  1. A program writes to its own code and dies. What caught it? The memory management unit,

because the text region is mapped read and execute but not write. It raised a trap and the kernel then killed the process; no software check took place.

Contents This chapter on its own page

munotes.in269

Chapter Sixty-Nine

Swapping

Syllabus topic Module 2, "Memory Management - Swapping"

In one line

Swapping is moving a whole process out of memory to a disk and bringing it back later, so that more processes can exist than fit in memory at once.

The idea

A process must be in memory to run. It does not have to be in memory while it is not running.

So: take a process that is not running, write the whole of it to a reserved area of the disk, and use its memory for somebody else. When it is to run again, read it back in.

WordMeans
Swap out, or roll outwrite a process from memory to the disk
Swap in, or roll inread it back from the disk into memory
Backing storethe disk area used, fast and large enough for every swapped process
Swap spacethe same thing, named as a part of the disk

Chapter seventeen of Module 1 named the scheduler that does this: the medium term scheduler. Swapping is a scheduling decision as much as a memory one, which is why it appears in both rows.

Swapping needs execution time binding if the process is to come back to a different place, which it usually must. With compile time or load time binding it can only be swapped back to exactly where it was, and if that memory is taken it cannot come back at all.

What it costs, worked

This is the sum a question wants, and the numbers make the whole chapter.

A process of 100 megabytes. A disk that transfers at 50 megabytes per second, with an average latency before the transfer starts of 8 milliseconds.

transfer time = 100 / 50 = 2 seconds

one way = 2 seconds + 8 milliseconds

out and back = 2 × (2 seconds + 8 milliseconds) = 4 seconds and 16 milliseconds

Four seconds to swap one process out and back, of which the latency is 16 milliseconds and the transfer is four whole seconds. The latency does not matter at all; the size does.

Compare that with the round robin quantum of Chapter fifty one, which was ten to a hundred milliseconds. The swap costs forty to four hundred times the quantum. So a system cannot swap a process out and in for every turn of the processor: swapping is something done to a process that will not run for a long time.

The one lever that matters

The time is proportional to the amount swapped, so swap less. If a process has 100 megabytes of address space but has only touched 10, swapping only what it has touched costs a tenth of the time. That single observation is why modern systems do not swap whole processes at all, and it is the answer to "why is swapping not used as described".

munotes.in270

Swapping

To swap only what is needed, the system must know how much a process is using, which means the process must tell it: requesting and releasing memory explicitly rather than being given a fixed block.

Two things that go wrong

A process with pending input or output cannot be swapped out naively. If a device is writing into the process's buffer and the process's memory is given to somebody else, the device writes into the wrong process's memory. Two answers:

  1. Never swap a process with pending input or output.
  2. Do all input and output into kernel buffers and copy to or from the process's memory only

when it is in memory. This is what real systems do, and it is the reason for the copies Chapter five of Module 1 called buffering.

The swapped out process's state must be complete. Everything the process owns must go, and its kernel records must stay, because the kernel has to know where on the disk it is.

Why modern systems do not do this

No general purpose system swaps whole processes as a normal operation, and a question that asks why has three answers.

  1. The time is proportional to the size, and processes are large. Four seconds is not a cost

that can be paid often.

  1. Most of a process is not needed. A word processor with a hundred megabytes of address

space may be running a loop in four kilobytes of it.

  1. Paging can do it a page at a time. Chapters eighty to ninety one are exactly this idea

applied to pieces of a process instead of the whole, and the piece is four kilobytes rather than a hundred megabytes.

So what survives is the name and a reduced version of the mechanism: a modern system pages out parts of processes, and calls the disk area the swap space. Some systems will still swap an entire process out when memory is desperately short, as a last measure before killing something.

The lab machine's own swap

$ free -m | awk 'NR==1 || /Swap/'
               total        used        free      shared  buff/cache   available
Swap:              0           0           0
$ grep -E 'SwapTotal|SwapFree' /proc/meminfo
SwapTotal:             0 kB
SwapFree:              0 kB
$ echo "this machine has no swap space at all"
this machine has no swap space at all

This machine has no swap space, and that is worth a paragraph rather than a footnote. A container normally has none: swap belongs to the host, not to the container. On a machine with no swap, a process's pages that are dirty, meaning changed since they were read in, have nowhere to go, so the system cannot free them. When memory runs out it must kill something instead.

munotes.in271

Swapping

That is a real operational fact and not an artefact of this book's lab: a Linux machine with no swap does not run slowly when memory is short, it kills a process. The mechanism has a name, the out of memory killer, and a system administrator's first question about a killed process is always whether the machine has swap.

So the ordinary contents of a swapped page cannot be shown on this machine. What can be shown is that the facility is named, sized and reported by the kernel, in the two files above, and that it is empty here.

Distinctions that carry marks

SwappingPaging (Chapter seventy four)
The unit movedthe whole processone page, typically 4 kilobytes
Time for one movesecondsmicroseconds to milliseconds
Can be done per quantumnoyes
Needs execution time bindingyes, to come back elsewhereyes
Used by modern systemsonly as a last measureconstantly
Swap outSwap in
Directionmemory to backing storebacking store to memory
Also calledroll outroll in
Chosen bythe medium term schedulerthe same, when the process is to run again

What it does not mean

Swapping is not virtual memory. Swapping moves whole processes and every process in memory is complete. Virtual memory keeps part of a process in memory and is Chapter eighty.

The backing store is not a file. It is a raw area of disk, used without a file system, because a file system's bookkeeping would be pure overhead on data with no name and no permanence.

Swapping out does not free the process's kernel records. The kernel must keep knowing about it, including where on the disk it now is.

A machine with no swap is not a machine with more memory. It is a machine with nowhere to put pages that must be kept, so it kills processes instead of slowing down.

Quick revision

  • Swapping writes a whole process out to a backing store and reads it back later, so more

processes can exist than fit in memory.

  • Swap out or roll out; swap in or roll in. The chooser is the medium term scheduler.
  • It needs execution time binding if the process is to come back somewhere else.
  • The cost is proportional to the size: 100 megabytes at 50 megabytes a second is 2 seconds

each way, so about 4 seconds out and back, against a quantum of tens of milliseconds. The latency is negligible; the size is everything.

  • Therefore: swap less, which means swapping pieces rather than whole processes, which is
munotes.in272

Swapping

paging.

  • A process with pending input and output must not be swapped naively; real systems do input and

output into kernel buffers to avoid it.

  • The lab machine has no swap space, which is normal for a container, and a Linux machine

with no swap kills a process when memory runs out rather than slowing down.

Test yourself

  1. What is swapping? Moving a whole process out of main memory to a backing store on disk and

bringing it back later, so that the total of the processes can exceed the memory available.

  1. What is the backing store and why is it not a file? A raw area of disk large and fast

enough to hold every swapped process. It is not a file because a file system's naming and bookkeeping would be pure overhead for data that has no name and does not survive. 3. A 100 megabyte process, a disk at 50 megabytes a second and 8 milliseconds of latency. What does a swap out and back cost? The transfer is 100 / 50 = 2 seconds each way, so about 4 seconds plus 16 milliseconds of latency.

  1. Compare that with a scheduling quantum, and say what follows. A quantum is tens of

milliseconds, so the swap costs forty to four hundred times as much. Swapping cannot be done for every turn of the processor; it is for processes that will not run for a long time.

  1. Why do modern systems not swap whole processes? The cost is proportional to the size and

processes are large; most of a process is not needed at any moment; and paging does the same thing one four kilobyte page at a time.

  1. Why is a process with pending input or output a problem, and what is the fix? A device

writing into its buffer would write into whatever process is given that memory. The fix is to do input and output into kernel buffers and copy to the process only while it is in memory.

  1. What happens on a Linux machine with no swap when memory runs out? It cannot free dirty

pages because there is nowhere to put them, so it kills a process rather than slowing down.

Contents This chapter on its own page

munotes.in273

Chapter Seventy

Contiguous Allocation, and the Holes It Leaves

Syllabus topic Module 2, "Memory Management - Contiguous Memory Allocation"

In one line

Give each process one unbroken block of memory, and after a while the free memory is a scatter of gaps that are each too small to be useful.

The starting point

Memory is divided in two: the operating system and the user processes. The kernel usually sits in low memory, with the interrupt vector, because that is where the hardware looks for the vector.

Each user process then gets one contiguous block. That is the whole of contiguous allocation, and everything about it follows from the word contiguous.

With base and limit registers from Chapter sixty six, protection is complete and free: one register pair per process, loaded at the context switch, and an address outside the block traps.

Fixed partitions, the first scheme

Divide the user memory into a fixed number of partitions, once, when the system is configured. One process per partition, so the number of partitions is the degree of multiprogramming.

Simpleextremely: a partition is free or it is not
Costa process smaller than its partition wastes the rest, and that waste has a name: internal fragmentation
Costa process larger than every free partition cannot run, even if the total free memory is plenty
Used byIBM's OS/360 MFT, and nothing modern

The waste inside a partition is the definition of internal fragmentation: memory allocated to a process that the process is not using, and which nobody else can use because it has been allocated. Chapter seventy two is this and its opposite.

Variable partitions, the real scheme

Keep a table of which parts of memory are free. Each free part is a hole. When a process arrives, find a hole big enough, give the process what it needs, and leave the remainder as a smaller hole.

EventWhat happens to the holes
A process arrives and is placedthe chosen hole shrinks by the process's size, or disappears if it fitted exactly
A process leavesits block becomes a hole
A process leaves next to a holethe two are merged into one larger hole
A process leaves between two holesall three are merged

The merging is not optional and it is where a naive implementation goes wrong. Two adjacent holes that are not merged are two holes of 100 and 200 where there should be one of 300, and a process of 250 is then refused memory that exists. Merging adjacent free blocks is called coalescing and every allocator does it.

The hole list, worked as a sequence

Memory of 2560 kilobytes, with the operating system in the first 400.

StepEventThe holes afterwards
0startone hole: 400 to 2560, size 2160
1P1 of 600 placed at 4001000 to 2560, size 1560
2P2 of 1000 placed at 10002000 to 2560, size 560
3P3 of 300 placed at 20002300 to 2560, size 260
4P2 finishes1000 to 2000 (1000) and 2300 to 2560 (260)
5P4 of 700 placed at 10001700 to 2000 (300) and 2300 to 2560 (260)
6P1 finishes, and 400 to 1000 is next to nothing free400 to 1000 (600), 1700 to 2000 (300), 2300 to 2560 (260)
7P4 finishes, and 1000 to 1700 is next to the hole at 1700400 to 1000 (600), 1000 to 2000 (1000), 2300 to 2560 (260)
munotes.in274

Contiguous Allocation, and the Holes It Leaves

Look at step 7. P4's block of 700 and the existing hole of 300 became one hole of 1000, not two. Had they not been coalesced, a later process of 800 would have been refused although the memory was there.

And look at the whole table's last line. There are 1860 kilobytes free in three pieces. A process of 1200 cannot run, although there is more than enough memory in total. That is external fragmentation and it is Chapter seventy two.

How the free memory is actually recorded

Two standard implementations, and a question may ask for them.

A bit mapA linked list of holes
Recordsone bit per unit of memory: free or usedone node per hole, with its address and size
Finding a hole of n unitsscan for n consecutive free bitswalk the list
Cost of a bit mapone bit per unit; a 1 kilobyte unit over 2 gigabytes is 256 kilobytes of map
Cost of a lista node per hole, and the nodes can live in the holes themselves, costing nothing
Coalescingautomatic: adjacent free bits are adjacent free bitsneeds the neighbours to be checked and merged

A list is usual, and the nodes go inside the holes, because a hole is by definition memory nobody is using. That is a neat trick worth knowing: the free list costs no memory at all.

Why contiguous allocation is worth studying at all

No modern general purpose system allocates a process's memory contiguously. So why is it in the syllabus?

  1. The fits of the next chapter are the standard sums on this paper, and they are asked every

year.

  1. The fragmentation it creates is the reason paging exists, and paging cannot be motivated

without it.

  1. It is still used inside things. The kernel's own allocator gives contiguous physical

memory for device buffers, because a device writing directly into memory cannot follow a page table. Choosing where those blocks go is exactly this problem.

munotes.in275

Contiguous Allocation, and the Holes It Leaves

  1. Every file system does it, in a different currency: Chapter one hundred seven of this

module allocates contiguous disk blocks to a file, with the same holes and the same fragmentation.

So the ideas transfer even though the scheme does not, and that is the honest reason it is taught.

Distinctions that carry marks

Fixed partitionsVariable partitions
Boundariesset once at configurationchange as processes come and go
A small process in a big spacewastes the rest: internal fragmentationleaves the rest as a hole
Degree of multiprogrammingthe number of partitionsas many as fit
Records keptwhich partitions are occupieda list of holes, with addresses and sizes
Main faultinternal fragmentationexternal fragmentation
CoalescingCompaction (Chapter seventy two)
Doesmerges adjacent free blocksmoves processes so all the free memory is together
Costsalmost nothingcopying every process's memory
Needsnothingexecution time binding

What it does not mean

Contiguous does not mean fixed. A process's block can be moved, with execution time binding, and moving them all is compaction.

A hole is not wasted memory. It is free memory in an awkward place. What is wasted is the part nobody can use because no hole is big enough.

A bit map is not always worse than a list. It coalesces for nothing and makes finding n consecutive units a simple scan, which is why disk free space in Chapter one hundred nine often uses one.

Fixed partitions are not merely historical. A system with a fixed number of identical jobs, such as some embedded controllers, is still configured this way on purpose, because it is predictable.

Quick revision

  • Contiguous allocation gives each process one unbroken block, so base and limit

registers are enough for protection.

  • Fixed partitions: memory divided once; the number of partitions is the degree of

multiprogramming; a process smaller than its partition wastes the rest, which is internal fragmentation.

  • Variable partitions: a list of holes; a placement shrinks a hole and a departure

creates one.

  • Adjacent holes must be coalesced, or memory that exists is refused.
  • Free memory is recorded either as a bit map, one bit per unit, which coalesces for nothing;

or as a linked list of holes, whose nodes can live inside the holes and so cost nothing.

  • The scheme is obsolete for process memory and the ideas are not: the fits are examined, the

fragmentation motivates paging, the kernel still allocates contiguous physical memory for devices, and every file system allocates contiguous disk blocks the same way.

Test yourself

  1. What does contiguous allocation mean, and what does it buy? Each process gets one unbroken
munotes.in276

Contiguous Allocation, and the Holes It Leaves

block of memory, so a base and a limit register are enough to translate and protect every access.

  1. What is a fixed partition scheme, and what does it waste? Memory divided into a set number

of partitions once, with one process per partition. A process smaller than its partition wastes the remainder, which is internal fragmentation.

  1. What happens to the holes when a process finishes? Its block becomes a hole, and if it is

adjacent to an existing hole the two are merged, which is coalescing.

  1. Why must adjacent holes be coalesced? Otherwise two adjacent holes of 100 and 200 stay two

holes, and a process of 250 is refused memory that is physically there.

  1. Give the two ways of recording free memory, with an advantage of each. A bit map, one bit

per unit, which coalesces automatically and turns the search into a scan; and a linked list of holes, whose nodes can be stored inside the holes and so cost no memory at all. 6. Memory is 2560 with the operating system in the first 400. P1 600, P2 1000 and P3 300 are placed in order, then P2 finishes. What are the holes? One of 1000, from 1000 to 2000, and one of 260, from 2300 to 2560.

  1. Give two reasons contiguous allocation is still worth studying. The fits are the standard

examination sums, and the fragmentation it creates is the reason paging exists; besides which the kernel still allocates contiguous physical memory for device buffers, and file systems allocate contiguous disk blocks in the same way.

Contents This chapter on its own page

munotes.in277

Chapter Seventy-One

First Fit, Best Fit and Worst Fit

Syllabus topic Module 2, "Memory Management - Contiguous Memory Allocation"

In one line

Given a list of holes and a process that needs memory, there are four standard ways to choose which hole to use, and they give different answers.

The four policies

PolicyChoosesSearch
First fitthe first hole big enough, scanning from the startstops as soon as one fits
Best fitthe smallest hole big enoughmust examine every hole, unless the list is kept sorted by size
Worst fitthe largest holemust examine every hole, unless sorted
Next fitthe first hole big enough, scanning from where the last search stoppedstops as soon as one fits

"Best" and "worst" are names, not verdicts. Best fit means the closest fit in size; it is not the policy that performs best. The measurements below are the reason to be careful about that.

The reasoning behind each is worth a line, because a question asks why anybody would choose worst fit:

  • First fit is fastest, because it stops early.
  • Best fit leaves the smallest remainder, on the argument that a small remainder wastes

least.

  • Worst fit leaves the largest remainder, on the opposite argument: a large remainder is

still useful, whereas best fit's tiny remainders are useless slivers.

  • Next fit is first fit that does not keep re-scanning the crowded start of the list.

The standard problem, worked four times

Five holes, in address order: 100, 500, 200, 300, 600. Four processes arrive in order, needing 212, 417, 112, 426.

First fit

Placement: first fit

Holes: 100 500 200 300 600

Processes: 212 417 112 426

ProcessSizeHoleHole sizeLeft over
P12122500288
P24175600183
P31122500176
P4426none

External fragmentation: 959

Follow it. P1 of 212 skips hole 1 (100 is too small) and takes hole 2, leaving 288. P2 of 417 no longer fits in hole 2, so it walks on to hole 5 of 600, leaving 183. P3 of 112 takes hole 2 again, which still has 288, leaving

  1. P4 of 426 fits nowhere: the holes are now 100, 176, 200, 300, 183, and the largest is

300.

Best fit

Placement: best fit

Holes: 100 500 200 300 600

Processes: 212 417 112 426

ProcessSizeHoleHole sizeLeft over
P1212430088
P2417250083
P3112320088
P44265600174

External fragmentation: 533

All four processes fit. P1 of 212 takes hole 4 of 300, the smallest that will do, rather than wasting hole 2. That one decision keeps hole 2's 500 intact for P2's 417, and hole 5's 600 stays free for P4's 426.

munotes.in278

First Fit, Best Fit and Worst Fit

Worst fit

Placement: worst fit

Holes: 100 500 200 300 600

Processes: 212 417 112 426

ProcessSizeHoleHole sizeLeft over
P12125600388
P2417250083
P31125600276
P4426none

External fragmentation: 959

P1 of 212 takes the largest hole, 600, leaving 388. P2 of 417 then needs the 500. P3 of 112 takes the 388, leaving 276. P4 of 426 fits nowhere.

Next fit

Placement: next fit

Holes: 100 500 200 300 600

Processes: 212 417 112 426

ProcessSizeHoleHole sizeLeft over
P12122500288
P24175600183
P3112560071
P4426none

External fragmentation: 959

Next fit starts each search where the last one stopped. P1 takes hole 2. P2 continues from hole 2 and reaches hole

  1. P3 continues from hole 5 and finds hole 5 still has 183, which is more than 112, so it takes

that. P4 of 426 fits nowhere.

The comparison

PolicyProcesses placedLeft waitingExternal fragmentation
First fit3P4959
Best fit4none533
Worst fit3P4959
Next fit3P4959

On this problem best fit wins outright, and that is why it is the example every textbook uses. Three things to say about it, and the third is the important one.

  1. Best fit placed every process and the other three each left P4 waiting for memory that

existed in total but not in one piece.

  1. Best fit left the least free memory because it left the least: it used the tightest hole

each time.

  1. This is one problem. Change the sizes and the ranking changes. The measured result on real

workloads, from decades of simulation, is that first fit and best fit are both better than worst fit, and first fit is usually faster, which is why first fit is what an implementation reaches for. Worst fit is worst at both speed and storage, and a question asking which to avoid wants it named.

So the honest summary, and the one to write: best fit and first fit are comparable in how well they use storage, first fit is faster, and worst fit is worse than both. The worked example above is evidence about this set of holes and not a proof about the policies.

Why best fit's remainders are a problem anyway

Best fit leaves the smallest remainder each time, and a remainder too small for any process to use is dead memory. After a thousand placements, best fit's free list is a long list of useless slivers, and walking it is slow as well as pointless.

That is precisely the argument for worst fit, and it is the argument that measurement rejected: worst fit's large remainders sound useful and in practice it runs out of large holes, which is exactly what a process needing a large block requires. The idea is sound and the numbers disagree with it, which is worth knowing as an example of why allocator design is measured rather than reasoned.

munotes.in279

First Fit, Best Fit and Worst Fit

How much of the memory can be lost

A result worth quoting, and it is Chapter seventy two's headline.

With first fit, statistical analysis shows that for N blocks allocated, another 0.5 N blocks are lost to fragmentation. A third of the memory may be unusable. That figure has a name: the fifty per cent rule.

Distinctions that carry marks

First fitBest fitWorst fit
Choosesthe first that fitsthe smallest that fitsthe largest
Search coststops earlythe whole list, or keep it sortedthe whole list, or keep it sorted
Remainder leftwhatever it happens to bethe smallestthe largest
Measured storage usegoodgoodworst
Measured speedbestslowerslower
External fragmentationThe fifty per cent rule
Isfree memory scattered in pieces too small to usethe measured size of that loss under first fit
Saysnothing about how muchfor N blocks allocated, about 0.5 N are lost

What it does not mean

Best fit is not the best policy. It is the policy that picks the closest fit in size. Measurement makes it roughly equal to first fit in storage and slower.

Worst fit is not a joke. Its reasoning is sound: large remainders stay useful. It is simply wrong in practice.

A process that fits nowhere is not out of memory. In three of the four runs above there was more free memory than P4 needed. It was in the wrong pieces.

These are not the algorithms a modern allocator uses. A real one keeps separate lists per size class so that finding a hole takes constant time, and that design exists because scanning a list was measured to be the cost.

Quick revision

  • First fit: the first hole big enough. Best fit: the smallest big enough. Worst fit:

the largest. Next fit: first fit continuing from where the last search stopped.

  • Best and worst must scan the whole list unless it is kept sorted; first and next stop early.
  • On holes 100, 500, 200, 300, 600 with processes 212, 417, 112, 426:

best fit places all four and leaves 533 free; first, worst and next each leave P4 waiting and 959 free.

  • That is one problem, not a proof. Measured over real workloads:

first fit and best fit are comparable on storage, first fit is faster, worst fit is worse at both.

munotes.in280

First Fit, Best Fit and Worst Fit

  • Best fit's small remainders become useless slivers; worst fit's large remainders sound useful

and run out of the large holes a big process needs.

  • The fifty per cent rule: under first fit, for N blocks allocated about 0.5 N blocks are

lost to fragmentation, so up to a third of memory can be unusable.

Test yourself

  1. Define the three main placement policies. First fit takes the first hole large enough,

scanning from the start. Best fit takes the smallest hole large enough. Worst fit takes the largest hole.

  1. Which must examine the whole list, and how can that be avoided? Best fit and worst fit,

unless the list of holes is kept sorted by size. 3. Holes 100, 500, 200, 300, 600 and processes 212, 417, 112, 426. Which policy places all four? Best fit. The others leave the last process, 426, with nowhere to go.

  1. Why does best fit succeed there? Because it puts 212 into the hole of 300 rather than into

the hole of 500, leaving 500 for 417 and 600 for 426.

  1. What do measurements say about the three policies? First fit and best fit are comparable

in how well they use storage and first fit is faster; worst fit is worse than both on storage and on speed.

  1. What is the reasoning behind worst fit and why does it fail? That a large remainder is

still useful whereas best fit's small remainders are useless. It fails because it consumes the large holes, which are exactly what a process needing a big block requires.

  1. State the fifty per cent rule. Under first fit, for every N blocks allocated about another

0.5 N blocks are lost to fragmentation, so as much as a third of memory may be unusable.

Contents This chapter on its own page

munotes.in281

Chapter Seventy-Two

Fragmentation, Internal and External

Syllabus topic Module 2, "Memory Management - Contiguous Memory Allocation"

In one line

External fragmentation is free memory in pieces too small to use; internal fragmentation is memory given to a process that the process does not need.

The two, defined

External fragmentationInternal fragmentation
The wasted memory isoutside every process, in the holesinside a process's own allocation
It is wasted becauseno single hole is large enoughit was given away and the owner is not using it
Who could use itsomebody, if it were in one piecenobody, because it is allocated
Caused byallocating exactly what is asked for, repeatedlyallocating in fixed sizes larger than what is asked for
Cured bycompaction, or by paginga smaller allocation unit

The word to hold on to is WHERE. External waste is between the processes; internal waste is within one. That is the whole distinction and it is what a two mark question is checking.

And note the last row. They are cured by opposite things, which is why no scheme has neither: making the allocation unit smaller reduces internal fragmentation and increases the number of pieces, and making it larger does the reverse. Chapter seventy four's paging chooses a unit and accepts the internal waste that comes with it.

External fragmentation, measured

Chapter seventy one's first fit run ended with holes of 100, 176, 200, 300 and 183, which is 959 kilobytes free, and a process needing 426 could not run.

free in total = 100 + 176 + 200 + 300 + 183 = 959

the largest single hole = 300

what P4 needed = 426

959 free, 426 wanted, and refused. That is external fragmentation in three numbers, and it is the way to show it in an answer: the total free, the largest hole, and the request.

The fifty per cent rule

Statistical analysis of first fit gives a figure worth quoting: for N blocks allocated, another 0.5 N blocks are lost to fragmentation. So of every three blocks' worth of memory, one may be unusable: a third of the machine.

That is an average over a long run of random sizes, not a promise about any moment. What it establishes is that external fragmentation is not a corner case to be dismissed: it is large.

Internal fragmentation, measured

Internal fragmentation appears the moment memory is given out in fixed sizes. Two examples, and the second is the one that matters for the rest of this module.

Fixed partitions. A partition of 512 kilobytes holding a process of 300 wastes 212 kilobytes. Nobody else can have it, because the partition is occupied.

Variable partitions with a minimum block size. Suppose an allocator refuses to leave a hole smaller than 4 kilobytes, because a hole that small is not worth recording. A hole of 18,464 bytes and a request for 18,462 leaves 2 bytes; rather than record a 2 byte hole, the allocator gives the process all 18,464.

munotes.in282

Fragmentation, Internal and External

given = 18464

asked for = 18462

internal fragmentation = 18464 - 18462 = 2

Two bytes is trivial and the principle is not: the allocator chose internal fragmentation over an unusably small hole, and every allocator makes that choice with a threshold of its own.

And the one that matters: paging

Chapter seventy four allocates memory in pages, all the same size. A process of 72,766 bytes with a page size of 2,048 bytes needs:

pages needed = 72766 / 2048 = 35.5302734375

which is 35 full pages and a part page, so 36 pages are allocated.

allocated = 36 × 2048 = 73728

internal fragmentation = 73728 - 72766 = 962

962 bytes wasted in the last page, and paging has no external fragmentation at all. That is the trade the whole of Chapter seventy four rests on: paging converts external fragmentation into internal fragmentation, and bounds it.

And the bound is the examinable part. The waste is in the last page only, so on average it is half a page per process. A question that asks for the expected internal fragmentation under paging wants exactly that: half a page, per process.

average internal fragmentation per process = 2048 / 2 = 1024

Which gives the argument about page size in one line. Small pages waste less memory and need a bigger page table; large pages waste more memory and need a smaller table. Four kilobytes is the usual compromise and Chapter seventy eight shows the table side of it.

Compaction

The cure for external fragmentation without changing the allocation scheme: move the processes so that all the free memory is in one piece.

Needsexecution time binding (Chapter sixty seven): the relocation register is simply changed
Costscopying the memory of every process that moves
When it can be doneonly while the process is not running, and not while a device is writing into its memory
How much to moveas little as possible, and choosing the cheapest set of moves is a real problem of its own

Compaction is impossible under compile time or load time binding, and that is worth saying in an answer: the whole point of late binding, from Chapter sixty seven, is what makes compaction available at all.

Worked, and the cost

Three processes of 600, 1000 and 300 kilobytes, with holes of 300 and 260 between and after them. Compacting means moving the last process down by 300 so the two holes become one of 560.

munotes.in283

Fragmentation, Internal and External

BeforeAfter
P1 600, hole 300, P3 300, hole 260P1 600, P3 300, hole 560

The cost is copying 300 kilobytes, which is P3's whole memory. At a memory copy rate of 1 gigabyte per second:

time = 300 / 1000000 = 0.0003 seconds

A third of a millisecond, which is cheap. Now compact 4 gigabytes of processes and it is four seconds, and the machine does nothing else while it happens. Compaction costs time proportional to the memory moved, exactly like Chapter sixty nine's swapping, and for the same reason it is not something a system can do often.

What each scheme suffers from

The summary table, and it is the one to reproduce.

SchemeExternalInternal
Fixed partitionsyes, between partitions that are free but too smallyes, inside every partition
Variable partitionsyes, and the fifty per cent rule measures ita little, from minimum block sizes
Pagingnoneyes, half a page per process on average
Segmentationyes: segments are variable sizednone, by itself

Read the paging row twice. No external fragmentation at all, because any free frame will do for any page: there is no such thing as a frame being in the wrong place. That is the single reason paging won.

What it does not mean

External fragmentation is not a shortage of memory. The memory is there; it is in the wrong shape.

Internal fragmentation is not a bug in the allocator. It is a price the allocator chose to pay, usually to avoid recording holes too small to use.

Compaction is not free because memory is fast. It is proportional to the memory moved, and while it happens the processes being moved cannot run.

Paging does not eliminate fragmentation. It eliminates the external kind and bounds the internal kind at half a page per process.

Quick revision

  • External fragmentation: free memory outside the processes, in pieces too small to use.

Internal: memory inside an allocation that its owner does not need.

  • The test is where the waste is: between processes, or within one.
  • They are cured by opposite things, so a smaller allocation unit trades one for the other.
  • The fifty per cent rule: under first fit, for N blocks allocated about 0.5 N are lost, so

up to a third of memory may be unusable.

  • Paging converts external fragmentation into internal fragmentation and bounds it: none

external, and on average half a page per process internal.

  • A process of 72,766 bytes with 2,048 byte pages takes 36 pages, 73,728 bytes, wasting 962.
  • Compaction moves the processes so the free memory is in one piece. It needs
munotes.in284

Fragmentation, Internal and External

execution time binding and costs time proportional to the memory moved.

Test yourself

  1. Distinguish internal from external fragmentation. Internal fragmentation is memory

allocated to a process that the process does not need, wasted inside its own allocation. External fragmentation is free memory outside every process, in pieces each too small to satisfy a request.

  1. Which does paging suffer from, and how much? Internal only, and on average half a page per

process, because the waste is confined to the last page. 3. A process of 72,766 bytes and a page size of 2,048. How many pages, and how much is wasted? 36 pages, allocating 73,728 bytes, wasting 962.

  1. Why does paging have no external fragmentation? Because every frame is the same size as

every page, so any free frame will satisfy any page: a frame cannot be in the wrong place.

  1. State the fifty per cent rule. Under first fit, for every N blocks allocated another 0.5 N

blocks are lost to fragmentation, so as much as a third of memory may be unusable.

  1. What is compaction, and what does it require? Moving the processes so that all the free

memory is in one block. It requires execution time binding, so that a process's relocation register can simply be changed, and it costs time proportional to the memory moved. 7. 959 kilobytes are free and a process needing 426 cannot run. What is happening, and what would fix it? External fragmentation: the largest single hole is 300. Compaction would fix it, or paging would prevent it.

Contents This chapter on its own page

munotes.in285

Chapter Seventy-Three

Segmentation

Syllabus topic Module 2, "Memory Management - Segmentation"

In one line

Divide a program into its natural pieces, give each piece its own base and limit, and let each grow independently.

Why a programmer does not think in one address space

A program is not a single run of bytes to the person who wrote it. It is a main routine, some functions, a stack, a symbol table, some arrays. Each of those is a separate thing with a separate length, and the programmer refers to the fifth element of an array, not to byte 26,472 of the program.

Contiguous allocation forces all of it into one block with one base and one limit, and then a growing stack runs into a growing heap. Segmentation is the scheme that keeps the pieces apart, and that is the one sentence answer to why it exists.

A program's segments, typicallyHolds
the codethe instructions
global variablesthe data section
the heapmemory asked for while running
the stack, one per threadcall frames
the standard libraryshared with every other process

The address

A logical address under segmentation is two parts, and the program supplies both.

logical address = < segment number, offset >

On a machine where the address is one word, the segment number is the top bits and the offset the bottom bits, and the compiler produces both. On the Intel architecture it was two separate registers, a segment selector and an offset, which is where the design comes from.

The segment table

One entry per segment, and every entry is two numbers.

FieldMeans
basethe physical address where the segment starts
limitthe length of the segment

Two registers point at the table itself: the segment table base register, which says where the table is, and the segment table length register, which says how many entries it has.

The limit is a length, not a top address, exactly as in Chapter sixty six, and it is the same off by one trap.

The translation, in three steps

Given a logical address of segment s and offset d:

  1. Is s a legal segment number? If s is not less than the segment table length register,

trap.

  1. Is d within the segment? If d is not less than limit[s], trap: this is an

addressing error.

  1. The physical address is base[s] + d.

Step 2 is the protection, and it is per segment. That is a real gain over one base and limit for the whole process: a stack overrunning its own segment is caught immediately, rather than quietly writing over the heap.

Worked, with every case

A process with five segments.

SegmentWhat it holdsbaselimit
0the code14001000
1the symbol table6300400
2a string table4300400
3the main routine32001100
4the stack47001000
munotes.in286

Segmentation

Now translate eight addresses. Each line is base plus offset, after the limit test.

Logical addressLimit testPhysical address
segment 0, offset 00 below 10001400 + 0 = 1400
segment 0, offset 999999 below 10001400 + 999 = 2399
segment 0, offset 10001000 is not below 1000trap
segment 1, offset 5353 below 4006300 + 53 = 6353
segment 2, offset 399399 below 4004300 + 399 = 4699
segment 2, offset 400400 is not below 400trap
segment 3, offset 852852 below 11003200 + 852 = 4052
segment 4, offset 12221222 is not below 1000trap

Three of the eight trap, and two of them trap on an offset exactly equal to the limit. The test is strictly less than, and a question will include exactly that case.

Notice the segments are not in order in physical memory: segment 1 is at 6300, above segment 4 at 4700. The operating system put each segment wherever there was a hole big enough, which is the freedom segmentation buys and the fragmentation it costs.

What segmentation is good at

Three things, and each is a real advantage over one block per process.

  1. Each segment grows on its own. The stack growing does not threaten the heap: they are

different segments with different bases, and only the stack's own limit constrains it.

  1. Protection is per segment, and can differ per segment. The code segment can be marked read

only and execute, the data segment read and write and not execute. A program that jumps into its data is caught. One base and limit for the whole process cannot express that at all.

  1. Sharing is natural. Two processes running the same program can have segment 0 entries with

the same base, so one copy of the code serves both. That is Chapter fourteen of Module 1's shared text, implemented.

The third has a catch worth knowing: a shared code segment must be the same segment number in every process that shares it, because a jump inside the code contains a segment number. If it is segment 0 in one process and segment 2 in another, the jump goes to the wrong place.

What segmentation is bad at, and why paging replaced it

Segments are variable sized, so segmentation has external fragmentation, exactly as Chapter seventy's variable partitions do, and for the same reason: a hole must be found that is big enough for the whole segment.

SegmentationVariable partitions
Allocation unitone segmentone whole process
Number of units per processseveralone
External fragmentationyesyes
Internal fragmentationnonea little
munotes.in287

Segmentation

Segmentation is better than one block per process, because the pieces are smaller and a smaller piece is easier to place. It is not a cure, because the pieces are still variable in size. Paging's insight is to make every piece the same size, and that removes external fragmentation entirely.

Which is why real machines did both: segmentation with paging, where each segment is itself paged. The Intel architecture worked exactly that way, and modern systems have dropped the segmentation half and kept the paging.

Distinctions that carry marks

SegmentationPaging (Chapter seventy four)
Pieces arevariable sized, and meaningful to the programmerfixed sized, and meaningless to the programmer
The address issegment number and offset, supplied by the programpage number and offset, computed from one number
External fragmentationyesnone
Internal fragmentationnonehalf a page per process
The programmer is aware of ityesno
Table entry holdsbase and limita frame number, and no limit

The last two rows are the deepest difference and a question about it wants them. A segment has a limit because segments differ in size; a page needs none because every page is full. And the programmer chooses the segments, while pages are cut by the hardware wherever the page size falls.

Segment tableThe relocation register of Chapter sixty six
Entriesone per segmentone
Protectionper segment, and can differone rule for the whole process
A growing stackconstrained only by its own limitconstrained by the whole process's limit

What it does not mean

A segment is not a page. A segment is a logical piece of a program, of whatever size that piece is. A page is a fixed sized piece of an address space with no meaning at all.

Segmentation is not obsolete as an idea. Every program still has a code segment, a data segment and a stack, and the permissions on them are still different. What is obsolete is implementing those divisions with a segment table rather than with page permissions.

The segment number is not computed by the hardware. The program supplies it. That is the property that makes segmentation visible to the programmer and paging invisible.

Per segment limits do not remove the need for a stack limit. They give the stack its own limit, which is a much better thing than sharing the process's.

Quick revision

  • Segmentation divides a program into its natural, variable sized pieces: code, data,

heap, stack, library.

  • A logical address is a segment number and an offset, and the program supplies both.
  • Each segment table entry holds a base and a limit, where the limit is a length.
munotes.in288

Segmentation

Two registers give the table's position and its length.

  • Translation: check the segment number against the table length; check the offset is

strictly less than the limit, or trap with an addressing error; then the physical address is base plus offset.

  • Gains: each segment grows independently, protection is per segment and can differ, and

sharing one segment between processes is natural.

  • A shared segment must have the same segment number in every process, because a jump carries

one.

  • Segments are variable sized, so segmentation has external fragmentation. That is why

paging, whose pieces are all the same size, replaced it.

Test yourself

  1. What is a segment, and what does a logical address look like? A logical, variable sized

piece of a program such as the code, the stack or an array. The address is a segment number together with an offset within that segment, both supplied by the program.

  1. What does each segment table entry hold? A base, the physical address where the segment

begins, and a limit, the length of the segment.

  1. Give the translation steps. Check the segment number is less than the segment table length

register; check the offset is strictly less than the segment's limit, trapping with an addressing error if not; then add the base to the offset.

  1. Segment 2 has base 4300 and limit 400. Translate offsets 399 and 400. 399 is below the

limit, so 4300 + 399 = 4699. 400 is not strictly below 400, so it traps.

  1. Give three advantages of segmentation over one block per process. Each segment grows

independently, so a growing stack cannot damage the heap; protection is per segment and can differ, so code can be read only and data non executable; and one segment can be shared between processes with the same base in each.

  1. What must be true of a shared segment's number, and why? It must be the same number in

every process that shares it, because a jump within the code carries a segment number with it.

  1. Why did paging replace segmentation? Segments are variable sized, so a hole big enough for

a whole segment must be found and external fragmentation results. Paging makes every piece the same size, so any free frame fits any page and external fragmentation disappears.

Contents This chapter on its own page

munotes.in289

Chapter Seventy-Four

Paging

Syllabus topic Module 2, "Memory Management - Paging"

In one line

Cut physical memory into fixed sized frames, cut every process's address space into pages of exactly the same size, and then any page can go in any frame.

The idea, from the problem

Chapter seventy three ended with the fault: segments are variable sized, so a hole big enough for the whole segment has to be found, and external fragmentation follows.

Paging's insight is one sentence: make every piece the same size, and then no piece can be in the wrong place. There is no such thing as a hole being too small or a frame being unsuitable. Every free frame fits every page.

And the price, from Chapter seventy two: the last page of a process is not full, so internal fragmentation appears, bounded at half a page per process on average. Paging trades an unbounded problem for a bounded one.

The two words

WordIs a fixed sized piece ofTypical size
Framephysical memory4 kilobytes
Pagea process's logical address spacethe same 4 kilobytes

A page and a frame are the same size, always, by definition. That is not a coincidence of implementation: it is what makes any page fit any frame, and it is the whole mechanism. A question that suggests they could differ has misunderstood the scheme.

Frames are physical and pages are logical, and getting the two words the right way round is worth marks. The sentence to hold: a page goes into a frame.

The page table

One entry per page of the process, and the entry holds the frame number that page is in.

One tableper process: each process has its own
One entryper page of that process's address space
The entry holdsthe frame number, and some bits: valid, read only, modified, referenced
Where the table livesin memory, pointed at by a register the kernel loads at a context switch

A page table entry has no limit field, unlike Chapter seventy three's segment table, and the reason is worth giving: every page is exactly full, so there is nothing to bound. That single difference is the cleanest way to tell the two schemes apart in an answer.

The register that points at the table is the page table base register, and loading it is what makes a context switch expensive: it changes every translation the machine will do next.

The translation

The logical address is split in two, and the split is decided by the page size and nothing else.

PartIsUsed to
page number, pthe high bitsindex the page table
offset, dthe low bitssay how far into the page, and it is not translated at all
munotes.in290

Paging

The three steps:

  1. Split the logical address into p and d.
  2. Look up entry p in the page table to get the frame number f.
  3. Join f to d: the physical address is f times the page size, plus d.

The offset passes through untouched, and that is the property that makes the whole thing work. The page size is a power of two, so the low bits of a logical address are already the offset within the page, and the low bits of the physical address are the same number. Nothing has to be added or carried. Chapter seventy five works the arithmetic.

The page size is always a power of two, and that is not a convention. It is what makes the split a matter of which bits rather than a division.

What the process knows

Nothing. A process's addresses are 0 upwards, contiguous as far as it can tell, and it has no way of discovering that page 3 is in frame 8 and page 4 in frame 1,207. Chapter sixty seven proved it cannot even read the mapping.

Compare segmentation, where the program supplies the segment number and therefore knows its program is in pieces. Paging is invisible to the program and segmentation is not, and that is the second great difference between them.

What the operating system must keep

Two structures, and a question asks for both.

StructureOne perHolds
Page tableprocesswhich frame each of its pages is in
Frame tablethe machinefor each frame: whether it is free, and if not, which process and which page is in it

The frame table is the free list of Chapter seventy, reduced to something trivial: every frame is the same size, so the free list is just a list of numbers, or a bit per frame. There is no "find a hole big enough" step at all, which is the second cost paging removes.

Worked: a tiny machine

The standard small example, which makes the whole scheme visible.

A logical address space of 16 bytes, a page size of 4 bytes, so 4 pages; and a physical memory of 32 bytes, so 8 frames.

Suppose the page table is:

PageFrame
05
16
21
32

Now translate. The page size is 4, so the page number is the address divided by 4 and the offset is the remainder.

Logicalp = address / 4d = address mod 4framephysical = frame x 4 + d
00055 × 4 + 0 = 20
30355 × 4 + 3 = 23
41066 × 4 + 0 = 24
92111 × 4 + 1 = 5
133122 × 4 + 1 = 9
munotes.in291

Paging

Read logical 3 and logical 4. They are next to each other in the program and land at 23 and 24, which are next to each other by luck. Now read logical 4 and logical 9: adjacent pages, landing at 24 and 5, thousands of bytes apart on a real machine. The process cannot tell, and does not care.

And notice the frame table: frames 5, 6, 1 and 2 are used, and frames 0, 3, 4 and 7 are free. A fifth page would go in any of them, with no search and no fit.

What paging costs

Three costs, and each is a later chapter.

CostSizeAnswered by
Internal fragmentationhalf a page per processaccepted, and it is why 4 kilobytes not 4 megabytes
A second memory access per access, to read the tabledoubles the cost of every accessthe translation cache, Chapter seventy six
The page table itself takes memory4 megabytes per process on a 32 bit machinehierarchical paging, Chapter seventy eight

The third is the one that surprises students and it is a standard sum. A 32 bit address space with 4 kilobyte pages has 2 to the power 20 pages, which is 1,048,576 entries; at 4 bytes each that is 4 megabytes of page table, per process. Chapter seventy eight works it out and fixes it.

Distinctions that carry marks

PageFrame
A piece ofthe logical address spacephysical memory
Belongs toa processthe machine
Same size asthe framethe page
Numbered from0, per process0, for the whole machine
PagingSegmentation
Piecesfixed sizevariable size
Address supplied by the programone number, split by the hardwaretwo numbers, segment and offset
Visible to the programnoyes
Table entrya frame number, no limita base and a limit
External fragmentationnoneyes
Internal fragmentationhalf a page per processnone
Page tableFrame table
One perprocessmachine
Entry sayswhich frame this page is inwhether this frame is free, and whose page is in it

What it does not mean

A page is not a piece of the program in any meaningful sense. The boundary falls wherever the page size falls, possibly in the middle of a function or an array. That is why the programmer is not told about it.

Paging does not need the process's pages to be in order in memory. They can be in any frames at all, which is the point.

munotes.in292

Paging

The offset is not translated. Only the page number is looked up. The offset is copied straight through.

Paging is not virtual memory. Paging is a way of placing a process's pages in frames, and under plain paging every page is in memory. Keeping only some of them in memory is virtual memory, Chapter eighty.

Quick revision

  • Paging cuts physical memory into frames and each address space into pages of

exactly the same size. Any page fits any frame, so there is no external fragmentation and no search for a hole.

  • A frame is physical, a page is logical, and a page goes into a frame.
  • Each process has its own page table, one entry per page, holding the frame number.

There is no limit field, because every page is full.

  • The machine has one frame table, saying which frames are free and whose page is in each.
  • The logical address splits into a page number and an offset; the page number is looked

up and the offset passes through untouched. The page size is a power of two so the split is a matter of bits.

  • Paging is invisible to the program; segmentation is not.
  • Three costs: internal fragmentation of half a page per process; a second memory access

per access, fixed by Chapter seventy six; and the page table's own size, 4 megabytes per process on a 32 bit machine, fixed by Chapter seventy eight.

Test yourself

  1. Define a page and a frame. A frame is a fixed sized piece of physical memory; a page is a

piece of a process's logical address space of exactly the same size. A page is placed into a frame.

  1. Why does paging have no external fragmentation? Every page is the same size as every

frame, so any free frame satisfies any page: no frame can be the wrong size or in the wrong place.

  1. What does a page table entry hold, and what does it not hold? The frame number that page

occupies, plus status bits. It holds no limit, because every page is exactly full.

  1. Give the three translation steps. Split the logical address into a page number and an

offset; look the page number up in the page table to get a frame number; join the frame number to the untranslated offset.

  1. Why must the page size be a power of two? So that the split into page number and offset is

a matter of which bits an address has, needing no division, and so the offset can be copied straight through. 6. A 16 byte address space, 4 byte pages, and page 2 is in frame 1. Where does logical address 9 land? 9 divided by 4 is page 2 with offset 1; frame 1 times 4 plus 1 gives physical address 5.

munotes.in293

Paging

  1. Name the two tables the operating system keeps and what each is for. A page table per

process, saying which frame each of its pages is in; and one frame table for the machine, saying which frames are free and which page of which process is in each of the rest.

Contents This chapter on its own page

munotes.in294

Chapter Seventy-Five

Splitting an Address, and Translating One

Syllabus topic Module 2, "Memory Management - Paging"

In one line

The page size decides how many bits are the offset, the address size decides the rest, and then a translation is one lookup and one join.

The one rule everything follows from

The number of offset bits is log to base 2 of the page size, and nothing else affects it.

Page sizeOffset bitsBecause
512 bytes92 to the power 9 is 512
1 kilobyte102 to the power 10 is 1024
2 kilobytes11
4 kilobytes12the usual size on every modern machine
8 kilobytes13
1 megabyte20a large page, used deliberately in some systems

And then:

page number bits = address bits - offset bits

number of pages = 2 to the power (page number bits)

In this book, and in every operating systems paper, 1 kilobyte means 1024 bytes. A page size is a power of two by definition, so it could not be 1000. A student who works in thousands gets every bit count wrong.

The standard sums, four of them

One: a 16 bit address space with 1 kilobyte pages

QuantityValueWorking
address bits16given
address space65536 bytes2 to the power 16
page size1024 bytesgiven
offset bits102 to the power 10 is 1024
page number bits616 - 10
number of pages642 to the power 6

Two: a 32 bit address space with 4 kilobyte pages

This is the one to know by heart, because every later chapter uses it.

QuantityValueWorking
address bits32given
address space4 gigabytes2 to the power 32
page size4096 bytesgiven
offset bits122 to the power 12 is 4096
page number bits2032 - 12
number of pages10485762 to the power 20

One million pages, and therefore one million page table entries, per process. Chapter seventy eight is what to do about it.

Three: given the number of pages, find the page size

A question sometimes runs the other way. A 24 bit address space is divided into 4096 pages. How big is a page?

page number bits = 12

offset bits = 24 - 12 = 12

page size = 4096

Because 4096 pages needs 12 bits to number them, and the remaining 12 bits are the offset.

Four: physical memory is a different size from the address space

The logical and physical address spaces need not be the same size, and this is the case students get wrong.

A machine with a 32 bit logical address, 4 kilobyte pages, and 1 gigabyte of physical memory.

QuantityValueWorking
offset bits12from the page size
page number bits2032 - 12
pages per process10485762 to the power 20
physical memory1073741824 bytes1 gigabyte
number of frames2621441073741824 / 4096
frame number bits182 to the power 18 is 262,144
physical address bits3018 + 12
munotes.in295

Splitting an Address, and Translating One

Twenty bits of page number and eighteen bits of frame number. The process can address four times more than the machine has, and every one of its pages must be in one of the 262,144 frames or not in memory at all. That gap is where virtual memory lives, and a question that gives you different logical and physical sizes is pointing at it.

Translating an address

Two steps of arithmetic, and it is always the same two.

page number = logical address / page size

offset = logical address mod page size

physical address = frame number x page size + offset

Worked

A page size of 4 kilobytes, and page 3 is in frame 8. Translate logical address 13452.

13452 = 3 × 4096 + 1164

page number = 3

offset = 1164

physical = 8 × 4096 + 1164 = 33932

Check it the other way as a habit: 33932 divided by 4096 is 8 remainder 1164. The frame is 8 and the offset came through unchanged, which is the property of Chapter seventy four.

The same thing in binary, which is what the hardware does

13452 in 16 bits is 0011 0100 1000 1100.

FieldBitsValue
page number, top 4 bits of a 16 bit address with 12 offset bits00113
offset, bottom 12 bits0100 1000 11001164

The frame number 8 is 1000. Join it to the same offset:

FieldBitsValue
frame number10008
offset, unchanged0100 1000 11001164

Together: 1000 0100 1000 1100, which is 33932.

No addition happened. The hardware replaced the top bits and left the bottom bits alone. That is why the page size must be a power of two, and it is why a translation costs nothing once the frame number is known.

How big is a page table

The other standard sum, and it is the one that motivates the next four chapters.

entries = number of pages

table size = entries x bytes per entry

For a 32 bit address space, 4 kilobyte pages and a 4 byte entry:

entries = 1048576

table size = 1048576 × 4 = 4194304

4,194,304 bytes, which is 4 megabytes of page table per process.

And the fact that makes it worse: the page table must itself be in memory, so at 4 kilobytes a page it takes 1,024 pages to hold it.

munotes.in296

Splitting an Address, and Translating One

pages to hold the table = 4194304 / 4096 = 1024

A hundred processes would then need 400 megabytes of page tables before a single byte of anybody's data. That is the problem Chapter seventy eight solves.

The trap in choosing a page size

Every question about page size is this trade, and it has four terms.

Larger pagesSmaller pages
fewer entries, so a smaller page tablemore entries, so a bigger table
more internal fragmentation, half a page per processless internal fragmentation
fewer page faults for a program that reads sequentiallymore faults
a page fault transfers more data, so each one costs moreeach fault is cheaper

4 kilobytes has been the usual answer for forty years, and modern machines also offer huge pages of 2 megabytes or 1 gigabyte, used deliberately for programs with very large data, precisely to shrink the page table and the translation cache pressure.

Distinctions that carry marks

Page numberOffset
Which bitsthe high bitsthe low bits
Decided bywhat is left after the offsetthe page size
Translatedyes, looked up in the page tableno, copied through
Its size fixeshow many pages there can behow big a page is
Logical address bitsPhysical address bits
Fixed bythe architecture, for example 32the amount of memory fitted
Split intopage number and offsetframe number and the same offset
Can be larger than the otheryes, and usually is

What it does not mean

The page number is not a memory address. It is an index into the page table.

The offset is not an address either. It is a distance into a page, and it is the same distance into the frame.

A larger page size is not simply better. It shrinks the page table and grows the internal fragmentation, and makes each page fault transfer more.

The physical address space is not the amount of memory a process can use. A process can address its whole logical space; how much of it is in memory at once is Chapter eighty.

Quick revision

  • Offset bits = log base 2 of the page size. 4 kilobytes gives 12. Everything else

follows.

  • Page number bits = address bits - offset bits, and the

number of pages = 2 to that power.

  • A 32 bit address with 4 kilobyte pages gives 12 offset bits, 20 page bits and

1,048,576 pages.

  • Frames = physical memory / page size, and the frame number bits follow from that. The

logical and physical spaces are usually different sizes.

  • Translate: page = address / page size, offset = address mod page size,
munotes.in297

Splitting an Address, and Translating One

physical = frame x page size + offset.

  • In binary the top bits are replaced and the bottom bits are untouched. No addition happens.
  • A page table for a 32 bit space with 4 byte entries is 4 megabytes per process, needing

1,024 pages to hold it.

  • Page size trade: larger means a smaller table and more internal fragmentation and dearer

faults; 4 kilobytes is the usual answer, with huge pages of 2 megabytes for large data.

Test yourself

  1. How many offset bits does a 2 kilobyte page need, and why? Eleven, because 2 to the power

11 is 2048. 2. A 32 bit address space with 4 kilobyte pages. Give the offset bits, the page number bits and the number of pages. Twelve offset bits, twenty page number bits, and 2 to the power 20, which is 1,048,576 pages.

  1. A 24 bit address space is divided into 4096 pages. How big is a page? 4096 pages need 12

bits to number, so 12 bits remain for the offset and the page size is 2 to the power 12, which is 4096 bytes. 4. A 32 bit logical address, 4 kilobyte pages, 1 gigabyte of memory. How many frames and how many frame number bits? 1,073,741,824 divided by 4096 gives 262,144 frames, which needs 18 bits.

  1. Page size 4 kilobytes and page 3 is in frame 8. Translate logical address 13452. 13452 /

4096 is page 3 with offset 1164; 8 times 4096 plus 1164 gives 33,932. 6. How big is a page table for a 32 bit address space with 4 kilobyte pages and 4 byte entries? 1,048,576 entries times 4 bytes, which is 4,194,304 bytes: 4 megabytes per process, needing 1,024 pages to hold it.

  1. Give two effects of making the page size larger. The page table has fewer entries and is

smaller; and the internal fragmentation, half a page per process, grows, while each page fault transfers more data.

Contents This chapter on its own page

munotes.in298

Chapter Seventy-Six

The TLB, and the Effective Access Time

Syllabus topic Module 2, "Memory Management - Paging"

In one line

A small associative cache of recent translations turns two memory accesses back into one for the pages a program is actually using, and the effective access time is the weighted average of the hit and the miss.

The problem, stated as a number

Paging doubles the cost of every memory reference. The page table is in memory, so reading one byte of your own data means:

StepCost
read the page table entry to find the frameone memory access
read the byteone memory access
totaltwo

With a memory access of 100 nanoseconds, a paged machine with no help would take 200 nanoseconds for every reference: twice as slow, on every instruction, for the whole life of the machine. No design survives that, and every real paged machine has the hardware in this chapter.

It is worse than twice for a two level page table, which Chapter seventy eight needs: three accesses, 300 nanoseconds, three times as slow.

The translation look aside buffer

The TLB is a small, fast, associative cache of page number to frame number pairs. It sits inside the memory management unit, it is searched for all its entries at once, and the search costs a fraction of a memory access.

PropertyTypical valueWhy it is that
number of entries64 to 1024associative search is expensive in hardware, so it must stay small
what one entry holdsa page number and its frame numberthe same pair the page table holds
how it is searchedall entries in parallela sequential search would cost what it saves
search time10 to 20 nanoseconds, a fraction of a memory accessit is registers, not memory

The rule for one reference:

The page number isThen
in the TLB (a hit)the frame number comes straight out, and one memory access follows
not in the TLB (a miss)the page table is read from memory as usual, and the pair is then added to the TLB, replacing an old entry if the TLB is full

A small table works because of locality of reference: a program spends its time in a few pages, not all of them. This is the same fact that makes Chapter eighty nine's working set idea work, and it is the reason 64 entries can cover 98 per cent of the references of a program with a million pages.

The effective access time

The formula, which is a weighted average and nothing cleverer.

hit cost = TLB search + one memory access

miss cost = TLB search + page table access + one memory access

EAT = hit ratio × hit cost + (1 - hit ratio) × miss cost

munotes.in299

The TLB, and the Effective Access Time

The TLB search is paid on a miss too. The hardware looks in the TLB, does not find the page, and then goes to memory: the search time is spent either way. A student who leaves it out of the miss cost gets every EAT question slightly wrong.

Worked, with a 20 nanosecond TLB and a 100 nanosecond memory

hit cost = 20 + 100 = 120

miss cost = 20 + 100 + 100 = 220

Hit ratioEATWorkingSlower than unpaged memory by
100 per cent1201.00 × 12020 per cent
99 per cent1210.99 × 120 + 0.01 × 220 = 118.8 + 2.2 = 12121 per cent
98 per cent1220.98 × 120 + 0.02 × 220 = 117.6 + 4.4 = 12222 per cent
90 per cent1300.90 × 120 + 0.10 × 220 = 108 + 22 = 13030 per cent
80 per cent1400.80 × 120 + 0.20 × 220 = 96 + 44 = 14040 per cent
0 per cent2201.00 × 220120 per cent

Read the last two rows together. Without the TLB the machine is 120 per cent slower; with a 98 per cent hit ratio it is 22 per cent slower. That is what the cache is worth, and real machines do reach 98 and 99 per cent.

The same sum with a faster TLB

A 10 nanosecond search, so a hit costs 110 and a miss 210.

Hit ratioEATWorking
80 per cent1300.80 × 110 + 0.20 × 210 = 88 + 42 = 130
98 per cent1120.98 × 110 + 0.02 × 210 = 107.8 + 4.2 = 112
99 per cent1110.99 × 110 + 0.01 × 210 = 108.9 + 2.1 = 111

Eleven per cent slower than a machine with no paging at all, for the whole benefit of paging. That is the bargain the chapter exists to explain.

Solving it backwards

A question can give the EAT and ask for the hit ratio. It is one line of algebra, and the marks are for doing it rather than guessing.

A 20 nanosecond TLB, a 100 nanosecond memory, and an EAT of 130 nanoseconds. What is the hit ratio?

130 = r × 120 + (1 - r) × 220

130 = 220 - 100r

100r = 90

r = 0.9

Ninety per cent, which the table above confirms.

The general form, worth memorising because it saves the algebra every time:

EAT = miss cost - hit ratio × (miss cost - hit cost)

munotes.in300

The TLB, and the Effective Access Time

The bracket is what one hit saves, which here is one memory access.

What the operating system has to do about the TLB

Two duties, and both are asked as short questions.

One: a context switch. The TLB holds one process's translations. The next process has different pages in different frames, so the entries are wrong for it.

ApproachWhat happensCost
Flush the TLBevery entry is thrown away at each switchthe new process starts with a miss on every page, so the first hundreds of references are slow
Tag each entry with an address space identifierentries for both processes live together, and a hit needs the tag to match tooneeds hardware for the tag, and it is what modern machines do

This is a real part of the cost of a context switch named in Chapter seventy two: not only the registers, but a cold translation cache.

Two: changing a page table entry. If the operating system moves a page to another frame, or takes it away, the TLB may still hold the old pair. The entry must be invalidated in the TLB, by hand. A kernel that forgets is a kernel where a process reads a frame that no longer belongs to it, which is the most serious class of bug in a memory manager.

Distinctions that carry marks

TLBPage table
Where it livesinside the MMU, in fast hardwarein main memory
How big64 to 1024 entriesone entry per page, up to a million
Holdssome translations, the recent onesevery translation for the process
Searchedall entries at once, associativelyindexed by page number
If the entry is absentgo to the page tablethe address is invalid, and it is a trap
Hit ratioEAT
What it isthe fraction of references found in the TLBthe average time one reference costs
Unitsnone, a fraction or a percentagenanoseconds
Given in a questionusuallyasked for, or given so the ratio can be found

What it does not mean

The TLB does not hold data. It holds translations. The data cache is a different cache, and a reference can hit in one and miss in the other.

A TLB miss is not a page fault. A miss means the translation was not cached, and the page table has it: the cost is one extra memory access. A fault means the page is not in memory at all, and the cost is a disk access, which Chapter eighty two measures at a hundred thousand times more.

The hit ratio is not a property of the TLB alone. It is a property of the TLB size, the page size, and how the program moves through memory. The same hardware gives different ratios for different programs.

munotes.in301

The TLB, and the Effective Access Time

Paging is not free even with a perfect TLB. A 100 per cent hit ratio still costs the search: 120 against 100, 20 per cent. It buys what Chapter seventy four listed, and it is never nothing.

Quick revision

  • Paging makes every reference two memory accesses; at 100 nanoseconds each that is 200,

twice as slow.

  • The TLB is a small associative cache of page-to-frame pairs inside the MMU,

64 to 1024 entries, searched all at once in 10 to 20 nanoseconds.

  • On a miss the page table is read and the pair is loaded into the TLB.
  • EAT = hit ratio × hit cost + miss ratio × miss cost, where the hit cost = TLB + memory

and the miss cost = TLB + memory + memory. The TLB time is paid on a miss as well.

  • With a 20 nanosecond TLB and 100 nanosecond memory: 80 per cent gives 140,

98 per cent gives 122, no TLB gives 220.

  • Solve backwards with EAT = miss cost - hit ratio × (miss cost - hit cost).
  • A context switch must flush the TLB or tag its entries with an address space identifier; a

changed page table entry must be invalidated in the TLB by hand.

  • The TLB works because of locality of reference, which is also why the working set of

Chapter eighty nine is small.

Test yourself

  1. Why does paging double the cost of a memory reference? The page table is in memory, so one

access reads the entry to find the frame and a second reads the byte.

  1. What is in one TLB entry, and how is the TLB searched? A page number and its frame number;

every entry is searched at the same time, associatively, which is why the TLB must be small.

  1. A 20 nanosecond TLB, a 100 nanosecond memory, and an 80 per cent hit ratio. Find the EAT.

A hit costs 120 and a miss 220, so 0.80 × 120 + 0.20 × 220 = 96 + 44 = 140 nanoseconds.

  1. The same machine at a 98 per cent hit ratio. 0.98 × 120 + 0.02 × 220 = 117.6 + 4.4 = 122

nanoseconds, which is 22 per cent slower than unpaged memory instead of 120 per cent. 5. The EAT is 130 nanoseconds on a machine where a hit costs 120 and a miss 220. What is the hit ratio? The EAT is 220 minus 100 times the ratio, so 100 times the ratio is 90 and the ratio is 0.9, ninety per cent.

munotes.in302

The TLB, and the Effective Access Time

  1. Why is the TLB search time counted in the miss cost? Because the hardware searches the TLB

first whatever the outcome; on a miss that time is spent and the memory accesses follow.

  1. What must happen to the TLB at a context switch, and what is the alternative? It must be

flushed, because its entries belong to the old process; or every entry carries an address space identifier so entries for several processes can coexist.

  1. Distinguish a TLB miss from a page fault. A miss costs one extra memory access and the

page table answers it; a fault means the page is not in memory and a disk access is needed.

Contents This chapter on its own page

munotes.in303

Chapter Seventy-Seven

Protection and Sharing in a Paged System

Syllabus topic Module 2, "Memory Management - Paging"

In one line

Protection in a paged system is a few bits stored beside each frame number, checked by the hardware on every reference, and the same mechanism lets two processes share one copy of a program.

Why the question has to be asked

Chapter sixty eight protected a process with a base and a limit register: one pair of registers, one continuous region, and any address outside it is a trap. Paging has no limit register. A process's pages are scattered through memory, and the hardware sees only a page number and an offset.

So where does protection come from? From the page table itself. The page table is the only route from a logical address to a frame, it is in kernel memory, and a process cannot change it. A frame that is not named in your page table cannot be reached by any address you can write. That is the whole of the protection, and it is stronger than a limit register because it is per page rather than per process.

The bits beside each entry

Each page table entry holds a frame number and, next to it, a few bits that the hardware checks before it lets the reference through.

BitSet whenWhat the hardware does if the reference breaks it
Valid and invalidthe page is in the process's legal address spacetraps to the operating system: this is the classic segmentation fault
Read and writethe page may be writtentraps on a store to a read-only page
Executeinstructions may be fetched from the pagetraps on a jump into data, which stops a large family of attacks
Modify, also called dirtyset by the hardware when the page is writtennothing; Chapter eighty four reads it to decide whether the page must be written back
Referenceset by the hardware when the page is usednothing; Chapter eighty seven's algorithms read and clear it

Two of these bits are set by the hardware and read by the operating system, not the other way round. The modify bit and the reference bit are the hardware's report on the program's behaviour, and the next module is built on them. The rest are set by the operating system and enforced by the hardware.

Read and write and execute are checked on every reference, by the hardware, in parallel with the translation. They cost nothing in time. A student who says the operating system checks them has described a machine that would run a thousand times slower: the kernel is not consulted on a memory reference.

The valid and invalid bit, worked

The standard sum. A 14 bit address space and a 2 kilobyte page size, so 8 pages, numbered 0 to 7.

munotes.in304

Protection and Sharing in a Paged System

address space = 16384

page size = 2048

pages = 16384 / 2048 = 8

A process whose program and data occupy addresses 0 to 10468. Which pages are valid?

10468 / 2048 = 5, remainder 228

PagesBitBecause
0, 1, 2, 3, 4validwholly inside the program
5validholds addresses 10240 to 12287, and the program ends at 10468 inside it
6, 7invalidwholly outside the program: a reference to any of them traps

Page 5 is the internal fragmentation of Chapter seventy two, made visible. Addresses 10469 to 12287 are inside a valid page, so the hardware cannot trap a reference to them. The process can read and write 1,819 bytes it never asked for. Paging protects at the granularity of a page, and that is its one weakness: it cannot protect part of a page.

Some machines add a page table length register, holding the number of entries the process actually has. A page number at or beyond it traps without the table being read at all, which saves the memory of entries 6 and 7 for a process with a small address space and is how a one million entry table stays small in practice for most processes.

Sharing

The second half of the chapter, and the part with a sum in it. If two processes have the same frame number in their page tables, they are using the same physical memory. Nothing else is needed.

What may be shared

Part of a programSharableWhy
Code, if it is reentrantyesreentrant code never modifies itself, so one copy serves every process
Read-only data, tables, fontsyesnobody writes to it
Data, stack, heapno, each process needs its owntwo processes writing one copy would corrupt each other
Data, deliberately sharedyes, and this is shared memory as a form of inter-process communicationChapter thirty seven of Module 1

Reentrant means the code does not change itself while it runs, so it can be marked read-only and entered by many processes at once. Each process has its own copy of the data pages and its own registers, so each is at its own place in the same instructions.

The sum

40 students editing at the same time, on a system with a 2 kilobyte page size. The editor is 150 kilobytes of reentrant code and each user needs 50 kilobytes of data.

WorkingMemory needed
Without sharing40 × 200 = 80008,000 kilobytes
With sharing150 + 40 × 50 = 150 + 2000 = 21502,150 kilobytes
Saved8000 - 2150 = 58505,850 kilobytes
munotes.in305

Protection and Sharing in a Paged System

Nearly six megabytes saved on one program, and the saving grows with every extra user: each new user costs 50 kilobytes instead of 200. In pages, the shared code is 75 pages held once, and each user's data is 25 pages of their own.

code pages = 150 / 2 = 75

data pages per user = 50 / 2 = 25

The shared pages need not be at the same page number in each process. Process A may have the editor at pages 0 to 74 and process B at pages 10 to 84. Only the frame numbers must agree. This is the answer to a question that sounds harder than it is.

What the operating system must keep to make this safe

Two records, and a question on shared memory wants both.

RecordWhy
the frame table, saying which processes hold each framea shared frame must not be freed when one of its holders exits
a count of how many page tables name the framethe frame is free only when the count reaches zero

This is the same reference count that Chapter eighty three's copy on write uses, and the same idea as the link count of a file in Chapter fifty six hundred one.

Distinctions that carry marks

Protection in contiguous allocationProtection in paging
Mechanismbase and limit registersbits in the page table entry
Granularitythe whole processone page
Read-only regionsnot possible without extra hardwarea bit per page
A reference past the endtrapped exactlytrapped only at the next page boundary
Sharinghard: one region, one ownereasy: the same frame number in two tables
Valid and invalid bitModify bit
Set bythe operating systemthe hardware, on a write
Read bythe hardware, on every referencethe operating system, when replacing the page
Purposeprotection, and later whether the page is in memoryto know whether the page must be written back

What it does not mean

An invalid page is not always an error. In Chapter eighty one the same bit means the page is legal but not in memory, and the trap is serviced instead of killing the process. The bit is one bit with two uses, and the operating system knows which by looking at its own records.

A read-only page is not a read-only file. The bit protects a frame in memory, not anything on disk.

Sharing does not mean two processes see each other's variables. Only the frames named in both tables are shared. Everything else is private, and the sharing is arranged deliberately.

Paging does not protect inside a page. The last page of a process is partly beyond its data and the hardware cannot tell, which is exactly the internal fragmentation the scheme accepts.

munotes.in306

Protection and Sharing in a Paged System

Quick revision

  • Paging has no limit register. Protection is that

a frame not named in your page table has no address you can write, plus the bits in each entry.

  • The bits: valid and invalid, read and write, execute, modify and reference.

The last two are set by the hardware and read by the operating system.

  • The valid bit is checked by the hardware on every reference, in parallel with the

translation, and costs nothing.

  • 14 bit space, 2 kilobyte pages, a program ending at 10468: pages 0 to 5 valid,

6 and 7 invalid, and 1,819 bytes of page 5 are reachable but unused. Paging cannot protect part of a page.

  • A page table length register traps a page number beyond the table without reading it.
  • Sharing is the same frame number in two page tables; reentrant code, which never

modifies itself, can be shared, and private data cannot.

  • 40 users, a 150 kilobyte editor, 50 kilobytes of data each:

8,000 kilobytes unshared against 2,150 shared, a saving of 5,850.

  • Shared pages need not have the same page number in each process; only the frame numbers

must agree.

  • A shared frame needs a count of how many tables name it, and is freed only at zero.

Test yourself

  1. Paging has no limit register. How is one process kept out of another's memory? Its page

table is the only route to a frame, it is kept by the kernel, and a frame not named in it cannot be reached by any address the process can form.

  1. Name the five bits beside a page table entry, and say which the hardware sets. Valid and

invalid, read and write, execute, modify and reference. The hardware sets the modify bit on a write and the reference bit on any use; the operating system sets the other three. 3. A 14 bit address space, 2 kilobyte pages, and a program occupying 0 to 10468. Which pages are invalid, and what is the flaw? Pages 6 and 7 are invalid. Page 5 is valid but the program ends inside it, so 1,819 bytes beyond the program can be touched without a trap: paging protects only whole pages.

  1. What does a page table length register do? It holds the number of entries the process has,

so a page number at or beyond it traps immediately, without reading the table.

  1. What makes code sharable, and what must each process still have of its own? The code must

be reentrant, that is it must never modify itself, so it can be read-only; each process keeps its own data pages and its own registers. 6. 40 users, a 150 kilobyte reentrant editor, 50 kilobytes of data each. Memory with and without sharing. Without: 40 × 200 = 8,000 kilobytes. With: 150 + 40 × 50 = 2,150 kilobytes. The saving is 5,850 kilobytes.

munotes.in307

Protection and Sharing in a Paged System

  1. Must a shared page have the same page number in both processes? No. Only the frame number

must be the same; each process may see it at a different page number.

  1. Why does the operating system count how many page tables name a frame? So that a shared

frame is freed only when the last process holding it has finished with it.

Contents This chapter on its own page

munotes.in308

Chapter Seventy-Eight

The Page Table Is Too Big: Hierarchical Paging

Syllabus topic Module 2, "Memory Management - Structure of the Page Table"

In one line

A page table too big to keep in one piece is paged itself, so that only the parts a process actually uses have to be in memory.

The problem, as a number

From Chapter seventy five: a 32 bit address space, 4 kilobyte pages, a 4 byte entry.

entries = 1,048,576

table size = 1,048,576 × 4 = 4,194,304

pages to hold it = 4,194,304 / 4096 = 1024

Four megabytes per process, and it must be 1,024 contiguous frames. Read the second half again. The page table is looked up by index, so it has to be one continuous array, and the whole point of Chapter seventy four was that the operating system no longer has to find a large continuous region for anything. A paging scheme whose own table needs a four megabyte hole has defeated itself.

And the waste: a process that uses 100 kilobytes of memory still has 1,048,576 entries, of which about 25 are valid. The table is nearly all zeroes.

The answer: page the page table

Split the page number itself. The outer part indexes a table of tables, the inner part indexes the table it finds, and only the inner tables for the parts of the address space the process is really using need to exist.

The Linux kernel documentation puts the general rule in one line: "The page tables are organized hierarchically."

The standard 32 bit split

Choose the inner field so that one inner table fits exactly in one page, because then every piece of the structure is a page and every piece can be paged.

entries in one page = 4096 / 4 = 1024

inner field = 10 bits

outer field = 32 - 12 - 10 = 10

FieldBitsWhat it indexes
p1, the outer page number10the outer page table, 1024 entries, each naming an inner table
p2, the inner page number10one inner page table, 1024 entries, each naming a frame
d, the offset12the byte inside the frame

This is called a forward mapped page table, and a two level table is also called a paged page table.

PieceSizeCovers
the outer table1024 × 4 = 4096 bytes, one pagethe whole 4 gigabyte space
one inner table1024 × 4 = 4096 bytes, one page1024 × 4096 = 4,194,304 bytes, 4 megabytes of address space

What it saves

A process using 8 megabytes of address space needs the outer table and two inner tables.

memory for the tables = 3 × 4096 = 12,288

12,288 bytes against 4,194,304. And none of the three pages has to be next to any other, which was the harder half of the problem.

munotes.in309

The Page Table Is Too Big: Hierarchical Paging

The saving is not free: the outer table is 4 kilobytes even for a process that uses one page, so a very small process pays a little more than it would under a flat table it could afford. The scheme is built for the common case, which is a large sparse address space.

Translating, worked

Address 4,206,596 on the two level machine above.

d = 4

page number = 1027

check = 1027 × 4096 + 4 = 4,206,596

p1 = 1

p2 = 3

StepWhat the hardware does
1take p1 = 1, read entry 1 of the outer table, and find the physical address of an inner table
2take p2 = 3, read entry 3 of that inner table, and find the frame number
3join the frame number to d = 4 and make the physical address

Two memory accesses before the data, so three in all. Every level added costs one more memory access, which is why the level count is kept as low as the arithmetic allows.

What the extra level costs

With a 100 nanosecond memory and a 20 nanosecond TLB:

Accesses to memoryCost with no TLB
no paging1100
flat page table2200
two level page table3300

And with a TLB at a 98 per cent hit ratio, where a hit still costs 120 and a miss now costs 320:

EAT = 0.98 × 120 + 0.02 × 320 = 117.6 + 6.4 = 124

124 nanoseconds against 122 for the flat table. The second level costs two nanoseconds on average, and saves four megabytes per process. That trade is why every real machine makes it, and it is worth saying in a question in exactly those terms.

Why two levels are not enough for a 64 bit machine

The same arithmetic, on a 64 bit address space with 4 kilobyte pages and an 8 byte entry, which is what a 64 bit machine needs.

offset bits = 12

page number bits = 64 - 12 = 52

entries in one page = 4096 / 8 = 512

inner field = 9 bits

outer field = 52 - 9 = 43

The outer table would then hold 2 to the power 43 entries:

outer entries = 8,796,093,022,208

outer table = 8,796,093,022,208 × 8 = 70,368,744,177,664

in terabytes = 70,368,744,177,664 / 1,099,511,627,776 = 64

Sixty four terabytes for the outer table alone. Two levels does not help a 64 bit address space at all: the outer table is now the thing that will not fit, and it has to be paged in its turn. Splitting again and again, nine bits at a time:

munotes.in310

The Page Table Is Too Big: Hierarchical Paging

levels = 7 + 9 + 9 + 9 + 9 + 9 = 52

Six levels, and therefore seven memory accesses for one unmatched reference. That is the argument that sends the next chapter looking for a different structure altogether.

What this machine actually does

A 64 bit machine does not have to implement all 64 bits, and none of them do. The lab machine says so itself: the highest address in its own address space is printed by the kernel, and counting its hex digits is the measurement.

$ end=$(awk -F'[- ]' '/\[stack\]/ {print $2}' /proc/self/maps)
$ printf 'the top of this address space takes %d hex digits, so %d bits\n' ${#end} $(( ${#end} * 4 ))
the top of this address space takes 12 hex digits, so 48 bits
$ getconf PAGESIZE
4096

Forty eight bits, not sixty four, and a 4096 byte page. The arithmetic for this machine is therefore:

page number bits = 48 - 12 = 36

levels = 36 / 9 = 4

Four levels, which is the shape of the page tables on the machine this book was checked on. The 64 bit address is a 64 bit register, not 64 bits of memory anybody can name.

Distinctions that carry marks

Flat page tableTwo level page table
Size for a 32 bit space4 megabytes, alwaysthe outer page plus one page per 4 megabytes used
Must be contiguousyes, 1024 frames in a rowno, every piece is one page
Memory accesses per reference23
A process using 100 kilobytespays for 1,048,576 entriespays for two pages
Parts can be swapped outnot usefullyyes, an unused inner table need not exist at all
Outer tableInner table
How manyone per processone per 4 megabytes of address space in use
An entry namesan inner tablea frame
Indexed byp1p2

What it does not mean

Hierarchical paging does not reduce the number of pages. The process still has 1,048,576 pages. It reduces the number of entries that have to exist.

It is not the same as segmentation. The fields are cut out of one address by the hardware and the program supplies nothing, which is the test from Chapter seventy four.

It does not make translation faster. It makes it slower, by one memory access per level, and the TLB is what makes that acceptable.

An inner table is not optional for a page that is in use. If a page is valid, the inner table covering it must exist. What is saved is the tables for the parts of the address space nothing lives in, which in a large sparse space is nearly all of it.

munotes.in311

The Page Table Is Too Big: Hierarchical Paging

Quick revision

  • A flat table for a 32 bit space with 4 kilobyte pages and 4 byte entries is 4 megabytes,

and it must be 1,024 contiguous frames. Both halves are the problem.

  • Page the page table. Split the page number into an outer and an inner field, choosing the

inner field so one inner table is exactly one page: 4096 / 4 = 1024 entries, so 10 bits.

  • The 32 bit split is 10, 10, 12. The outer table is one page and covers the whole space; one

inner table is one page and covers 4 megabytes.

  • A process using 8 megabytes needs three pages of tables, 12,288 bytes, none of them

adjacent.

  • Translating takes three memory accesses: outer, inner, data. Each level adds one.
  • The cost with a TLB at 98 per cent is 124 nanoseconds against 122 for a flat table: two

nanoseconds for four megabytes.

  • On a 64 bit space with 8 byte entries, an inner table holds 512 entries, so the fields

are 9 bits and two levels leave an outer table of 64 terabytes. Six levels are needed, and seven accesses.

  • Real 64 bit machines implement fewer bits: the lab machine's own addresses take 48 bits,

which gives four levels.

Test yourself

  1. Give the two reasons a four megabyte page table is unacceptable. It is four megabytes per

process; and because it is indexed as an array it must be 1,024 contiguous frames, which is exactly the large hole paging was meant to stop needing. 2. How is the inner field width chosen, and what is it for a 32 bit machine with 4 kilobyte pages and 4 byte entries? So that one inner table fills one page: 4096 / 4 = 1024 entries, which needs 10 bits.

  1. Give the three fields of a 32 bit two level address and their widths. p1 of 10 bits, p2 of

10 bits, and an offset of 12 bits.

  1. How much address space does one inner table cover? 1024 entries of 4096 bytes each,

4,194,304 bytes, four megabytes.

  1. A process uses 8 megabytes of address space. How much memory do its page tables take? The

outer table and two inner tables, 3 × 4096 = 12,288 bytes.

  1. Translate 4,206,596 on that machine. The offset is 4 and the page number 1027, so p1 is 1

and p2 is 3: read entry 1 of the outer table, entry 3 of the inner table it names, and join the frame to the offset. 7. How many memory accesses does a reference take on a two level machine, and what does the second level cost in time at a 98 per cent TLB hit ratio? Three; and 0.98 × 120 + 0.02 × 320 = 124 nanoseconds, two more than the flat table's 122.

munotes.in312

The Page Table Is Too Big: Hierarchical Paging

  1. Why is two level paging useless on a 64 bit address space? With 4 kilobyte pages and 8

byte entries the inner field is 9 bits, leaving 43 bits of outer field: an outer table of 64 terabytes. Six levels are needed instead.

Contents This chapter on its own page

munotes.in313

Chapter Seventy-Nine

Hashed and Inverted Page Tables

Syllabus topic Module 2, "Memory Management - Structure of the Page Table"

In one line

Both structures make the page table proportional to the memory a machine has rather than to the addresses a process could name, and both pay for it with a search.

The one idea behind both

A forward mapped table, flat or hierarchical, is indexed by the page number, so its size is fixed by the size of the address space. Chapter seventy eight showed where that ends on a 64 bit machine.

Turn it round. Store an entry for each page that is actually in memory, and find it by searching rather than by indexing. Then the table is as big as the memory, not as big as the address space, and a 64 bit address space costs nothing extra. The cost moves from space to time, and both structures in this chapter are ways of making that search cheap.

Hashed page tables

The virtual page number is hashed, and the hash value indexes a table of chains.

PartWhat it is
the hash tablean array of slots, usually as many slots as the machine has frames
one chain elementthree fields: the virtual page number, the frame number, and a pointer to the next element in the chain
a chainevery page whose number hashed to that slot

The steps for one reference:

StepWhat happens
1hash the virtual page number to get a slot
2walk the chain at that slot, comparing the virtual page number in each element
3on a match, take the frame number from that element

The comparison in step 2 is the whole point, and a student who leaves it out has described something that gives wrong answers. Two different page numbers can hash to the same slot, so finding a chain is not finding your page: the page number stored in the element must be checked against the one you asked for.

The chains are short if the table is big enough, so a lookup is typically the hash, one memory access for the slot, and one or two for the chain. Any structure here is a trade of memory for chain length.

Clustered page tables

The variant MU's textbook names, and the reason is worth one line in an answer. Each entry maps several consecutive pages, not one. A 64 bit address space is used sparsely but in runs: a program's code is a run of pages, its heap is a run, its stack is a run. One entry covering sixteen pages makes the table sixteen times smaller for a run and no worse for scattered pages, so it suits sparse address spaces, which is what a 64 bit space always is.

munotes.in314

Hashed and Inverted Page Tables

Inverted page tables

One entry per frame of physical memory, for the whole machine, not one table per process.

Forward mapped tableInverted page table
Entry number i meanspage i of this processframe i of the machine
The entry holdsthe frame numberthe process identifier and the page number in that process
How many tablesone per processone, for the whole machine
Size is set bythe size of the address spacethe amount of physical memory

Read the second row again, because it is the definition. The entry for frame 9 says which process and which of its pages is sitting in frame 9. The table records what is in memory, which is why it is called inverted: it is the frame table of Chapter seventy four, doing the translation as well.

The sum that makes the case

1 gigabyte of memory, 4 kilobyte frames, and an entry of 8 bytes holding a process identifier and a page number.

frames = 1,073,741,824 / 4096 = 262,144

table = 262,144 × 8 = 2,097,152

2,097,152 bytes, two megabytes, for the entire machine. Set that beside a hundred processes with a flat table each:

flat tables = 100 × 4,194,304 = 419,430,400

Four hundred megabytes of page tables against two. That is the whole argument for an inverted table, and it is worth writing out with the arithmetic in a question rather than asserting.

What it costs

The table cannot be indexed by the page number, so a reference has to search it. The entry for your page may be anywhere in 262,144 entries.

The fixWhat it costs
a linear searchunusable: a quarter of a million comparisons per memory reference
a hash table on the page number, whose entries point into the inverted tableone extra memory access, and every real inverted table has one

So a reference with a TLB miss costs: the hash table, the inverted table, then the data. Three memory accesses, which is what two level paging cost in Chapter seventy eight: the structure saved the space and not the time.

The two things it cannot do

Both are asked as short questions, and both follow from one entry per frame.

One: shared memory is hard. An inverted table has one page number per frame. If two processes share a frame, the entry can name only one of them, so a reference from the other process finds nothing. Systems that use inverted tables restrict sharing to one mapping at a time: the entry names the process that is using the frame now, and a reference from the other process is a fault which the operating system resolves by rewriting the entry.

munotes.in315

Hashed and Inverted Page Tables

Two: it says nothing about pages that are not in memory. There is no entry for a page on disk, because there is no frame. So the operating system still keeps a per process table on the side, not consulted by the hardware, to find where a page is stored when it faults. That table can itself live in swap space, since it is only needed when a fault has already happened.

All three compared

The table a question on "Structure of the Page Table" is asking for.

Flat or hierarchicalHashedInverted
Found byindexing by page numberhashing the page numbersearching, in practice by a hash
Size set bythe address spacethe memory in usethe physical memory
Tablesone per processone per processone per machine
32 bit cost4 megabytes flat, or a page per 4 megabytes usedproportional to the pages in use2 megabytes for a gigabyte of memory
Accesses on a TLB miss2 flat, 1 per level plus the datathe slot and the chainthe hash table and the table, then the data
Sharingeasy, the same frame number in two tableseasy, two elements naming one framehard, one page number per frame
Pages not in memoryrecorded in the same entryrecorded in the chainnot recorded, a separate table is needed
Suitsany 32 bit machinea sparse 64 bit spacea machine with much more address space than memory

What it does not mean

Hashing does not remove the page table. It replaces the index with a hash and a comparison; the pairs still have to be stored.

An inverted page table is not one per process. It is one per machine, and the process identifier in each entry is what keeps processes apart.

Neither structure makes a reference faster. Both cost at least as much as two level paging on a miss. What they save is memory, and the TLB is still what makes any of it fast enough.

A hash collision is not an error. It is expected, which is why a chain and a comparison exist. Only a wrong comparison is an error.

Quick revision

  • Both structures are sized by the memory in use, not by the address space, and both pay with

a search.

  • A hashed table hashes the page number to a slot and walks a chain of

page number, frame number, next, comparing the page number in every element.

  • Clustered page tables put several consecutive pages in one entry, which suits

sparse 64 bit spaces.

  • An inverted table has one entry per frame for the whole machine, and each entry
munotes.in316

Hashed and Inverted Page Tables

holds the process identifier and page number in it.

  • 1 gigabyte of memory with 4 kilobyte frames gives 262,144 entries; at 8 bytes that is

2 megabytes for the machine, against 400 megabytes for a hundred flat tables.

  • An inverted table needs a hash table in front of it, so a TLB miss costs three memory

accesses.

  • Its two weaknesses: sharing, because one entry names one page; and

no record of pages not in memory, so a separate per process table is still needed.

Test yourself

  1. What do a hashed and an inverted page table have in common? Both are sized by the memory

in use rather than by the address space, and both must search for a translation instead of indexing. 2. What are the three fields of a chain element in a hashed page table, and why is the first one needed? The virtual page number, the frame number and a pointer to the next element. The page number is needed because several page numbers hash to one slot and the right element must be identified by comparison.

  1. What is a clustered page table for? Each entry maps several consecutive pages, which makes

the table much smaller for the runs of pages a sparse 64 bit address space is used in.

  1. What does entry i of an inverted page table describe, and what does it hold? Frame i of

physical memory; it holds the process identifier and the page number of whatever is in that frame. 5. A machine with 1 gigabyte of memory, 4 kilobyte frames and 8 byte entries. How big is its inverted page table? 262,144 frames times 8 bytes, which is 2,097,152 bytes: two megabytes for the whole machine.

  1. Why does an inverted page table need a hash table, and what does a reference then cost?

Because it cannot be indexed by page number and a linear search of a quarter of a million entries is unusable; with the hash table a TLB miss costs three memory accesses.

  1. Why is sharing hard with an inverted page table? There is one entry per frame and it can

name only one process and page, so two processes cannot both be recorded as using the frame.

  1. Where is a page that is not in memory recorded under an inverted scheme? In a separate per

process table kept by the operating system, which the hardware never reads and which may itself be in swap space.

Contents This chapter on its own page

munotes.in317

Chapter Eighty

Virtual Memory: Running What Will Not Fit

Syllabus topic Module 2, "Virtual Memory Management - Background"

In one line

Virtual memory separates the addresses a program may use from the memory the machine has, so a program can be larger than memory and only the parts in use need a frame.

The observation it is built on

A program is never all needed at once. Look at any real program and the same four things are true.

What is in the programHow often it runs
error handling for cases that almost never happenalmost never
options and features a given user does not usenever, for that user
a table sized for the worst casepartly, and the rest is untouched
a large array, used in one cornerpartly

If every byte had to be in memory before the program could start, the machine would hold the whole of every program in order to run the small part of each that is actually working. Virtual memory is what happens when you stop doing that.

The definition

Virtual memory is the separation of a process's logical address space from physical memory. The program uses virtual addresses from 0 upwards; the operating system keeps some of its pages in frames and the rest on disk; and the program cannot tell the difference.

Logical, or virtual, address spacePhysical memory
How bigas big as the architecture allowsas big as the machine has fitted
Who sees itthe programthe operating system and the hardware
Set bythe address widththe price of memory
In Chapter seventy five32 bits, 4 gigabytes1 gigabyte, 262,144 frames

The gap in that last row is the whole subject. The process may name four gigabytes; the machine has one; and the operating system makes the difference up with a disk.

What it buys

Five benefits, and a question asking "what is virtual memory for" wants at least three of them.

BenefitBecause
A program can be larger than physical memoryit never has to be in memory all at once
More processes fit at once, so the processor is better usedeach one needs only its working part, so the degree of multiprogramming rises and Chapter twenty of Module 1's argument applies
Less input and output to start or swap a processonly the pages actually used are ever read from disk
Sharing and copy on write become cheapa frame can be named by two page tables, and a copy can be put off until somebody writes
A sparse address space costs nothinga hole in the middle of the address space has no frames behind it, so it takes no memory at all

The second one is the reason a system administrator cares. The processor is better used because more processes are resident, not because any one program got faster. Virtual memory makes a single program slightly slower and the machine as a whole far more useful, and an answer that says otherwise has the trade backwards.

munotes.in318

Virtual Memory: Running What Will Not Fit

The shape of a virtual address space

The standard picture, and the hole in the middle is the part worth understanding.

RegionWhere it isWhich way it grows
the program's codeat the bottom, at a low addressit does not grow
initialised and uninitialised dataabove the codeit does not grow
the heap, which malloc hands outabove the dataupwards
the holebetween the heap and the stackit is what both grow into
shared librariesmapped into the holethey do not grow
the stackat the top, at a high addressdownwards

The hole is the point. It is enormous, nothing is stored in it, and it costs no memory because no page in it is valid. That is why the heap and the stack can both grow for as long as the program needs without either of them having to guess how much the other will want.

Measured on the lab machine

The kernel prints the map of a process's own address space, so the hole can be measured rather than described.

$ heap=$(awk '/\[heap\]/ {print $1}' /proc/self/maps | cut -d- -f2)
$ stack=$(awk '/\[stack\]/ {print $1}' /proc/self/maps | cut -d- -f1)
$ echo "the heap ends at 0x$heap and the stack starts at 0x$stack" | sed 's/0x[0-9a-f]*/0xNNNN/g'
the heap ends at 0xNNNN and the stack starts at 0xNNNN
$ echo "the two addresses take ${#heap} and ${#stack} hex digits"
the two addresses take 8 and 12 hex digits
$ echo "the hole between them is $(( (0x$stack - 0x$heap) / 1024 / 1024 / 1024 / 1024 )) terabytes"
the hole between them is 255 terabytes

Two hundred and fifty five terabytes of address space, in a container with a couple of gigabytes of memory. The addresses change on every run, because the kernel places the regions differently each time for safety, but the size of the hole does not. Nothing is in it, and nothing needs to be.

An address space is not memory

The claim of the chapter, checked on the machine. This program asks for one gigabyte of address space and then reports two numbers the kernel keeps about it: the virtual size, which is address space, and the resident set size, which is memory.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <string.h>
#include <unistd.h>
#include <sys/mman.h>

/* One field of /proc/self/status, in kilobytes. */
static long kb(const char *name)
{
    FILE *f = fopen("/proc/self/status", "r");
    char line[256];
    long value = -1;
    while (fgets(line, sizeof line, f))
        if (strncmp(line, name, strlen(name)) == 0)
        sscanf(line + strlen(name), " %ld", &value);
    fclose(f);
    return value;
}

int main(void)
{
    size_t want = 1024UL * 1024 * 1024;          /* one gigabyte */
    char *p = mmap(NULL, want, PROT_READ | PROT_WRITE,
        MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
    if (p == MAP_FAILED) {
        perror("mmap");
        return 1;
    }
    long size = kb("VmSize:");
    long resident = kb("VmRSS:");
    kb("VmRSS:");                    /* again: the reading itself costs pages the first time */
    long before = kb("VmRSS:");
    p[0] = 1;
    p[want - 1] = 1;
    long after = kb("VmRSS:");
    printf("address space asked for: %zu megabytes\n", want / (1024 * 1024));
    printf("virtual size of the whole process: %ld megabytes\n", size / 1024);
    printf("of that, in memory: %ld kilobytes\n", resident);
    printf("in memory as a whole percentage of the virtual size: %ld\n", 100 * resident / size);
    printf("one byte written at each end cost %ld kilobytes, and a page is %ld\n",
        after - before, sysconf(_SC_PAGESIZE) / 1024);
    return 0;
}
munotes.in319

Virtual Memory: Running What Will Not Fit

$ gcc -std=c17 -Wall -Wextra -o bigmap bigmap.c
$ ./bigmap
address space asked for: 1024 megabytes
virtual size of the whole process: 1026 megabytes
of that, in memory: 1324 kilobytes
in memory as a whole percentage of the virtual size: 0
one byte written at each end cost 8 kilobytes, and a page is 4

Read the numbers one at a time.

The machine saysWhat it means
virtual size 1026 megabytesthe process may name a gigabyte and a little more
in memory 1324 kilobytesabout one megabyte of it exists
0 per cent, rounded downthe address space is more than a thousand times the memory behind it
two bytes written cost 8 kilobytes, and a page is 4exactly two pages were brought in, one for each byte

The last line is demand paging caught in the act, and it is the whole of the next chapter: the memory appeared when it was written to and not before. Asking for a gigabyte cost nothing at all.

How it is implemented

Two ways, and MU's textbook names both.

SchemeUnit brought in on demandUsed by
Demand paginga pageevery general purpose system, and the rest of this module
Demand segmentationa segmentsystems whose hardware has no paging; harder, because segments differ in size, which is Chapter seventy three's problem back again

Distinctions that carry marks

Swapping, Chapter sixty nineVirtual memory
Unit movedthe whole processone page
Whenwhen the process is not runningwhile it runs, on demand
A process larger than memorycannot run at allruns
Who decidesthe medium term schedulerthe page fault, as it happens
Virtual sizeResident set size
What it countsaddress space, whether or not anything is behind itthe pages actually in frames
Costs memorynoyes
In the measurement above1026 megabytesabout 1 megabyte
munotes.in320

Virtual Memory: Running What Will Not Fit

What it does not mean

Virtual memory is not the disk. The disk holds the pages that are not in memory. Virtual memory is the scheme that lets a program use addresses whose pages may be in either place.

It is not the same as swap space. Swap space is where pages go; virtual memory is why they can.

It is not infinite. A process's address space is limited by the address width, and the total of all the pages everybody has must fit in memory plus swap.

It does not make a program faster. It makes each reference slightly dearer and lets far more work be resident at once. Chapter eighty two prices the dearness exactly.

A hole is not memory. Two hundred and fifty five terabytes of hole cost nothing, because no page in it is valid.

Quick revision

  • A program is never all needed at once: error paths, unused options, worst case tables,

partly used arrays.

  • Virtual memory is the separation of the logical address space from physical memory. The

program uses virtual addresses; the operating system decides which pages have frames.

  • Five benefits: a program larger than memory; more processes resident, so better

processor use; less input and output to start or swap; cheap sharing and copy on write; and sparse address spaces cost nothing.

  • The address space is code, data, heap growing up, a hole, shared libraries, and the

stack growing down. The hole costs nothing.

  • Measured: the hole between heap and stack on the lab machine is 255 terabytes.
  • Measured: a gigabyte of address space, 1026 megabytes virtual and about 1 megabyte

resident, 0 per cent; and two bytes written brought in 8 kilobytes, exactly two pages.

  • Implemented by demand paging, or rarely by demand segmentation.
  • It makes one program a little slower and the machine much more useful.

Test yourself

  1. Why is it possible to run a program without all of it in memory? Because parts of a

program are rarely or never used: error handling, unused options, tables sized for the worst case, and the untouched parts of large arrays.

  1. Define virtual memory in one sentence. It is the separation of a process's logical address

space from physical memory, so that a process can run with only some of its pages in frames.

  1. Give four benefits of virtual memory. A program can be larger than memory; more processes

can be resident, so the processor is better used; less input and output is needed to load or swap a process; and sharing, copy on write and sparse address spaces all become cheap.

munotes.in321

Virtual Memory: Running What Will Not Fit

  1. Why does a hole in an address space cost no memory? Because no page in it is valid, so no

frame is allocated for any of it, and nothing about it is stored except the absence in the page table.

  1. What is the difference between the virtual size of a process and its resident set size?

The virtual size is how much address space it may name; the resident set size is how much of that is in frames. Only the second costs memory. 6. A process asks for a gigabyte and writes one byte at each end. How much memory does it use? Two pages, eight kilobytes on this machine, and the rest of the gigabyte has no frames at all.

  1. Distinguish swapping from virtual memory. Swapping moves a whole process out and in while

it is not running, and a process bigger than memory can never run. Virtual memory moves single pages while the process runs, so a process bigger than memory runs normally.

  1. Does virtual memory make programs faster? No. Each reference costs slightly more; what

improves is how much work the machine can keep resident and therefore how well the processor is used.

Contents This chapter on its own page

munotes.in322

Chapter Eighty-One

Demand Paging, and the Page Fault

Syllabus topic Module 2, "Virtual Memory Management - Demand Paging"

In one line

A page is brought into memory the first time the program touches it, and the trap that happens when it is touched is called a page fault.

The bit that does the work

The valid and invalid bit of Chapter seventy seven is used for a second purpose, and nothing else has to be added.

The bit saysThe operating system's own tables sayWhat it means
validthe page is in frame fread or write it; the hardware does not trap
invalidthe page is not part of this processan illegal reference: kill the process
invalidthe page is part of this process, and is on diska page fault: bring it in and carry on

One bit, two meanings, and only the operating system can tell them apart. The hardware traps on invalid and knows nothing more; the kernel then looks at its own record of the process's address space to decide whether this is a fault to service or a program to kill. A question asking how the system distinguishes a legal fault from an illegal reference is asking for exactly this.

Servicing a fault

The sequence, which is worth learning in this order because a question asks for the steps.

StepWhat happens
1the hardware finds the invalid bit and traps to the operating system
2the kernel saves the registers and the process state
3it establishes that the trap was a page fault, and for which address
4it checks the reference was legal, and finds where the page is on disk
5it finds a free frame, and issues a read of the page into it
6while the disk works, the processor is given to another process: the faulting one waits
7the disk finishes and interrupts
8the kernel corrects the page table: the page is now in frame f, and the bit becomes valid
9the process waits for its turn, and then the faulting instruction is restarted

Step 6 is the one students leave out, and it is the reason demand paging works at all on a busy machine: a page fault is not idle time, it is a chance to run somebody else. Step 9 is the one examiners like, because restarting the instruction is where the difficulty is.

Pure demand paging

Start the process with no pages in memory at all. The first instruction faults, the page holding it comes in, and the program goes on to fault its way up to the set of pages it actually needs.

It sounds dreadful and it is not, for the reason Chapter seventy six gave: locality of reference. A program that has faulted in the six pages it is working in will run for a long time in those six pages. The fault rate is high for a moment at the start and then small, which is what Chapter eighty nine's working set describes.

munotes.in323

Demand Paging, and the Page Fault

Some systems help a little by reading a few nearby pages at the same time, since the disk arm is already there. That is prepaging, and it is a gamble: pages read and not used are work wasted.

The hardware it needs

RequirementWhy
a page table with a valid and invalid bitto trap the reference at all
secondary storage, the swap spacesomewhere to keep the pages that are not in memory
the ability to restart any instruction exactlybecause a fault can happen in the middle of one

The instruction that makes it hard

The third requirement is the hard one, and it is a favourite question. Consider an instruction that moves a block of memory and faults part way through, after it has already overwritten some bytes. Restarting it from the beginning would move the block again, and the source has already been partly destroyed.

The way outHow it works
touch both ends firstbefore doing anything, read the first and last byte of the source and of the destination, so every page is brought in before a single byte is written
undo the changeskeep the old values in registers and put them back before restarting

An instruction that increments a register as a side effect, as some addressing modes do, is the same problem in miniature: the register must be put back before the restart, or the instruction runs with the wrong address.

Counted on the lab machine

The kernel counts page faults per process, and separates the two kinds.

KindWhat it costsName
a minor faultno disk: the frame is found, or already in memory, and the table is correctedreclaiming a frame
a major faulta disk reada real page-in

This program asks for 64 megabytes, writes one byte in every page, and prints the faults the kernel charged it.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <unistd.h>
#include <sys/mman.h>
#include <sys/resource.h>

static long minor_faults(void)
{
    struct rusage r;
    getrusage(RUSAGE_SELF, &r);
    return r.ru_minflt;
}

static long major_faults(void)
{
    struct rusage r;
    getrusage(RUSAGE_SELF, &r);
    return r.ru_majflt;
}

int main(void)
{
    long page = sysconf(_SC_PAGESIZE);
    size_t want = 64UL * 1024 * 1024;
    size_t pages = want / (size_t) page;
    long before_map = minor_faults();
    char *p = mmap(NULL, want, PROT_READ | PROT_WRITE,
        MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
    if (p == MAP_FAILED) {
        perror("mmap");
        return 1;
    }
    long after_map = minor_faults();
    printf("asked the kernel for %zu pages of memory\n", pages);
    printf("page faults the asking cost: %ld\n", after_map - before_map);
    for (size_t i = 0; i < pages; i++)
        p[i * (size_t) page] = 1;
    long after_touch = minor_faults();
    printf("page faults the first write to every page cost: %ld\n", after_touch - after_map);
    printf("faults per page touched: %ld\n", (after_touch - after_map) / (long) pages);
    printf("major faults, the ones that go to disk: %ld\n", major_faults());
    return 0;
}
munotes.in324

Demand Paging, and the Page Fault

$ gcc -std=c17 -Wall -Wextra -o faults faults.c
$ ./faults
asked the kernel for 16384 pages of memory
page faults the asking cost: 0
page faults the first write to every page cost: 16392
faults per page touched: 1
major faults, the ones that go to disk: 1
$ grep -E '^SwapTotal' /proc/meminfo
SwapTotal:             0 kB

Four things are proved by those five lines.

The machine saysWhat it proves
asking for 16,384 pages cost 0 faultsnothing is brought in until it is used. The memory was promised, not given
writing to every page cost 16,392 faultsone fault per page, and the eight extra are the program's own pages
one fault per page touchedthe unit of demand paging is exactly one page
1 major fault, and SwapTotal 0the only disk read in the whole run was the program's own text; with no swap space configured, a page of this program's data can never be sent to disk, so the faults above are all minor

Be careful with the last row in an answer. This machine cannot demonstrate a major fault on data, because it has no swap. What it demonstrates exactly is the first half of demand paging: memory arrives one page at a time, when it is written to, and not before.

Distinctions that carry marks

SwappingDemand paging
Movesa whole processone page
Triggered bythe schedulera reference by the program
Process state while it happensnot runningwaiting, and another process runs
Minor faultMajor fault
Disk readnoyes
Costmicrosecondsmilliseconds
Causea first touch, or a page already in memory that this process had not mappedthe page is on disk
Counted by the kernel asminor, or reclaiming a framemajor
Page faultIllegal reference
The bit the hardware sawinvalidinvalid
What the kernel's records saythe page is legal and on diskthe address is not in the address space
What happensthe page is read in and the instruction restartedthe process is killed

What it does not mean

A page fault is not an error. It is the normal way memory arrives. The error is the other meaning of the same bit.

Demand paging does not need a special bit. It reuses the valid and invalid bit that protection already needed.

munotes.in325

Demand Paging, and the Page Fault

A fault does not stop the machine. It stops the faulting process, and the processor goes to another one.

Restarting the instruction is not always simple. An instruction that has already written to memory must be able to undo or to prefetch, which is a hardware design requirement, not a software choice.

Pure demand paging is not slow in the long run. The fault rate is high while the working set comes in and small afterwards, by locality.

Quick revision

  • Demand paging brings a page in on the first reference to it, using the

valid and invalid bit.

  • Invalid has two meanings: on disk, which is a fault to service; or not in the address

space, which kills the process. Only the kernel's own records tell them apart.

  • Servicing a fault:

trap, save state, identify the fault, check it is legal, find a free frame, read the page while another process runs, take the disk interrupt, correct the page table, restart the instruction.

  • Pure demand paging starts a process with no pages at all; locality makes the fault rate

fall quickly. Prepaging reads neighbours in advance and risks wasted work.

  • Hardware needed: the valid bit, swap space, and the ability to

restart any instruction, which is hard for a block move that has already written: touch both ends first, or undo.

  • Measured: asking for 16,384 pages cost 0 faults; writing to each cost 16,392, which is

one per page.

  • A minor fault costs no disk read; a major fault does. The lab machine has no swap,

so it cannot show a major fault on data.

Test yourself

  1. Which bit makes demand paging possible, and what are its two meanings? The valid and

invalid bit. Invalid means either that the page is legal but on disk, which is serviced, or that the address is outside the process's address space, which kills it.

  1. How does the operating system tell a page fault from an illegal reference? By its own

record of the process's address space; the hardware reports only that the bit was invalid.

  1. List the steps of servicing a page fault. Trap to the kernel, save the process state,

identify the fault and the address, check the reference is legal and find the page on disk, find a free frame and start the read, run another process while the disk works, take the disk interrupt, correct the page table and set the bit valid, then restart the faulting instruction.

  1. What is pure demand paging, and why is it not ruinous? Starting a process with no pages in
munotes.in326

Demand Paging, and the Page Fault

memory, so that even the first instruction faults. Locality of reference means the program soon has the few pages it is working in and then faults rarely.

  1. What three things must the hardware provide? A page table with a valid and invalid bit,

secondary storage for the pages that are out, and the ability to restart any instruction exactly.

  1. Why is a block move instruction a problem, and what are the two fixes? It may fault after

it has already overwritten part of the destination, so restarting it would work on damaged data. Either touch both ends of both blocks first so every page is present, or record the old values and undo the change before restarting. 7. A program asks for 64 megabytes and then writes one byte in each of its 16,384 pages. How many faults, and of what kind? About 16,384, one per page, and all minor: the asking itself cost none.

  1. Distinguish a minor fault from a major fault. A minor fault is resolved without a disk

read, in microseconds; a major fault reads the page from disk and costs milliseconds.

Contents This chapter on its own page

munotes.in327

Chapter Eighty-Two

The Effective Access Time Under Demand Paging

Syllabus topic Module 2, "Virtual Memory Management - Performance of Demand Paging"

In one line

A page fault costs about eighty thousand times what a memory access costs, so the fault rate, not the fault, is what decides whether a paged system is usable.

The two numbers the sum rests on

EventTimeIn nanoseconds
a memory access100 nanoseconds100
a page fault serviced from diskabout 8 milliseconds8,000,000

A fault costs 80,000 memory accesses.

8,000,000 / 100 = 80,000

That single ratio is the reason this chapter exists. Nothing else in a computer has a gap of that size between two things that happen in the same program, and every consequence in the next five chapters follows from it.

Where the 8 milliseconds goes

Part of servicing a faultTypical time
service the page fault interrupt: save state, decide, start the read1 to 100 microseconds
read the page in from diskabout 8 milliseconds
restart the process: take the interrupt, restore state, resume1 to 100 microseconds

The first and third parts are software and can be tuned. The middle part is a disk, and Chapter ninety two explains why it cannot be made much faster: it is a seek, a rotation, and a transfer, and two of those are mechanical.

The formula

With p as the probability that a reference faults:

EAT = (1 - p) × memory access + p × fault service time

p is a probability, not a percentage, and not a count. It is the fraction of all memory references that fault. A question that says "one fault every thousand references" means p = 0.001, and a question that says "a fault rate of 1 per cent" means p = 0.01, which as the table below shows is a machine nobody could use.

Worked, at four fault rates

With a 100 nanosecond memory and an 8 millisecond fault:

Fault rate pWorkingEATTimes slower
0.001, one in a thousand0.999 × 100 + 0.001 × 8,000,000 = 99.9 + 8000 = 8,099.98,099.9about 81
0.0001, one in ten thousand0.9999 × 100 + 0.0001 × 8,000,000 = 99.99 + 800 = 899.99899.99about 9
0.00001, one in a hundred thousand0.99999 × 100 + 0.00001 × 8,000,000 = 99.999 + 80 = 179.999179.999about 1.8
0.000001, one in a million0.999999 × 100 + 0.000001 × 8,000,000 = 99.9999 + 8 = 107.9999107.9999about 1.08

Read the first row again. One fault in a thousand references, which sounds rare, makes the program eighty one times slower. A program that should take one second takes a minute and twenty one seconds. That is the sentence to write in an answer.

And read the last row. One fault in a million costs 8 per cent. The whole art of the next five chapters is getting the fault rate down to something with six zeroes in it.

munotes.in328

The Effective Access Time Under Demand Paging

Solving it backwards

The other way the question comes: what fault rate can we afford? Allow the program to be 10 per cent slower than an unpaged machine, so an effective access time of 110 nanoseconds.

110 = (1 - p) × 100 + p × 8,000,000

110 = 100 + 7,999,900p

7,999,900p = 10

p = 10 / 7,999,900 = 1 / 799,990

Fewer than one fault in 799,990 memory references. That is the answer, and it is worth saying what it means: the program may fault about once in every eight hundred thousand references, which for a program touching memory every few nanoseconds is a handful of faults a second. Demand paging is only usable because programs have locality, and Chapter eighty nine is the measure of that.

When the page thrown out is dirty

The variant a question uses to see whether the formula was understood or memorised. If the frame chosen for the incoming page holds a page that has been modified, it has to be written out before the new one can be read in, so that fault costs two transfers.

Suppose 70 per cent of the pages chosen for replacement are dirty.

average fault = 8,000,000 + 0.7 × 8,000,000 = 8,000,000 + 5,600,000 = 13,600,000

EAT = 0.999 × 100 + 0.001 × 13,600,000 = 99.9 + 13,600 = 13,699.9

13,699.9 nanoseconds, about 137 times slower, against 81 times when nothing has to be written back.

That is what the modify bit of Chapter seventy seven is worth: it lets the operating system pick a clean page when it can, and the difference between 81 and 137 is the whole reason the bit exists. Chapter eighty four uses it again.

Distinctions that carry marks

The TLB sum, Chapter seventy sixThis sum
What is missinga translation, which is in memorythe page itself, which is on disk
The penaltyone extra memory access, 100 nanosecondsa disk access, 8 milliseconds
Ratio of penalty to a normal accessabout 2about 80,000
A rate of 2 per centcosts 22 per centcosts 160,000 per cent
What the fix isa small cache of translationskeeping the right pages in memory, which is the rest of this module
p1 - p
Namethe page fault ratethe hit rate
Multipliesthe fault service timethe memory access time
A good value0.000001 or smaller0.999999

What it does not mean

The 8 milliseconds is not the fault handler's code. Nearly all of it is the disk. The software part is microseconds.

munotes.in329

The Effective Access Time Under Demand Paging

A fault rate of 1 per cent is not nearly as good as 0.1 per cent. It is ten times worse, and both are unusable: 800 times slower against 81.

The effective access time is not what one fault costs. It is the average cost of a reference, which is what decides how long the program takes.

A smaller page size does not reduce the fault cost. It reduces the transfer a little and increases the number of faults, which Chapter seventy five compared.

Locality is not an assumption in this sum. The sum takes the fault rate as given. Locality is the reason the rate can be as small as 0.000001 in practice.

Quick revision

  • A memory access is 100 nanoseconds; a fault serviced from disk is about 8 milliseconds,

which is 8,000,000 / 100 = 80,000 memory accesses.

  • EAT = (1 - p) × memory access + p × fault service time.
  • p = 0.001 gives 8,099.9 nanoseconds, about 81 times slower; p = 0.0001 gives

899.99, about 9 times; p = 0.00001 gives 179.999; p = 0.000001 gives 107.9999, only 8 per cent slower.

  • For a 10 per cent degradation, 110 nanoseconds is 100 plus 7,999,900p, so

p = 1 / 799,990, fewer than one fault in eight hundred thousand references.

  • If a fraction of the replaced pages are dirty, the fault costs two transfers: at 70 per

cent dirty the average fault is 13,600,000 nanoseconds and the EAT at p = 0.001 is 13,699.9, about 137 times slower.

  • The modify bit is what lets the system avoid the write-back, and the gap between 81 and 137

times is its worth.

  • The fault service time splits into interrupt service and restart, both microseconds, and

the disk read, which is milliseconds.

Test yourself

  1. How many memory accesses does one page fault cost? 8,000,000 divided by 100, which is

80,000.

  1. Write the formula for the effective access time under demand paging. EAT = (1 - p) times

the memory access time, plus p times the fault service time.

  1. Find the EAT for p = 0.001 with a 100 nanosecond memory and an 8 millisecond fault.

0.999 × 100 + 0.001 × 8,000,000 = 99.9 + 8000 = 8,099.9 nanoseconds, about 81 times slower.

  1. The same machine at one fault in a million. 99.9999 + 8 = 107.9999 nanoseconds, about 8

per cent slower.

  1. What fault rate keeps the slowdown below 10 per cent? Set 110 against 100 plus 7,999,900p

to get p = 10 / 7,999,900 = 1 / 799,990: fewer than one fault in 799,990 references. 6. Why does a dirty replaced page make a fault dearer, and by how much if 70 per cent are dirty? It must be written out before the new page is read in, so the average fault is 8,000,000 + 0.7 × 8,000,000 = 13,600,000 nanoseconds, and at p = 0.001 the EAT becomes 13,699.9, about 137 times slower.

munotes.in330

The Effective Access Time Under Demand Paging

  1. Where does the 8 milliseconds actually go? Almost entirely into the disk read; servicing

the interrupt and restarting the process are microseconds each.

  1. Why is this sum so much harsher than the TLB sum? A TLB miss costs one extra memory

access, about twice a normal reference; a page fault costs a disk access, about eighty thousand times a normal reference.

Contents This chapter on its own page

munotes.in331

Chapter Eighty-Three

Copy on Write

Syllabus topic Module 2, "Virtual Memory Management - Copy-on-Write"

In one line

A child process shares its parent's pages read-only, and a page is copied only when one of them writes to it.

The problem it solves

Chapter five of Module 1 said that fork gives the child a copy of the parent's address space. Take that literally and the cost is dreadful.

The parent hasCopying it literally costs
64 megabytes of data16,384 pages copied, one at a time
1 gigabyte of data262,144 pages copied

And in the commonest case every byte of that work is thrown away, because the very next thing the child does is exec, which replaces the whole address space with a different program. A shell runs fork and then exec for every command you type.

The mechanism

On fork, do not copy anything. Let both processes point at the same frames, and mark every one of those pages read-only in both page tables.

ThenWhat happens
either process reads a pageit finds the frame through its own page table and reads it. No copy, no cost
either process writes to a pagethe read-only bit traps. The kernel sees the page is a copy-on-write page, makes a copy into a free frame, points the writer's page table at the copy, marks both writable, and restarts the instruction
the child calls execthe shared pages are dropped without ever having been copied

So the copying is done one page at a time, only for the pages that are actually written, and only when they are written. A child that reads everything and writes nothing costs no copying at all.

The trap is a protection fault, not a page fault in the sense of Chapter eighty one. The page is in memory the whole time. Nothing is read from disk, which is why the cost is microseconds rather than milliseconds, and why this chapter belongs beside demand paging rather than inside it.

Where the free frame comes from

The kernel keeps a free frame list, and for copy on write it wants a frame whose contents cannot leak: frames handed out for a copy are zero filled on demand, that is, wiped before they are given away. The same rule applies to a new page of stack or heap, and it is a protection requirement, not an optimisation: the previous owner's data must not be readable by the next.

Proved on the lab machine

The claim is that a read costs nothing and a write costs one copy. The kernel publishes, for any process, how much of its memory is shared with another process and how much is its own, so the claim can be watched happening.

munotes.in332

Copy on Write

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <string.h>
#include <unistd.h>
#include <sys/mman.h>
#include <sys/resource.h>
#include <sys/wait.h>

static long faults(void)
{
    struct rusage r;
    getrusage(RUSAGE_SELF, &r);
    return r.ru_minflt;
}

/* One field of /proc/self/smaps_rollup, in kilobytes: how much of this
process's memory is shared with another process, and how much is its own. */
    static long rollup(const char *name)
{
    FILE *f = fopen("/proc/self/smaps_rollup", "r");
    char line[256];
    long v = -1;
    while (fgets(line, sizeof line, f))
        if (strncmp(line, name, strlen(name)) == 0)
        sscanf(line + strlen(name), " %ld", &v);
    fclose(f);
    return v;
}

int main(void)
{
    long page = sysconf(_SC_PAGESIZE);
    size_t want = 64UL * 1024 * 1024;
    size_t pages = want / (size_t) page;
    char *p = mmap(NULL, want, PROT_READ | PROT_WRITE,
        MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
    if (p == MAP_FAILED) {
        perror("mmap");
        return 1;
    }
    for (size_t i = 0; i < pages; i++)      /* the parent fills it */
        p[i * (size_t) page] = 1;
    printf("the parent holds %zu pages, and %ld kilobytes of them are its own\n",
        pages, rollup("Private_Dirty:"));
    long parent_before = faults();
    fflush(stdout);                          /* or the child flushes our line too */

    pid_t child = fork();
    if (child < 0) {
        perror("fork");
        return 1;
    }
    if (child == 0) {
        long start = faults();
        long sum = 0;
        for (size_t i = 0; i < pages; i++)  /* read every page */
            sum += p[i * (size_t) page];
        long after_read = faults();
        long shared = rollup("Shared_Dirty:");
        long mine = rollup("Private_Dirty:");
        for (size_t i = 0; i < pages; i++)  /* now write to every page */
            p[i * (size_t) page] = 2;
        long after_write = faults();
        printf("the child read all %zu pages: %ld faults\n", pages, after_read - start);
        printf("  after reading, shared with the parent: %ld kilobytes, its own: %ld\n",
            shared, mine);
        printf("the child wrote to all %zu pages: %ld faults\n", pages, after_write - after_read);
        printf("  after writing, shared with the parent: %ld kilobytes, its own: %ld\n",
            rollup("Shared_Dirty:"), rollup("Private_Dirty:"));
        printf("  the sum it read back was %ld\n", sum);
        fflush(stdout);
        _exit(0);
    }
    int status = 0;
    waitpid(child, &status, 0);
    printf("the parent's own faults while all that happened: %ld\n", faults() - parent_before);
    printf("the parent's first byte is still %d\n", p[0]);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o cow cow.c
$ ./cow
the parent holds 16384 pages, and 65616 kilobytes of them are its own
the child read all 16384 pages: 16384 faults
  after reading, shared with the parent: 65584 kilobytes, its own: 40
the child wrote to all 16384 pages: 16405 faults
  after writing, shared with the parent: 44 kilobytes, its own: 65576
  the sum it read back was 16384
the parent's own faults while all that happened: 4
the parent's first byte is still 1
munotes.in333

Copy on Write

Read it line by line, because every line is one claim of the chapter.

The machine saysWhat it proves
the parent's 16,384 pages are 65,616 kilobytes of its ownbefore the fork, all of it belongs to the parent alone
after the child read everything: 65,584 kilobytes shared, 40 its ownthe fork copied nothing. Sixty four megabytes are in one set of frames that both processes are reading
after the child wrote to everything: 44 kilobytes shared, 65,576 its ownnow the copy has been made, and it was made because of the writing
the parent's own faults: 4the parent did no work at all while the child copied sixty four megabytes
the parent's first byte is still 1the copy is private: the child wrote 2 everywhere and the parent's memory did not change

One honest detail about this kernel. The child's reads cost 16,384 faults too, one per page, and that is not a contradiction: this kernel does not copy the page tables at fork either, so the child's first touch of a page, read or write, costs a minor fault to fill in its own table entry. The shared figure of 65,584 kilobytes is the proof that those faults mapped the existing frames rather than copying them. The writes then cost a second fault each, and that one is the copy.

What it is used for besides fork

UseHow copy on write helps
fork, then execnothing is copied before the address space is thrown away
fork, then a child that only readsnothing is copied at all
two processes mapping the same file privatelyeach sees its own changes, and unchanged pages stay shared
a snapshot of memory or of a file systemthe snapshot shares the pages and copies only what changes afterwards

vfork, an older call, went further: the child shares the parent's address space and the parent is suspended, so nothing is copied even on a write. It is safe only if the child does nothing but exec, and copy on write has made it unnecessary.

Distinctions that carry marks

Sharing, Chapter seventy sevenCopy on write
The page is markedread-only because it is read-only, being reentrant coderead-only only to catch the write
A write to itis an error, and the process is killedcauses a copy and then succeeds
Both processes see each other's changesnot applicable, nobody writesno, each has its own copy afterwards
Purposeone copy of code for many usersa cheap fork
Page fault, Chapter eighty oneCopy-on-write fault
What was missingthe page, which was on disknothing: the page is in memory
The bit that trappedvalid and invalidread and write
Costa disk access, millisecondsa frame copy, microseconds
After itthe page table names a frame that was emptythe page table names a new copy
munotes.in334

Copy on Write

What it does not mean

Copy on write does not avoid the copy. It defers it, and avoids it only for pages nobody writes to. A child that writes to everything copies everything, as the measurement above shows.

It is not sharing. After the write the two processes have separate pages, and neither can see the other's changes.

The read-only marking is not protection. The pages are logically writable; the bit is there so that the hardware will call the kernel at the first write.

It is not only for fork. Any time two page tables may name one frame until somebody writes, the same trick works.

The pages are not copied lazily to disk. Nothing here touches a disk.

Quick revision

  • On fork, the pages are not copied. Both page tables name the same frames, marked

read-only in both.

  • A read costs nothing. A write traps, the kernel copies that one page into a free

frame, marks it writable for the writer, and restarts the instruction.

  • The commonest case is fork then exec, where a literal copy would be thrown away entirely.
  • Frames for a copy are zero filled on demand, so the previous owner's data cannot leak.
  • Measured: after the child read all 16,384 pages, 65,584 kilobytes were shared and

40 its own; after it wrote to them all, 65,576 kilobytes were its own. The parent took 4 faults and its data was unchanged.

  • The trap is on the read and write bit, the page is in memory, and the cost is a

frame copy in microseconds, not a disk access.

  • vfork shared the address space outright and is made unnecessary by copy on write.

Test yourself

  1. What does fork copy, under copy on write? The page tables, not the pages: both processes

name the same frames, which are marked read-only in both.

  1. What happens when one of them writes? The read-only marking traps, the kernel copies that

single page into a free frame, gives the writer the copy and marks it writable, and restarts the instruction.

  1. Why is fork followed by exec the case that matters most? Because a literal copy of the

whole address space would be discarded immediately by the exec, and copy on write means none of it is ever made.

  1. Why must a frame given out for a copy be zero filled? So that whatever the previous owner

left in it cannot be read by the new one.

munotes.in335

Copy on Write

  1. Is a copy-on-write fault a page fault? Not in the disk sense. The page is already in

memory; the trap comes from the read and write bit and is resolved by copying a frame, in microseconds. 6. The measurement shows 65,584 kilobytes shared after the child read everything, and 65,576 kilobytes of its own after it wrote. What does that pair of numbers prove? That the fork copied nothing and the reads only mapped the existing frames, and that the copying was caused by the writes.

  1. Does copy on write ever cost more than a plain copy? Marginally, for a child that writes

to every page: it pays one trap per page as well as the copy. The gain is in the common case where most pages are never written.

  1. How is copy on write different from the sharing of Chapter seventy seven? Shared reentrant

code is genuinely read-only and stays one copy; a copy-on-write page is logically writable and becomes two copies at the first write.

Contents This chapter on its own page

munotes.in336

Chapter Eighty-Four

Page Replacement: The Problem

Syllabus topic Module 2, "Virtual Memory Management - Page Replacement"

In one line

When a page must come in and no frame is free, the operating system chooses a page already in memory, writes it out if it has been changed, and takes its frame.

How the machine gets into this state

Demand paging lets the operating system promise more memory than it has, and the promise is usually safe because no process uses all of its pages at once. It is called over allocation, and it is deliberate.

The machine hasThe processes have been promisedWhy it works
262,144 framesforty processes of a million pages eacheach is using a few hundred pages at any moment

It works until a moment when everybody's working part grows at once. Then a page is needed, and every frame is taken.

The three things that can be done

OptionWhy it is or is not acceptable
terminate the processunacceptable. Paging is meant to be invisible to the program, and a program that dies because the machine is busy is not invisible
swap a whole process out, as in Chapter sixty nineacceptable, and still used when memory is desperately short. It is crude: a whole process is stopped to free frames a few of which are needed
page replacementthe answer: free one frame by choosing a page nobody will need soon

The fault service routine, with replacement

Chapter eighty one's list, with the new steps in it. This is the sequence a question asks for.

StepWhat happens
1find where the wanted page is on disk
2look for a free frame
3if there is one, use it
4if there is not, run a page replacement algorithm to choose a victim
5if the victim has been modified, write it out to disk
6mark the victim invalid in its owner's page table, and remove it from the TLB
7read the wanted page into the freed frame
8set the entry valid, and restart the faulting instruction

Steps 5 and 6 are where the marks are.

Step 5 can double the cost of the fault, because it is a second disk transfer: one out, one in. Chapter eighty two priced it exactly: with 70 per cent of victims dirty the average fault went from 8 milliseconds to 13.6, and the program from 81 to 137 times slower.

Step 6 is the one that is forgotten and the one that corrupts memory. The victim's owner must not be able to reach that frame any more, and a stale TLB entry is a way to reach it. Chapter seventy six said the same thing about any change to a page table entry.

munotes.in337

Page Replacement: The Problem

The modify bit earns its keep

The hardware sets the modify bit, also called the dirty bit, when a page is written to. The replacement code reads it.

The victim isWhat has to happenCost
clean: never written since it came innothing. The copy on disk is still correct, so the frame is simply takenone transfer, the read
dirty: written toit must be written out before the frame is reusedtwo transfers

So an algorithm that can choose between two equally good victims should choose the clean one, and a page of read-only code is always clean, which makes program text the cheapest thing in memory to throw away: it can always be read again from the program file.

What makes one algorithm better than another

The lowest page fault rate on the same reference string with the same number of frames. Nothing else, and in particular not how clever it sounds.

TermWhat it means
reference stringthe sequence of page numbers a program touches, in order
frameshow many frames this process is allowed. It is the other half of every question
page fault ratefaults divided by references, which is what Chapter eighty two's p is

A reference string is page numbers, not addresses. The offset inside the page makes no difference to whether the page is in memory, so the addresses 0, 100 and 4000 with a 4096 byte page are all page 0 and are one reference for this purpose.

And a repeat costs nothing: two references in a row to the same page can cause at most one fault, so a string 1 1 1 2 behaves exactly like 1 2. That is worth knowing in the hall, because it shortens a long string before the work starts.

More frames, fewer faults

The relationship every question assumes, computed here on the reference string this book uses throughout the next four chapters:

Faults: frames 1 to 7

Reference string: 7 0 1 2 0 3 0 4 2 3 0 3 2 1 2 0 1 7 0 1

Twenty references over six distinct pages. The faults, counted by sim/pagerepl.py:

FramesFIFOOptimalLRU
1202020
2151317
315912
41088
5977
6666
7666

Three things to take from that table, and each is a question in itself.

One: with one frame every reference faults. Twenty references, twenty faults, because no page survives to the next reference.

Two: with six frames all three algorithms give six faults, which is the number of distinct pages. Once there are as many frames as the program has pages, the algorithm stops mattering: every page is brought in once and never thrown out. That is why a question always gives a frame count smaller than the number of distinct pages.

munotes.in338

Page Replacement: The Problem

Three: more frames is not a guarantee. FIFO gives 15 faults at two frames and 15 again at three: the extra frame bought nothing at all on this string. Chapter eighty five shows a string where an extra frame makes FIFO worse, which is Belady's anomaly, and it is found by search rather than asserted.

The five algorithms this module needs

ChapterAlgorithmChooses as victim
30FIFOthe page that came in earliest
31Optimalthe page that will not be used for the longest time to come
32LRUthe page that was used longest ago
33Second chance, and the counting algorithms LFU and MFUthe oldest page that has not been used since last time, or the least or most frequently used
34(comparison)

Every one of them is the same three steps, and only the choice of victim changes.

Distinctions that carry marks

Swapping out a processPage replacement
Freesevery frame the process hasone frame
Stopsthe whole processnothing: the faulting process was waiting anyway
Decided bythe medium term schedulerthe replacement algorithm, at the fault
Clean victimDirty victim
The modify bit is01
Transfers neededonetwo
Examplea page of program codea page of the heap that has been written

What it does not mean

Replacement is not swapping. One page moves, not a process.

The victim need not belong to the faulting process. Whether it may belong to another one is the global against local question of Chapter ninety.

A page fault is not the algorithm's fault. The algorithm decides only which page leaves, and therefore how many faults happen later.

More frames does not always mean fewer faults. It usually does, and FIFO can break the rule.

The reference string is not the addresses. It is the page numbers, and repeats next to each other collapse.

Quick revision

  • Demand paging over allocates memory on purpose; when every frame is taken and a page is

needed, something must give.

  • Of the three options, terminating the process is unacceptable, swapping a whole process out is

crude, and page replacement frees one frame.

  • The service routine: find the page on disk, look for a free frame, else choose a victim,

write it out if dirty, invalidate it in its page table and in the TLB, read the page in, restart the instruction.

  • A dirty victim costs two transfers and a clean one costs a single read; program
munotes.in339

Page Replacement: The Problem

text is always clean.

  • An algorithm is judged by the

page fault rate on the same reference string with the same number of frames.

  • A reference string is page numbers, and consecutive repeats of a page cause at most one

fault.

  • On the standard string of twenty references over six pages: one frame gives 20 faults

for every algorithm; six frames give 6 for every algorithm; three frames give 15 for FIFO, 12 for LRU and 9 for optimal.

  • FIFO gives 15 faults at both two and three frames, and Chapter eighty five shows it getting

worse with more frames.

Test yourself

  1. How can the operating system run out of frames if it is managing memory properly? Because

demand paging lets it promise more memory than exists, which is safe while no process needs all its pages, and stops being safe when several processes need more at the same moment.

  1. Why is terminating a process not an acceptable answer? Paging is meant to be invisible to

the program; a program that dies because other programs are busy is not.

  1. Give the steps of servicing a fault when no frame is free. Find the page on disk; find no

free frame; choose a victim by the replacement algorithm; write the victim out if it is dirty; mark it invalid in its page table and remove it from the TLB; read the wanted page in; set the entry valid; restart the instruction.

  1. Why does the modify bit matter, and by how much? A clean victim needs one transfer and a

dirty one needs two, so the bit can halve the cost of a fault; Chapter eighty two's sum put the difference at 81 against 137 times slower.

  1. What is forgotten at step 6, and why is it serious? Removing the victim from the TLB. A

stale entry lets its former owner reach a frame that now belongs to somebody else.

  1. On what basis is a replacement algorithm judged? The page fault rate it gives on the same

reference string with the same number of frames. 7. The addresses 0, 100 and 4000 are touched in turn, with a 4096 byte page. How long is the reference string? One reference: all three addresses are in page 0.

  1. With six frames and six distinct pages, which algorithm is best? None of them: every page

is loaded once and never replaced, so all give six faults. The algorithm matters only when the frames are fewer than the pages in use.

Contents This chapter on its own page

munotes.in340

Chapter Eighty-Five

FIFO Replacement, and Belady's Anomaly

Syllabus topic Module 2, "Virtual Memory Management - Page Replacement: FIFO"

In one line

Throw out the page that has been in memory longest, whatever it is being used for.

The algorithm

First in, first out. The operating system keeps the resident pages in a queue in the order they arrived; the victim is always the page at the head, and the new page joins the tail.

What is neededHow little it costs
a queue of the resident pagesone pointer per frame, or a list
no information about usenothing is recorded when a page is read or written
the choice itselfone step: take the head

Nothing about how the page has been used is recorded, and that is the whole objection to FIFO. The page that came in first may be the most heavily used page the program has: a table of constants read on every instruction is as old as the program itself.

An equivalent implementation keeps the time each page came in and picks the oldest. It is the same algorithm and the same answer; the queue is cheaper.

Worked on the standard string

Twenty references over six pages, three frames, starting with all three frames empty.

Replacement: first in first out, 3 frames

Reference string: 7 0 1 2 0 3 0 4 2 3 0 3 2 1 2 0 1 7 0 1

ReferenceFrame 1Frame 2Frame 3FaultEvicted
77F
070F
1701F
2201F7
0201H
3231F0
0230F1
4430F2
2420F3
3423F0
0023F4
3023H
2023H
1013F2
2012F3
0012H
1012H
7712F0
0702F1
1701F2

Page faults: 15

Hits: 5

Fifteen faults and five hits out of twenty references.

hit ratio = 5 / 20 = 0.25

fault rate = 15 / 20 = 0.75

A hit ratio of 25 per cent, which Chapter eighty nine sets beside the other algorithms on the same string.

How to do this in the hall

The reliable method, and the one mistake that costs the marks.

StepWhat to do
1draw the frames as rows or columns and keep them in a fixed order: frame 1 is frame 1 all the way down
2fill the empty frames first, in order, and mark each of those a fault
3once full, replace the oldest arrival, not the frame you last wrote to
4mark every row a fault or a hit, and count the faults at the end
munotes.in341

FIFO Replacement, and Belady's Anomaly

The mistake is losing track of which page is oldest after a few replacements, because the oldest page is not in the frame you would guess. Write the arrival order in the margin, in a line: after the tenth reference above, the order of arrival is 4, then 2, then 3, so the next victim is 4 and not the page in frame 1 by position.

Note also what happened at references 12 and 13: page 3 and page 2 are hits. A hit changes nothing at all in FIFO: not the queue, not the order, not the contents. That is exactly what LRU changes in Chapter eighty seven.

Belady's anomaly

Chapter eighty four said more frames usually means fewer faults. For FIFO it can be the other way round, and that fact is named after the engineer who published it.

The string below is a different one, chosen because the anomaly is on it. sim/pagerepl.py searched every frame count from one to eight and reported one pair where more memory did more harm: three frames against four.

FIFO with three frames

Replacement: first in first out, 3 frames

Reference string: 1 2 3 4 1 2 5 1 2 3 4 5

ReferenceFrame 1Frame 2Frame 3FaultEvicted
11F
212F
3123F
4423F1
1413F2
2412F3
5512F4
1512H
2512H
3532F1
4534F2
5534H

Page faults: 9

Hits: 3

FIFO with four frames, on the same string

Replacement: first in first out, 4 frames

Reference string: 1 2 3 4 1 2 5 1 2 3 4 5

ReferenceFrame 1Frame 2Frame 3Frame 4FaultEvicted
11F
212F
3123F
41234F
11234H
21234H
55234F1
15134F2
25124F3
35123F4
44123F5
54523F1

Page faults: 10

Hits: 2

Nine faults with three frames, ten faults with four. The extra frame made the program worse. Nothing was changed except the amount of memory.

munotes.in342

FIFO Replacement, and Belady's Anomaly

Why it can happen

The reason is worth understanding because it also says which algorithms are safe.

FIFOLRU and optimal
The set of pages in memory with n + 1 framesneed not contain the set with n framesalways contains it
Name for the propertyFIFO is not a stack algorithmthey are stack algorithms
Can more frames mean more faultsyesno, never

In FIFO the queue order itself changes when the number of frames changes, so a page that survived with three frames can be thrown out earlier with four. An algorithm that chooses its victim by use cannot do this, because with more frames it simply keeps everything it kept before, and more. That is proved for LRU in Chapter eighty seven.

What to say in an answer. Belady's anomaly is that with FIFO replacement, increasing the number of frames can increase the number of page faults; it happens because FIFO is not a stack algorithm; and the standard demonstration is the string above at three and four frames, nine faults against ten.

Distinctions that carry marks

FIFOLRU
Victimthe oldest arrivalthe page used longest ago
A hit changesnothingthe page's position: it becomes the most recent
Information keptarrival orderorder of use, updated on every reference
Belady's anomalypossibleimpossible
Faults on the standard string, 3 frames1512

What it does not mean

FIFO is not the same as swapping out the oldest process. It is about pages, and the age is the age of the page in memory, not of the process.

The oldest page is not the one in the first frame. After a few replacements the arrival order and the frame numbering have nothing to do with each other.

A hit does not protect a page under FIFO. A page used on every reference is thrown out as soon as it reaches the head of the queue.

Belady's anomaly is not a bug. It is a property of the algorithm, and it is the reason the property of being a stack algorithm is worth naming.

The anomaly is not common. It is possible, which is enough to disqualify FIFO from being trusted; on the standard string of this chapter more frames did help.

Quick revision

  • FIFO throws out the page that arrived earliest, kept in a queue; the victim is the head

and the arrival is the tail.

  • It records nothing about use, so a heavily used old page is thrown out; a

hit changes nothing.

  • On the standard string with three frames: 15 faults, 5 hits, a hit ratio of
munotes.in343

FIFO Replacement, and Belady's Anomaly

5 / 20 = 0.25.

  • In the hall: keep the frames in a fixed order, fill the empty ones first, and

track the arrival order in the margin.

  • Belady's anomaly: with FIFO, more frames can mean more faults. On the string 1 2 3 4 1

2 5 1 2 3 4 5, three frames give 9 faults and four frames give 10.

  • It happens because FIFO is not a stack algorithm: the pages resident with n + 1 frames need

not include those resident with n. LRU and optimal are stack algorithms and cannot show the anomaly.

Test yourself

  1. State the FIFO rule and the data structure it needs. Replace the page that has been in

memory longest; a queue of resident pages, with the victim at the head and arrivals at the tail.

  1. What does FIFO record about how a page is used? Nothing, which is its weakness: a page

used constantly is evicted as soon as it is the oldest.

  1. Work FIFO on 7 0 1 2 0 3 0 4 2 3 0 3 2 1 2 0 1 7 0 1 with three frames. How many faults?

Fifteen, with five hits, a hit ratio of 0.25.

  1. What changes in a FIFO queue when a reference hits? Nothing at all.
  2. State Belady's anomaly. With FIFO replacement, increasing the number of frames can

increase the number of page faults.

  1. Give a string and two frame counts that show it. 1 2 3 4 1 2 5 1 2 3 4 5 gives nine faults

with three frames and ten with four.

  1. Why can LRU not show the anomaly? Because it is a stack algorithm: the set of pages

resident with n + 1 frames always contains the set resident with n, so extra memory can never cause a fault that smaller memory avoided.

  1. Is FIFO ever a reasonable choice? It is the cheapest to implement and needs no hardware

support, so it appears where simplicity matters more than the fault rate; for a general purpose system its fault rate and the anomaly rule it out.

Contents This chapter on its own page

munotes.in344

Chapter Eighty-Six

Optimal Replacement

Syllabus topic Module 2, "Virtual Memory Management - Page Replacement: Optimal"

In one line

Throw out the page that will not be needed for the longest time to come, which gives the fewest possible faults and cannot be implemented because it needs the future.

The rule

Replace the page that will not be used for the longest time. Among the resident pages, look forward through the rest of the reference string; the victim is the one whose next use is furthest away, and a page that is never used again is the best victim of all.

It is provably the best possible. No algorithm can have a lower fault rate on the same string with the same number of frames. That is why it is called optimal, and it is also called MIN and OPT in the literature.

And it cannot be implemented, because at the moment of the fault the operating system does not know what the program will do next. A program's future references are not knowable in general: they depend on its input.

So what is it for

Two real uses, and a question sometimes asks exactly this.

UseHow
a yardstickrun the real algorithm and the optimal one on the same recorded reference string, and the gap is how much there is left to win
a bound in an argumentif optimal takes nine faults, no algorithm anybody invents will take eight

The comparison this chapter sets up: on the string below, optimal takes 9 faults and FIFO takes 15. FIFO is therefore doing two thirds again as much work as the best possible, and Chapter eighty seven's LRU comes in between at 12.

Worked on the standard string

The same twenty references and the same three frames as Chapter eighty five.

Replacement: optimal, 3 frames

Reference string: 7 0 1 2 0 3 0 4 2 3 0 3 2 1 2 0 1 7 0 1

ReferenceFrame 1Frame 2Frame 3FaultEvicted
77F
070F
1701F
2201F7
0201H
3203F1
0203H
4243F0
2243H
3243H
0203F4
3203H
2203H
1201F3
2201H
0201H
1201H
7701F2
0701H
1701H

Page faults: 9

Hits: 11

Nine faults and eleven hits.

hit ratio = 11 / 20 = 0.55

fault rate = 9 / 20 = 0.45

munotes.in345

Optimal Replacement

How to do it in the hall

The method is mechanical, and writing the three distances down is what stops the mistakes.

StepWhat to do
1at a fault, list the pages in the frames
2for each, look forward in the string and write down how many references away its next use is
3evict the one with the largest distance; a page that never appears again has an infinite distance and wins
4if two are never used again, either may go: the fault count is the same

Work the fourth reference of the table above as the example. The frames hold 7, 0 and 1, and the reference is 2. Looking forward from there: 0 is used at the very next reference, 1 is used much later, and 7 is not used again until the last three references. So 7 is the victim, which is what the table shows.

Look forward, never backward. Optimal is about the future; the algorithm that looks backward is LRU, and it is the next chapter. Writing the previous uses down by mistake is the commonest way to lose the marks in this question.

Optimal is a stack algorithm

Chapter eighty five needed this property to explain why FIFO can get worse with more memory. Optimal has it, so it never can.

On Chapter eighty five's anomaly string, optimal at one to six frames gives:

Faults: frames 1 to 6

Reference string: 1 2 3 4 1 2 5 1 2 3 4 5

Frames123456
Optimal1297655
FIFO121291055

Never rising. FIFO on the same string went 12, 12, 9, 10, 5, 5, and the rise from 9 to 10 is the anomaly.

Distinctions that carry marks

OptimalFIFO
Looksforwardat arrival order
Needsthe future, so it cannot be builta queue
Fault ratethe lowest possiblehigh
Belady's anomalyimpossiblepossible
On the standard string, 3 frames9 faults15 faults
OptimalLRU, Chapter eighty seven
Directionthe next usethe last use
Implementablenoyes, with hardware help
Relationshipthe idealthe mirror image: LRU assumes the recent past predicts the near future

What it does not mean

Optimal is not the best algorithm to use. It is the best result any algorithm could get. It cannot be used at all.

It is not clairvoyance about the program's data. It is defined on a given reference string; that is why it can be computed afterwards from a recording and not while the program runs.

Zero faults is not what optimal means. The first reference to a page must always fault, so on the standard string even optimal pays six faults for the six distinct pages; its nine is six unavoidable faults and three more.

munotes.in346

Optimal Replacement

A tie does not change the answer. When two resident pages are never used again, either choice gives the same total.

Quick revision

  • Optimal replaces the page whose next use is furthest in the future, and a page never used

again first of all.

  • It has the lowest possible fault rate, which is why it is the yardstick, and it

cannot be implemented because the future is unknown.

  • On the standard string with three frames: 9 faults, 11 hits, a hit ratio of

11 / 20 = 0.55, against FIFO's 15 faults.

  • Method: at each fault, write the forward distance of each resident page and evict the

largest.

  • Look forward. Looking backward is LRU.
  • Optimal is a stack algorithm: on the anomaly string it gives 12, 9, 7, 6, 5, 5 as frames

rise, never more faults with more frames.

  • Its floor is the number of distinct pages: six of its nine faults on the standard string

are first references.

Test yourself

  1. State the optimal rule. Replace the resident page that will not be used for the longest

time to come.

  1. Why is it called optimal, and why can it not be used? No algorithm can fault less on the

same string with the same frames; and it needs to know the program's future references, which the operating system cannot.

  1. What are its two real uses? As a yardstick against which a real algorithm is measured on a

recorded reference string, and as a bound in an argument about what any algorithm could achieve.

  1. Work optimal on 7 0 1 2 0 3 0 4 2 3 0 3 2 1 2 0 1 7 0 1 with three frames. Nine faults and

eleven hits, a hit ratio of 0.55. 5. At the fourth reference the frames hold 7, 0 and 1 and page 2 is wanted. Which page goes, and why? Page 7: 0 is needed at the very next reference and 1 later, but 7 is not needed again until near the end, so its forward distance is the largest.

  1. Can optimal show Belady's anomaly? No: it is a stack algorithm, and on the anomaly string

its faults fall steadily as frames rise.

  1. Why can optimal not reach zero faults? Every page must be brought in the first time it is

referenced, so the number of distinct pages is a floor.

  1. How does optimal relate to LRU? LRU is its mirror image: optimal looks forward to the next

use, LRU looks back to the last use and assumes the recent past predicts the near future.

Contents This chapter on its own page

munotes.in347

Chapter Eighty-Seven

LRU Replacement

Syllabus topic Module 2, "Virtual Memory Management - Page Replacement: LRU"

In one line

Throw out the page that was used longest ago, on the argument that a page not used for a long time will probably not be used soon.

The rule, and why it is reasonable

Least recently used: replace the resident page whose last use is furthest in the past.

It is optimal with the string reversed. Optimal looks forward to the next use; LRU looks backward to the last use. The justification is locality of reference: a program's recent past is the best cheap guess at its near future, and the same fact makes the TLB of Chapter seventy six work and the working set of Chapter eighty nine small.

LRU is therefore the algorithm that a question means by "a good, implementable algorithm", and the comparison worth remembering is the three numbers on one string: optimal 9, LRU 12, FIFO 15.

Worked on the standard string

Replacement: least recently used, 3 frames

Reference string: 7 0 1 2 0 3 0 4 2 3 0 3 2 1 2 0 1 7 0 1

ReferenceFrame 1Frame 2Frame 3FaultEvicted
77F
070F
1701F
2201F7
0201H
3203F1
0203H
4403F2
2402F3
3432F0
0032F4
3032H
2032H
1132F0
2132H
0102F3
1102H
7107F2
0107H
1107H

Page faults: 12

Hits: 8

Twelve faults and eight hits.

hit ratio = 8 / 20 = 0.4

fault rate = 12 / 20 = 0.6

How to do it in the hall

StepWhat to do
1keep a list of the resident pages in order of use, most recent first
2on a hit, move that page to the front of the list. This is the step FIFO does not have
3on a fault with no free frame, the victim is the page at the back of the list
4put the new page at the front

Take the eighth reference of the table above as the example. The frames hold 2, 0 and 1 and page 4 is wanted. The last uses were: 0 at the reference just before, 2 four references back, 1 five references back. So 1 is the victim, which is what the table shows, and notice that FIFO would have evicted 2.

munotes.in348

LRU Replacement

The two implementations

A question asks how LRU is implemented, and there are exactly two answers.

One: counters

PartWhat it does
a logical clock in the processor, incremented on every memory referencegives an order to references
a time of use field in each page table entrythe hardware writes the clock into it on every reference
the replacement stepsearch the table for the smallest time of use

Two costs, and both are the reason this is not free. Every memory reference now writes to memory as well, to update the field; and every replacement searches the whole table. The clock can also overflow, which has to be handled.

Two: a stack

PartWhat it does
a doubly linked list of the page numberskeeps the order of use directly
on every referencethe page is taken out of the middle and put on the top
the replacement stepthe victim is the page at the bottom: no search at all

Moving a page to the top changes up to six pointers: the page's own two, and one each in the entries before and after it in two places. That is cheap in hardware and hopeless in software.

Neither implementation can be done by the operating system alone. Both need work on every memory reference, and software cannot see a memory reference: to intervene it would have to trap, which costs a fault on every instruction. LRU needs hardware support, and real machines do not provide it. They provide one reference bit per page, and Chapter eighty eight is what can be built out of that.

LRU is a stack algorithm

The proof is short enough to write in an answer. With n frames, the pages in memory under LRU are exactly the n most recently used pages. With n + 1 frames they are the n + 1 most recently used, which contains the n most recently used. So a page present with n frames is present with n + 1 frames, and no fault can be created by adding memory: Belady's anomaly is impossible.

Chapter eighty five's anomaly string, with both rows for comparison:

Faults: frames 1 to 6

Reference string: 1 2 3 4 1 2 5 1 2 3 4 5

Frames123456
LRU121210855
FIFO121291055

Two things in that table, and both are worth saying. LRU never rises, as the proof requires. And at three frames LRU is worse than FIFO on this string, 10 against 9: being a better algorithm in general does not mean being better on every string.

munotes.in349

LRU Replacement

Distinctions that carry marks

FIFOLRUOptimal
Victimoldest arrivaloldest usefurthest next use
A hit changesnothingthe order of usenothing
Information neededarrival orderthe pastthe future
Implementableyes, triviallyyes, with hardwareno
Belady's anomalypossibleimpossibleimpossible
Standard string, 3 frames15129
Counter implementationStack implementation
Per referencewrite the clock into the entrymove the page to the top, up to six pointers
Per replacementsearch the whole table for the smallesttake the bottom, no search
Extra problemthe clock can overflowmore work on every reference

What it does not mean

LRU is not least frequently used. It counts recency, not how often. A page used a thousand times an hour ago goes before a page used once a moment ago. The frequency algorithms are in the next chapter, and they are worse.

LRU does not need the future. That is optimal. LRU needs a record of the past, which is why it can be built.

Software cannot implement LRU. Both implementations need hardware action on every reference.

LRU is not always better than FIFO. On the anomaly string at three frames it is worse. It is better on average and it cannot show the anomaly, which is why it is preferred.

A hit is not free under LRU. It costs the bookkeeping: a field written, or pointers moved. It is free under FIFO, which is FIFO's only advantage.

Quick revision

  • LRU replaces the page whose last use is furthest in the past, justified by

locality of reference.

  • It is optimal looking backwards, and on the standard string the three compare as

optimal 9, LRU 12, FIFO 15.

  • On the standard string with three frames: 12 faults, 8 hits, a hit ratio of

8 / 20 = 0.4.

  • Method: keep the resident pages in order of use, move a page to the front on a hit, and

evict from the back.

  • Counters: a clock written into the entry on every reference, and a search of the table at

every replacement; the clock can overflow.

  • Stack: a doubly linked list, the page moved to the top on every reference at a cost of up

to six pointers, and the victim is the bottom with no search.

  • Both need hardware on every reference; software cannot do it. Real machines give only a

reference bit.

  • LRU is a stack algorithm: with n frames it holds the n most recently used pages, so more

frames can never cause more faults.

Test yourself

  1. State the LRU rule and its justification. Replace the page whose last use was longest ago;
munotes.in350

LRU Replacement

by locality of reference the recent past is a good guess at the near future.

  1. Work LRU on 7 0 1 2 0 3 0 4 2 3 0 3 2 1 2 0 1 7 0 1 with three frames. Twelve faults and

eight hits, a hit ratio of 0.4. 3. At the eighth reference the frames hold 2, 0 and 1 and page 4 is wanted. Which page goes under LRU, and which under FIFO? Under LRU page 1, whose last use is furthest back; under FIFO page 2, which arrived earliest.

  1. Describe the counter implementation and its two costs. A logical clock is written into

each page table entry on every reference; every reference therefore also writes to memory, and every replacement searches the whole table for the smallest value.

  1. Describe the stack implementation and its cost. A doubly linked list of page numbers, with

the referenced page moved to the top on every reference, up to six pointers changed; replacement is free because the victim is the bottom.

  1. Why can the operating system not implement LRU by itself? Both implementations need work

on every memory reference, and software can only act on a reference by trapping, which would cost a fault per instruction.

  1. Prove that LRU cannot show Belady's anomaly. With n frames it holds the n most recently

used pages; with n + 1 it holds the n + 1 most recently used, which include them all. Adding a frame can therefore never remove a page that would have been present.

  1. Is LRU always better than FIFO? No. On 1 2 3 4 1 2 5 1 2 3 4 5 with three frames LRU takes

10 faults and FIFO 9. It is better on average and it is safe from the anomaly.

Contents This chapter on its own page

munotes.in351

Chapter Eighty-Eight

Second Chance, and the Counting Algorithms

Syllabus topic Module 2, "Virtual Memory Management - Page Replacement: Second Chance, Counting Algorithms"

In one line

Real machines give one reference bit per page, and second chance is FIFO that skips a page whose bit is set, which is as close to LRU as cheap hardware gets.

The one bit there is

Chapter eighty seven ended with the problem: LRU needs hardware on every reference and machines do not provide it. What they do provide is a reference bit.

WhoDoes what to the reference bit
the hardwaresets it to 1 on any read or write of the page, and never clears it
the operating systemmay clear it to 0 whenever it likes

One bit says only whether the page has been used since the operating system last cleared it. It does not say when, and it does not say how often. Everything in this chapter is built from that one bit, and the whole difficulty is that it is so little.

In this book a page enters memory with its reference bit 0, and the bit is set by any later reference to it. The convention is stated because a book that loads a page with the bit already set gets a different trace from the same string, and a question must be answered in the convention it is set in.

Additional reference bits

Keep a history instead of a single bit. Give each page an eight bit byte in a table in memory. At a regular interval, say every hundred milliseconds, the operating system shifts every byte one place right and puts the page's reference bit into the high bit, then clears the reference bit.

The byte readsWhat it means
00000000not used in any of the last eight intervals
11111111used in every one of the last eight intervals
11000100used in the two most recent intervals, and once four intervals ago
01110111used in seven of the eight, but not the most recent

Read the bytes as unsigned numbers and the smallest is the page used least recently. 11000100 is 196 and 01110111 is 119, so the second page is the better victim even though it has been used more often, because the high bits are the recent intervals.

It is an approximation of LRU, and how good it is depends on how many bits are kept and how often they are shifted. Eight bits and eight intervals is the usual example; one bit is second chance below.

Second chance

FIFO, but a page whose reference bit is set is given another turn instead of being thrown out. It is also called the clock algorithm, because the resident pages are held in a circular queue with a hand that points at the next candidate.

munotes.in352

Second Chance, and the Counting Algorithms

At a fault, look at the page under the handWhat happens
its reference bit is 0it is the victim. Replace it, move the hand on one
its reference bit is 1clear the bit to 0, move the hand on one, and look again. The page keeps its frame: that is the second chance

The hand keeps going until it finds a zero, and it will, because it clears every bit it passes. In the worst case, where every bit is set, it goes all the way round, clears everything, and comes back to where it started: second chance degenerates into FIFO and evicts the page under the hand.

The cost is one bit per frame, a pointer, and no work at all on a hit beyond the bit the hardware sets by itself. That is why this is the algorithm real systems actually use.

Worked on the standard string

Replacement: second chance, 3 frames

Reference string: 7 0 1 2 0 3 0 4 2 3 0 3 2 1 2 0 1 7 0 1

ReferenceFrame 1Frame 2Frame 3FaultEvicted
77F
070F
1701F
2201F7
0201H
3203F1
0203H
4403F2
2402F3
3302F4
0302H
3302H
2302H
1312F0
2312H
0012F3
1012H
7017F2
0017H
1017H

Page faults: 11

Hits: 9

Eleven faults and nine hits, against FIFO's fifteen and LRU's twelve on the same string and the same three frames.

hit ratio = 9 / 20 = 0.45

fault rate = 11 / 20 = 0.55

Worth noticing: second chance beat LRU here. It is an approximation of LRU, so on average it is a little worse; on a particular string either may win, exactly as Chapter eighty seven found for LRU against FIFO.

How to do it in the hall

StepWhat to do
1draw the frames in a ring, and mark where the hand is
2beside each page write its reference bit
3on a hit, set that page's bit to 1 and change nothing else
4on a fault, walk the hand: clear a 1 and move on, replace at the first 0
5after replacing, leave the hand on the next frame, not on the one you just filled
munotes.in353

Second Chance, and the Counting Algorithms

Step 5 is where the marks go. The hand does not go back to the start after a replacement.

Enhanced second chance

The refinement that uses the modify bit of Chapter seventy seven as well as the reference bit, because Chapter eighty two showed a dirty victim costs two transfers instead of one.

Each page is then in one of four classes, written as (reference, modify):

ClassMeaningHow good a victim
(0, 0)not used recently, not modifiedthe best: not wanted, and free to throw away
(0, 1)not used recently, but modifiedgood, but it must be written out first
(1, 0)used recently, not modifiedprobably wanted again soon
(1, 1)used recently and modifiedthe worst: probably wanted, and dear to remove

The algorithm walks the circular queue looking for a page in the lowest numbered non empty class, clearing reference bits as it goes, and it may go round more than once: up to four passes in the worst case. The gain is that it prefers a clean page, which halves the cost of the fault.

This is the scheme the Macintosh virtual memory of the period used, and MU's textbook names it for that reason. The principle to remember is the ordering of the four classes.

The counting algorithms

Both keep a count of how many times each page has been referenced, and they differ only in which end they take.

AlgorithmVictimThe argument for it
LFU, least frequently usedthe smallest counta page hardly used is probably not needed
MFU, most frequently usedthe largest countthe page with the smallest count has probably just been brought in and not had its turn yet

The objection to LFU, and it is the one a question wants. A page used heavily during the program's start up keeps its high count for the whole run, long after it has stopped being needed, and can never be evicted. The fix is aging: shift every count one place right at intervals, so old use fades.

The objection to MFU is simpler: it throws out the page the program is using most. Both are uncommon, dear to implement, and poor approximations of optimal, and that sentence is the answer to "comment on the counting algorithms".

In this book a tie between two equal counts is broken by last use, oldest first. A convention is needed because ties are common, and a question should say which it wants.

All six on one string

Faults: frames 3

Reference string: 7 0 1 2 0 3 0 4 2 3 0 3 2 1 2 0 1 7 0 1

munotes.in354

Second Chance, and the Counting Algorithms

Frames3
Optimal9
Second chance11
LFU11
LRU12
MFU14
FIFO15

Optimal is the floor at nine and FIFO the ceiling at fifteen, and everything implementable sits between them. Chapter eighty nine reads this table properly, with the hit ratios and the frame counts varied.

Distinctions that carry marks

Reference bitModify bit
Set by the hardware onany use, read or writea write only
Cleared bythe operating systemthe operating system, after writing the page out
Used forwhich page to replacewhether the victim must be written out
Together they givethe four classes of enhanced second chance
FIFOSecond chance
Structurea queuea circular queue with a hand
A hitchanges nothingsets the page's reference bit
The oldest pagealways goesgoes only if its bit is 0
Worst caseit is the algorithmdegenerates to FIFO
Standard string, 3 frames15 faults11 faults
LRULFU
Countsrecency: when last usedfrequency: how often used
A page used a thousand times an hour agois the victimis nearly immortal, without aging
Implementable cheaplynono

What it does not mean

The reference bit does not say when the page was used. Only whether it has been used since the bit was cleared.

Second chance is not LRU. It is a one bit approximation of it, and on a given string it may be better or worse.

The hand does not reset. It stays where the last replacement left it, which is what makes the algorithm circular.

Enhanced second chance does not always find a (0, 0) page. It may pass round several times and settle for a worse class; the point is the preference, not a guarantee.

LFU is not LRU with counting done properly. It answers a different question, and without aging it keeps pages that were busy long ago.

Quick revision

  • The hardware gives one reference bit per page: set on any use, cleared only by the

operating system.

  • Additional reference bits: an eight bit history shifted right at each interval with the

reference bit put in the high end; the smallest byte read as a number is the least recently used.

  • Second chance, the clock algorithm: a circular queue and a hand; a bit of 0 is

replaced, a bit of 1 is cleared and skipped. With every bit set it degenerates to FIFO.

  • On the standard string with three frames: 11 faults, 9 hits, a hit ratio of

9 / 20 = 0.45.

  • After a replacement the hand stays on the next frame.
  • Enhanced second chance uses reference and modify together:
munotes.in355

Second Chance, and the Counting Algorithms

(0, 0) best, (0, 1), (1, 0), (1, 1) worst, up to four passes, and it prefers a clean victim because that halves the fault's cost.

  • LFU replaces the smallest count and keeps start up pages for ever unless the counts are

aged; MFU replaces the largest count on the argument that a small count means a new arrival. Both are dear and poor.

  • On the standard string at three frames: optimal 9, second chance 11, LFU 11, LRU 12, MFU 14,

FIFO 15.

Test yourself

  1. What does the reference bit tell the operating system, and what does it not? That the page

has been used since the bit was last cleared; not when, and not how often.

  1. How do additional reference bits approximate LRU? Each page keeps a byte of history,

shifted right at each interval with the reference bit put into the high bit; read as unsigned numbers, the smallest byte is the page used least recently.

  1. Which is the better victim, a page reading 11000100 or one reading 01110111, and why? The

second: 119 is smaller than 196, and the high bits are the most recent intervals, so it has not been used lately.

  1. State the second chance rule. In a circular queue, examine the page under the hand: if its

reference bit is 0 replace it; if 1, clear the bit, move on and examine the next.

  1. What happens if every reference bit is set? The hand goes right round clearing bits and

replaces the page it started on: the algorithm degenerates to FIFO.

  1. Work second chance on the standard string with three frames. Eleven faults and nine hits,

a hit ratio of 0.45.

  1. Give the four classes of enhanced second chance in order of preference. (0, 0), then (0,

1), then (1, 0), then (1, 1): not recently used and clean is best, recently used and dirty is worst.

  1. What is wrong with LFU, and what is the fix? A page used heavily at start up keeps a high

count and cannot be evicted long after it is needed; the fix is to age the counts by shifting them right at intervals.

Contents This chapter on its own page

munotes.in356

Chapter Eighty-Nine

Comparing the Algorithms, and the Hit Ratio

Syllabus topic Module 2, "Virtual Memory Management - Page Replacement"

In one line

On the same string with the same frames, optimal is the floor and no implementable algorithm is far above it except FIFO and MFU.

The table

Faults: frames 1 to 7

Reference string: 7 0 1 2 0 3 0 4 2 3 0 3 2 1 2 0 1 7 0 1

Frames1234567
Optimal201398766
LRU2017128766
Second chance2017118766
LFU2015118766
MFU20201412966
FIFO20151510966

Six things to read off it, and each of them is a sentence worth writing in an answer.

What the table showsWhy
one frame: 20 faults for everythingno page survives to its next reference, so the algorithm cannot matter
six frames: 6 faults for everythingthere are six distinct pages, so every page is loaded once and never replaced
optimal is never beatenit is the lowest possible at every frame count, which is what optimal means
six is the floor at every frame countthe six first references must fault whatever is done
MFU and FIFO are the two bad onesMFU throws out what the program is using; FIFO ignores use altogether
the middle three are closeLRU, second chance and LFU are within one fault of each other at three frames, and equal from four frames up

And the anomaly is visible. MFU at 5 frames takes 9 faults and FIFO at 5 frames takes 9, but look along FIFO's row: 15 at three frames and 15 at two. An extra frame bought nothing. Chapter eighty five's other string shows the row actually rising.

The hit ratio

Two definitions, and a question can ask for either.

hit ratio = hits / references

fault rate = faults / references

hit ratio + fault rate = 1

At three frames on this string, out of twenty references:

AlgorithmFaultsHitsHit ratioWorking
Optimal9110.5511 / 20 = 0.55
Second chance1190.459 / 20 = 0.45
LFU1190.459 / 20 = 0.45
LRU1280.48 / 20 = 0.4
MFU1460.36 / 20 = 0.3
FIFO1550.255 / 20 = 0.25

Give the ratio as the division, then the number. A hit ratio written as a bare 0.45 has no marks in it; written as 9 / 20 = 0.45 it shows where it came from, and the same rule applied to every average in Module 1.

munotes.in357

Comparing the Algorithms, and the Hit Ratio

The warning, and it matters

A fault rate of 0.55 in an exercise is not a fault rate of 0.55 on a machine. Chapter seventy five seven showed that p = 0.001 makes a program eighty one times slower. If a real program faulted on 55 per cent of its memory references it would run about forty four thousand times slower than one that never faulted, and no such machine has ever been built.

p = 0.55

EAT = 0.45 × 100 + 0.55 × 8,000,000 = 45 + 4,400,000 = 4,400,045

slowdown = 4,400,045 / 100 = 44,000.45

The two numbers measure different things.

The p of Chapter eighty twoThe fault rate of a reference string exercise
Denominatorevery memory reference the program makes, billions of themthe twenty page references the question gives
What a reference isone instruction fetch or data accessa page, already reduced, with repeats collapsed
Realistic value0.0000010.25 to 0.55
What it is forpricing the slowdowncomparing algorithms

So the right sentence about an exercise is comparative: FIFO faults three times as often as optimal on this string, not that any machine faults on a quarter of its references.

Choosing one

The summary a "compare and comment" question is asking for.

AlgorithmFault rateCost to implementAnomalyVerdict
Optimalthe best possibleimpossible: needs the futurenothe yardstick
LRUclose to optimalneeds hardware on every referencenothe standard to aim at
Second chanceclose to LRUone bit per page and a pointernowhat real systems use
Enhanced second chanceas good, and cheaper faultsreference and modify bitsnoused where the write back matters
LFUmixeda counter per page, and agingnorare
MFUpoora counter per pagenorare, and hard to defend
FIFOpoora queue: the cheapestyesonly where simplicity rules

What makes a comparison fair

A question that gives two algorithms different conditions is not comparing them, and marks are given for saying so.

Must be the sameWhy
the reference stringdifferent strings favour different algorithms, as Chapter eighty seven's string favoured FIFO over LRU
the number of framesthe whole table above is a demonstration of that
the starting stateall frames empty, so the first references fault for everybody
whether repeats are collapsed1 1 2 and 1 2 give different fault rates although the same fault count

What it does not mean

A lower fault count on one string does not make an algorithm better. Second chance beat LRU here and would lose on another string. The claim to make is about the average and about the anomaly.

munotes.in358

Comparing the Algorithms, and the Hit Ratio

The hit ratio is not the TLB hit ratio. Chapter seventy six's ratio is about translations in a cache; this one is about pages in memory.

Optimal is not an algorithm anyone can choose. It is the bound.

Equal fault counts at six frames do not make the algorithms equal. They make the memory sufficient, which is a different statement and the reason a question never gives you enough frames.

A better algorithm is not a substitute for enough frames. Every row of the table falls faster with more memory than any algorithm change achieves, which is what Chapter ninety and Chapter ninety one are about.

Quick revision

  • On the standard string with one frame every algorithm gives 20 faults; with six

frames every algorithm gives 6, the number of distinct pages.

  • Six faults is the floor at every frame count, because a first reference must fault.
  • At three frames: optimal 9, second chance 11, LFU 11, LRU 12, MFU 14, FIFO 15.
  • Hit ratio = hits / references, and fault rate = faults / references; they sum to 1.

Write the division, then the number: 9 / 20 = 0.45.

  • The fault rate of an exercise is not the p of Chapter eighty two: that p is over every

memory reference and is about 0.000001, while an exercise's is a quarter or a half over twenty page references.

  • Verdicts: optimal is the yardstick; LRU the standard, needing hardware;

second chance what systems use; FIFO the cheapest and the only one with the anomaly; LFU and MFU rare.

  • A fair comparison needs the same string, the same frames and the same starting state.

Test yourself

  1. Why do all the algorithms give twenty faults with one frame? No page can survive until its

next reference, so every reference faults whatever the rule is.

  1. Why do they all give six faults with six frames? There are six distinct pages, so each is

loaded once and nothing ever has to be replaced.

  1. What is the lowest number of faults any algorithm can have on this string, and why? Six:

every page must be brought in the first time it is referenced.

  1. Give the six fault counts at three frames in order. Optimal 9, second chance 11, LFU 11,

LRU 12, MFU 14, FIFO 15.

  1. Define the hit ratio and give it for LRU at three frames. Hits divided by references:

8 / 20 = 0.4. 6. A student writes that a program faulting on 45 per cent of its references is normal. What is wrong? That figure is the fault rate of a twenty reference exercise, counted over page references with repeats collapsed. A real fault rate is counted over every memory reference and must be about one in a million, or the program becomes tens of thousands of times slower.

munotes.in359

Comparing the Algorithms, and the Hit Ratio

  1. Second chance beat LRU on this string. Does that make it the better algorithm? No. On one

string either may win; LRU is the better approximation on average, and second chance is preferred because it is far cheaper.

  1. Name the three things that must be the same for a comparison to be fair. The reference

string, the number of frames, and the starting state of the frames.

Contents This chapter on its own page

munotes.in360

Chapter Ninety

Allocation of Frames

Syllabus topic Module 2, "Virtual Memory Management - Allocation of Frames"

In one line

The frames have to be divided among the processes, and the two questions are how many each one gets and whose page may be taken when one of them faults.

The two limits

Before any scheme, the arithmetic has two ends.

LimitWhat sets it
the maximum a process can havethe number of frames in the machine
the minimum a process must havethe instruction set

The minimum is set by the hardware, not by policy, and this is the part students miss. A process must have enough frames to hold every page one instruction can touch at once, or the instruction can never complete: it faults, the fault brings in one page and throws out another the same instruction still needs, and it faults again for ever.

The instructionPages it can touch at onceFrames needed
a simple load from memorythe instruction, and the operand2
an instruction that spans a page boundarytwo pages of instruction, and the operand3
a memory to memory movethe instruction, the source, the destination3, and more if any of them spans a boundary
one level of indirect addressingthe instruction, the address word, and the word it points at3

A machine that allows indirect addressing to any depth would need a whole address space of frames in the worst case, so architectures limit the depth. That limit is what makes a minimum possible to state at all.

Equal allocation

Divide the frames equally. With m frames and n processes, each gets m / n, and the remainder goes into the free frame pool.

Worked: 64 frames, 5 processes.

64 / 5 = 12, remainder 4

Twelve frames each, and four in the free pool for whoever faults next.

The objection is obvious once stated: a process of 10 pages and a process of 127 pages get the same twelve frames. The small one cannot use them all and the large one is starved.

Proportional allocation

Divide the frames in proportion to the size of each process.

allocation for process i = (size of process i / total size) × m

Worked, and this is the standard sum: 62 frames, one process of 10 pages and one of 127 pages.

total size = 10 + 127 = 137

ProcessSizeWorkingValueFrames given
P11010 / 137 × 62about 4.534
P2127127 / 137 × 62about 57.4757

frames given out = 4 + 57 = 61

left in the pool = 62 - 61 = 1

Round down and keep the remainder in the pool. A fraction of a frame does not exist, and a question expects the integers with the division shown. Every allocation must also be at least the minimum of the first section and at most the number of pages the process has: a process of 10 pages is given 10 frames at most, however large its share comes out.

munotes.in361

Allocation of Frames

The proportion changes when the degree of multiprogramming changes. A new process arriving takes frames from everybody, and a process leaving gives frames back, so the allocation is recomputed rather than fixed for ever.

Priority allocation

The same arithmetic with a different weight: allocate in proportion to priority, or to a combination of size and priority, so that an important process gets more memory and therefore faults less. A high priority process that is starved of frames is not high priority in any way that matters.

Global against local replacement

The second question, and the one with the interesting answer.

Local replacementGlobal replacement
The victim may beonly a page of the faulting processany page, including another process's
A process's allocationfixedgrows and shrinks as the machine runs
A process's fault rate depends onits own behaviour onlyother processes' behaviour too
Run time of the same program on the same datarepeatablevaries from run to run
Throughput of the machinelower: a process cannot use frames another one is wastinghigher, and this is why it is the usual choice

Global replacement is the common one, and its defect is that a program's performance is no longer its own property. The same program with the same input can take twice as long because of what else was running, which makes it very hard to reason about. Local replacement gives each process a private, predictable world and wastes memory doing it.

A compromise used in practice: global replacement with a reserved minimum for each process, so no process can be squeezed below the number of frames it needs to make progress. That is the arrangement that leads directly to the next chapter, because a process squeezed too far does not merely go slowly, it stops doing useful work at all.

Distinctions that carry marks

Equal allocationProportional allocation
Rulem / n eachsize divided by total size, times m
62 frames, sizes 10 and 12731 and 314 and 57
Fault rate of the small processvery low, and frames are wastedfair
Fault rate of the large processvery highfair
Changes when a process arrivesonly the divisorthe whole calculation
Minimum framesMaximum frames
Set bythe instruction setthe memory fitted
What happens below itan instruction can never complete
Typical value2 or 3thousands
munotes.in362

Allocation of Frames

What it does not mean

The minimum is not one frame. One frame cannot hold an instruction and its operand.

Proportional allocation is not by need. It is by size, which is a cheap guess at need. A large process using a small part of itself gets more than it can use.

Global replacement does not mean processes are unprotected. A process still cannot read another's pages; the operating system moves a frame from one to the other and fixes both page tables.

A fixed allocation is not the same as local replacement. Local replacement makes an allocation fixed in effect, but a system can also recompute local allocations as processes come and go.

More frames for one process is not free. Every frame given to one is taken from another, which is why the next chapter is about the point where the machine as a whole collapses.

Quick revision

  • The maximum is the frames in the machine; the minimum is set by the

instruction set: enough for every page one instruction can touch, typically 2 or 3, more with indirect addressing.

  • Below the minimum an instruction can never complete.
  • Equal allocation: m / n each, remainder to the free pool. 64 frames and 5 processes give

12 each with 4 spare. A 10 page process and a 127 page process get the same.

  • Proportional allocation: size over total size, times m. 62 frames with sizes 10 and 127

give 10 / 137 × 62, about 4.53, so 4 and 127 / 137 × 62, about 57.47, so 57, with 1 left in the pool.

  • Every allocation must be at least the minimum and at most the process's page count, and

it is recomputed when the degree of multiprogramming changes.

  • Priority allocation weights by priority instead of, or as well as, size.
  • Local replacement: only your own pages, repeatable performance, memory wasted.

Global: anybody's page, higher throughput, and a program's speed depends on what else is running. Global with a reserved minimum is the practical compromise.

Test yourself

  1. What sets the minimum number of frames a process needs? The instruction set: a process

must be able to hold every page a single instruction can touch, or that instruction can never finish.

  1. Give an instruction that needs three frames. A memory to memory move needs the

instruction, the source and the destination; so does one level of indirect addressing.

  1. Divide 64 frames equally among 5 processes. Twelve each, with four left in the free frame

pool.

  1. What is wrong with equal allocation? A small process gets frames it cannot use while a

large process is starved, because size is ignored.

  1. Allocate 62 frames proportionally between processes of 10 and 127 pages. The total is 137,
munotes.in363

Allocation of Frames

so 10 / 137 × 62, about 4.53, which rounds down to 4 frames, and 127 / 137 × 62, about 57.47, which rounds down to 57, leaving one frame in the pool.

  1. What two bounds must an allocation respect? Not below the minimum set by the instruction

set, and not above the number of pages the process actually has.

  1. Distinguish global from local replacement, with one advantage of each. Global lets a

faulting process take any frame, which uses memory better and raises throughput; local restricts it to its own pages, which makes a program's performance repeatable and independent of what else is running.

  1. Why is a program's run time unpredictable under global replacement? Because its fault rate

depends on how many frames other processes take from it, so the same program on the same data can take different times on different runs.

Contents This chapter on its own page

munotes.in364

Chapter Ninety-One

Thrashing, and the Working Set

Syllabus topic Module 2, "Virtual Memory Management - Thrashing"

In one line

A process is thrashing when it spends more time waiting for pages than executing, and a machine thrashes when the total memory the processes are actively using is more than the memory there is.

The definition

Thrashing is high paging activity with no work getting done: a process spends more time paging than executing.

Remember Chapter eighty two's ratio. A fault costs 80,000 memory accesses. A process that faults every few hundred instructions is doing almost nothing but waiting for the disk, and the disk is the slowest thing in the machine.

How a machine gets there, and why it gets worse

The feedback loop, and this is the part a question wants in order.

StepWhat happens
1a process is given too few frames for the part of itself it is using
2it faults, and faults again immediately, because every page it needs has just been replaced
3it spends its time in the wait for disk queue, so it uses almost no processor time
4the processor is idle, so CPU utilisation is low
5the long term scheduler sees low utilisation and starts more processes to fill the machine
6the new processes need frames, which are taken from the processes already there
7more processes are now short of frames, so go to step 2

The scheduler makes it worse, and that is the whole point of the chapter. Everything the operating system is designed to do about an idle processor is exactly the wrong thing here. The machine collapses, and it collapses quickly.

The graph

The shape to be able to draw: CPU utilisation against the degree of multiprogramming.

RegionWhat the curve doesWhy
few processesrises steeplyeach new process gives the processor more to do while others wait for input and output
the peakflattensthe processor is nearly fully used
past the peakfalls off sharplythe processes between them need more memory than there is, and all of them start paging

The fall is not gentle. Past the point where the memory is exhausted, adding one more process can halve the work the machine gets done.

Global replacement spreads it

Under global replacement, a thrashing process takes frames from processes that were not thrashing, and they begin to thrash. Under local replacement a thrashing process can only take its own pages, so it suffers alone; but it still fills the disk queue, and everybody else waits longer for their own faults. Local replacement contains thrashing, it does not cure it.

The locality model

A program does not reference its pages at random: at any moment it is working inside a locality, a small set of pages it uses together, and from time to time it moves to another one.

munotes.in365

Thrashing, and the Working Set

Where a locality comes fromExample
a procedure being executedits code page, its local variables, a few globals
a loop over a data structurethe loop's code and the part of the structure it is walking
a phase of the programreading input, then sorting, then writing output: three different sets of pages

So the localities are what the memory has to hold. If every process can have the pages of its current locality, faults happen only when localities change. If it cannot, it thrashes. The whole of demand paging works because localities exist; this chapter is what happens when they do not fit.

The working set model

The model that turns locality into a number the operating system can measure.

SymbolNameWhat it is
the Greek letter deltathe working set windowa number of references, for example the last 10,000
WSS for a processits working set sizehow many distinct pages it referenced in the last window
Dthe total demandthe sum of every process's WSS

The rule: if D is greater than m, the number of frames in the machine, the system will thrash.

Worked

A window of 10 references on this string, with the references numbered from one:

Working set: window 10

Reference string: 1 2 3 4 1 2 3 4 5 6 7 8 7 8 7 8 9 1 9 1 9 1 2 3

At referenceThe last 10 references areDistinct pagesWSS
101 2 3 4 1 2 3 4 5 61, 2, 3, 4, 5, 66
123 4 1 2 3 4 5 6 7 81, 2, 3, 4, 5, 6, 7, 88
163 4 5 6 7 8 7 8 7 83, 4, 5, 6, 7, 86
207 8 7 8 7 8 9 1 9 11, 7, 8, 94
247 8 9 1 9 1 9 1 2 31, 2, 3, 7, 8, 96

Read the working set shrinking from 8 to 4 between references 12 and 20. The program has settled into a tight loop over pages 7, 8, 9 and 1, and needs four frames instead of eight. That is a locality change, measured.

The demand, and what the system does about it

Three processes with working sets of 8, 4 and 5, on a machine with 12 frames.

D = 8 + 4 + 5 = 17

17 is greater than 12, so the system will thrash. The operating system's answer is to suspend one process: swap it out entirely, give its frames to the others, and let it run later. Suspending the process with the working set of 5 leaves:

munotes.in366

Thrashing, and the Working Set

D = 8 + 4 = 12

Twelve, which the machine has. Two processes now run at full speed instead of three running at none, and the third is restarted when a working set shrinks or a process finishes.

Choosing the window

The window is the whole difficulty of the model, and a question asks about it.

The window isWhat goes wrong
too smallit misses part of the current locality, so WSS is under-estimated and the process is starved
too largeit spans several localities, so WSS is over-estimated and frames are wasted
infinitethe working set is the whole program, which is where the module started

How it is implemented

The model needs the distinct pages of the last ten thousand references, which nothing counts exactly. It is approximated with the reference bit of Chapter eighty eight: a timer interrupts every 5,000 references, the reference bits are copied into two history bits per page and cleared, and a page counts as in the working set if any of its history bits is set. More bits and more frequent interrupts give a better approximation and cost more, which is the same trade as the whole chapter.

The other way out: page fault frequency

The direct method, and the neater answer to "how does a system control thrashing".

StepWhat happens
1set an upper and a lower bound on the acceptable fault rate of a process
2measure each process's actual fault rate
3above the upper bound: give it another frame
4below the lower bound: take a frame away
5above the upper bound with no free frame to give: suspend a process and hand its frames out

It measures the thing that matters, the fault rate itself, instead of estimating the working set and inferring the fault rate from it. That is why it is simpler, and it needs no window at all.

What the lab machine can and cannot show

This machine cannot be made to thrash, and the book says so rather than pretending. Chapter eighty one measured SwapTotal as 0 kilobytes: with no swap space there is nowhere for a page of data to be written, so the kernel cannot page anonymous memory out at all. A program that asks for more than there is does not thrash here, it is killed.

What the earlier chapters did measure stands: faults counted by the kernel one per page, the 8 kilobytes two bytes cost, and the copy on write accounting. Thrashing needs a machine with a disk to thrash onto, and the arithmetic above is how it is reasoned about on one.

munotes.in367

Thrashing, and the Working Set

Distinctions that carry marks

PagingThrashing
What it isthe normal way memory is usedpaging so heavily that no work is done
Fault rateabout one in a million referenceshigh enough that the process is always waiting
Curenone neededreduce the degree of multiprogramming
Working set modelPage fault frequency
What is measuredthe distinct pages in a window of referencesthe fault rate of each process
Needsa window, and an approximation with reference bitsa counter and two bounds
Acts bysuspending a process when D is greater than mgiving or taking single frames, and suspending as a last resort

What it does not mean

Thrashing is not a lot of paging. It is paging that has crowded out the computing.

It is not caused by a bad replacement algorithm. A better algorithm helps a little; the cause is not enough frames for the localities, and the cure is fewer processes.

The working set is not the process's size. It is the pages used in the recent window, which is usually a small part of it.

D greater than m is not a prediction about one process. It is a statement about the machine: some process must be suspended.

Suspending a process is not killing it. It is swapped out and resumed later, which is Chapter sixty nine's medium term scheduler doing exactly the job it exists for.

Quick revision

  • Thrashing: a process spends more time paging than executing; the machine does almost no

work.

  • The loop: too few frames, constant faults, idle processor,

the scheduler starts more processes, fewer frames each, worse. The scheduler's normal response is the wrong one.

  • The graph of CPU utilisation against degree of multiprogramming rises, peaks, and

falls off sharply.

  • Global replacement spreads thrashing between processes; local replacement contains

it but does not cure it, because the disk queue is shared.

  • Locality model: a program works in a small set of pages at a time and moves between such

sets.

  • Working set: with a window of the last delta references, WSS is the number of

distinct pages in it and D is the sum over all processes. If D is greater than m, the system will thrash, and a process is suspended.

  • Worked: window 10, WSS 8 at reference 12 and 4 at reference 20. With working sets 8, 4

and 5 and 12 frames, D = 17, so one process is suspended and D = 12.

  • The window: too small starves, too large wastes, infinite is the whole program. It
munotes.in368

Thrashing, and the Working Set

is approximated with the reference bit and a timer.

  • Page fault frequency: bounds on the fault rate; above the upper bound give a frame, below

the lower take one, and suspend a process if there is no frame to give.

  • The lab machine has no swap, so it cannot thrash: it kills the program instead.

Test yourself

  1. Define thrashing. A process is thrashing when it spends more time paging than executing; a

system is thrashing when this is true of enough processes that the machine gets almost no work done.

  1. Why does low CPU utilisation make thrashing worse? Because the long term scheduler

responds to an idle processor by starting more processes, which takes frames from the processes that are already short of them.

  1. Draw and describe the CPU utilisation graph. Against the degree of multiprogramming,

utilisation rises steeply, flattens at a peak, and then falls off sharply once the processes need more memory than exists.

  1. Does local replacement cure thrashing? No. It stops one process stealing another's frames,

but the thrashing process still fills the disk queue and slows everyone's faults.

  1. What is the working set of a process? The set of distinct pages it has referenced in the

last delta references, and its size is WSS. 6. For the string 1 2 3 4 1 2 3 4 5 6 7 8 7 8 7 8 9 1 9 1 9 1 2 3 with a window of 10, give WSS at references 12 and 20. At 12 the window is 3 4 1 2 3 4 5 6 7 8, so WSS is 8; at 20 it is 7 8 7 8 7 8 9 1 9 1, so WSS is 4.

  1. Three processes have working sets of 8, 4 and 5 on a machine with 12 frames. What happens?

D = 17, which is more than 12, so the system would thrash; the operating system suspends one process, and suspending the one with a working set of 5 brings D to 12.

  1. What goes wrong if the working set window is too large? It covers more than one locality,

so the working set is over-estimated and frames are held that the process is not using.

  1. Describe page fault frequency control. Keep upper and lower bounds on a process's fault

rate: above the upper bound give it a frame, below the lower bound take one away, and if there is no frame to give, suspend a process and share out its frames.

Contents This chapter on its own page

munotes.in369

Chapter Ninety-Two

What a Disk Is, and What It Costs

Syllabus topic Module 2, "Mass-Storage Structure - Overview of Mass-Storage Structure"

In one line

A disk access is a mechanical movement, a wait for the platter to come round, and only then a transfer, which is why it is measured in milliseconds while memory is measured in nanoseconds.

The parts

The vocabulary, and a question can ask for the diagram.

PartWhat it is
plattera flat circular plate with a magnetic surface, on both sides
trackone circle of the surface, at one distance from the centre
sectora piece of a track: the smallest unit that can be read or written, usually 512 bytes
cylinderthe set of tracks at the same distance from the centre on every platter
read write headflies just above one surface; there is one per surface
arm assemblyholds every head and moves them together, so all the heads are on the same cylinder at once
spindleturns every platter together, at a constant speed

The heads move together, which is why the cylinder is the unit that matters. Reading two tracks of the same cylinder costs no arm movement; reading two tracks of different cylinders costs a seek.

QuantityTypical value
rotation speed5,400 to 15,000 revolutions per minute
sector size512 bytes, and 4,096 on newer disks
transfer ratetens to hundreds of megabytes a second
seek timea few milliseconds

The three times

One access is three things, and the marks are for naming all three in order.

NameWhat is happeningOrder of size
seek timethe arm moves to the right cylindermilliseconds: the largest part
rotational latencythe machine waits for the right sector to come under the headmilliseconds
transfer timethe bytes go bymicroseconds for one block

access time = seek time + rotational latency + transfer time

The first two together are called the positioning time, or the random access time. They are where a disk's reputation comes from: nothing in them depends on how much data you want, only on where it is.

Rotational latency is half a rotation

The sum that is always available in a question, because it needs nothing but the rotation speed.

one rotation in milliseconds = 60,000 / revolutions per minute

The average latency is half a rotation, because on average the sector wanted is half a turn away.

SpeedOne rotationAverage latency
5,400 rpmabout 11.11 msabout 5.56 ms
7,200 rpmabout 8.33 msabout 4.17 ms
10,000 rpm6 ms3 ms
15,000 rpm4 ms2 ms

Half a rotation, not a whole one. The worst case is a whole rotation, when the sector has just gone past, and the best case is none at all. The average is what a question means unless it says otherwise.

munotes.in370

What a Disk Is, and What It Costs

What one block costs, worked

A 7,200 rpm disk with a 5 millisecond average seek and a 100 megabyte a second transfer rate, reading one 4 kilobyte block.

The transfer, exactly, in milliseconds:

4 / 102,400 × 1000 = 0.0390625

PartTime in milliseconds
seek5
rotational latencyabout 4.17
transferabout 0.04
totalabout 9.21

About nine milliseconds, of which the transfer is less than half a per cent. That is where the eight milliseconds of Chapter eighty two came from, and it is why the fault service time there was a disk read and nothing else.

Why scattered is so much worse than sequential

The sum that motivates the next three chapters and Chapter one hundred seven's allocation methods. Read 100 blocks of 4 kilobytes from the same disk, two ways.

Scattered, a seek and a latency for every block:

WorkingTotal
100 separate accesses100 × about 9.21about 921 ms

Sequential, one seek and one latency, then 100 transfers:

transfers = 100 × 0.0390625 = 3.90625

PartTime in milliseconds
seek5
rotational latencyabout 4.17
100 transfersabout 3.91
totalabout 13.07

Nine hundred and twenty one milliseconds against thirteen: the same 400 kilobytes, seventy times the time. Every disk decision in the rest of this module is an attempt to make an access look more like the second case than the first.

Solid state storage

MU's overview names it, and one paragraph is what a question wants.

Magnetic diskSolid state drive
Moving partsyes: the arm and the spindlenone
Seek timemillisecondsnone
Rotational latencymillisecondsnone
Access timeabout 9 ms as abovetens of microseconds
Cost per bytelowhigher
Wearthe surface lastseach cell takes a limited number of writes

The consequence for this module is large: disk scheduling, which the next two chapters are about, is worth much less on a drive with no arm to move. The algorithms are still asked for, and they are still what a magnetic disk needs, but a modern machine gains far less from them.

What the lab machine says about its own disk

The kernel publishes what it has been told about the device, and it is worth seeing because of what it gets wrong.

$ awk '$4 ~ /^vda$/ {print $4, $3 " blocks of 1 kilobyte"}' /proc/partitions
vda 20971520 blocks of 1 kilobyte
$ cat /sys/block/vda/queue/rotational
1
$ cat /sys/block/vda/queue/logical_block_size
512

The flag says 1, meaning a rotating disk, and nothing in this machine rotates. The storage is a file on a host with solid state storage, presented to the container through a virtual disk driver, and the driver reports the default.

munotes.in371

What a Disk Is, and What It Costs

Two lessons, and the second is the useful one. The 512 is real: that is the sector size the rest of this module assumes. And the 1 is a reminder that a flag is a claim by a driver, not a measurement: the scheduling decisions the kernel makes on the strength of that flag are being made about a disk that does not exist.

Distinctions that carry marks

Seek timeRotational latencyTransfer time
What movesthe arm, across cylindersthe platter, under the headnothing: the bytes go by
Depends onhow far the arm must gothe rotation speedthe amount of data and the rate
Typical, 7,200 rpma few msabout 4.17 ms0.04 ms for 4 kilobytes
Reduced bydisk schedulinga faster spindlea faster interface
SectorTrackCylinder
What it isthe smallest readable piece, 512 bytesone circle on one surfacethe same track on every surface
Reading two of these costsa latencya latencyno seek within one cylinder

What it does not mean

The transfer rate is not the access time. A disk quoting 100 megabytes a second still takes nine milliseconds to give you the first byte.

Average latency is not the rotation time. It is half of it.

A sector is not a block. A block is what the file system chooses to use, usually several sectors, which is the next chapter.

Cylinders are not physical objects. A cylinder is the set of tracks the heads can reach without moving.

The rotational flag is not a measurement. The lab machine reports a rotating disk and has none.

Quick revision

  • Platter, track, sector, cylinder, head, arm assembly, spindle. The heads move together,

so a cylinder is free of seeks.

  • A sector is the smallest unit read or written, usually 512 bytes.
  • Access time = seek time + rotational latency + transfer time, and the first two are the

positioning time.

  • One rotation = 60,000 / rpm milliseconds; the average latency is half of it. At

7,200 rpm the rotation is about 8.33 ms and the latency about 4.17 ms.

  • A 4 kilobyte block on that disk with a 5 ms seek and 100 megabytes a second: transfer

4 / 102,400 × 1000 = 0.0390625 ms, total about 9.21 ms, of which the transfer is under half a per cent.

  • 100 scattered blocks cost about 921 ms; the same 100 in sequence cost about 13.07 ms.

Seventy times, for the same data.

  • A solid state drive has no seek and no latency, so access is tens of microseconds and

disk scheduling is worth far less on one.

munotes.in372

What a Disk Is, and What It Costs

  • The lab machine's virtual disk reports a 512 byte sector, which is real, and

rotational 1, which is not.

Test yourself

  1. Name the seven parts of a disk. Platter, track, sector, cylinder, read write head, arm

assembly and spindle.

  1. Why does the cylinder matter more than the track? Because all the heads move together, so

every track of one cylinder can be read without moving the arm.

  1. Give the three components of an access time and say which is largest. Seek time,

rotational latency and transfer time; the seek is usually the largest, and the transfer of one block is the smallest by far.

  1. Find the average rotational latency of a 7,200 rpm disk. One rotation is 60,000 / 7,200,

about 8.33 milliseconds, and the average latency is half of it, about 4.17 milliseconds.

  1. Why half a rotation? Because on average the sector wanted is half a turn away: the worst

case is a full rotation and the best is none. 6. How long does one 4 kilobyte block take on a 7,200 rpm disk with a 5 ms seek and 100 megabytes a second? About 9.21 milliseconds: 5 for the seek, about 4.17 of latency, and 4 / 102,400 × 1000 = 0.0390625 for the transfer.

  1. Compare 100 scattered blocks with 100 sequential ones on that disk. Scattered costs a seek

and a latency each, about 921 milliseconds in all; sequential costs one seek, one latency and 100 transfers, about 13.07 milliseconds.

  1. Why is disk scheduling less valuable on a solid state drive? It exists to reduce arm

movement, and a solid state drive has no arm: its access time is tens of microseconds with no seek or rotational component.

Contents This chapter on its own page

munotes.in373

Chapter Ninety-Three

Disk Structure, and the Logical Block

Syllabus topic Module 2, "Mass-Storage Structure - Disk Structure, Disk Attachment"

In one line

The operating system does not address cylinders and heads: it treats the disk as one long array of numbered blocks, and the disk itself hides the geometry.

The array of logical blocks

A disk is presented as a one dimensional array of logical blocks, numbered 0 upwards, and the logical block is the smallest unit the operating system transfers.

SectorLogical block
Whose unitthe disk's, in hardwarethe operating system's
Size512 bytes usuallyone or more sectors: 512, 1024, 4096 bytes
Numberedby cylinder, track and position0 to n minus 1, one long line

Everything above this chapter uses the block number and nothing else. That is the point of the abstraction: the file system of Chapter ninety eight asks for block 91,482 and neither knows nor cares where it is.

The mapping

The order in which the blocks are laid out, and a question asks for it in exactly this order.

LevelRule
1sector by sector along a track
2then track by track within the same cylinder, because those need no arm movement
3then cylinder by cylinder, from the outermost inwards

Track before cylinder, and that is the whole reason the order is worth learning. Block n and block n + 1 are almost always in the same cylinder, so reading in block order costs almost no seeking. It is why Chapter ninety two's sequential read was seventy times faster than the scattered one, and it is why Chapter one hundred seven's contiguous allocation works.

Converting a block number to a place, worked

A disk with 4 surfaces and 100 sectors per track. So one cylinder holds:

blocks per cylinder = 4 × 100 = 400

Where is logical block 1234?

1234 = 3 × 400 + 34

cylinder = 3

34 = 0 × 100 + 34

surface = 0

sector = 34

Cylinder 3, surface 0, sector 34.

And back again

Which block is cylinder 5, surface 2, sector 17?

block = 5 × 400 + 2 × 100 + 17 = 2000 + 200 + 17 = 2217

The method both ways in one line: divide by the blocks per cylinder, then by the sectors per track, and multiply in the same order to go back.

The capacity sum

The other standard sum, and it needs four numbers.

capacity = cylinders × surfaces × sectors per track × bytes per sector

Worked on the geometry that old machines were limited to, 16,383 cylinders, 16 heads, 63 sectors of 512 bytes:

16,383 × 16 × 63 × 512 = 8,455,200,768

8,455,200,768 bytes, which is the eight and a half gigabyte limit that once stopped disks from growing.

munotes.in374

Disk Structure, and the Logical Block

Why the sums are a simplification

Real disks do not have a constant number of sectors per track, and a question that gives you one is giving you a model.

The recording schemeWhat it doesConsequence
constant angular velocity, the old waythe platter turns at a fixed speed and every track holds the same number of sectorsthe outer tracks waste space, because their sectors are longer than they need to be
zoned recording, the modern wayouter tracks hold more sectors than inner onesmore capacity, but sectors per track is not one number
constant linear velocity, on optical discsthe sectors are the same length and the rotation speed changes with the trackmore capacity again, and a variable spindle speed

Two more reasons the mapping is not exact: bad sectors are remapped to spares elsewhere on the disk, so a block's neighbour may be far away, and the outermost tracks are sometimes reserved. The operating system cannot see any of it, which is the abstraction working as intended.

Disk attachment

MU's section names three ways a disk reaches a machine, and a question wants the distinction.

Host attachedNetwork attached, NASStorage area network, SAN
How it is reacheda local bus: SATA, SAS, USBover the local network, by a file protocol such as NFSa separate network dedicated to storage
What the client asks fora blocka filea block
Who runs the file systemthis machinethe serverthis machine
Used forone machine's own diskssharing files between machinesmany servers sharing a pool of storage

The distinction that carries the marks: NAS serves files and a SAN serves blocks. A machine using a SAN thinks it has its own disk; a machine using NAS knows it is talking to a file server.

Measured on the lab machine

The kernel publishes the block count of a device, which is all that is needed to work out its capacity.

$ cat /sys/block/vda/size
41943040
$ cat /sys/block/vda/queue/hw_sector_size
512
$ echo "capacity = $(( 41943040 * 512 )) bytes"
capacity = 21474836480 bytes
$ echo "that is $(( 41943040 * 512 / 1024 / 1024 / 1024 )) gigabytes"
that is 20 gigabytes

Forty one million sectors of 512 bytes: exactly twenty gigabytes. And the partition inside it is a range of block numbers and nothing more:

$ echo "starts at sector $(cat /sys/block/vda/vda1/start), and is $(cat /sys/block/vda/vda1/size) sectors long"
starts at sector 2099200, and is 39843807 sectors long

A partition is arithmetic on block numbers. Nothing about the geometry appears anywhere, and Chapter ninety two showed that the machine's claim to be a rotating disk is not even true.

munotes.in375

Disk Structure, and the Logical Block

Distinctions that carry marks

Physical addressLogical block address
Formcylinder, surface, sectorone number
Who uses itthe disk's own controllerthe operating system
Changes if a sector goes badyes, the spare is elsewhereno: the number keeps working
Blocks per cylinderSectors per track
Formulasurfaces × sectors per trackgiven, and on a real disk not constant
Used to findthe cylinder of a blockthe surface and sector within it

What it does not mean

A logical block is not a sector. It is one or more sectors, chosen by the file system.

Block order is not cylinder order. Blocks fill a whole cylinder, across all its surfaces, before moving to the next cylinder.

The geometry in a question is not a real disk's geometry. Zoned recording means sectors per track varies, and remapped bad sectors break the neighbour relation.

NAS is not a SAN. One serves files over the ordinary network; the other serves blocks over a network of its own.

A partition is not a physical division. It is a first block and a length.

Quick revision

  • A disk is an array of logical blocks numbered from 0, and the block is the unit the

operating system transfers.

  • The layout order:

sector by sector, then track by track inside the cylinder, then cylinder by cylinder outwards in. Track before cylinder, which is why sequential block numbers cost no seeking.

  • Blocks per cylinder = surfaces × sectors per track. With 4 surfaces and 100 sectors, that

is 400.

  • Block 1234 on that disk: 1234 = 3 × 400 + 34, so cylinder 3, surface 0, sector

34.

  • Cylinder 5, surface 2, sector 17: 5 × 400 + 2 × 100 + 17 = 2217.
  • Capacity = cylinders × surfaces × sectors per track × bytes per sector.

16,383 × 16 × 63 × 512 = 8,455,200,768 bytes.

  • Real disks use zoned recording, so sectors per track varies; bad sectors are remapped;

the sums are a model.

  • Host attached by SATA or SAS; NAS serves files over the network; a SAN serves

blocks over its own network.

  • Measured: the lab disk is 41,943,040 sectors of 512 bytes, exactly 20 gigabytes,

and a partition is a start sector and a length.

Test yourself

  1. What does the operating system use to address a disk? A logical block number, from 0 to

one less than the number of blocks; the geometry is hidden by the disk.

  1. Give the order in which logical blocks are laid out. Sector by sector along a track, then

track by track within the same cylinder, then cylinder by cylinder from the outside inwards.

munotes.in376

Disk Structure, and the Logical Block

  1. Why is track before cylinder the right order? Because the heads move together, so the

tracks of one cylinder need no arm movement: consecutive block numbers are cheap to read.

  1. A disk has 4 surfaces and 100 sectors per track. Where is logical block 1234? 400 blocks

per cylinder, so 1234 = 3 × 400 + 34: cylinder 3, surface 0, sector 34.

  1. Which block is cylinder 5, surface 2, sector 17 on that disk?

5 × 400 + 2 × 100 + 17 = 2217.

  1. Find the capacity of a disk of 16,383 cylinders, 16 heads and 63 sectors of 512 bytes.

16,383 × 16 × 63 × 512 = 8,455,200,768 bytes.

  1. Why is a constant sectors per track a simplification? Modern disks use zoned recording,

where outer tracks hold more sectors than inner ones, and bad sectors are remapped elsewhere.

  1. Distinguish NAS from a SAN. Network attached storage serves whole files over the ordinary

network and runs the file system itself; a storage area network serves blocks over a dedicated network, and the client runs the file system.

Contents This chapter on its own page

munotes.in377

Chapter Ninety-Four

Disk Scheduling: FCFS and SSTF

Syllabus topic Module 2, "Mass-Storage Structure - Disk Scheduling: FCFS, SSTF"

In one line

Several requests are waiting, the arm can serve them in any order, and the order decides how far the arm travels.

What is being minimised, and why it can be

Chapter ninety two's arithmetic said the seek is the expensive part of an access. A busy disk has a queue of requests for different cylinders, so the operating system may choose the order, and the total head movement is what it is choosing.

GivenThe standard question
the current head positionwhere the arm starts
the queue of cylinder numberswhat is waiting
the range of cylinders, and sometimes a directionwhat the arm may do
asked forthe order of service and the total head movement

Total head movement is counted in cylinders, and it is a distance, so it is always positive. Add the absolute difference at every hop. A student who subtracts and keeps the sign gets a smaller and wrong answer.

The queue used in this chapter and the next is the standard one, and it is worth knowing by heart because every textbook uses it: head at 53, cylinders 0 to 199, and the queue

98 183 37 122 14 124 65 67

FCFS

First come, first served: serve the queue in the order it arrived. It is not really a scheduling algorithm at all, which is the point of starting with it.

Disk: FCFS, head 53, cylinders 0 to 199, moving up

Request queue: 98 183 37 122 14 124 65 67

StepFromToMovement
1539845
29818385
318337146
43712285
512214108
614124110
71246559
865672

Order of service: 98 183 37 122 14 124 65 67

Total head movement: 640

640 cylinders of movement. Look at steps 2, 3 and 5: 183 down to 37, and 122 down to 14. The arm crosses most of the disk, comes back, and crosses it again, because 183 and 37 happened to be asked for in that order.

FCFS
Fairevery request is served in turn, so no request can be starved
Simplea queue and nothing else
Wastefulthe arm swings across the disk whenever the queue happens to alternate
Used whenthe queue is usually one request long, where there is nothing to schedule

SSTF

Shortest seek time first: serve the waiting request whose cylinder is nearest the head, then repeat.

In practice "shortest seek time" means "smallest difference in cylinder number", because that is what the operating system can see. A question means exactly that.

Disk: SSTF, head 53, cylinders 0 to 199, moving up

munotes.in378

Disk Scheduling: FCFS and SSTF

Request queue: 98 183 37 122 14 124 65 67

StepFromToMovement
1536512
265672
3673730
4371423
5149884
69812224
71221242
812418359

Order of service: 65 67 37 14 98 122 124 183

Total head movement: 236

236 cylinders, against FCFS's 640: less than two fifths of the movement on the same queue.

Follow the first four hops, because they are the whole algorithm. From 53, the nearest waiting cylinder is 65, then 67; from 67 the choice is between 98 above and 37 below, and 37 is nearer by one hop of 30 against 31; from 37 the nearest is 14; and only then does the arm turn round and climb through 98, 122, 124 to 183.

SSTF is not optimal

Shortest is not best, and this is the trap. A greedy choice at every step does not give the shortest total. On this queue a better order exists:

Order: 37, 14, 65, 67, 98, 122, 124, 183

total = 16 + 23 + 169 = 208

The orderTotal movement
SSTF236
the order above208

It goes down first, to 37 and 14, and then climbs once to 183 without turning round again. SSTF went down to 14 as well but reached it by way of 65 and 67, which it then had to cross again.

Say this in an answer: SSTF gives a substantial improvement over FCFS, but it is a greedy algorithm and not optimal. 208 is in fact the optimum for this queue, which Chapter ninety six proves by working out all 40,320 possible orders; finding it is not worth doing for a queue that changes every few milliseconds.

SSTF can starve a request

The serious objection. A request for a distant cylinder can wait for ever if requests near the head keep arriving. The queue is not static: new requests join while the arm works, and SSTF always prefers the newcomer that is closer.

FCFSSSTF
Total movement on the standard queue640236
Starvationimpossiblepossible
Choice depends onarrival orderthe head position
Optimalnono, and it looks as though it should be
Cost to decidenonea scan of the queue per request

Distinctions that carry marks

Seek timeTotal head movement
Unitsmillisecondscylinders
What a question asks foroccasionally, with a time per cylinderusually this
How to get one from the othermultiply by the time per cylinder, or use a seek time formula if one is given
munotes.in379

Disk Scheduling: FCFS and SSTF

What it does not mean

SSTF does not minimise total movement. It minimises the next hop, which is not the same thing, as the order above shows.

FCFS is not always wrong. With one request in the queue every algorithm is FCFS, and on a lightly loaded disk that is the normal case.

Starvation is not slowness. A starved request is never served at all while the stream of nearer requests continues.

Head movement is not measured in time unless the question gives a rate. Answer in cylinders unless told otherwise.

The head's own position is not a request. The arm starts there; it is not counted as served, and no movement is counted to reach it.

Quick revision

  • Disk scheduling chooses the order of service to reduce total head movement, counted in

cylinders as a sum of absolute differences.

  • The standard problem: head 53, cylinders 0 to 199, queue

98 183 37 122 14 124 65 67.

  • FCFS serves the arrival order: 640 cylinders. Fair, simple, no starvation, and it

swings across the disk.

  • SSTF serves the nearest waiting cylinder each time: 236 cylinders, order

65 67 37 14 98 122 124 183.

  • SSTF is greedy, not optimal: the order 37 14 65 67 98 122 124 183 costs

16 + 23 + 169 = 208.

  • SSTF can starve a distant request while nearer ones keep arriving; FCFS cannot.

Test yourself

  1. What quantity does disk scheduling minimise, and in what units is it answered? The total

head movement, in cylinders, as the sum of the absolute differences between successive positions.

  1. Work FCFS for head 53 and the queue 98 183 37 122 14 124 65 67. The order is the queue's

own, and the total is 45 + 85 + 146 + 85 + 108 + 110 + 59 + 2 = 640 cylinders.

  1. Give one advantage and one disadvantage of FCFS. It is fair, so no request starves; the

arm swings wildly whenever successive requests are far apart.

  1. Work SSTF on the same problem. 65, 67, 37, 14, 98, 122, 124, 183, and the total is 236

cylinders.

  1. At cylinder 67 with 98 and 37 both waiting, which does SSTF choose and why? 37: it is 30

cylinders away against 31 for 98.

  1. Show that SSTF is not optimal on this queue. The order 37, 14, 65, 67, 98, 122, 124, 183

costs 16 + 23 + 169 = 208 cylinders, less than SSTF's 236.

  1. Why can SSTF starve a request? Because the queue keeps changing: if requests near the head

keep arriving, the distant one is never the nearest and is never chosen.

  1. When does the choice of algorithm not matter? When the queue holds one request, which is
munotes.in380

Disk Scheduling: FCFS and SSTF

the usual state of a lightly loaded disk.

Contents This chapter on its own page

munotes.in381

Chapter Ninety-Five

SCAN, C-SCAN, LOOK and C-LOOK

Syllabus topic Module 2, "Mass-Storage Structure - Disk Scheduling: SCAN, C-SCAN, LOOK, C-LOOK"

In one line

Sweep the arm steadily from one end towards the other, serving what it passes, and the only questions are where to turn round and whether to serve on the way back.

The elevator idea

A lift does not go to the nearest waiting passenger. It continues in one direction, stopping for everybody on the way, and then turns round. SCAN is that, applied to the arm, and it is why SCAN is called the elevator algorithm.

The gain over SSTF is not less movement: it is that no request can be starved. The arm is coming, and the only question is how soon.

The four algorithms are two independent choices:

Turn round at the end of the diskTurn round at the last request
Serve on the way backSCANLOOK
Return without servingC-SCANC-LOOK

The C is for circular: the arm goes back to the beginning and starts the same way again, so every cylinder is approached from the same direction. That is the whole reason C-SCAN exists, and it is in the next section.

SCAN

Continue in the current direction to the end of the disk, serving everything passed, then reverse and do the same.

Disk: SCAN, head 53, cylinders 0 to 199, moving up

Request queue: 98 183 37 122 14 124 65 67

StepFromToMovement
1536512
265672
3679831
49812224
51221242
612418359
718319916
819937162
9371423

Order of service: 65 67 98 122 124 183 199 37 14

Total head movement: 331

331 cylinders. The arm goes up from 53 through 65, 67, 98, 122, 124, 183 to 199, where there is nothing waiting, and then turns and comes down to 37 and 14.

Notice step 7: 183 to 199 with no request there. SCAN goes to the end of the disk because it is SCAN, and those 16 cylinders are wasted. That is exactly what LOOK removes.

C-SCAN

Go to the end serving everything, then jump straight back to the other end serving nothing, and sweep the same way again.

Disk: C-SCAN, head 53, cylinders 0 to 199, moving up

Request queue: 98 183 37 122 14 124 65 67

StepFromToMovement
1536512
265672
3679831
49812224
51221242
612418359
718319916
81990199
901414
10143723

Order of service: 65 67 98 122 124 183 199 0 14 37

Total head movement: 382

munotes.in382

SCAN, C-SCAN, LOOK and C-LOOK

382 cylinders, the largest of the four, because the return trip from 199 to 0 is counted and serves nobody.

So why use it? Because the waiting time becomes uniform. Under SCAN a cylinder near the middle is passed twice per sweep and a cylinder at the end once, so the wait depends on where you are. Under C-SCAN every cylinder is visited once per sweep, in the same direction, and the longest possible wait is the same everywhere. C-SCAN trades total movement for fairness, and that sentence is the answer to why it exists.

LOOK

SCAN, but turn round at the last request instead of at the end of the disk. It "looks" to see whether anything is waiting further on, and if not, it reverses.

Disk: LOOK, head 53, cylinders 0 to 199, moving up

Request queue: 98 183 37 122 14 124 65 67

StepFromToMovement
1536512
265672
3679831
49812224
51221242
612418359
718337146
8371423

Order of service: 65 67 98 122 124 183 37 14

Total head movement: 299

299 cylinders, the smallest of the four. It is still more than SSTF's 236, and that is the trade: LOOK saves SCAN's wasted trip to 199 and keeps SCAN's freedom from starvation, which SSTF does not have.

C-LOOK

C-SCAN, but the jump back goes only as far as the lowest request, not to cylinder 0.

Disk: C-LOOK, head 53, cylinders 0 to 199, moving up

Request queue: 98 183 37 122 14 124 65 67

StepFromToMovement
1536512
265672
3679831
49812224
51221242
612418359
718314169
8143723

Order of service: 65 67 98 122 124 183 14 37

Total head movement: 322

322 cylinders. The return jump is from 183 down to 14, the lowest waiting request, rather than to 0.

All four together

AlgorithmTurns round atServes on the way backOrder of serviceTotal
SCANthe end of the diskyes65 67 98 122 124 183 199 37 14331
C-SCANthe end of the diskno65 67 98 122 124 183 199 0 14 37382
LOOKthe last requestyes65 67 98 122 124 183 37 14299
C-LOOKthe last requestno65 67 98 122 124 183 14 37322

Read the two columns on the left and the totals follow from them. Turning at the last request always saves movement; returning without serving always costs it and buys uniform waiting.

munotes.in383

SCAN, C-SCAN, LOOK and C-LOOK

The convention a question must state

The answer changes with the direction, and a question that does not give one is incomplete. Every trace above starts moving up. Starting down, SCAN would serve 37 and 14, go to 0, and then climb to 183: a different order and a different total.

What must be givenWhy
the starting directionup and down give different answers
the range of cylindersSCAN and C-SCAN go to the ends, so 0 to 199 and 0 to 999 differ
whether the end of the disk is reachedthat is SCAN against LOOK
whether the return trip is countedin C-SCAN the jump from 199 to 0 is 199 cylinders of movement and it is counted here

Say your convention in the answer. A line saying that the arm starts by moving up and that the return from 199 to 0 is counted as movement turns a disagreement with the examiner's arithmetic into a difference of stated assumptions.

Distinctions that carry marks

SCANC-SCAN
At the end of the diskreversesjumps to the other end
On the returnserves requestsserves nothing
Waiting timedepends where you areuniform
Total on the standard queue331382
SCANLOOK
Goes to cylinder 0 and to the last cylinderyes, alwaysonly if a request is there
Total on the standard queue331299
Starvationimpossibleimpossible

What it does not mean

SCAN is not SSTF with a direction. SSTF may turn round at any moment; SCAN cannot until it reaches the end.

The elevator algorithm is not about lifts. It is a name for the sweep.

C-SCAN is not faster than SCAN. It moves the arm further, on purpose, to make the waiting even.

LOOK is not always better than SCAN in a question's terms. It moves less, but if the examiner's convention is that the arm reaches the end of the disk, the expected answer is SCAN's.

None of the four is optimal. SSTF moves less than all of them on this queue and starves requests; the optimal order moves less than SSTF.

Quick revision

  • SCAN, the elevator: sweep to the end of the disk, serving everything, then reverse.

Standard queue: 331.

  • C-SCAN: sweep to the end, jump back without serving, sweep the same way. 382, and

the waiting time is uniform.

  • LOOK: SCAN, but turn at the last request. 299, the least of the four.
  • C-LOOK: C-SCAN, but the jump goes back only to the lowest request. 322.
  • The two questions that define all four: turn at the end of the disk or at the last request,
munotes.in384

SCAN, C-SCAN, LOOK and C-LOOK

and serve on the way back or not.

  • None can starve a request, which is their advantage over SSTF.
  • A question must give the starting direction and the cylinder range, and must say

whether the return trip is counted. State the convention in the answer.

Test yourself

  1. Why is SCAN called the elevator algorithm? Because like a lift it continues in one

direction, stopping for everything on the way, and reverses only at the end.

  1. Work SCAN for head 53, cylinders 0 to 199, queue 98 183 37 122 14 124 65 67, moving up.

65, 67, 98, 122, 124, 183, 199, then 37 and 14: 331 cylinders.

  1. What does C-SCAN do differently, and what does it cost and buy? It jumps from the end of

the disk back to the start without serving anything and sweeps the same way again: it costs more movement, 382 against 331, and it buys a uniform waiting time for every cylinder.

  1. Work LOOK on the same problem. 65, 67, 98, 122, 124, 183, then 37 and 14: 299 cylinders,

because the arm never goes to 199.

  1. Work C-LOOK on the same problem. 65, 67, 98, 122, 124, 183, then a jump down to 14 and 37:

322 cylinders.

  1. Fill the two by two table that defines the four algorithms. Turning at the end of the disk

and serving on the way back is SCAN; at the end of the disk without serving is C-SCAN; at the last request serving on the way back is LOOK; at the last request without serving is C-LOOK. 7. Which of the six algorithms so far moves the arm least on this queue, and which cannot starve a request? SSTF moves least at 236; SCAN, C-SCAN, LOOK, C-LOOK and FCFS cannot starve a request.

  1. What must a disk scheduling question state, and what should the answer state? The starting

direction, the cylinder range, and whether the return trip counts as movement; the answer should state the convention it used.

Contents This chapter on its own page

munotes.in385

Chapter Ninety-Six

Random Scheduling, and All Six Compared

Syllabus topic Module 2, MU prints RSS among the disk scheduling algorithms

In one line

Serving the queue in a random order is the baseline every other algorithm has to beat, and on the standard queue it beats FCFS.

What RSS is

Random scheduling, which MU writes as RSS: pick a waiting request at random and serve it, then pick again.

Nobody builds it. It is in the list for two reasons, and both are worth writing in an answer.

Why it is taughtWhat it gives
a baselinean algorithm that does not beat a random order is doing nothing useful
an analysis toolthe average, the best and the worst possible orders bound what any algorithm can do

Every possible order of the standard queue

The queue of Chapter ninety four and Chapter ninety five: head at 53, cylinders 0 to 199, and eight requests. Eight requests can be served in

orders = 8 × 7 × 6 × 5 × 4 × 3 × 2 × 1 = 40,320

40,320 different orders, and sim/disk.py works out the head movement of every one of them.

Over all 40,320 ordersTotal head movement
the best order208
the mean, which is what RSS gives on average504.5
the worst order702

208 is therefore the optimum, not merely a better order than SSTF's. Chapter ninety four showed the order 37, 14, 65, 67, 98, 122, 124, 183 costing 16 + 23 + 169 = 208; the enumeration proves no order does better, because every order was tried.

And the mean is 504.5, while FCFS came to 640. On this queue, serving the requests in the order they arrived is worse than serving them in a random order. FCFS is not a scheduling algorithm and this is the arithmetic that says so.

All seven results

AlgorithmTotalCompared with randomCan a request starve
the best possible order208296 lessnot an algorithm
SSTF236268 lessyes
LOOK299205 lessno
C-LOOK322182 lessno
SCAN331173 lessno
C-SCAN382122 lessno
RSS, on average504.5the baselinein principle no, in practice unbounded
FCFS640135 moreno
the worst possible order702197 morenot an algorithm

Three things to take from the table, and they are the three sentences a comparison question wants.

One: every real algorithm except FCFS beats the random baseline, and by a wide margin: SSTF by more than half.

Two: SSTF is the closest to the optimum and the only one that can starve a request. Everything else in the table is safe from starvation and pays for it in movement.

Three: the range is 208 to 702, a factor of more than three between the best and worst ways of serving the same eight requests. That factor is what disk scheduling is worth on a rotating disk, and Chapter ninety two's note about solid state drives is what has happened to it since.

munotes.in386

Random Scheduling, and All Six Compared

Where RSS sits on fairness

The property a question can ask about, and the answer is a careful one.

FCFSRSSSSTF
Order depends onarrivalchancethe head position
Longest waitbounded: the queue lengthunbounded, though each request is chosen eventually with probability oneunbounded, and a distant request may never be chosen
Fairyes, strictlyon averageno

Unbounded and never are not the same thing. Under RSS a request may be unlucky for a long time, but every draw gives it a chance, so it is served eventually. Under SSTF a request that is always the furthest away is never chosen at all while nearer requests keep arriving. That distinction is worth a mark.

What it does not mean

RSS is not an algorithm anybody uses. It is a baseline and a way of bounding what is possible.

The mean of 504.5 is not a measurement of one run. A single random order might give 208 or

  1. It is the average over every order, which is what a random choice gives in the long run.

208 is not SSTF done properly. It is the optimum, found by enumeration, and no greedy rule finds it in general.

A factor of three is not available on every queue. This queue was chosen because the algorithms differ on it; a queue already in cylinder order would give the same answer for all of them.

Beating the random baseline is not a high standard. It is the minimum for an algorithm to be worth its code, and FCFS does not clear it.

Quick revision

  • RSS serves a randomly chosen waiting request. It is a baseline, not a design.
  • Eight requests have 8 × 7 × 6 × 5 × 4 × 3 × 2 × 1 = 40,320 orders, and every one was

enumerated for the standard queue.

  • Best 208, mean 504.5, worst 702. The 208 order is the optimum, which proves Chapter

ninety four's point about SSTF.

  • FCFS at 640 is worse than the random mean of 504.5.
  • The order on the standard queue:

best 208, SSTF 236, LOOK 299, C-LOOK 322, SCAN 331, C-SCAN 382, random 504.5, FCFS 640, worst 702.

  • SSTF is nearest the optimum and the only one that can starve a request; the four sweeps are

safe and pay in movement.

  • Under RSS a request's wait is unbounded but it is served eventually; under SSTF it
munotes.in387

Random Scheduling, and All Six Compared

may never be served.

Test yourself

  1. What is RSS and why is it taught? Serving a randomly chosen waiting request; it is the

baseline an algorithm must beat and a way of bounding the best and worst possible totals. 2. How many orders are there for eight requests, and what were the best, mean and worst totals on the standard queue? 40,320 orders; 208, 504.5 and 702 cylinders.

  1. What does the enumeration prove about the order 37, 14, 65, 67, 98, 122, 124, 183? That

its 208 cylinders is the optimum for that queue, because every order was tried and none is shorter.

  1. Which algorithm fails to beat the random baseline, and by how much? FCFS: 640 against

504.5, which is 135 more than a random order costs on average.

  1. Put the six algorithms in order of head movement on the standard queue. SSTF 236, LOOK

299, C-LOOK 322, SCAN 331, C-SCAN 382, FCFS 640.

  1. Which algorithm is closest to the optimum, and what does it cost? SSTF, at 236 against

208, and it is the only one of the six that can starve a request.

  1. Distinguish an unbounded wait from starvation. Under RSS a request may wait a long time

but every draw gives it a chance, so it is served eventually; under SSTF a request that is always furthest from the head is never chosen while nearer ones keep arriving.

  1. Why was this particular queue chosen for all three chapters? Because the algorithms give

different answers on it: a queue already in cylinder order would give the same total for every one of them and prove nothing.

Contents This chapter on its own page

munotes.in388

Chapter Ninety-Seven

Disk Management

Syllabus topic Module 2, "Mass-Storage Structure - Disk Management"

In one line

A disk arrives as magnetic surface and has to be divided into sectors, given a file system, made bootable, protected from its own bad spots, and partly reserved for paging.

Low level formatting

Low level, or physical, formatting divides each track into sectors the controller can read and write. It is usually done at the factory.

Each sector is written as three parts:

PartWhat is in itWhy
headerthe sector numberso the controller knows which sector it is reading
data areathe bytes, usually 512the sector's contents
traileran error correcting codeto detect, and often repair, a sector that has decayed

The error correcting code is recomputed on every write and checked on every read. If the check fails but the code can repair the damage, the controller repairs it silently and the operating system never knows: that is a soft error. If it cannot, the read fails, and that is a hard error.

The header and trailer are overhead, so a disk holds less than cylinders times sectors times 512 bytes of usable data, and a larger sector size wastes proportionally less of the surface on headers.

Partitioning and logical formatting

StepWhat it does
partitioningdivides the array of blocks into ranges, each treated as a separate disk. Chapter ninety three measured one: a start block and a length
logical formatting, or making a file systemwrites the initial structures into a partition: the free space map, an empty root directory, and the superblock of Chapter fifty six hundred five

The operating system may also group blocks into clusters of several blocks: the disk transfers blocks, but the file system allocates clusters, which reduces the size of the structures that track free space and makes transfers longer and more sequential.

Some programs, chiefly databases, ask for a partition with no file system at all and do their own organisation. That is called raw disk access, and it exists because a database knows its own access pattern better than a general purpose file system does.

The boot block

The sequence a question asks for, and it is a chain of four steps because each stage can only hold so much code.

StepWhere the code isWhat it does
1ROM, on the motherboarda tiny bootstrap loader. It cannot be changed, so it is kept as small as possible
2the boot block, the first blocks of the boot diskthe ROM loads the full bootstrap program from it
3the boot partitionthe full bootstrap knows the file system, finds the kernel, and loads it
4the kernelstarts, and the machine is running
munotes.in389

Disk Management

On a system with a master boot record, the first block holds the partition table and a small loader, and the partition marked bootable holds the next stage. The reason for the indirection is that ROM is fixed at manufacture and everything after it can be replaced: a new kernel does not need a new motherboard.

A disk with no bootstrap in its boot block is a perfectly good disk that cannot start the machine. That is why the boot block is written by a separate step and not by the file system.

Bad blocks

Disks develop bad sectors, and there are three ways of living with them. A question asks for the difference between the last two.

MethodWhat happensCost
mark them at format timethe format program tests each sector and writes a value into the file system's structures meaning "unusable"; the file system never allocates itsimple, and it must be done again if a sector fails later
sector sparing, also called forwardingthe controller keeps a pool of spare sectors. A bad sector's number is remapped to a spare, and the operating system keeps using the same block numberthe spare is somewhere else on the disk, so that block now costs a seek
sector slippingthe sectors are renumbered so the whole run slides along by one, moving past the bad sector and keeping sequential orderthe data from the last sector onwards must be copied one place along

Sector sparing defeats disk scheduling, and that is the point worth making. Chapters ninety four to ninety six spent their effort on the order of block numbers; a remapped block is not where its number says it is, so the arm goes where the schedule did not intend. Controllers reduce the damage by keeping spare sectors in every cylinder, so the substitute is at least nearby.

And the honest part: the data in a bad sector is usually lost. Remapping gives you a working block number, not the bytes that were in it, which is why a system with important data has backups rather than confidence.

Swap space management

The next section of the same chapter in MU's textbook, and it finishes the module's memory story.

Where swap space lives

In a raw partitionIn a file in the file system
Speedfaster: no file system structures, no directory lookup, and large contiguous runsslower: the file system's own allocation and indirection are in the way
Flexibilityfixed size, decided when the disk is partitionedeasy to grow or add, no repartitioning
Used bysystems that care about paging speedsystems that value convenience

Swap space is optimised for speed, not for space efficiency, which is the sentence to remember. The data in it lives only as long as the process does, so nothing is worth the overhead that protects a file.

munotes.in390

Disk Management

How much

The estimate isWhat happens
too smallprocesses are aborted, or the system stops
too largedisk is wasted, and nothing else

The traditional rule of twice physical memory dates from machines whose memory was small; a modern figure depends entirely on what the machine runs. The asymmetry is what matters in an answer: too little is fatal and too much merely wasteful.

Two designs worth naming

SystemWhat it does
an older designcopies the whole process image into swap space when the process starts, and pages from there
a newer oneswaps nothing at start: pages are read from the program file on demand, and space in swap is used only when a page is replaced and must be kept

The second is better for the reason Chapter eighty gave: most of a program is never touched, so copying it all into swap is work thrown away. Note also that pages of program text need no swap at all, because they can always be read again from the program file, which is Chapter eighty four's clean victim.

What the lab machine has

$ swapon --show | wc -l
0
$ grep -c . /proc/swaps
1
$ grep -E '^Swap' /proc/meminfo
SwapCached:            0 kB
SwapTotal:             0 kB
SwapFree:              0 kB

No swap areas at all: swapon lists none, /proc/swaps has only its header line, and the totals are zero. This is the limitation Chapter eighty one and Chapter ninety one ran into and named. With no swap space there is nowhere to write a page of data, so this machine cannot page anonymous memory out, cannot show a major fault on data, and cannot be made to thrash. It kills a process that asks for too much instead.

What it can still do is everything the earlier measurements used: pages arriving on first touch, faults counted one per page, and copy on write. A missing swap area is a fact about this container, stated here rather than worked around.

Distinctions that carry marks

Low level formattingLogical formatting
Createssectors, with header, data and error correcting codethe file system's structures in a partition
Done bythe manufacturer, usuallythe operating system, on demand
After it the disk holdsnumbered sectorsa usable file system
Sector sparingSector slipping
The bad sector's numberis remapped to a spare elsewhereis skipped by renumbering
Sequential orderbroken: that block needs a seekpreserved
Costa seek on every use of that blockcopying the data along by one sector once
munotes.in391

Disk Management

Soft errorHard error
The error correcting coderepairs itcannot repair it
The operating systemis not toldgets a failed read

What it does not mean

Formatting a disk does not erase it securely. Logical formatting writes new structures; the old blocks are still there until something writes over them.

The boot block is not the kernel. It holds the bootstrap that finds the kernel.

A remapped sector does not recover the data. It gives a working block number.

Swap space is not virtual memory. It is the place pages go; Chapter eighty is why they can.

More swap space does not make a machine faster. It stops it failing when memory runs out; the paging is the cost.

Quick revision

  • Low level formatting makes sectors of header, data and an error correcting code. The

code fixes soft errors silently; a hard error is a failed read.

  • Partitioning divides the block array into ranges; logical formatting writes the file

system's initial structures. Raw disk access skips the file system, and databases use it.

  • Boot: ROM bootstrap, then the boot block's full bootstrap, then the kernel from the

boot partition. ROM is fixed, so everything after it can be replaced.

  • Bad blocks: marked at format time; sector sparing, which remaps to a spare and

defeats disk scheduling, so spares are kept per cylinder; or sector slipping, which renumbers and keeps sequential order.

  • The data in a bad sector is usually lost either way.
  • Swap space: a raw partition is faster, a file is more flexible; it is optimised for

speed, not space. Too little is fatal, too much is merely wasteful.

  • Better design: nothing copied at start, pages read from the program file, swap used

only for pages that are replaced; program text needs no swap.

  • The lab machine has no swap area, which is why it cannot show a major fault on data and

cannot thrash.

Test yourself

  1. What are the three parts of a formatted sector, and what is the third for? A header with

the sector number, the data area of usually 512 bytes, and a trailer holding an error correcting code, which detects and often repairs decay.

  1. Distinguish a soft error from a hard error. A soft error is repaired by the error

correcting code and never reported; a hard error cannot be repaired and the read fails.

  1. What does logical formatting do? It writes a file system's initial structures into a

partition: the free space map and an empty root directory among them.

  1. Why do some databases ask for a raw partition? Because they know their own access pattern
munotes.in392

Disk Management

and can organise the blocks better than a general purpose file system.

  1. Give the boot sequence in four steps. The ROM bootstrap runs; it loads the full bootstrap

program from the boot block; that finds and loads the kernel from the boot partition; the kernel starts.

  1. Why is the bootstrap split between ROM and the boot block? ROM cannot be changed after

manufacture, so only a tiny loader lives there and everything replaceable lives on the disk.

  1. Distinguish sector sparing from sector slipping. Sparing remaps a bad sector's number to a

spare elsewhere on the disk, which costs a seek and breaks sequential order; slipping renumbers the sectors to move past the bad one and preserves the order, at the cost of copying the data along by one place.

  1. Why is sector sparing bad for disk scheduling? Because the schedule assumes block numbers

reflect position, and a remapped block is somewhere else; controllers keep spares in each cylinder to limit the damage.

  1. Compare swap space in a raw partition with swap space in a file. The partition is faster,

with no file system structures in the way and large contiguous runs; the file is easier to create, grow and remove.

  1. Why is it worse to have too little swap space than too much? Too little means processes

are aborted or the system stops; too much only wastes disk.

Contents This chapter on its own page

munotes.in393

Chapter Ninety-Eight

What a File Is

Syllabus topic Module 2, "File System Interface - File Concept"

In one line

A file is a named sequence of bytes on secondary storage, and everything the system knows about it besides its contents is called its attributes.

The definition

A file is a named collection of related information recorded on secondary storage. It is the smallest logical unit of storage: the operating system offers nothing smaller with a name.

From the operating system's point of view a file is a sequence of bytes and nothing more. It does not know whether it holds a program, a picture or a letter. That is the whole reason a file system is usable by programs nobody had thought of when it was written, and it is the answer to whatever a question asks about the structure the operating system imposes on a file.

Who gives a file meaningHow
the program that wrote itby the format it chose
the userby the name and often the extension
the operating systemit does not, beyond a few types it must recognise, such as a program it can execute

The attributes

The list a question asks for, and it is worth learning all nine.

AttributeWhat it is
namethe only part a person uses; kept in the directory
identifiera number that names the file inside the file system, with no name in it: the inode number
typefor systems that distinguish them: regular file, directory, device, and so on
locationa pointer to the file's blocks on the device
sizein bytes, and sometimes also in blocks
protectionwho may read, write, and execute it
timesof creation, last modification and last access
ownerthe user, and usually a group as well
link counthow many names in the directory tree refer to it

The attributes are not in the file. They are kept in the directory structure, on the disk, beside the file rather than inside it. That is why the size of a file and the bytes of a file are two separate things, and it is why reading a file's name costs a directory access and not a read of the file.

What the lab machine keeps

Every attribute above can be asked for by name, and the figures here are worth reading twice.

$ printf 'hello\n' > greeting.txt
$ stat -c 'size %s bytes, allocated %b blocks of 512, links %h, type %F, mode %A, owner %U' greeting.txt
size 6 bytes, allocated 8 blocks of 512, links 1, type regular file, mode -rw-r--r--, owner student
$ echo "$(stat -c %s greeting.txt) bytes of data occupy $(( $(stat -c %b greeting.txt) * 512 )) bytes of disk"
6 bytes of data occupy 4096 bytes of disk
$ od -c greeting.txt
0000000   h   e   l   l   o  \n
0000006
munotes.in394

What a File Is

Three things are proved there.

The machine saysWhat it means
six bytes of data occupy 4096 bytes of diskthe file system allocates a block at a time, so a six byte file takes a whole 4 kilobyte block. 4090 bytes are wasted, and that is Chapter seventy two's internal fragmentation in a file system
h e l l o \n and then 0000006the file really is a sequence of bytes: five letters and the newline printf wrote, and no structure of any kind around them
type regular file, mode -rw-r--r--, owner studentthe attributes are kept by the system and answered without reading the file at all

The 4,090 wasted bytes are the reason a file system chooses its block size carefully, and Chapter one hundred returns to it: a large block wastes more on small files and transfers large files faster.

The operations

The six MU's textbook names, and the two that surround them.

OperationWhat the system call does
createfind space, and make a directory entry
writewrite at the current position, and advance it
readread from the current position, and advance it
reposition, or seekmove the current position without transferring anything
deleterelease the space and remove the directory entry
truncatekeep the attributes, throw the contents away, set the size to zero
openChapter ninety nine: find the file and set up the state the other calls use
closewrite the state back and release it

Every one of read, write and seek works on one current position per open file, and that position is the subject of the next chapter.

Truncate is not delete. The file keeps its name, its owner, its permissions and its identifier: only the contents go. It exists because a program that rewrites a file wants the name and permissions preserved.

Internal structure

Three ways a system can see the inside of a file, and a question about "file structure" wants the distinction.

ModelWhat the system knowsWho uses it
no structure: a stream of bytesnothing at allUNIX, Linux, and everything modern
records: a sequence of fixed or variable length recordswhere each record startsolder systems, and databases inside their own files
a structured file, a defined format the system understandsthe file's whole layoutsystems that support only a few file types

A system that understands file formats must be changed for every new format; a system that sees bytes need never change. The byte stream won, and the cost is that every program has to know its own formats.

munotes.in395

What a File Is

Packing records into blocks

Since the disk deals in blocks and a program may deal in records, something has to put one into the other.

IfThen
a record is smaller than a blockseveral records go in one block, and the leftover space is internal fragmentation
a record is larger than a blockthe record spans blocks, and reading it needs two transfers
records do not divide the block evenlythe remainder is wasted, or a record is split

The 4,090 bytes measured above is this arithmetic at its worst: one record of six bytes in a block of 4096.

Distinctions that carry marks

A file's contentsA file's attributes
Where they livein data blocksin the directory structure and the inode
Read byreadstat
Needed to list a directorynoyes
SizeAllocated space
What it countsthe bytes of datathe blocks given to the file
The six byte file above64096
The difference isinternal fragmentation
DeleteTruncate
The directory entrygoesstays
The attributesgostay
The contentsgogo
The size afterwardsthere is no filezero

What it does not mean

A file is not a stream of characters with a meaning. It is bytes; the meaning is the program's.

An extension is not a type. On most systems it is a convention the programs agree on, and the operating system does not enforce it.

The size is not the space used. A six byte file used 4,096 bytes on the lab machine.

The identifier is not the name. It is a number, and Chapter one hundred one shows one file with two names and one number.

Attributes are not stored in the file. Reading them does not read the file.

Quick revision

  • A file is a named collection of related information on secondary storage, and the

smallest logical unit of storage.

  • To the operating system it is a sequence of bytes with no structure; the meaning belongs to

the program.

  • Attributes: name, identifier, type, location, size, protection, times, owner, link count,

and they live in the directory structure, not in the file.

  • Measured: a 6 byte file occupies 4096 bytes of disk, so 4,090 bytes are wasted:

internal fragmentation in a file system.

  • Operations: create, write, read, reposition, delete, truncate, with open and close

around them.

  • Truncate keeps the name and the attributes and throws away only the contents.
  • File structure models: no structure (the byte stream, and it won), records, or a

structured file the system understands, which needs changing for every new format.

  • Records are packed into blocks, wasting the remainder, or span blocks and cost two
munotes.in396

What a File Is

transfers.

Test yourself

  1. Define a file. A named collection of related information recorded on secondary storage,

and the smallest logical unit of storage the operating system names.

  1. What structure does the operating system impose on a file's contents? None: it sees a

sequence of bytes, and the meaning belongs to the program that wrote it.

  1. List six file attributes. Name, identifier, type, location, size, protection, times, owner

and link count are nine; any six of those.

  1. Where are the attributes kept, and why does it matter? In the directory structure and the

inode, not inside the file, so they can be read without reading the file. 5. A file holds 6 bytes and the machine reports 8 blocks of 512 bytes. How much disk does it use and what is the waste called? 4,096 bytes, so 4,090 are wasted: internal fragmentation.

  1. Name the six file operations. Create, write, read, reposition or seek, delete and

truncate.

  1. Distinguish delete from truncate. Delete removes the directory entry and the attributes

with the contents; truncate keeps the name, the owner, the permissions and the identifier, and sets the size to zero.

  1. Give the three models of internal file structure, and say which is used today. No

structure, a byte stream; records; and a structured file the system understands. The byte stream is what modern systems use, and the cost is that every program must know its own formats.

Contents This chapter on its own page

munotes.in397

Chapter Ninety-Nine

Opening a File, and What the Kernel Keeps

Syllabus topic Module 2, "File System Interface - File Concept: open files"

In one line

Opening a file searches the directory once and keeps the answer, so that every later read is arithmetic instead of a search.

Why open exists

Every read could take the file's name. It does not, and the reason is cost.

If read took a nameWhat would happen
the directory would be searched on every reada directory search is one or more disk accesses
the permissions would be checked on every readmore work for a result that cannot have changed
the file's block locations would be found again every timemore work again

So open does all of it once and returns a small integer, the file descriptor, which indexes the state it set up. Every later call quotes the number.

The permission check happens at open, not at read. A file whose permissions are taken away while it is open stays readable through the descriptor already held. That is a favourite question and the answer follows from this design.

The two tables

There are two, and getting them the right way round is the whole of the topic.

Per process open file tableSystem wide open file table
One perprocessmachine
Indexed bythe file descriptornothing: entries are pointed at
Holdsthe current position, the access mode, and a pointer into the system wide tablea copy of the file's inode, its location on disk, and an open count
Entry exists whilethat process has the file openany process has the file open

So a descriptor is a number into a per process table, whose entry points at a shared entry, which describes the file. The open count in the shared entry is what makes close cheap and correct: it is decremented, and only when it reaches zero are the file's attributes written back and the entry released.

DescriptorConvention
0standard input
1standard output
2standard error
3 upwardswhatever the process opens, lowest free number first

Where the position lives

The current position is in the per process table, so it belongs to the open and not to the file. Two processes reading the same file each have their own position, and one program can open the same file twice and read two places at once.

But fork does not copy the position: it copies the descriptor, and both descriptors point at the same shared entry. So a parent and child that inherit a descriptor share one position, and a read by one moves it for the other. That is how a shell makes two commands write to the same output file without overwriting each other.

This program does both halves, and the second half is the one worth watching.

munotes.in398

Opening a File, and What the Kernel Keeps

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/wait.h>

int main(void)
{
    int fd = open("alphabet.txt", O_WRONLY | O_CREAT | O_TRUNC, 0644);
    write(fd, "abcdefghijklmnopqrstuvwxyz", 26);
    close(fd);

    int a = open("alphabet.txt", O_RDONLY);
    int b = open("alphabet.txt", O_RDONLY);
    char buf[4] = {0};
    read(a, buf, 3);
    printf("read %s through the first descriptor\n", buf);
    printf("position of the first descriptor: %ld\n", (long) lseek(a, 0, SEEK_CUR));
    printf("position of the second descriptor: %ld\n", (long) lseek(b, 0, SEEK_CUR));
    fflush(stdout);

    pid_t child = fork();
    if (child == 0) {
        char c[4] = {0};
        read(a, c, 3);
        printf("the child read %s through the inherited descriptor\n", c);
        fflush(stdout);
        _exit(0);
    }
    waitpid(child, NULL, 0);
    printf("position of the parent's descriptor after the child read: %ld\n",
        (long) lseek(a, 0, SEEK_CUR));
    char d[4] = {0};
    read(a, d, 3);
    printf("so the parent reads %s next\n", d);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -o twoopens twoopens.c
$ ./twoopens
read abc through the first descriptor
position of the first descriptor: 3
position of the second descriptor: 0
the child read def through the inherited descriptor
position of the parent's descriptor after the child read: 6
so the parent reads ghi next

Two claims, each proved by two lines.

The machine saysWhat it proves
first descriptor at 3, second at 0, after reading through the firsttwo opens of one file have two positions: the position is per open, kept in the per process table
the child read def, and the parent's position is then 6, so the parent reads ghian inherited descriptor shares one position: fork copied the descriptor, not the position, and both point at the same system wide entry

Read the last row once more. The parent never moved its own position, and the position moved. Nothing else in the file interface behaves like that, and it is the reason a shell can write ls > out; date >> out and also the reason two processes sharing a descriptor must not both read blindly.

What close does

StepWhat happens
1the per process entry is released, and the descriptor number becomes free
2the open count in the system wide entry is decremented
3if the count is now zero, the file's attributes are written back to disk and the shared entry is released

A file is not finished with until every process has closed it, which is also why deleting an open file on UNIX removes the name and leaves the contents alive until the last descriptor is closed.

And the reverse: when a process exits, the kernel closes everything it had open. A program that forgets to close is tidied up, but it may have used up its limit on open descriptors first.

munotes.in399

Opening a File, and What the Kernel Keeps

What else the open file entry holds

Kept thereWhy
the access modeso a file opened for reading cannot be written through that descriptor
the location of the file's blocksso a read needs no directory work
the open countso close and delete behave correctly
a record of locksChapter one hundred three: whether anybody has locked the file or part of it

Distinctions that carry marks

File descriptorSystem wide entry
Belongs toa processthe machine
Countsnothingthe number of processes with the file open
Two processes opening one file havetwo descriptors and two positionsone entry
A forked childshares the entry and the positionthe same entry
OpenRead
Searches the directoryyes, onceno
Checks the permissionsyesno: already done
Moves the positionsets it to 0, or to the end for appendadvances it

What it does not mean

A file descriptor is not a pointer to the file. It is an index into a table of the process's own.

The position is not an attribute of the file. It is part of the open, which is why two opens have two of them.

Fork does not give the child its own position. It gives it a second descriptor onto the same shared entry.

Close does not always write the file to disk. It writes the attributes back when the last holder closes; the data may have been written earlier by the buffer cache of Chapter one hundred four.

Removing an open file does not free its blocks. The name goes at once, the blocks when the last descriptor closes.

Quick revision

  • open searches the directory, checks the permissions once, finds the file's location, and

returns a file descriptor.

  • The permission check is at open, not at read: a descriptor keeps working after the

permissions change.

  • Per process table: indexed by the descriptor, holds the current position and the access

mode. System wide table: one entry per open file, holding the inode, the location and the open count.

  • Descriptors 0, 1 and 2 are standard input, output and error; new ones take the lowest free

number.

  • Measured: two opens of one file gave positions 3 and 0; an inherited descriptor gave the

parent position 6 after the child read, so the parent read ghi.

  • close frees the descriptor, decrements the open count, and writes the attributes back

only at zero.

  • Deleting an open file removes the name; the blocks survive until the last descriptor

closes.

Test yourself

  1. Why does read take a descriptor rather than a name? Because taking a name would mean
munotes.in400

Opening a File, and What the Kernel Keeps

searching the directory and checking permissions on every read; open does that work once.

  1. When is a file's permission checked? At open. A descriptor already held keeps working even

if the permissions change afterwards.

  1. Name the two open file tables and say what each holds. The per process table, indexed by

the file descriptor, holding the current position and the access mode; and the system wide table, one entry per open file, holding the inode, the location on disk and the open count. 4. Two processes open the same file. How many positions are there, and how many system wide entries? Two positions, one for each open, and one system wide entry. 5. A process opens a file and forks. The child reads three bytes. What has happened to the parent's position? It has advanced by three: the descriptors were copied but both point at the same system wide entry, so they share one position.

  1. What does close do? Frees the descriptor, decrements the open count in the shared entry,

and if the count reaches zero writes the attributes back and releases the entry.

  1. What are descriptors 0, 1 and 2? Standard input, standard output and standard error.
  2. What happens when an open file is deleted on a UNIX system? The directory entry goes

immediately; the contents remain until the last process closes the file.

Contents This chapter on its own page

munotes.in401

Chapter One Hundred

Access Methods

Syllabus topic Module 2, "File System Interface - Access Methods"

In one line

Read it from the beginning, jump straight to the record you want, or look the key up in an index first.

Sequential access

Read or write the next record, in order. It is the commonest pattern by a wide margin, and it is what every editor, compiler and command does unless it says otherwise.

OperationWhat it does
read nextread the record at the current position, then advance
write nextwrite at the end, and advance
reset, or rewindgo back to the beginning
skip forward n recordssome systems offer it: n read next operations, done cheaply

It is the model a magnetic tape forces, and that is where it comes from: on a tape there is nothing else you can do. A file on disk can be read sequentially just as well, and reading in block order is exactly what Chapter ninety two measured as seventy times cheaper than scattered reading.

Direct access, also called relative access

Read or write record n, for any n, in any order, with no cost for skipping.

OperationWhat it does
read nread record number n
write nwrite record number n
position to n, then read nextthe same thing in two steps: this is what lseek and a read do

It needs fixed length records, and the reason is the arithmetic.

byte position of record n = n × record length

That single multiplication is what direct access is. Records of different lengths would mean the position of record n could not be worked out without reading every record before it, which is sequential access again.

Note the word relative: the record numbers are relative to the start of the file, and the first is usually 0. The program never deals in disk block numbers, so the file system is free to put the blocks anywhere, which is Chapter one hundred seven's business.

Worked

A file of 80 byte records. Where does record 1000 begin, and which block of a 4096 byte block file system is it in?

byte position = 1000 × 80 = 80,000

block = 80,000 / 4096 = 19, remainder 2176

Byte 80,000, which is block 19, at offset 2,176 inside it. And since 2176 + 80 is less than 4096, the record does not span two blocks, so reading it costs one block access.

The follow-up a question likes: which record is the first to span two blocks? A record spans a boundary when its offset plus 80 passes 4096. With 51 records in a block, 51 × 80 = 4080, so record 51 starts at 4080 and runs to 4160, crossing the boundary at 4096.

51 × 80 = 4080

4080 + 80 = 4160

munotes.in402

Access Methods

So record 51 spans two blocks and costs two accesses, and so does every 51st record after it. That is the reason a system may pad records to divide the block evenly, wasting 16 bytes in each block to save the second access.

Demonstrated on the lab machine

$ for i in $(seq 0 99); do printf 'record %03d %-68s\n' "$i" "of a file of fixed length records"; done > records.dat
$ echo "$(stat -c %s records.dat) bytes for 100 records, so $(( $(stat -c %s records.dat) / 100 )) bytes each"
8000 bytes for 100 records, so 80 bytes each
$ dd if=records.dat bs=80 skip=42 count=1 status=none
record 042 of a file of fixed length records
$ echo "record 42 begins at byte $(( 42 * 80 ))"
record 42 begins at byte 3360
$ dd if=records.dat bs=1 skip=3360 count=10 status=none; echo
record 042

Record 42 was read without reading records 0 to 41, twice over: once by asking for the 42nd block of 80 bytes, and once by asking for byte 3,360 directly. 3,360 is 42 × 80, which is the formula above doing the work.

Indexed access

Keep an index: a table of keys, each with the record number the key is at. Search the index, then use direct access.

StepWhat happens
1search the index for the key
2the index gives a record number or a byte position
3one direct access fetches the record

The index is itself a file, and if it is large it gets an index of its own: a multi level index. The scheme a question names is ISAM, indexed sequential access method, where a small master index in memory names blocks of a secondary index on disk, which names the records.

LevelSizeWhere it lives
master indexsmallmemory
secondary indexlargerdisk, one access
the data recordsthe whole filedisk, one access

So a lookup costs two disk accesses however large the file is, which is why the scheme was worth inventing. The cost is that the index must be kept up to date: every insertion changes it, and a file whose records are inserted often spends its time maintaining indexes.

Note that the index makes access by key possible. Direct access needs the record number, which a program using names or account codes does not have; the index is the translation.

Simulating one with another

A favourite question: which can be built on which, and at what cost.

WantedBuilt onHowCost
sequentialdirectkeep a counter and read record cp, then cp + 1none: it is the natural use
directsequentialread forward, discarding records, until record n is reachedruinous: n reads to get one record
munotes.in403

Access Methods

Direct access cannot be built usefully on sequential access, and that is the whole reason the distinction exists. A tape file is sequential and nothing can be done about it; a disk file can be either.

Distinctions that carry marks

SequentialDirectIndexed
Order of accessin order onlyany orderany, and by key
Needs fixed length recordsnoyesusually
Cost of getting record nn readsone, by multiplicationtwo: the index, then the record
Given by the file system asread and writeseek and readthe program's own, or a database's
Natural ontape and diskdiskdisk
Relative record numberDisk block number
Counted fromthe start of the filethe start of the device
Known tothe programthe file system
Changes if the file moves on the disknoyes

What it does not mean

Sequential access is not slow. It is the fastest way to read a whole file, because the blocks are in order.

Direct access does not need the blocks to be contiguous. It needs fixed length records; the file system finds the block.

An index is not part of the file system. In these schemes the program or the database keeps it, in a file of its own.

Indexed access is not a third kind of hardware operation. It is an index lookup followed by a direct access.

A relative record number is not an address. It is a count, and the file system turns it into a block number.

Quick revision

  • Sequential: read next, write next, reset. The tape model, the commonest pattern, and the

cheapest way to read a whole file.

  • Direct, or relative: read or write record n in any order. It needs fixed length records

because position = n × record length.

  • Worked: 80 byte records, record 1000 starts at 1000 × 80 = 80,000, which is block 19

offset 2,176 in a 4096 byte block system, and does not span blocks.

  • Record 51 is the first to span a 4096 byte block: 51 × 80 = 4080 and

4080 + 80 = 4160.

  • Measured: a file of 100 records of 80 bytes, and record 42 read directly at byte 3,360,

which is 42 × 80.

  • Indexed: search an index of keys for a record number, then one direct access. ISAM

keeps a master index in memory and a secondary index on disk, so a lookup is two accesses; the cost is keeping the index current.

  • Sequential can be built on direct at no cost; direct on sequential costs n reads and is
munotes.in404

Access Methods

not worth having.

Test yourself

  1. Name the three access methods. Sequential, direct or relative, and indexed.
  2. What operations does sequential access offer? Read next, write next and reset, and on some

systems skipping forward by n records.

  1. Why does direct access need fixed length records? Because the position of record n is

found by multiplying n by the record length; with variable lengths it could not be calculated without reading everything before it.

  1. A file has 80 byte records and 4096 byte blocks. Where does record 1000 begin? At byte

1000 × 80 = 80,000, which is block 19 at offset 2,176, and it does not cross into block 20.

  1. Which is the first record to span two blocks, and why does it matter? Record 51:

51 × 80 = 4080 and it runs to 4160, past the 4096 boundary, so reading it costs two block accesses instead of one.

  1. Describe indexed access and the ISAM scheme. An index of keys gives the record number,

then one direct access fetches the record; ISAM keeps a small master index in memory pointing at blocks of a secondary index on disk, so a lookup costs two disk accesses.

  1. What does indexed access add that direct access lacks? Access by key rather than by record

number.

  1. Can direct access be built on sequential access? Only by reading and discarding every

record before the one wanted, which costs n reads for one record, so in practice no.

Contents This chapter on its own page

munotes.in405

Chapter One Hundred One

Directories, and the Shapes They Take

Syllabus topic Module 2, "File System Interface - Directory and Disk Structure"

In one line

A directory is a file whose contents are a table of names and the file identifiers they stand for, and the interesting question is what shape the collection of directories has.

What a directory is

A directory is a symbol table: it translates a name into the information the file system needs about a file. And it is itself a file, stored in blocks like any other, which is why Chapter one hundred six is about how to search one efficiently.

Operation on a directoryWhat it must do
search for a namethe operation everything else is built on
create a fileadd a name and its identifier
delete a fileremove the entry
list the directoryprint the names, and usually some attributes
rename a filechange the name, keeping the identifier
traverse the file systemvisit every directory and every file, for a backup or a search

Shape one: a single level directory

One directory, for everybody and everything.

Simpleone table, and the search is obvious
names must be unique across the whole machinetwo users cannot both have a file called notes
no groupinga user with forty files sees forty files

The name collision is fatal on any machine with more than one user, and it was fatal even on single user machines as soon as the disk was large enough to hold a few hundred files.

Shape two: a two level directory

One directory per user, and a master directory that names the users.

NameWhat it is
the master file directory, MFDone entry per user, pointing at that user's directory
a user file directory, UFDthat user's files

A path is then the user name and the file name, and two users may both have notes because the path distinguishes them.

Solves the collision problem between users
no grouping within a userone user's forty files are still forty files in one list
system files are a problemevery user would need a copy of the compiler, unless the system looks in a special directory too: a search path

That search path is the ancestor of the PATH variable of Chapter nine of Module 1, and it exists for exactly this reason.

Shape three: a tree structured directory

A directory may contain directories, to any depth. This is what every system has used for fifty years.

IdeaWhat it means
rootthe directory every path starts from
absolute path namethe path from the root, for example /home/student/notes/os.txt
relative path namethe path from the current working directory, for example notes/os.txt
current directoryone per process, changed by cd, and inherited by a child
the bit that says directoryin the entry, so the system knows which entries can be descended into
munotes.in406

Directories, and the Shapes They Take

Deleting a directory needs a policy, and a question asks for both options. Either a directory must be empty before it can be deleted, and the user clears it first, or the system deletes everything inside it recursively, which is convenient and dangerous. Systems offer both, under different names.

Shape four: an acyclic graph directory

Let two directories share one file, or one subdirectory, so the structure is no longer a tree. Two programmers on one project need the same file to be in both their directories, and a copy is not the same thing: a copy does not change when the other person edits it.

Way to shareHow it works
a hard link: a second directory entry naming the same file identifierboth names are the file; neither is the original
a symbolic link: a small file holding a paththe name is a pointer to a path, resolved when it is used
a duplicate entry copying the file's detailsnever done: the two copies of the attributes get out of step

The three problems, and this is what a question asks for

ProblemWhat goes wrongThe answer
several absolute path names for one filea traversal visits the same file twice, so a backup copies it twicethe traversal must recognise it, usually by identifier
deletionremoving one name must not leave the other pointing at nothingkeep a reference count in the file, and free the file when it reaches zero
a symbolic link left danglingthe target is deleted and the link still existsaccept it: using the link fails, and the link may be deleted or left

The reference count is the answer to the deletion problem, and it is the same counter as the open count of Chapter ninety nine and the shared frame count of Chapter seventy seven. The idea appears three times in this module for the same reason: something is shared, and the last user must be the one who frees it.

Demonstrated on the lab machine

$ printf 'the contents\n' > report.txt
$ ln report.txt copy-of-report.txt
$ stat -c '%n has %h name(s) and inode %i' report.txt copy-of-report.txt | sed 's/inode [0-9]*/inode N/'
report.txt has 2 name(s) and inode N
copy-of-report.txt has 2 name(s) and inode N
$ stat -c %i report.txt copy-of-report.txt | sort -u | wc -l
1
$ ln -s report.txt shortcut.txt
$ stat -c '%n is a %F of %s bytes' shortcut.txt
shortcut.txt is a symbolic link of 10 bytes
$ rm report.txt
$ echo "after removing the first name, the second still reads:"; cat copy-of-report.txt
after removing the first name, the second still reads:
the contents
$ stat -c '%n now has %h name(s)' copy-of-report.txt
copy-of-report.txt now has 1 name(s)
$ cat shortcut.txt
cat: shortcut.txt: No such file or directory
munotes.in407

Directories, and the Shapes They Take

Every line of the theory above is in that transcript.

The machine saysWhat it proves
both names report 2 names and the inode numbers reduce to onea hard link is not a copy: one file, two names, and the reference count is the 2
the symbolic link is a symbolic link of 10 bytesit is a file of its own, holding the ten characters report.txt and nothing else
after rm report.txt the other name still reads the contents, and the count is now 1rm removed a name and decremented the count; the file is freed only at zero
cat shortcut.txt failsthe symbolic link is now dangling: it holds a path that no longer resolves

The link count of a directory

The cleanest proof that . and .. are ordinary entries.

$ mkdir project
$ stat -c 'a new directory has a link count of %h' project
a new directory has a link count of 2
$ mkdir project/chapter1 project/chapter2
$ stat -c 'with two subdirectories it has %h' project
with two subdirectories it has 4
$ ls -a project
.  ..  chapter1  chapter2

An empty directory has two links, not one. Its name in the parent is one, and its own . entry is the other. Each subdirectory then adds one more, because each holds a .. entry pointing back. Two subdirectories therefore make four, and the arithmetic is exact:

links = 1 + 1 + 2 = 4

So a directory's link count is 2 plus the number of subdirectories, which is a fact worth knowing: it tells you how many subdirectories a directory has without listing it.

Shape five: a general graph directory

Allow a link from a directory to one of its own ancestors, and the structure has a cycle.

Problem a cycle causesWhy it is serious
a traversal never endssearching for a file can go round the cycle for ever
reference counts never reach zeroa cycle's entries refer to each other, so the count stays positive after the last real name is gone, and the space is never freed
The way outWhat it costs
allow links only to files, never to directoriesthe structure stays acyclic and the reference count works
limit the number of links a traversal will followsimple, and it makes a legal deep structure fail
run a cycle detection algorithm when a link is addedexpensive, and it must run on every link
garbage collection: traverse everything, mark what is reachable, then free the restcorrect, and it is extremely slow on a large disk, so it is done rarely
munotes.in408

Directories, and the Shapes They Take

UNIX takes the first way out: a hard link to a directory is not allowed, so the graph of hard links is acyclic and the link count can be trusted. Symbolic links can make cycles, and the answer there is the second way out: the system refuses after following a fixed number of links.

Distinctions that carry marks

Hard linkSymbolic link
What it isa directory entry naming the same file identifiera file holding a path
Uses disk space of its ownno, only the entryyes: the path is its contents, 10 bytes above
Counts towards the link countyesno
Survives deletion of the other nameyes: the file lives while the count is positiveno: it dangles
Can point at another file systemnoyes
Can point at a directorynot on UNIXyes
Single levelTwo levelTree
Name collisionsbetween all usersbetween one user's own filesavoided by the path
Groupingnoneper user onlyto any depth
System filesin the same directoryneed a search pathan ordinary directory
Absolute pathRelative path
Starts fromthe rootthe current working directory
Same meaning in every processyesno

What it does not mean

A directory is not a list of files. It is a list of names and identifiers; the files are elsewhere.

A hard link is not a copy. There is one file. Editing through either name changes the same bytes.

Deleting a name is not deleting a file. The file goes when the reference count reaches zero.

A symbolic link is not free. It is a file with its own inode and its own block, holding the path.

A general graph is not a richer tree. It is a structure whose traversal may not end and whose space may never be freed, which is why systems prevent it.

Quick revision

  • A directory is a symbol table from names to file identifiers, and it is itself a file.
  • Operations: search, create, delete, list, rename, traverse.
  • Single level: one directory, so names must be unique machine wide and there is no

grouping.

  • Two level: a master file directory of users, each with a user file directory.

Solves collisions between users; needs a search path for system files.

  • Tree: directories inside directories, with a root, absolute and relative paths,

and a current directory per process. Deleting a directory: empty only, or recursive.

munotes.in409

Directories, and the Shapes They Take

  • Acyclic graph: sharing by hard links or symbolic links. Three problems:

several path names, deletion, and dangling links. Deletion is solved by a reference count freed at zero.

  • Measured: a hard link gives one inode with two names; a symbolic link is a 10 byte file

holding the path; after removing one name the count is 1 and the contents survive; the symbolic link then dangles.

  • A directory's link count is 2 plus its number of subdirectories: an empty one has 2,

and with two subdirectories 1 + 1 + 2 = 4.

  • General graph: cycles make traversals endless and reference counts never zero. Ways out:

links to files only, a link limit, cycle detection, or garbage collection. UNIX forbids hard links to directories.

Test yourself

  1. What is a directory, and what is it stored as? A symbol table translating file names into

the identifiers the file system uses, stored as a file like any other.

  1. Why is a single level directory unusable? Every name must be unique across the whole

machine and there is no way to group files.

  1. What does a two level directory solve, and what does it still lack? It stops users

colliding over names by giving each a directory under a master file directory; it still offers no grouping within a user, and system files need a search path.

  1. Distinguish an absolute path from a relative path. An absolute path starts at the root and

means the same thing in every process; a relative path starts at the process's current working directory.

  1. Give the two policies for deleting a non empty directory. Refuse unless it is empty, or

delete its contents recursively.

  1. Name the three problems of an acyclic graph directory and the answer to the second.

Several absolute path names for one file, deletion, and dangling links; deletion is handled by a reference count in the file, freed only when it reaches zero.

  1. How do a hard link and a symbolic link differ? A hard link is another directory entry for

the same file and counts towards its link count; a symbolic link is a separate small file holding a path, does not count, and dangles if the target goes. Only the symbolic link can cross file systems or point at a directory.

  1. A directory has a link count of 5. How many subdirectories does it have? Three: the count

is 2 plus the number of subdirectories, one for its name in the parent, one for its own dot entry, and one for each child's dot dot.

  1. Why are cycles in a directory structure dangerous, and what does UNIX do about it? A
munotes.in410

Directories, and the Shapes They Take

traversal may never end and reference counts may never reach zero, so space is never freed; UNIX forbids hard links to directories and limits how many symbolic links it will follow.

Contents This chapter on its own page

munotes.in411

Chapter One Hundred Two

Mounting

Syllabus topic Module 2, "File System Interface - File-System Mounting"

In one line

Mounting attaches the root of a file system on a device to a chosen directory in the tree already there, so that one tree spans many devices.

What a mount is

A file system must be mounted before any file in it can be used, and mounting means naming the place in the existing tree where its root is to appear. That place is the mount point.

The operating system is givenExample
the device holding the file systema partition such as the vda1 of Chapter ninety three
the mount point: a directory in the tree already mounted/home, or /mnt/backup

What it does before it agrees

StepWhat happens
1ask the device driver to read the device's directory structure, the superblock of Chapter one hundred five
2check that it is a file system of a format the system understands, and that it is consistent
3note in the kernel's own mount table that a file system is mounted at that directory
4from then on, any path that reaches that directory continues into the mounted file system

Step 2 is the reason a mount can fail. A device holding no file system, or one left inconsistent by a crash, is refused. That refusal is a feature: it stops the system reading rubbish as though it were directories.

What happens to what was already there

The question every examiner asks, and there are two answers because systems differ.

PolicyWhat happens
hide the old contentsthe directory's own files become inaccessible for as long as the file system is mounted, and reappear when it is unmounted. Nothing is lost
refuse a non empty mount pointthe mount fails unless the directory is empty, which prevents the surprise

The first is what UNIX does. A file at /mnt/notes.txt is invisible while a file system is mounted at /mnt, and nothing has happened to it: the path now resolves into the other file system.

The same reasoning covers unmounting. A file system that is busy, because some process has a file open in it or has its current directory inside it, cannot be unmounted: the system refuses rather than pull the tree out from under a running program.

Who mounts, and when

PolicyHow it works
at boot, automaticallya table on disk lists devices and their mount points, and the system mounts each one as it starts
on demandthe system mounts a file system the first time a path into it is used, which is how a network file system is often arranged
by handa command, usually restricted to the administrator, because mounting affects every user
munotes.in412

Mounting

Mounting is a privileged operation for a simple reason: whoever mounts a file system chooses the owners and permissions recorded in it, so an ordinary user who could mount could bring a file owned by anybody into the tree.

One tree or many

The design difference worth knowing, because it changes what a path looks like.

SystemHow several devices appear
UNIX and Linuxone tree. Every file system is mounted somewhere inside it, and a path never says which device it is on
Windowsone namespace per volume, named by a drive letter: C: and D: are separate trees, and a path begins with the device
macOSone tree, with removable volumes mounted automatically under a standard directory

The UNIX arrangement means a program never knows or cares which device its file is on, and an administrator can move a directory to a new disk by mounting it in the same place. The Windows arrangement makes the device visible in every path, which is simpler to explain and harder to change.

The lab machine's own tree

Four different file systems, four different types, in one tree, and one of them is not a disk at all.

$ stat -f -c '%n is a file system of type %T' / /proc /sys /dev/shm
/ is a file system of type overlayfs
/proc is a file system of type proc
/sys is a file system of type sysfs
/dev/shm is a file system of type tmpfs
$ awk '$2 == "/proc" || $2 == "/sys" {print $2, "is mounted, and holds a", $3, "file system"}' /proc/mounts
/proc is mounted, and holds a proc file system
/sys is mounted, and holds a sysfs file system
$ test "$(stat -c %d /)" != "$(stat -c %d /proc)" && echo "the root and /proc are on different devices, in one tree"
the root and /proc are on different devices, in one tree

Read the last line first. /proc is on a different device from /, and a path walks from one to the other without saying so. That is mounting doing its work: /proc/self/status, which Chapter eighty read, is a path through two file systems.

The mounted file systemWhat is really behind it
overlayfs at /two directories on the host, layered so that writes go to one of them
proc at /procno storage at all: the kernel makes up the files as they are read
sysfs at /systhe kernel's own objects, which is how Chapter ninety three read the disk's block count
tmpfs at /dev/shmmemory, not disk: files in it disappear when the machine stops

Three of the four store nothing on a disk. Mounting is not a disk operation; it is a namespace operation, and a file system is anything that can answer the file system interface of Chapter one hundred four. That is worth one sentence in an answer about mounting.

munotes.in413

Mounting

Distinctions that carry marks

MountingFormatting
What it doesattaches an existing file system to the treecreates a file system on a device
How oftenevery bootonce
Destroys datanoyes
Mount pointRoot of the mounted file system
Belongs tothe existing treethe new file system
After mountingits own contents are hiddenit is what the path now reaches
UNIXWindows
Several devicesone treeone namespace per volume
A path says which devicenoyes, the drive letter
Moving a directory to another diskmount it in the same placethe paths change

What it does not mean

Mounting does not copy anything. It records where a file system's root is to appear.

A hidden mount point directory is not deleted. Its contents come back when the file system is unmounted.

A file system is not always a disk. On the lab machine three of four are not.

Unmounting is not always possible. A busy file system is refused.

Mounting is not something an ordinary user does. It is privileged, because the mounted file system carries its own owners and permissions.

Quick revision

  • Mounting attaches a file system on a device to a mount point in the tree already there;

nothing in it can be used before that.

  • Before agreeing, the system reads the superblock, checks the format and the

consistency, and records the mount in its mount table.

  • Files already in the mount point directory are hidden while the mount lasts, or the mount

is refused if the directory is not empty. Nothing is lost either way.

  • A busy file system, one with an open file or a process's current directory inside it,

cannot be unmounted.

  • Mounting happens at boot from a table, on demand, or by hand, and it is

privileged because the mounted file system carries its own owners and permissions.

  • UNIX has one tree; Windows has a namespace per volume, named by a drive letter.
  • Measured on the lab machine: / is overlayfs, /proc is proc, /sys is sysfs,

/dev/shm is tmpfs, and / and /proc are on different devices in one tree. Three of the four store nothing on a disk.

Test yourself

  1. What does mounting do? It attaches the root of a file system on a device to a directory,

the mount point, in the tree that is already mounted, so that paths reaching that directory continue into the new file system.

munotes.in414

Mounting

  1. What does the operating system check before it mounts? That the device holds a file system

of a format it understands and that it is consistent, by reading its superblock through the device driver.

  1. What happens to files already in the mount point directory? They become inaccessible while

the file system is mounted and reappear when it is unmounted; some systems instead refuse to mount onto a non empty directory.

  1. Why can a file system be impossible to unmount? Because it is busy: a process has a file

open in it or has its current directory inside it.

  1. Give three policies for when a file system is mounted. Automatically at boot from a table

of devices and mount points; on demand when a path into it is first used; or by hand with a privileged command.

  1. Why is mounting privileged? Because a mounted file system carries its own record of owners

and permissions, so anyone who could mount could introduce files owned by anybody.

  1. How do UNIX and Windows differ in handling several devices? UNIX mounts every file system

into one tree, so a path never names a device; Windows gives each volume its own namespace under a drive letter, so every path begins with the device. 8. The lab machine has four file systems of four types and three store nothing on disk. What does that show about mounting? That it is an operation on the namespace, not on a disk: anything that can answer the file system interface can be mounted, including the kernel's own made up files and a file system in memory.

Contents This chapter on its own page

munotes.in415

Chapter One Hundred Three

File Sharing, and Locking

Syllabus topic Module 2, "File System Interface - File Sharing"

In one line

Sharing a file means deciding who may touch it, when each of them sees the others' changes, and how two writers are kept from writing at once.

Several users: owner, group and everybody

Sharing needs the system to know who is asking, and the scheme every UNIX uses has three classes and three rights.

ClassWho it is
ownerthe user who created the file
groupa named set of users the owner belongs to
others, or the universeeverybody else
RightOn a file it meansOn a directory it means
readread the contentslist the names in it
writechange the contentscreate and delete entries in it
executerun it as a programpass through it in a path

On a directory, execute means the right to use it in a path, and it is separate from the right to list it. A directory with execute but not read lets a user open a file whose name they already know and refuses to tell them what else is there. That distinction is a favourite question and it follows from the table above.

The nine bits are written as three octal digits, read, write and execute being 4, 2 and 1:

rw- r-- r-- = 644

rwx r-x r-x = 755

Chapter ninety eight's demonstration printed -rw-r--r--, which is 644: the owner may read and write, and everybody else may read.

A file on another machine

The three things a remote file system needs, and a question on file sharing usually wants the middle one.

NeedHow it is met
get at the filethe client server model: the server exports a directory and the client mounts it, so Chapter one hundred two's mechanism does the work
know who the user isa distributed information system holds the user names and identities for every machine, so that the same person is the same owner everywhere
survive failurea local disk fails rarely and visibly; a remote file system fails when the server stops, the network breaks, or either is merely slow

Stateful and stateless service

The distinction to be able to give, because the recovery behaviour follows from it.

Stateful serviceStateless service
The server remembersthat the client has the file open, and its positionnothing between requests
A request carriesa short identifiereverything needed: the file, the position, the length
Fasteryes: the server has the state to handrequests are longer
Recovery after the server restartshard: the state is gone and the clients believe they have files openeasy: the next request works as though nothing happened

The system a question names is NFS, which chose to be stateless for exactly that reason. The client keeps retrying and the user sees a pause rather than an error.

munotes.in416

File Sharing, and Locking

Consistency semantics: when the others see your changes

Three named answers, and they are the heart of the topic.

SemanticsWhen a write becomes visible to the othersUsed by
UNIX semanticsimmediately. Writers and readers sharing the file see one ordering of the writes, and a shared position moves for everybody, exactly as Chapter ninety nine measuredUNIX, for local files
Session semanticsnot until the writer closes the file, and then only to opens that begin afterwards. A session already reading sees the old contents to the endthe Andrew File System
Immutable shared filea shared file cannot be changed at all: its name and its contents are fixed, so there is nothing to make consistentsystems that share by publishing a new file instead

Session semantics exists because of the network. Sending every write to the server at once is too expensive over a network, so the file is copied to the client, written there, and sent back at close. The price is that two people editing one file do not see each other, and the last to close wins.

Locking

Consistency semantics say when changes are seen. Locks are how a program says not yet.

KindWho may hold itWhat it prevents
shared, or readerseveral processes at oncea writer from getting in
exclusive, or writerone processeverybody else
KindWhat the system does
mandatorythe system refuses a read or write that breaks a lock: the lock cannot be ignored
advisorythe system records the lock and leaves it to the processes to ask. A process that does not ask is not stopped

The lock is recorded in the open file entry of Chapter ninety nine, which is why it is the kernel and not the program that knows about it, and why closing the file releases it.

Locks can deadlock, and the first half of this module is about that. Two processes each holding one lock and waiting for the other's is exactly Chapter fifty seven's four conditions, on files instead of resources. The banker's algorithm is not used here: systems either make the request non blocking, so it fails instead of waiting, or leave the programmer to take locks in a fixed order.

Demonstrated on the lab machine

$ printf 'the shared file\n' > data.txt
$ flock -x data.txt sh -c 'flock -n -x data.txt echo "the second process took the lock" || echo "the second process could not take the lock"'
the second process could not take the lock
$ flock -s data.txt sh -c 'flock -n -s data.txt echo "two readers can hold a shared lock at the same time"'
two readers can hold a shared lock at the same time
$ flock -s data.txt sh -c 'flock -n -x data.txt echo "a writer got in" || echo "a writer cannot take the lock while a reader holds it"'
a writer cannot take the lock while a reader holds it
$ flock -x data.txt sh -c 'printf "written anyway\n" >> data.txt; echo "a process that never asks for the lock wrote to the file"'
a process that never asks for the lock wrote to the file
$ cat data.txt
the shared file
written anyway
munotes.in417

File Sharing, and Locking

Four claims, each with its line.

The machine saysWhat it proves
the second process could not take the lockan exclusive lock excludes another exclusive one
two readers held it at the same timea shared lock is shared
a writer cannot take it while a reader holds itshared and exclusive are incompatible, which is the readers and writers problem of Module 1
a process that never asks wrote to the file anywaythese locks are advisory. The lock is a convention between programs that agree to ask; it is not a wall

That last row is the one to carry into an answer about UNIX. Nothing stops an unco-operative program, and a system that needs to stop it must use mandatory locking, which few systems now offer because a stuck lock then makes a file unusable.

Distinctions that carry marks

Consistency semanticsLocking
Answerswhen other processes see a changewhether a process may proceed now
Chosen bythe file system's designerthe program, at run time
ExampleUNIX semantics, session semanticsshared and exclusive locks
Shared lockExclusive lock
Holders at oncemanyone
Blocks readersnoyes
Blocks writersyesyes
Mandatory lockingAdvisory locking
Enforced bythe system, on every read and writenobody: the programs co-operate
A program that does not askcannot touch the filecan
Riska forgotten lock makes the file unusablea careless program corrupts the file
Stateful serviceStateless service
Server keepsopen files and positionsnothing
Request sizesmalllarge: self contained
After a server crashclients must recover their statethe next request just works

What it does not mean

Execute permission on a directory is not the right to list it. It is the right to pass through it in a path.

UNIX semantics does not mean two writers are safe. It means their writes are seen at once and in one order; keeping them out of each other's way is what locks are for.

munotes.in418

File Sharing, and Locking

Session semantics is not a bug. It is a deliberate choice for a network, and the cost is that the last close wins.

An advisory lock does not protect a file. It coordinates programs that ask.

A stateless server is not a slower server in every way. Its requests are larger, and it needs no recovery at all.

Quick revision

  • Three classes, owner, group, others, and three rights, read, write, execute, written in

octal as 4, 2, 1: rw-r--r-- is 644.

  • On a directory, read is listing and execute is passing through in a path; they are

separate.

  • A remote file system needs access by the client server model and a mount, identity from

a distributed information system, and a plan for failure.

  • Stateful servers remember open files and positions and recover badly; stateless servers

remember nothing, send bigger requests, and recover trivially. NFS is stateless.

  • Consistency semantics: UNIX, visible immediately; session, visible only after

close and only to later opens; immutable shared files, which cannot change.

  • Locks are shared (many readers) or exclusive (one writer), and mandatory or

advisory. They live in the open file entry and are released by close.

  • Measured: an exclusive lock excludes another; two shared locks coexist; a writer is refused

while a reader holds it; and a process that never asks writes anyway, because these locks are advisory.

  • Locks can deadlock: the four conditions of the first half of this module apply to files as

well as devices.

Test yourself

1. Give the three classes and three rights of the UNIX protection scheme, and write rw-r--r-- in octal. Owner, group and others; read, write and execute, worth 4, 2 and 1; rw-r--r-- is 644.

  1. What do read and execute mean on a directory? Read is the right to list the names in it;

execute is the right to pass through it in a path, so a user with execute but not read can open a file whose name they already know.

  1. What three problems does a remote file system have to solve? Reaching the file, usually by

mounting an exported directory; agreeing on who the user is, through a distributed information system; and coping with failures of the server or the network.

  1. Distinguish stateful from stateless service, and say which NFS chose. A stateful server

remembers that a client has a file open and where its position is, which is faster but hard to recover after a crash; a stateless server remembers nothing and each request carries everything. NFS is stateless.

  1. Give the three consistency semantics. UNIX semantics, where a write is visible immediately
munotes.in419

File Sharing, and Locking

to everyone sharing the file; session semantics, where it is visible only after the writer closes and only to opens that start afterwards; and immutable shared files, which cannot be changed at all.

  1. Why does session semantics exist? Because sending every write across a network as it

happens is too expensive: the file is copied to the client, written there, and returned at close. 7. Distinguish a shared lock from an exclusive one, and mandatory locking from advisory locking. A shared lock may be held by many readers and keeps writers out; an exclusive lock is held by one and keeps everyone out. Mandatory locking is enforced by the system on every read and write; advisory locking is only recorded, and a program that does not ask is not stopped.

  1. On the lab machine a process wrote to a locked file. What does that prove? That the locks

are advisory: they coordinate programs that agree to ask for them and do not stop one that does not.

  1. Can file locks deadlock? Yes: two processes each holding one lock and waiting for the

other's satisfy the four conditions of deadlock. Systems avoid it with non blocking requests or by taking locks in a fixed order.

Contents This chapter on its own page

munotes.in420

Chapter One Hundred Four

The Layers a Read Passes Through

Syllabus topic Module 2, "File System Implementation - File-System Structure"

In one line

A read passes through a layer that knows names, a layer that knows files, a layer that knows blocks and caches them, and a layer that knows the hardware.

The layers

The diagram a question asks for, from the top down.

LayerWhat it knowsWhat it does
application programsfiles and namescall open, read, write
logical file systemnames, directories, protectionresolves a path, checks the permissions, and finds the file control block
file organisation modulefiles and their logical blockstranslates logical block 3 of this file into physical block 91,482, and manages the free space
basic file systemphysical blocksissues generic commands to the driver, and keeps the buffers and caches of blocks in memory
input and output controlthe hardwaredevice drivers and interrupt handlers: turns "read block 91,482" into the commands this controller understands
devices

The translation from a file's own block numbering to the device's is the middle layer's whole job, and it is the one students merge with the others. Chapter one hundred seven is how it is done.

What each layer would have to change for

The test of whether a layering is right: what has to be rewritten when something changes.

If this changesOnly this layer changes
a new kind of disk controllerinput and output control: a driver
a new file system format, with different allocationthe file organisation module and the logical file system's structures
a new directory structure or protection schemethe logical file system
a better caching policythe basic file system

That is why systems can support many file systems at once: everything below the file organisation module is shared, and Chapter one hundred two's four file systems in one tree is that sharing at work.

One read, step by step

The path of a single read(fd, buf, 4096), which is the answer to "trace a read through the file system".

StepLayerWhat happens
1logical file systemthe descriptor of Chapter ninety nine gives the open file entry: no path is resolved, because open did that
2logical file systemthe file control block, held in memory, gives the file's size and where its blocks are described
3file organisation modulethe position is divided by the block size to get a logical block number, and the file's own structures turn that into a physical block
4basic file systemis that block already in memory? If yes, copy it to the program and stop
5basic file systemif not, ask the driver for it, and keep it in the cache afterwards
6input and output controlthe driver programs the controller, waits for the interrupt, and returns the block
7back up the layersthe bytes are copied into the program's buffer and the position advances
munotes.in421

The Layers a Read Passes Through

Step 4 is where most reads end, and the next section measures it.

The caches, and there are three

A question about file system performance wants these named.

CacheHoldsWhy it works
the buffer cache, or page cacheblocks of filesa block read once is often read again, and a block written is often written again
the directory entry cachethe name to inode translationsresolving a path touches every directory in it, and the same paths are used over and over
the open file tablesthe file's location and attributesso no directory work happens on a read at all

The Linux kernel's own documentation says of the directory entry cache that its entries are "never saved to disc: they exist only for performance."

Read ahead belongs here too: when a file is being read in order, the basic file system fetches the next blocks before they are asked for, which is why Chapter ninety two's sequential read is as fast as it is.

Measured: the read that never reached the device

The kernel counts, per process, both the bytes a program reads and the bytes actually fetched from the device. The two are not the same number.

#define _POSIX_C_SOURCE 200809L
#include <stdio.h>
#include <string.h>
#include <fcntl.h>
#include <unistd.h>

/* One counter from /proc/self/io: what this process has read. */
static long io(const char *name)
{
    FILE *f = fopen("/proc/self/io", "r");
    char line[128];
    long v = -1;
    while (fgets(line, sizeof line, f))
        if (strncmp(line, name, strlen(name)) == 0)
        sscanf(line + strlen(name), " %ld", &v);
    fclose(f);
    return v;
}

static long slurp(const char *path)
{
    static char buf[65536];
    int fd = open(path, O_RDONLY);
    long total = 0, n;
    while ((n = read(fd, buf, sizeof buf)) > 0)
        total += n;
    posix_fadvise(fd, 0, 0, POSIX_FADV_DONTNEED);   /* ask for its pages back */
    close(fd);
    return total;
}

int main(void)
{
    long before = io("read_bytes:");
    long bytes = slurp("big.dat");
    long after = io("read_bytes:");
    printf("the program read %ld bytes of the file\n", bytes);
    printf("the device was asked for %ld bytes of it\n", after - before);

    long before2 = io("read_bytes:");
    slurp("big.dat");
    long after2 = io("read_bytes:");
    printf("after the cache was dropped, the same read fetched %ld bytes from the device\n",
        after2 - before2);
    return 0;
}
$ dd if=/dev/zero of=big.dat bs=1M count=32 status=none
$ gcc -std=c17 -Wall -Wextra -o layers layers.c
$ ./layers
the program read 33554432 bytes of the file
the device was asked for 0 bytes of it
after the cache was dropped, the same read fetched 0 bytes from the device
munotes.in422

The Layers a Read Passes Through

Thirty two megabytes read by the program, and nothing at all asked of the device. Every block was already in the basic file system's cache, because dd had just written them; the read was a copy from memory to memory and the layers below were never entered.

And the third line is the honest part. The program asks the kernel to give the file's pages back, and the next read still fetches nothing from the device. Two reasons, and both are worth knowing: the request is advice, which the kernel may ignore, and the storage under this container is itself a file on the host, cached there as well. This machine cannot be made to read a block from a real device, which is the same kind of limit as the missing swap space of Chapter ninety seven, and it is stated rather than worked around.

What the measurement does prove is the layering: read_bytes is the basic file system's counter, rchar is the logical file system's, and the gap between 33,554,432 and 0 is the cache.

Distinctions that carry marks

Logical file systemFile organisation moduleBasic file system
Deals innames and directoriesa file's logical blocksthe device's physical blocks
Knows about protectionyesnono
Knows where a file's blocks arenoyesno
Keeps the cachesnonoyes
Changed by a new disk formatthe structuresyesno
Logical block numberPhysical block number
Counted fromthe start of the filethe start of the device
Produced bydividing the file position by the block sizethe file organisation module

What it does not mean

The layers are not processes. They are parts of the kernel, called one from another.

A read does not always reach the disk. Most do not: the cache answers them.

The file organisation module is not the directory. The directory is the logical layer's; this one deals in block numbers.

Read ahead is not guesswork about the program. It is a response to what the program has already done, which is Chapter seventy six's locality again.

A layered design is not free. Every layer copies and checks, and the gain is that each can be replaced alone.

Quick revision

  • The layers, top down: application programs, logical file system (names, directories,

protection, the file control block), file organisation module (logical block to physical block, and free space), basic file system (physical blocks, buffers and caches), input and output control (drivers and interrupt handlers), devices.

  • The logical to physical block translation is the middle layer's job and belongs to no

other.

  • A new controller needs only a driver; a new format needs the organisation module; a new
munotes.in423

The Layers a Read Passes Through

directory scheme needs the logical layer. That is why one kernel supports many file systems.

  • One read: descriptor, file control block, position divided by block size,

is the block cached, else the driver and an interrupt, then copy up.

  • Three caches: the buffer or page cache of blocks, the directory entry cache of name to

inode translations, and the open file tables. Read ahead fetches the next blocks of a sequential read.

  • Measured: a program read 33,554,432 bytes and the device was asked for 0. The lab

machine cannot be made to fetch a block from a real device, because the request to drop the cache is advice and the storage is itself cached on the host.

Test yourself

  1. Name the layers of a file system from the top down. Application programs, the logical file

system, the file organisation module, the basic file system, input and output control, and the devices.

  1. Which layer translates a logical block number into a physical one? The file organisation

module, which also manages free space.

  1. Which layer knows about protection and directories? The logical file system, which also

holds the file control block.

  1. Which layer keeps the caches, and what does it deal in? The basic file system, which deals

in physical blocks and issues generic commands to the driver.

  1. What has to change for a new disk controller, and what for a new file system format? A new

controller needs only a driver in the input and output control layer; a new format needs the file organisation module and the logical layer's structures.

  1. Trace a read of one block through the layers. The descriptor gives the open file entry;

the file control block gives the file's details; the position divided by the block size gives a logical block, which the organisation module turns into a physical block; the basic file system checks its cache and, if the block is absent, asks the driver, which programs the controller and waits for the interrupt; the bytes are then copied into the program's buffer.

  1. Name the three caches a file system uses. The buffer or page cache of file blocks, the

directory entry cache of name to inode translations, and the open file tables.

  1. A program reads 32 megabytes and the device is asked for nothing. What has happened? Every

block was already in the basic file system's cache, so the read was a copy within memory and the lower layers were never reached.

Contents This chapter on its own page

munotes.in424

Chapter One Hundred Five

On the Disk: Superblock, Inode, Data Blocks

Syllabus topic Module 2, "File System Implementation - File-System Implementation"

In one line

A file system on disk is a boot block, a description of the whole volume, a table of file descriptions, and the blocks of the files themselves.

The four structures on disk

The list a question asks for, in the order they appear on the volume.

StructureAlso calledWhat it holds
boot control blockboot block, partition boot sectorwhat is needed to boot an operating system from this volume. Empty if the volume is not bootable
volume control blocksuperblock in UNIX, master file table in NTFSthe number of blocks, the block size, the count of free blocks and where the free list is, the count of free file control blocks and where they are
directory structurethe names in the file system and, with each, the number of its file control block
file control block, one per fileinode in UNIXpermissions, dates, owner, size, and pointers to the file's data blocks

The name is in the directory and everything else is in the inode. That division is the whole reason a file can have two names, as Chapter one hundred one measured: the two entries hold one inode number, and the inode holds the file.

What the operating system keeps in memory

Four more structures, and a question sometimes asks for "in memory" rather than "on disk".

In memoryWhy
the mount tableone entry per mounted file system: Chapter one hundred two
the directory entry cacheso a path already resolved need not be read again
the system wide open file tablea copy of each open file's inode, and the open count: Chapter ninety nine
the per process open file tablethe position and mode for each descriptor
the buffersblocks on their way to or from the disk: Chapter one hundred four

Note what these have in common: every one of them is a copy of something on the disk, or a note about who is using it. The disk is the truth and memory is the working copy, which is why an unmounted file system is consistent and a crashed one may not be.

Building one, on the lab machine

mkfs writes those structures into a file, and dumpe2fs reads them back. Two file systems are made, the second with a journal, because the difference is instructive.

$ mkfs.ext4 -q -F -b 1024 -I 128 -O ^has_journal plain.fs 4096 2>&1 | grep -v deprecated
Creating regular file plain.fs
$ mkfs.ext4 -q -F -b 1024 -I 128 plain-with-journal.fs 4096 2>&1 | grep -v deprecated
Creating regular file plain-with-journal.fs
$ dumpe2fs -h plain.fs 2>/dev/null | grep -E '^(Block count|Inode count|Block size|Inode size|Free inodes|First inode)' | tr '\t' ' '
Inode count:              1024
Block count:              4096
Free inodes:              1013
Block size:               1024
First inode:              11
Inode size:           128
$ for f in plain.fs plain-with-journal.fs; do echo "$f: metadata uses $(( 4096 - $(dumpe2fs -h $f 2>/dev/null | awk '/^Free blocks/ {print $3}') )) blocks of 4096"; done
plain.fs: metadata uses 178 blocks of 4096
plain-with-journal.fs: metadata uses 1202 blocks of 4096
munotes.in425

On the Disk: Superblock, Inode, Data Blocks

Read the superblock's own words. 4096 blocks of 1024 bytes, which is the four megabyte file system asked for; 1024 inodes of 128 bytes each; and First inode 11, meaning inodes 1 to 10 are reserved by the format.

The cost of the structures

inode table = 1024 × 128 = 131,072

in blocks of 1024 = 131,072 / 1024 = 128

128 of the 4,096 blocks are the inode table alone, and the measured total is 178: the rest is the superblock, the group descriptors, the two bitmaps of Chapter one hundred nine, and the root directory.

journal cost = 1202 - 178 = 1024

The journal costs 1,024 blocks, a quarter of this file system. It is a log of the changes the file system is about to make, so that a crash can be recovered by replaying or discarding it, and on a small volume its fixed size is enormous in proportion. On a large volume the same megabyte is nothing, which is why the default is to have one.

The general arithmetic a question can ask for:

usable blocks = total blocks - metadata blocks

BlocksShare of the volume
total4096
metadata, without a journal178about 4 per cent
metadata, with a journal1202about 29 per cent
usable, with a journal2894about 71 per cent

Reading the structures back

debugfs walks the same structures the kernel would, from outside the kernel.

$ printf 'a file inside the image\n' > payload.txt
$ debugfs -w -R "write payload.txt notes.txt" plain.fs 2>&1 | grep -v "^debugfs"
Allocated inode: 12
$ debugfs -R "ls -l /" plain.fs 2>/dev/null | awk '{print $1, $2, $NF}'
2 40755 .
2 40755 ..
11 40700 lost+found
12 100644 notes.txt
$ debugfs -R "stat <12>" plain.fs 2>/dev/null | grep -E '^(Inode:|User|Links)' | tr '\t' ' '
Inode: 12   Type: regular    Mode:  0644   Flags: 0x80000
User:     0   Group:     0   Size: 24
Links: 1   Blockcount: 2
$ echo "and the free blocks are now $(dumpe2fs -h plain.fs 2>/dev/null | awk '/^Free blocks/ {print $3}')"
and the free blocks are now 3917

Five facts, each from one line.

The image saysWhat it shows
the root directory is inode 2the root's inode number is fixed by the format, so the system can find it without searching
lost+found is inode 11the first non reserved inode, exactly as First inode 11 promised
the new file got inode 12inodes are handed out from a free list in the superblock, next free first
the directory holds name and inode number and nothing elsenotes.txt and 12: the attributes are in the inode, not here
free blocks fell from 3918 to 3917a 24 byte file took one whole block of 1024: internal fragmentation, and the superblock's free count kept up to date
munotes.in426

On the Disk: Superblock, Inode, Data Blocks

The inode's own Size: 24 and Blockcount: 2 say the same thing twice: 24 bytes of data, and 2 units of 512 bytes allocated, which is the one 1024 byte block.

Distinctions that carry marks

SuperblockInode
How manyone per file system, with copies for safetyone per file
Describesthe whole volume: block size, counts, free listsone file: owner, size, times, block pointers
Holds a namenono
Read atmount timeevery time the file is used
Directory entryInode
Holds the nameyesno
Holds the attributesnoyes
Two of them can share oneyes: that is a hard link
On diskIn memory
Which structuresboot block, superblock, directories, inodes, data blocksmount table, directory entry cache, the two open file tables, buffers
Which is the truththe diska working copy
After a crashmay be inconsistentgone

What it does not mean

The superblock is not the boot block. One describes the file system, the other starts the machine.

An inode does not hold the file's name. Names are in directories, which is why one file can have several.

First inode 11 does not mean ten inodes are wasted. They are reserved for the format's own use, the root directory among them.

Metadata is not overhead you can remove. 178 blocks of this file system are what make the other 3,918 usable.

A journal is not a backup. It is a log of pending changes, so that a crash leaves the structures consistent.

Quick revision

  • On disk: the boot control block, the volume control block or superblock, the

directory structure, and one file control block or inode per file, plus the data blocks.

  • The superblock holds the block size, the block and inode counts, and the free lists. The

inode holds the owner, permissions, times, size and the pointers to the data blocks. The name is in the directory.

  • In memory: the mount table, the directory entry cache, the system wide and

per process open file tables, and the buffers. All copies: the disk is the truth.

  • Measured on a 4 megabyte ext4 image: 4096 blocks of 1024, 1024 inodes of 128 bytes,
munotes.in427

On the Disk: Superblock, Inode, Data Blocks

First inode 11.

  • The inode table is 1024 × 128 = 131,072 bytes, which is 128 blocks; all metadata is

178 blocks, about 4 per cent.

  • With a journal the metadata is 1202 blocks, so the journal costs 1,024 of them, about

29 per cent of a small volume and nothing on a large one.

  • Measured: the root directory is inode 2, lost+found is 11, a new file got 12, the

directory holds only name and inode number, and a 24 byte file took one 1024 byte block.

Test yourself

  1. Name the four on disk structures of a file system. The boot control block, the volume

control block or superblock, the directory structure, and the file control blocks or inodes.

  1. What does a superblock hold? The number of blocks, the block size, the count of free

blocks and where the free list is, and the same for free file control blocks.

  1. What does an inode hold, and what does it not? Permissions, dates, owner, size and

pointers to the data blocks; it does not hold the file's name.

  1. Where is a file's name kept, and why there? In the directory, so that several names can

refer to one inode and so that renaming a file does not touch it.

  1. Name four structures the operating system keeps in memory. The mount table, the directory

entry cache, the system wide open file table and the per process open file table, with the block buffers as a fifth. 6. A file system has 1024 inodes of 128 bytes with 1024 byte blocks. How many blocks does the inode table take? 1024 × 128 = 131,072 bytes, which is 128 blocks. 7. A 4 megabyte image uses 178 metadata blocks without a journal and 1202 with one. What does the journal cost, and is it worth it? 1,024 blocks, about a quarter of this volume; on a large volume the same megabyte is negligible, and it buys recovery of a consistent file system after a crash.

  1. What did the free block count show when a 24 byte file was created? It fell by exactly one

block of 1024 bytes, so the file wasted 1,000 bytes and the superblock's count was kept up to date.

Contents This chapter on its own page

munotes.in428

Chapter One Hundred Six

Directory Implementation

Syllabus topic Module 2, "File System Implementation - Directory Implementation"

In one line

The names in a directory are kept either in a list, which is simple and slow to search, or in a hash table, which is fast and fixed in size.

Why the choice matters

Every path resolution is a directory search. Opening /home/student/notes/os.txt searches four directories, and that happens on every open, every stat, every ls of a path. The directory search is the most frequent operation in the file system, which is why the choice of structure is worth a chapter.

A linear list

A list of entries, each holding a name and the inode number. It is what the ext4 image of Chapter one hundred five had: the debugfs listing printed a name and a number for each entry and nothing else.

OperationWhat it costs
search for a namea linear scan: on average half the entries, and all of them if the name is absent
create a filea full search first, to be sure the name is not already there, then add an entry
delete a filea search, then release the entry
list the directoryread it through: this one is ideal

Creating a file costs a search of the whole directory, not half. The system must prove the name is absent, and proving absence means looking everywhere. That is the answer to why creating files in a huge directory is slow.

What to do with a deleted entry

MethodWhat happens
mark it unused, with a special name or a zero inode numbersimple; the directory never shrinks and the holes are searched over
move the last entry into the holekeeps the list packed; loses any ordering
keep a free list of the holesa new entry goes into the first hole that fits

A sorted list allows a binary search, which turns 500 comparisons into about 10, but every insertion has to move the entries after it. Keeping the list as a linked list in sorted order, or as a B tree, is how real systems get the search without the insertion cost.

The cost as a number

A directory of 1,000 entries in a file system with 1024 byte blocks, each entry taking about 16 bytes:

entries per block = 1024 / 16 = 64

blocks in the directory = 1000 / 64, about 16

OperationComparisonsBlock reads
a successful search, on average500about 8
an unsuccessful search, or a create1000about 16
a binary search of a sorted listabout 10about 4, one per probe

The comparisons are cheap and the block reads are not: Chapter ninety two priced a block read at about nine milliseconds. A linear directory search of a large directory is a disk operation, and that is the real cost.

munotes.in429

Directory Implementation

A hash table

Hash the name to a number, and keep the entry in the bucket with that number. The search is then one hash and a short chain instead of a scan.

PartWhat it does
the hash functionturns a name into a bucket number
the table of bucketseach bucket holds the entries whose names hashed to it
a chained overflow listfor names that collide, which they will
OperationWhat it costs
searchone hash, then the length of one chain
createone hash, a scan of that chain to check the name, then an insertion
list in orderbadly: the entries are in hash order, so listing sorted means reading everything and sorting it

The weakness is the fixed size of the table. The number of buckets is chosen when the directory is made; as the directory grows, the chains grow with it, and the search becomes linear again. The answer is to rehash into a larger table, which means recomputing every entry's bucket, so it is done rarely and it is expensive when it is.

The same directory, with 64 buckets

average chain = 1000 / 64 = 15.625

About 16 comparisons instead of 500, and usually one block read instead of eight, because a bucket and its chain are small enough to sit together.

Real systems do better again: ext4's directories are hashed B trees, which keep the fast lookup and also grow gracefully. The feature is recorded in the superblock, and it is in the image built in Chapter one hundred five:

$ mkfs.ext4 -q -F -b 1024 -I 128 -O ^has_journal dir.fs 4096 2>&1 | grep -v deprecated
Creating regular file dir.fs
$ dumpe2fs -h dir.fs 2>/dev/null | grep '^Filesystem features' | tr -s ' ' '\n' | grep -c dir_index
1

The 1 is the dir_index feature being present: this file system indexes its directories rather than scanning them.

A directory is a file that grows

The measurement that makes the whole chapter concrete. 120 files are created in the root directory of the image, and the directory's own size is read from its inode before and after.

$ printf 'x\n' > one.txt
$ debugfs -R "stat <2>" dir.fs 2>/dev/null | awk '/^User/ {print "the root directory is", $6, "bytes"}'
the root directory is 1024 bytes
$ { for i in $(seq 1 120); do printf 'write one.txt file-%03d\n' "$i"; done; } > script.debugfs
$ debugfs -w -f script.debugfs dir.fs > /dev/null 2>&1; echo done
done
$ debugfs -R "stat <2>" dir.fs 2>/dev/null | awk '/^User/ {print "after 120 files the root directory is", $6, "bytes"}'
after 120 files the root directory is 2048 bytes
$ debugfs -R "ls /" dir.fs 2>/dev/null | tr -s ' ' '\n' | grep -c '^file-'
120
munotes.in430

Directory Implementation

One block became two. The root directory is inode 2, as Chapter one hundred five showed, and it is an ordinary file: it was 1024 bytes, one block; 120 more names did not fit, so the file system gave it a second block and it is now 2048.

entries = 120 + 3 = 123

2048 / 123, about 17 bytes an entry

123 entries, counting ., .. and lost+found, in 2,048 bytes: about 17 bytes each, which is a name of a few characters, an inode number, and the lengths that go with them. The arithmetic of the earlier section was not a model of something else: it is this.

Distinctions that carry marks

Linear listHash table
Searchlinear: half the entries on averageone hash and a chain
Createa full search, then insertone hash, one chain, then insert
1,000 entries, 64 buckets500 comparisonsabout 16
Listing in ordereasyhard: hash order is not name order
Fixed sizenoyes, until it is rehashed
Programming difficultytrivialsome, and the hash function matters
Successful searchUnsuccessful search
Comparisons in a list of nn / 2 on averagen, always
Happens whenopening a filecreating one

What it does not mean

A directory is not a special kind of object. It is a file, and it grows in blocks like any other, as the measurement shows.

A hash table does not remove the search. It makes it short.

A linear list is not always wrong. Most directories hold a few dozen entries, where a scan of one block is cheaper than a hash.

Rehashing is not something the user sees. It is the file system's own work, and it is why the fixed table size is tolerable.

The 16 bytes an entry is not a rule. It depends on the length of the names, and the image measured about 17.

Quick revision

  • Every path resolution is a directory search, so the structure matters more than any other

choice in the file system.

  • A linear list of name and inode number: a search costs n / 2 on average and n when

the name is absent, so creating a file costs a full search.

  • Deleted entries are marked unused, filled by the last entry, or kept on a

free list. A sorted list allows a binary search but costs on insertion; real systems use a B tree.

munotes.in431

Directory Implementation

  • Worked: 1,000 entries of 16 bytes in 1024 byte blocks is 64 entries a block, about

16 blocks; a search averages 500 comparisons and about 8 block reads.

  • A hash table costs one hash and one chain: with 64 buckets, 1000 / 64 = 15.625,

about 16 comparisons. It lists badly in order, and its fixed size means rehashing as it grows.

  • ext4 records dir_index in the superblock: its directories are hashed B trees.
  • Measured: the root directory was 1024 bytes, and after 120 files it is 2048. With

., .. and lost+found that is 123 entries in two blocks, about 17 bytes each.

Test yourself

  1. Why does the directory structure matter so much? Because every path resolution searches

one directory per component, which makes it the most frequent operation in the file system.

  1. What does a linear list cost to search, and why is creating a file worse? Half the entries

on average for a successful search; creating a file must prove the name is absent, which means searching all of them.

  1. Give the three ways of handling a deleted entry in a linear list. Mark it unused, move the

last entry into the hole, or keep a free list of holes. 4. A directory holds 1,000 entries of about 16 bytes in 1024 byte blocks. How many blocks is it, and what does a search cost? 64 entries a block, so about 16 blocks; a successful search averages 500 comparisons and about 8 block reads.

  1. How does a hash table change the cost, with 64 buckets and 1,000 entries? The search is

one hash and one chain of 1000 / 64 = 15.625 entries, about 16 comparisons and usually one block read.

  1. Give two disadvantages of a hash table for a directory. Listing the entries in name order

needs a sort, and the table is a fixed size, so the chains lengthen as the directory grows until it is rehashed.

  1. What does ext4 do, and how can you tell? Its directories are hashed B trees; the

dir_index feature is recorded in the superblock and printed by dumpe2fs. 8. The root directory of an image was 1024 bytes and became 2048 after 120 files were added. What does that show? That a directory is an ordinary file: the names did not fit in one block, so the file system allocated a second one.

Contents This chapter on its own page

munotes.in432

Chapter One Hundred Seven

Contiguous and Linked Allocation

Syllabus topic Module 2, "File System Implementation - Allocation Methods"

In one line

Either a file's blocks are together and the inode holds a start and a length, or they are scattered and each one points at the next.

What has to be recorded

The inode of Chapter one hundred five holds pointers to the file's data blocks, and there are three ways to arrange them. Each one is judged by three questions.

QuestionWhy it matters
how much space does the record of the blocks takeit is in the inode, and the inode is small
what does direct access to block n costChapter one hundred needs it
what happens when the file growsmost files do

Contiguous allocation

Put the whole file in consecutive blocks. The inode then holds two numbers: the first block and the length.

Space in the inodetwo numbers, whatever the size of the file
Direct accessone access: physical = start + logical, and nothing else
Sequential accessthe best possible: the blocks are in order, so Chapter ninety two's sequential figure applies
Growththe file cannot grow unless the block after it is free
Free spacesuffers external fragmentation, which is Chapter seventy one's problem again

Worked

A file starting at block 19 with a length of 6 blocks. Where is its logical block 4?

physical = 19 + 4 = 23

That single addition is contiguous allocation, and it is why the method survives inside modern file systems as extents.

The two objections in full

External fragmentation. The free space becomes a set of holes, and a new file needs a hole big enough: first fit, best fit and worst fit reappear exactly as in Chapter seventy one, with the same answers. The cure is the same too: compaction, which for a disk means copying the whole volume, offline, and is what defragmenting a disk is.

The size must be known in advance. A program that says how large its file will be is asking the user to guess; guessing high wastes space, and guessing low means the file must be copied to a larger hole when it fills. Extents are the answer used in practice: the file is a list of contiguous runs, so it grows by adding another run rather than by moving.

Linked allocation

Each block holds a pointer to the next block of the file. The inode holds the first, and usually the last.

Space in the inodeone or two numbers
Free spaceno external fragmentation at all: any free block will do
Growthfree: take any block and link it on
Direct accessruinous: to read logical block n, read every block before it
Space in each blockthe pointer is part of the block, so the data in a block is not a power of two
Reliabilityone damaged pointer loses the rest of the file
munotes.in433

Contiguous and Linked Allocation

The sums

To read logical block 100:

accesses = 100 + 1 = 101

A hundred and one disk accesses to read one block, against one for contiguous allocation. At Chapter ninety two's nine milliseconds each that is most of a second for four kilobytes of data.

And the pointer's share of a 512 byte block with a 4 byte pointer:

overhead = 4 / 512

Four bytes in every 512, which is about 0.78 per cent, and the more awkward consequence is that a block holds 508 bytes of data. A program reading in powers of two now has every read straddling a block boundary, which is why systems that use linked allocation allocate in clusters of several blocks: fewer pointers, at the cost of more internal fragmentation.

FAT: the links taken out of the blocks

Keep the pointers in a table at the beginning of the volume instead of inside the blocks. One entry per block of the disk, holding the number of the next block of whatever file it belongs to. The name is the file allocation table, and the directory entry holds the number of the file's first block.

Linked, pointers in the blocksFAT
Where the pointers arein each data blockin one table at the start of the volume
A data block holds508 of 512 bytesall 512
Direct access, table not cachedn + 1 accesses2 accesses: the table, then the block
Direct access, table cachedn + 11
Costthe table itself must be held in memory, or every access reads it

FAT entries for 40 gigabytes in 4 kilobyte clusters = 10,485,760

table size at 4 bytes an entry = 10,485,760 × 4 = 41,943,040

Forty megabytes of table for a forty gigabyte volume, which has to be in memory for the scheme to be fast. That is the trade FAT makes, and it is why the cluster size grows with the volume.

Where the real file system put a file

ext4 does not use either method as described: it uses extents, which are contiguous runs. The image of Chapter one hundred five is asked where it put a hundred kilobyte file.

$ mkfs.ext4 -q -F -b 1024 -I 256 -O ^has_journal alloc.fs 8192 2>&1 | grep -v deprecated
Creating regular file alloc.fs
$ dd if=/dev/urandom of=solid.dat bs=1024 count=100 status=none
$ debugfs -w -R "write solid.dat solid.dat" alloc.fs 2>&1 | grep Allocated
Allocated inode: 12
$ debugfs -R "stat <12>" alloc.fs 2>/dev/null | sed -n '/EXTENTS/{n;p;}'
(0-1):80-81, (2-16):83-97, (17-99):611-693
munotes.in434

Contiguous and Linked Allocation

That one line is the chapter's summary, written by a real allocator.

The extentWhat it says
(0-1):80-81the file's logical blocks 0 and 1 are physical blocks 80 and 81
(2-16):83-97logical 2 to 16 are physical 83 to 97. Physical block 82 was not free, so the run had to break
(17-99):611-693logical 17 to 99 are physical 611 to 693: a run of 83 blocks, far away from the first two

A hundred block file in three pieces, on a file system four minutes old. The break at 82 and the jump to 611 are external fragmentation happening in front of us, and the extent list is the compromise: contiguous where it can be, a new run where it cannot.

And direct access still works by addition, one run at a time. Logical block 50 is in the third run, which starts at logical 17 and physical 611:

physical = 611 + 50 - 17 = 644

That is the file organisation module of Chapter one hundred four doing its one job, with the numbers this machine actually wrote.

Distinctions that carry marks

ContiguousLinkedFAT
The inode holdsstart and lengththe first blockthe first block
External fragmentationyesnono
Direct access to block n1 accessn + 11 if the table is cached, else 2
Growthhard: needs the next block freefreefree
Space lostto holes between filesto a pointer in every blockto the table
Reliabilitygoodone lost pointer loses the tailthe table can be duplicated
Compaction of memory, Chapter seventy oneDefragmenting a disk
What movesprocessesfiles
How longmicrosecondshours, for a large volume
Done while runningpossibleusually offline

What it does not mean

Contiguous allocation is not obsolete. It survives as the extent, which is the unit modern file systems allocate in.

Linked allocation is not slow for sequential reading. It is slow for direct access; read in order, it is one block after another, though the blocks may be far apart on the disk.

FAT is not a different allocation method. It is linked allocation with the links moved out of the data blocks.

A pointer in every block does not cost only its four bytes. It makes the usable size of a block not a power of two, which is worse than the space.

Extents do not remove fragmentation. The measurement above shows a file in three runs on a nearly empty file system.

Quick revision

  • Contiguous: the inode holds a start and a length; direct access is start + logical,
munotes.in435

Contiguous and Linked Allocation

one access; sequential access is ideal. It suffers external fragmentation and the file cannot grow; the cures are compaction and extents.

  • Worked: a file at block 19, logical block 4 is physical 19 + 4 = 23.
  • Linked: each block points at the next. No external fragmentation and free growth.

Direct access to block 100 costs 101 accesses; the pointer takes 4 of 512 bytes, leaving 508; one lost pointer loses the rest.

  • Clusters cut the pointer cost and raise internal fragmentation.
  • FAT moves the links into one table at the start of the volume: a data block holds all 512

bytes, and direct access is one access if the table is cached. A 40 gigabyte volume with 4 kilobyte clusters needs 10,485,760 × 4 = 41,943,040 bytes of table.

  • Measured: ext4 put a 100 block file in three extents,

(0-1):80-81, (2-16):83-97, (17-99):611-693. Physical block 82 was taken, so the run broke: external fragmentation on a new file system.

  • Logical block 50 of that file is physical 611 + 50 - 17 = 644.

Test yourself

  1. What does an inode hold under contiguous allocation, and what does direct access cost? The

first block and the length; direct access to logical block n is one addition and one access.

  1. A file starts at block 19. Where is its logical block 4? Physical block 19 + 4 = 23.
  2. Give the two objections to contiguous allocation and the cure for each. External

fragmentation, cured by compaction, which for a disk means defragmenting it offline; and the need to know the size in advance, worked around by extents.

  1. What are the advantages of linked allocation? No external fragmentation, since any free

block will do, and free growth. 5. How many accesses does linked allocation need to read logical block 100, and how many does contiguous need? 101 against 1.

  1. A 512 byte block holds a 4 byte pointer. State the two costs. About 0.78 per cent of the

space, and, more awkwardly, only 508 bytes of data in a block, so reads in powers of two straddle block boundaries.

  1. What does FAT change, and what does it cost? It moves the links out of the data blocks

into one table at the start of the volume, so a block holds all its bytes and direct access is one access when the table is cached; the cost is holding the table, 40 megabytes for a 40 gigabyte volume with 4 kilobyte clusters. 8. A real file system put a 100 block file in three extents, with a break at physical block 82. What does that demonstrate? External fragmentation: block 82 was already taken, so the contiguous run had to end and a new one begin elsewhere.

munotes.in436

Contiguous and Linked Allocation

  1. Logical block 50 of a file whose third extent is (17-99):611-693 is which physical block?

611 + 50 - 17 = 644.

Contents This chapter on its own page

munotes.in437

Chapter One Hundred Eight

Indexed Allocation, and What a Real File System Does

Syllabus topic Module 2, "File System Implementation - Allocation Methods"

In one line

Put all of a file's block pointers in a block of their own, and when one block of pointers is not enough, point at blocks of pointers.

Indexed allocation

Give each file an index block, holding the pointers to all its data blocks. The inode points at the index block.

Direct accesstwo accesses: the index block, then the data block. One if the index is already in memory
External fragmentationnone: the data blocks may be anywhere
Growthfree, while the index block has room
Spacea whole block of pointers for every file, however small
Size limita file can have no more blocks than the index block has pointers

The last two objections pull in opposite directions, and that tension is the rest of the chapter. A small index block wastes little and limits the file; a large one allows a large file and wastes a block on every tiny one.

With 1024 byte blocks and 4 byte pointers:

pointers in one index block = 1024 / 4 = 256

largest file = 256 × 1024 = 262,144

256 kilobytes, which is not a file size anybody would accept as a limit.

Making the index bigger

Three schemes, and MU's textbook names all three.

SchemeHow it worksThe largest file
linked index blocksthe last pointer of an index block names the next index blockunlimited, but finding a late block means walking the chain of index blocks
multilevel indexa first level index block points at second level index blocks, each of which points at datawith 256 pointers a block, 256 × 256, which is 65,536 blocks, which is 64 megabytes
combined schemea few direct pointers for small files, then single, double and triple indirect pointersthe UNIX answer, sized in the next section

two level index = 256 × 256 = 65,536

in bytes = 65,536 × 1024 = 67,108,864

Sixty four megabytes, at the cost of three accesses for a direct read: first level, second level, data.

The combined scheme, which is the UNIX inode

The scheme to be able to draw and to size. An inode holds fifteen pointers: 12 direct, then one single indirect, one double indirect and one triple indirect.

With 4096 byte blocks and 4 byte pointers, an index block holds:

pointers per block = 4096 / 4 = 1024

PointerReachesBytesAccesses for a direct read
12 direct12 blocks12 × 4096 = 49,152, 48 kilobytes1
single indirect1024 blocks1024 × 4096 = 4,194,304, 4 megabytes2
double indirect1024 × 1024 blocks4,294,967,296, 4 gigabytes3
triple indirect1024 × 1024 × 1024 blocks4,398,046,511,104, 4 terabytes4
munotes.in438

Indexed Allocation, and What a Real File System Does

largest file = 49,152 + 4,194,304 + 4,294,967,296 + 4,398,046,511,104 = 4,402,345,721,856

About four terabytes, and the shape of the scheme is the point: a small file needs no index block at all, because its blocks fit in the twelve direct pointers, and a file of forty eight kilobytes or less is read in one access per block. The cost rises only for the parts of a file that are large.

Which pointer holds logical block 5,000

The standard sum. The ranges first:

Logical blocksReached by
0 to 11the direct pointers
12 to 1035the single indirect, since 12 + 1024 - 1 = 1035
1036 to 1,049,611the double indirect, since 1036 + 1024 × 1024 - 1 = 1,049,611
beyond thatthe triple indirect

Block 5,000 is in the double indirect range. Its position inside it:

offset = 5000 - 1036 = 3964

3964 = 3 × 1024 + 892

Entry 3 of the double indirect block, then entry 892 of the index block it names, then the data block: three accesses, and the arithmetic is a subtraction and a division.

Note the direction of the sum, because it is the mistake to avoid: subtract the start of the range first, then divide. Dividing 5,000 by 1024 without subtracting gives the wrong entry.

What a real file system does

ext4 does not keep a list of block pointers at all. It keeps extents: a run of consecutive blocks recorded as a start and a length, exactly Chapter one hundred seven's contiguous allocation, in a tree when there are many of them.

Block pointersExtents
Recordsone pointer per blockone entry per run
A 100 megabyte file in one run needs25,600 pointersone entry
A badly fragmented filethe same 25,600 pointersone entry per run, in a tree
Direct accessindex levelswalk the extent tree, then add

The four inode pointers of the combined scheme are still the right thing to learn, because they are what the paper asks for and because the arithmetic is the same idea: the cost of reaching a block rises with how far into the file it is.

A file that occupies no blocks at all

One last measurement, and it shows what a pointer of zero means. Two files of a hundred kilobytes are written into the image: one of random bytes, one of zeros.

$ mkfs.ext4 -q -F -b 1024 -I 256 -O ^has_journal sparse.fs 8192 2>&1 | grep -v deprecated
Creating regular file sparse.fs
$ dd if=/dev/urandom of=solid.dat bs=1024 count=100 status=none
$ dd if=/dev/zero of=holey.dat bs=1024 count=100 status=none
$ debugfs -w -R "write solid.dat solid.dat" sparse.fs > /dev/null 2>&1; debugfs -w -R "write holey.dat holey.dat" sparse.fs > /dev/null 2>&1; echo both written
both written
$ debugfs -R "stat <12>" sparse.fs 2>/dev/null | awk '/Size:/ {print "the file of random bytes is", $NF, "bytes"; exit}'
the file of random bytes is 102400 bytes
$ debugfs -R "stat <12>" sparse.fs 2>/dev/null | awk '/^Links/ {print "and it holds", $4, "units of 512 bytes"}'
and it holds 200 units of 512 bytes
$ debugfs -R "stat <13>" sparse.fs 2>/dev/null | awk '/Size:/ {print "the file of zeros is", $NF, "bytes"; exit}'
the file of zeros is 102400 bytes
$ debugfs -R "stat <13>" sparse.fs 2>/dev/null | awk '/^Links/ {print "and it holds", $4, "units of 512 bytes"}'
and it holds 0 units of 512 bytes
munotes.in439

Indexed Allocation, and What a Real File System Does

Two files of 102,400 bytes: one occupies 200 units of 512 bytes, which is the hundred kilobytes, and the other occupies none.

The file of random bytesThe file of zeros
size102,400102,400
blocks used200 units of 5120

A file with no blocks that is a hundred kilobytes long is called a sparse file, and the hole is a pointer of zero. A read of a hole returns zeros without any disk access; a write into a hole allocates a block then. It follows directly from indexed allocation: a pointer can say "no block", and nothing in the scheme has to change. It is how a database file of a terabyte can sit on a disk of a gigabyte, and it is why the size of a file and the space it uses are two different numbers, as Chapter ninety eight first showed with six bytes in 4,096.

Distinctions that carry marks

ContiguousLinkedIndexed
Where the pointers arenowhere: a start and a lengthin the data blocksin an index block
External fragmentationyesnono
Direct access to block n1n + 12, or one per level
Space wasted on a one block filenoneone pointera whole index block
Growthhardfreefree, within the index
Single indirectDouble indirectTriple indirect
Blocks reached, 4 KB blocks10241,048,5761,073,741,824
Bytes4 MB4 GB4 TB
Accesses234

What it does not mean

An index block is not the inode. The inode points at it; on UNIX the first fifteen pointers are in the inode itself.

The combined scheme is not three schemes. It is one, with the cost of reaching a block rising with its position.

Indexed allocation does not give constant time access. It gives one access per level, and the deepest level is four.

A sparse file is not compressed. Nothing is encoded: the pointers simply say there is no block.

munotes.in440

Indexed Allocation, and What a Real File System Does

Extents are not a fourth method. They are contiguous allocation applied run by run, recorded in a tree.

Quick revision

  • Indexed allocation: an index block of pointers, so no external fragmentation, free

growth, and direct access in two accesses. It wastes a whole block on a small file and limits the file to the block's pointer count.

  • 1024 byte blocks and 4 byte pointers give 256 pointers, so one index block limits a file to

256 × 1024 = 262,144 bytes.

  • Bigger indexes: linked index blocks, a multilevel index (256 × 256, which is

65,536 blocks, 64 megabytes, three accesses), or the combined scheme.

  • The UNIX inode: 12 direct, single, double and triple indirect. With 4096

byte blocks and 4 byte pointers, 1024 pointers a block: 48 kilobytes, 4 megabytes, 4 gigabytes and 4 terabytes, at 1, 2, 3 and 4 accesses.

  • Ranges: 0 to 11 direct, 12 to 1035 single indirect, 1036 to 1,049,611 double

indirect.

  • Block 5,000: 5000 - 1036 = 3964, then dividing 3964 by 1024 gives 3 remainder 892,

so entry 3 then entry 892. Subtract the start of the range before dividing.

  • ext4 uses extents, one entry per run of consecutive blocks, in a tree.
  • Measured: two files of 102,400 bytes, one holding 200 units of 512 bytes and the other

0. A sparse file's hole is a pointer of zero: a read returns zeros with no disk access.

Test yourself

  1. What does indexed allocation keep, and what does direct access cost? An index block

holding the pointers to all the file's data blocks; direct access costs two accesses, the index and the data, or one if the index is cached.

  1. Give the two objections to a single index block. A whole block of pointers is used even by

a tiny file, and the file can have no more blocks than the index block has pointers.

  1. With 1024 byte blocks and 4 byte pointers, how large a file can one index block describe?

256 pointers, so 256 × 1024 = 262,144 bytes. 4. Name the three ways of enlarging the index, and the largest file a two level index gives with 256 pointers a block. Linked index blocks, a multilevel index, and the combined scheme; two levels give 256 × 256, which is 65,536 blocks, 64 megabytes. 5. Give the four parts of the UNIX combined scheme and what each reaches with 4 kilobyte blocks. Twelve direct pointers reaching 48 kilobytes; a single indirect reaching 4 megabytes; a double indirect reaching 4 gigabytes; a triple indirect reaching 4 terabytes.

munotes.in441

Indexed Allocation, and What a Real File System Does

  1. How many accesses does a read from the double indirect part need? Three: the double

indirect block, the index block it names, and the data block.

  1. Which pointer reaches logical block 5,000, and which entries? The double indirect:

5000 - 1036 = 3964, and dividing 3964 by 1024 gives 3 remainder 892, so entry 3 of the double indirect block and entry 892 of the index block it names.

  1. What does ext4 keep instead of block pointers? Extents: one entry for each run of

consecutive blocks, held in a tree when there are many.

  1. Two files are both 102,400 bytes and one occupies no blocks. Explain. The second is

sparse: its pointers say there is no block, so a read of those parts returns zeros without any disk access and a write allocates a block at that moment.

Contents This chapter on its own page

munotes.in442

Chapter One Hundred Nine

Free Space Management

Syllabus topic Module 2, "File System Implementation - Free-Space Management"

In one line

The file system keeps a free space list, and the choice is between a bit for every block, a chain through the free blocks themselves, and two schemes that record runs.

Why there has to be a list

Allocation needs a free block, and deletion returns one. The list of free blocks is consulted on every create and every extension of every file, so it must be quick to search, quick to update, and small enough to keep in memory.

And it must survive a crash consistently. A block that is in a file and on the free list will be handed out twice and two files will share it, which destroys data; a block in neither is merely lost. The first is a catastrophe and the second is a leak, which is why a checking program repairs the free list towards losing space rather than sharing it.

The bit vector

One bit for every block on the disk: 1 if free, 0 if in use. This is what nearly every modern file system uses.

Finding a free blockscan for the first 1 bit, which processors do a word at a time
Finding n consecutive free blocksscan for n consecutive 1 bits: possible, and this is the method's great advantage
Updatingflip one bit
Spaceone bit per block of the whole disk, and it must be in memory to be fast

The first free block, as a formula

Scanning word by word, the answer a question wants:

block number = (number of whole words that were all zeroes) × bits per word + offset of the first 1 bit

Worked, with 32 bit words, where the first three words are zero and the next has its first 1 bit at offset 22:

block = 3 × 32 + 22 = 118

How large is the bitmap

bits = disk size / block size

bytes = bits / 8

DiskBlock sizeBlocksBitmap
1 gigabyte4096262,14432,768 bytes, 32 kilobytes
40 gigabytes409610,485,7601,310,720 bytes, 1.25 megabytes
1 terabyte4096268,435,45633,554,432 bytes, 32 megabytes

A terabyte needs 32 megabytes of bitmap, which is why a large file system does not keep all of it resident: it is divided into groups, each with its own bitmap, and only the groups in use are read. That is the next section, and the image measures it.

A linked list of free blocks

The superblock points at the first free block; each free block holds the number of the next.

Spacenone wasted: the pointers are inside blocks that are free anyway
Finding one free blocktake the first: one access, and the superblock must be rewritten
Finding n consecutive blockshopeless: the list is in no order, and walking it means reading one block per step
Traversalreading the whole list is one disk access per free block
munotes.in443

Free Space Management

The free list is never traversed in normal use, and the method survives on that: a file system only ever wants the first block. It is the scheme FAT used, where the table of Chapter one hundred seven already held the links.

Grouping

The first free block holds the addresses of the next n free blocks, and the last of those addresses points at another block of addresses.

Finding n free blocksone access gives a whole block of addresses
Spacestill free blocks holding the information
Compared with the plain listn times fewer accesses to collect n blocks

Counting

Keep pairs: the first block of a run, and how many blocks the run has.

The reason it works is the observation that space is freed and allocated in runs, because files are allocated contiguously where possible: a file of 500 blocks that is deleted returns 500 consecutive blocks, and one pair records them all.

Size of the listmuch shorter: one entry per run rather than per block
Fitsextents exactly, which is what Chapter one hundred eight showed a real file system using
Where it is kepta balanced tree, so that a run of the right size is found quickly

Space maps, the scheme MU's textbook names from ZFS, are this idea taken further: each group of blocks has a log of allocations and frees, which is replayed into a tree in memory when the group is first used, and condensed when it grows too long. Writing a log is sequential and writing a bitmap is not, which on a disk is the difference Chapter ninety two measured.

The real file system's free space

The image of Chapter one hundred five, asked where its bitmaps are and what they say.

$ mkfs.ext4 -q -F -b 1024 -I 128 -O ^has_journal free.fs 8192 2>&1 | grep -v deprecated
Creating regular file free.fs
$ dumpe2fs free.fs 2>/dev/null | grep -E '^(Block size|Blocks per group|Inodes per group)' | tr '\t' ' '
Block size:               1024
Blocks per group:         8192
Inodes per group:         2048
$ dumpe2fs free.fs 2>/dev/null | sed -n '/^Group 0/,/^Group 1/p' | grep -E 'itmap at|Free blocks:|Free inodes:' | sed 's/, csum.*//' | head -4
  Block bitmap at 66 (+65)
  Inode bitmap at 82 (+81)
  Free blocks: 80-81, 83-97, 355-8191
  Free inodes: 12-2048
$ echo "one block of 1024 bytes holds $(( 1024 * 8 )) bits, and there are 8192 blocks in a group"
one block of 1024 bytes holds 8192 bits, and there are 8192 blocks in a group
munotes.in444

Free Space Management

Four things, and the third is the one to remember.

The image saysWhat it means
Block bitmap at 66, Inode bitmap at 82the free space list is a bit vector, and there are two: one for blocks and one for inodes. Each is one block
Blocks per group 8192, and one 1024 byte block holds 8192 bitsthe group is exactly as large as one bitmap block can track. That is why block groups exist and why they are that size: 8 × 1024 = 8192
Free blocks: 80-81, 83-97, 355-8191the free space is printed as runs, which is the counting method of this chapter used as a way of reporting: three entries instead of 7,854 block numbers
Free inodes: 12-2048inodes are tracked the same way, and 1 to 11 are the reserved ones of Chapter one hundred five

And the list is kept up to date

$ dumpe2fs -h free.fs 2>/dev/null | awk '/^Free blocks/ {print "free blocks to begin with:", $3}'
free blocks to begin with: 7854
$ dd if=/dev/urandom of=chunk.dat bs=1024 count=500 status=none
$ debugfs -w -R "write chunk.dat chunk.dat" free.fs > /dev/null 2>&1
$ dumpe2fs -h free.fs 2>/dev/null | awk '/^Free blocks/ {print "free blocks after a 500 block file:", $3}'
free blocks after a 500 block file: 7354
$ echo "the difference is $(( 7854 - 7354 )) blocks"
the difference is 500 blocks

Five hundred blocks of file cost exactly five hundred blocks of free space, and the count in the superblock was adjusted as they were taken. No extra block was needed for the file's own bookkeeping, because its extents fitted in the inode: Chapter one hundred eight's arithmetic, confirmed by subtraction.

Distinctions that carry marks

Bit vectorLinked listGroupingCounting
Space usedone bit a block, in memorynone: inside free blocksnoneone entry a run
One free blockscan for a 1 bitthe first one, immediatelythe first address in the blockthe first entry
n consecutive free blockseasy: n consecutive 1 bitshopelessn addresses, but not consecutiveeasy: a run of length n
Size of the structurefixed by the diskgrows with the free spaceas the listgrows with the fragmentation
Used bymost modern file systemsFATZFS and extent based systems
A block in a file and on the free listA block in neither
What happensit is handed out twice and two files share itit is lost until the file system is checked
Severitydata is destroyedspace is wasted
So a repair program prefersthis
munotes.in445

Free Space Management

What it does not mean

The free list is not a list of files. It is a list of blocks nobody owns.

A bitmap is not searched bit by bit. It is searched a machine word at a time, which is why the formula multiplies by the bits in a word.

Grouping is not the same as counting. Grouping records addresses in bulk; counting records runs.

The counting method is not shorter on a fragmented disk. Its length grows with the number of runs, so a badly fragmented file system has a long list.

The free count in the superblock is not the free list. It is a total, kept beside the list so that the size of the file system can be reported without reading it.

Quick revision

  • The free list is consulted on every create and every extension, and after a crash it must

not claim a block that is also in a file: a shared block destroys data, a lost block only wastes space.

  • Bit vector: one bit a block, 1 for free. Easy to find one free block

and n consecutive ones. Must be in memory: 32 kilobytes for a gigabyte, 1.25 megabytes for 40 gigabytes, 32 megabytes for a terabyte.

  • The first free block is (whole zero words) × bits per word + offset of the first 1 bit;

with 32 bit words and three zero words, 3 × 32 + 22 = 118.

  • Linked list: the superblock points at the first free block, each free block at the next. No

space wasted; finding n consecutive blocks is hopeless and a traversal costs one access a block.

  • Grouping: one free block holds n addresses and the last points at another such block,

so n blocks are found in one access.

  • Counting: first block and length pairs, because space is freed in runs; a short list,

and it matches extents. Space maps log allocations and frees instead, because a log is written sequentially.

  • Measured: the image keeps a block bitmap and an inode bitmap, each one block;

8192 blocks a group because 8 × 1024 = 8192 bits fit in one 1024 byte block; and the free space is reported as runs.

  • Measured: a 500 block file reduced the free count from 7854 to 7354, exactly 500.

Test yourself

  1. Why must the free space list be consistent after a crash, and which error is worse?

Because a block recorded as both free and in a file will be given to a second file and the data destroyed; a block in neither is only lost space, so a repair program prefers to lose space.

munotes.in446

Free Space Management

  1. Describe the bit vector and its main advantage. One bit per block, set if the block is

free; whole words can be scanned at once, and consecutive free blocks are easy to find, which the other methods cannot do.

  1. How large is the bitmap for a 40 gigabyte disk with 4 kilobyte blocks? 10,485,760 blocks,

so 10,485,760 bits, which is 1,310,720 bytes or about 1.25 megabytes. 4. Give the formula for the first free block when scanning a bitmap by words, and work it for three zero words of 32 bits and a first set bit at offset 22. The number of whole zero words times the bits per word, plus the offset of the first set bit: 3 × 32 + 22 = 118.

  1. Describe the linked list method and its weakness. The superblock points at the first free

block and each free block holds the number of the next; it wastes no space, but finding several consecutive free blocks is impractical and walking the list costs one disk access per block.

  1. How does grouping improve on the linked list? The first free block holds the addresses of

n free blocks, with the last pointing at another such block, so one access yields many blocks.

  1. What does the counting method record, and why is it short? Pairs of a first block and a

run length, because blocks are allocated and freed in runs, so one entry covers many blocks.

  1. A file system has 1024 byte blocks and 8192 blocks in a group. Why that number? Because

one block of 1024 bytes holds 8 × 1024 = 8192 bits, so a group is exactly what one bitmap block can track.

  1. What did the free block count do when a 500 block file was written? It fell from 7,854 to

7,354, exactly 500, with no extra block needed for the file's own bookkeeping.

Contents This chapter on its own page

munotes.in447

Chapter One Hundred Ten

Designing a Small File System

Syllabus topic Module 2, the course outcome MU prints as "Design file system"

In one line

Choose the block size, count the inodes, size the bitmaps and the inode table, lay them out in order, and the rest of the volume is for files.

The brief

Design a file system for a volume of four megabytes. MU's course outcome for this paper is to design a file system, and this is what that means in practice: ten decisions, each with its arithmetic, and a layout that adds up.

The design sheet

One: the block size

ChoiceEffect
a small block, 512 byteslittle internal fragmentation, more blocks to track, more pointers per file
a large block, 4096 bytesfewer blocks and pointers, more waste in the last block of every file

1024 bytes is chosen, because the volume is small and the files on it are expected to be small: Chapter ninety eight measured six bytes of data in a 4,096 byte block and 4,090 bytes wasted, and a 1024 byte block wastes a quarter of that.

blocks = 4,194,304 / 1024 = 4096

Two: how many inodes

The inode count is fixed when the file system is made and cannot be changed afterwards, so it is the decision to get right. Too few and the volume fills with space left but no inodes; too many and the inode table wastes the volume.

The usual rule is one inode for every 4 kilobytes of volume, on the assumption that the average file is a few kilobytes.

inodes = 4,194,304 / 4096 = 1024

Three: the inode size, and the table

128 bytes an inode is enough for the attributes of Chapter one hundred five and fifteen block pointers.

inode table = 1024 × 128 = 131,072

in blocks = 131,072 / 1024 = 128

128 blocks of the 4,096 are the inode table, and that is the largest single piece of the design.

Four: the free space method

A bit vector for the blocks and another for the inodes, because Chapter one hundred nine showed it is the only method that finds consecutive free blocks easily.

block bitmap = 4096 / 8 = 512

inode bitmap = 1024 / 8 = 128

512 bytes and 128 bytes, so one block each, with room to spare. A volume whose bitmap needed more than one block would be divided into groups, as Chapter one hundred nine measured on the real file system.

Five: the directory

A linear list of entries, each a name and an inode number, as in Chapter one hundred six. The volume is small, so a directory will be one or two blocks and a scan of it is one or two accesses. A hash table would cost more than it saves at this size.

munotes.in448

Designing a Small File System

Six: the allocation method

The combined scheme of Chapter one hundred eight: twelve direct pointers, then single, double and triple indirect. With 1024 byte blocks and 4 byte pointers there are 256 pointers in an index block, so the scheme reaches:

pointer bound = 12 + 256 + 65,536 + 16,777,216 = 16,843,020

Sixteen million blocks, which is four thousand times the volume. So the largest file is not limited by the pointers at all: it is limited by the volume, and a design should say which limit binds.

Seven: the layout

The map, in block order, which is what a question asking for a design wants to see.

BlockWhat is in it
0the boot block
1the superblock: block size, block count, inode count, free counts
2the block bitmap, 4,096 bits
3the inode bitmap, 1,024 bits
4 to 131the inode table, 128 blocks
132the root directory, inode 2
133 to 4095data blocks

Eight: the overhead

metadata = 1 + 1 + 1 + 1 + 128 = 132

with the root directory = 133

usable = 4096 - 133 = 3963

usable bytes = 3963 × 1024 = 4,058,112

Blocks
the volume4096
metadata and the root directory133about 3.2 per cent
available for files39634,058,112 bytes

Nine: the largest file

From the two bounds: the pointers allow 16,843,020 blocks and the volume has 3,963 free. The largest file is 3,963 blocks, and even that only if it is the only file. A design states the binding limit rather than the impressive one.

Ten: what a crash does

The part a design is judged on. After a crash the structures may disagree, and the checking program must be able to repair them:

What can be wrongHow it is foundHow it is repaired
a block is in a file and free in the bitmapwalk every inode, build a bitmap, compareclear the bit: the file keeps the block
a block is in no file and not freethe same walkset the bit: the space comes back
an inode's link count is wrongcount the directory entries pointing at itwrite the true count
a file has no name in any directorythe walk finds an inode with a positive link count and no entryput it in lost+found

That last row is what the lost+found directory of Chapter one hundred five is for, and it is why every file system has one: somewhere to put a file whose name was lost.

Building it, and checking every number

The design is now built with the same parameters, by a professional implementation, and the predictions are compared with what it did.

munotes.in449

Designing a Small File System

$ mkfs.ext4 -q -F -b 1024 -I 128 -O ^has_journal design.fs 4096 2>&1 | grep -v deprecated
Creating regular file design.fs
$ dumpe2fs design.fs 2>/dev/null | grep -E '^(Block count|Inode count|Inode size|Block size|Free blocks|Overhead|Reserved GDT)' | sed 's/, csum.*//' | tr '\t' ' '
Inode count:              1024
Block count:              4096
Overhead clusters:        164
Free blocks:              3918
Block size:               1024
Reserved GDT blocks:      31
Inode size:           128
$ dumpe2fs design.fs 2>/dev/null | sed -n '/^Group 0/,/^$/p' | sed -E 's/,? csum.*//' | head -7
Group 0: (Blocks 1-4095)
  Primary superblock at 1, Group descriptors at 2-2
  Reserved GDT blocks at 3-33
  Block bitmap at 34 (+33)
  Inode bitmap at 50 (+49)
  Inode table at 66-193 (+65)
  3918 free blocks, 1013 free inodes, 2 directories, 1013 unused inodes

Now the comparison, line by line.

The design predictedThe real file system didAgreed
4096 blocks of 1024 bytesBlock count 4096, Block size 1024yes
1024 inodes of 128 bytesInode count 1024, Inode size 128yes
an inode table of 128 blocksInode table at 66-193, which is 128 blocksyes
a block bitmap of one blockBlock bitmap at 34yes
an inode bitmap of one blockInode bitmap at 50yes
block 0 outside the file system, for bootingGroup 0: (Blocks 1-4095), so block 0 is not in ityes
metadata of 132 blocksOverhead clusters: 164no: 32 more

The thirty two blocks the design did not have

The whole point of building it. The difference is not a mistake in the arithmetic: it is two features the design does not provide.

132 + 1 + 31 = 164

The extra blocksWhat they are for
1 block of group descriptorsext4 divides a volume into groups and describes each one. A four megabyte volume has one group, and the descriptor block is there anyway, because the format is the same at every size
31 reserved GDT blocksroom to grow the file system later without moving anything: more groups will need more descriptors, and the space for them is reserved now

132 plus 1 plus 31 is exactly 164, which is the number the implementation printed. The design was right; it was answering a smaller question.

And the blocks actually used

used = 4096 - 3918 = 178

164 + 1 + 12 = 177

The used blocksCount
overhead, as above164
the root directory, one block1
lost+found, which ext4 creates at 12,288 bytes12
accounted for177
actually used178

One block of the 178 is not accounted for by name. It is ext4's own padding between its bitmaps, which the group listing shows as the gaps at blocks 35 to 47 and 51 to 65. The honest form of this exercise is to say so: a design reconciles to within a block of somebody else's implementation, and the unexplained block is a difference in their format, not an error in the arithmetic.

munotes.in450

Designing a Small File System

The design in one table

What to write in an answer that asks for a file system design.

DecisionChoiceBecause
block size1024 bytessmall files, so little internal fragmentation
blocks40964,194,304 / 1024
inodes1024, one per 4 kilobytesfixed for ever at format time
inode size128 bytesthe attributes and fifteen pointers
inode table128 blocks1024 × 128 / 1024
free spacetwo bit vectors, one block eachthe only method that finds consecutive blocks easily
directorylinear list of name and inodeone or two blocks at this size
allocationcombined scheme, 12 direct and three indirectsmall files need no index block
layoutboot, superblock, bitmaps, inode table, root, datametadata first, so it is found without searching
overhead133 blocks, about 3.2 per cent
largest file3963 blocks, the volumethe pointers allow 16,843,020
after a crashwalk every inode, rebuild the bitmaps, fix link counts, put nameless files in lost+found

What it does not mean

The inode count is not adjustable later. That is why the rule of thumb matters.

The layout order is not arbitrary. The superblock must be at a known block, because nothing can be found before it is read.

A design is not judged by the largest file it allows. It is judged by whether the arithmetic adds up and the binding limit is stated.

The 32 extra blocks are not waste. They buy the ability to grow the file system, which this design cannot do at all.

Reconciling to within one block is not a failure. It is what comparing two formats looks like, and the unexplained block is named as unexplained.

Quick revision

  • Block size 1024, so 4,194,304 / 1024 = 4096 blocks.
  • Inodes: one per 4 kilobytes, so 1024, and the count is fixed at format time.
  • Inode table: 1024 × 128 = 131,072 bytes, which is 128 blocks, the largest single piece.
  • Bitmaps: 4096 / 8 = 512 bytes for blocks and 1024 / 8 = 128 for inodes,

one block each.

  • Directory: a linear list; allocation: the combined scheme, whose pointers reach

16,843,020 blocks, so the volume is the binding limit.

  • Layout: 0 boot, 1 superblock, 2 block bitmap, 3 inode bitmap, 4 to 131 inode table, 132
munotes.in451

Designing a Small File System

root, 133 to 4095 data.

  • Overhead 133 blocks, about 3.2 per cent; 3963 blocks, 4,058,112 bytes, for

files.

  • After a crash: rebuild the bitmaps from a walk of every inode, fix link counts, and put

nameless files in lost+found.

  • Built for real: every prediction agreed except the overhead, and 132 + 1 + 31 = 164

reconciles it exactly: one block of group descriptors and 31 reserved for growing the file system.

  • 164 + 1 + 12 = 177 of the 178 blocks used are accounted for, and the last one is ext4's

padding, named as unexplained rather than argued away.

Test yourself

  1. A volume of four megabytes with 1024 byte blocks. How many blocks?

4,194,304 / 1024 = 4,096.

  1. How many inodes would you give it, and why does the decision matter so much? About one per

four kilobytes, so 1,024; the count is fixed when the file system is made and cannot be changed, so too few wastes the volume's space and too many wastes the volume.

  1. How large is an inode table of 1,024 inodes of 128 bytes, in blocks of 1024? 131,072

bytes, which is 128 blocks.

  1. Size the two bitmaps. 4,096 bits for the blocks, 512 bytes; 1,024 bits for the inodes, 128

bytes. One block each.

  1. Give the layout in block order. Boot block, superblock, block bitmap, inode bitmap, the

128 block inode table, the root directory, then the data blocks.

  1. What is the overhead of this design, and how much is left for files? 133 blocks, about 3.2

per cent; 3,963 blocks or 4,058,112 bytes.

  1. Which limit decides the largest file here, and what is the other one? The volume, at 3,963

blocks; the combined scheme's pointers would allow 16,843,020 blocks. 8. A real ext4 file system with the same parameters reported 164 overhead blocks against the design's 132. Account for the difference. One block of group descriptors and 31 reserved descriptor blocks, which let the file system be grown later: 132 + 1 + 31 = 164.

  1. What must a checking program do after a crash? Walk every inode to rebuild both bitmaps,

correct link counts from the directory entries, and place any file with a positive link count and no directory entry into lost+found.

Contents This chapter on its own page

munotes.in452

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!