The Memory Management Unit
Chapter Sixty-Eight
Syllabus topic Module 2, "Memory Management - MMU"
Pages 266 to 269 of 452
In one line
The memory management unit is the hardware that turns every logical address the processor produces into a physical address, on every access, before the memory ever sees it.
Where it sits
The processor produces an address. The memory takes an address. The memory management unit is between them, and nothing gets past it.
| In order | What happens |
|---|---|
| 1 | the instruction produces a logical address |
| 2 | the memory management unit checks it and translates it |
| 3 | the physical address goes on the address bus |
| 4 | the memory answers |
It is hardware, and it is not part of the operating system. The operating system programs it, by loading its registers during a context switch, and then it works unaided on every one of the millions of accesses a process makes per second. A question that says the operating system translates each address is wrong: it could not possibly, because it is not running while the process is.
And it is not optional. Remove it and there is no protection, no relocation and no virtual memory, because all three are the same mechanism.
Its simplest form: the relocation register
The base and limit of Chapter sixty six is a memory management unit, and the simplest one there is.
if address < limit then physical = base + address else trap
In that form the base register is called the relocation register, because loading a different value into it relocates the whole process without touching a byte of it.
Worked, and the arithmetic written out
Base 14000, limit 100000. A process refers to addresses 0, 346, and 100000.
| Logical | Test | Physical |
|---|---|---|
| 0 | 0 is below 100000 | 14000 + 0 = 14000 |
| 346 | 346 is below 100000 | 14000 + 346 = 14346 |
| 99999 | just below | 14000 + 99999 = 113999 |
| 100000 | not below | trap |
Now change the relocation register to 90000 and run the same process.
| Logical | Physical |
|---|---|
| 0 | 90000 + 0 = 90000 |
| 346 | 90000 + 346 = 90346 |
Nothing in the program changed. That is relocation, and it is one register.
What it does on every access, in full
A real unit does more than add. Whatever the scheme, it performs some subset of these four things, and a question about what the unit does wants them.
- Translate: turn the logical address into a physical one, by adding a base, or by looking
up a segment table, or by looking up a page table.
- Check the bounds: refuse an address outside what the process owns, and trap.
- Check the permissions: refuse a write to a read only region, or an execute of data, and
trap.
- Cache the translation: remember recent lookups so that most accesses need no table read at
The Memory Management Unit
all. That is Chapter seventy six's translation lookaside buffer.
Steps 2 and 3 are the reason the unit is a protection device and not merely an adder, and they are why a program that reads address 1 dies (Chapter two) and a program that writes to its own code dies (Chapter fourteen of Module 1 showed the text section mapped r-xp, read and execute and not write).
The three schemes it can implement
Everything in the rest of this row is one of these, and they differ only in what the unit looks the address up in.
| Scheme | The unit holds | The translation is | Chapter |
|---|---|---|---|
| Base and limit | two registers | add the base | eleven, this one |
| Segmentation | a pointer to a segment table | look the segment up, check its limit, add its base | eighteen |
| Paging | a pointer to a page table | split the address, look the page up, join the frame to the offset | nineteen, twenty |
So this chapter is the shape of the next eight. They are all the same three steps, with a different table and a different arithmetic in the middle.
What it costs, and why that matters
A table lookup is a memory access. So a naive paged unit makes two memory accesses for every one the program asked for: one to read the table, one to get the data.
That would halve the speed of the machine, and it is the reason the fourth thing in the list above exists. Chapter seventy six works the arithmetic and shows how a cache of translations brings the cost back to a few per cent.
The bounds and permission checks cost nothing: they happen in parallel with the translation, in hardware, and a legal access is not slowed at all by being checked.
Seeing the permission check work
The text section of a program is mapped read and execute but not write. Writing to it is not a mistake the compiler can catch, and the unit catches it on the instruction.
#include <stdio.h>
#include <string.h>
int main(void)
{
/* a pointer to this function's own first instruction */
void *code = (void *)main;
printf("about to write one byte over my own code\n");
memset(code, 0, 1);
printf("this line is never reached\n");
return 0;
}$ gcc -std=c17 -Wall -Wextra -o writecode writecode.c
$ sh -c ./writecode
about to write one byte over my own code
Segmentation fault (core dumped)
$ echo $?
139The program was allowed to compute that address, hold it in a pointer, and pass it to a function. Nothing stopped it until the instruction that actually wrote, and what stopped it then was the memory management unit noticing that the region is not writable. No software check happened: the kernel only found out because the hardware trapped.
The Memory Management Unit
That is the whole argument for doing protection in hardware rather than in the operating system: the check has to happen on every access, and only the hardware is present on every access.
Distinctions that carry marks
| The memory management unit | The operating system | |
|---|---|---|
| Is | hardware | software |
| Acts on | every memory access | a context switch, a trap, a system call |
| Translates addresses | yes | no: it programs the unit that does |
| Runs while a process runs | yes | no |
| Bounds check | Permission check | |
|---|---|---|
| Refuses | an address outside the process's own memory | a write to read only memory, or executing data |
| Costs | nothing: parallel with the translation | nothing |
| Failure gives | a trap, then the kernel kills the process | the same |
| Relocation register | Page table | |
|---|---|---|
| Size | one register | one entry per page, thousands of them |
| Translation | add | look up, then join |
| Process must be contiguous | yes | no |
What it does not mean
The unit is not a cache. It has one inside it, the translation lookaside buffer, but its job is translation and protection.
It does not know about processes. It knows what is in its registers, which the kernel changes at every context switch. That change is the expensive part of Chapter sixteen's switch.
A trap is not the unit killing the process. The unit refuses and raises a trap; the kernel decides what to do, and what it usually does is send a signal.
Translation is not slow. A single translation is done in hardware in parallel with the access. What is slow is reading a table from memory to do it, which is why the cache exists.
Quick revision
- The memory management unit is hardware between the processor and the memory. It
translates every logical address to a physical one, on every access.
- The operating system does not translate addresses: it programs the unit, at a context
switch. It is not running while the process is.
- Four jobs: translate, check the bounds, check the permissions, and
cache the translation.
- Its simplest form is a relocation register plus a limit:
if address < limit then base + address else trap. Changing the relocation register moves the whole process.
- Three schemes, all the same three steps with a different table: base and limit,
segmentation, paging.
- A table lookup is itself a memory access, so a naive paged unit doubles the cost of every
access. That is what Chapter seventy six's cache is for.
- The bounds and permission checks are free: they happen in parallel with the translation.
- Writing to your own code compiles, runs, and dies on the instruction: the hardware refused, not
The Memory Management Unit
the operating system.
Test yourself
- What does the memory management unit do? It converts every logical address the processor
generates into a physical address, checking the bounds and the permissions as it does so, and traps if either fails.
- Is it hardware or software, and why does that matter? Hardware. The check must happen on
every memory access, and the operating system is not running while the process is, so only hardware can be present.
- Give the translation and check for a relocation register. If the logical address is
strictly less than the limit, the physical address is the base plus the address; otherwise the hardware traps.
- How is a process relocated under that scheme? By loading a different value into the
relocation register. Nothing in the program changes.
- Name the four things a memory management unit does on an access. Translate the address,
check it is within the process's bounds, check the permissions for that region, and cache the translation for next time.
- Why would a naive paged unit halve the speed of the machine? Because reading the page
table is itself a memory access, so every access the program makes costs two.
- A program writes to its own code and dies. What caught it? The memory management unit,
because the text region is mapped read and execute but not write. It raised a trap and the kernel then killed the process; no software check took place.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.