munotes®

Memory Layout: the Stack, the Heap and the Frame

Get access to whole semester resourcesSemester Pass

Chapter Ninety-Nine

Syllabus topic Module 2, "Buffer Overflow and Memory Exploitation Concepts: Study ... memory layout"

Pages 459 to 462 of 578

In one line

A running program organises memory into regions: the stack for the short-lived data of function calls, the heap for data that lives longer, and, within the stack, a frame for each function call that holds its local data next to the information needed to return. Understanding this layout is what makes the buffer overflow's danger comprehensible.

In examination wording: a process organises its memory into regions including the stack, which stores function call frames in a last-in-first-out manner, and the heap, which stores dynamically allocated data; each function call creates a stack frame containing its local variables and control information such as the return address, and the adjacency of these within the frame is what gives memory-corruption bugs their security significance.

Why a security topic needs a memory model

The attacks earlier in this book abused logic: a query, a session, a permission, an access-control check. This block abuses the machine underneath the logic, and to understand it a student needs a simple model of how a program uses memory. The chapter builds only that model, at concept level, because without it the next chapter's overflow is a mystery and the defences are arbitrary; with it, both are obvious.

The reason memory safety is a security issue at all: in languages that let the programmer manage memory directly (notably C and C++), the program is responsible for not writing past the space it allocated, and nothing automatically stops it if it does. When a program gets that wrong on data an attacker controls, the attacker can reach beyond the intended space into the program's own control information. So the layout, and specifically what sits next to what, is where the security significance lies.

The regions of memory

A running program's memory is divided into regions, each for a different kind of data. The two that matter here:

The stack. Used for the short-lived data of function calls: the local variables of the function currently running, and the bookkeeping needed to return to whoever called it. It is called a stack because it works last-in-first-out: when a function is called, a new block is added on top; when it returns, that block is removed. This makes it fast and automatic, and it is where the buffer overflow of the next chapter occurs.

The heap. Used for data that must live longer than a single function call, or whose size is not known in advance: data the program explicitly allocates while running and frees when done. It is managed differently from the stack, and it has its own kind of corruption bug (the heap overflow, mentioned in a later chapter).

Other regions hold the program's own code and its global data, but the stack and the heap are where the memory-corruption attacks live, and the stack is the classic case.

munotes.in459

Memory Layout: the Stack, the Heap and the Frame

The stack frame

When a program calls a function, it sets aside a block of stack memory for that call, called a stack frame (or activation record). The frame holds the data that call needs, and its contents are the crux of the whole block:

  • the function's local variables, including any buffers (fixed-size areas for data such as a string or an array);
  • saved control information needed to return correctly, most importantly the return address: the location in the calling code to jump back to when this function finishes.

The single most important fact, and the reason the buffer overflow is dangerous, is that these sit next to each other in the frame. A local buffer and the saved return address are neighbours in memory. The exact arrangement varies, but the essential point holds: a buffer's memory is adjacent to control information, including the return address.

So writing past the end of a buffer does not write into empty space; it writes into the neighbouring memory, which can include the saved control information. That adjacency is what turns a memory mistake into a security problem, and it is the fact the next chapter builds on.

A picture of one frame

A simplified view of a single frame, from the buffer outward toward the caller, is worth holding as a mental image:

Region in the frameHolds
the bufferthe function's local fixed-size array
other saved local stateother saved registers or values
the saved return addresswhere to continue when the function returns
the caller's framethe calling function's data, further along

The direction matters: writing more into "the buffer" than it can hold flows into "other saved local state" and then toward "the saved return address". This is the geography the buffer overflow exploits, and it is why the next chapter can say that overflowing a buffer writes toward the return address.

The stack's discipline, and why overflow breaks it

The stack works because of a strict discipline: each function's frame is added on entry and removed on return, and the return address in each frame tells the program where to go when the function finishes. When a function returns, the program reads the return address from the current frame and jumps there, continuing the caller's work.

That discipline depends on the return address being correct. If the return address in a frame is overwritten, the program, on returning, jumps to wherever the overwritten address now points, not to the caller. The next chapter explains how a buffer overflow can overwrite it and why that is dangerous; this chapter's contribution is the model that makes it clear: the return address is a piece of data in the frame, adjacent to a buffer, that controls where the program goes next.

munotes.in460

Memory Layout: the Stack, the Heap and the Frame

A worked example, at concept level

Consider a function that copies some data into a local buffer, then returns to its caller.

  • On entry, a frame is created for the function, containing the buffer and, beyond it, the saved return address pointing back to the caller.
  • The function fills the buffer with data. If the data fits, all is well: the buffer holds the data, the return address is untouched, and on return the program jumps back to the caller correctly.
  • If the data is larger than the buffer and the function does not check, the copy writes past the buffer's end, into the adjacent saved state and toward the return address. The return address may be overwritten.
  • On return, the program reads the (now possibly overwritten) return address and jumps there. If it was overwritten with nonsense, the program jumps to a nonsense location and crashes; the next chapter explains what happens when the overwriting is controlled.

At this concept level the point is complete: the buffer and the return address are neighbours, so overflowing the buffer can corrupt the return address, which controls where the program goes. Nothing here is an exploit; it is the memory model that makes the security problem comprehensible, which is what MU's "Concepts" asks for.

What beginners get wrong

  • Thinking memory layout is irrelevant to security. The adjacency of a buffer and the return address in a stack frame is precisely what gives buffer overflows their significance; without the layout the attack is inexplicable.
  • Confusing the stack and the heap. The stack holds short-lived call frames last-in-first-out; the heap holds longer-lived, explicitly-allocated data. Each has its own corruption bug.
  • Missing that the return address is data. It is a value stored in the frame, adjacent to buffers, that determines where the program goes on return, which is why overwriting it matters.
  • Assuming writing past a buffer hits empty space. It writes into neighbouring memory, which can include control information such as the return address.
  • Thinking this applies to every language. It is specific to languages that let the program manage memory directly and do not check bounds automatically; memory-safe languages prevent it, as a later chapter explains.

Quick revision

  • A program's memory has regions: the stack (short-lived function-call data, last-in-first-out) and the heap (longer-lived, explicitly-allocated data), plus code and globals.
  • Each function call creates a stack frame holding the call's local variables (including buffers) and saved control information, notably the return address (where to go on return).
  • The crucial fact: within the frame, a buffer is adjacent to the return address, so writing past the end of a buffer writes into neighbouring memory and toward the return address.
  • The stack's discipline depends on the return address being correct; overwriting it redirects where the program goes on return, which the next chapter develops.
  • This is specific to languages that manage memory directly without automatic bounds checking.
munotes.in461

Memory Layout: the Stack, the Heap and the Frame

Test yourself

  1. What are the stack and the heap, and what kind of data does each hold?

The stack holds the short-lived data of function calls, the local variables and control information of the currently running functions, organised last-in-first-out so a frame is added on call and removed on return. The heap holds data that must live longer than a single function call or whose size is not known in advance, allocated explicitly while the program runs and freed when done. The stack is where the classic buffer overflow occurs.

  1. What does a stack frame contain, and which adjacency gives it security significance?

A stack frame contains the function call's local variables, including any fixed-size buffers, and saved control information needed to return correctly, most importantly the return address, which is where the program continues when the function finishes. Its security significance comes from the adjacency of the buffer and the return address within the frame: writing past the end of a buffer writes into neighbouring memory and toward the return address.

  1. Why is the return address described as "data in the frame that controls where the program goes"?

Because it is a value stored in the stack frame, alongside the local variables, that the program reads when the function returns in order to jump back to the caller and continue. Since it determines the location the program jumps to on return, and it sits adjacent to buffers that a function writes into, corrupting it by writing past a buffer changes where the program goes, which is the basis of the buffer-overflow attack.

  1. Why does writing past the end of a buffer not simply hit empty space?

Because a buffer is part of a stack frame in which local variables and control information sit next to one another, so the memory immediately beyond a buffer belongs to other saved state and, further along, the saved return address. Writing more into the buffer than it holds therefore flows into that neighbouring memory rather than into unused space, which can corrupt the control information.

  1. Why is this memory model specific to certain languages?

Because the buffer overflow depends on a language that lets the programmer manage memory directly and does not automatically check that data fits the space allocated for it, as C and C++ do not. In memory-safe languages, which check bounds automatically or prevent the mistake by design, writing past a buffer raises an error rather than corrupting adjacent memory, so the layout does not have the same security significance.

munotes.in462

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Issue
Done!