The objective of this book is to develop a particular way of understanding memory-safety bugs in the C programming language. For the purposes of this book, memory safety concerns whether accesses remain valid with respect to objects, bounds, storage, and lifetimes.
C emerged from a design context in which close control over machine resources was a feature rather than an anomaly. Its relatively low-level memory model allows programmers to manipulate objects through pointers and pointer arithmetic without generally requiring the language to enforce that each access remains within the bounds and lifetime of the object concerned. [1] [2]
Consider:
// Allocate enough storage for n struct item objects and make p point to the first one.
struct item *p = malloc(n * sizeof *p);
The compiler can establish facts about the types involved:
p has type struct item *;sizeof *p is the size of a struct item;p is scaled according to that type.But those facts are only part of the reasoning that makes the statement meaningful. The language determines what the compiler can be required to establish, and C leaves much of the reasoning connecting values and operations to valid memory implicit.
The compiler therefore generally does not establish that:
n * sizeof *p did not overflow;malloc succeeded;n objects;p is being used to designate those objects;n is correct;C's types help the compiler determine how operations on values should be translated, but they do not require it to prove the entire chain of reasoning that connects those operations to valid memory. That responsibility remains, to a significant extent, with the programmer.
Memory-safety problems arise when that reasoning ceases to hold. A value may produce the wrong allocation size; a pointer may consequently designate the wrong location; a range may extend beyond its intended object; or an object may no longer exist when it is accessed.
The central difficulty is that the programmer's model of the memory being accessed can diverge from the memory that the program actually accesses. This book seeks to make that divergence visible and to develop a way of reasoning backwards from an invalid access to the assumptions about values, objects, bounds, locations, and lifetimes that made it possible.
Memory-safety bugs in C are often described using categories such as:
integer overflow
buffer overflow
out-of-bounds access
use-after-free
double free
memory corruption
These categories are useful for describing failures, but they can hide the relationships between them. A vulnerability may begin as an error in one part of a program and become visible only after that error has propagated through several subsequent operations.
Consider a program that receives a count, calculates a size from that count, allocates memory, walks through the resulting object, and eventually copies data into it.
The count might be wrong.
The calculation might be wrong.
The allocation might be too small.
The pointer might be calculated from the wrong value.
The bounds might be wrong.
The copy might be too large.
The object might no longer exist.
The eventual failure may therefore occur several steps away from the original mistake.
A useful starting point for following such a failure is:
integer arithmetic
↓
allocation size
↓
pointer arithmetic
↓
bounds
↓
memory corruption
This is not a rigid sequence. Real programs may skip steps, repeat them, or introduce other operations between them. The memory chain is a way of thinking about the dependencies that connect one operation to the next.
When a number is used to describe memory, we should ask where that number came from.
When an address is calculated, we should ask what object it is intended to identify.
When memory is allocated, we should ask what logical object the allocation is supposed to contain.
When a range is accessed, we should ask why the program believes that range is valid.
And when memory is freed, we should ask whether the pointer still identifies the intended object.
The answers to these questions form a chain of reasoning.
C code often looks deceptively local. Consider:
size = count * element_size; // Calculate the total number of bytes needed
char *p = malloc(size); // Allocate that amount of memory
char *q = p + offset; // Compute a new pointer
memcpy(q, src, length); // Copy length bytes to that location
Each statement is short. The interesting question is not whether each statement looks reasonable by itself. It is whether the assumptions connecting them are true.
What does count represent?
What does size represent?
What object does p point to?
What does offset measure?
What is the relationship between length and the remaining
space?
The answers to these questions turn a collection of statements into a chain of reasoning. They require us to identify the invariants that connect the operations and to treat the code as an argument that memory is being allocated and addressed correctly.
A memory-safety bug occurs when an assumption required for that reasoning no longer holds.
Memory can be wrong in two fundamental ways: the program can access the wrong amount, or it can access the wrong location.
A program can use the correct location but the wrong amount of memory. It can use the correct amount but the wrong location. It can get both wrong.
These cases lead to different manifestations, but the underlying question is the same:
What memory did the programmer believe this operation referred to,
and what memory did it actually refer to?
That question is often more useful than simply asking which bug category applies.
Suppose a program crashes during a copy.
The copy may be where the problem becomes visible, but it may not be where the problem began.
Perhaps an input value was converted incorrectly.
Perhaps an arithmetic calculation overflowed.
Perhaps an allocation was therefore too small.
Perhaps a later pointer calculation still used the original logical size.
Perhaps the copy finally crossed the boundary of the allocation.
Looking only at the final copy tells us what went wrong at the end of the chain. Following the chain tells us why.
That distinction matters both for debugging and for security. It determines whether we merely patch the operation that happened to fail or repair the assumption that made the failure possible.
The chapters that follow use this idea as a recurring method of analysis. We will not treat integer arithmetic, allocation, pointers, bounds, and memory corruption as unrelated topics. Instead, we will follow values as they move through a program.
A value may begin as:
input
become:
count
then:
size
then:
allocation
then:
offset
then:
address
and finally:
memory access
At each stage, the program relies on an assumption about what that value means. Our task is to determine whether that assumption remains true.
This approach also explains why apparently minor C details can become security vulnerabilities. A small numerical mistake is not necessarily dangerous by itself. It becomes dangerous when the resulting value is trusted to describe memory.
Throughout the book, return to one question whenever code manipulates memory:
What memory did the programmer believe this operation referred to,
and what memory did it actually refer to?
Then ask:
Which value determines it?
How was that value calculated?
Which object is involved?
What bounds apply?
What happens next?
These questions provide a way to trace an invalid access backwards through the assumptions that made it possible.
The goal is not to memorize a longer list of dangerous C constructs. It is to develop a way of seeing the relationships between numbers, objects, addresses, and memory.
Once those relationships become visible, many apparently different C memory vulnerabilities start to look like variations of the same problem:
the program believed it was operating on one region of memory,
but the calculation, address, allocation, copy, or lifetime
described another.
That is the wrong memory.
The Wrong Memory is for C and systems programmers, security engineers, code reviewers, students, and others interested in understanding how memory-safety failures arise.
It assumes familiarity with C, pointers, arrays, structures, and dynamic allocation, but does not require experience in vulnerability research. The book is not a catalogue of C vulnerabilities; it develops a way of reasoning about the relationships between values, objects, bounds, locations, and lifetimes that make those vulnerabilities possible.
Experienced programmers will likely recognise many of the individual rules and failure modes discussed here. The aim is to see them not as isolated problems, but as different manifestations of the same underlying reasoning.
The question at its centre is:
What did the program believe about memory, and where did that belief stop being true?
References:
The Development of the C Language
ISO/IEC 9899 — Programming Language C