Article revision Version 1 of 1 Current

Designing Memory Management for a Kernel That Can Grow

View current article
Initial version

Original article publication

This is the article as it appeared before any published revisions.

Snapshot

Article at version 1

Operating Systems

Memory management often begins as a handful of functions and ends up influencing nearly every kernel subsystem. The early version may only need to mark frames as used, create a few page-table entries, and provide a basic allocator. Later, the same design must support isolated processes, file mappings, device buffers, multiple CPUs, large pages, memory pressure, and useful diagnostics.

The most important design decision is therefore not which allocator algorithm to choose. It is where to draw the boundaries.

A durable design separates memory management into three cooperating layers:

  1. Object allocation divides usable virtual memory into blocks for data structures.
  2. Address-space management controls which objects appear at which virtual addresses and with what permissions.
  3. Physical-memory management tracks frames and physical address ranges, including resources that are not ordinary RAM.

Each layer solves a different problem. Keeping their contracts explicit makes the simple kernel easier to understand and gives the advanced kernel somewhere to add new behavior without rewriting everything above it.

Start with three different questions

When code asks for memory, the request is often ambiguous. It helps to translate it into three separate questions:

  • How many bytes does the caller need? This is an allocation question.
  • What should be visible at this virtual address? This is an address-space question.
  • Which physical frames or ranges satisfy the request? This is a physical-resource question.

Consider a request for a 200-byte kernel object. A small-object allocator may satisfy it from an existing slab. If that slab is full, the allocator asks the address-space layer for more usable pages. The address-space layer reserves an address range, obtains frames from the physical layer, installs mappings, and applies permissions. The caller does not need to know which frames were chosen, and the physical allocator does not need to know what a 200-byte object is.

That direction of dependency should remain clear:

kernel subsystem or runtime
          |
          v
   object allocator
          |
          v
 address-space manager
          |
          v
 physical-memory manager

Requests flow downward; policy and failures are reported upward. Avoid shortcuts that let arbitrary code reach around one layer into another. Those shortcuts are convenient during bootstrapping but become hidden coupling later.

Layer one: object and region allocation

The highest layer answers requests expressed in bytes, alignment, lifetime, and sometimes object type. In userspace this work usually belongs to a language runtime or C library. The kernel normally provides larger virtual regions, while the process runtime turns those regions into individual allocations. Crossing into the kernel for every small allocation would add unnecessary overhead and contention.

The kernel has the same pattern internally. A general-purpose allocator is useful, but one allocator rarely serves every workload well. A practical kernel may combine:

  • a page-granular region allocator for large requests;
  • size-class caches for frequently allocated small objects;
  • dedicated pools for descriptors, messages, or scheduler structures;
  • per-CPU magazines or caches to reduce shared locking;
  • temporary arenas for data that can be discarded as a group.

The interface should describe requirements instead of exposing an implementation. Useful request fields include:

struct allocation_request {
    size_t bytes;
    size_t alignment;
    enum lifetime_hint lifetime;
    enum zeroing_policy zeroing;
    unsigned flags;
};

A lifetime hint can distinguish short-lived scratch data from long-lived kernel objects. A nonblocking flag can tell the allocator that the caller is in an interrupt or another context where sleeping is forbidden. These are not promises that a particular algorithm will be used; they are information the allocator can use to choose safely.

Fragmentation has two forms

Internal fragmentation is wasted space inside an allocated block, such as placing a 33-byte object in a 64-byte size class. External fragmentation is free space split into pieces that cannot satisfy a larger contiguous request. Both matter, but they matter at different layers.

Small-object allocators primarily manage byte-level fragmentation. The physical layer deals with frame-level fragmentation when contiguous physical memory is required. Virtual contiguity does not imply physical contiguity, so most ordinary buffers should be virtually contiguous and assembled from unrelated frames. Reserving physically contiguous memory for every large allocation needlessly makes the hardest physical requests more difficult.

Debugging belongs in the design

Allocator diagnostics are much easier to add when metadata and failure modes were considered early. Useful development features include:

  • guard pages around selected allocations;
  • red zones or canaries near object boundaries;
  • poisoning newly freed memory;
  • allocation-site tracking;
  • detection of double frees and invalid frees;
  • quarantine before reuse to expose use-after-free bugs;
  • per-cache counts and high-water marks.

These features may be too expensive for every production build, but the allocator should be structured so they can be enabled without replacing its public interface.

Layer two: address spaces and mappings

The address-space manager controls virtual ranges. Its job is larger than manipulating page tables. Page tables are a hardware representation; the manager also needs a software model of what each region means.

For every mapped or reserved range, the kernel may need to know:

  • the start and end addresses;
  • whether the range is private or shared;
  • read, write, and execute permissions;
  • whether pages are anonymous, file-backed, or device-backed;
  • whether physical storage has already been committed;
  • the required cache policy;
  • page-size constraints;
  • ownership and accounting information.

This metadata lets the page-fault handler make policy decisions instead of treating every missing translation as a fatal error.

Reservation, commitment, and mapping are different operations

These terms should not be collapsed into one function:

  • Reserve claims part of a virtual address space so another mapping cannot use it.
  • Commit promises backing storage under a defined policy.
  • Map installs or prepares a translation to a particular backing object.

Separating them enables lazy allocation. A process can reserve a large address range without immediately consuming the same amount of RAM. On first access, a fault handler obtains a frame, clears it, and installs the translation. Untouched pages cost address space and metadata but not physical memory.

The same machinery supports file mappings. A region describes a file and offset; individual pages are loaded only when accessed. Clean file-backed pages can later be discarded because their contents can be read again. Dirty pages require writeback or other preservation.

Page faults are part of the normal control flow

A fault is not automatically an error. The handler must classify it:

  1. Is the address inside a valid region?
  2. Does the attempted access satisfy that region's permissions?
  3. Is the page absent because allocation or I/O was deferred?
  4. Is this a copy-on-write access?
  5. Is the backing object available, or has a real failure occurred?

Only invalid addresses, forbidden accesses, and unrecoverable backing failures should take the fatal path. This classification also makes diagnostics better: "write to a read-only mapping" is much more useful than "page fault."

Isolation depends on permissions and stale-data control

Virtual memory is a security boundary. The kernel must prevent one process from reading or modifying another process's data, and it must prevent ordinary code from writing executable pages without an explicit policy.

At minimum:

  • apply user/supervisor permissions consistently;
  • prefer non-executable data mappings;
  • avoid writable-and-executable mappings unless strictly necessary;
  • clear a frame before exposing it to a different security domain;
  • validate mapping arithmetic for overflow and overlap;
  • keep device mappings out of ordinary applications unless capability checks allow them.

Mapping every physical address permanently into the kernel may seem convenient, but it expands the impact of memory-corruption bugs, consumes translation structures, and weakens control over locality and exposure. Prefer deliberate mappings for the resources the kernel is actively using. If a direct-map window is part of the architecture, define its threat model and access rules rather than treating it as invisible plumbing.

Multiprocessor costs are easy to underestimate

Changing a page table is not the end of an unmap operation. Other CPU cores may still hold the old translation in their translation lookaside buffers. A correct design needs a shootdown protocol or another generation-based invalidation scheme before a frame can be reused unsafely.

Batching helps. Unmapping 1,000 pages one at a time and forcing a cross-CPU invalidation for each page can be dramatically more expensive than changing the whole range and issuing one coordinated invalidation. The interface should therefore support ranges and deferred completion, not only single-page primitives.

Multiple page sizes add another policy choice. Large pages can reduce translation overhead and page-table memory for stable, aligned regions, but they increase internal waste and make fine-grained protection or reclamation harder. Treat them as an optimization selected by measured workload, not as the default answer to every large mapping.

Layer three: physical frames and address ranges

The physical-memory manager owns the machine's physical address map. RAM is only one kind of entry. Firmware tables, device windows, persistent memory, reserved regions, and holes can occupy the same address space.

Build a normalized map early in boot. Resolve overlaps, align boundaries carefully, and retain attributes such as:

  • usable or reserved;
  • volatile or persistent;
  • normal memory or memory-mapped I/O;
  • cacheability requirements;
  • hot-plug capability;
  • NUMA node or proximity domain;
  • firmware reclaimability.

Do not discard this information after placing usable RAM into an allocator. Drivers may later need to find an unused physical range for a device window, and the virtual-memory layer needs the correct cache policy when mapping it.

Optimize the common request

The dominant request is usually simple: allocate or free one ordinary frame with no special physical address requirement. This path should be fast and scalable. A bitmap can provide compact tracking, a buddy allocator can manage blocks of power-of-two frames, and per-CPU frame caches can keep common operations away from a global lock. The exact combination matters less than measuring the real path and keeping special cases out of it.

Free-frame metadata does not require every free frame to be mapped. The manager needs to track the frames, not read their contents. A frame can remain unmapped until some subsystem has a reason to access it.

Special requests need explicit constraints

Devices can impose restrictions that ordinary code should never have to care about. A buffer may need to be:

  • below a maximum physical address;
  • physically contiguous;
  • aligned to a device-specific boundary;
  • confined so it does not cross a boundary;
  • allocated from a particular NUMA node;
  • mapped with a specific cache policy.

Represent these conditions in a constraint structure and route them through a slower path. Dividing memory into zones is one common strategy: scarce low-address memory stays available for hardware that truly requires it, while unrestricted allocations use the general pool.

An input-output memory management unit can remove some contiguity and address-range restrictions by giving devices their own translated address spaces. It does not eliminate the need for a coherent physical-resource model, but it can greatly reduce pressure on special zones.

Reclamation and failure are policies, not afterthoughts

When no free frame is immediately available, the physical allocator alone cannot decide which data should disappear. It needs cooperation from higher layers that understand meaning:

  • clean file-cache pages may be dropped;
  • dirty pages may be written back;
  • anonymous pages may be moved to secondary storage;
  • optional subsystem caches may be shrunk;
  • compressed memory may postpone slower I/O;
  • a process may ultimately need to be terminated under a documented policy.

This is why memory pressure should be signaled through a reclaim framework instead of handled by an arbitrary allocator retry loop. Reclaim can re-enter filesystems, storage drivers, and allocators, so recursion and lock ordering must be designed deliberately.

Hardware also fails. Marking a frame as unusable, migrating data away from a failing region, adding or removing hot-plugged memory, and evacuating a NUMA node all become manageable when ownership is tracked through explicit layers.

Define contracts before choosing algorithms

Algorithms are replaceable; confused ownership is not. Write down the invariants each layer guarantees.

The physical layer might guarantee:

  • a returned frame is not simultaneously owned elsewhere;
  • its NUMA node and zone are known;
  • its previous contents are either inaccessible or sanitized before exposure;
  • freeing a frame updates accounting exactly once.

The address-space layer might guarantee:

  • regions never overlap unless the interface explicitly permits aliasing;
  • installed mappings obey the requested permissions and cache policy;
  • an unmapped frame is not reused until stale translations are handled;
  • every mapping holds a reference to its backing object.

The object-allocation layer might guarantee:

  • returned pointers meet size and alignment requirements;
  • allocation-context rules are honored;
  • freed objects are not returned twice;
  • statistics distinguish requested bytes from reserved capacity.

These invariants should appear in assertions, tests, and interface documentation. They are more valuable than comments describing the current data structure.

Bootstrapping without trapping the final design

Memory management has a circular startup problem: allocators need metadata, but metadata itself needs memory. Solve it with a deliberately temporary boot allocator.

A useful sequence is:

  1. Parse and normalize the firmware-provided physical map.
  2. Reserve the kernel image, boot data, early page tables, and device-owned ranges.
  3. Use a simple monotonic allocator for early metadata.
  4. Initialize the physical frame allocator with every remaining usable region.
  5. Transfer early allocations into permanent accounting.
  6. Bring up the kernel address-space manager.
  7. Initialize the general and specialized object allocators.
  8. Disable or seal the boot allocator so accidental late use is visible.

The temporary allocator can be intentionally unsophisticated. Its essential property is that every byte it consumes can be identified and reconciled later.

A staged implementation roadmap

Build in layers that are useful on their own.

Stage 1: correct ownership

  • Normalize the physical map.
  • Track free and reserved frames.
  • Create and destroy mappings with explicit permissions.
  • Provide one simple kernel allocator.
  • Add assertions, counters, and a map-dump command.

At this stage, favor clarity over cleverness. A slower allocator with strong invariants is a better foundation than a fast allocator whose ownership rules are unclear.

Stage 2: process isolation

  • Create independent user address spaces.
  • Validate user-accessible ranges.
  • Handle anonymous demand allocation.
  • Tear down mappings without leaks.
  • Sanitize frames crossing security boundaries.

Stress-test process creation and destruction. Repeatedly allocate, map, unmap, and exit under constrained RAM so leaks appear quickly.

Stage 3: efficient sharing and I/O

  • Add shared mappings and reference-counted backing objects.
  • Implement copy-on-write.
  • Add file-backed mappings and a page cache.
  • Introduce constrained allocations for devices.
  • Batch translation invalidations.

Stage 4: scalability and pressure handling

  • Add per-CPU caches where profiling shows contention.
  • Introduce NUMA-aware placement.
  • Support large pages for suitable regions.
  • Build reclaim, writeback, and pressure notifications.
  • Add memory hot-plug or fault isolation if the platform requires it.

Each optimization should preserve the same public contracts. If adding NUMA support requires every caller to understand physical nodes, the layering has probably leaked.

Make memory state observable

Memory bugs often surface far from their cause. A kernel that can explain its own state is much easier to develop.

Expose at least:

  • total, free, reserved, reclaimable, and pinned physical memory;
  • usage per zone and NUMA node;
  • virtual regions and permissions for each address space;
  • page-fault counts by reason;
  • allocation failures by size, flags, and call site;
  • object-cache occupancy and fragmentation;
  • translation shootdown counts and latency;
  • reclaim attempts, pages scanned, and pages recovered.

Counters need precise definitions. "Used memory" might mean committed bytes, resident frames, mapped pages, or allocator capacity. Report these separately instead of publishing one misleading number.

Fault injection is equally valuable. Force selected frame allocations to fail, make a page-in operation return an error, or constrain the test system to a tiny memory budget. Failure paths that are never exercised are rarely correct.

Design principles to keep

A memory manager that lasts is built less from one brilliant allocator than from clear boundaries and testable ownership.

  • Keep byte allocation, virtual mapping, and physical-resource tracking separate.
  • Optimize ordinary single-frame and small-object operations without forcing rare constraints into the fast path.
  • Treat page faults, lazy commitment, and reclamation as planned mechanisms.
  • Make permissions, cache policy, and stale-data handling part of the security model.
  • Preserve the full physical address map, not only the list of free RAM.
  • Use temporary boot mechanisms deliberately and reconcile them with permanent accounting.
  • Instrument everything early enough that failures can be explained.

Start simple, but make the simple version an honest subset of the architecture you intend to grow. That choice prevents early convenience from becoming a permanent limitation.

kernelpagingvirtual-memorymemory-managementallocators