From ELF File to Running Process: Building an ELF64 Loader for Your Kernel
Your kernel boots. Paging works. The scheduler can switch between tasks. Then you reach a question that changes the project: how does a program stored in a file become a process the CPU can execute?
An executable loader is the bridge. It interprets a file, creates a memory image, and prepares a starting context. Get that bridge right and applications can become separate programs, built independently of the kernel. Get it wrong and an innocent-looking executable can overwrite memory, start with corrupt data, or fail before its first useful instruction.
This guide develops the design for a first ELF64 user-program loader. It includes a worked memory layout, validation helpers, an implementation sequence, and a tiny assembly program for testing. The loader code is architectural pseudocode: adapt its memory and process operations to your own kernel.
1. Start with a deliberately small executable contract
Before writing the parser, decide what your kernel accepts. A practical first contract is:
- x86-64, little-endian ELF64.
- Fixed-address
ET_EXECexecutables built for your operating system. - Fully linked freestanding code, with no runtime relocation requirements.
- Eager loading into fresh, private 4 KiB pages.
- No dynamic linker, thread-local storage, or executable stack.
- A documented user address range and a small initial stack contract.
Reject ET_REL object files and ET_DYN images for now. Reject PT_INTERP, PT_DYNAMIC, and PT_TLS, rather than silently pretending to support the services they need. Define an allowlist for any other accepted program-header types. For PT_GNU_STACK, reject executable-stack requests and otherwise keep your non-executable-stack policy; use the same default when that header is absent. This is your kernel's policy, informed by the Linux ELF stack-header documentation.
An ELF container does not make a program compatible with your OS. A Linux executable may expect Linux system calls, startup data, TLS, and runtime services that your kernel does not provide. Even a statically linked program can have those expectations.
Your kernel should already support separate address spaces, user-accessible mappings, permission changes, a user stack, and a controlled transition into user mode. A loader does not implement those foundations for you. If you are still building paging, begin with Paging for Beginners.
2. Load segments, not section names
When inspecting a binary, .text, .data, and .bss are familiar names. They help developers and linkers organize the program. The execution plan comes from the program header table, especially its PT_LOAD entries. One segment can contain several sections.
The important load-segment fields are:
| Field | What the loader uses it for |
|---|---|
p_offset |
Beginning of the segment's bytes in the file |
p_vaddr |
Destination virtual address |
p_filesz |
Number of file bytes to copy |
p_memsz |
Size of the segment in memory |
p_flags |
Requested access permissions |
p_align |
Alignment constraint |
These fields are documented in the ELF reference manual.
Require p_filesz <= p_memsz. Copy the file portion, then make the remaining memory bytes zero. p_filesz can be zero: a segment may need memory without containing initialized file data. These are the rules described by the ELF program-loading specification.
For a user-process loader, allocate physical frames through your memory manager. Do not treat p_paddr as an instruction to claim a particular physical address. The program's address space is your responsibility.
3. Turn one header into a concrete memory layout
Consider this example segment:
p_offset = 0x2123
p_vaddr = 0x402123
p_filesz = 0x0600
p_memsz = 0x1200
The required result is:
Virtual address Contents
0x402000 .. 0x402122 Zeroed page padding
0x402123 .. 0x402722 0x600 bytes from file offset 0x2123
0x402723 .. 0x403322 0xC00 bytes of zero-filled memory
0x403323 .. 0x403FFF Zeroed page padding
The segment starts partway into a page and spans two pages. Its actual memory interval is [0x402123, 0x403323), while its page coverage is [0x402000, 0x404000).
That difference matters. A mapping routine that insists every segment starts at a page boundary will reject useful layouts. A copy routine that starts writing at the rounded page address will put the program's data in the wrong place.
Zeroing the complete newly allocated frames also prevents bytes outside the segment from exposing memory left behind by previous owners. The loader promises a particular process image; the allocator must not accidentally add someone else's data to it.
4. Validate arithmetic before trusting addresses
A loader parses input. Treat the file as potentially malformed even when you compiled the first test program yourself.
This check is unsafe:
if (offset + length > file_size)
return ELF_BAD_FILE;
The addition can wrap before the comparison. Use subtraction after checking the starting point:
#include <stdbool.h>
#include <stdint.h>
static bool file_span_ok(uint64_t offset,
uint64_t length,
uint64_t file_size)
{
return offset <= file_size && length <= file_size - offset;
}
/* USER_END is an exclusive limit within the lower canonical range. */
static bool user_span_ok(uint64_t start,
uint64_t length,
uint64_t user_begin,
uint64_t user_end)
{
return start >= user_begin && start < user_end &&
length <= user_end - start;
}
These helpers check numeric intervals. They do not check whether a range collides with your user stack, a reserved mapping, or another segment. Those are additional checks.
Decode fields from bounded byte ranges. Avoid casting an arbitrary file pointer straight to an ELF structure: that hides assumptions about alignment, structure layout, and byte order. If you use copied structures for this specific format, assert their layout and decode only after the corresponding bytes have been validated.
Header checks for this loader
Require at least a complete 64-byte ELF64 header, then validate the magic bytes, class, encoding, architecture, executable type, and both version fields. For this narrow implementation, require e_ehsize == 64 and e_phentsize == 56; choose which OSABI and ABI-version values your toolchain contract accepts.
Require a nonzero, bounded program-header count. Reject extended program-header numbering rather than accidentally interpreting its sentinel as a normal count. Check the count-times-entry-size multiplication before calculating the table's span, and validate that span against the complete file size. The ELF header reference defines the header fields.
Keep the file bytes stable between validation and copying. A simple approach is to read the file into an immutable kernel-owned buffer; another is a filesystem interface that guarantees a stable image for the load operation.
Segment checks for this loader
For each nonempty load segment:
- Check its file span and require
p_filesz <= p_memsz. - Check its actual virtual span against your permitted user interval.
- Validate alignment and checked page rounding.
- Reject collisions with reserved regions using the rounded page span.
- Reject overlapping actual segment byte ranges under this first-loader policy.
- Bound total image size, page count, and planning metadata before enumerating pages or allocating.
Validate the headers of empty segments too, but let them contribute no pages to the plan. Accept nonempty p_memsz with zero p_filesz.
The offset and virtual address must have matching residues modulo the page size. If p_align exceeds one, require a power of two and matching residues modulo that alignment as well. The gABI alignment rules describe these separate constraints.
Do not forget rounding overflow. Even after checking start + size, an expression such as (end + 4095) & ~4095 needs a checked addition and a final range check.
5. Plan the image by page before allocating
The tempting implementation is:
for each segment:
allocate its pages
zero its pages
copy its bytes
It has a subtle bug: distinct segment byte ranges can share a page after rounding. When the second segment is processed, allocating again can replace the first mapping, and zeroing again can erase the first segment's contents.
Create a plan indexed by virtual page instead:
For every validated, nonempty load segment:
Find every page touched by its memory interval
Add each page to the plan once
Record the permissions requested on that page
Record the segment's copy operation
Validate the completed plan
Allocate and zero each planned page once
Copy every initialized segment range
Apply final page permissions
Combine the permission requests of all segments touching each page. For this implementation, reject any planned page that needs both writing and execution. Checking only whether one segment has both flags misses the case where an executable segment and a writable segment share a rounded page. Rejecting that layout is a deliberate limitation of this educational loader; document it and produce compatible binaries.
Common layouts use read/execute code pages and read/write data pages. Your mapping layer must translate the requested permissions into what the architecture can enforce. For this tutorial, require working execute-disable support before claiming writable data pages cannot execute.
6. Build the process without making it runnable
Use a fresh address space that the scheduler cannot see yet. Populate it through controlled kernel mappings or a copy-to-address-space helper; do not blindly dereference p_vaddr in the currently running kernel context.
The overall operation can look like this:
load_elf(file):
header = decode_and_validate_header(file)
segments = decode_and_validate_program_headers(file, header)
plan = build_and_validate_page_plan(segments)
validate_entry_inside_executable_segment(header.entry, segments)
image = create_unpublished_address_space()
try:
for page in plan:
frame = allocate_owned_frame(image)
zero_frame(frame)
map_for_loading(image, page.address, frame)
for segment in segments.loadable:
copy_to_image(image, segment.vaddr,
file[segment.offset : segment.offset + segment.filesz])
apply_final_permissions(image, plan)
stack = create_zeroed_user_stack_with_guard_page(image)
context = prepare_user_context(header.entry, stack)
remove_loading_aliases_and_finalize(image)
return create_ready_process(image, context)
catch load_failure:
destroy_image_and_all_owned_resources(image)
return failure
Every helper that can fail needs an explicit error path in real code. Make allocations owned as soon as they succeed, including frames that have not yet been mapped. If loading fails halfway through, cleanup must reclaim those frames, mappings, page tables, stack pages, and process metadata.
All frames were zeroed before copying, and overlapping segment byte ranges were rejected, so the uninitialized tails remain zero without a second clearing pass.
Keep temporary loading mappings unavailable to user execution. Finalize permissions and the architecture's required translation updates before publishing the process to the scheduler.
7. The entry point is the start of a process
Require e_entry to lie inside the actual memory interval of an executable load segment and on a final executable page. Landing somewhere in executable page padding is insufficient.
Do not call the address as a kernel C function. That would use the kernel's privilege, stack, and calling context. Enter through your established user-mode transition with valid selectors, sanitized flags, and a kernel stack ready for the next trap or interrupt.
The executable entry is normally _start, which prepares runtime state before calling main. If you adopt the AMD64 System V process-entry convention, the initial stack pointer is 16-byte aligned; ordinary called-function entry has a different alignment because call pushes a return address. The ABI also specifies argument, environment, and auxiliary-vector startup data. See the AMD64 ABI process-initialization description.
Your first assembly program can use a simpler documented contract: a mapped non-executable user stack, aligned stack pointer, known register policy, and no arguments. Establish the remaining CPU state required by your chosen ABI before running compiled programs. Do not claim Linux runtime compatibility from a working _start alone.
If you later supply AT_PHDR, it must name the program header table in the user image, not a pointer to the kernel's file buffer. The test layout below deliberately does not map its ELF headers and does not provide a full System V startup environment.
8. Build a tiny executable with an observable result
Start without a C library or system calls. Save this as probe.asm:
bits 64
default rel
global _start
section .text
_start:
cmp dword [marker], 0x11223344
jne .failed
cmp qword [scratch], 0
jne .failed
mov qword [scratch], 42
mov eax, 0x600D600D
.passed:
pause
jmp .passed
.failed:
mov eax, 0xBAD0BAD0
.failed_loop:
pause
jmp .failed_loop
section .data
marker: dd 0x11223344
section .bss
alignb 8
scratch: resq 1
Save this linker script as user.ld:
OUTPUT_FORMAT(elf64-x86-64)
OUTPUT_ARCH(i386:x86-64)
ENTRY(_start)
PHDRS
{
text PT_LOAD FLAGS(5);
data PT_LOAD FLAGS(6);
}
SECTIONS
{
. = 0x400000;
.text : { *(.text .text.*) *(.rodata .rodata.*) } :text
. = ALIGN(0x1000);
.data : { *(.data .data.*) } :data
.bss (NOLOAD) : { *(.bss .bss.*) *(COMMON) } :data
/DISCARD/ : { *(.comment) *(.note.GNU-stack) }
}
Here FLAGS(5) requests read/execute and FLAGS(6) requests read/write. GNU ld's PHDRS documentation explains how explicit program headers control the output. The page boundary separates code from writable data.
Build with NASM and an ELF-targeting linker:
nasm -f elf64 probe.asm -o probe.o
x86_64-elf-ld -z max-page-size=0x1000 -T user.ld -o probe.elf probe.o
x86_64-elf-readelf -h -l -W probe.elf
A native x86-64 Linux GNU ld can also produce this ELF target. A Windows PE-only linker cannot replace the ELF linker merely because both are named ld.
Before loading the file, inspect the output. Expect EXEC, machine x86-64, an entry in the code segment, two LOAD segments, and no interpreter or dynamic segment. The writable segment should have a larger memory size than file size because it includes scratch.
Load it into your test kernel and observe it with your emulator's debugger. Once it reaches a loop, EAX == 0x600D600D and scratch == 42 indicate that initialized data was copied, the tested BSS word began at zero, and writable memory worked. EAX == 0xBAD0BAD0 identifies a failed data check. A page fault before either loop points to mapping, permissions, or transition setup.
This program deliberately spins. It tests loading without assuming a particular syscall ABI; add process exit once your kernel defines that service. It checks one BSS word, so expand your memory assertions to cover complete segments before declaring the loader correct.
9. Test failures as carefully as successful loads
Keep one known-good ELF fixture, then change one property at a time. Run parser and planning tests outside the kernel where practical; reserve emulator tests for mappings and CPU behavior.
| Test | Expected result |
|---|---|
| Truncated header or program table | Rejected before allocations |
File offset near UINT64_MAX |
Rejected without arithmetic wrap |
p_filesz > p_memsz |
Rejected |
| Zero file size, nonzero memory size | Entire segment begins zeroed |
| Misaligned start with valid congruence | Copied at the exact byte address |
| Two nonoverlapping ranges share one page | Contents preserved; permissions checked together |
| Writable/executable page requirement | Rejected under this loader's policy |
| Entry outside executable segment bytes | Rejected |
| Image touches a stack or reserved page | Rejected |
| Allocation fails partway through | All acquired resources reclaimed |
| Two processes load the same fixed addresses | Distinct private physical frames |
| User code writes to a protected code page | Process fault through the kernel's normal fault path |
Log the rejected field and a useful reason, such as segment 1: file span exceeds image or page 0x401000: writable and executable. A precise loader error is much more helpful than an unexplained reset.
Once this subset is reliable, extend the contract deliberately: arguments and environment, richer startup metadata, demand paging, shared file-backed pages, position-independent executables, relocation, TLS, and dynamic linking. Each extension should add its own validation and failure tests.
Continue with Daniel McCarthy's kernel development courses
If you want to build the surrounding operating system as one connected project, I recommend Daniel McCarthy's kernel development courses. An ELF loader becomes much easier to understand when you can trace how its pages, processes, filesystem, and system calls fit together.
Start with Developing a Multithreaded Kernel From Scratch. It builds the 32-bit PeachOS foundations and includes paging, filesystems, multitasking, and ELF loading. The architecture differs from this ELF64 example, but the connection between executable files and process memory is directly relevant.
For the 64-bit continuation, follow Part Two - Module One, then Part Two - Module Two. The series moves into UEFI and long mode, graphics foundations, and an interactive operating system with user applications.
You can explore all three together in The Complete PeachOS Course Series. Choose the starting point that matches your C, assembly, and kernel experience, then build and inspect each milestone as you go.
Daniel McCarthy