Article Community improvements enabled

Kernel Development Roadmap: From First Boot to a Graphical x86-64 OS

Daniel McCarthy published 4 weeks ago 14 min read 190 views
New post
Daniel McCarthy remains the original author. Improvements are attributed to their editors and reviewed by the original author before publication.

Kernel Development Roadmap: From First Boot to a Graphical x86-64 OS

Kernel development becomes manageable when you stop treating an operating system as one enormous program and start treating it as a sequence of testable milestones.

The practical path is:

toolchain
   -> first boot
   -> CPU setup and exceptions
   -> memory management
   -> interrupts and scheduling
   -> user mode and system calls
   -> filesystems and drivers
   -> 64-bit graphics and applications

This roadmap explains what each stage is for, what a successful result looks like, and which DragonZap Community guide to read next. You do not need to buy anything to use it. The course recommendations near the end are for readers who want to build the connected PeachOS project with Daniel McCarthy in a structured sequence.

Before you begin: are you ready?

You do not need previous kernel experience, but kernel development is not a first programming project. You should be able to:

  • Read and write C functions.
  • Follow pointers and work with structs.
  • Understand hexadecimal and bitwise operations.
  • Recognise basic x86 assembly instructions and registers.
  • Run build commands in a terminal.
  • Use Git to save known-good checkpoints.
  • Spend time debugging assumptions that ordinary application code can take for granted.

If pointers, structs, stack frames, or assembly function calls are still unfamiliar, strengthen those skills first. A kernel has no host operating system underneath it to catch a bad pointer, provide printf, or recover a broken memory map.

For the path below, use an emulator such as QEMU before trying physical hardware. An emulator gives you repeatable hardware, quick resets, debug logs, and a clean connection to GDB. Keep the exact compiler, assembler, linker, emulator, and firmware versions in your project README.

Stage 0: choose a small, specific target

Decide what your first operating system is meant to teach you. An educational x86 kernel is a different project from a production desktop OS, a real-time embedded kernel, or a Linux-compatible system.

A good first target is:

Boot one x86 machine configuration in QEMU, enter the intended CPU mode, print useful diagnostics, and halt safely.

Choose one architecture and one boot route. Legacy BIOS is excellent for seeing the early boot process in a small amount of assembly. UEFI is the modern route and becomes important when moving into 64-bit development. Learning the legacy path first does not make it UEFI; they are separate firmware environments with different interfaces.

Write down three project rules:

  1. Every milestone must have visible expected output.
  2. Every working milestone gets a Git commit or tag.
  3. When the kernel fails, reduce the test before adding another subsystem.

For a broader orientation before writing code, read Getting Started with Operating System Development.

Stage 1: build a freestanding toolchain

Your normal compiler is configured for programs that run on your current operating system. It may assume the host ABI, startup objects, C library, executable format, and linker behaviour. A kernel supplies those foundations itself.

A cross-compiler separates the machine on which the compiler runs from the target for which it emits code. Names such as i686-elf-gcc and x86_64-elf-gcc make that boundary explicit. You will also need target-aware assembler, linker, objdump, readelf, and related binary tools.

Work through Build a GCC Cross-Compiler for Kernel Development. Record the versions you use, and verify the result instead of stopping when the build command exits successfully:

i686-elf-gcc --version
i686-elf-ld --version
i686-elf-objdump --version

Your first toolchain checkpoint is complete when a tiny freestanding object can be compiled and inspected without linking against your host operating system.

Common failure: using -ffreestanding with the host compiler and assuming that this alone creates a complete kernel toolchain. That flag changes language assumptions, but it does not remove every host-specific assembler, linker, target, and runtime assumption.

Stage 2: make the machine execute your code

The first boot is the smallest end-to-end proof in the project. It demonstrates that the firmware or bootloader, disk image, assembly, linker layout, and emulator agree about where execution begins.

For the legacy BIOS route, follow How to Create a 16-bit MS-DOS-Style Bootloader with NASM. The guide starts with a 512-byte sector and then loads a second stage.

The important model is:

BIOS reads sector 0 to physical address 0x7C00
   -> validates the 0x55AA signature
   -> jumps to 16-bit code
   -> your loader establishes a known stack and segment state
   -> your code prints a message or loads the next stage

A useful first exercise

Build the one-sector example, then change the message and rebuild it. Verify all three results:

  1. The output file is exactly 512 bytes.
  2. Its final two bytes are 55 AA.
  3. QEMU displays your changed message.

Expected output might be:

My kernel journey starts here.

If QEMU says the disk is not bootable, check the file size and signature. If nothing prints, inspect org 0x7C00, the segment registers, the stack, and the terminating zero in the string. If stage two does not run, verify its byte offset in the disk image and the number of sectors the first stage reads.

This is a legacy introduction. Modern UEFI boot uses firmware protocols and executable images rather than BIOS interrupts. Keep that distinction explicit as your project evolves.

Stage 3: establish a controlled CPU environment

On 32-bit x86, moving from real mode to protected mode gives the kernel access to a larger address space, privilege levels, descriptor tables, and the environment in which the later kernel can run.

Read Protected Mode in OSDEV for the complete transition. The essential sequence is:

  1. Disable interrupts while the interrupt environment is incomplete.
  2. Enable the A20 line.
  3. Define and load a valid Global Descriptor Table.
  4. Set the protected-mode enable bit in CR0.
  5. Perform a far jump so the processor reloads the code segment and instruction stream correctly.
  6. Reload the data segment registers.
  7. Establish a known 32-bit stack.
  8. Transfer control to the kernel entry point.

Protected mode does not automatically enable paging, install exception handlers, or preserve access to BIOS services. Those are separate jobs.

Your checkpoint is not merely “the emulator did not reset.” Print a message from 32-bit code, inspect the register state in a debugger, and deliberately trigger a controlled fault after an Interrupt Descriptor Table is installed. A useful exception handler should report the vector and enough CPU state to locate the failure.

A sudden reboot commonly means a triple fault: the original exception could not be delivered, handling that failure caused another fault, and the processor reset. Check the GDT, IDT, selectors, stack, and handler entry convention before changing unrelated code.

Stage 4: add interrupts and a timer

Exceptions describe events caused by the current instruction or CPU state. Hardware interrupts let devices and timers request attention asynchronously. Your kernel needs both.

Start with CPU exceptions, then configure the appropriate interrupt controller for your target. Add one timer source and prove that its count increases while the kernel continues to run. Keep the interrupt handler small: acknowledge the source correctly, capture necessary state, update minimal bookkeeping, and defer larger work.

Before attempting task switching, test these checkpoints independently:

  • A breakpoint or deliberate fault reaches the correct exception handler.
  • A timer produces repeated interrupts, not just one.
  • Unexpected vectors are logged rather than silently ignored.
  • The stack remains aligned according to the calling convention used by your C code.
  • Interrupts are not enabled until the IDT and handler stacks are valid.

The debugging question is always concrete: which vector arrived, what state did the CPU save, what did your assembly stub add, and what exact frame does the C handler expect?

Stage 5: build memory management in layers

“Memory management” is at least three related problems:

  • Physical memory management decides which page frames are free or owned.
  • Virtual memory management decides which virtual addresses map to which frames and with what permissions.
  • Dynamic allocation gives kernel subsystems convenient regions or objects during execution.

Do not hide all three behind one vague allocate() function.

Start with Paging for Beginners: How Your Kernel Controls Memory. Learn to trace a virtual address through its page-table indexes and identify the final physical frame. Then read Designing Memory Management for a Kernel That Can Grow to separate ownership, mapping, reservation, commitment, reclamation, and failure policy.

For a concrete first dynamic allocator, work through Build a Kernel Heap from Scratch: The Block Table Method. A fixed-block table is not the final word in allocator design, but its metadata and invariants are visible enough to test thoroughly.

Memory milestones

  1. Discover or receive a physical memory map.
  2. Reserve the kernel, boot data, page tables, firmware regions, and device memory.
  3. Allocate and free individual physical frames without duplicates.
  4. Map one virtual page, read and write it, then unmap it.
  5. Enforce read/write, user/supervisor, and execute permissions where supported.
  6. Report page-fault addresses and error bits.
  7. Allocate several kernel-heap blocks and validate their metadata.
  8. Stress allocation and freeing in different orders.

Test the allocator's internal table as well as the pointers it returns. Detect double free, invalid free, overlap, out-of-range access, arithmetic overflow, and broken allocation chains. Poison freed memory in a debug build so stale uses become visible.

Stage 6: create tasks and preempt them safely

A task switch is a controlled replacement of one saved CPU context with another. The hardest part is not choosing the next task; it is defining exactly who saves every register and what the restore code expects on the new stack.

Read From Timer Tick to Task Switch: A Preemptive x86-64 Kernel Scheduler. Even if your first kernel is 32-bit, the guide's core invariant transfers: the interrupt entry code and task-switching code must agree on one exact frame layout.

Build scheduling in this order:

  1. Verify timer interrupts without switching.
  2. Save and restore the same context.
  3. Manufacture the initial stack frame for one new kernel thread.
  4. Switch cooperatively between two threads.
  5. Preempt two CPU-bound threads.
  6. Stress general-purpose registers and stack alignment.
  7. Add blocking, waking, and explicit run-queue invariants.

Printing from a timer interrupt can distort timing or deadlock on a logging lock. Use counters, trace buffers, or rate-limited diagnostics when testing preemption.

Stage 7: cross the user/kernel boundary

Kernel threads all execute with kernel privilege. User processes require separate address spaces, restricted page permissions, a defined executable format, controlled entry into kernel services, and a safe return path.

A practical order is:

  1. Create an address space with kernel mappings protected from user access.
  2. Load a small ELF program into mapped user pages.
  3. Enter user mode with a known stack.
  4. Implement one system call, such as writing a character.
  5. Validate every user pointer and length before dereferencing it in the kernel.
  6. Terminate a faulty process without bringing down the whole system.
  7. Add process creation, waiting, and resource cleanup.

Treat the system-call interface as an ABI. Document register use, argument sizes, error results, pointer rules, and which side owns every buffer. User input is untrusted even when you wrote the first user program yourself.

Stage 8: filesystems and storage

Separate these layers:

user API
   -> system call
   -> file descriptor
   -> virtual filesystem
   -> filesystem implementation
   -> block cache or disk streamer
   -> storage driver
   -> hardware or emulated device

Start with an initial RAM filesystem or a small, well-understood disk format. FAT is useful for learning directory entries, allocation tables, sector I/O, and the difference between a generic file interface and one on-disk implementation.

Your first filesystem milestone should mount one known image, list a directory, open one known file, read exact bytes, close it, and report useful errors. Test truncated images, invalid clusters, reads across sector boundaries, end-of-file behaviour, and repeated open/close cycles.

Do not begin with every storage technology at once. A simple emulated ATA or VirtIO device is easier to observe than a complex physical controller. Add caching only after the uncached path is correct and measurable.

Stage 9: move into 64-bit UEFI and graphics

The 64-bit path adds long mode, four-level paging, a different calling convention, wider registers, and a modern firmware environment. UEFI can provide a framebuffer, memory map, files, and services during boot, but the kernel must stop relying on boot services after the defined handover.

Use the PeachOS 64-Bit project showcase to see the larger destination: a 64-bit kernel with graphical foundations and, in the later project, interactive windows and user applications.

Build visible graphics in small steps:

  1. Capture and validate the framebuffer description.
  2. Draw one pixel and one filled rectangle.
  3. Render a bitmap font.
  4. Build a graphical terminal.
  5. Separate graphics buffers, windows, and screen composition.
  6. Deliver mouse and keyboard events through defined queues.
  7. Let a user-space application draw through a controlled interface.

A graphical desktop is not a single feature. It sits on memory allocation, processes, system calls, input drivers, filesystem access, event delivery, and redraw rules. If a window fails, test the lowest failing layer instead of debugging the whole desktop at once.

Stage 10: drivers and real hardware

Drivers turn bus discovery, registers, interrupts, DMA, and device-specific protocols into kernel services. Begin with devices whose behaviour is easy to observe: serial, timer, keyboard, framebuffer, and a simple emulated storage device.

For an advanced example of how a modern device changes the problem, read Building an Intel Dual-Band Wi-Fi Driver: Firmware, DMA, and Rings. It demonstrates why serious drivers need explicit initialization states, bounded timeouts, validated firmware, DMA ownership rules, ring invariants, interrupt discipline, and staged bring-up.

The right order for a complex driver is not “make networking work.” It is:

identify device
   -> map registers
   -> reset predictably
   -> prove interrupts
   -> establish valid DMA memory
   -> load and validate firmware
   -> exchange one command
   -> receive one event
   -> add one end-to-end feature

Never put a CPU virtual address into a DMA descriptor unless your platform's DMA API explicitly says that address is valid for the device. Track CPU and device addresses separately.

The debugging discipline that makes the roadmap work

Add observability before you need it:

  • Serial logging that works before graphics.
  • Panic output with vector, instruction pointer, stack pointer, and fault address.
  • objdump, readelf, and linker-map inspection in the build workflow.
  • A QEMU launch mode that waits for GDB.
  • Assertions for page, allocator, run-queue, and descriptor invariants.
  • Stack canaries and guard pages when the memory manager can support them.
  • Small tests for parsers and allocators that can run outside the kernel when practical.

When the system resets or freezes, write down the last confirmed checkpoint. Check generated machine code and hardware state rather than assuming the source code expresses what the CPU received.

A compact milestone checklist

  • A versioned cross-toolchain produces freestanding objects.
  • A boot image prints a message in QEMU.
  • Protected mode or long mode is entered deliberately and verified.
  • Exceptions report useful register state.
  • A timer produces repeated interrupts.
  • Physical frames can be allocated without overlap.
  • Virtual pages can be mapped, protected, faulted, and unmapped.
  • Kernel heap metadata survives stress tests.
  • Two tasks switch contexts reliably.
  • One user program enters through an ELF loader.
  • One validated system call crosses the privilege boundary.
  • One file can be read through a filesystem abstraction.
  • One storage driver survives repeated I/O and reset tests.
  • Framebuffer output becomes a terminal or window.
  • Input reaches a user application through a defined event path.

Do not treat this as a race. Each checked box represents understanding and a reproducible result, not merely code that happened to run once.

Free paths and reference material

There are excellent free ways to learn OS development. The OSDev Wiki is a broad reference for prerequisites, toolchains, booting, memory, debugging, and architecture topics. Nanobyte's OS tutorial repository accompanies a practical video series. If you prefer Rust, Philipp Oppermann's Writing an OS in Rust builds a small x86-64 kernel through focused tutorials.

Use these resources, processor manuals, specifications, and the DragonZap Community articles together. The value of a paid course is not secret information; it is a coherent project sequence, explanation, maintained learning context, and a clear continuation when you would rather not assemble the path alone.

Where Daniel McCarthy's kernel courses fit

Daniel's PeachOS courses follow the same progression as this roadmap, but as one connected C and x86 assembly project.

If you have working C fundamentals, some assembly familiarity, and want the complete route from the first boot sector to a 64-bit graphical desktop, see The Complete PeachOS Course Series. It connects three courses and lets you inspect each curriculum before choosing.

If you prefer a smaller first commitment, start with Developing a Multithreaded Kernel From Scratch - Part 1. Part 1 builds the 32-bit foundations: boot, protected mode, interrupts, paging, a heap, filesystems, processes, system calls, ELF loading, and multitasking.

Already completed or purchased Part 1? Continue with Part 2 - Module 1, which moves PeachOS into UEFI and 64-bit long mode and develops its graphics, memory, and partition foundations. Then use Part 2 - Module 2 for the interactive window system, user-space graphics, PCI/PCIe, NVMe, applications, and later project work.

Check the prerequisites and your existing DragonZap courses before buying. Recorded video duration is viewing time, not the total time needed to type, test, debug, and experiment. The project is educational; physical-hardware support depends on the drivers implemented.

Your next action

Choose the earliest unchecked milestone in the list. Make its expected output precise, finish only that step, and save the working state.

If you are starting today, build the 512-byte boot sector, change its message, and verify the result in QEMU. If you already have a booting kernel, post your current milestone, architecture, toolchain, and the exact failure you are investigating in the community. A reproducible technical question is much easier for other kernel developers to help with than “my OS crashes.”

Build one layer. Verify it. Then build the next.

Discussion 0

No comments yet. Start a thoughtful discussion.

Join the discussion

You need an account to contribute.

Sign in