Kernel, User Mode, and System Calls
A CPU refuses to let ordinary code touch a disk, a page table, or another process. Programs get those operations by asking the kernel through a system call — a trap instruction that switches privilege level and lands on one registered entry point, never on an address the caller chose.
User Mode and Kernel Mode
Processors run in at least two privilege levels. x86-64 calls them ring 3 (user) and ring 0 (kernel); ARM calls them EL0 and EL1. The level is a couple of bits in a control register, and the hardware checks them on every instruction and every memory access.
In user mode the privileged instructions simply fail. Code cannot load a page-table base register, mask interrupts, execute I/O port instructions, or halt the CPU — attempting any of them raises an exception instead of executing. Page-table entries additionally carry a supervisor bit, so kernel memory is not merely off-limits by convention: a user-mode read of a kernel address faults before the load completes.
In kernel mode none of those restrictions apply, which is exactly why the entry points into it are so tightly controlled. The whole security model rests on there being no way to reach ring 0 except at an address the kernel registered in advance.
This is enforcement by hardware, not by the operating system. An OS cannot protect itself with software checks alone, because the code doing the checking would run at the same privilege as the code being checked. The MMU and the privilege bits are what make the boundary real.
- Ring 3 / EL0 for applications, ring 0 / EL1 for the kernel
- Privileged instructions fault rather than execute in user mode
- A supervisor bit in each PTE hides kernel memory from user reads
- Only the CPU can enforce a boundary the kernel sits behind
What a System Call Actually Does
A system call is not a function call. There is no address to jump to, because letting a user process choose a kernel address would defeat the entire boundary. Instead the program executes a dedicated trap instruction — syscall on x86-64, svc on ARM — and the CPU transfers control to whatever handler the kernel installed at boot.
The call number goes in a register (rax on Linux x86-64) and the arguments in a fixed sequence of others (rdi, rsi, rdx, r10, r8, r9). Six arguments is the practical limit; anything larger is passed by pointer to a struct. The trap switches to ring 0, swaps to a per-thread kernel stack, and dispatches through a table indexed by the call number — so the call number is validated as a bounds check before a single handler byte runs.
Then the kernel distrusts everything it was handed. A user pointer might be null, unmapped, point into kernel memory, or be revoked by another thread mid-call, so the kernel never dereferences it directly — it uses checked accessors (copy_from_user on Linux) that fault safely. This is why a buffer is copied rather than used in place: a TOCTTOU race, where the user rewrites the buffer between the check and the use, is otherwise trivial.
On return the kernel places a result in rax, restores the user register state and stack, and executes sysret. Negative small values encode errors, which the C library turns into a −1 return and an errno. A round trip costs roughly 100–300 ns — far more than a function call, mostly from the privilege switch and the cache and branch-predictor state it disturbs.
- A trap instruction, not a jump — the target address is fixed by the kernel
- Call number in a register, up to six arguments, dispatched via a table
- User pointers are copied through checked accessors, never dereferenced raw
- ~100–300 ns per round trip, dominated by the privilege transition
Libraries, Wrappers, and Why Buffering Matters
Programs almost never issue trap instructions themselves. They call a library — fopen in C, Files.readString in Java, Path.read_text in Python — and the library issues the call. The API is the programmer-facing contract; the system-call interface is the privilege boundary underneath it. They are not the same layer and do not correspond one to one.
One library call can become zero, one, or many system calls. fread of 10 bytes from a buffered stream usually becomes zero, because a previous call already read 4 KB. printf to a pipe buffers until a flush. fopen becomes one open. And getaddrinfo may issue dozens across files, sockets, and timers.
The performance consequence is large and easy to demonstrate: writing a million lines unbuffered costs a million traps at ~200 ns each, roughly 0.2 seconds of pure overhead, while a 4 KB buffer cuts that to a few thousand traps. This is why C stdio buffers by default, and why the buffering mode is chosen by destination — line-buffered to a terminal so output appears promptly, fully buffered to a file or pipe where nobody is watching.
It is also why output disappears when a program crashes. Buffered bytes live in user-space memory that never reached the kernel, so a SIGSEGV loses them, while anything already written with an unbuffered write survives. strace and dtruss show the real calls, and the count is usually far lower than the number of library calls in the source.
- API and system-call interface are different layers, not one-to-one
- Buffering turns a million traps into a few thousand
- Line-buffered to terminals, fully buffered to files and pipes
- Lost output after a crash is unflushed user-space buffer
Terms, operations, and practical uses
Privilege levels
- User modeRing 3 / EL0. Privileged instructions fault and kernel pages are unreadable.
- Kernel modeRing 0 / EL1. Full access to hardware, page tables, and interrupt state.
- Trap instructionsyscall on x86-64, svc on ARM. The only way into ring 0, at a fixed entry point.
- Kernel stackA per-thread stack switched to on entry, so kernel state never touches user memory.
Call mechanics
- Call numberSelects the handler through a bounds-checked dispatch table (rax on Linux x86-64).
- Argument registersrdi, rsi, rdx, r10, r8, r9 — six maximum, larger sets pass a struct pointer.
- copy_from_userChecked accessor that faults safely rather than trusting a user pointer.
- errnoA small negative return value the C library converts into -1 plus an error code.
Types of system calls
- Process controlfork, execve, exit, wait.
- File managementopen, read, write, close, lseek.
- Communicationpipe, socket, send, shmget.
- Protectionchmod, umask, setuid.
Control transfers
- TrapSynchronous and deliberate; resumes at the instruction after.
- InterruptAsynchronous and external; unrelated to the running instruction.
- ExceptionSynchronous and unintended; restarts the faulting instruction.
- vDSOKernel code mapped into user space so gettimeofday needs no trap at all.
Kernel architecture
- MonolithicDrivers and file systems in ring 0. Fast; one bad driver corrupts the kernel.
- MicrokernelOnly IPC, threads, and address spaces privileged. Restartable drivers, more overhead.
- HybridXNU pairs a Mach core with a monolithic BSD layer.
- eBPFVerified user-supplied programs run safely inside a monolithic kernel.
Why buffering collapses 6 writes into 2 system calls
# A library write() is not a system call. It appends to a user-space
# buffer, and only a full buffer (or a flush) traps into the kernel.
BUF = 16
buffer = ''
syscalls = 0
def flush():
global buffer, syscalls
if buffer:
syscalls += 1 # this is the only place we enter the kernel
buffer = ''
def write(line):
# the library call
global buffer
if len(buffer) + len(line) > BUF:
flush()
buffer += line
lines = ['one\n', 'two\n', 'three\n', 'four\n', 'five\n', 'six\n']
for line in lines:
write(line)
flush() # fclose/exit flushes what is left
print(f'{len(lines)} library calls -> {syscalls} system calls')#include <iostream>
#include <string>
#include <vector>
using namespace std;
// A library write() is not a system call. It appends to a user-space
// buffer, and only a full buffer (or a flush) traps into the kernel.
const size_t BUF = 16;
string buffer;
int syscalls = 0;
void flush() {
if (!buffer.empty()) {
syscalls++; // the only place we enter the kernel
buffer.clear();
}
}
void write(const string& line) { // the library call
if (buffer.size() + line.size() > BUF) flush();
buffer += line;
}
int main() {
vector<string> lines = {"one\n", "two\n", "three\n", "four\n", "five\n", "six\n"};
for (const string& line : lines) write(line);
flush(); // fclose/exit flushes what is left
cout << lines.size() << " library calls -> " << syscalls << " system calls\n";
}import java.util.List;
class Main {
// A library write() is not a system call. It appends to a user-space
// buffer, and only a full buffer (or a flush) traps into the kernel.
static final int BUF = 16;
static StringBuilder buffer = new StringBuilder();
static int syscalls = 0;
static void flush() {
if (buffer.length() > 0) {
syscalls++; // the only place we enter the kernel
buffer.setLength(0);
}
}
static void write(String line) { // the library call
if (buffer.length() + line.length() > BUF) flush();
buffer.append(line);
}
public static void main(String[] args) {
List<String> lines = List.of("one\n", "two\n", "three\n", "four\n", "five\n", "six\n");
for (String line : lines) write(line);
flush(); // fclose/exit flushes what is left
System.out.println(lines.size() + " library calls -> " + syscalls + " system calls");
}
}Step through it
Running on write 6 lines through a buffered stream, buffer holds 16 bytes
Read all 9 Steps
- Nothing has entered the kernel yet The buffer is empty and no system call has been made. Every write below is an ordinary function call in user mode until the buffer fills.
- write("one") — library only 4 bytes are appended to the user-space buffer. The privilege boundary is not crossed, so this costs a few nanoseconds rather than a few hundred.
- write("two") — still no trap 8 of 16 bytes used. The kernel still has no idea this program is producing output.
- write("three") fills the buffer 8 + 6 = 14 bytes still fit. The buffer is nearly full, and the next line is what forces the kernel entry.
- write("four") overflows — trap #1 14 + 5 would exceed 16, so the library flushes first. The trap instruction switches to kernel mode and hands over all 14 buffered bytes in one call.
- Back in user mode, buffer restarted The kernel returns, the process resumes in ring 3, and "four" goes into the now-empty buffer. One system call carried three lines.
- write("five") and write("six") 5 + 5 + 4 = 14 bytes accumulate without another trap. Two more library calls, still zero additional system calls.
- Final flush — trap #2 There is no seventh line to force an overflow, so the remaining bytes would sit in user memory forever. This is why a crash loses output: fclose, exit, or an explicit flush is what makes the second and last system call.
- 6 library calls became 2 system calls Unbuffered, this would have been 6 traps at roughly 200 ns each. Buffering cut the privilege transitions by two thirds — and at a million lines it is the difference between 0.2 s of overhead and almost none.
Kernel Architectures
Where the boundary is drawn is a design choice. A monolithic kernel puts schedulers, memory management, file systems, network stacks, and drivers all in ring 0. Calls between subsystems are ordinary function calls, so it is fast — but a bug in any driver can corrupt any kernel structure. Linux and the BSDs are monolithic, with loadable modules that run at full privilege.
A microkernel keeps only address spaces, threads, and message passing in ring 0 and pushes file systems and drivers into user-space servers. A crashed driver takes down one restartable process, not the machine, and the trusted computing base shrinks to something formally verifiable — seL4 has a machine-checked proof of correctness. The cost is that a single file read becomes several IPC round trips instead of a function call.
The famous 1992 Tanenbaum–Torvalds exchange over exactly this trade-off was decided by benchmarks rather than argument: monolithic won general-purpose computing on speed. Microkernels won everywhere correctness dominates — QNX in cars, seL4 in avionics and security hardware, and Apple's XNU as a hybrid with a Mach microkernel core and a monolithic BSD layer above it.
The trend now blurs the line from both directions. Windows moved graphics into the kernel for speed, then moved drivers back out into a user-mode framework for stability; Linux runs FUSE file systems in user space and sandboxes verified eBPF programs inside the kernel. The question stopped being which architecture and became which subsystems earn ring 0.
- Monolithic: everything in ring 0, fast, one bad driver kills the kernel
- Microkernel: IPC-based, restartable drivers, verifiable, slower
- seL4 has a machine-checked correctness proof; QNX ships in cars
- eBPF and FUSE move the line per-subsystem rather than wholesale