8.1 Demand Paging & The Page Fault Handling Cycle
💡 Core Intuition
🍳 The Everyday Analogy: The University Lending Library
Imagine a university student conducting comprehensive research for a master's thesis:
The Library Book Request Pipeline
Contrasting borrowing an entire warehouse with checking out chapters on demand
Hauling the Entire Library
You attempt to haul all 500 reference books to your tiny desk before reading a single line.
Reading at the Desk
You sit down with only the chapter open that you are actively reading right now.
The Missing Book Request
When a book is not on your desk, you place a reservation ticket with the librarian.
- Traditional Swapping: Requires loading an entire process into physical RAM before execution can begin.
- Demand Paging: Loads pages strictly on demand—only when the CPU attempts to execute code or access data residing on that specific page.
- The Page Fault: A hardware trap generated when the CPU attempts to access a page marked invalid (not currently resident in physical RAM).
💻 Bridging to Computer Science
In earlier memory systems, a program could not execute unless the entire executable image fitted completely inside physical DRAM. This imposed severe restrictions:
- Program size was strictly bounded by the physical RAM installed on the motherboard.
- Inactive code (such as error-handling routines, initialization scripts, and rarely used menu options) permanently squandered precious DRAM.
Virtual Memory shatters this constraint by separating the Logical Address Space perceived by user programs from the physical DRAM installed in the machine:
Virtual Memory Architectural Gains
Key operating system capabilities unlocked by decoupling virtual address space from physical DRAM
Programs Exceed Physical RAM
Capacity Decoupling- A 64 GB machine-learning model can execute on a system with only 16 GB of physical DRAM.
- Non-resident pages remain safely on backing swap storage until actively referenced.
Higher Degree of Multiprogramming (DoM)
Memory Efficiency- Because each process consumes only a fraction of its total pages in physical RAM, the kernel can host dozens of concurrent processes simultaneously.
- Maximizes CPU utilization and prevents memory starvation.
Faster Launch Latency (Zero-I/O Cold Start)
Demand Paging- The OS does not read gigabytes of binaries into RAM before calling main().
- Execution begins immediately via Demand Paging, faulting in pages only on demand.
Seamless Memory Sharing
Inter-Process Optimization- Shared libraries (e.g., libc.so) and inter-process shared memory regions (POSIX shm) map to identical physical frames across disparate page tables.
- Dramatically reduces physical DRAM consumption across processes.
📚 Core Deep-Dive & Concepts
1. Pure Demand Paging & The Lazy Swapper (Pager)
Under Demand Paging, the operating system uses a Lazy Swapper (Pager):
- A traditional swapper manipulates entire process address spaces.
- A Pager manipulates individual pages. A page is never brought into physical memory unless the CPU actively references an instruction or operand on that page.
- In Pure Demand Paging, a process begins execution with zero pages in physical RAM.
- The OS loads the PCB, sets the instruction pointer (PC) to the first instruction, and dispatches the process.
- The very first memory fetch immediately triggers a Page Fault!
- The OS pages in Page 0, restarts the instruction, and execution proceeds, faulting only as new un-cached pages are needed.
- Thanks to the Principle of Locality of Reference, after an initial cluster of page faults, execution runs at full DRAM wire speed with minimal faults.
2. The Valid / Invalid Bit Mechanics
To support demand paging, every Page Table Entry (PTE) contains a hardware-monitored Valid-Invalid Bit ():
| Valid Bit () | Target Memory Residence | Hardware MMU & OS Action |
|---|---|---|
1 (Valid) | Physical DRAM Frame () | Direct translation; memory access completes at full DRAM wire speed. |
0 (Invalid) | Backing Store (Swap/Disk) or Unmapped | Hardware Page Fault Exception raised. OS checks VMA: loads page from disk if legal, or raises SIGSEGV if invalid. |
When the CPU executes an instruction referencing page :
- The MMU consults the process's page table.
- If
Valid == 1: The translation succeeds instantly; frame is accessed in DRAM. - If
Valid == 0: The MMU halts execution and triggers a hardware exception: Page Fault Trap.
3. The 7-Step Page Fault Handling Cycle
Handling a page fault requires a coordinated sequence between hardware MMU circuits and OS kernel interrupt routines:
4. Mathematical Modeling of Demand Paging: EMAT Performance
A standard DRAM access takes roughly ().
In contrast, servicing a page fault requires disk access, typically taking ()!
Because disk latency is slower than DRAM, even a tiny fraction of page faults degrades performance catastrophically.
The Effective Memory Access Time Formula
Let:
- = Page Fault Rate (Probability of a page fault per memory access, ).
- = Main Memory (DRAM) access time (e.g. ).
- = Page Fault Service Time (e.g. ).
Numerical Derivation: Tolerable Page Fault Rate
Suppose an operating system targets a maximum performance degradation of less than :
Critical Performance Axiom:
To keep system slowdown under , fewer than out of every memory references may trigger a page fault! ().
5. The Modify (Dirty) Bit Optimization
When memory is completely full and a page fault occurs, the kernel must evict an existing victim page from physical DRAM to make room for the incoming page:
🏭 In The Real World: Production Case Study
Linux kswapd, Anonymous Pages & Memory Compaction
In production Linux kernels, waiting for a memory allocation to fail before evicting pages (Direct Reclaim) introduces massive tail-latency spikes. The kernel prevents this using asynchronous memory daemons:
Linux Memory Pressure Watermarks
Hierarchical memory thresholds governing background eviction and synchronous reclaim
Normal Wire-Speed Allocation (> 15% Free RAM)
Sufficient free memory available. Memory allocations complete directly from free page lists with zero swap stalls.
Asynchronous Background Reclaim (< 10% Free RAM)
Kernel activates kswapd daemon in background. Asynchronously scans LRU lists and flushes dirty pages to disk.
Direct Reclaim & OOM Killer (< 3% Free RAM)
Allocating threads stall synchronously in Direct Reclaim. If allocations still fail, kernel OOM Killer terminates processes.
- Background Eviction via
kswapd:- When free memory crosses below the
lowwatermark,kswapdwakes up in the background and writes dirty pages out to storage before RAM is completely exhausted.
- When free memory crosses below the
- File-Backed vs Anonymous Memory:
- File-backed pages (such as code binaries and mapped database files) are clean and can be dropped instantly without writing to swap space.
- Anonymous pages (stack, heap, and
mallocbuffers) have no file on disk and must be explicitly written to swap partitions or compressed in RAM using zRAM.
🎯 Exam & Interview Pitfall Check
Question 1: In a computer system, memory access time is . It takes () to service a page fault if the victim page is clean, and if the victim page is dirty. If of victim pages are dirty, and the page fault rate is (), calculate the Effective Memory Access Time (EMAT). Answer:
- Average Page Fault Service Time ():
- Calculate EMAT:
- Observation: A tiny page fault rate of only slows down average memory access from to nearly (an almost performance collapse)!
Question 2: Why must the CPU restart the exact instruction that triggered a page fault, rather than proceeding to the next instruction? Answer:
- A page fault is an interrupt / trap that occurs mid-instruction execution (during the instruction fetch phase or operand read/write phase).
- The instruction never finished executing; its destination register or memory location was never updated.
- If the CPU simply executed the next instruction, the program state would become corrupted (e.g. an un-fetched operand would be treated as garbage data).
- Therefore, the CPU must roll back any micro-architectural side-effects (such as autoincremented index registers) and restart the identical instruction from scratch once the missing page is resident in RAM.
- The Unit Conversion Trap: Always convert units to nanoseconds or milliseconds before evaluating EMAT. Mixing with without converting () is the single most common mathematical mistake on university and technical exams.
- Confusing Page Fault with Hardware Crash: A page fault is NOT an error; it is a routine, planned hardware-software coordination mechanism that allows programs to execute seamlessly without requiring all pages to be pre-loaded into physical DRAM.