Skip to main content

7.6 Translation Lookaside Buffer (TLB) & Effective Memory Access Time (EMAT)

📚Module 07: Main Memory ManagementTopic 7.6⏱️20 min read
🎯High-Yield For:Computer Science Foundations • Systems Engineering • Microarchitecture

💡 Core Intuition​

🍳 The Everyday Analogy: The Receptionist's Speed-Dial Directory​

Imagine a busy corporate receptionist who must route phone calls to hundreds of corporate extensions:

Architecture Flow

The Phone Extension Lookup Pipeline

Contrasting master filing cabinet lookups with a high-speed desktop speed-dial ledger

💡 Hover or click any card for deep-dive operational details
🗄️The Problem

Massive Filing Cabinet

Page Table in DRAM

For every incoming caller, the receptionist must stand up, walk to the filing cabinet, find the employee folder, and read their extension.

→
Hardware Optimization
⚡High-Speed Cache

Desktop Speed-Dial Rolodex

Translation Lookaside Buffer (TLB)

The receptionist keeps a 16-card index on top of their desk with the most frequently called extensions.

→
Fallback on Miss
🚶Fallback Path

Filing Cabinet Fallback

TLB Miss & Refill

If an obscure employee is called, the receptionist walks to the cabinet, fetches the number, and writes it onto the desk index.

  • The Two-Access Penalty: Without a cache, pure paging requires two physical memory accesses for every logical reference (one for the Page Table, one for the operand).
  • The TLB Solution: An on-chip associative hardware cache holding the most active page-to-frame translations.
  • Effective Memory Access Time (EMAT): The weighted average latency of memory operations governed by the TLB hit ratio.

💻 Bridging to Computer Science​

Pure paging solves external fragmentation, but doubles physical memory latency:

  1. Access 1: Read the Page Table Entry (PTE) from RAM to resolve Frame Number ff.
  2. Access 2: Fetch the actual program instruction or operand from physical RAM address (f,d)(f, d).

Because DRAM access times hover around 50–100 ns50\text{–}100\text{ ns}, doubling access time halves overall CPU throughput. Computer architects eliminate this bottleneck by placing a specialized hardware cache on the CPU die: the Translation Lookaside Buffer (TLB).

TLB Fast-Path vs Slow-Path Address Translation

Hardware associative search resolving virtual page translations in sub-nanosecond time

CPU Core
Hardware TLB (CAM)
RAM Page Table
Physical DRAM Bus
1
CPU Core→Hardware TLB (CAM)

Issue Logical Address [p | d]

2
Hardware TLB (CAM)→Physical DRAM Bus

Fast-Path TLB Hit: Dispatch Frame f Directly

3
Hardware TLB (CAM)→RAM Page Table

Slow-Path TLB Miss: Hardware Table Walk

4
RAM Page Table→Hardware TLB (CAM)

Return PTE & Reload TLB Cache


📚 Core Deep-Dive & Concepts​

1. Hardware Architecture of the TLB​

The TLB is constructed using Content Addressable Memory (CAM) / Associative Registers:

  • Associative Search: All TLB entries are compared against the target Page Number (pp) simultaneously in parallel within a fraction of a clock cycle (<1 ns< 1\text{ ns}).
  • Key-Value Structure:
    • Key: Page Number (pp).
    • Value: Physical Frame Number (ff), along with protection and status bits.
  • Limited Capacity: Because CAM circuits require dense transistor logic for parallel comparators, TLB sizes are strictly bounded (typically 6464 to 20482048 entries).

2. Effective Memory Access Time (EMAT) Derivation​

Let:

  • hh = TLB Hit Ratio (fraction of accesses where page translation is found in TLB, 0≤h≤10 \le h \le 1).
  • (1−h)(1 - h) = TLB Miss Ratio.
  • tTLBt_{\text{TLB}} = TLB lookup time / access latency (typically 1–2 ns1\text{–}2\text{ ns}).
  • tMMt_{\text{MM}} = Main Memory (DRAM) access latency (typically 50–100 ns50\text{–}100\text{ ns}).

Scenario A: Sequential Hardware Lookup (Industry Standard)​

The MMU first interrogates the TLB. If a miss occurs, it subsequently navigates to Main Memory:

Time on TLB Hit=tTLB+tMM\text{Time on TLB Hit} = t_{\text{TLB}} + t_{\text{MM}}

Time on TLB Miss=tTLB+tMM (Read Page Table)+tMM (Read Data)=tTLB+2⋅tMM\text{Time on TLB Miss} = t_{\text{TLB}} + t_{\text{MM}} \text{ (Read Page Table)} + t_{\text{MM}} \text{ (Read Data)} = t_{\text{TLB}} + 2 \cdot t_{\text{MM}}

Combining both branches into the weighted expectation formula:

EMAT=h⋅(tTLB+tMM)+(1−h)⋅(tTLB+2⋅tMM)\text{EMAT} = h \cdot (t_{\text{TLB}} + t_{\text{MM}}) + (1 - h) \cdot (t_{\text{TLB}} + 2 \cdot t_{\text{MM}})

EMAT=tTLB+tMM+(1−h)⋅tMM\mathbf{\text{EMAT} = t_{\text{TLB}} + t_{\text{MM}} + (1 - h) \cdot t_{\text{MM}}}


Scenario B: Parallel Hardware Search (Simultaneous Lookup)​

On architectures where TLB lookup and initial cache/memory addressing are triggered concurrently:

Time on TLB Hit=tMM\text{Time on TLB Hit} = t_{\text{MM}}

Time on TLB Miss=tTLB+2⋅tMM\text{Time on TLB Miss} = t_{\text{TLB}} + 2 \cdot t_{\text{MM}}

EMAT=h⋅tMM+(1−h)⋅(tTLB+2⋅tMM)\mathbf{\text{EMAT} = h \cdot t_{\text{MM}} + (1 - h) \cdot (t_{\text{TLB}} + 2 \cdot t_{\text{MM}})}


Scenario C: Multi-Level Paging EMAT with TLB​

If the architecture uses nn-level paging (e.g. 2-level or 4-level paging), a TLB miss requires traversing all nn page table levels before accessing the final data operand:

Time on TLB Hit=tTLB+tMM\text{Time on TLB Hit} = t_{\text{TLB}} + t_{\text{MM}}

Time on TLB Miss=tTLB+n⋅tMM (Traverse n levels)+tMM (Target Data)=tTLB+(n+1)⋅tMM\text{Time on TLB Miss} = t_{\text{TLB}} + n \cdot t_{\text{MM}} \text{ (Traverse n levels)} + t_{\text{MM}} \text{ (Target Data)} = t_{\text{TLB}} + (n + 1) \cdot t_{\text{MM}}

EMAT=h⋅(tTLB+tMM)+(1−h)⋅[tTLB+(n+1)⋅tMM]\mathbf{\text{EMAT} = h \cdot (t_{\text{TLB}} + t_{\text{MM}}) + (1 - h) \cdot \big[ t_{\text{TLB}} + (n + 1) \cdot t_{\text{MM}} \big]}


3. Worked Numerical Examples​

Numerical Example 1: Basic EMAT Calculation​

  • Given:
    • TLB Access Time tTLB=20 nst_{\text{TLB}} = 20\text{ ns}.
    • Main Memory Access Time tMM=100 nst_{\text{MM}} = 100\text{ ns}.
    • TLB Hit Ratio h=80%=0.8h = 80\% = 0.8.
    • Sequential lookup model.
  • Calculation: Hit Latency=20+100=120 ns\text{Hit Latency} = 20 + 100 = 120\text{ ns} Miss Latency=20+2×100=220 ns\text{Miss Latency} = 20 + 2 \times 100 = 220\text{ ns} EMAT=(0.8×120)+(0.2×220)=96+44=140 ns\text{EMAT} = (0.8 \times 120) + (0.2 \times 220) = 96 + 44 = \mathbf{140\text{ ns}}
  • Impact of High Hit Ratio:
    • If hit ratio improves to 98%=0.9898\% = 0.98: EMAT=(0.98×120)+(0.02×220)=117.6+4.4=122 ns\text{EMAT} = (0.98 \times 120) + (0.02 \times 220) = 117.6 + 4.4 = \mathbf{122\text{ ns}}
    • Compare 122 ns122\text{ ns} to raw non-paged access (100 ns100\text{ ns}): the TLB achieves a 98%98\% hit rate with only an effective 22%22\% slowdown over unpaged bare metal!

Numerical Example 2: Target Hit Ratio for Performance Threshold​

  • Problem: A computer system has a memory access time of 120 ns120\text{ ns} and a TLB access time of 15 ns15\text{ ns}. What hit ratio hh is necessary to ensure the effective access time does not exceed 140 ns140\text{ ns}?
  • Solution: EMAT≤140\text{EMAT} \le 140 tTLB+tMM+(1−h)⋅tMM≤140t_{\text{TLB}} + t_{\text{MM}} + (1 - h) \cdot t_{\text{MM}} \le 140 15+120+(1−h)⋅120≤14015 + 120 + (1 - h) \cdot 120 \le 140 135+120−120h≤140135 + 120 - 120h \le 140 255−140≤120h255 - 140 \le 120h 115≤120h  ⟹  h≥115120≈95.83%115 \le 120h \implies h \ge \frac{115}{120} \approx \mathbf{95.83\%}

4. Advanced Microarchitecture: ASID & Set Associativity​

Wired-Down Entries​

Operating system kernels require guaranteed non-evictable TLB lines for:

  • Interrupt Service Routines (ISRs): Handling timer ticks, page faults, and hardware interrupts.
  • Kernel Memory Mapping Core: If the page fault handler itself suffered a TLB miss that faulted again, the CPU would encounter a fatal Double / Triple Fault.
  • These critical kernel translations are marked "wired down" (pinned in hardware) so replacement algorithms never evict them.

Set-Associative TLB Tag Bit Derivations​

If a TLB is organized as an nn-way set associative cache with mm total entries:

Number of Sets S=mn\text{Number of Sets } S = \frac{m}{n}

Set Index Bits s=log⁡2(S)\text{Set Index Bits } s = \log_2(S)

TLB Tag Bits =p−s=p−log⁡2(mn)\text{TLB Tag Bits } = p - s = p - \log_2\left(\frac{m}{n}\right)


🏭 In The Real World: Production Case Study​

HugePages in Linux (2 MB & 1 GB Pages) & Database Performance​

In database engines (such as PostgreSQL, MySQL, and Oracle) managing terabytes of buffer pools, standard 4 KB4\text{ KB} pages cause severe TLB Thrashing:

Page ArchitectureSingle Page Size1024-Entry TLB ReachCoverage for 64 GB DatabaseProduction Impact
Standard Paging4 KB4\text{ KB}4 MB4\text{ MB}0.006%0.006\% residentMassive TLB thrashing; continuous multi-level page table walks
2 MB HugePages2 MB2\text{ MB}2 GB2\text{ GB}3.125%3.125\% resident500×500\times reach expansion; 90% reduction in page table walk overhead
1 GB HugePages1 GB1\text{ GB}1024 GB1024\text{ GB} (1 TB1\text{ TB})100%100\% residentZero runtime TLB misses; buffer pool fully cached in TLB
  1. The TLB Reach Problem:
    • TLB Reach=TLB Entries×Page Size\text{TLB Reach} = \text{TLB Entries} \times \text{Page Size}.
    • If a server has 1024 TLB entries and uses 4 KB4\text{ KB} pages, the TLB can only cache translations for 4 MB4\text{ MB} of memory simultaneously.
  2. Transparent Huge Pages (THP) Solution:
    • Modern Linux distributions automatically merge adjacent 4 KB4\text{ KB} pages into 2 MB2\text{ MB} HugePages.
    • A single TLB entry now covers 2 MB2\text{ MB} instead of 4 KB4\text{ KB}, slashing TLB misses by over 90%90\% and accelerating database query throughput by 15–30%15\text{–}30\%.

🎯 Exam & Interview Pitfall Check​

Core Conceptual Questions

Question 1: A system employs 2-level paging. The TLB access time is 10 ns10\text{ ns} and main memory access time is 80 ns80\text{ ns}. If the TLB hit ratio is 90%90\%, calculate the Effective Memory Access Time (EMAT). Answer:

  1. Identify the levels: n=2n = 2 levels of page tables.
  2. TLB Hit: Requires TLB access (10 ns10\text{ ns}) + 11 Memory access for data (80 ns80\text{ ns}) =90 ns= 90\text{ ns}.
  3. TLB Miss:
    • Interrogate TLB: 10 ns10\text{ ns}.
    • Access Outer Page Table in RAM: 80 ns80\text{ ns}.
    • Access Inner Page Table in RAM: 80 ns80\text{ ns}.
    • Access Actual Target Data in RAM: 80 ns80\text{ ns}.
    • Total Miss Time =10+(2+1)×80=10+240=250 ns= 10 + (2 + 1) \times 80 = 10 + 240 = \mathbf{250\text{ ns}}.
  4. Calculate EMAT: EMAT=(0.90×90)+(0.10×250)=81+25=106 ns\text{EMAT} = (0.90 \times 90) + (0.10 \times 250) = 81 + 25 = \mathbf{106\text{ ns}}

Question 2: Explain the role of the dirty bit (modify bit) in a TLB entry. Answer:

  1. When the CPU performs a write operation to a page, hardware sets the dirty bit to 1 in the TLB (and propagates it to the page table).
  2. When the operating system evicts that page during page replacement:
    • If Dirty Bit = 0: The page was only read, never modified. The kernel can discard the frame immediately without writing to secondary storage.
    • If Dirty Bit = 1: The page's contents were modified in RAM. The kernel must write the frame back to swap space / disk before repurposing the frame, preventing silent data corruption.
Common Interview Traps
  • The Multi-Level Miss Multiplier Trap: On a TLB miss with kk-level paging, you do NOT perform 2 memory accesses—you perform k+1k + 1 memory accesses! Every extra page table level adds a full DRAM access cycle on a TLB miss.
  • The "TLB Stores the Entire Page Table" Trap: The TLB does NOT store the entire page table. It is a tiny, expensive on-chip cache storing only a handful (e.g. 64 to 1024) of the most recently referenced translation entries.
  • Forgetting Sequential TLB Addition: Unless a question explicitly states "parallel lookup", always add tTLBt_{\text{TLB}} to both the hit path and the miss path (tTLB+2⋅tMMt_{\text{TLB}} + 2 \cdot t_{\text{MM}}).

💬

Discussion & Doubts