What this lesson covers
The cache memory lesson treated main memory as "the slow thing below the caches". This lesson opens that box. You will see how a DRAM chip stores bits, why it must be refreshed, how rows and columns shape its latency, and how DDR generations raised bandwidth. Then you will follow a memory access through the memory management unit (MMU), the hardware that translates the virtual addresses a program uses into the physical addresses the DRAM understands, including page tables, the translation lookaside buffer (TLB) and the cost of a page walk. The lesson ends with how caches and TLBs cooperate (VIPT caches) and with non-volatile memory: ROM and flash.
Interviewers commonly ask: SRAM versus DRAM, why DRAM needs refresh, what a TLB is and what happens on a TLB miss, how much memory a TLB covers (TLB reach), why huge pages help, how multi-level page tables save space, and how to compute effective access time with a TLB. The operating-system view of the same topics (page replacement, demand paging, thrashing) lives in virtual memory; here we focus on the hardware.
SRAM versus DRAM
Both are volatile memories: they lose their contents when power is removed. They differ in how a single bit is stored.
SRAM: a latch made of transistors
SRAM (static random-access memory) stores each bit in a small circuit of usually six transistors: two cross-coupled inverters that hold the value, plus two access transistors connecting the cell to the bit lines. As long as power is on, the inverters keep reinforcing each other, so the value stays put with no maintenance. "Static" means exactly this: no refresh.
SRAM cell (6 transistors), conceptual
bit line word line bit line (inverted)
| | |
+--[access]--+ | +--[access]----+
| | |
+---v---+---v---+
| inv A <-> inv B | two inverters feed each
+------------------+ other: a stable latch
SRAM is fast (it is read by sensing the latch directly) but each bit costs six transistors, so it is large and expensive per bit. It is used for registers, caches, TLBs and other on-chip buffers.
DRAM: a charge in a capacitor
DRAM (dynamic RAM) stores each bit as charge on a tiny capacitor, with one transistor acting as a switch: the 1T1C cell. Charged means 1, discharged means 0 (or the reverse; it is a convention).
DRAM cell (1 transistor + 1 capacitor)
word line ----+
|
bit line ---[transistor]---||--- ground
capacitor
Two consequences follow:
- Leakage. The capacitor slowly leaks charge, so the cell forgets within milliseconds unless it is periodically read and rewritten. That periodic rewrite is refresh, and it is why the memory is called "dynamic".
- Destructive reads. Reading connects the capacitor to a long bit line, which shares and disturbs the charge. A sense amplifier detects the tiny voltage change, amplifies it to a full 0 or 1, and writes it back.
| Property | SRAM | DRAM |
|---|---|---|
| Cell | 6 transistors (typical) | 1 transistor + 1 capacitor |
| Density | low | high (several times more bits per area) |
| Cost per bit | high | low |
| Access time | about a nanosecond on chip | tens of nanoseconds from the core |
| Refresh | not needed | needed (every few tens of ms per row) |
| Read | non-destructive | destructive, restored by sense amps |
| Typical use | registers, caches, TLBs | main memory |
Interview tip
A crisp answer to "SRAM vs DRAM": SRAM holds a bit in a six-transistor latch, needs no refresh, is fast but large, and is used for caches. DRAM holds a bit as charge in one capacitor, is dense and cheap, but leaks so it must be refreshed, and reads are destructive and slower. That is why caches are SRAM and main memory is DRAM.
How DRAM is organised: banks, rows and columns
A DRAM chip is not one long list of bits. It is built from several banks, each a two-dimensional array of cells arranged in rows (driven by word lines) and columns (bit lines). Each bank has a row buffer: a row of sense amplifiers that holds one whole row (typically a few kilobytes across the chips of a module) after it has been read.
One DRAM bank
column address
|
row +-------------------------+
addr ->| row 0 ................ |
| row 1 ................ |
| row r ################ | <- activated row
| ... |
+-------------------------+
| row buffer (sense amps) | holds row r
+-------------------------+
|
column mux --> data pins
Reading data takes three steps, each with a named timing parameter:
- Activate (sometimes called RAS, row address strobe): send the row address; the whole row is copied into the row buffer. Time: tRCD (row-to-column delay).
- Read (CAS, column address strobe): send the column address; the selected part of the row buffer is sent out. Time: CL (CAS latency).
- Precharge: close the row, writing it back and preparing the bit lines for a different row. Time: tRP.
This leads to three cases for a request:
- Row hit: the needed row is already open in the row buffer. Cost: CL only.
- Row empty (closed): no row open. Cost: tRCD + CL.
- Row conflict: a different row is open. Cost: tRP + tRCD + CL, the slowest.
Worked example: CAS latency in nanoseconds
Timings are quoted in clock cycles of the memory bus clock. DDR4-3200 transfers 3,200 million times per second on two clock edges, so its I/O clock is 1,600 MHz and one cycle is 0.625 ns.
- A module with CL22: 22 × 0.625 = 13.75 ns.
- A faster-binned DDR4-3200 module with CL16: 16 × 0.625 = 10 ns.
- DDR5-4800 with CL40: clock 2,400 MHz, so 40 ÷ 2.4 GHz = about 16.7 ns.
Notice that the newer DDR5 part has a larger CL in cycles and a similar or higher latency in nanoseconds. Each DDR generation has raised bandwidth far more than it has cut latency. The absolute access latency of DRAM has improved only slowly for many years.
Common mistake
Do not compare CAS latencies across speeds by cycle count. CL40 on DDR5-4800 is not "2.5 times slower" than CL16 on DDR4-3200; convert both to nanoseconds first. Also remember that CL is only part of the full load latency the core sees, which also includes the caches, the on-chip network and the memory controller queue, which is how you end up at the 50 to 100 ns quoted for a DRAM access.
DRAM refresh
Each DRAM cell must be refreshed before its charge leaks below the level the sense amplifiers can read. The JEDEC standards for DDR3 and DDR4 specify that every row must be refreshed within 64 ms at normal operating temperature (the window halves at high temperature).
Refresh is done a row (or group of rows) at a time. The memory controller issues refresh commands at a regular interval; for DDR4 there are 8,192 refresh commands per 64 ms window.
Worked example: refresh interval
Refresh interval (tREFI) = 64 ms / 8,192 = 7.8125 us
So roughly every 7.8 microseconds the controller issues a refresh command, and during the refresh time (tRFC, a few hundred nanoseconds for large chips) the affected bank or rank cannot serve reads or writes. The fraction of time lost is small but grows with chip capacity, since bigger chips have more rows to refresh. That is one reason DDR5 added same-bank refresh, which refreshes one bank while others keep working.
Refresh also connects to security: Rowhammer is a known effect where repeatedly activating one row disturbs charge in neighbouring rows enough to flip bits before they are refreshed. Mitigations include more frequent refresh of victim rows and target-row-refresh logic inside the chips.
Modules, channels and DDR generations
A DIMM (dual inline memory module) is a circuit board carrying several DRAM chips. A classic DDR4 DIMM has a 64-bit data bus (72 bits with ECC, error-correcting code). A rank is a set of chips that together fill that bus width and respond to the same command. A channel is an independent bus between the memory controller and one or more DIMMs. More channels mean more parallel bandwidth.
SDRAM (synchronous DRAM) synchronised the memory with a clock. DDR (double data rate) SDRAM transfers data on both the rising and falling clock edges, so a 1,600 MHz I/O clock yields 3,200 mega-transfers per second (MT/s). Each generation since has raised transfer rates, lowered voltage and increased internal parallelism.
| Generation | Typical transfer rates | Supply voltage | Notable change |
|---|---|---|---|
| DDR | 200-400 MT/s | 2.5 V | data on both clock edges |
| DDR2 | 400-1066 MT/s | 1.8 V | 4n prefetch |
| DDR3 | 800-2133 MT/s | 1.5 V (1.35 V low-voltage) | 8n prefetch |
| DDR4 | 1600-3200 MT/s | 1.2 V | bank groups |
| DDR5 | 4800 MT/s and up | 1.1 V | two independent 32-bit subchannels per DIMM, on-die ECC, same-bank refresh |
The prefetch figure is how many bits each data pin's internal array delivers per access, which lets a slow core array feed a fast interface. Exact rates vary by vendor and by year; treat the table as a high-level map.
Two related families: LPDDR (low-power DDR) for phones and laptops, soldered close to the processor; and HBM (high-bandwidth memory), which stacks DRAM dies vertically and connects them with a very wide interface, used in GPUs and AI accelerators.
Worked example: peak bandwidth
Peak bandwidth = transfer rate × bus width in bytes.
- One DDR4-3200 channel, 64 bits = 8 bytes: 3,200 × 10^6 × 8 = 25.6 GB/s.
- Two such channels (dual channel): 51.2 GB/s.
- One DDR5-4800 DIMM: two 32-bit subchannels = 64 bits = 8 bytes total, so 4,800 × 10^6 × 8 = 38.4 GB/s.
- DDR5-6400: 6,400 × 10^6 × 8 = 51.2 GB/s per DIMM.
These are peaks. Real sustained bandwidth is lower because of refresh, row conflicts, read-write turnarounds and command overhead.
Memory interleaving
A single DRAM bank can only work on one row at a time. Interleaving spreads consecutive addresses across several banks (or channels) so that their accesses overlap.
- Low-order interleaving: the low-order bits of the block address choose the bank. Consecutive blocks land in different banks, so a sequential stream keeps all banks busy. This is what memory systems do for bandwidth.
- High-order interleaving: the high-order bits choose the bank, so each bank holds one contiguous region. It is simpler to expand (add a bank, add a region) but a sequential stream hits one bank at a time.
Low-order interleaving across 4 banks (word address mod 4)
Bank 0: words 0, 4, 8, 12 ...
Bank 1: words 1, 5, 9, 13 ...
Bank 2: words 2, 6, 10, 14 ...
Bank 3: words 3, 7, 11, 15 ...
Worked example: three memory organisations
A cache miss fetches a 32-byte block = four 8-byte words. Assume:
- 1 bus cycle to send the address,
- 20 cycles of DRAM access time per access,
- 1 cycle to transfer one 8-byte word on the bus.
(a) One-word-wide memory, no interleaving. Each word needs its own access and transfer.
cycles = 1 + 4 × (20 + 1) = 1 + 84 = 85
bandwidth = 32 bytes / 85 cycles = 0.38 bytes per cycle
(b) Four-word-wide memory and bus. One access returns all four words at once.
cycles = 1 + 20 + 1 = 22
bandwidth = 32 / 22 = 1.45 bytes per cycle
(c) Four-way low-order interleaved, one-word bus. All four banks start their access at the same time; then the words are transferred one after another.
cycles = 1 + 20 + 4 × 1 = 25
bandwidth = 32 / 25 = 1.28 bytes per cycle
Interleaving gets most of the benefit of a wide bus (25 versus 22 cycles) without the cost of a bus four times wider. Modern systems combine both: a 64-bit channel, many banks and bank groups, and several channels.
The memory controller
The memory controller translates cache-line requests into DRAM commands. On modern processors it is integrated on the processor die (it used to sit in a separate "northbridge" chip), which cut latency.
Its jobs:
- Address mapping: decide which channel, rank, bank, row and column each physical address goes to, usually interleaving at a fine grain to spread traffic.
- Command scheduling: obey dozens of timing constraints (tRCD, tRP, CL, and many more) while reordering requests for throughput. A common policy, FR-FCFS (first-ready, first-come-first-served), prefers requests that hit an already open row, then falls back to the oldest request.
- Page policy: an open-page policy leaves a row open after access, betting on another row hit; a closed-page policy precharges immediately, betting on random access.
- Refresh: issue refresh commands on schedule.
- Reliability: check and correct ECC, scrub memory in the background.
- Write buffering: batch writes to reduce read-write turnaround penalties.
Virtual memory in hardware: the MMU
Every program believes it has its own large, private address space starting at zero. That illusion is virtual memory. Programs issue virtual addresses (VAs); the hardware turns each into a physical address (PA) before it reaches memory. The unit that does this, on every instruction fetch, load and store, is the MMU (memory management unit), part of each core.
Memory is divided into fixed-size pages (commonly 4 KiB). A virtual address is split into a virtual page number (VPN) and a page offset. Translation replaces the VPN with a physical frame number (PFN); the offset passes through unchanged.
Virtual address (4 KiB pages)
+--------------------------------+--------------+
| virtual page number (VPN) | offset (12) |
+--------------------------------+--------------+
| |
translate | unchanged
v v
+--------------------------------+--------------+
| physical frame number (PFN) | offset (12) |
+--------------------------------+--------------+
Physical address
Page table entries
The mapping is stored in a page table in ordinary memory, maintained by the operating system but read directly by hardware. Each page table entry (PTE) contains the PFN plus control bits. On x86-64 and similar designs these include:
| Bit | Meaning |
|---|---|
| Present / valid | the page is in physical memory; if clear, access raises a page fault |
| Read/write | writes allowed |
| User/supervisor | user-mode code may access |
| Accessed | set by hardware when the page is used (helps the OS approximate LRU) |
| Dirty | set by hardware when the page is written (must be saved before reuse) |
| No-execute (NX / XD) | instruction fetch from this page is forbidden |
A special register holds the physical address of the current process's top-level page table: CR3 on x86, satp on RISC-V, TTBR0/TTBR1 on ARM. A context switch loads a new value, which switches the entire address space.
Why page tables have several levels
A single flat table would be enormous.
- 32-bit addresses, 4 KiB pages, 4-byte PTEs: 2^32 ÷ 2^12 = 2^20 entries × 4 bytes = 4 MiB per process, even for a tiny program.
- 48-bit addresses, 4 KiB pages, 8-byte PTEs: 2^36 entries × 8 bytes = 512 GiB per process. Impossible.
A multi-level (hierarchical) page table is a tree. Only the parts of the tree covering addresses the process actually uses need to exist.
On x86-64 with 4-level paging, a 48-bit virtual address is split 9 + 9 + 9 + 9 + 12. Each table has 2^9 = 512 entries of 8 bytes, exactly one 4 KiB page.
47 39 38 30 29 21 20 12 11 0
+-----------+-----------+-----------+-----------+------------+
| PML4 (9) | PDPT (9) | PD (9) | PT (9) | offset (12)|
+-----------+-----------+-----------+-----------+------------+
CR3 --> [PML4] --> [PDPT] --> [PD] --> [PT] --> frame + offset
Newer x86 processors also support 5-level paging with 57-bit virtual addresses. ARMv8 and RISC-V Sv48 use similar 4-level trees.
Worked example: splitting a virtual address
Translate the indexes for VA 0x7F123456789A under 4-level paging.
- Offset = low 12 bits =
0x89A(2,202). - PT index = bits 12-20 = 359.
- PD index = bits 21-29 = 418.
- PDPT index = bits 30-38 = 72.
- PML4 index = bits 39-47 = 254.
The walker reads entry 254 of the table pointed to by CR3, uses its PFN to find the PDPT and reads entry 72, then entry 418 of the PD, then entry 359 of the PT, which gives the frame. The final address is that frame with offset 0x89A.
For a 32-bit two-level scheme (10 + 10 + 12), VA 0x00403ABC splits into directory index 1, table index 3 and offset 0xABC.
Hardware page walkers versus software-managed TLBs
On x86, ARM and RISC-V, a TLB miss is handled entirely in hardware by a page-table walker state machine that reads the page-table levels from memory (through the caches). Only if it finds a non-present or permission-violating entry does it raise a page fault for the operating system.
Some older architectures, notably classic MIPS, used a software-managed TLB: a TLB miss trapped to an OS handler that walked whatever data structure the OS liked and wrote the TLB entry with special instructions. This is flexible but slower per miss.
The translation lookaside buffer (TLB)
Walking four levels of page table on every load would multiply memory traffic by five. The TLB is a small cache of recent translations (VPN to PFN plus permission bits). It sits in the MMU and is checked on every access.
TLB organisation
- Size and associativity. First-level TLBs are small (tens of entries) and often fully associative or highly associative, so they can be checked in the same cycle as the L1 cache. Second-level TLBs hold around a thousand to a few thousand entries and are set-associative. Exact sizes vary a lot between processor generations.
- Split I-TLB and D-TLB at the first level, mirroring split L1 caches; a unified L2 TLB behind them.
- Multiple page sizes. Separate arrays or tagged entries for 4 KiB, 2 MiB and 1 GiB pages (on x86-64).
- Address space identifiers. Without a tag, every context switch would have to flush the TLB, since VA 0x1000 means different things in different processes. An ASID (ARM, RISC-V) or PCID (x86) tags each entry with its process, so entries from several processes coexist.
TLB entry (conceptual)
+---+------+-----------+-----------+---------------------+
| V | ASID | VPN | PFN | perm: R W X U, D, G |
+---+------+-----------+-----------+---------------------+
The G (global) bit marks kernel mappings that are identical in every process and need not be flushed.
TLB reach
TLB reach is the amount of memory a TLB can map without missing: number of entries × page size. If a program's working set exceeds the reach, it suffers TLB misses even when its data fits in the caches.
Worked example: TLB reach
- A 64-entry L1 D-TLB with 4 KiB pages: 64 × 4 KiB = 256 KiB.
- A 1,536-entry L2 TLB with 4 KiB pages: 1,536 × 4 KiB = 6,144 KiB = 6 MiB.
- The same 1,536 entries holding 2 MiB pages: 1,536 × 2 MiB = 3,072 MiB = 3 GiB.
- A 32-entry array for 2 MiB pages: 64 MiB. Four 1 GiB-page entries: 4 GiB.
A database or JVM heap of tens of gigabytes overwhelms a 6 MiB reach. This is the reason operating systems offer huge pages (Linux transparent huge pages and hugetlbfs, Windows large pages): each TLB entry then covers 512 times more memory, and the page walk is one level shorter. The trade-offs are internal fragmentation and the cost of finding contiguous physical memory.
What happens on a memory access
CPU issues virtual address
|
v
+---------+ hit -> physical address -> cache / memory
| TLB |
+---------+
| miss
v
hardware page walk (reads up to 4 PTEs, through the caches)
|
PTE present and permitted?
yes -> fill TLB, retry access
no -> page fault exception -> OS handler
(load page from disk or kill process)
The cost of a page walk
Each level of the walk is a dependent memory read: you cannot read the PD until you have the PDPT entry. In the worst case a 4-level walk costs four memory accesses before the data access itself.
Worked example: effective access time
Assume TLB lookup = 1 ns, memory access = 100 ns, TLB hit ratio = 98 percent, 4-level page table, and (pessimistically) no page-table entries in the caches.
- On a TLB hit: 1 + 100 = 101 ns.
- On a TLB miss: 1 (lookup) + 4 × 100 (walk) + 100 (data) = 501 ns.
- EAT (effective access time) = 0.98 × 101 + 0.02 × 501 = 98.98 + 10.02 = 109 ns.
With a 99 percent hit ratio: 0.99 × 101 + 0.01 × 501 = 105 ns.
Even a 2 percent miss rate adds about 8 percent to every access. Real processors soften this in three ways:
- Page-walk caches (also called paging-structure caches) keep recently used upper-level entries, so a miss often needs only the last level or two.
- Page-table entries are ordinary data and are usually found in the L1, L2 or L3 caches, costing tens of cycles instead of a DRAM access.
- Several page walks can be in flight at once, and out-of-order cores keep doing other work during a walk.
Common mistake
In EAT questions, read carefully whether the TLB lookup time is added on a miss (it usually is: you search the TLB first, then walk), and how many levels the page table has. A single-level table gives a miss cost of lookup + 2 memory accesses; a 4-level table gives lookup + 5.
TLB shootdowns
When the OS changes a mapping (unmapping memory, changing permissions, migrating a page), stale copies may sit in other cores' TLBs. TLBs are generally not kept coherent by hardware the way data caches are, so the OS sends an inter-processor interrupt asking each relevant core to invalidate the entry: a TLB shootdown. They are expensive, which is one reason munmap-heavy workloads scale poorly on many cores.
How caches and the TLB interact
The cache needs an address to index into and a tag to compare. Should those come from the virtual or the physical address?
- PIPT (physically indexed, physically tagged): translate first, then access the cache. Simple and free of aliasing problems, but the TLB lookup sits on the critical path of every hit. Common for L2 and L3.
- VIVT (virtually indexed, virtually tagged): use the VA for both. Fast, but suffers homonyms (the same VA in two processes means different data, so you need ASIDs or flushes) and synonyms or aliases (two VAs mapping to the same PA can end up in two cache lines, so a write to one is not seen in the other). Rare in modern general-purpose L1s.
- VIPT (virtually indexed, physically tagged): start indexing the cache with the virtual address at the same time as the TLB lookup; when the TLB produces the PFN, compare it with the physical tags read from the set. This is what most L1 caches do.
VIPT L1 lookup (in parallel)
VA: | VPN ................ | page offset (12 bits) |
| | index | block offset |
v |
[ TLB ] -> PFN v
| [ L1 set: N ways of tags ]
+------> compare PFN with each physical tag --> hit?
VIPT avoids the alias problem when the index and block-offset bits lie entirely inside the page offset, because those bits are identical in the virtual and physical address. That gives the classic constraint:
cache size / associativity must be at most page size
Worked example: why L1 caches are often 32 KiB and 8-way
With 4 KiB pages, the page offset is 12 bits. A 32 KiB, 8-way cache with 64-byte lines has 64 sets: 6 index bits + 6 offset bits = 12 bits, exactly the page offset. Each way is 32 KiB ÷ 8 = 4 KiB. A 48 KiB L1 must therefore be 12-way to keep each way at 4 KiB. Growing an L1 under VIPT means adding ways, which costs power and hit time. That is one reason L1 sizes grew slowly on 4 KiB-page systems, while processors that use 16 KiB base pages (such as Apple's M-series under macOS) can have larger L1s with the same associativity.
Non-volatile memory: ROM and flash
Non-volatile memory keeps data without power. Computers need it for firmware (the code that runs at power-on, such as BIOS or UEFI) and for storage.
ROM types
| Type | Full name | How it is written | Erase |
|---|---|---|---|
| Mask ROM | read-only memory | contents fixed in the chip's photomask at manufacture | never |
| PROM | programmable ROM | once, by the user, by blowing fuses (or antifuses) | never |
| EPROM | erasable PROM | electrically, with a programmer | whole chip, by ultraviolet light through a quartz window |
| EEPROM | electrically erasable PROM | electrically, in circuit | electrically, byte by byte; slow, limited cycles |
| Flash | (a type of EEPROM) | electrically, in pages | electrically, in large blocks; fast |
The name "ROM" persists for firmware chips even though they are now almost always flash and can be updated.
Flash basics
Flash stores bits as charge trapped in a floating gate (or a charge-trap layer) of a transistor. The trapped charge shifts the transistor's threshold voltage, which the chip reads as a bit value. Key properties:
- NOR flash allows random byte reads and can run code directly from the chip (execute in place); it is used for small firmware chips. NAND flash is far denser and cheaper, read and written in pages (a few KB to tens of KB) and erased in blocks of many pages; it is used in SSDs, phones and USB drives.
- Erase before write. A page cannot be overwritten in place; its block must be erased first. So SSDs write new data to fresh pages and mark old ones invalid, and a garbage collector later copies valid pages out of a block and erases it.
- Limited endurance. Each block survives a limited number of program/erase cycles. Wear levelling spreads writes over all blocks.
- Bits per cell. SLC (1 bit), MLC (2), TLC (3), QLC (4). More bits per cell means more capacity and lower cost but slower writes and fewer program/erase cycles.
- FTL (flash translation layer): firmware inside the SSD that maps logical block addresses to physical flash pages, doing all of the above invisibly. In effect, an SSD contains its own small virtual memory system.
- Modern NAND is stacked vertically (3D NAND) with many layers of cells.
How the SSD connects to the processor (SATA versus NVMe over PCIe) is covered in the I/O organization lesson; the OS side is in storage and I/O.
Interview questions
Q1. Why does DRAM need refresh, and SRAM does not?
A DRAM bit is charge on a tiny capacitor that leaks away within milliseconds, and reading it also disturbs it, so every row must be read and rewritten periodically (within 64 ms for DDR4 at normal temperature). An SRAM bit is held by two cross-coupled inverters that actively maintain the value as long as power is applied, so it never needs refreshing.
Q2. What is a row buffer, and what is a row hit?
Each DRAM bank copies an entire row into a row of sense amplifiers, the row buffer, when the row is activated. A subsequent access to the same row is a row hit and only pays the CAS latency. Access to a different row must precharge (close) the current row and activate the new one, which is much slower. Memory controllers reorder requests to maximise row hits.
Q3. What does "DDR" mean, and what changed across DDR generations?
DDR stands for double data rate: data moves on both rising and falling clock edges. Each generation raised transfer rates (from a few hundred MT/s for DDR to 4,800 MT/s and beyond for DDR5), lowered voltage (2.5 V to 1.1 V) and increased internal parallelism through larger prefetch, bank groups and, in DDR5, two subchannels per DIMM. Latency in nanoseconds has improved much less than bandwidth.
Q4. Compute the peak bandwidth of dual-channel DDR4-3200.
Each channel transfers 8 bytes per transfer at 3,200 million transfers per second, giving 25.6 GB/s. Two channels give 51.2 GB/s peak; sustained bandwidth is lower because of refresh, row conflicts and bus turnarounds.
Q5. What is memory interleaving and why does it help?
Interleaving spreads consecutive addresses across multiple banks or channels, usually by the low-order address bits. Sequential accesses then go to different banks whose access times overlap, so bandwidth approaches that of a wider memory without the cost of a wider bus.
Q6. What does the MMU do?
The memory management unit translates every virtual address issued by the core into a physical address, using the TLB and, on a TLB miss, a hardware walk of the page tables. It also enforces protection (read, write, execute, user/kernel) and raises a page fault when a page is not present or an access is not permitted.
Q7. Why are page tables multi-level?
A flat table must have an entry for every virtual page whether used or not: 4 MiB per process for 32-bit addresses and 512 GiB for 48-bit addresses with 8-byte entries. A multi-level tree only allocates lower-level tables for regions that are actually mapped, so sparse address spaces cost very little. The price is that a walk needs one memory access per level.
Q8. What is a TLB, and what happens on a TLB miss on x86?
The TLB is a small, fast cache of virtual-to-physical translations with permission bits. On a miss, a hardware page walker reads the page-table levels (often hitting in the data caches or page-walk caches), installs the translation in the TLB and retries the access. If the entry is not present or the access is not permitted, it raises a page fault for the OS.
Q9. What is TLB reach, and how do huge pages improve it?
TLB reach is entries × page size: the memory you can access without a TLB miss. 1,536 entries with 4 KiB pages cover 6 MiB; with 2 MiB pages they cover 3 GiB. Huge pages also shorten each page walk by one level. The cost is coarser allocation, possible internal fragmentation, and the need for contiguous physical memory.
Q10. With a 98 percent TLB hit ratio, 1 ns TLB, 100 ns memory and a 4-level page table, what is the effective access time?
A hit costs 1 + 100 = 101 ns. A miss costs 1 + 4 × 100 + 100 = 501 ns. EAT = 0.98 × 101 + 0.02 × 501 = 109 ns, assuming no page-table entries are cached.
Q11. What are ASIDs or PCIDs for?
They tag TLB entries with the address space they belong to. Without them, the OS must flush the TLB on every context switch because the same virtual address means different things in different processes. With them, entries from several processes can coexist, so switching back to a process finds its translations still warm.
Q12. What is a VIPT cache, and what constraint does it impose?
A virtually indexed, physically tagged cache uses page-offset bits of the virtual address to select the set while the TLB translates the page number in parallel, then compares the physical tag. Because the index bits come from the page offset, which is identical in the VA and PA, there are no aliasing problems, provided cache size ÷ associativity is at most the page size. That is why a 32 KiB L1 with 4 KiB pages is typically 8-way.
Q13. What is the difference between NAND and NOR flash?
NOR flash supports fast random reads at byte granularity and can execute code in place, so it suits small firmware chips. NAND flash is much denser and cheaper, reads and writes in pages and erases in blocks, so it suits bulk storage like SSDs. Both wear out after a limited number of erase cycles.
Q14. Why can't an SSD overwrite a page in place?
Flash cells must be erased before they are reprogrammed, and erase works only on whole blocks containing many pages. So the SSD's flash translation layer writes new data to an already erased page, remaps the logical address to it and marks the old page invalid; garbage collection later reclaims blocks. This also lets the FTL spread wear evenly.
Key takeaways
- SRAM (6-transistor latch, no refresh, fast, costly) builds caches; DRAM (1 transistor + 1 capacitor, dense, refreshed, destructive reads) builds main memory.
- DRAM access = activate a row into the row buffer, read a column, precharge; row hits are much cheaper than row conflicts.
- DDR4 refreshes every row within 64 ms using 8,192 commands, one about every 7.8 microseconds.
- DDR generations raised bandwidth far more than they cut latency; peak bandwidth = MT/s × bytes per transfer (25.6 GB/s for one DDR4-3200 channel).
- Low-order interleaving overlaps bank accesses and nearly matches a wide bus for sequential streams.
- The MMU translates VPN to PFN via a TLB and, on a miss, a hardware walk of a multi-level page table (9+9+9+9+12 on x86-64).
- TLB reach = entries × page size; huge pages multiply it by 512 and shorten walks.
- EAT with a TLB must include the walk: each level is one dependent memory access.
- VIPT L1 caches overlap TLB lookup with cache indexing, constraining size ÷ ways to at most the page size.
- Flash must be erased in blocks before writing; SSD firmware hides this with an FTL, wear levelling and garbage collection.
Next lesson
Continue with I/O organization.

