Why I/O organisation matters
A processor that cannot talk to the outside world is useless. Keyboards, network cards, SSDs, GPUs and sensors all need to move data into and out of memory, and they run at wildly different speeds: a keyboard produces a few bytes per second, a modern SSD several gigabytes per second. I/O organisation is the set of hardware mechanisms that connect these devices to the processor and memory without wasting the processor's time.
This lesson covers how the processor addresses a device (memory-mapped versus isolated I/O), the three ways of moving data (programmed I/O, interrupts and DMA), how interrupts are prioritised, how buses carry signals and who gets to use them, and the interfaces you meet today: PCI Express, USB, SATA and NVMe. Interviewers commonly ask you to compare polling, interrupts and DMA, to explain what a DMA controller does and the difference between burst mode and cycle stealing, and to explain why NVMe is faster than SATA. The worked example puts real numbers on how much CPU time each I/O method costs.
The operating-system side (device drivers, I/O scheduling, buffering) is covered in storage and I/O.
The I/O interface: controllers and ports
A device is never wired straight to the processor. Between them sits a device controller (also called an I/O interface or adapter): a piece of hardware that understands the device's electrical signals and timing on one side and the system bus on the other. A disk controller, a USB host controller and a network interface card (NIC) are all examples.
The processor talks to a controller through a small set of registers, usually called ports:
| Register | Direction | Purpose |
|---|---|---|
| Data register (in/out) | both | holds the byte or word being transferred |
| Status register | device to CPU | flags such as busy, ready, error, data available |
| Control (command) register | CPU to device | start an operation, set a mode, enable interrupts |
+--------+ system bus +------------------+ +--------+
| CPU |<=========================>| device controller|<--->| device |
+--------+ address / data / | data register | +--------+
^ control lines | status register |
| | control register |
+--------+ +------------------+
| memory |
+--------+
So the question "how does the CPU do I/O?" really means "how does the CPU read and write these registers, and how does it know when to?"
Addressing devices: memory-mapped versus isolated I/O
Memory-mapped I/O
In memory-mapped I/O (MMIO), device registers are assigned addresses in the same physical address space as memory. Ordinary load and store instructions access them. The address decoder routes, say, addresses 0xFE000000-0xFE000FFF to a device instead of DRAM.
Physical address space with memory-mapped I/O
0x00000000 +--------------------+
| DRAM |
| |
0xC0000000 +--------------------+
| PCIe device BARs | <- device registers and buffers
0xFEC00000 +--------------------+
| interrupt ctrl etc |
0xFFFFFFFF +--------------------+
Advantages: no special instructions; every addressing mode and the whole register file can be used; the operating system can protect devices with ordinary page tables and even map a device into a user process. ARM and RISC-V use MMIO exclusively, and PCIe devices on x86 use it too.
Two cautions. Device registers must not be cached (a cached status register would never show the device's update), so those pages are marked uncacheable in the page tables or memory-type registers. And the compiler must not remove or reorder accesses, which in C is why device registers are accessed through volatile pointers.
Isolated (port-mapped) I/O
In isolated I/O, devices live in a separate, smaller address space reached only by special instructions. x86 has a 64K-entry I/O port space accessed with IN and OUT. A control signal on the bus (conceptually an "I/O versus memory" line) tells the system which space an address belongs to.
Advantages: the full memory address space stays free for memory, and I/O instructions are easy to spot and restrict (x86 lets the OS control port access by privilege level). Disadvantages: fewer addressing modes and extra instructions in the ISA. On modern x86 PCs, port I/O is mostly legacy (old serial ports, some configuration mechanisms); high-speed devices use MMIO.
| Memory-mapped I/O | Isolated I/O | |
|---|---|---|
| Address space | shared with memory | separate I/O space |
| Instructions | ordinary loads and stores | special (IN/OUT on x86) |
| Protection | page tables | I/O privilege checks |
| Uses up memory addresses | yes | no |
| Where used | ARM, RISC-V, PCIe devices everywhere | x86 legacy devices |
/* Memory-mapped I/O in C: a hypothetical UART at a fixed address */
#define UART_BASE 0x10000000u
#define UART_DATA (*(volatile unsigned char *)(UART_BASE + 0))
#define UART_STATUS (*(volatile unsigned char *)(UART_BASE + 5))
#define TX_READY 0x20
void uart_putc(char c) {
while ((UART_STATUS & TX_READY) == 0) {
/* busy-wait: this loop is programmed I/O by polling */
}
UART_DATA = (unsigned char)c;
}
This fragment is for a bare-metal system (for example QEMU's RISC-V virt board places a 16550-style UART at 0x10000000); on an ordinary OS a user program would crash touching that address.
Three ways to move data
1. Programmed I/O (polling)
In programmed I/O (PIO), the CPU does everything. It repeatedly reads the status register until the device is ready (polling, or busy-waiting), then moves each word between the data register and memory with its own loads and stores. The uart_putc function above is exactly this.
Programmed I/O for one word
CPU: read status -> not ready -> read status -> not ready -> ...
-> ready -> read data register -> store to memory -> repeat
PIO is simple and has the lowest latency when the device is almost always ready, because there is no interrupt overhead. Its problem is waste: while the CPU polls, it does no useful work, and for a slow device it may poll millions of times for one byte.
2. Interrupt-driven I/O
In interrupt-driven I/O, the CPU starts the operation and goes off to do other work. When the device is ready, its controller raises an interrupt request (IRQ) signal. At the end of the current instruction, the processor:
- Notices the pending interrupt (if interrupts are enabled).
- Saves enough state to resume later: at least the program counter and status flags.
- Identifies the source and jumps to the matching interrupt service routine (ISR), also called an interrupt handler.
- The ISR reads or writes the data register, acknowledges the interrupt and returns.
- A return-from-interrupt instruction restores the state and the interrupted program continues.
Program: ---instr---instr---instr [IRQ] instr---instr--->
| ^
v |
ISR: save state -> move data -> restore, return
The CPU no longer wastes time polling, but it still handles every word itself, and each interrupt costs hundreds to thousands of cycles of saving state, cache and pipeline disruption and returning. For a high-speed device that delivers millions of words per second, interrupt-per-word is unaffordable.
Identifying the source: polled versus vectored interrupts
With many devices sharing one interrupt line, how does the CPU know who interrupted?
- Polled (software) identification: one common handler reads each device's status register in turn to find the one requesting service. Simple but slow, and the order of checking sets the priority.
- Vectored interrupts: the interrupting device (or an interrupt controller) supplies a number, the interrupt vector, which indexes a table of handler addresses: the interrupt vector table (on x86 in protected and long mode, the IDT, interrupt descriptor table). The CPU jumps straight to the right handler.
Vectored interrupt
device -> interrupt controller --(vector 33)--> CPU
|
interrupt vector table v
+----+----------------+
| 32 | timer_isr |
| 33 | keyboard_isr | <- jump here
| 34 | ... |
+----+----------------+
Priority and nesting
Devices differ in urgency: a timer tick or a network card about to overflow its buffer matters more than a keyboard. Interrupt systems therefore assign priorities. A higher-priority interrupt may interrupt a running lower-priority handler (nested interrupts), while lower ones wait. An interrupt mask register lets software temporarily disable selected interrupts; non-maskable interrupts (NMI) cannot be disabled and are reserved for critical events such as hardware errors or watchdogs.
Priority is resolved in one of three ways:
- Software polling order (as above).
- A priority interrupt controller in hardware: the classic Intel 8259 PIC, and on modern systems the APIC (advanced programmable interrupt controller) on x86 or the GIC (generic interrupt controller) on ARM, which also route interrupts to different cores.
- Daisy chaining, described next.
Daisy chaining
In a daisy chain, all devices share one interrupt request line, and an interrupt acknowledge signal from the CPU passes from device to device in series. Electrical position equals priority.
shared INTR line (wired-OR)
+-----------+-------------+-------------+
| | | |
+--+--+ ACK +-+---+ ACK +-+---+ ACK +-+---+
| CPU |------>| D1 |------>| D2 |------>| D3 |
+-----+ +-----+ +-----+ +-----+
highest lowest priority
- One or more devices assert the shared request line.
- The CPU, when ready, sends acknowledge to the first device.
- A device that is requesting keeps the acknowledge (does not pass it on) and puts its vector on the data bus.
- A device that is not requesting passes the acknowledge to the next device.
Daisy chaining needs very little wiring and makes it easy to add devices, but priority is fixed by position, a device near the CPU can starve those further down, and a failed device can break the chain. Parallel priority (a separate request line per device into a priority encoder) is faster and flexible but needs more wires.
3. Direct memory access (DMA)
For bulk transfers, the CPU should not touch every word at all. Direct memory access (DMA) lets a DMA controller (or, on modern systems, the device itself acting as a bus master) move data between the device and memory directly, without the CPU.
The steps:
- The CPU (driver) programs the DMA controller with: the memory address, the byte count, the direction (read or write) and the device. Modern devices read a list of such requests, called descriptors, from memory, which allows scatter-gather transfers across non-contiguous pages.
- The CPU returns to other work.
- The DMA controller requests the bus (DMA request), the CPU or bus arbiter grants it (bus grant, sometimes HOLD/HLDA in older processors), and the controller moves the data, incrementing the address and decrementing the count.
- When the count reaches zero, the controller raises one interrupt to tell the CPU the whole block is done.
CPU --(1) program: addr, count, dir--> DMA controller
CPU does other work ...
DMA controller <==== data moves directly ====> memory
^ |
+--------- device ------------------+
DMA controller --(4) one interrupt: "done"--> CPU
DMA modes
- Burst (block) mode: the DMA controller takes the bus and keeps it until the whole block is transferred. Fastest for the transfer, but the CPU cannot use the bus for memory access during the burst.
- Cycle stealing: the controller grabs the bus for one word (one bus cycle) at a time, then gives it back. The CPU is delayed by a cycle here and there but never blocked for long. Good when the device is slow relative to the bus.
- Transparent (hidden) mode: the controller only uses the bus during cycles when the CPU is not using it. No CPU slowdown at all, but the transfer is slowest and the detection logic is more complex.
With caches, the CPU often does not need the bus anyway: most of its loads hit in the cache, so stolen cycles hurt less than they did on early machines.
DMA and caches
DMA writes go to memory, not to the CPU's cache. If the cache holds an old copy of that memory, the CPU may read stale data; if the cache holds a dirty line, DMA out of memory may send old data. Systems solve this either with coherent DMA (the device's accesses snoop the caches, as on most x86 servers) or by having the driver explicitly flush or invalidate cache lines around each transfer (common on smaller ARM systems). Interviewers like this follow-up.
Modern systems also put an IOMMU between devices and memory, translating device addresses and restricting each device to the memory it is allowed to touch, much as the MMU does for processes.
Worked example: CPU cost of each I/O method
A device delivers data at 2 MB/s (2 × 10^6 bytes per second) in 4-byte words. The CPU runs at 1 GHz (10^9 cycles per second). How much CPU time does each method use?
Words per second = 2,000,000 ÷ 4 = 500,000.
(a) Polling. Each poll (read status, test, branch, then move the word) costs 100 cycles. In the best case the CPU polls exactly once per word, at just the right moment.
cycles/s = 500,000 × 100 = 50,000,000
fraction = 50,000,000 / 1,000,000,000 = 5%
That 5 percent is a lower bound. A real busy-wait loop cannot predict when the device is ready, so while a transfer is in progress the CPU spends 100 percent of its time polling.
(b) Interrupt-driven. Each interrupt (entry, handler, moving one word, return) costs 500 cycles, one per word.
cycles/s = 500,000 × 500 = 250,000,000
fraction = 25%
The CPU is free between words, but a quarter of it goes to interrupt overhead. A faster device would need more than 100 percent: impossible.
(c) DMA. Transfers are in 4 KB blocks (4,096 bytes). Per block, the CPU spends 1,000 cycles setting up the DMA and 500 cycles handling the completion interrupt: 1,500 cycles.
blocks/s = 2,000,000 / 4,096 = 488.28
cycles/s = 488.28 × 1,500 = 732,422
fraction = 732,422 / 10^9 = 0.073%
| Method | CPU cycles per second | CPU fraction |
|---|---|---|
| Polling (ideal) | 50,000,000 | 5% (100% while busy-waiting) |
| Interrupt per word | 250,000,000 | 25% |
| DMA, 4 KB blocks | about 732,000 | about 0.07% |
DMA uses roughly 340 times less CPU than interrupts here, because the CPU's cost is paid once per block instead of once per word.
Bus cost of DMA. The memory bus still carries the data. Suppose each 4-byte word takes one 10 ns bus cycle in cycle-stealing mode: 500,000 × 10 ns = 5 ms of bus time per second, or 0.5 percent of the bus. In burst mode with an 8-byte, 100 MHz bus, one 4 KB block is 512 transfers × 10 ns = 5.12 microseconds during which the CPU cannot use the bus.
Interview tip
Summarise the comparison in one sentence each. Polling: CPU checks the device repeatedly; simplest, lowest latency, wastes CPU. Interrupts: device signals the CPU when ready; no wasted polling but the CPU still moves each word and pays a per-interrupt cost. DMA: a controller moves the whole block and interrupts once at the end; best for bulk transfers. Then add that very fast devices (NVMe, 100 Gb NICs) mix approaches, polling completion queues under high load to avoid interrupt storms.
Buses
A bus is a shared set of wires (or, in modern point-to-point links, a protocol over dedicated wires) that carries information between components.
Data, address and control lines
A classic system bus has three groups of lines:
- Data bus: carries the values being transferred. Its width (8, 16, 32, 64 bits) sets how many bits move per transfer.
- Address bus: carries the address of the memory location or device register. Its width sets how much can be addressed: 32 lines address 2^32 bytes = 4 GiB.
- Control bus: carries commands and timing: read, write, memory versus I/O, clock, interrupt request and acknowledge, bus request and grant, reset, wait/ready.
+-------+ +--------+ +-----------+
| CPU | | memory | | I/O ctrl |
+-+-+-+-+ +-+-+-+--+ +-+-+-+-----+
| | | | | | | | |
data =+=|=|=========+=|=|==========+=|=|====
address ===+=|===========+=|============+=|====
control =====+=============+==============+====
Bus bandwidth = width in bytes × transfers per second. A 64-bit (8-byte) bus at 100 MHz with one transfer per cycle carries 8 × 100 × 10^6 = 800 MB/s.
Buses are also classified by what they connect: the processor-memory bus (short, fast), I/O buses (longer, many device types, standardised), and a backplane that connects everything in older designs. Modern PCs have replaced the shared front-side bus with point-to-point links and on-chip networks, but the vocabulary is still examined.
Synchronous versus asynchronous buses
A synchronous bus includes a clock line, and every transaction follows a fixed protocol relative to the clock edges ("put the address on in cycle 1, data appears in cycle 3"). It is simple and fast, but every device must keep up with the clock, and clock skew (the clock arriving at different times at different points) limits the bus length and speed.
An asynchronous bus has no shared clock. It uses a handshake protocol, so each device runs at its own speed.
Asynchronous read handshake (four-phase)
master: put address, assert ReadReq ----+
slave: sees ReadReq, reads address, |
puts data, asserts Ack <---+
master: sees Ack, latches data,
deasserts ReadReq ----+
slave: sees ReadReq low, |
removes data, deasserts Ack <---+
| Synchronous | Asynchronous | |
|---|---|---|
| Timing | shared clock | request/acknowledge handshake |
| Speed | fast if all devices fast | adapts to each device |
| Length | limited by clock skew | can be longer |
| Logic | simple | more complex (handshake) |
| Example | memory buses, PCI (parallel, older) | older peripheral buses; handshake ideas live on in many protocols |
Bus arbitration
Only one device can drive a shared bus at a time. A device that wants to start transfers is a bus master; deciding among competing masters is bus arbitration.
- Daisy-chain arbitration: a grant line passes through devices in order, like the interrupt daisy chain. Cheap; fixed priority; possible starvation.
- Centralised parallel arbitration: each device has its own request and grant lines to a central arbiter, which can implement fixed priority, round-robin or other fairness rules. Used by classic PCI.
- Distributed arbitration by self-selection: each requesting device puts its identification code on shared arbitration lines; the highest code wins and every device can see who won. No central arbiter.
- Distributed arbitration by collision detection: devices transmit and detect collisions, then back off and retry, the scheme classic Ethernet used.
The goals are to give high-priority devices low latency while guaranteeing that every device eventually gets the bus (fairness).
PCI Express
PCI Express (PCIe) replaced the old shared parallel PCI bus. Instead of one bus shared by all, each device has a point-to-point serial link to a switch or to the processor's root complex. Shared parallel buses hit limits from clock skew and the electrical load of many devices; fast serial links avoid both.
- A lane is two differential pairs: one pair transmits, one receives, so each lane is full duplex.
- A link has x1, x2, x4, x8 or x16 lanes; bandwidth scales with lane count. Graphics cards typically use x16, NVMe SSDs x4.
- Data is sent as packets (transaction layer packets), with a layered protocol for reliability: transaction layer, data link layer (sequence numbers, CRC, retries) and physical layer.
- Each generation roughly doubles the per-lane rate while staying backward compatible.
| Generation | Raw rate per lane | Encoding | Usable per lane, each direction | x16 |
|---|---|---|---|---|
| 1.0 | 2.5 GT/s | 8b/10b | 250 MB/s | 4 GB/s |
| 2.0 | 5 GT/s | 8b/10b | 500 MB/s | 8 GB/s |
| 3.0 | 8 GT/s | 128b/130b | about 0.98 GB/s | about 15.8 GB/s |
| 4.0 | 16 GT/s | 128b/130b | about 1.97 GB/s | about 31.5 GB/s |
| 5.0 | 32 GT/s | 128b/130b | about 3.94 GB/s | about 63 GB/s |
GT/s means gigatransfers per second on the wire. The encoding overhead explains the gap: 8b/10b sends 10 bits for every 8 data bits (20 percent overhead), 128b/130b only about 1.5 percent. Later generations (6.0 and beyond) double the rate again using multi-level signalling. Packet headers reduce real throughput a little further.
Worked example: PCIe bandwidth
PCIe 4.0 x4: 16 GT/s × (128 ÷ 130) ÷ 8 bits per byte = 1.97 GB/s per lane; × 4 lanes = 7.88 GB/s each direction. That is the ceiling for an NVMe SSD on a Gen4 x4 slot, which is why fast Gen4 drives advertise sequential reads around 7 GB/s.
PCIe devices expose their registers through BARs (base address registers) that the firmware or OS maps into the physical address space: PCIe is memory-mapped I/O. Devices use DMA as bus masters and signal completion with MSI/MSI-X (message-signalled interrupts): an interrupt is delivered as a special memory write instead of a dedicated wire, which allows many vectors per device and routing to specific cores.
USB
USB (Universal Serial Bus) connects external peripherals with one standard connector family. Key ideas:
- Host-controlled: a host controller schedules all traffic; devices speak only when the host polls them. This keeps devices simple and cheap.
- Tiered star topology via hubs, up to 127 devices per host controller.
- Hot-plugging and enumeration: when a device is plugged in, the host assigns it an address and reads its descriptors to load the right driver.
- Transfer types: control (configuration), bulk (large reliable data, for example storage), interrupt (small periodic data with bounded latency, for example keyboards and mice, despite the name, these are polled by the host), and isochronous (guaranteed bandwidth, no retries, for example audio and video streams).
- Power delivered over the cable; USB Power Delivery negotiates higher power for charging.
| Version (marketing name) | Signalling rate |
|---|---|
| USB 1.1 Full Speed | 12 Mb/s |
| USB 2.0 High Speed | 480 Mb/s |
| USB 3.2 Gen 1 (formerly USB 3.0) | 5 Gb/s |
| USB 3.2 Gen 2 | 10 Gb/s |
| USB 3.2 Gen 2x2 | 20 Gb/s |
| USB4 | 20 or 40 Gb/s (USB4 version 2 adds 80 Gb/s) |
Note that USB rates are in bits per second, and the connector shape (Type-A, Type-C) is separate from the protocol version.
Storage interfaces: SATA versus NVMe
SATA (Serial ATA) was designed for hard disks. Its third revision runs at 6 Gb/s with 8b/10b encoding, so the payload ceiling is 6 × 0.8 ÷ 8 = 600 MB/s, and real SSDs top out around 550 MB/s. The host side uses AHCI (advanced host controller interface), which provides one command queue of 32 commands. That was ample for a spinning disk that can serve only one request at a time, but it throttles flash, which can serve many requests in parallel across its chips.
NVMe (Non-Volatile Memory Express) is a protocol designed for flash. It runs directly over PCIe (typically x4) and allows up to about 64K I/O queues, each up to 64K commands deep. Each CPU core can have its own submission and completion queue pair, so cores do not contend on a shared lock, and command submission needs very few register writes.
NVMe queue pairs (per core)
core 0: [submission queue] --> SSD --> [completion queue] --> core 0
core 1: [submission queue] --> SSD --> [completion queue] --> core 1
driver writes command to SQ in memory, rings a "doorbell" register
SSD fetches it by DMA, executes, writes result to CQ, sends MSI-X
| SATA (AHCI) | NVMe | |
|---|---|---|
| Physical link | SATA cable, 6 Gb/s | PCIe lanes (usually x4) |
| Practical peak | about 550 MB/s | several GB/s (about 7.9 GB/s ceiling on Gen4 x4) |
| Queues | 1 queue, 32 commands | up to about 64K queues, 64K commands each |
| Designed for | hard disks | flash |
| Typical latency overhead | higher (more register accesses, single queue) | lower |
| Form factors | 2.5-inch drive, mSATA, M.2 SATA | M.2, U.2, add-in card |
Common mistake
"M.2" is a connector and card shape, not a protocol. An M.2 SSD may be SATA or NVMe; the slot and drive must agree. Similarly, an NVMe drive in a PCIe 3.0 slot is limited to about 3.9 GB/s on x4.
Interview questions
Q1. What is the difference between memory-mapped and isolated I/O?
In memory-mapped I/O, device registers occupy addresses in the normal physical address space and are accessed with ordinary loads and stores; protection comes from page tables, and the registers must be marked uncacheable. In isolated I/O, devices have a separate address space accessed with special instructions such as x86 IN/OUT, which keeps memory addresses free but needs extra instructions. Modern high-speed devices, including all PCIe devices, use memory-mapped I/O.
Q2. Compare polling, interrupt-driven I/O and DMA.
Polling has the CPU repeatedly check a status register; it is simple and low-latency but wastes CPU time. Interrupt-driven I/O lets the device signal readiness so the CPU can do other work, but the CPU still moves each word and pays an overhead per interrupt. DMA lets a controller move an entire block between device and memory and interrupts once at the end, making it the choice for bulk transfers.
Q3. What steps does a processor take when an interrupt arrives?
It finishes the current instruction, checks that the interrupt is enabled and not masked, saves the program counter and status (and switches to kernel mode and stack), identifies the source, typically via a vector, and jumps to the handler through the vector table. The handler services the device, acknowledges the interrupt, and a return-from-interrupt instruction restores state.
Q4. What is a vectored interrupt?
The interrupting device or interrupt controller supplies a vector number identifying the source, and the CPU uses it as an index into a table of handler addresses. This avoids the slow alternative of polling every device to find the one that interrupted.
Q5. Explain daisy chaining and its drawbacks.
Devices share one interrupt request line, and the acknowledge signal passes through them in series; the first requesting device in the chain keeps the acknowledge and supplies its vector. It is cheap and easy to extend, but priority is fixed by physical position, distant devices can starve, and one faulty device can block those behind it.
Q6. What does a DMA controller need to be told to start a transfer?
The memory address (or a descriptor list for scatter-gather), the number of bytes, the direction (device to memory or memory to device) and which device or channel. It then arbitrates for the bus, transfers the data while updating its address and count, and raises an interrupt when the count reaches zero.
Q7. What is the difference between burst mode and cycle stealing in DMA?
In burst mode the DMA controller holds the bus until the whole block is transferred, which is fastest for the transfer but blocks the CPU's memory accesses for that time. In cycle stealing it takes the bus for one word at a time and releases it, slightly slowing the CPU over a longer period but never locking it out for long.
Q8. Why can DMA cause cache coherence problems, and how are they solved?
DMA reads and writes memory directly, while the CPU may hold stale or dirty copies of the same lines in its caches. Hardware-coherent systems make device accesses snoop the caches; otherwise the driver must flush dirty lines before a device reads memory and invalidate cached lines after a device writes memory.
Q9. Compare synchronous and asynchronous buses.
A synchronous bus uses a shared clock and fixed timing, making it simple and fast but requiring all devices to keep pace and limiting length because of clock skew. An asynchronous bus uses a request/acknowledge handshake so devices of different speeds can coexist, at the cost of more complex control logic and handshake delays.
Q10. What is bus arbitration, and name two schemes?
Arbitration decides which of several potential bus masters may use a shared bus next. Daisy-chain arbitration passes a grant signal through devices in priority order; centralised parallel arbitration has a request and grant line per device to an arbiter that can apply priority or round-robin. Distributed schemes such as self-selection or collision detection need no central arbiter.
Q11. What is a PCIe lane, and how is bandwidth calculated?
A lane is a pair of differential serial links, one in each direction, so it is full duplex. Bandwidth per direction is transfer rate × encoding efficiency ÷ 8 per lane, multiplied by the lane count: PCIe 4.0 x4 is 16 GT/s × 128/130 ÷ 8 × 4, about 7.9 GB/s.
Q12. Why is NVMe faster than SATA for SSDs?
SATA is limited to 6 Gb/s (about 600 MB/s of data) and its AHCI interface has a single 32-command queue designed for hard disks. NVMe runs over several PCIe lanes with multiple gigabytes per second of bandwidth and supports thousands of deep queues, one per core, with fewer register accesses per command. That matches flash's internal parallelism and cuts software overhead.
Q13. Why do very fast devices sometimes poll instead of using interrupts?
At very high request rates, the per-interrupt overhead dominates and can overwhelm a core (an interrupt storm). Polling a completion queue when work is known to be arriving avoids that overhead and lowers latency. Linux's NAPI for network cards and polled I/O modes for NVMe switch between interrupts at low load and polling at high load.
Q14. What are the USB transfer types?
Control transfers configure devices; bulk transfers carry large amounts of data reliably without timing guarantees, as for storage; interrupt transfers carry small, periodic data with bounded latency, as for keyboards; isochronous transfers reserve bandwidth for real-time streams like audio, with no retries. All are scheduled by the host controller.
Key takeaways
- Devices sit behind controllers that expose data, status and control registers.
- Memory-mapped I/O uses normal loads and stores (and must be uncached and
volatile); isolated I/O uses a separate space and special instructions. - Polling wastes CPU, interrupts pay a per-event cost, and DMA pays once per block; in the worked example, 25 percent versus 0.07 percent CPU.
- Vectored interrupts jump straight to the right handler; priority comes from controllers (APIC, GIC) or daisy-chain position; NMIs cannot be masked.
- DMA modes: burst (fast, blocks the bus), cycle stealing (one word at a time), transparent (idle cycles only); DMA must be kept coherent with caches.
- A bus has data, address and control lines; bandwidth = width × rate; synchronous buses use a clock, asynchronous ones a handshake; arbitration picks the master.
- PCIe is point-to-point serial lanes; usable rate roughly doubles per generation (about 2 GB/s per lane on Gen4).
- USB is host-scheduled with four transfer types; NVMe over PCIe beats SATA through bandwidth and many deep queues.
Next lesson
Continue with Parallelism and multicore.

