What this lesson covers
A processor has two halves. The datapath is the hardware that holds and transforms data: registers, the ALU, memories, and the wires and multiplexers connecting them. The control unit is the "brain" that reads each instruction and, step by step, tells the datapath what to do by setting control signals (for example "ALU, add now", "register file, write this value").
In the previous lessons you saw the building blocks (adders, MUXes, flip-flops) and the instruction set they must implement. This lesson connects the two: how an instruction actually flows through hardware. Interviewers and written tests typically probe:
- the roles of PC, IR, MAR, MDR, the accumulator and the flags register,
- micro-operations written in register transfer language (RTL), especially the fetch cycle,
- why a single-cycle design is slow and how a multi-cycle design differs,
- hardwired versus microprogrammed control, and horizontal versus vertical microcode, including control-memory size calculations,
- what the hardware does when an interrupt arrives.
Registers inside the CPU
Registers are the fastest storage in the machine: small groups of flip-flops inside the processor. Some are visible to programs; others exist only to make the hardware work.
| Register | Full name | What it holds |
|---|---|---|
| PC | Program counter | Address of the next instruction to fetch |
| IR | Instruction register | The instruction currently being decoded and executed |
| MAR | Memory address register | The address being sent to memory for the current read or write |
| MDR (or MBR) | Memory data register (memory buffer register) | The data just read from memory, or about to be written |
| AC (ACC) | Accumulator | Implicit operand and result in one-address machines |
| General-purpose registers | e.g. x0 to x31 in RISC-V, rax/rbx... in x86-64 | Operands and results chosen by the program |
| Flags / status register (PSW) | Program status word | Condition flags (zero, negative, carry, overflow), interrupt enable bit, current privilege mode |
| SP | Stack pointer | Top of the call stack (a general-purpose register by convention in RISC ISAs) |
MAR and MDR are the CPU's interface to memory. To read memory, the CPU places an address in MAR, asserts a read signal, and the data arrives in MDR. To write, it places the address in MAR, the data in MDR, and asserts write. In real modern processors these are buried inside the load/store unit and caches, but they remain the standard model in textbooks and exams.
Visible and invisible registers
Programmers can name general-purpose registers, the PC (indirectly, through jumps), the stack pointer and often the flags. IR, MAR and MDR are invisible: they are part of the organization, not the architecture.
Register transfer language (RTL)
Register transfer language is a compact notation for describing what happens in each clock cycle. The basic statement is a transfer:
R2 <- R1 copy R1 into R2
R3 <- R1 + R2 add R1 and R2, put the sum in R3
MDR <- M[MAR] read memory at address MAR into MDR
M[MAR] <- MDR write MDR to memory at address MAR
PC <- PC + 4 increment PC by 4
IR[addr] the address field of IR
A micro-operation is one such elementary transfer or operation that happens in one clock cycle. Transfers separated by commas happen in the same cycle, which is possible only if they use different hardware:
T1: MDR <- M[MAR], PC <- PC + 1
Conditional transfers use a colon: If (Z = 1): PC <- target (sometimes written Z: PC <- target).
Micro-operations for fetch and execute
Consider a simple accumulator machine with word-addressed memory (each instruction is one word, so the PC increments by 1). The time steps T0, T1, ... are clock cycles within one instruction.
The fetch cycle
T0: MAR <- PC
T1: MDR <- M[MAR], PC <- PC + 1
T2: IR <- MDR
Step by step:
- T0: copy the PC into MAR so memory knows which address to read.
- T1: memory returns the word into MDR. In the same cycle, increment the PC; it uses the PC's own incrementer, not the memory path, so the two do not conflict.
- T2: move the instruction into IR, where the decoder can see the opcode.
Why can't T0 and T1 be merged? Memory needs a stable address in MAR before it can start the read. Why not increment the PC in T0? Some designs do (T0: MAR <- PC, PC <- PC + 1 is valid if the PC update does not disturb MAR's input that cycle); exam answers usually use the three-step form above.
Execute: ADD X (accumulator machine)
ADD X means AC <- AC + M[X].
T3: MAR <- IR[addr]
T4: MDR <- M[MAR]
T5: AC <- AC + MDR
The address field of the instruction goes to MAR, memory supplies the operand, and the ALU adds it to the accumulator.
Other examples
LOAD X (AC <- M[X]):
T3: MAR <- IR[addr]
T4: MDR <- M[MAR]
T5: AC <- MDR
STORE X (M[X] <- AC):
T3: MAR <- IR[addr]
T4: MDR <- AC
T5: M[MAR] <- MDR
JUMP X: T3: PC <- IR[addr].
BRANCH-IF-ZERO X: T3: If (Z = 1): PC <- IR[addr].
Indirect addressing adds an extra memory read to fetch the real address: ADD @X needs MAR <- IR[addr], MDR <- M[MAR], MAR <- MDR, MDR <- M[MAR], then the add.
Register-to-register example
On a machine with general-purpose registers and an internal bus, ADD R1, R2, R3 (R1 = R2 + R3) might be:
T3: Y <- R2 (temporary register feeding the ALU)
T4: Z <- Y + R3 (ALU output latched in Z)
T5: R1 <- Z
With only one internal bus, only one register can drive it per cycle, so the operands must be moved one at a time through temporaries Y and Z. A three-bus datapath (two read buses and one write bus) could do R1 <- R2 + R3 in a single cycle. This is a direct example of organization trading hardware for time.
Interview tip
When asked to write micro-operations, say which register each step uses and why steps cannot be merged: "each bus carries one value per cycle, and memory needs a stable address before it can return data." That shows you understand the constraint, not just the memorized sequence.
Datapath components
A modern RISC datapath is built from a handful of components:
- Program counter: a 32- or 64-bit register.
- Instruction memory: given an address, outputs the instruction (in reality, the L1 instruction cache).
- Adder for PC + 4.
- Register file: 32 registers, two read ports (supply rs1 and rs2 values at once) and one write port (rd), written at the clock edge when
RegWriteis 1. - Immediate generator (sign extender): extracts the immediate bits from the instruction and sign-extends them to full width.
- ALU: performs the operation selected by the ALU control lines; outputs the result and a Zero flag.
- Data memory: read when
MemReadis 1, written whenMemWriteis 1 (the L1 data cache). - Multiplexers that choose between alternative sources.
- Branch target adder: computes PC + offset.
The single-cycle datapath
In a single-cycle implementation, every instruction completes in exactly one (long) clock cycle. Every component is used at most once per instruction, which is why there are separate instruction and data memories (a Harvard arrangement) and separate adders for PC + 4 and the branch target.
+-----+ +----------+ +----------+ +-----+ +--------+
+-->PC--+ IM +-->| Register |-->| MUX -> |-->| ALU |-->| Data |
| | | | file | | (ALUSrc) | | | | memory |
| +-----+ +----------+ +----------+ +--+--+ +---+----+
| ^ imm ---> ^ | |
| | v v
| +---------- MUX (MemtoReg) <--+---------+
|
+--- MUX (PCSrc) <-- PC + 4
<-- PC + imm (taken when Branch AND Zero)
The main multiplexers and the signals that control them:
| Signal | Meaning when 1 | Meaning when 0 |
|---|---|---|
RegWrite | Write the result into register rd | Do not write |
ALUSrc | ALU's second input is the sign-extended immediate | Second input is register rs2 |
MemRead | Read data memory | No read |
MemWrite | Write data memory | No write |
MemtoReg | Value written back comes from data memory | Value comes from the ALU |
Branch | Instruction is a branch; take it if ALU Zero is 1 | Not a branch |
ALUOp (2 bits) | Tells ALU control: add (loads/stores), subtract (branch compare), or "look at funct fields" (R-type) |
Control signal settings for four instruction types (X = don't care):
| Instruction | RegWrite | ALUSrc | MemRead | MemWrite | MemtoReg | Branch | ALU operation |
|---|---|---|---|---|---|---|---|
R-type (add) | 1 | 0 | 0 | 0 | 0 | 0 | from funct fields |
lw | 1 | 1 | 1 | 0 | 1 | 0 | add |
sw | 0 | 1 | 0 | 1 | X | 0 | add |
beq | 0 | 0 | 0 | 0 | X | 1 | subtract |
Why is MemtoReg a don't care for sw and beq? Because RegWrite is 0, so whatever value the MUX picks is never written.
Tracing lw x5, 8(x2) through the single-cycle datapath
- PC sends the address to instruction memory; the instruction comes out.
- Bits for rs1 (x2) go to the register file, which outputs x2's value. The immediate generator produces 8.
ALUSrc= 1 selects the immediate; the ALU adds x2 + 8 to form the address.MemRead= 1: data memory reads the word at that address.MemtoReg= 1 selects the memory output;RegWrite= 1 writes it into x5 at the clock edge.- Meanwhile PC + 4 has been computed;
Branch= 0, so the PC MUX selects PC + 4 for the next cycle.
All of this happens within one clock period, flowing through combinational logic, and the state elements (PC, register file, data memory) update at the clock edge.
Worked example: the cost of a single cycle
Assume these component delays (ignoring MUX and wiring delays for simplicity):
| Component | Delay |
|---|---|
| Instruction memory read | 200 ps |
| Register file read | 100 ps |
| ALU | 200 ps |
| Data memory access | 200 ps |
| Register file write (setup) | 100 ps |
Each instruction type uses a different path:
| Instruction | Path | Total |
|---|---|---|
| Load | IM + Reg read + ALU + DM + Reg write = 200 + 100 + 200 + 200 + 100 | 800 ps |
| Store | IM + Reg read + ALU + DM = 200 + 100 + 200 + 200 | 700 ps |
| R-type | IM + Reg read + ALU + Reg write = 200 + 100 + 200 + 100 | 600 ps |
| Branch | IM + Reg read + ALU = 200 + 100 + 200 | 500 ps |
| Jump | IM (plus address formation) | 200 ps |
The clock period must accommodate the slowest instruction, the load: 800 ps (1.25 GHz). Every instruction takes 800 ps, even a jump that needs only 200 ps.
How much is wasted? Suppose the program mix is 25% loads, 10% stores, 45% R-type, 15% branches and 5% jumps. If each instruction could take exactly as long as it needed:
Average = 0.25x800 + 0.10x700 + 0.45x600 + 0.15x500 + 0.05x200
= 200 + 70 + 270 + 75 + 10
= 625 ps
So a single fixed 800 ps cycle is 800 / 625 = 1.28 times slower than an ideal variable-length cycle.
Limits of single-cycle design
- The slowest instruction sets the clock. Common, fast instructions pay for rare, slow ones. Adding a slow instruction (say a multiply or floating-point operation) would slow down everything.
- Duplicated hardware. Each unit can be used once per cycle, so you need separate instruction and data memories and extra adders.
- No overlap. Most of the hardware sits idle most of the cycle: while data memory is being accessed, instruction memory and the ALU have finished and are idle.
Single-cycle designs are used for teaching, not for real processors. The two ways forward are multi-cycle execution and pipelining.
Multi-cycle implementation
A multi-cycle implementation breaks each instruction into steps, each taking one short clock cycle. Different instructions take different numbers of cycles. The clock period is set by the slowest step, not the slowest instruction.
A typical RISC breakdown uses up to five steps:
| Step | Name | Work |
|---|---|---|
| 1 | IF | IR <- Mem[PC], PC <- PC + 4 |
| 2 | ID | Read registers into temporaries A and B; compute branch target speculatively |
| 3 | EX | ALU operation, address calculation, or branch comparison (branch completes here) |
| 4 | MEM | Load reads memory into MDR, or store writes memory (store completes here); R-type writes back here |
| 5 | WB | Load writes MDR into the register file |
Cycles per instruction type: load 5, store 4, R-type 4, branch 3, jump 3.
Key differences from single-cycle:
- Hardware is reused across cycles. A single memory can serve both instruction fetch and data access (in different cycles), and the one ALU can compute PC + 4 in one cycle and the operation result in another.
- Extra internal registers (IR, MDR, A, B, ALUOut) hold values between cycles, because combinational results disappear when the inputs change.
- The control unit becomes a finite state machine: each state corresponds to one step and asserts that step's control signals; the next state depends on the opcode.
Worked example: multi-cycle versus single-cycle
Using the same component delays, the longest step is 200 ps, so the multi-cycle clock is 200 ps (ignoring the extra register overhead, which would add a little).
With the same mix (25% loads, 10% stores, 45% R-type, 15% branches, 5% jumps):
CPI = 0.25x5 + 0.10x4 + 0.45x4 + 0.15x3 + 0.05x3
= 1.25 + 0.40 + 1.80 + 0.45 + 0.15
= 4.05
Average time per instruction = 4.05 x 200 ps = 810 ps
Surprise: 810 ps is slightly worse than the single-cycle 800 ps. With these balanced step times, every instruction pays for whole 200 ps steps even when a step (register read or write, 100 ps) needs only half.
Now add a slow instruction. Suppose the ISA includes a multiply whose ALU operation takes 1600 ps, and 10% of instructions are multiplies (mix: 20% loads, 10% stores, 40% R-type, 15% branches, 5% jumps, 10% multiplies).
- Single-cycle: the multiply path is IM + Reg read + 1600 + Reg write = 200 + 100 + 1600 + 100 = 2000 ps. The clock must be 2000 ps for every instruction.
- Multi-cycle: multiply takes IF + ID + 8 EX cycles + WB = 11 cycles; the clock stays 200 ps.
CPI = 0.20x5 + 0.10x4 + 0.40x4 + 0.15x3 + 0.05x3 + 0.10x11
= 1.00 + 0.40 + 1.60 + 0.45 + 0.15 + 1.10
= 4.70
Time per instruction = 4.70 x 200 ps = 940 ps
Speedup over single-cycle = 2000 / 940 = 2.13
The lesson: multi-cycle wins when instruction latencies vary widely, because slow instructions no longer slow down fast ones. It also saves hardware. But neither design overlaps instructions. Pipelining keeps the short clock of multi-cycle while starting a new instruction every cycle, pushing CPI toward 1; with a 200 ps clock that approaches 200 ps per instruction, four times better than the single-cycle 800 ps. That is the next lesson.
| Aspect | Single-cycle | Multi-cycle | Pipelined |
|---|---|---|---|
| CPI | 1 | 3 to 5 (more for complex instructions) | Close to 1 |
| Clock period | Slowest instruction | Slowest step | Slowest stage plus latch overhead |
| Hardware reuse | None; duplicates units | Units reused across steps | Units used by different instructions at once |
| Control | Combinational decoder | FSM or microprogram | Combinational, with hazard logic |
| Overlap of instructions | No | No | Yes |
The control unit
The control unit turns each instruction into the right sequence of control signals. There are two ways to build it.
Hardwired control
In hardwired control, the control signals are produced by fixed combinational logic driven by the opcode, the current step (from a step counter or FSM state), and status flags.
+------------------+
IR opcode -->| |
step count ->| combinational |--> control signals
flags -->| logic (gates, | (RegWrite, ALUOp,
| decoders) | MemRead, ...)
+------------------+
For example, in the accumulator machine, the signal "load AC from ALU" might be:
LoadAC = T5 · (ADD + SUB + AND) + T5 · LOAD
meaning: at step T5, if the instruction is ADD, SUB, AND or LOAD, enable the accumulator's load input.
- Advantages: fast (signals come straight out of gates), efficient for a small, regular instruction set.
- Disadvantages: hard to design and verify for a large, complex ISA; any change means redesigning the circuit.
Microprogrammed control
In microprogrammed control, the control signals for each step are stored as words in a small, fast memory inside the CPU, the control memory (or control store). Each word is a microinstruction; a sequence of them that implements one machine instruction is a microroutine; the whole contents are the microprogram (or microcode). Maurice Wilkes proposed this idea in 1951.
+----------------+
opcode ------> | mapping logic |---+
+----------------+ |
v
+-------+ +--------------------------+
| μPC |-----> | Control memory (ROM) |
+-------+ +--------------------------+
^ |
| microinstruction
| +---------+-----------+
| | control | next |
| | fields | address |
| +-----+-----+----+----+
| | |
| v v
| control signals sequencing logic
+---------------------------- (uses flags)
How it runs:
- A micro-program counter (μPC) points to the current microinstruction.
- The microinstruction's control bits drive the datapath for one cycle.
- Sequencing logic picks the next microinstruction: the next one in order, a branch within the microroutine (possibly depending on flags), or, after fetch, a jump to the start of the microroutine for the opcode in IR (found by a mapping table or logic).
- Advantages: systematic and flexible. Complex instructions are just longer microroutines. Bugs can be fixed or instructions added by changing the microcode (modern x86 processors accept microcode updates from the BIOS or operating system, used, for example, to apply fixes for hardware bugs and security vulnerabilities).
- Disadvantages: slower, since each step needs a control-memory read; costs chip area for the control store.
Hardwired versus microprogrammed
| Aspect | Hardwired | Microprogrammed |
|---|---|---|
| How signals are produced | Combinational logic | Read from control memory |
| Speed | Faster | Slower (control memory access each step) |
| Flexibility | Changes require redesign | Change the microcode |
| Design complexity for big ISAs | Very high | Manageable, systematic |
| Hardware cost | Grows with ISA complexity | Mostly the control store |
| Suits | RISC (simple, regular instructions) | CISC (many complex instructions) |
| Examples | Classic RISC pipelines (MIPS, RISC-V cores) | IBM System/360 models, older x86; modern x86 for complex or rare instructions |
Modern x86 processors are a hybrid: common instructions are decoded by hardwired decoders into one or a few micro-operations, while complex, rare ones (some string and system instructions) are handled by a microcode sequencer.
Horizontal versus vertical microcode
There are two ways to lay out the bits of a microinstruction.
Horizontal microinstructions have one bit per control signal. If the datapath has 40 control signals, the control field is 40 bits wide. Any combination of signals can be active in the same cycle, giving maximum parallelism and no decoding delay, but the words are wide.
Vertical microinstructions encode groups of signals. Signals that can never be active together (for example "ALU add" and "ALU subtract", or the different sources that can drive one bus) form a group, and the group is stored as a binary code that a decoder expands. A group of k mutually exclusive signals needs ceil(log2(k + 1)) bits (the +1 reserves a code for "none of them").
| Aspect | Horizontal | Vertical |
|---|---|---|
| Encoding | One bit per signal | Encoded fields, decoded on use |
| Word width | Wide | Narrow |
| Control memory size | Larger | Smaller |
| Parallelism per microinstruction | High | Lower (one signal per group) |
| Speed | Faster (no decoding) | Slower (decoders add delay) |
| Microcode length | Fewer microinstructions | Often more microinstructions |
Worked example: control memory size
A CPU has 40 control signals and a control memory of 256 microinstructions. Each microinstruction also holds a next-address field.
- Next-address field: 256 words need
log2(256)= 8 bits. - Horizontal: control field = 40 bits. Word = 40 + 8 = 48 bits. Control memory = 256 × 48 = 12,288 bits.
- Vertical: suppose the 40 signals split into 4 groups of mutually exclusive signals with 7, 12, 3 and 18 signals.
- Group of 7:
ceil(log2 8)= 3 bits - Group of 12:
ceil(log2 13)= 4 bits - Group of 3:
ceil(log2 4)= 2 bits - Group of 18:
ceil(log2 19)= 5 bits - Control field = 3 + 4 + 2 + 5 = 14 bits. Word = 14 + 8 = 22 bits. Control memory = 256 × 22 = 5,632 bits.
- Group of 7:
Vertical encoding cuts the control store by more than half here, at the cost of four decoders and less parallelism (only one signal per group can be active per microinstruction).
Common mistake
Forgetting the "none" code when encoding a group. A group of 7 signals fits in 3 bits only because 7 + 1 = 8 codes are needed; a group of 8 signals needs 4 bits, since 9 codes are required. Many exam questions turn on this detail; read whether the question says to reserve a no-operation code.
Interrupt handling in hardware
An interrupt is a signal that asks the CPU to stop what it is doing, run a special routine (the interrupt service routine (ISR) or handler), and then resume. Interrupts let devices get the CPU's attention without the CPU constantly polling them.
Terminology varies between textbooks and ISAs, but the common categories are:
- Hardware interrupts (external, asynchronous): from devices or timers (a disk transfer finished, a key was pressed, the timer tick for the scheduler). They are unrelated to the instruction currently executing.
- Exceptions (internal, synchronous): caused by the current instruction: divide by zero, page fault, illegal opcode.
- Traps / software interrupts: deliberately requested by the program, such as
ecallon RISC-V orsyscallon x86-64 to request an operating-system service.
What the hardware does
- Request: a device raises its interrupt-request line (or sends a message on the interconnect).
- Check at an instruction boundary: after finishing the current instruction (in the simple model), the CPU checks whether an interrupt is pending and whether interrupts are enabled (not masked by the interrupt-enable flag or a priority mask).
- Save state: the hardware saves at least the PC (the address to resume at) and the status register. Depending on the ISA, this goes to the stack or to dedicated registers. RISC-V saves the PC in
mepc/sepcand the reason inmcause/scause; x86 pushes the instruction pointer, flags and more onto the stack. - Disable further interrupts (at least of the same or lower priority) and switch to privileged mode.
- Identify the source: with vectored interrupts, the device or interrupt controller supplies a number that indexes an interrupt vector table holding handler addresses. With non-vectored interrupts, one common handler polls devices to find out which one interrupted.
- Jump: load the handler address into PC.
- Handler runs (software): saves any additional registers it will use, services the device, acknowledges the interrupt, restores registers.
- Return: a special instruction (
mret/sreton RISC-V,ireton x86) restores the saved PC and status register, re-enabling interrupts, and the interrupted program continues as if nothing happened.
In RTL, the hardware's interrupt cycle on a simple machine (saving PC to a fixed memory location) might look like:
T0: MDR <- PC
T1: MAR <- SAVE_ADDRESS, PC <- HANDLER_ADDRESS
T2: M[MAR] <- MDR, IEN <- 0 (disable interrupts)
Then the normal fetch cycle begins, fetching the first handler instruction.
Priority and nesting
When several devices interrupt at once, an interrupt controller decides priority, typically with a priority encoder or a daisy chain (devices wired in series, where the closest device to the CPU gets the acknowledgement first). With nested interrupts, a handler can re-enable higher-priority interrupts so that urgent events (for example a timer) are not delayed by a long, low-priority handler.
Non-maskable interrupts (NMI) cannot be disabled; they are reserved for critical events such as hardware errors.
Precise interrupts
In a pipelined or out-of-order processor, several instructions are in flight at once. An interrupt is precise if the saved state looks as if all instructions before the saved PC completed and none after it started. Precise interrupts are essential for page faults (the faulting instruction must be restartable after the OS loads the page) and are a major source of complexity in modern CPU design; the reorder buffer in out-of-order processors exists partly to provide them.
Interview tip
Distinguish interrupt from exception: an interrupt is asynchronous (comes from outside, unrelated to the current instruction), an exception is synchronous (caused by the current instruction). Then walk through "finish instruction, check enabled, save PC and status, vector to handler, return with a special instruction".
Interview questions
Q1. What are the roles of PC, IR, MAR and MDR?
The PC holds the address of the next instruction; the IR holds the instruction being executed so the control unit can decode it. MAR holds the address sent to memory, and MDR holds the data read from or written to memory. Together, MAR and MDR form the CPU side of the memory interface.
Q2. Write the micro-operations of the instruction fetch cycle.
T0: MAR <- PC; T1: MDR <- M[MAR], PC <- PC + 1 (or + 4 for byte-addressed 32-bit instructions); T2: IR <- MDR. The PC increment can share a cycle with the memory read because it uses different hardware. After T2 the opcode in IR selects the execute sequence.
Q3. What is the difference between the datapath and the control unit?
The datapath contains the components that store and process data: registers, ALU, memories, MUXes and buses. The control unit decodes instructions and generates the signals that select MUX inputs, ALU operations, and register and memory reads and writes in each cycle. The datapath is the muscle and control is the nervous system.
Q4. Why is a single-cycle implementation inefficient?
The clock period must fit the slowest instruction (usually a load), so faster instructions waste most of the cycle. It also duplicates hardware, needing separate instruction and data memories and extra adders, since each unit can be used only once per cycle. And no instructions overlap, so most hardware is idle most of the time.
Q5. How does a multi-cycle datapath differ, and is it always faster?
It splits each instruction into short steps, one per clock cycle, with the clock set by the slowest step; instructions take 3 to 5 or more cycles. It reuses one ALU and one memory across steps and needs intermediate registers. It is not always faster: with balanced steps its average time per instruction can match or exceed single-cycle, but it wins clearly when some instructions are much slower than others.
Q6. Compare hardwired and microprogrammed control.
Hardwired control produces signals from combinational logic driven by opcode, step and flags; it is fast but hard to change or scale to a large ISA. Microprogrammed control stores the signals for each step in a control memory and steps through microroutines; it is slower but systematic and patchable. RISC designs use hardwired control; CISC designs traditionally used microcode, and modern x86 uses both.
Q7. What is the difference between horizontal and vertical microinstructions?
Horizontal microinstructions have one bit per control signal, so they are wide, fast and highly parallel. Vertical microinstructions encode groups of mutually exclusive signals into short fields that are decoded, so they are narrow and save control memory but add decoding delay and allow less parallelism. A group of k signals needs ceil(log2(k + 1)) bits when a "none" code is reserved.
Q8. A CPU has 32 control signals and 512 microinstructions. What is the horizontal control memory size?
The next-address field needs log2(512) = 9 bits. Each horizontal microinstruction is 32 + 9 = 41 bits. The control memory is 512 × 41 = 20,992 bits. If the design uses sequential addressing without a next-address field, it would be 512 × 32 = 16,384 bits, so state your assumption.
Q9. What are control signals like RegWrite, ALUSrc and MemtoReg?
RegWrite enables writing the register file. ALUSrc selects whether the ALU's second operand is a register or the sign-extended immediate. MemtoReg selects whether the value written back comes from data memory (loads) or the ALU (arithmetic). Branch instructions set Branch so the PC takes the target when the ALU's zero output says the operands are equal.
Q10. What happens in hardware when an interrupt occurs?
At an instruction boundary, if interrupts are enabled, the CPU saves the PC and status register, disables further interrupts and switches to privileged mode. It identifies the source, usually through a vector table, and loads the handler's address into the PC. The handler services the device, and a return-from-interrupt instruction restores the saved PC and status.
Q11. What is the difference between an interrupt, an exception and a trap?
An interrupt is asynchronous and comes from outside the current instruction, such as a device or timer. An exception is a synchronous error caused by the current instruction, such as a page fault or divide by zero. A trap is a deliberate synchronous transfer to the OS, such as a system call instruction. Exact naming differs between ISAs, so define your terms.
Q12. What is a vectored interrupt?
The interrupting device or the interrupt controller supplies an identifier that indexes an interrupt vector table, giving the handler address directly. That avoids polling every device to find the source. x86's interrupt descriptor table and ARM's vector table are examples.
Q13. What are precise interrupts and why do they matter?
An interrupt is precise when the saved state reflects all instructions before the saved PC completed and none after it took effect. This lets the OS fix the cause (for example load a missing page) and restart the faulting instruction correctly. Pipelined and out-of-order CPUs need extra hardware, such as reorder buffers, to guarantee it.
Q14. Why do single-bus datapaths need temporary registers such as Y and Z?
A bus carries only one value per cycle, but the ALU needs two operands and produces a result. One operand is first parked in Y, the second arrives over the bus while the ALU computes, and the result is captured in Z before being moved to its destination. Multi-bus datapaths remove these extra steps at the cost of more wiring.
Key takeaways
- The datapath stores and transforms data; the control unit generates the signals that drive it each cycle.
- PC, IR, MAR, MDR, the accumulator and the status register each have a fixed role; MAR and MDR are the memory interface.
- RTL describes micro-operations per cycle; fetch is
MAR <- PC,MDR <- M[MAR], PC <- PC + 1,IR <- MDR. - Single-cycle designs use CPI 1 but a clock set by the slowest instruction and duplicated hardware.
- Multi-cycle designs use a short clock set by the slowest step and reuse hardware; they pay off when instruction latencies vary widely, and lead naturally to pipelining.
- Hardwired control is fast and suits RISC; microprogrammed control is flexible and suits CISC; modern x86 combines them.
- Horizontal microcode uses one bit per signal; vertical encodes exclusive groups in ceil(log2(k + 1)) bits, trading speed for a smaller control store.
- On an interrupt the hardware finishes the current instruction, checks the enable mask, saves PC and status, vectors to the handler, and later returns with a special instruction; pipelined CPUs must keep interrupts precise.
Next lesson
Continue with Pipelining.

