What an ISA is and why it matters
The instruction set architecture (ISA) is the contract between software and hardware. It defines everything a compiler or assembly programmer must know to produce a working program, and nothing about how the chip is built inside. It specifies:
- the instructions (add, load, branch, and so on) and what each does,
- the registers visible to programs, and their sizes,
- the data types (8/16/32/64-bit integers, floats),
- the instruction formats: how each instruction is encoded as bits,
- the addressing modes: how instructions name their operands,
- the memory model: byte addressing, alignment, endianness,
- how exceptions and interrupts are reported, and privilege levels.
Because the contract is fixed, a binary compiled for x86-64 in 2010 still runs on a 2026 x86-64 chip with a completely different organization. The ISA also shapes performance through the CPU time equation: it influences instruction count (how much work each instruction does) and how easy it is to build a fast, pipelined implementation (CPI and cycle time).
Interviewers typically probe addressing modes (very commonly with a numerical effective-address question), instruction formats and expanding opcodes, RISC versus CISC with real examples, and above all what happens at the machine level when a function is called: the stack, stack frames, saved registers and return addresses. That last topic links directly to recursion depth, stack overflow and buffer-overflow attacks, so it shows up in software interviews too.
Anatomy of an instruction
A machine instruction is a binary word divided into fields:
- Opcode: which operation to perform.
- Operand specifiers: where the inputs come from and where the result goes. Each can be a register number, a memory address, or a constant (an immediate value) embedded in the instruction.
- Sometimes extra fields that refine the opcode (RISC-V calls them
funct3andfunct7).
Example: RISC-V add x7, x5, x6 (x7 = x5 + x6) is a 32-bit R-type (register) instruction:
31 25 24 20 19 15 14 12 11 7 6 0
+----------+-------+-------+------+-------+---------+
| funct7 | rs2 | rs1 |funct3| rd | opcode |
| 0000000 | 00110 | 00101 | 000 | 00111 | 0110011 |
+----------+-------+-------+------+-------+---------+
x6 x5 add x7 OP (ALU)
Concatenated: 0000 0000 0110 0010 1000 0011 1011 0011 = 0x006283B3. Note that register fields are 5 bits, because RISC-V has 32 integer registers (2^5 = 32).
A load, lw x5, 0(x2), uses the I-type format, where a 12-bit signed immediate replaces rs2 and funct7:
+--------------+-------+------+-------+---------+
| imm[11:0] | rs1 |funct3| rd | opcode |
| 000000000000 | 00010 | 010 | 00101 | 0000011 |
+--------------+-------+------+-------+---------+
That encodes to 0x00012283. A 12-bit signed immediate covers −2048 to 2047, which is why constants outside that range need two instructions (lui to set the upper 20 bits, then addi).
Keeping fields in the same positions across formats (rs1, rs2 and rd sit in the same bits whenever present) lets the hardware start reading registers before it has fully decoded the instruction. This is a deliberate RISC design choice.
Instruction formats by number of addresses
How many operands does an instruction name explicitly? Classic textbooks classify machines as 3-, 2-, 1- and 0-address. The best way to see the difference is to compute the same expression in each.
Task: X = (A + B) * (C - D), where A, B, C, D and X are memory variables.
Three-address
Each instruction names two sources and a destination.
ADD R1, A, B ; R1 = A + B
SUB R2, C, D ; R2 = C - D
MUL X, R1, R2 ; X = R1 * R2
Short programs, but each instruction is long because it must hold three addresses.
Two-address
One operand is both a source and the destination.
MOV R1, A ; R1 = A
ADD R1, B ; R1 = R1 + B
MOV R2, C ; R2 = C
SUB R2, D ; R2 = R2 - D
MUL R1, R2 ; R1 = R1 * R2
MOV X, R1 ; X = R1
This is the x86 style: add eax, ebx means eax = eax + ebx. Shorter instructions, but the original value of the first operand is destroyed, so extra moves are needed.
One-address (accumulator)
An implicit register, the accumulator (AC), is always one source and the destination. Instructions name only one memory operand.
LOAD A ; AC = A
ADD B ; AC = AC + B
STORE T ; T = AC (temporary in memory)
LOAD C ; AC = C
SUB D ; AC = AC - D
MUL T ; AC = AC * T
STORE X ; X = AC
Early computers and simple microcontrollers used this style. Instructions are short, but there is lots of memory traffic for temporaries.
Zero-address (stack)
Operands are implicit: the top of a stack. PUSH and POP name memory; arithmetic pops its inputs and pushes the result. Binary operators compute second-from-top OP top.
PUSH A ; stack: A
PUSH B ; stack: A B
ADD ; stack: (A+B)
PUSH C ; stack: (A+B) C
PUSH D ; stack: (A+B) C D
SUB ; stack: (A+B) (C-D)
MUL ; stack: (A+B)*(C-D)
POP X ; X = result, stack empty
This is exactly the postfix (reverse Polish) form A B + C D - *. The Java Virtual Machine bytecode and the WebAssembly instruction set are stack-based, because stack code is compact and easy to generate; the JIT compiler later maps it to registers.
Load-store (RISC)
Real RISC machines are three-address but only for registers; memory is reached only by loads and stores.
LW R1, A
LW R2, B
ADD R3, R1, R2
LW R4, C
LW R5, D
SUB R6, R4, R5
MUL R7, R3, R6
SW R7, X
More instructions, but each is simple, fixed-length and easy to pipeline.
| Format | Instructions for our task | Instruction length | Memory traffic |
|---|---|---|---|
| Three-address (memory) | 3 | Longest | Operands come straight from memory |
| Two-address | 6 | Medium | Moderate |
| One-address | 7 | Short | High (temporaries in memory) |
| Zero-address (stack) | 8 | Shortest | Stack accesses |
| Load-store RISC | 8 | Fixed 32 bits | Only explicit loads and stores |
Expanding opcodes
When instructions have fixed length, an ISA can give instructions with fewer address fields longer opcodes, using the bits the missing addresses would have occupied. This is called an expanding opcode.
Worked example. A machine has 16-bit instructions and 4-bit address fields. It needs 14 three-address instructions. How many two-address and one-address instructions are possible?
- A three-address instruction is
4-bit opcode + 3 × 4-bit addresses= 16 bits. The opcode field has2^4= 16 patterns. - Use 14 of them for three-address instructions. 2 patterns remain.
- Each remaining pattern can be extended by the first address field (4 bits), since a two-address instruction does not need it: 2 × 16 = 32 two-address instructions possible.
- If only 30 are used, 2 patterns remain, each extendable by another 4 bits: 2 × 16 = 32 one-address instructions possible.
Common mistake
In expanding-opcode questions, first check that the formats fit in the instruction length, then count unused patterns at each level and multiply by 2^(freed bits). Do not simply add bit counts.
Addressing modes
An addressing mode is a rule for finding an operand. The effective address (EA) is the actual memory address the operand is read from or written to, after the mode's calculation.
The setup for our worked examples
Every mode below uses this same machine state, so you can compare them directly:
Instruction: LOAD R0, <mode> with address field = 500
Instruction is at address 200 and is 2 words long,
so while it executes, PC = 202.
Registers: R1 = 400 XR (index register) = 100
Memory: M[399] = 450
M[400] = 700
M[500] = 800
M[600] = 900
M[702] = 325
M[800] = 300
1. Immediate
The operand is the constant inside the instruction. No memory access for data.
- Operand = 500. No EA.
- Example:
addi x5, x5, 1. Used for constants.
2. Register
The operand is in a register named by the instruction.
- Operand = R1 = 400. No memory access.
- Fastest mode; RISC arithmetic uses only register and immediate operands.
3. Direct (absolute)
The address field is the effective address.
- EA = 500, operand = M[500] = 800.
- One memory access. Reach limited by the address-field width.
4. Indirect (memory indirect)
The address field points to a memory word that holds the effective address.
- EA = M[500] = 800, operand = M[800] = 300.
- Two memory accesses. Acts like a pointer stored in memory.
5. Register indirect
A register holds the effective address.
- EA = R1 = 400, operand = M[400] = 700.
- One memory access and a short instruction. This is how a pointer in a register (
*pin C) is dereferenced.
6. Displacement: indexed and base-plus-offset
EA = address field + register contents.
- Indexed: EA = 500 + XR = 500 + 100 = 600, operand = M[600] = 900.
- The address field is the base of an array and the register is the index. In base-register addressing it is the other way round (register = base, field = offset), as in RISC-V
lw x5, 8(x2), which accesses a struct field or a local variable at a fixed offset from a pointer or the stack pointer. The arithmetic is the same addition.
7. PC-relative
EA = PC + address field (the field is a signed offset).
- EA = 202 + 500 = 702, operand = M[702] = 325.
- Used for branches and for position-independent code (code that works wherever it is loaded, as shared libraries must). Note that PC already points past the current instruction; that detail is a frequent trap in exam questions.
8. Auto-increment and auto-decrement
Register indirect, plus an automatic update of the register.
- Auto-increment (post-increment): EA = R1 = 400, operand = M[400] = 700, then R1 becomes 401.
- Auto-decrement (pre-decrement): R1 first becomes 399, then EA = 399, operand = M[399] = 450.
- Ideal for walking through arrays and for pushing and popping a stack. C's
*p++maps onto auto-increment directly. ARM provides pre- and post-indexed modes; RISC-V omits them to keep instructions simple.
Summary table
| Mode | EA calculation | EA | Operand loaded into R0 |
|---|---|---|---|
| Immediate | none | none | 500 |
| Register | none | none | 400 |
| Direct | 500 | 500 | 800 |
| Indirect | M[500] | 800 | 300 |
| Register indirect | R1 | 400 | 700 |
| Indexed | 500 + XR | 600 | 900 |
| PC-relative | PC + 500 | 702 | 325 |
| Auto-increment | R1, then R1 = R1 + 1 | 400 | 700 |
| Auto-decrement | R1 = R1 − 1, then R1 | 399 | 450 |
Interview tip
Tie each mode to a programming construct: immediate for constants, register for local variables in registers, register indirect for pointer dereference, displacement for struct fields, array elements and stack locals, PC-relative for branches and position-independent code, auto-increment for array traversal and stacks, and indirect for pointer-to-pointer access.
Why support many modes, or few?
More addressing modes reduce instruction count (one instruction can do address arithmetic and a load) but complicate decoding and pipelining, and a mode such as memory indirect needs several memory accesses in one instruction. RISC ISAs keep just immediate, register, base-plus-offset and PC-relative, and let the compiler build anything else from simple instructions.
Types of instructions
Every ISA has instructions in these categories:
| Category | Purpose | RISC-V examples | x86 examples |
|---|---|---|---|
| Data transfer | Move data between registers and memory | lw, sw, lb, ld, sd | mov, push, pop |
| Arithmetic | Integer math | add, sub, addi, mul, div | add, sub, imul, idiv |
| Logical | Bitwise operations | and, or, xor, andi | and, or, xor, not |
| Shift and rotate | Move bits | sll, srl, sra | shl, shr, sar, rol |
| Comparison | Produce a flag or a 0/1 value | slt, sltu | cmp, test |
| Control transfer | Change the flow | beq, bne, blt, jal, jalr | jmp, je, call, ret |
| System | Privileged and OS interaction | ecall, ebreak, CSR instructions | syscall, int, hlt |
| Floating point | FP math | fadd.s, fmul.d | SSE/AVX: addss, mulsd |
Two styles of conditional branch are worth knowing:
- Condition codes (flags): x86 and classic ARM compare first (
cmpsets flags), then branch on the flags (je,jl). The flags are implicit state shared between instructions. - Compare-and-branch: RISC-V and MIPS compare registers inside the branch instruction itself (
beq x5, x6, label). No hidden flag state, which simplifies pipelining and out-of-order execution.
x86, ARM and RISC-V compared
| Aspect | x86-64 | ARMv8-A (AArch64) | RISC-V (RV64) |
|---|---|---|---|
| Style | CISC | RISC | RISC |
| Ownership | Intel and AMD | Arm Ltd. licenses it | Open standard, free to implement |
| Instruction length | Variable, 1 to 15 bytes | Fixed 32-bit | 32-bit base; optional 16-bit compressed |
| Integer registers | 16 general-purpose | 31 general-purpose plus a zero register/stack pointer encoding | 32 (x0 is hardwired to 0) |
| Memory operands in ALU instructions | Yes | No (load-store) | No (load-store) |
| Condition flags | Yes | Yes (NZCV) | No; compare-and-branch |
| Byte order | Little-endian | Bi-endian, little-endian in practice | Little-endian |
| Design approach | Backward compatible back to the 1970s 8086 | Clean 64-bit redesign | Small base ISA plus optional extensions (M, A, F, D, C, V) |
| Typical use | Desktops, laptops, servers | Phones, tablets, embedded, growing in laptops and servers | Embedded and microcontrollers, research, growing elsewhere |
RISC-V is popular in teaching (and in this lesson) because it is open, small and regular: the base integer ISA has around 40 instructions.
From C to assembly: a RISC-V example
The RISC-V calling convention (its ABI, application binary interface) names registers by role:
| Register | ABI name | Role | Preserved across calls? |
|---|---|---|---|
| x0 | zero | Always 0 | n/a |
| x1 | ra | Return address | No (caller saves if needed) |
| x2 | sp | Stack pointer | Yes |
| x5 to x7, x28 to x31 | t0 to t6 | Temporaries | No (caller-saved) |
| x8 to x9, x18 to x27 | s0 to s11 | Saved registers (s0 doubles as frame pointer fp) | Yes (callee-saved) |
| x10 to x11 | a0, a1 | Arguments and return values | No |
| x12 to x17 | a2 to a7 | Arguments | No |
Translating a loop
Here is a C function that sums an array:
int sum(int *a, int n) {
int s = 0;
for (int i = 0; i < n; i++)
s += a[i];
return s;
}
A straightforward hand translation into RV32I assembly (the pointer arrives in a0, n in a1, and the result returns in a0):
sum:
li t0, 0 # s = 0
li t1, 0 # i = 0
loop:
bge t1, a1, done # if i >= n, exit loop
slli t2, t1, 2 # t2 = i * 4 (int is 4 bytes)
add t2, a0, t2 # t2 = &a[i] = a + i*4
lw t3, 0(t2) # t3 = a[i]
add t0, t0, t3 # s += a[i]
addi t1, t1, 1 # i++
j loop
done:
mv a0, t0 # return value in a0
ret # jump back to the address in ra
Points to notice:
- Array indexing is address arithmetic.
a[i]becomesa + i × 4, computed with a shift (slliby 2 multiplies by 4) and an add, then loaded with base-plus-offset addressing. li,mv,j,bgewith a label andretare pseudo-instructions the assembler expands:li t0, 0becomesaddi t0, zero, 0;mv a0, t0becomesaddi a0, t0, 0;retbecomesjalr zero, 0(ra).- Because
sumcalls no other function and uses only temporaries, it needs no stack frame at all. Such functions are called leaf functions. - An optimizing compiler would do better: increment a pointer instead of recomputing
a + i*4, and test the loop condition at the bottom. Each iteration then needs about 5 instructions instead of 7.
The call stack and stack frames
Why a stack?
Functions can call other functions, including themselves. Each active call needs its own private storage for its return address, local variables and saved registers. Calls finish in the reverse order they started (last in, first out), so a stack is the natural structure.
The call stack is a region of memory that, on almost all modern ISAs, grows downward from high addresses toward low ones. The stack pointer (SP) register holds the address of the current top. Pushing means decrementing SP and storing; popping means loading and incrementing SP.
The stack frame
Each call's slice of the stack is its stack frame (or activation record). A typical frame holds:
- the return address (where to continue in the caller),
- saved registers the function must restore before returning (callee-saved registers it uses, and the caller's frame pointer),
- local variables that do not fit in registers, and local arrays,
- arguments beyond those passed in registers, and space to save arguments across further calls.
high addresses
+--------------------------+
| caller's frame |
| ... |
+--------------------------+ <- caller's sp (= callee's fp)
| return address (ra) |
| saved fp (s0) |
| saved s1..s11 if used |
| local variables |
| outgoing args (if any) |
+--------------------------+ <- sp (current top of stack)
| free space |
v (stack grows down) v
low addresses
A frame pointer (FP) optionally points to a fixed spot in the current frame, so locals have constant offsets even while SP moves. Optimizing compilers often omit it to free a register; debuggers and profilers then rely on extra metadata to walk the stack.
Calling conventions
A calling convention is the agreement between caller and callee on:
- Where arguments go: RISC-V passes the first eight integer arguments in
a0toa7; x86-64 on Linux (System V ABI) usesrdi,rsi,rdx,rcx,r8,r9; Windows x64 usesrcx,rdx,r8,r9. Older 32-bit x86 conventions passed arguments mostly on the stack. - Where the return value goes:
a0(anda1) on RISC-V,raxon x86-64. - Who saves which registers:
- Caller-saved (volatile) registers, such as
t0tot6anda0toa7: the callee may overwrite them freely. If the caller needs a value in one after the call, the caller must save it first. - Callee-saved (non-volatile) registers, such as
s0tos11andsp: the callee must restore them before returning. If it wants to use one, it saves the old value in its frame and restores it at the end.
- Caller-saved (volatile) registers, such as
- Stack alignment: RISC-V and x86-64 both require the stack pointer to be 16-byte aligned at calls.
Without an agreed convention, code compiled by different compilers, or library code and your code, could not call each other.
How a function call works at machine level
Step by step, for result = f(x) on RISC-V:
- Caller prepares arguments: puts
xina0. Saves any caller-saved registers whose values it needs later. - Caller executes
jal ra, f(jump and link): stores the address of the next instruction (PC + 4) inraand jumps tof. On x86,call fpushes the return address onto the stack instead. - Callee prologue: decrements
spto allocate its frame, savesra(if it will call other functions) and any callee-saved registers it will use. - Callee body: does its work, using
a0and others. - Callee epilogue: places the result in
a0, restores saved registers andra, incrementsspto free the frame. - Callee executes
ret(jalr zero, 0(ra)): jumps back to the address inra. - Caller continues: reads the result from
a0and restores any caller-saved registers it saved.
Worked example: recursive factorial
int fact(int n) {
if (n <= 1) return 1;
return n * fact(n - 1);
}
RV32IM assembly (the M extension provides mul):
fact:
addi sp, sp, -16 # prologue: allocate a 16-byte frame
sw ra, 12(sp) # save return address
sw a0, 8(sp) # save n (we need it after the call)
li t0, 1
bgt a0, t0, recurse # if n > 1, recurse
li a0, 1 # base case: return 1
addi sp, sp, 16 # free the frame
ret
recurse:
addi a0, a0, -1 # argument: n - 1
jal ra, fact # a0 = fact(n - 1); overwrites ra
lw t1, 8(sp) # reload n
mul a0, t1, a0 # a0 = n * fact(n - 1)
lw ra, 12(sp) # restore our return address
addi sp, sp, 16 # epilogue: free the frame
ret
Why must ra be saved? The inner jal overwrites ra with a return address inside fact itself. Without saving the original, the outer call could never return to its caller. Why save n? Because a0 is caller-saved and gets overwritten by the recursive call's result.
Tracing fact(3) called from main, with each frame 16 bytes:
Call sp after prologue Frame holds
main->fact(3) S - 16 ra = (in main), n = 3
fact(3)->fact(2) S - 32 ra = (in fact), n = 2
fact(2)->fact(1) S - 48 ra = (in fact), n = 1 -> returns 1
Unwinding:
fact(1) returns 1
fact(2): reloads n = 2, returns 2 * 1 = 2
fact(3): reloads n = 3, returns 3 * 2 = 6, ret to main
At the deepest point three frames (48 bytes) are on the stack. In general, recursion depth n needs about n × frame size bytes: fact(100000) would need about 1.6 MB. The default main-thread stack is commonly 8 MB on Linux and 1 MB on Windows (both configurable), so that call would overflow on a default Windows stack but not on Linux. Exceeding the limit gives a stack overflow.
Stack frames and security
Local arrays live in the frame below the saved return address. If a program copies more bytes into a local buffer than it holds (for example with C's unsafe gets or an unchecked strcpy), it can overwrite the return address, and ret then jumps wherever the attacker chose. This is the classic stack buffer overflow. Defenses include stack canaries (a check value placed before the return address), non-executable stacks and address-space layout randomization.
Tail calls
If a function's last action is to call another function and return its result unchanged (return g(x);), the compiler can reuse the current frame and jump instead of call. This tail-call optimization makes some recursion run in constant stack space. C compilers do it at higher optimization levels; the JVM does not guarantee it.
Interview questions
Q1. What is an ISA and what does it include?
It is the programmer-visible specification of a processor: instructions, registers, data types, instruction encodings, addressing modes, memory model and exception behavior. It is a contract that lets the same binary run on any implementation. It does not specify pipelines, caches or clock speed.
Q2. Compare zero-, one-, two- and three-address instruction formats.
Three-address instructions name two sources and a destination, giving few but long instructions. Two-address instructions reuse one source as the destination (x86 style). One-address machines use an implicit accumulator, and zero-address machines use an operand stack (JVM bytecode). Fewer addresses give shorter instructions but more of them and more data movement.
Q3. What is the effective address, and compute it for indexed and PC-relative modes.
The effective address is the memory location actually accessed after applying the addressing mode. For indexed mode, EA = address field + index register, so with field 500 and index 100 it is 600. For PC-relative, EA = PC + offset, where PC usually already points to the next instruction, so with PC = 202 and offset 500 it is 702.
Q4. Difference between direct, indirect and register indirect addressing?
Direct: the instruction holds the operand's address (one memory access). Indirect: the instruction holds the address of a memory word containing the operand's address (two accesses). Register indirect: a register holds the operand's address (one access, short instruction), which is how pointer dereferences are implemented.
Q5. Why do RISC ISAs support few addressing modes?
Each extra mode complicates decoding, can need multiple memory accesses per instruction, and makes pipelining harder. Compilers can build complex addressing from a few simple instructions, and the simple instructions run at near one per cycle. RISC-V keeps only immediate, register, base-plus-offset and PC-relative.
Q6. What is a load-store architecture?
An ISA in which only load and store instructions access memory; arithmetic and logic operate only on registers and immediates. It simplifies instruction formats and pipelines, since each instruction does at most one memory access in a fixed stage. ARM, RISC-V and MIPS are load-store; x86 is not.
Q7. What happens at the machine level when a function is called?
The caller places arguments in argument registers (or on the stack), saves any caller-saved registers it needs, and executes a call instruction that records the return address and jumps. The callee's prologue allocates a stack frame and saves the return address and callee-saved registers it uses. The epilogue puts the result in the return register, restores those registers, frees the frame and returns to the saved address.
Q8. What is the difference between caller-saved and callee-saved registers?
Caller-saved registers may be overwritten by any call, so the caller saves them if it needs their values afterward. Callee-saved registers must hold the same values after the call returns, so a callee that uses one saves and restores it. Splitting registers this way reduces total save/restore work compared with saving everything on every call.
Q9. What is stored in a stack frame?
The return address (when the function makes further calls), the saved caller frame pointer, callee-saved registers the function uses, local variables and arrays that do not live in registers, and arguments or spill space for calls it makes. The frame is created in the prologue and released in the epilogue.
Q10. Why does deep recursion cause a stack overflow?
Each active call has its own frame on the call stack, and the stack has a fixed maximum size set by the OS or thread creation. Very deep recursion allocates frames until the stack crosses its limit and hits a guard page, which the OS reports as a segmentation fault or the runtime as StackOverflowError. Iteration, tail-call optimization or a larger stack size avoid it.
Q11. What is a stack buffer overflow attack?
A local buffer sits in the stack frame below the saved return address. Writing past its end can overwrite the return address, so the function's ret jumps to attacker-chosen code. Stack canaries, non-executable stacks, ASLR and bounds-checked functions mitigate it.
Q12. Compare x86 and ARM/RISC-V instruction encoding.
x86 instructions vary from 1 to 15 bytes with prefixes and complex modes, making decoding hard and requiring translation to micro-operations. ARM and RISC-V use fixed 32-bit instructions with register fields in fixed positions (RISC-V adds optional 16-bit compressed forms), which makes decoding simple and parallel. Variable length gives better code density; fixed length gives simpler, more power-efficient front ends.
Q13. What is an expanding opcode?
In fixed-length instruction sets, instructions that need fewer operand fields can use the freed bits to extend the opcode. Some opcode patterns are reserved at each level as "escape" patterns that mean "look further". This allows many more instructions without lengthening the instruction word.
Q14. What is PC-relative addressing used for?
Branches and jumps use PC-relative offsets, so targets need only a small signed field. It is also the basis of position-independent code: shared libraries reference their own code and data relative to the PC, so they work at any load address. That also enables address-space layout randomization.
Key takeaways
- The ISA is the hardware/software contract: instructions, registers, formats, addressing modes and memory model.
- Fewer explicit addresses give shorter instructions but more instructions; stack (0-address) code is postfix, and RISC is 3-address over registers with separate loads and stores.
- Know every addressing mode's EA formula: direct, indirect (two accesses), register indirect, indexed/displacement, PC-relative (PC already incremented) and auto-increment/decrement.
- Instruction types: data transfer, arithmetic, logical, shift, compare, control transfer, system and floating point; branches use either flags or compare-and-branch.
- x86 is variable-length CISC with flags; ARM and RISC-V are fixed-length load-store RISC; RISC-V is open and modular.
- Array indexing compiles to address arithmetic (
base + i × size) plus a base-plus-offset load. - A call saves the return address, the callee builds a stack frame in its prologue, saves callee-saved registers it uses, and tears the frame down in its epilogue.
- Calling conventions fix argument registers, return registers, caller- versus callee-saved registers and stack alignment.
- Stack frames explain recursion limits, stack overflow and stack buffer-overflow attacks.
Next lesson
Continue with CPU datapath and control.

