Why virtualization and security belong together
Two questions sit at the heart of this lesson. How can one physical machine safely run many operating systems or applications at once? And how does an operating system stop one program, or one user, from doing things it should not?
Both are about isolation. Virtualization isolates whole operating systems from each other using a hypervisor; containers isolate groups of processes using kernel features; and the OS itself isolates user programs from the kernel and from each other with hardware protection rings, virtual memory and access control. Attacks on systems are, almost always, attempts to break one of these isolation boundaries.
Interviewers commonly ask you to explain type 1 versus type 2 hypervisors, compare VMs with containers, describe user mode and kernel mode, decode a permission string like rwxr-x--- or a chmod 754, and explain how a buffer overflow works and which defences (ASLR, NX, stack canaries) stop it. This lesson covers each with worked examples.
What virtualization is
Virtualization means presenting software with a virtual version of some resource that behaves like the real thing. Virtual memory (earlier lessons) virtualises RAM for each process. Hardware virtualization, the subject here, virtualises an entire computer, CPU, memory, disks and network cards, so that a complete operating system can run on it unmodified.
- The software that creates and runs virtual machines is the hypervisor, also called the virtual machine monitor (VMM).
- The physical machine is the host.
- Each virtual machine (VM) runs a guest operating system.
A hypervisor must provide three properties, first stated formally by Popek and Goldberg in 1974:
- Fidelity (equivalence): software in a VM behaves the same as on real hardware, apart from timing.
- Safety (resource control): the hypervisor stays in full control of the real hardware; a guest cannot escape or take resources it was not given.
- Performance (efficiency): most guest instructions run directly on the CPU at native speed, with the hypervisor stepping in only for sensitive operations.
Why bother? Server consolidation (many lightly used servers become VMs on one machine), isolation between customers (every cloud VM you rent), running different operating systems side by side, snapshots and live migration of running machines between hosts, and safe testing environments.
Type 1 versus type 2 hypervisors
Type 1 (bare metal) Type 2 (hosted)
+--------+--------+--------+ +--------+--------+
| Guest | Guest | Guest | | Guest | Guest | other apps
| OS | OS | OS | | OS | OS | +--------+
+--------+--------+--------+ +--------+--------+ | browser|
| Hypervisor | | Hypervisor | +--------+
+--------------------------+ | (an application) |
| Hardware | +----------------------------+
+--------------------------+ | Host OS |
+----------------------------+
| Hardware |
+----------------------------+
- A type 1 (bare-metal or native) hypervisor runs directly on the hardware. It is effectively a small specialised operating system whose job is to run VMs. Examples: VMware ESXi, Microsoft Hyper-V, Xen. Used in data centres and clouds because it has less overhead and a smaller attack surface.
- A type 2 (hosted) hypervisor runs as an application on top of a normal host operating system, which handles the hardware. Examples: Oracle VirtualBox, VMware Workstation, Parallels Desktop. Used on laptops and desktops for development and testing; easy to install, but every I/O passes through the host OS.
KVM (Kernel-based Virtual Machine) blurs the line: it is a Linux kernel module that turns Linux itself into a hypervisor, with each VM running as a Linux process (usually managed by QEMU for device emulation). It is generally classed as type 1 because the hypervisor is part of the kernel running on bare metal. KVM underpins much of the public cloud, and AWS's Nitro hypervisor is based on KVM. Hyper-V is also a type 1 hypervisor even though you enable it from Windows: once enabled, Windows itself runs as a privileged partition on top of it.
| Aspect | Type 1 | Type 2 |
|---|---|---|
| Runs on | Bare hardware | A host OS |
| Performance | Near native | Lower; host OS adds a layer |
| Attack surface | Small | Includes the whole host OS |
| Typical use | Servers, cloud | Desktops, development |
| Examples | ESXi, Hyper-V, Xen, KVM | VirtualBox, VMware Workstation, Parallels |
How a hypervisor runs a guest kernel
The core difficulty: a guest kernel expects to run in the CPU's most privileged mode and execute privileged instructions (change page tables, disable interrupts, talk to devices). The hypervisor cannot let it do that on the real hardware, or the guest would control the whole machine. So the guest kernel is run with reduced privilege, and its privileged operations must be intercepted.
Trap and emulate
The classic approach: run the guest kernel in user mode. Whenever it executes a privileged instruction, the CPU traps (raises a fault) into the hypervisor, which emulates the effect on the VM's virtual state, then resumes the guest. Ordinary instructions (arithmetic, memory accesses through its own pages) run directly at full speed.
This works only if every sensitive instruction (one that reads or changes machine state) is also privileged (traps in user mode). The original x86 architecture broke this rule: about 17 instructions, such as POPF (which silently ignores changes to the interrupt flag in user mode instead of trapping), were sensitive but did not trap. So pure trap-and-emulate was impossible on x86, and three techniques arose.
Full virtualization with binary translation
Full virtualization runs unmodified guest operating systems. Before 2005, VMware achieved it on x86 with binary translation: the hypervisor scans guest kernel code just before it runs and rewrites problem instructions into safe sequences that call the hypervisor, caching the translated code. User-mode guest code, which has no sensitive instructions, runs directly. It worked well but is complex.
Paravirtualization
Paravirtualization modifies the guest OS so that it knows it is virtualised and cooperates. Instead of executing privileged instructions, the guest kernel calls the hypervisor directly with hypercalls (like system calls, but from guest kernel to hypervisor). Xen popularised this approach.
- Advantages: avoids expensive traps and translation; can be very efficient, especially for I/O.
- Disadvantage: the guest kernel must be modified, so closed-source operating systems could not originally be paravirtualised.
Today the idea survives mainly in paravirtual device drivers: rather than emulating a real network card register by register, the guest uses a driver designed for virtualization, such as virtio (on KVM) or the Xen PV drivers, which passes batches of requests through shared memory rings. Even fully virtualised guests use them for fast I/O.
Hardware-assisted virtualization (Intel VT-x, AMD-V)
From around 2005 to 2006, Intel (VT-x) and AMD (AMD-V, also called SVM) added CPU support that makes x86 virtualizable directly:
- A new root mode for the hypervisor and non-root mode for guests. The guest kernel runs in non-root mode at its normal ring 0, so it sees what it expects.
- Sensitive operations in non-root mode cause a VM exit to the hypervisor, configured per VM in a control structure (Intel's VMCS, AMD's VMCB). The hypervisor handles it and performs a VM entry back to the guest.
- Second-level address translation: Intel EPT (extended page tables) and AMD NPT/RVI (nested page tables). The guest manages its own page tables (guest virtual to guest physical); the hardware walks a second set maintained by the hypervisor (guest physical to host physical). Before this, hypervisors had to maintain shadow page tables in software, which was slow.
guest virtual addr --(guest page tables)--> guest physical addr
--(EPT/NPT, hypervisor)--> host physical addr
Additional hardware features: IOMMU support (Intel VT-d, AMD-Vi) lets a VM use a physical device directly (device passthrough) safely, and SR-IOV lets one network card present itself as many virtual cards, one per VM.
| Technique | Guest modified? | How sensitive operations are handled | Example |
|---|---|---|---|
| Trap and emulate | No | Privileged instructions trap to hypervisor | Classic mainframes (IBM VM/370) |
| Binary translation | No | Hypervisor rewrites guest kernel code on the fly | Early VMware on x86 |
| Paravirtualization | Yes | Guest calls hypervisor via hypercalls | Xen PV guests; virtio drivers |
| Hardware assisted | No | CPU exits to hypervisor (VM exit); EPT/NPT for memory | KVM, Hyper-V, modern ESXi |
Interview tip
If asked "full versus paravirtualization", give the one-line definitions (full: unmodified guest, hypervisor intercepts sensitive instructions; para: guest is modified to call the hypervisor directly) and then add that modern systems mix them: hardware-assisted full virtualization for the CPU and memory, paravirtual drivers like virtio for disks and networks.
Virtual machines versus containers
The Linux internals lesson showed that a container is a group of host processes isolated with namespaces and limited with cgroups. A VM runs a whole guest kernel on virtual hardware.
Virtual machines Containers
+-------+ +-------+ +-------+ +-------+ +-------+ +-------+
| App A | | App B | | App C | | App A | | App B | | App C |
| libs | | libs | | libs | | libs | | libs | | libs |
|GuestOS| |GuestOS| |GuestOS| +-------+ +-------+ +-------+
+-------+ +-------+ +-------+ | Container runtime |
| Hypervisor | | Host OS kernel (shared) |
| Hardware | | Hardware |
+---------------------------+ +---------------------------+
| Aspect | Virtual machine | Container |
|---|---|---|
| What is isolated | A full machine with its own kernel | Processes sharing the host kernel |
| Isolation mechanism | Hypervisor and hardware virtualization | Namespaces, cgroups, seccomp, capabilities, LSMs |
| Guest OS | Any OS (Linux VM on Windows host, etc.) | Same kernel as host (Linux containers need a Linux kernel) |
| Image size | Gigabytes (full OS) | Megabytes (app plus libraries) |
| Start-up time | Seconds to minutes (boot a kernel) | Milliseconds to seconds (start a process) |
| Density per host | Tens | Hundreds or more |
| Isolation strength | Strong; escaping needs a hypervisor bug | Weaker; a kernel bug can break out of every container |
| Typical use | Multi-tenant cloud, different OSes, strong boundaries | Packaging and deploying microservices, CI jobs |
Docker on a Mac or Windows machine actually runs a small Linux VM and runs the containers inside it, because containers need a Linux kernel.
The two are often combined: cloud providers run your containers inside VMs, so the hypervisor separates customers while containers separate services. Sandboxed container runtimes narrow the gap further: gVisor intercepts container system calls in a user-space kernel, and Kata Containers and AWS Firecracker run each container or function inside a lightweight microVM that boots in a fraction of a second.
Protection rings and user/kernel isolation
CPUs enforce privilege levels so that ordinary programs cannot execute dangerous instructions. x86 defines four rings, 0 (most privileged) to 3 (least).
+-----------------------------+
| Ring 3: user applications |
| +---------------------+ |
| | Rings 1, 2: unused | |
| | +-------------+ | |
| | | Ring 0: | | |
| | | kernel | | |
| | +-------------+ | |
| +---------------------+ |
+-----------------------------+
(hypervisor in VT-x root mode, sometimes called "ring -1")
Mainstream operating systems use just two: ring 0 for the kernel (kernel mode, supervisor mode) and ring 3 for applications (user mode). ARM has a similar idea with exception levels: EL0 for applications, EL1 for the kernel, EL2 for a hypervisor and EL3 for secure firmware.
What user mode cannot do:
- Execute privileged instructions: disable interrupts, change page-table registers (
CR3), access I/O ports, halt the CPU. Trying raises a fault, and the kernel typically kills the process. - Access kernel memory: kernel pages are marked supervisor-only in the page tables (the user/supervisor bit), so a user-mode access faults.
- Access other processes' memory: each process has its own page tables.
Crossing the boundary safely
User programs need kernel services, so there are controlled entry points:
- System calls: the program places a call number and arguments in registers and executes a special instruction (
syscallon x86-64,svcon ARM64). The CPU switches to kernel mode and jumps to a fixed kernel entry point; the program cannot choose where to land. - Exceptions: faults such as division by zero or page faults transfer control to kernel handlers.
- Interrupts: devices and timers.
The kernel must treat every system call argument as untrusted: check that pointers point into user memory (Linux uses copy_from_user and copy_to_user), validate lengths, and check permissions. A missing check here is one of the most common kinds of kernel vulnerability.
Hardware adds more guards: SMEP (supervisor mode execution prevention) stops the kernel from executing code in user pages, and SMAP (supervisor mode access prevention) stops it from accidentally reading or writing user pages except through the explicit copy routines. After the 2018 Meltdown attack, which used speculative execution to read kernel memory from user mode, Linux added KPTI (kernel page-table isolation), keeping most kernel memory unmapped while user code runs.
Access control: DAC, MAC and RBAC
Access control decides whether a subject (a user or process) may perform an operation (read, write, execute, delete) on an object (a file, socket, device). Conceptually this is an access matrix: rows are subjects, columns are objects, cells are allowed rights.
report.txt payroll.db /usr/bin/ls
asha read,write - execute
ravi read - execute
payroll - read,write execute
A real system stores the matrix either by column, as an access control list (ACL) attached to each object ("who may do what to me"), or by row, as capabilities held by each subject ("what I may do", like an unforgeable ticket). File permissions are a compact form of ACL; Unix file descriptors behave like capabilities once a file is open.
Three policy models decide who sets the rules:
- Discretionary access control (DAC): the owner of an object decides who can access it. Unix permissions and Windows NTFS ACLs are DAC. Flexible, but a user (or malware running as that user) can give access away, and a process has all the rights of the user who runs it.
- Mandatory access control (MAC): a system-wide policy set by an administrator decides, and users cannot override it, not even owners. Objects and subjects carry security labels. Linux implements MAC through Linux Security Modules (LSMs): SELinux (labels on every process and file; default on Red Hat family and Android) and AppArmor (path-based profiles; default on Ubuntu and SUSE). A web server confined by SELinux cannot read users' home directories even if file permissions would allow it.
- Role-based access control (RBAC): permissions are assigned to roles (such as "DBA" or "auditor"), and users are assigned roles. Managing roles is much easier than managing per-user rights in large organisations. Common in databases, cloud IAM and Kubernetes (
RoleandRoleBinding). SELinux also includes a role component.
| Model | Who decides | Can owners override? | Examples |
|---|---|---|---|
| DAC | The owner | Yes | Unix permissions, NTFS ACLs |
| MAC | Central policy | No | SELinux, AppArmor, military classification levels |
| RBAC | Administrators, via roles | Depends on policy | Kubernetes RBAC, database roles, cloud IAM |
The guiding principle across all models is least privilege: give every subject only the rights it needs, for only as long as it needs them.
Unix permissions and chmod
Every Unix file has an owner (user), a group, and three sets of three permission bits:
- rwx r-x r--
| | | |
| | | +-- others: everyone else
| | +------- group: members of the file's group
| +------------ user: the owner
+---------------- type: - file, d directory, l symlink,
c char device, b block device, s socket, p pipe
| Bit | On a file | On a directory |
|---|---|---|
| r (read) | Read contents | List the names in it |
| w (write) | Modify contents | Create, delete and rename entries in it |
| x (execute) | Run as a program | Enter it (cd) and access entries by name |
Note the directory rules: deleting a file needs write permission on the directory, not on the file. And r without x on a directory lets you list names but not open anything inside.
The kernel checks one class only: if you are the owner, only the user bits apply; else if you are in the group, only the group bits; else the others bits. So a file with ---rwxrwx cannot be read by its own owner. The superuser (root) bypasses these checks (except that execute requires at least one x bit).
Octal notation
Each set of three bits becomes one octal digit: r = 4, w = 2, x = 1, added together.
| Permission | Calculation | Digit |
|---|---|---|
| rwx | 4 + 2 + 1 | 7 |
| rw- | 4 + 2 + 0 | 6 |
| r-x | 4 + 0 + 1 | 5 |
| r-- | 4 + 0 + 0 | 4 |
| --- | 0 | 0 |
Worked examples
1. What does chmod 754 deploy.sh give?
- 7 = 4 + 2 + 1 =
rwxfor the user (owner). - 5 = 4 + 1 =
r-xfor the group. - 4 =
r--for others. - Result:
-rwxr-xr--. The owner can edit and run it, the group can read and run it, others can only read it.
2. Convert rw-r----- to octal.
- user
rw-= 4 + 2 = 6; groupr--= 4; others---= 0. - Answer: 640. A typical mode for a config file containing secrets readable by a service's group.
3. Symbolic changes. Starting from -rw-r--r-- (644), run chmod u+x,g+w,o-r file:
- u+x: user
rw-becomesrwx(7). - g+w: group
r--becomesrw-(6). - o-r: others
r--becomes---(0). - Result:
-rwxrw----= 760.
4. umask. The umask removes bits from the default mode of newly created files. Programs usually request 666 for files and 777 for directories; the kernel clears the bits set in the umask (mode AND NOT umask).
- umask 022: files 666 minus write for group and others = 644 (
rw-r--r--); directories 755 (rwxr-xr-x). - umask 027: files 666 AND NOT 027 = 640 (
rw-r-----); directories 777 AND NOT 027 = 750 (rwxr-x---).
Note that umask is a bit mask, not subtraction: with umask 027, the file case is 666 AND NOT 027. In binary, 027 clears the group write bit and all of the others' bits, giving 640, while plain subtraction 666 - 027 would give the wrong answer 637 (which would wrongly include execute bits).
$ umask 027
$ touch a.txt; mkdir d
$ ls -ld a.txt d
-rw-r----- 1 asha devs 0 Oct 10 11:20 a.txt
drwxr-x--- 2 asha devs 4096 Oct 10 11:20 d
Special bits: setuid, setgid and sticky
A fourth, leading octal digit holds three special bits:
| Bit | Value | On a file | On a directory | Shown as |
|---|---|---|---|---|
| setuid | 4 | Program runs with the file owner's user ID | (no effect on Linux) | s in user execute position |
| setgid | 2 | Program runs with the file's group ID | New files inherit the directory's group | s in group execute position |
| sticky | 1 | (historical; ignored on Linux) | Only a file's owner (or the directory owner or root) may delete or rename it | t in others execute position |
Examples:
/usr/bin/passwdis-rwsr-xr-x(4755), owned by root. Any user can run it, and it runs as root so it can update/etc/shadow. That is why setuid-root programs are prime targets for privilege escalation: a bug in them gives an attacker root./tmpisdrwxrwxrwt(1777): everyone may create files, but nobody can delete other users' files.chmod 2775 /srv/sharedmakes a shared team directory where new files belong to the team group.
Authentication basics
Authentication proves who you are; authorisation (access control, above) decides what you may do. The OS authenticates users at login using one or more factors:
- something you know: a password or PIN
- something you have: a phone, a hardware security key, a smart card
- something you are: a fingerprint or face (biometrics)
Using two or more different factors is multi-factor authentication (MFA).
How Unix-like systems store passwords:
- The system never stores the password itself. It stores a hash produced by a deliberately slow, salted algorithm such as yescrypt (the default on recent Debian, Ubuntu and Fedora), SHA-512-crypt, or bcrypt.
- A salt is a random value stored with each hash and mixed into it, so two users with the same password get different hashes and precomputed tables (rainbow tables) are useless.
- The hashes live in
/etc/shadow, readable only by root./etc/passwd, which every program reads to map user IDs to names, is world-readable but contains only anxplaceholder for the password. - At login, the system hashes the entered password with the stored salt and compares.
/etc/passwd: asha:x:1000:1000:Asha Rao:/home/asha:/bin/bash
/etc/shadow: asha:$y$j9T$Fq0...salt...$Wm3...hash...:20371:0:99999:7:::
| algorithm id ($y$ = yescrypt)
On Linux, the authentication steps are pluggable through PAM (pluggable authentication modules), so the same login program can use local passwords, LDAP, Kerberos or MFA. For remote access, SSH public-key authentication is preferred over passwords: the private key never leaves your machine, and the server stores only your public key in ~/.ssh/authorized_keys.
After authentication, each process carries the user's credentials: a real user ID, an effective user ID (used for permission checks; changed by setuid programs), group IDs and, on Linux, capabilities.
Linux capabilities
Traditionally, root (UID 0) bypasses all permission checks. Linux splits root's power into about 40 capabilities that can be granted separately: CAP_NET_BIND_SERVICE (bind ports below 1024), CAP_NET_ADMIN (configure networking), CAP_SYS_ADMIN (a very broad grab bag), CAP_SYS_PTRACE (trace other processes), CAP_CHOWN, and others. A web server can be given only CAP_NET_BIND_SERVICE instead of full root. Container runtimes drop most capabilities by default; docker run --privileged gives them all back, which removes most container isolation.
Common OS-level attacks
Buffer overflow
A buffer overflow happens when a program writes more data into a fixed-size buffer than it can hold, overwriting adjacent memory. In C and C++ there are no automatic bounds checks, so functions like strcpy, gets, sprintf and hand-written loops are classic sources.
A stack buffer overflow is the textbook case. A function's local array sits in its stack frame, below the saved frame pointer and the return address (where execution resumes when the function returns). Writing past the array's end overwrites them.
higher addresses
+------------------------+
| caller's frame |
+------------------------+
| return address | <- attacker overwrites this
+------------------------+
| saved frame pointer |
+------------------------+
| (stack canary) | <- defence: checked before return
+------------------------+
| char buf[16] | <- strcpy writes upward from here
+------------------------+
lower addresses (stack grows down)
The classic exploit: send input that fills the buffer, then overwrites the return address with the address of attacker-supplied code (shellcode) placed in the same input. When the function returns, the CPU jumps to the shellcode, which typically starts a shell with the program's privileges. If the program was a setuid-root binary or a network service running as root, the attacker now has root.
#include <stdio.h>
#include <string.h>
static void greet_unsafe(const char *name) {
char buf[16];
strcpy(buf, name); /* no length check: overflow if name >= 16 bytes */
printf("hello, %s\n", buf);
}
static void greet_safe(const char *name) {
char buf[16];
snprintf(buf, sizeof buf, "%s", name); /* truncates instead of overflowing */
printf("hello, %s\n", buf);
}
int main(int argc, char **argv) {
const char *name = argc > 1 ? argv[1] : "world";
greet_safe(name);
greet_unsafe(name);
return 0;
}
With a short name both functions print a greeting. With a 40-character name, greet_safe prints the first 15 characters, while greet_unsafe corrupts its stack frame. With modern compiler defences the process is stopped rather than hijacked: on Linux with GCC's stack protector you typically see *** stack smashing detected ***: terminated and the process aborts; with _FORTIFY_SOURCE enabled, the compiler may replace strcpy with a checked version that aborts even earlier. Either way, the program dies instead of running attacker code, which is exactly the point of the defences below.
Other memory-safety bugs in the same family: heap overflows, use-after-free (using memory after free, which may now hold attacker-controlled data), integer overflows leading to undersized allocations, and format-string bugs (printf(user_input)).
Privilege escalation
Privilege escalation means gaining more rights than you were given.
- Vertical: an ordinary user becomes root (or a process becomes the kernel).
- Horizontal: one user gains access to another user's data at the same level.
Common OS-level routes:
- A memory-safety bug in a setuid-root program or a root daemon (as above).
- A kernel vulnerability, reachable through a system call, a driver or a file system, which gives code execution in ring 0. Real examples: Dirty COW (CVE-2016-5195), a race in copy-on-write handling that let users write to read-only files, and Dirty Pipe (CVE-2022-0847), which allowed overwriting data in read-only files through the pipe page cache.
- Misconfiguration: world-writable scripts run by root's cron jobs, overly broad
sudorules (sudo vimcan spawn a root shell), setuid bits on binaries that can run arbitrary commands, secrets in readable files. - Path and environment tricks: a privileged script that runs
lswithout a full path can be fooled by a user-controlledPATH. - Container escapes: privileged containers, a mounted Docker socket (
/var/run/docker.sock, which effectively grants root on the host), or kernel bugs.
Defences
No single defence stops every attack, so systems layer them (defence in depth). Each one below breaks a specific step of the classic exploit.
Stack canaries
The compiler places a random value, the canary, between local buffers and the saved return address when the function starts, and checks it just before the function returns. A linear overflow that reaches the return address must overwrite the canary first, so the check fails and the program aborts. Enabled with GCC and Clang flags such as -fstack-protector-strong (default on many distributions). The name comes from the canaries miners carried to detect gas.
Limits: an attacker who can leak the canary value, or who overwrites a function pointer without touching the canary, can bypass it.
DEP / NX (non-executable memory)
Data execution prevention (DEP), implemented with the CPU's NX (no-execute) bit (AMD's name; Intel calls it XD), marks pages as either writable or executable, not both (W^X, "write xor execute"). The stack and heap are non-executable, so shellcode injected there cannot run: jumping to it causes a fault.
Attackers responded with return-oriented programming (ROP): instead of injecting new code, they chain together short existing instruction sequences ("gadgets") ending in ret from the program and its libraries. Hence the next defence.
ASLR (address space layout randomization)
ASLR loads the stack, heap, shared libraries, the memory-mapped region and (for position-independent executables, PIE) the program itself at random addresses each run. An exploit needs to know where its target code or gadgets are; with ASLR it must guess or first find an information leak. On Linux, /proc/sys/kernel/randomize_va_space controls it (2 = full randomisation, the default). The kernel randomises its own location too (KASLR).
ASLR is weaker on 32-bit systems, where there are few random bits and brute force is feasible, and it is defeated entirely by a single pointer leak, which is why it is combined with the others.
Other compiler and hardware defences
_FORTIFY_SOURCE: replaces risky libc calls with checked versions when buffer sizes are known at compile time.- RELRO: makes the dynamic linking tables read-only after start-up, so they cannot be overwritten to redirect calls.
- Control-flow integrity (CFI) and hardware support such as Intel CET shadow stacks and indirect-branch tracking, and ARM pointer authentication (PAC) and BTI: make it much harder to redirect returns and indirect jumps.
- Memory-safe languages (Rust, Go, Java, Python) remove whole classes of these bugs by checking bounds and managing memory automatically, which is why new systems code increasingly uses them.
Sandboxing
A sandbox runs code with deliberately restricted rights, so even a fully compromised process can do little damage. Building blocks:
- separate unprivileged user accounts per service
- namespaces and cgroups (containers)
- MAC policies (SELinux, AppArmor)
- seccomp (below)
- dropping capabilities, and
chrootorpivot_rootto a minimal file system
Browsers are the best-known example: Chrome and Firefox run each site's renderer in a heavily sandboxed process, so a bug in the JavaScript engine still cannot read your files.
seccomp
seccomp (secure computing mode) is a Linux feature that restricts which system calls a process may make. Since the system call interface is the main path from a process into the kernel, shrinking it shrinks the attack surface.
- Strict mode (the original): only
read,write,_exitandsigreturnare allowed. - Filter mode (seccomp-bpf): the process installs a small BPF program that inspects each system call number and its arguments and decides to allow it, fail it with an error, kill the process, or notify a supervisor. Once installed, a filter cannot be removed, and it is inherited by children.
Docker applies a default seccomp profile that blocks several dozen rarely needed, risky system calls (for example kexec_load and reboot, and mount unless the container is granted CAP_SYS_ADMIN). systemd services can use SystemCallFilter=. Chrome, Firefox, OpenSSH's pre-authentication process and Android's app zygote all use seccomp filters.
process --- syscall --> [ seccomp BPF filter ] -- allowed --> kernel
|
+-- denied --> EPERM or SIGSYS (killed)
| Defence | Stops | Bypassed by |
|---|---|---|
| Stack canary | Linear stack overflow reaching the return address | Canary leak; overwriting other pointers |
| NX / DEP | Running injected shellcode | Return-oriented programming (ROP) |
| ASLR / KASLR | Jumping to known addresses (gadgets, libraries) | Information leaks; brute force on 32-bit |
| CFI, CET, PAC | Hijacking returns and indirect calls | Still an active research area |
| Sandboxing, seccomp, MAC | Damage after compromise | Kernel bugs in allowed system calls |
| Least privilege | Escalation through over-privileged services | Misconfiguration |
Common mistake
Do not present ASLR, NX and canaries as alternatives. They defend different steps of an exploit and are meant to be used together: canaries detect the overwrite, NX stops injected code, ASLR hides the addresses that code-reuse attacks need. An exploit today typically needs an information leak plus a ROP chain to beat all three.
Interview questions
Q1. What is the difference between a type 1 and a type 2 hypervisor?
A type 1 (bare-metal) hypervisor runs directly on the hardware and hosts VMs itself, as with ESXi, Hyper-V, Xen and KVM; it is used in servers and the cloud for performance and a small attack surface. A type 2 (hosted) hypervisor runs as an application on a normal OS, such as VirtualBox or VMware Workstation, and suits desktops and development.
Q2. What is the difference between full virtualization and paravirtualization?
Full virtualization runs unmodified guest operating systems; the hypervisor intercepts sensitive instructions through traps, binary translation or hardware support. Paravirtualization modifies the guest to call the hypervisor directly through hypercalls, avoiding costly traps. Modern systems combine hardware-assisted full virtualization with paravirtual I/O drivers such as virtio.
Q3. What do Intel VT-x and AMD-V add?
They add a root mode for the hypervisor and a non-root mode for guests, so the guest kernel can run at ring 0 while sensitive operations cause VM exits to the hypervisor. Extended or nested page tables (EPT/NPT) let hardware translate guest physical to host physical addresses, replacing slow shadow page tables.
Q4. Compare containers and virtual machines.
VMs virtualise hardware and each runs its own kernel, giving strong isolation and any guest OS, at the cost of size and boot time. Containers are isolated processes sharing the host kernel through namespaces and cgroups, so they are small and start in milliseconds, but a kernel vulnerability can break isolation. Clouds often run containers inside VMs, or in microVMs like Firecracker, to get both.
Q5. What are user mode and kernel mode, and how does a program switch between them?
Kernel mode can execute privileged instructions and access all memory; user mode cannot, and attempts trap to the kernel. Programs enter the kernel only at fixed entry points through system calls (the syscall instruction), exceptions or interrupts, and the kernel returns to user mode afterwards. This prevents applications from taking over the machine.
Q6. Compare DAC, MAC and RBAC.
In DAC the object's owner decides who has access, as with Unix permissions. In MAC a central policy decides and users cannot override it, as with SELinux and AppArmor. In RBAC permissions are attached to roles and users receive roles, which simplifies administration in large organisations.
Q7. What permissions does chmod 640 set, and why use it?
6 = rw- for the owner, 4 = r-- for the group, 0 = --- for others, so rw-r-----. It suits a secret-bearing config file: the owner can edit, a service's group can read, and nobody else can see it.
Q8. What does the setuid bit do, and why is it dangerous?
A setuid executable runs with the file owner's effective user ID rather than the caller's; passwd uses it to run as root. Any bug in a setuid-root program can let an ordinary user execute code as root, so such programs must be minimal and carefully audited.
Q9. To delete a file in a directory, which permissions do you need?
Write and execute permission on the directory, because deletion modifies the directory's entries; the file's own permissions do not matter. In a sticky directory such as /tmp, you additionally must own the file or the directory (or be root).
Q10. How are passwords stored on Linux?
As salted hashes from a deliberately slow algorithm (yescrypt, SHA-512-crypt or bcrypt) in /etc/shadow, readable only by root. The salt makes identical passwords hash differently and defeats precomputed tables; the slowness makes brute force expensive. Login hashes the entered password with the stored salt and compares.
Q11. Explain a stack buffer overflow attack.
A program copies input into a fixed-size stack buffer without checking its length, so excess bytes overwrite the saved frame pointer and return address. The attacker sets the return address to point at injected code or existing code gadgets, and the function's return jumps there. The program then runs attacker-chosen code with its privileges.
Q12. How do ASLR, NX and stack canaries defend against it?
Stack canaries detect the overwrite before the function returns and abort. NX makes the stack and heap non-executable so injected code cannot run. ASLR randomises where code and libraries are loaded so the attacker cannot predict addresses for return-oriented programming. Together they force an attacker to find an information leak and build a ROP chain.
Q13. What is seccomp?
A Linux mechanism that restricts the system calls a process may make. In filter mode a BPF program checks each call and allows, denies or kills; filters cannot be removed once installed. Docker, browsers and systemd use it to shrink the kernel attack surface available to a compromised process.
Q14. What is privilege escalation? Give OS-level examples.
Gaining more privileges than granted, vertically (user to root) or horizontally (one user to another). Examples: exploiting a bug in a setuid-root binary, a kernel vulnerability such as Dirty COW, overly broad sudo rules, root cron jobs running world-writable scripts, or a container with access to the Docker socket.
Q15. Why does running docker with --privileged weaken security?
It grants all Linux capabilities, disables the default seccomp and LSM confinement, and exposes host devices. A process in such a container can typically mount host disks or load kernel modules, effectively becoming root on the host. Grant only the specific capabilities a workload needs instead.
Key takeaways
- A hypervisor provides fidelity, safety and performance; type 1 runs on bare metal (ESXi, Hyper-V, Xen, KVM), type 2 runs on a host OS (VirtualBox).
- Full virtualization runs unmodified guests, paravirtualization uses hypercalls, and VT-x/AMD-V with EPT/NPT made x86 efficiently virtualizable; virtio carries the paravirtual idea forward for I/O.
- VMs isolate kernels and are heavier; containers share the host kernel and are lighter but weaker boundaries.
- Rings separate kernel mode (ring 0) from user mode (ring 3); system calls are the only controlled way in, and the kernel must validate every argument.
- DAC lets owners decide, MAC enforces central policy (SELinux, AppArmor), RBAC grants rights through roles; always apply least privilege.
- chmod digits are r = 4, w = 2, x = 1 per class; 754 =
rwxr-xr--; umask clears bits (022 gives 644 files and 755 directories); setuid 4, setgid 2, sticky 1. - Buffer overflows overwrite return addresses; canaries, NX, ASLR, CFI and memory-safe languages each block a different step.
- Sandboxing with seccomp, capabilities, namespaces and MAC limits the damage a compromised process can do.
Next lesson
Continue with OS interview questions.

