A thread is a single sequence of instructions being executed inside a process. A process is a container of resources (an address space, open files, credentials); a thread is a path of execution through that container. A process can contain one thread or thousands, all sharing the same memory. Web servers handle many requests at once with threads, browsers keep the UI responsive while loading pages, and databases use threads to run queries in parallel.
Interviewers commonly ask you to compare threads with processes, list exactly what threads share, explain the threading models (one-to-one, many-to-one, many-to-many), write a small multithreaded program, and discuss thread pools, the Python GIL, async code and thread safety. This lesson covers all of these, with working C (pthreads), Java and Python examples. Synchronization (locks, semaphores) builds directly on this and has its own lesson, synchronization.
Why threads exist
Imagine a web server written as one process with one thread. While it waits for a slow database query for client A, it cannot do anything for client B. You could fork a process per client, as early servers did, but processes are expensive to create, each has its own address space, and sharing data between them (a cache, a connection pool) needs IPC.
Threads solve this:
- Responsiveness: one thread can block on I/O while others keep working. A GUI keeps a dedicated thread for user input so the window never freezes.
- Resource sharing: threads share memory by default, so passing data between them is just reading and writing variables (with care).
- Economy: creating a thread and switching between threads of the same process is cheaper than doing the same with processes, because there is no new address space to set up and no page-table switch.
- Scalability: on a multi-core CPU, threads of one process can run truly in parallel on different cores.
Thread versus process
| Aspect | Process | Thread |
|---|---|---|
| Definition | A running program with its own resources | A path of execution inside a process |
| Address space | Own, isolated | Shared with other threads of the process |
| Creation cost | Higher (new address space, PCB, copies) | Lower (new stack, registers, small kernel record) |
| Context switch | Includes page-table switch, more TLB and cache impact | No address-space change within a process |
| Communication | IPC: pipes, sockets, shared memory | Direct reads and writes to shared memory |
| Fault isolation | A crash usually affects only that process | A crash (for example a segfault) kills the whole process |
| Synchronization need | Low by default (nothing shared) | High: shared data needs locks |
| Example | Each Chrome site runs in separate renderer processes | Threads inside one Java server handling requests |
The fault-isolation row is why browsers like Chrome use multiple processes for security and stability, while servers like a Java or Go service use many threads inside one process for efficiency. Many real systems mix both, for example a pool of worker processes, each running several threads.
Interview tip
A good one-line answer: "A process is the unit of resource ownership; a thread is the unit of scheduling. Threads in a process share its memory and files, but each has its own stack, registers and program counter." Then give a trade-off: threads are cheaper and share data easily, but a bug in one can corrupt or crash all of them.
What threads share and what they do not
This is one of the most frequently asked questions in the topic. Get it exactly right.
+------------------------- process ---------------------------+
| SHARED by all threads: |
| code (text) | globals (data, BSS) | heap |
| open files | signal handlers | current dir, PID |
| |
| +--- thread 1 ---+ +--- thread 2 ---+ +--- thread 3 ---+ |
| | thread ID | | thread ID | | thread ID | |
| | registers, PC | | registers, PC | | registers, PC | |
| | stack | | stack | | stack | |
| | signal mask | | signal mask | | signal mask | |
| | errno, TLS | | errno, TLS | | errno, TLS | |
| +----------------+ +----------------+ +----------------+ |
+-------------------------------------------------------------+
Shared (one copy per process):
- Code segment.
- Global and static variables (data and BSS segments).
- The heap: memory from
malloc/newin one thread is visible to all. - Open file descriptors: if one thread closes a file, it is closed for all.
- Signal handlers (dispositions), the current working directory, user and group IDs, and the PID.
Private (one per thread):
- Thread ID.
- Registers, including the program counter and stack pointer.
- Stack: local variables and call frames. Each thread needs its own stack because each is in a different place in its call chain.
- Signal mask: which signals this thread blocks.
- Thread-local storage (TLS): variables declared
_Thread_localin C11,thread_localin C++, orThreadLocalin Java, which give each thread its own copy.errnois thread-local in modern C libraries, so one thread's failed call does not overwrite another's error code. - Scheduling priority and state.
Common mistake
"Each thread has its own heap." No: the heap is shared. Some allocators keep per-thread caches or arenas for speed (glibc's malloc, tcmalloc, jemalloc), but any thread can access any heap object through a pointer. Also note that stacks are private by convention, not by protection: all thread stacks live in the same address space, so passing a pointer to a local variable to another thread works, and becomes a bug if the first function returns.
In the kernel, each thread needs a small control block of its own, a thread control block (TCB), holding its registers, state and stack pointer. Linux makes no deep distinction: every thread is a task_struct, created by the clone system call with flags saying which resources to share with the creator. A "process" is simply a group of tasks sharing an address space and a thread-group ID, which is what getpid() returns.
User-level threads and kernel-level threads
Who manages threads, the kernel or a library in user space?
Kernel-level threads are known to, created by and scheduled by the kernel. Each one can block independently and can run on a different core. Creating them and switching between them needs system calls, so they cost more than user-level threads. Linux (NPTL), Windows and macOS threads are kernel-level.
User-level threads are implemented entirely by a library or language runtime in user space. The kernel sees only one execution context (or a few). The runtime keeps its own thread table and switches by saving and restoring registers itself.
| User-level threads | Kernel-level threads | |
|---|---|---|
| Managed by | Library or runtime | OS kernel |
| Creation and switch cost | Very low (no system call) | Higher (system call, kernel scheduling) |
| Blocking system call | Can block every thread mapped to the same kernel thread | Blocks only the calling thread |
| True parallelism on multiple cores | Only if mapped onto several kernel threads | Yes |
| Scheduling policy | Customizable by the runtime | Fixed by the OS |
| Kernel awareness | None | Full |
Multithreading models
The relationship between user threads and kernel threads defines the threading model.
Many-to-one
Many user threads map to one kernel thread.
user: U1 U2 U3 U4
\ | | /
kernel: K1
- Very fast switching, all in user space.
- One blocking system call blocks every thread, because the kernel only sees one.
- No parallelism: only one kernel thread, so only one core.
Examples: early Java "green threads" on Solaris and GNU Portable Threads. Rarely used today as the only model.
One-to-one
Each user thread maps to its own kernel thread.
user: U1 U2 U3 U4
| | | |
kernel: K1 K2 K3 K4
- True parallelism and independent blocking.
- Each thread costs a kernel thread (kernel memory and a stack, typically with a default size of several megabytes of virtual address space on Linux), so creating hundreds of thousands is impractical.
This is the model of Linux (NPTL), Windows and macOS, and of traditional Java platform threads.
Many-to-many
Many user threads are multiplexed onto a smaller or equal number of kernel threads.
user: U1 U2 U3 U4 U5 U6
\ | \ / | / /
kernel: K1 K2 K3
- Can create huge numbers of cheap user threads while still using all cores.
- When one user thread blocks, the runtime moves the others onto a different kernel thread.
- Complex to implement well. Some operating systems tried it at kernel level (older Solaris versions, FreeBSD's KSE) and later went back to one-to-one.
The model is very much alive in language runtimes: Go schedules goroutines onto a small set of OS threads (its "M:N" scheduler), Erlang schedules lightweight processes this way, and Java 21 virtual threads mount millions of lightweight threads on a small pool of carrier OS threads. A two-level model is a variant of many-to-many that also allows binding a specific user thread to its own kernel thread.
Interview tip
If asked "which model does Linux use?", say one-to-one, via NPTL, with every thread being a kernel task created with clone. Then add that languages build many-to-many on top: Go goroutines and Java virtual threads. That shows you know both the textbook and the present.
pthreads in C
POSIX threads (pthreads) is the standard C threading API on Unix-like systems. The core calls:
pthread_create(&tid, attr, func, arg)starts a thread runningfunc(arg).pthread_join(tid, &result)waits for a thread to finish and collects its return value, likewaitfor processes.pthread_exit(value)ends the calling thread.pthread_detach(tid)says nobody will join this thread, so its resources are freed automatically when it ends.
This program sums the numbers from 1 to 1,000,000 using four threads, each summing a quarter of an array:
#include <pthread.h>
#include <stdio.h>
#define N 1000000
#define THREADS 4
static long data[N];
struct range {
int start, end; /* half-open interval [start, end) */
long sum; /* each thread writes only its own result */
};
static void *partial_sum(void *arg) {
struct range *r = arg;
long s = 0;
for (int i = r->start; i < r->end; i++)
s += data[i];
r->sum = s;
return NULL;
}
int main(void) {
for (int i = 0; i < N; i++)
data[i] = i + 1;
pthread_t tid[THREADS];
struct range parts[THREADS];
int chunk = N / THREADS;
for (int t = 0; t < THREADS; t++) {
parts[t].start = t * chunk;
parts[t].end = (t == THREADS - 1) ? N : (t + 1) * chunk;
pthread_create(&tid[t], NULL, partial_sum, &parts[t]);
}
long total = 0;
for (int t = 0; t < THREADS; t++) {
pthread_join(tid[t], NULL);
total += parts[t].sum;
}
printf("total = %ld\n", total);
return 0;
}
Compile with gcc -O2 -pthread sum.c -o sum. Output: total = 500000500000, which matches the formula n(n + 1) / 2 for n = 1,000,000.
Notice the design: the threads read shared data (data) but each writes only to its own struct range, and the main thread combines results after pthread_join. Because no two threads write the same memory, no lock is needed. This "partition the work, combine at the end" pattern is the simplest correct way to use threads.
What happens if you skip the pthread_join? main returns, which calls exit, which ends the whole process, killing all threads mid-calculation. A thread that is neither joined nor detached also leaks its resources, much like a zombie process.
Forking a multithreaded process
If one thread calls fork(), POSIX says the child gets a copy of the address space but only the calling thread. Locks held by other threads at that moment are copied in the locked state, and their owners do not exist in the child, so the child can deadlock if it touches them (a common example is a lock inside malloc). The safe rule: in a multithreaded program, the child should call only async-signal-safe functions until it calls exec.
Threads in Java
In Java, every thread is a java.lang.Thread object. You give it work as a Runnable (or a lambda) and call start(). For anything beyond toy programs, you use an ExecutorService thread pool instead of creating threads by hand.
import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
public class ThreadDemo {
public static void main(String[] args) throws Exception {
// 1. A plain thread with a Runnable lambda.
Thread worker = new Thread(() ->
System.out.println("hello from " + Thread.currentThread().getName()));
worker.start(); // start() creates the new thread; run() would not
worker.join(); // wait for it to finish
// 2. A fixed thread pool running tasks that return values.
ExecutorService pool = Executors.newFixedThreadPool(4);
List<Future<Long>> results = new ArrayList<>();
for (int t = 0; t < 4; t++) {
final long start = t * 250_000L + 1;
final long end = start + 250_000L; // exclusive
results.add(pool.submit(() -> {
long s = 0;
for (long i = start; i < end; i++) s += i;
return s;
}));
}
long total = 0;
for (Future<Long> f : results) total += f.get(); // get() blocks until done
pool.shutdown();
System.out.println("total = " + total);
}
}
Output:
hello from Thread-0
total = 500000500000
Points interviewers commonly check:
start()versusrun():start()creates a new thread that then callsrun(). Callingrun()directly just runs the method on the current thread, with no concurrency at all.RunnableversusCallable:Runnable.run()returns nothing and cannot throw checked exceptions;Callable.call()returns a value and can throw.submitwith a lambda that returns a value creates aCallable.Future: a handle to a result that is not ready yet.get()blocks until it is.- Extending
Threadversus implementingRunnable: preferRunnable(or lambdas): Java has single inheritance, and separating the task from the thread lets the same task run in a pool. - Daemon threads: the JVM exits when only daemon threads remain; user threads keep it alive.
- Virtual threads (Java 21 and later):
Thread.ofVirtual().start(task)orExecutors.newVirtualThreadPerTaskExecutor()create lightweight threads managed by the JVM, scheduled many-to-many on carrier threads. They make the "one thread per request" style cheap for I/O-heavy servers.
Java thread states, if asked: NEW, RUNNABLE (covers both ready and running), BLOCKED (waiting to enter a synchronized block), WAITING, TIMED_WAITING and TERMINATED.
Thread pools
Creating a thread for every task has costs: creating and destroying kernel threads, stack memory for each, and with thousands of tasks, so many threads that the machine spends its time switching between them. A thread pool creates a fixed (or bounded) set of worker threads once and feeds them tasks through a queue.
submit(task) --> +----------------------------+
| task queue (bounded) |
| [t7][t6][t5][t4] |
+-------------+--------------+
| workers take tasks
+-------------------+-------------------+
v v v
+---------+ +---------+ +---------+
| worker1 | | worker2 | | worker3 |
| runs t1 | | runs t2 | | runs t3 |
+---------+ +---------+ +---------+
Benefits:
- No per-task creation cost: threads are reused.
- Bounded concurrency: the pool size caps how many tasks run at once, protecting the CPU and downstream services such as a database.
- Backpressure: a bounded queue lets you reject or slow down new work when the system is overloaded, instead of running out of memory.
How big should the pool be?
A common rule of thumb:
- CPU-bound tasks (computation, compression): about one thread per core, perhaps cores + 1. More threads only add switching overhead.
- I/O-bound tasks (waiting on network or disk): more threads than cores, because most are waiting. A well-known estimate is
threads = cores × (1 + wait time / compute time).
Worked example: an 8-core server handles requests that spend 90 ms waiting on a database and 10 ms computing. Then wait/compute = 90/10 = 9, so threads ≈ 8 × (1 + 9) = 80. Treat this as a starting point and confirm with load testing, since the database may be the real limit.
Common mistake
Using an unbounded queue in production. Java's Executors.newFixedThreadPool uses an unbounded LinkedBlockingQueue, so under sustained overload tasks pile up until the process runs out of memory. Production code often builds a ThreadPoolExecutor with a bounded queue and an explicit rejection policy.
Another classic pool deadlock: a task running in a pool submits a sub-task to the same pool and waits for it. If every worker is doing this, all are waiting for sub-tasks that no free worker can run. Java's ForkJoinPool uses work stealing (idle workers take tasks from the back of busy workers' queues) and is designed for this divide-and-conquer pattern.
Concurrency versus parallelism
These words are often used interchangeably, but they mean different things.
- Concurrency is about structure: several tasks are in progress during overlapping time periods. They may take turns on a single core.
- Parallelism is about execution: several tasks literally run at the same instant on different cores.
Concurrency on 1 core (interleaved):
core0: [A][B][A][C][B][A][C]
Parallelism on 3 cores (simultaneous):
core0: [A A A A A A A]
core1: [B B B B B B B]
core2: [C C C C C C C]
You can have concurrency without parallelism (a single-core machine running many threads, or a JavaScript event loop handling many requests), and parallelism without much concurrency structure (a GPU applying the same operation to a million pixels at once). Concurrency is a way of dealing with many things; parallelism is doing many things at once.
Multithreading versus multiprocessing
Multithreading runs several threads inside one process; multiprocessing runs several processes, possibly with a few threads each.
| Multithreading | Multiprocessing | |
|---|---|---|
| Memory | Shared | Separate (shared only by explicit IPC) |
| Startup and memory cost | Lower | Higher |
| Communication | Direct, needs locks | IPC, serialization |
| Failure isolation | One crash takes down all threads | Crash is contained |
| Best for | Shared-state servers, I/O concurrency | Isolation, untrusted code, avoiding a GIL |
The Python GIL
CPython (the standard Python interpreter) has a Global Interpreter Lock (GIL): a lock that only one thread can hold at a time while executing Python bytecode. It simplifies the interpreter's memory management (reference counting) but means that CPU-bound Python threads do not run in parallel. I/O-bound threads still help, because a thread releases the GIL while it waits on I/O, and many C extensions such as NumPy release it during heavy computation.
This script runs the same CPU-bound function four times, once with a thread pool and once with a process pool:
import time
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
def count(n):
total = 0
for i in range(n):
total += i
return total
def timed(executor_cls, workers=4, n=5_000_000):
start = time.perf_counter()
with executor_cls(max_workers=workers) as ex:
list(ex.map(count, [n] * workers))
return time.perf_counter() - start
if __name__ == "__main__":
print(f"threads: {timed(ThreadPoolExecutor):.2f} s")
print(f"processes: {timed(ProcessPoolExecutor):.2f} s")
On a standard CPython build on a multi-core machine, the process version finishes noticeably faster, because each process has its own interpreter and its own GIL. Exact times depend on your hardware. Python 3.13 added an optional free-threaded build without the GIL (PEP 703); the default build still has one, so in interviews say "CPython has a GIL by default" and mention the free-threaded option as recent.
Other languages differ: Java, C#, Go, C and C++ threads run truly in parallel. Node.js runs your JavaScript on a single thread with an event loop, using a background thread pool for some I/O, and worker_threads for parallel computation.
Hyper-threading (SMT)
Simultaneous multithreading (SMT) is a hardware feature in which one physical core presents itself as two (or more) logical processors. Intel's name for it is Hyper-Threading; AMD Zen cores also support two-way SMT.
one physical core with 2-way SMT
+-------------------------------------------+
| logical CPU 0 logical CPU 1 |
| [registers, PC] [registers, PC] | duplicated
|-------------------------------------------|
| shared: execution units, L1/L2 caches, | shared
| branch predictor, TLB |
+-------------------------------------------+
Each logical processor has its own architectural state (registers, program counter), so the OS schedules a thread on each as if they were separate cores. But they share the core's execution units and caches. When one hardware thread stalls, for instance waiting on a cache miss, the other can use the idle execution units.
The gain depends heavily on the workload. Code that often stalls on memory can benefit a lot; code that already keeps the execution units busy may gain little or even slow down because the two threads compete for cache. Two logical processors are not two cores. Some security-sensitive environments disable SMT because shared core resources have been used for side-channel attacks. On Linux, lscpu shows "Thread(s) per core".
Green threads, coroutines and async
Not all concurrency uses OS threads. Several lighter-weight models are common in modern languages.
- Green threads: user-level threads scheduled by a runtime instead of the OS. The name comes from early Java. Go goroutines, Erlang processes and Java virtual threads are modern descendants; they start with small stacks (a few kilobytes for goroutines, growing as needed) so a program can run hundreds of thousands of them.
- Coroutines: functions that can suspend themselves at defined points and be resumed later, keeping their local state. Switching is cooperative: a coroutine gives up control voluntarily, typically when it waits on I/O, rather than being preempted by a timer. Python generators and
async deffunctions, Kotlin coroutines and C++20 coroutines are examples. - async/await with an event loop: an event loop is a single thread that keeps a list of tasks and uses OS readiness notifications (
epollon Linux,kqueueon BSD and macOS, IOCP on Windows) to learn which I/O is ready. When a coroutine hitsawaiton I/O that is not ready, it is suspended and the loop runs another one.
import asyncio
async def fetch(name, delay):
await asyncio.sleep(delay) # stands in for a network call
return f"{name} done after {delay}s"
async def main():
results = await asyncio.gather(fetch("a", 1), fetch("b", 1), fetch("c", 1))
print(results) # all three finish in about 1s, not 3s
asyncio.run(main())
All three "requests" wait concurrently on one thread, so the program takes about one second instead of three. This is concurrency with no parallelism.
The main hazard of cooperative models: a coroutine that does heavy computation, or calls a blocking function like time.sleep instead of await asyncio.sleep, never yields, so the entire event loop stalls and every other task waits. CPU-heavy work belongs in a thread or process pool.
| Model | Scheduled by | Switching | Parallel on many cores? | Typical use |
|---|---|---|---|---|
| OS (kernel) threads | Kernel | Preemptive | Yes | General purpose |
| Green threads / goroutines / virtual threads | Runtime, many-to-many | Runtime decides, at blocking points or safe points | Yes, via several OS threads | Massive I/O concurrency |
| Coroutines with event loop | Program / event loop | Cooperative, at await | No (one loop thread) | Network servers, I/O-heavy scripts |
Thread safety and reentrancy
Code is thread-safe if it behaves correctly when called from several threads at the same time, with no additional coordination by the caller. The usual problem is shared mutable state. This function is not thread-safe:
static int counter = 0;
int next_id(void) {
return ++counter; /* read, add, write: three steps that can interleave */
}
Two threads can both read 5 and both write 6, so two callers get the same ID. The synchronization lesson walks through the exact interleaving. Ways to make code thread-safe:
- Avoid shared state: use local variables, or give each thread its own data (as in the pthread sum example).
- Immutability: data that is never modified after creation can be shared freely.
- Mutual exclusion: protect shared state with a lock.
- Atomic operations: for simple counters,
atomic_fetch_addin C11 orAtomicIntegerin Java. - Thread-local storage: each thread gets its own copy.
A function is reentrant if it can be safely interrupted in the middle and called again (re-entered) before the first call finishes, for example by a signal handler on the same thread or through recursion. A reentrant function uses only its parameters and local variables; it does not use static or global state, does not return pointers to static buffers and does not call non-reentrant functions.
The classic example is strtok, which remembers its position in a hidden static variable. Two threads (or a nested call) using strtok at once corrupt each other's position. The fix is strtok_r, where the caller supplies the position variable:
#include <stdio.h>
#include <string.h>
int main(void) {
char text[] = "a,b,c";
char *saveptr; /* caller-owned state */
for (char *tok = strtok_r(text, ",", &saveptr); tok != NULL;
tok = strtok_r(NULL, ",", &saveptr))
printf("%s\n", tok);
return 0;
}
How are the two ideas related?
- Reentrant but not thread-safe: rare in practice. A function that only touches its arguments is reentrant, but if two threads pass it pointers to the same data, the caller must still synchronize.
- Thread-safe but not reentrant: common. A function that protects a global with a mutex is thread-safe, but if a signal handler interrupts it while it holds the lock and calls it again on the same thread, the handler blocks on a lock its own thread holds: a self-deadlock.
| Property | Means | Typical technique | Example |
|---|---|---|---|
| Thread-safe | Correct when called concurrently from many threads | Locks, atomics, no shared state | malloc in glibc |
| Reentrant | Correct when re-entered before the previous call finishes, even on one thread | No static state, only args and locals | strtok_r |
| Neither | Uses hidden static state | strtok |
Interview tip
The clean distinction: "Thread safety is about many threads calling a function at once; reentrancy is about a function being interrupted and called again before it finishes, possibly on the same thread. A mutex-protected function is thread-safe but not reentrant, because re-entering on the same thread deadlocks on its own lock."
Interview questions
Q1. What is a thread, and how is it different from a process?
A thread is a single flow of execution inside a process, with its own program counter, registers and stack. A process owns resources such as the address space and open files, and all its threads share them. Threads are cheaper to create and switch between and can share data directly, but a crash or memory bug in one thread affects the whole process, while processes are isolated from each other.
Q2. What do threads of the same process share, and what is private?
They share the code, global and static data, the heap, open file descriptors, signal handlers, the working directory and the PID. Each thread has its own thread ID, registers and program counter, stack, signal mask and thread-local storage, including its own errno. Stacks are private by convention but live in the shared address space.
Q3. Why is switching between threads cheaper than switching between processes?
Threads of the same process share an address space, so a switch does not change the page-table base and need not flush or invalidate TLB entries. Only registers, the stack pointer and some kernel bookkeeping change. The new thread is also likely to find shared code and data still warm in the cache.
Q4. Compare user-level and kernel-level threads.
User-level threads are managed by a library or runtime; they are very cheap to create and switch, but the kernel does not know about them, so under a many-to-one mapping a blocking system call blocks them all and they cannot use multiple cores. Kernel-level threads are scheduled by the OS, block independently and run in parallel, but each operation involves the kernel and costs more. Modern runtimes like Go combine both through many-to-many scheduling.
Q5. Explain the many-to-one, one-to-one and many-to-many models.
Many-to-one maps all user threads to one kernel thread: fast switching but no parallelism, and one blocking call blocks all. One-to-one gives every user thread its own kernel thread: real parallelism and independent blocking, but each thread is a kernel resource; Linux, Windows and macOS use this. Many-to-many multiplexes many user threads onto fewer kernel threads, combining cheap threads with parallelism; Go goroutines and Java virtual threads work this way.
Q6. What is the difference between start() and run() in Java?
start() asks the JVM to create a new thread, which then executes run(). Calling run() directly is an ordinary method call executed on the current thread, so nothing runs concurrently. Calling start() twice on the same Thread object throws IllegalThreadStateException.
Q7. What is a thread pool, and why use one?
A thread pool keeps a set of worker threads that take tasks from a queue. It avoids the cost of creating and destroying a thread per task, caps how many tasks run at once to protect the CPU and downstream systems, and with a bounded queue provides backpressure under overload. For CPU-bound work, size it near the number of cores; for I/O-bound work, larger, roughly cores times (1 + wait/compute).
Q8. What is the difference between concurrency and parallelism?
Concurrency means multiple tasks make progress in overlapping time periods, possibly by interleaving on one core. Parallelism means multiple tasks execute at the same instant on different cores. A single-core machine can be concurrent but not parallel; an async event loop is concurrent on one thread.
Q9. What is the GIL and how does it affect Python threads?
The Global Interpreter Lock in CPython allows only one thread at a time to execute Python bytecode. CPU-bound Python threads therefore do not speed up on multiple cores, while I/O-bound threads still help because the GIL is released during blocking I/O. For CPU-bound work you use multiprocessing or a process pool, or libraries that release the GIL; Python 3.13 also offers an optional free-threaded build.
Q10. What is hyper-threading, and does it double performance?
Hyper-threading is Intel's implementation of simultaneous multithreading: one physical core exposes two logical processors, each with its own registers, sharing execution units and caches. It improves throughput when one thread stalls, for example on memory, by letting the other use idle units. It does not double performance; the benefit varies by workload and can be close to zero or negative when both threads compete for the same resources.
Q11. What are coroutines, and how do they differ from threads?
Coroutines are functions that can suspend and later resume while keeping their local state. They switch cooperatively, only at explicit points like await, whereas OS threads are preempted by the kernel at any instruction. Coroutines are extremely cheap and avoid many data races because switches happen only at known points, but a coroutine that blocks or computes for a long time stalls everything on its event loop.
Q12. What does thread-safe mean, and how do you make code thread-safe?
Thread-safe code behaves correctly when called from multiple threads at the same time without extra coordination by the caller. You achieve it by avoiding shared mutable state, using immutable data, protecting shared state with locks, using atomic operations for simple updates, or using thread-local storage.
Q13. What is a reentrant function, and how does it relate to thread safety?
A reentrant function can be interrupted and called again before the first call completes, because it relies only on its arguments and local variables. strtok is not reentrant because it keeps hidden static state; strtok_r is. A function made thread-safe with a mutex is usually not reentrant, since re-entering on the same thread, say from a signal handler, would deadlock on its own lock.
Q14. What happens if a multithreaded program calls fork()?
The child gets a copy of the whole address space but only the thread that called fork. Any mutex another thread held at that moment is copied in the locked state with no owner to release it, so the child may deadlock if it uses it, for example inside malloc. The safe practice is to call only async-signal-safe functions in the child until it calls exec.
Q15. If threads are cheaper, why do browsers use multiple processes?
Isolation. Separate processes have separate address spaces, so a crash or memory-corruption exploit in one site's renderer cannot directly read or crash another site's data or the browser itself. The OS can also apply a strict sandbox to each renderer process. The browser accepts the extra memory and IPC cost for this security and stability.
Key takeaways
- A process owns resources; a thread is a unit of execution inside it with its own stack, registers, program counter and thread-local storage.
- Threads share code, globals, heap and open files, so communication is cheap but shared data needs synchronization, and one crash kills them all.
- Linux, Windows and macOS use the one-to-one model; Go goroutines and Java virtual threads are many-to-many runtimes on top.
- Use
pthread_create/pthread_joinin C; in Java preferExecutorServicepools, and rememberstart()versusrun(). - Thread pools reuse threads and bound concurrency; size them by CPU-bound versus I/O-bound work and keep queues bounded.
- Concurrency is overlapping progress; parallelism is simultaneous execution. CPython's GIL blocks CPU-bound thread parallelism by default.
- SMT (hyper-threading) gives two logical CPUs per core sharing execution units; it is not two cores.
- Thread safety concerns concurrent callers; reentrancy concerns re-entry before completion. A locked function is thread-safe but not reentrant.
Next lesson
Continue with CPU scheduling.

