Why We Test
Automated tests are executable specifications. They catch regressions before users do, document how code is meant to behave, and give you the confidence to refactor aggressively. The goal is not 100% coverage — it is a suite that fails when behaviour breaks and stays green when you only change implementation details.
Good tests share four properties: they are fast (run in milliseconds so you run them constantly), isolated (no shared state between tests), deterministic (same result every run — no flakiness from time, network, or randomness), and readable (a failing test tells you exactly what broke).
The Testing Pyramid
The pyramid describes the ideal ratio of test types. Most of your tests should be cheap, fast unit tests at the base. Fewer integration tests sit in the middle. A thin layer of slow, expensive end-to-end tests sits at the top. Inverting this — many e2e tests, few unit tests — is the "ice-cream cone" anti-pattern that produces slow, flaky suites.
| Aspect | Unit | Integration | End-to-End |
|---|---|---|---|
| Scope | One function/module in isolation | Several modules together (e.g. route + DB) | Whole system via the UI/API |
| Speed | Milliseconds | Tens–hundreds of ms | Seconds |
| Dependencies | Mocked / none | Real (in-memory or test DB) | Real browser + backend |
| Count | Many (thousands) | Some (hundreds) | Few (critical flows only) |
| Confidence | Logic is correct | Pieces work together | User can actually do the thing |
| Failure clarity | Pinpoints the bug | Narrows it down | "Something broke" |
Testing Trophy
For front-end and API-heavy apps, Kent C. Dodds' "testing trophy" reshapes the pyramid — it puts integration tests as the widest layer because they give the best confidence-per-effort. Static analysis (TypeScript, ESLint) forms the base. Either way, the principle is the same: prefer the cheapest test that gives you real confidence.
Anatomy of a Test: Arrange–Act–Assert
Every good test has three phases. Arrange sets up inputs and state. Act runs the one thing under test. Assert checks the outcome. Keeping these visually separated makes tests scannable. A test should ideally have a single logical assertion — if you are asserting many unrelated things, it is probably several tests.
import { describe, it, expect } from 'vitest';
import { applyDiscount } from './pricing';
describe('applyDiscount', () => {
it('subtracts a percentage from the price', () => {
// Arrange
const price = 100;
const percentOff = 20;
// Act
const result = applyDiscount(price, percentOff);
// Assert
expect(result).toBe(80);
});
it('throws when the discount exceeds 100%', () => {
expect(() => applyDiscount(100, 150)).toThrow(/invalid discount/i);
});
});
Unit Tests with Vitest & Jest
Vitest is the modern default for Vite/ESM projects — it is fast, TypeScript-native, and its API is a superset of Jest. Both share the same core structure: describe groups tests, it/test defines a case, and expect makes assertions. Lifecycle hooks let you share setup and guarantee cleanup.
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { Cart } from './cart';
describe('Cart', () => {
let cart;
beforeEach(() => { cart = new Cart(); }); // fresh state per test
afterEach(() => { cart.clear(); }); // cleanup even on failure
it('starts empty', () => {
expect(cart.items).toHaveLength(0);
expect(cart.total()).toBe(0);
});
it('sums line items', () => {
cart.add({ id: 1, price: 30, qty: 2 });
cart.add({ id: 2, price: 40, qty: 1 });
expect(cart.total()).toBe(100);
});
});
Use it.each (or test.each) for table-driven tests instead of copy-pasting near-identical cases:
it.each([
[0, 'Free'],
[50, 'Standard'],
[200, 'Priority'],
])('tier for $%i is %s', (spend, expected) => {
expect(shippingTier(spend)).toBe(expected);
});
Assertions & Matchers
Matchers express intent and produce readable failure messages. Prefer the most specific matcher — toEqual over checking each field, toThrow over try/catch.
| Matcher | Use for |
|---|---|
toBe(x) | Primitives & reference identity (===) |
toEqual(x) | Deep structural equality of objects/arrays |
toStrictEqual(x) | Deep equality + type & undefined-key checks |
toContain(x) | Array membership or substring |
toThrow(err) | A function throwing (wrap the call in a fn) |
toHaveBeenCalledWith | Asserting how a mock was invoked |
resolves / rejects | Awaiting a promise inside the assertion |
expect(user).toEqual({ id: 1, name: 'Ada' }); // deep, ignores instance
expect([1, 2, 3]).toContain(2);
expect(() => parse('')).toThrow(TypeError);
await expect(fetchUser(1)).resolves.toMatchObject({ id: 1 });
await expect(fetchUser(-1)).rejects.toThrow(/not found/);
Test Doubles: Mocks, Stubs & Spies
"Test double" is the umbrella term for anything you swap in for a real dependency. Knowing the distinctions keeps tests honest — over-mocking couples tests to implementation and lets bugs slip through.
| Double | What it does |
|---|---|
| Dummy | Passed but never used — fills a parameter slot |
| Stub | Returns canned values; no behaviour verification |
| Spy | Wraps a real function and records how it was called |
| Mock | Pre-programmed with expectations you assert against |
| Fake | Working lightweight implementation (e.g. in-memory DB) |
import { describe, it, expect, vi } from 'vitest';
import { notifyUser } from './notify';
// Mock the whole module — factory returns fake implementations
vi.mock('./emailClient', () => ({
send: vi.fn().mockResolvedValue({ ok: true }),
}));
import { send } from './emailClient';
it('emails the user on signup', async () => {
await notifyUser({ email: 'ada@example.com' });
expect(send).toHaveBeenCalledOnce();
expect(send).toHaveBeenCalledWith(
expect.objectContaining({ to: 'ada@example.com' })
);
});
it('spies on a real method without replacing it', () => {
const logger = { warn: (m) => m };
const spy = vi.spyOn(logger, 'warn');
logger.warn('careful');
expect(spy).toHaveBeenCalledWith('careful');
spy.mockRestore();
});
Mock at the boundary
Mock the network, filesystem, clock, and randomness — the things that are slow or non-deterministic. Do not mock the code you own and want to test. Use vi.useFakeTimers() for time-based logic and a library like MSW to intercept HTTP at the network layer rather than stubbing your own fetch wrappers.
Code Coverage
Coverage measures how much of your code the tests execute. Vitest uses V8 or Istanbul under the hood. The four metrics are statements, branches (each side of every if/ternary), functions, and lines. Branch coverage is the most informative — 100% line coverage can still miss an untested else.
// vitest.config.ts
export default {
test: {
coverage: {
provider: 'v8',
reporter: ['text', 'html', 'lcov'],
thresholds: { statements: 80, branches: 75, functions: 80, lines: 80 },
},
},
};
// run: vitest run --coverage
Treat coverage as a smoke detector, not a goal. High coverage with weak assertions is worse than moderate coverage that actually checks behaviour. Chasing 100% often produces brittle tests of trivial getters.
Integration Testing an API with Supertest
Integration tests exercise real collaborators together. For an HTTP API, supertest spins up your Express/Fastify app in-process and lets you make requests without opening a real port. Pair it with an in-memory or throwaway test database so tests stay fast and isolated.
import request from 'supertest';
import { describe, it, expect, beforeEach } from 'vitest';
import { app } from '../app';
import { resetDb, seedUser } from './helpers';
describe('POST /api/login', () => {
beforeEach(async () => {
await resetDb();
await seedUser({ email: 'ada@example.com', password: 'hunter2' });
});
it('returns 200 and a token for valid credentials', async () => {
const res = await request(app)
.post('/api/login')
.send({ email: 'ada@example.com', password: 'hunter2' });
expect(res.status).toBe(200);
expect(res.body).toHaveProperty('token');
});
it('returns 401 for a wrong password', async () => {
const res = await request(app)
.post('/api/login')
.send({ email: 'ada@example.com', password: 'wrong' });
expect(res.status).toBe(401);
});
});
End-to-End Tests with Playwright
E2E tests drive a real browser through a real user journey against a running app. Playwright is the modern default: it runs across Chromium, Firefox, and WebKit, auto-waits for elements (killing most flakiness), and encourages accessible, user-facing locators like getByRole. Reserve e2e for a handful of critical paths — login, checkout, signup.
import { test, expect } from '@playwright/test';
test('user can log in and reach the dashboard', async ({ page }) => {
await page.goto('/login');
await page.getByLabel('Email').fill('ada@example.com');
await page.getByLabel('Password').fill('hunter2');
await page.getByRole('button', { name: 'Sign in' }).click();
// auto-waits for navigation and the element to appear
await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();
await expect(page).toHaveURL(/\/dashboard/);
});
Test-Driven Development (TDD)
TDD flips the order: write the test first, watch it fail, then write just enough code to pass. The cycle is Red → Green → Refactor.
- Red — write a small failing test that describes the next behaviour you want. Run it and confirm it fails for the right reason.
- Green — write the simplest code that makes the test pass, even if it is ugly. Do not add anything the test does not demand.
- Refactor — with a green safety net, clean up names, remove duplication, and improve structure. Re-run tests; they must stay green.
TDD's payoff is design pressure: writing the test first forces you to think about the interface before the implementation, and you end up with a suite by construction.
Snapshot Testing
A snapshot test serialises output (a rendered component, a config object) and stores it in a .snap file. On later runs it diffs the current output against the saved snapshot. It is excellent for catching unintended changes, but easy to abuse — a huge snapshot no one reads gets rubber-stamped with --update on every failure.
it('formats an invoice', () => {
expect(formatInvoice(order)).toMatchSnapshot();
});
// inline snapshot — lives right in the test, reviewed in PRs
it('builds a slug', () => {
expect(slugify('Hello World!')).toMatchInlineSnapshot(`"hello-world"`);
});
// update after an intentional change: vitest -u
Snapshot discipline
Keep snapshots small and focused, prefer inline snapshots so diffs show up in code review, and never blindly run -u to make a failure disappear — read the diff first. A snapshot that changes on every run is testing nothing.
Testing Best Practices
- Test behaviour, not implementation. Assert on public outputs and observable effects, not private internals — so refactors don't break tests.
- One reason to fail per test. A focused test tells you exactly what broke.
- Descriptive names. "returns 401 for a wrong password" beats "test login 2".
- Keep tests independent. No test should depend on another running first; reset state in
beforeEach. - Kill flakiness at the source. Fake time, seed randomness, and use auto-waiting locators rather than
sleep. - Run tests in CI on every push, and gate merges on a green suite.
Practice Exercises
- Write unit tests for a
slugifyfunction covering spaces, punctuation, accents, and empty input usingit.each. - Refactor a function that calls
Date.now()so it is testable, then verify it with fake timers. - Build a "retry with backoff" helper using TDD — write the failing test for each behaviour before implementing it.
- Write a supertest integration test for a
POST /todosendpoint covering the happy path, validation error (400), and duplicate (409). - Add a Playwright e2e test for a signup flow and make it resilient by using role-based locators and auto-waiting assertions.
- Introduce a coverage threshold in CI, find an uncovered branch, and add a test that exercises it.