Intermediate~20 min read

Testing

The testing pyramid, AAA structure, unit tests with Vitest/Jest, mocking and test doubles, coverage, API integration tests with supertest, Playwright e2e, TDD, and snapshot testing.

Unit TestsVitestTDDPlaywright

Why We Test

Automated tests are executable specifications. They catch regressions before users do, document how code is meant to behave, and give you the confidence to refactor aggressively. The goal is not 100% coverage — it is a suite that fails when behaviour breaks and stays green when you only change implementation details.

Good tests share four properties: they are fast (run in milliseconds so you run them constantly), isolated (no shared state between tests), deterministic (same result every run — no flakiness from time, network, or randomness), and readable (a failing test tells you exactly what broke).

The Testing Pyramid

The pyramid describes the ideal ratio of test types. Most of your tests should be cheap, fast unit tests at the base. Fewer integration tests sit in the middle. A thin layer of slow, expensive end-to-end tests sits at the top. Inverting this — many e2e tests, few unit tests — is the "ice-cream cone" anti-pattern that produces slow, flaky suites.

AspectUnitIntegrationEnd-to-End
ScopeOne function/module in isolationSeveral modules together (e.g. route + DB)Whole system via the UI/API
SpeedMillisecondsTens–hundreds of msSeconds
DependenciesMocked / noneReal (in-memory or test DB)Real browser + backend
CountMany (thousands)Some (hundreds)Few (critical flows only)
ConfidenceLogic is correctPieces work togetherUser can actually do the thing
Failure clarityPinpoints the bugNarrows it down"Something broke"

Testing Trophy

For front-end and API-heavy apps, Kent C. Dodds' "testing trophy" reshapes the pyramid — it puts integration tests as the widest layer because they give the best confidence-per-effort. Static analysis (TypeScript, ESLint) forms the base. Either way, the principle is the same: prefer the cheapest test that gives you real confidence.

Anatomy of a Test: Arrange–Act–Assert

Every good test has three phases. Arrange sets up inputs and state. Act runs the one thing under test. Assert checks the outcome. Keeping these visually separated makes tests scannable. A test should ideally have a single logical assertion — if you are asserting many unrelated things, it is probably several tests.

javascript
import { describe, it, expect } from 'vitest';
import { applyDiscount } from './pricing';

describe('applyDiscount', () => {
  it('subtracts a percentage from the price', () => {
    // Arrange
    const price = 100;
    const percentOff = 20;

    // Act
    const result = applyDiscount(price, percentOff);

    // Assert
    expect(result).toBe(80);
  });

  it('throws when the discount exceeds 100%', () => {
    expect(() => applyDiscount(100, 150)).toThrow(/invalid discount/i);
  });
});

Unit Tests with Vitest & Jest

Vitest is the modern default for Vite/ESM projects — it is fast, TypeScript-native, and its API is a superset of Jest. Both share the same core structure: describe groups tests, it/test defines a case, and expect makes assertions. Lifecycle hooks let you share setup and guarantee cleanup.

javascript
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { Cart } from './cart';

describe('Cart', () => {
  let cart;

  beforeEach(() => { cart = new Cart(); });   // fresh state per test
  afterEach(() => { cart.clear(); });         // cleanup even on failure

  it('starts empty', () => {
    expect(cart.items).toHaveLength(0);
    expect(cart.total()).toBe(0);
  });

  it('sums line items', () => {
    cart.add({ id: 1, price: 30, qty: 2 });
    cart.add({ id: 2, price: 40, qty: 1 });
    expect(cart.total()).toBe(100);
  });
});

Use it.each (or test.each) for table-driven tests instead of copy-pasting near-identical cases:

javascript
it.each([
  [0,   'Free'],
  [50,  'Standard'],
  [200, 'Priority'],
])('tier for $%i is %s', (spend, expected) => {
  expect(shippingTier(spend)).toBe(expected);
});

Assertions & Matchers

Matchers express intent and produce readable failure messages. Prefer the most specific matcher — toEqual over checking each field, toThrow over try/catch.

MatcherUse for
toBe(x)Primitives & reference identity (===)
toEqual(x)Deep structural equality of objects/arrays
toStrictEqual(x)Deep equality + type & undefined-key checks
toContain(x)Array membership or substring
toThrow(err)A function throwing (wrap the call in a fn)
toHaveBeenCalledWithAsserting how a mock was invoked
resolves / rejectsAwaiting a promise inside the assertion
javascript
expect(user).toEqual({ id: 1, name: 'Ada' });     // deep, ignores instance
expect([1, 2, 3]).toContain(2);
expect(() => parse('')).toThrow(TypeError);
await expect(fetchUser(1)).resolves.toMatchObject({ id: 1 });
await expect(fetchUser(-1)).rejects.toThrow(/not found/);

Test Doubles: Mocks, Stubs & Spies

"Test double" is the umbrella term for anything you swap in for a real dependency. Knowing the distinctions keeps tests honest — over-mocking couples tests to implementation and lets bugs slip through.

DoubleWhat it does
DummyPassed but never used — fills a parameter slot
StubReturns canned values; no behaviour verification
SpyWraps a real function and records how it was called
MockPre-programmed with expectations you assert against
FakeWorking lightweight implementation (e.g. in-memory DB)
javascript
import { describe, it, expect, vi } from 'vitest';
import { notifyUser } from './notify';

// Mock the whole module — factory returns fake implementations
vi.mock('./emailClient', () => ({
  send: vi.fn().mockResolvedValue({ ok: true }),
}));
import { send } from './emailClient';

it('emails the user on signup', async () => {
  await notifyUser({ email: 'ada@example.com' });

  expect(send).toHaveBeenCalledOnce();
  expect(send).toHaveBeenCalledWith(
    expect.objectContaining({ to: 'ada@example.com' })
  );
});

it('spies on a real method without replacing it', () => {
  const logger = { warn: (m) => m };
  const spy = vi.spyOn(logger, 'warn');
  logger.warn('careful');
  expect(spy).toHaveBeenCalledWith('careful');
  spy.mockRestore();
});

Mock at the boundary

Mock the network, filesystem, clock, and randomness — the things that are slow or non-deterministic. Do not mock the code you own and want to test. Use vi.useFakeTimers() for time-based logic and a library like MSW to intercept HTTP at the network layer rather than stubbing your own fetch wrappers.

Code Coverage

Coverage measures how much of your code the tests execute. Vitest uses V8 or Istanbul under the hood. The four metrics are statements, branches (each side of every if/ternary), functions, and lines. Branch coverage is the most informative — 100% line coverage can still miss an untested else.

typescript
// vitest.config.ts
export default {
  test: {
    coverage: {
      provider: 'v8',
      reporter: ['text', 'html', 'lcov'],
      thresholds: { statements: 80, branches: 75, functions: 80, lines: 80 },
    },
  },
};
// run:  vitest run --coverage

Treat coverage as a smoke detector, not a goal. High coverage with weak assertions is worse than moderate coverage that actually checks behaviour. Chasing 100% often produces brittle tests of trivial getters.

Integration Testing an API with Supertest

Integration tests exercise real collaborators together. For an HTTP API, supertest spins up your Express/Fastify app in-process and lets you make requests without opening a real port. Pair it with an in-memory or throwaway test database so tests stay fast and isolated.

javascript
import request from 'supertest';
import { describe, it, expect, beforeEach } from 'vitest';
import { app } from '../app';
import { resetDb, seedUser } from './helpers';

describe('POST /api/login', () => {
  beforeEach(async () => {
    await resetDb();
    await seedUser({ email: 'ada@example.com', password: 'hunter2' });
  });

  it('returns 200 and a token for valid credentials', async () => {
    const res = await request(app)
      .post('/api/login')
      .send({ email: 'ada@example.com', password: 'hunter2' });

    expect(res.status).toBe(200);
    expect(res.body).toHaveProperty('token');
  });

  it('returns 401 for a wrong password', async () => {
    const res = await request(app)
      .post('/api/login')
      .send({ email: 'ada@example.com', password: 'wrong' });

    expect(res.status).toBe(401);
  });
});

End-to-End Tests with Playwright

E2E tests drive a real browser through a real user journey against a running app. Playwright is the modern default: it runs across Chromium, Firefox, and WebKit, auto-waits for elements (killing most flakiness), and encourages accessible, user-facing locators like getByRole. Reserve e2e for a handful of critical paths — login, checkout, signup.

javascript
import { test, expect } from '@playwright/test';

test('user can log in and reach the dashboard', async ({ page }) => {
  await page.goto('/login');
  await page.getByLabel('Email').fill('ada@example.com');
  await page.getByLabel('Password').fill('hunter2');
  await page.getByRole('button', { name: 'Sign in' }).click();

  // auto-waits for navigation and the element to appear
  await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();
  await expect(page).toHaveURL(/\/dashboard/);
});

Test-Driven Development (TDD)

TDD flips the order: write the test first, watch it fail, then write just enough code to pass. The cycle is Red → Green → Refactor.

  1. Red — write a small failing test that describes the next behaviour you want. Run it and confirm it fails for the right reason.
  2. Green — write the simplest code that makes the test pass, even if it is ugly. Do not add anything the test does not demand.
  3. Refactor — with a green safety net, clean up names, remove duplication, and improve structure. Re-run tests; they must stay green.

TDD's payoff is design pressure: writing the test first forces you to think about the interface before the implementation, and you end up with a suite by construction.

Snapshot Testing

A snapshot test serialises output (a rendered component, a config object) and stores it in a .snap file. On later runs it diffs the current output against the saved snapshot. It is excellent for catching unintended changes, but easy to abuse — a huge snapshot no one reads gets rubber-stamped with --update on every failure.

javascript
it('formats an invoice', () => {
  expect(formatInvoice(order)).toMatchSnapshot();
});

// inline snapshot — lives right in the test, reviewed in PRs
it('builds a slug', () => {
  expect(slugify('Hello World!')).toMatchInlineSnapshot(`"hello-world"`);
});
// update after an intentional change:  vitest -u

Snapshot discipline

Keep snapshots small and focused, prefer inline snapshots so diffs show up in code review, and never blindly run -u to make a failure disappear — read the diff first. A snapshot that changes on every run is testing nothing.

Testing Best Practices

  • Test behaviour, not implementation. Assert on public outputs and observable effects, not private internals — so refactors don't break tests.
  • One reason to fail per test. A focused test tells you exactly what broke.
  • Descriptive names. "returns 401 for a wrong password" beats "test login 2".
  • Keep tests independent. No test should depend on another running first; reset state in beforeEach.
  • Kill flakiness at the source. Fake time, seed randomness, and use auto-waiting locators rather than sleep.
  • Run tests in CI on every push, and gate merges on a green suite.

Practice Exercises

  1. Write unit tests for a slugify function covering spaces, punctuation, accents, and empty input using it.each.
  2. Refactor a function that calls Date.now() so it is testable, then verify it with fake timers.
  3. Build a "retry with backoff" helper using TDD — write the failing test for each behaviour before implementing it.
  4. Write a supertest integration test for a POST /todos endpoint covering the happy path, validation error (400), and duplicate (409).
  5. Add a Playwright e2e test for a signup flow and make it resilient by using role-based locators and auto-waiting assertions.
  6. Introduce a coverage threshold in CI, find an uncovered branch, and add a test that exercises it.

Section navigation