Learn/system design/Microservices
Advanced~25 min read

Microservices

Decomposing monoliths into services — bounded contexts, sync vs async comms, service discovery, mesh, resilience patterns, sagas, and when not to bother.

DDDService MeshSagaResilience

What Microservices Actually Are

A microservices architecture structures an application as a suite of independently deployable services, each owning a single business capability, its own data, and its own release cadence. The defining property is not "small" — it is independent deployability. If you cannot deploy service A without coordinating a release of service B, you have a distributed monolith, which combines the worst of both worlds.

Microservices trade local complexity (a big codebase) for distributed complexity (a network of processes). That trade is only worth it when organizational and scale pressures make the monolith the bottleneck.

Monolith vs Microservices

DimensionMonolithMicroservices
DeploymentOne artifact, all-or-nothingIndependent per service
ScalingScale the whole appScale hot services only
CallsIn-process, nanosecondsNetwork, milliseconds + failure
Data consistencyACID transactionsEventual, sagas
Team autonomyCoupled release trainsTeams ship on their own
DebuggingSingle stack traceDistributed tracing needed
OnboardingWhole codebaseOne service

Rule of thumb

Start with a well-modularized monolith. Extract services only when a module has a distinct scaling profile, a distinct team, or a distinct release cadence. Premature decomposition is the most expensive mistake in the field.

Service Decomposition with DDD

The best decomposition boundary is the bounded context from Domain-Driven Design: a region of the domain where a term has one unambiguous meaning. "Customer" in Billing (payment methods, invoices) is a different model from "Customer" in Shipping (address, delivery slots). Each bounded context becomes a candidate service.

Decompose along business capabilities, not technical layers. A "UI service", "logic service", and "database service" is a layered monolith spread over a network. Good services are vertical slices: orders, inventory, payments.

Aim for high cohesion within a service and loose coupling between services. If two services always change together, they belong together.

Inter-Service Communication

Synchronous: REST and gRPC

Request/response calls are simple to reason about but create temporal coupling — the caller blocks and the callee must be up. REST/JSON is universal and debuggable; gRPC (HTTP/2 + Protobuf) is faster, strongly typed, and supports streaming, favored for internal east-west traffic.

Asynchronous: Events

Services publish events to a broker (Kafka, Pulsar, NATS, RabbitMQ); consumers react on their own schedule. This removes temporal coupling and improves resilience, at the cost of eventual consistency and harder debugging. Prefer async for cross-service workflows; reserve sync for read paths that need an immediate answer.

StyleUse WhenDownside
RESTPublic APIs, simple CRUDVerbose, no typing
gRPCInternal, low-latency, streamingPoor browser support
EventsWorkflows, fan-out, decouplingEventual consistency

API Gateway & Service Discovery

An API gateway is the single entry point for clients. It handles TLS termination, authn/authz, rate limiting, request routing, response aggregation, and protocol translation (e.g. external REST to internal gRPC). It stops clients from needing to know your internal topology.

Service discovery answers "where is service X right now?" when instances come and go dynamically.

Client-side discovery Client --> [Service Registry] --> instance list Client picks an instance and calls it directly (LB in client) Server-side discovery Client --> [Load Balancer] --> [Registry] --> instance LB does the lookup; client stays dumb (k8s Service works this way)

A service registry (Consul, etcd, Eureka, or the Kubernetes control plane) holds the live map of service name to healthy instances, kept fresh by health checks and heartbeats.

Service Mesh

A service mesh (Istio, Linkerd) moves cross-cutting network concerns — mTLS, retries, timeouts, traffic splitting, telemetry — out of your code and into a sidecar proxy (Envoy) deployed next to each service. Application code just calls localhost; the sidecar handles the mesh.

Pod A Pod B +----------+ +-------+ +-------+ +----------+ | app A |--| Envoy |==mTLS==| Envoy |--| app B | +----------+ +-------+ +-------+ +----------+ data plane (sidecars) <---- controlled by ----> control plane (Istiod)

Mesh cost

Sidecars add ~0.5–2 ms latency per hop and double your process count. Adopt a mesh only when you have enough services that duplicated retry/mTLS/observability logic is a real burden — usually past a dozen services.

Resilience Patterns

In a distributed system, dependencies fail constantly. Design for it.

  1. Timeout — never wait forever; a hung call ties up a thread and cascades.
  2. Retry with backoff + jitter — only for idempotent ops; add jitter to avoid thundering herds.
  3. Circuit breaker — after N failures, "open" the circuit and fail fast for a cooldown, then "half-open" to test recovery. Prevents hammering a sick dependency.
  4. Bulkhead — isolate resources (thread pools, connection pools) per dependency so one slow downstream can't drown the whole service.
  5. Rate limiting / load shedding — reject or queue excess load instead of collapsing.
  6. Fallback — degrade gracefully (cached data, default response) when a dependency is down.

Data per Service & Distributed Transactions

Database-per-service is non-negotiable: each service owns its store, and no other service touches it directly. Sharing a database re-couples deployments and destroys autonomy. The cost is that a single business operation may span multiple services — and you lose ACID across them.

Two-Phase Commit (2PC)

A coordinator asks all participants to "prepare", then "commit". It gives atomicity but is blocking, slow, and fragile — a coordinator crash can lock resources indefinitely. Avoid 2PC in microservices.

Sagas

A saga is a sequence of local transactions; if one fails, earlier steps are undone by compensating transactions. Consistency is eventual, not atomic.

Saga typeHowTrade-off
ChoreographyServices react to each other's events, no central brainDecoupled but flow is implicit, hard to trace
OrchestrationA central orchestrator issues commands step by stepExplicit & observable but orchestrator is a coupling point
Order saga (orchestrated): 1. Order: create (PENDING) 2. Payment: charge card -- fail? --> compensate: cancel order 3. Inventory: reserve stock -- fail? --> compensate: refund + cancel 4. Shipping: schedule 5. Order: mark CONFIRMED

Transactional outbox

To publish an event AND commit a DB change atomically, write the event to an outbox table in the same local transaction, then a relay (or CDC via Debezium) ships it to the broker. This avoids the dual-write problem where a crash leaves DB and broker out of sync.

Observability

You cannot attach a debugger to a request that hops 8 services. You need the three pillars, tied together by a correlation / trace ID propagated in headers (W3C traceparent) across every hop.

  • Logs — structured, centralized (Loki, ELK), stamped with trace ID.
  • Metrics — RED (Rate, Errors, Duration) per service; Prometheus + Grafana.
  • Traces — end-to-end request spans (OpenTelemetry, Jaeger, Tempo) to find the slow hop.

Deployment & Config

Package each service as a container and orchestrate with Kubernetes, which provides scheduling, self-healing, service discovery, and rolling updates. Externalize configuration (ConfigMaps, Secrets, or a config service) — never bake environment-specific values into the image.

StrategyBehaviorRisk
RollingReplace instances graduallyMixed versions live briefly
Blue-greenFull parallel env, flip traffic2x resources during cutover
CanaryRoute 1–5% first, watch metricsNeeds good automated metrics

Fallacies of Distributed Computing

Peter Deutsch's eight fallacies — each is a false assumption that will burn you: (1) the network is reliable, (2) latency is zero, (3) bandwidth is infinite, (4) the network is secure, (5) topology doesn't change, (6) there is one administrator, (7) transport cost is zero, (8) the network is homogeneous. Every microservices design must actively defend against all eight.

When NOT to Use Microservices

Skip microservices when: the team is small (fewer than ~15 engineers), the domain is not yet well understood (boundaries will be wrong), you lack CI/CD, containers, and observability maturity, or the product is early-stage and pivoting. In all these cases a modular monolith ships faster and refactors more easily. Microservices are an answer to organizational scaling as much as technical scaling — use them when Conway's Law is fighting you, not before.

Bottom line

Microservices buy independent deployability and team autonomy at the price of distributed-systems complexity. Pay that price deliberately, with the operational tooling to survive it.

Section navigation