Start with user-visible behavior
A service can have low CPU and still time out waiting for a saturated connection pool. Track request latency, error ratio, throughput and saturation together. Separate successful and failed request latency and inspect tails, not just averages. Observe latency, traffic, errors and saturation as a connected picture.
An SLI is a measured indicator. An SLO sets its target over a window. An SLA is an external agreement with consequences. Define eligible events and excluded maintenance explicitly. A 99.9% request-success objective permits a 0.1% error fraction; it is not automatically a monthly time-based outage allowance.
Logs, metrics, traces and profiles
Metrics aggregate behavior; logs describe discrete events; traces join work across service boundaries; profiles locate CPU/allocation cost. Correlate them with trace/request IDs and deployment versions. Do not put secrets or raw personal documents in logs. OpenTelemetry defines signals and context propagation for instrumenting systems.
High-cardinality metric labels such as user IDs can make a metrics backend expensive. Keep business IDs in controlled logs/traces where appropriate. Sampling reduces cost but can hide rare failures; choose sampling and retention according to diagnostic needs, and preserve meaningful error evidence.
Alert on symptoms and urgency
Page for actionable user harm, not every transient CPU spike. Error-budget burn measures how rapidly allowed failures are being consumed. Fast and slow windows help distinguish urgent outages from sustained degradation. Add runbook links, current context and an owner. Periodically test the alert path itself.
Dashboards should include queue age, consumer lag, pool wait time, DB execution time, cache hit rate, admission rejections and dependency errors. A queue depth of 100 can be fine at 10,000 tasks/second and disastrous at one task/minute.
Latency vs throughput
Latency is time per operation. Throughput is operations per unit time. Batching often increases throughput but adds waiting latency. Parallel fan-out can reduce elapsed time while increasing resource use and the probability that one slow dependency dominates completion.
For independent parallel calls, elapsed time is roughly the slowest branch plus overhead, not their sum. For sequential calls it is roughly the sum. Percentiles cannot generally be added as if each request hit the same percentile in every dependency.
Little's Law and queueing intuition
For a stable system, average items in a boundary equal average arrival rate times average time in that boundary. At 500 requests/second and 0.2 seconds average residence, about 100 requests are in flight. Use consistent boundaries and units.
As a bottleneck approaches saturation, wait time can rise sharply. Increasing app threads may just increase DB queueing. Measure where work waits, bound concurrency and tune the bottleneck. Autoscaling on CPU alone misses I/O saturation and cold-start delay.
Concurrency vs parallelism
Concurrency means multiple tasks make progress within overlapping periods. Parallelism means they execute simultaneously. An event loop can handle concurrent I/O on one thread, but CPU-heavy work blocks it unless delegated. Threads, processes and async tasks have different memory, scheduling and isolation costs.
Protect shared state with appropriate synchronization or ownership/message passing. Deadlocks require a design response: consistent lock order, small critical sections and cancellation. Race detection and load tests help, but tests do not prove all schedules safe.
Separate throughput, bandwidth and execution overlap
Bandwidth describes a path's capacity; throughput measures the useful work actually completed per time unit. Latency measures one operation's delay. A wide network link can still deliver poor throughput under packet loss or an overloaded receiver, and a high-throughput service can have unacceptable tail latency.
Concurrency means multiple activities make progress during overlapping periods. Parallelism means simultaneous execution. One core can interleave tasks while one waits for I/O; two cores can execute independent work simultaneously. More workers can improve overlap until a shared dependency saturates, after which queues and switching overhead can make latency worse. Measure the bottleneck instead of equating a higher thread count with a faster system.
Worked investigation
After deployment, p99 latency rises from a hypothetical 250 ms to 1.8 seconds while average CPU remains 40%. Trace data shows 1.5 seconds waiting for DB connections. Inspect pool size, query duration and transaction lifetime. A new query scans many rows and holds a connection longer. Fix the query/index or access pattern, then verify pool wait and end-to-end latency. Increasing frontend servers can worsen the connection demand.
Benchmark responsibly
Define representative reads/writes, payload sizes and hot-key distributions. Warm and cold caches behave differently. Record client-side latency including queue waiting and avoid coordinated omission, where a blocked load generator stops sending the very requests that would reveal the worst latency.
Run stepped load, spikes, sustained soak and failure experiments. Compare throughput, p50/p95/p99, errors and saturation. Include cost per successful operation and headroom. State hardware/configuration with results; a benchmark number without workload context is not a capacity plan.
Next: scalability and recovery.