Why Observability Matters
Monitoring tells you whether a system is working; observability lets you ask arbitrary questions about why it is not — without shipping new code. Modern distributed systems fail in ways you did not predict, so you instrument them to expose their internal state through three complementary signal types.
The Three Pillars of Observability
| Pillar | What it answers | Typical tooling |
|---|---|---|
| Metrics | Aggregated numbers over time (rates, counts, gauges). Cheap, good for alerting and trends. | Prometheus, VictoriaMetrics, Grafana |
| Logs | Discrete, timestamped events with rich context. Good for forensic detail. | Loki, Elasticsearch/ELK, OpenSearch |
| Traces | The path of a single request across services, with per-span latency. | OpenTelemetry, Tempo, Jaeger |
2026 convention
OpenTelemetry (OTel) is the vendor-neutral standard for all three signals. Instrument your app with the OTel SDK, export via the OTLP protocol to an OTel Collector, and fan out to Prometheus/Tempo/Loki. This decouples instrumentation from backends.
Prometheus: Metrics Collection
Prometheus is a pull-based time-series database. It scrapes HTTP /metrics endpoints on an interval and stores samples keyed by a metric name plus labels. There are four core metric types:
| Type | Meaning |
|---|---|
counter | Monotonically increasing (requests, errors). Use rate() to read it. |
gauge | Value that goes up and down (memory, queue depth, temperature). |
histogram | Bucketed observations; enables quantiles via histogram_quantile(). |
summary | Client-side computed quantiles; cannot be aggregated across instances. |
Scrape Configuration
Targets are defined in prometheus.yml. In Kubernetes you normally use the Prometheus Operator with ServiceMonitor CRDs instead of static config, but the file below shows the fundamentals.
global:
scrape_interval: 15s
evaluation_interval: 15s
rule_files:
- /etc/prometheus/rules/*.yml
alerting:
alertmanagers:
- static_configs:
- targets: ["alertmanager:9093"]
scrape_configs:
- job_name: "node"
static_configs:
- targets: ["node-exporter:9100"]
- job_name: "api"
metrics_path: /metrics
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [__meta_kubernetes_pod_label_app]
regex: api
action: keep
Exporters
Software that does not expose Prometheus metrics natively is scraped through an exporter — a sidecar process that translates system state into the metrics format:
| Exporter | Exposes |
|---|---|
node_exporter | Host CPU, memory, disk, network |
cadvisor | Per-container resource usage |
blackbox_exporter | HTTP/TCP/ICMP probes (uptime, latency) |
postgres_exporter | Database connections, locks, query stats |
PromQL: Querying Metrics
PromQL is the query language for both dashboards and alert rules. A few patterns cover the vast majority of real queries:
# Per-second request rate over the last 5 minutes
rate(http_requests_total[5m])
# Request rate grouped by route
sum by (route) (rate(http_requests_total[5m]))
# Error ratio (errors / total)
sum(rate(http_requests_total{code=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
# 95th percentile latency from a histogram
histogram_quantile(
0.95,
sum by (le) (rate(http_request_duration_seconds_bucket[5m]))
)
# Memory saturation: used / total, per instance
1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)
# Predict disk full within 4 hours (linear extrapolation)
predict_linear(node_filesystem_avail_bytes[6h], 4 * 3600) < 0
Golden rule of counters
Never graph a raw counter — always wrap it in rate() or increase(). Counters reset to zero on restart, and rate() automatically corrects for those resets.
Grafana Dashboards
Grafana is the visualization layer. It connects to Prometheus (and Loki, Tempo, and many others) as data sources and renders PromQL/LogQL queries into panels. Dashboards should be stored as JSON in git and provisioned automatically rather than clicked together by hand.
# grafana/provisioning/datasources/datasources.yml
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://prometheus:9090
isDefault: true
- name: Loki
type: loki
access: proxy
url: http://loki:3100
Use template variables (e.g. $namespace, $instance) so one dashboard serves every service, and set panel units correctly so 0.95 renders as a percentage, not a raw float.
Alerting with Alertmanager
Prometheus evaluates alerting rules and pushes firing alerts to Alertmanager, which handles grouping, deduplication, silencing, and routing to receivers (Slack, PagerDuty, email, Opsgenie).
# rules/api.yml
groups:
- name: api-availability
rules:
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{code=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m])) > 0.05
for: 10m
labels:
severity: page
annotations:
summary: "5xx error rate above 5% for 10m"
description: "Current ratio: {{ $value | humanizePercentage }}"
The for clause is critical: it requires the condition to hold continuously before firing, which suppresses transient spikes and flapping.
Log Aggregation
Centralizing logs lets you search across every instance from one place. Two dominant stacks:
| Stack | Approach | Trade-off |
|---|---|---|
| Loki + Promtail/Alloy | Indexes only labels, not log body. LogQL mirrors PromQL. | Cheap storage; less powerful full-text search |
| ELK / OpenSearch | Elasticsearch + Logstash/Beats + Kibana. Full inverted index. | Rich search; heavier storage and ops cost |
# LogQL: 5xx lines from the api container in the last hour
{app="api"} |= "status=5" | json | status >= 500
# Count error log lines per second, by pod
sum by (pod) (rate({app="api"} |= "level=error" [1m]))
Structured Logging
Emit logs as JSON with consistent fields instead of free-form text. This makes them queryable and correlatable with traces via a shared trace_id.
{"ts":"2026-06-01T10:22:01Z","level":"error","service":"api",
"trace_id":"4bf92f...","route":"/checkout","status":500,
"latency_ms":842,"msg":"payment gateway timeout"}
SLIs, SLOs, and SLAs
| Term | Definition |
|---|---|
| SLI | Service Level Indicator — a measured number, e.g. "% of requests under 300ms". |
| SLO | Service Level Objective — your internal target for the SLI, e.g. "99.9% over 30 days". |
| SLA | Service Level Agreement — a contractual promise with financial penalties, usually looser than the SLO. |
The gap between 100% and your SLO is the error budget. A 99.9% SLO permits roughly 43 minutes of downtime per 30 days. When the budget is spent, you freeze risky releases and prioritize reliability work.
Golden Signals, RED, and USE
These are frameworks for deciding what to measure so you are not drowning in metrics:
| Method | Signals | Best for |
|---|---|---|
| Four Golden Signals | Latency, Traffic, Errors, Saturation | Any user-facing service (Google SRE) |
| RED | Rate, Errors, Duration | Request-driven microservices |
| USE | Utilization, Saturation, Errors | Resources: CPU, disk, network, queues |
Practice Exercises
- Run Prometheus, node_exporter, and Grafana with Docker Compose. Add node_exporter as a scrape target and confirm samples appear in the Prometheus expression browser.
- Write a PromQL query that returns the 99th percentile request latency per route from a histogram metric, then build a Grafana panel from it with correct time-series units.
- Create an Alertmanager rule that fires when the 5xx error ratio exceeds 2% for 5 minutes, and route it to a Slack webhook receiver.
- Instrument a small HTTP service with the OpenTelemetry SDK to emit a request counter and a latency histogram, then scrape it with Prometheus.
- Deploy Loki with Promtail (or Grafana Alloy), ship your service logs, and write a LogQL query that graphs error-log rate per pod.
- Define an SLO of 99.9% availability for one endpoint, compute the monthly error budget in minutes, and write the PromQL that tracks remaining budget.