contentintech
Learn/devops/Monitoring & Logging
Intermediate~20 min read

Monitoring & Logging

The three pillars of observability, Prometheus metrics and PromQL, exporters and scrape configs, Grafana dashboards, Alertmanager, log aggregation with Loki and ELK, structured logging, SLIs/SLOs/SLAs, the four golden signals, and the RED and USE methods.

PrometheusGrafanaObservabilitySLO

Why Observability Matters

Monitoring tells you whether a system is working; observability lets you ask arbitrary questions about why it is not — without shipping new code. Modern distributed systems fail in ways you did not predict, so you instrument them to expose their internal state through three complementary signal types.

The Three Pillars of Observability

PillarWhat it answersTypical tooling
MetricsAggregated numbers over time (rates, counts, gauges). Cheap, good for alerting and trends.Prometheus, VictoriaMetrics, Grafana
LogsDiscrete, timestamped events with rich context. Good for forensic detail.Loki, Elasticsearch/ELK, OpenSearch
TracesThe path of a single request across services, with per-span latency.OpenTelemetry, Tempo, Jaeger

2026 convention

OpenTelemetry (OTel) is the vendor-neutral standard for all three signals. Instrument your app with the OTel SDK, export via the OTLP protocol to an OTel Collector, and fan out to Prometheus/Tempo/Loki. This decouples instrumentation from backends.

Prometheus: Metrics Collection

Prometheus is a pull-based time-series database. It scrapes HTTP /metrics endpoints on an interval and stores samples keyed by a metric name plus labels. There are four core metric types:

TypeMeaning
counterMonotonically increasing (requests, errors). Use rate() to read it.
gaugeValue that goes up and down (memory, queue depth, temperature).
histogramBucketed observations; enables quantiles via histogram_quantile().
summaryClient-side computed quantiles; cannot be aggregated across instances.

Scrape Configuration

Targets are defined in prometheus.yml. In Kubernetes you normally use the Prometheus Operator with ServiceMonitor CRDs instead of static config, but the file below shows the fundamentals.

yaml
global:
  scrape_interval: 15s
  evaluation_interval: 15s

rule_files:
  - /etc/prometheus/rules/*.yml

alerting:
  alertmanagers:
    - static_configs:
        - targets: ["alertmanager:9093"]

scrape_configs:
  - job_name: "node"
    static_configs:
      - targets: ["node-exporter:9100"]

  - job_name: "api"
    metrics_path: /metrics
    kubernetes_sd_configs:
      - role: pod
    relabel_configs:
      - source_labels: [__meta_kubernetes_pod_label_app]
        regex: api
        action: keep

Exporters

Software that does not expose Prometheus metrics natively is scraped through an exporter — a sidecar process that translates system state into the metrics format:

ExporterExposes
node_exporterHost CPU, memory, disk, network
cadvisorPer-container resource usage
blackbox_exporterHTTP/TCP/ICMP probes (uptime, latency)
postgres_exporterDatabase connections, locks, query stats

PromQL: Querying Metrics

PromQL is the query language for both dashboards and alert rules. A few patterns cover the vast majority of real queries:

text
# Per-second request rate over the last 5 minutes
rate(http_requests_total[5m])

# Request rate grouped by route
sum by (route) (rate(http_requests_total[5m]))

# Error ratio (errors / total)
sum(rate(http_requests_total{code=~"5.."}[5m]))
  /
sum(rate(http_requests_total[5m]))

# 95th percentile latency from a histogram
histogram_quantile(
  0.95,
  sum by (le) (rate(http_request_duration_seconds_bucket[5m]))
)

# Memory saturation: used / total, per instance
1 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)

# Predict disk full within 4 hours (linear extrapolation)
predict_linear(node_filesystem_avail_bytes[6h], 4 * 3600) < 0

Golden rule of counters

Never graph a raw counter — always wrap it in rate() or increase(). Counters reset to zero on restart, and rate() automatically corrects for those resets.

Grafana Dashboards

Grafana is the visualization layer. It connects to Prometheus (and Loki, Tempo, and many others) as data sources and renders PromQL/LogQL queries into panels. Dashboards should be stored as JSON in git and provisioned automatically rather than clicked together by hand.

yaml
# grafana/provisioning/datasources/datasources.yml
apiVersion: 1
datasources:
  - name: Prometheus
    type: prometheus
    access: proxy
    url: http://prometheus:9090
    isDefault: true
  - name: Loki
    type: loki
    access: proxy
    url: http://loki:3100

Use template variables (e.g. $namespace, $instance) so one dashboard serves every service, and set panel units correctly so 0.95 renders as a percentage, not a raw float.

Alerting with Alertmanager

Prometheus evaluates alerting rules and pushes firing alerts to Alertmanager, which handles grouping, deduplication, silencing, and routing to receivers (Slack, PagerDuty, email, Opsgenie).

yaml
# rules/api.yml
groups:
  - name: api-availability
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{code=~"5.."}[5m]))
            / sum(rate(http_requests_total[5m])) > 0.05
        for: 10m
        labels:
          severity: page
        annotations:
          summary: "5xx error rate above 5% for 10m"
          description: "Current ratio: {{ $value | humanizePercentage }}"

The for clause is critical: it requires the condition to hold continuously before firing, which suppresses transient spikes and flapping.

Log Aggregation

Centralizing logs lets you search across every instance from one place. Two dominant stacks:

StackApproachTrade-off
Loki + Promtail/AlloyIndexes only labels, not log body. LogQL mirrors PromQL.Cheap storage; less powerful full-text search
ELK / OpenSearchElasticsearch + Logstash/Beats + Kibana. Full inverted index.Rich search; heavier storage and ops cost
text
# LogQL: 5xx lines from the api container in the last hour
{app="api"} |= "status=5" | json | status >= 500

# Count error log lines per second, by pod
sum by (pod) (rate({app="api"} |= "level=error" [1m]))

Structured Logging

Emit logs as JSON with consistent fields instead of free-form text. This makes them queryable and correlatable with traces via a shared trace_id.

json
{"ts":"2026-06-01T10:22:01Z","level":"error","service":"api",
 "trace_id":"4bf92f...","route":"/checkout","status":500,
 "latency_ms":842,"msg":"payment gateway timeout"}

SLIs, SLOs, and SLAs

TermDefinition
SLIService Level Indicator — a measured number, e.g. "% of requests under 300ms".
SLOService Level Objective — your internal target for the SLI, e.g. "99.9% over 30 days".
SLAService Level Agreement — a contractual promise with financial penalties, usually looser than the SLO.

The gap between 100% and your SLO is the error budget. A 99.9% SLO permits roughly 43 minutes of downtime per 30 days. When the budget is spent, you freeze risky releases and prioritize reliability work.

Golden Signals, RED, and USE

These are frameworks for deciding what to measure so you are not drowning in metrics:

MethodSignalsBest for
Four Golden SignalsLatency, Traffic, Errors, SaturationAny user-facing service (Google SRE)
REDRate, Errors, DurationRequest-driven microservices
USEUtilization, Saturation, ErrorsResources: CPU, disk, network, queues

Practice Exercises

  1. Run Prometheus, node_exporter, and Grafana with Docker Compose. Add node_exporter as a scrape target and confirm samples appear in the Prometheus expression browser.
  2. Write a PromQL query that returns the 99th percentile request latency per route from a histogram metric, then build a Grafana panel from it with correct time-series units.
  3. Create an Alertmanager rule that fires when the 5xx error ratio exceeds 2% for 5 minutes, and route it to a Slack webhook receiver.
  4. Instrument a small HTTP service with the OpenTelemetry SDK to emit a request counter and a latency histogram, then scrape it with Prometheus.
  5. Deploy Loki with Promtail (or Grafana Alloy), ship your service logs, and write a LogQL query that graphs error-log rate per pod.
  6. Define an SLO of 99.9% availability for one endpoint, compute the monthly error budget in minutes, and write the PromQL that tracks remaining budget.

Section navigation