contentintech
Intermediate

Monitoring & Logging Cheatsheet

Quick reference for observability — PromQL queries, Prometheus scrape config, common exporters, Alertmanager rules, LogQL for Loki, structured logging fields, and SLI/SLO formulas.

PrometheusGrafanaObservabilitySLO
NotesCheatsheet

Observability Pillars

Signals

Signal Tool Query lang
MetricsPrometheusPromQL
LogsLoki / ELKLogQL / KQL
TracesTempo / JaegerTraceQL

PromQL Essentials

Rates & aggregation

# per-second rate of a counter
rate(http_requests_total[5m])

# total increase over a window
increase(http_requests_total[1h])

# sum grouped by label
sum by (route) (rate(http_requests_total[5m]))

# average, max, min
avg(node_cpu_seconds_total)
max by (instance) (node_load1)

Percentiles & ratios

# p95 latency from a histogram
histogram_quantile(0.95,
  sum by (le) (rate(http_request_duration_seconds_bucket[5m])))

# error ratio
sum(rate(http_requests_total{code=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m]))

# predict disk full in 4h
predict_linear(node_filesystem_avail_bytes[6h], 4*3600) < 0

Operators

Op Use
=~ / !~Regex label match / no match
by / withoutGroup aggregation by/excluding labels
offset 1hShift query back in time

Prometheus Config

Scrape job

global:
  scrape_interval: 15s
scrape_configs:
  - job_name: "app"
    metrics_path: /metrics
    static_configs:
      - targets: ["app:8080"]

Common exporters

Exporter Port
node_exporter9100
cadvisor8080
blackbox_exporter9115
postgres_exporter9187
Prometheus / Alertmanager9090 / 9093

Alerting

Alert rule

groups:
  - name: api
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(http_requests_total{code=~"5.."}[5m]))
            / sum(rate(http_requests_total[5m])) > 0.05
        for: 10m
        labels: { severity: page }
        annotations:
          summary: "5xx > 5% for 10m"

Logs (LogQL)

Queries

# filter lines
{app="api"} |= "error"

# parse JSON then filter
{app="api"} | json | status >= 500

# error rate per pod
sum by (pod) (rate({app="api"} |= "level=error" [1m]))

Structured log fields

ts        RFC3339 timestamp
level      debug|info|warn|error
service    logical service name
trace_id   correlate with traces
msg        human-readable message

SLO Reference

Availability → downtime

SLO Downtime / 30 days
99%~7.2 hours
99.9%~43 minutes
99.99%~4.3 minutes

What to measure

Method Signals
Golden SignalsLatency, Traffic, Errors, Saturation
REDRate, Errors, Duration
USEUtilization, Saturation, Errors

Section navigation