Intermediate
Quick reference for observability — PromQL queries, Prometheus scrape config, common exporters, Alertmanager rules, LogQL for Loki, structured logging fields, and SLI/SLO formulas.
PrometheusGrafanaObservabilitySLO
Observability Pillars
Signals
| Signal |
Tool |
Query lang |
| Metrics | Prometheus | PromQL |
| Logs | Loki / ELK | LogQL / KQL |
| Traces | Tempo / Jaeger | TraceQL |
PromQL Essentials
Rates & aggregation
# per-second rate of a counter
rate(http_requests_total[5m])
# total increase over a window
increase(http_requests_total[1h])
# sum grouped by label
sum by (route) (rate(http_requests_total[5m]))
# average, max, min
avg(node_cpu_seconds_total)
max by (instance) (node_load1)
Percentiles & ratios
# p95 latency from a histogram
histogram_quantile(0.95,
sum by (le) (rate(http_request_duration_seconds_bucket[5m])))
# error ratio
sum(rate(http_requests_total{code=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# predict disk full in 4h
predict_linear(node_filesystem_avail_bytes[6h], 4*3600) < 0
Operators
| Op |
Use |
=~ / !~ | Regex label match / no match |
by / without | Group aggregation by/excluding labels |
offset 1h | Shift query back in time |
Prometheus Config
Scrape job
global:
scrape_interval: 15s
scrape_configs:
- job_name: "app"
metrics_path: /metrics
static_configs:
- targets: ["app:8080"]
Common exporters
| Exporter |
Port |
| node_exporter | 9100 |
| cadvisor | 8080 |
| blackbox_exporter | 9115 |
| postgres_exporter | 9187 |
| Prometheus / Alertmanager | 9090 / 9093 |
Alerting
Alert rule
groups:
- name: api
rules:
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{code=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m])) > 0.05
for: 10m
labels: { severity: page }
annotations:
summary: "5xx > 5% for 10m"
Logs (LogQL)
Queries
# filter lines
{app="api"} |= "error"
# parse JSON then filter
{app="api"} | json | status >= 500
# error rate per pod
sum by (pod) (rate({app="api"} |= "level=error" [1m]))
Structured log fields
ts RFC3339 timestamp
level debug|info|warn|error
service logical service name
trace_id correlate with traces
msg human-readable message
SLO Reference
Availability → downtime
| SLO |
Downtime / 30 days |
| 99% | ~7.2 hours |
| 99.9% | ~43 minutes |
| 99.99% | ~4.3 minutes |
What to measure
| Method |
Signals |
| Golden Signals | Latency, Traffic, Errors, Saturation |
| RED | Rate, Errors, Duration |
| USE | Utilization, Saturation, Errors |