Kubernetes is the de facto standard for orchestrating containers at scale. You declare the desired state of your application in YAML, and Kubernetes continuously reconciles the actual state to match: scheduling containers onto machines, restarting failed ones, rolling out new versions, and scaling under load. This guide covers the architecture, the core objects, and the operational patterns you need to run workloads confidently.
Cluster Architecture
A cluster has a control plane that makes global decisions and a set of worker nodes that run your workloads. The control plane is the brain; nodes are the muscle.
| Component | Role |
|---|---|
| kube-apiserver | Front door; all reads/writes go through the API |
| etcd | Consistent key-value store holding all cluster state |
| scheduler | Assigns Pods to suitable nodes |
| controller-manager | Runs reconciliation loops (e.g. keeps replicas at target) |
| kubelet | Agent on each node; starts and monitors containers |
| kube-proxy | Programs node networking rules for Services |
Declarative reconciliation
You never tell Kubernetes to "start three containers." You declare "I want three replicas," and controllers work continuously to make reality match. Delete a Pod and one is recreated; this self-healing is the core idea.
Pods
A Pod is the smallest deployable unit: one or more containers that share a network namespace and storage. You rarely create Pods directly; higher-level controllers manage them for you. But understanding the Pod spec is essential.
apiVersion: v1
kind: Pod
metadata:
name: web
labels:
app: web
spec:
containers:
- name: nginx
image: nginx:1.27
ports:
- containerPort: 80
Deployments and ReplicaSets
A Deployment manages a ReplicaSet, which in turn keeps a target number of identical Pods running. Deployments add rolling updates and rollbacks on top. This is the workhorse object for stateless services.
apiVersion: apps/v1
kind: Deployment
metadata:
name: web
spec:
replicas: 3
selector:
matchLabels:
app: web
template:
metadata:
labels:
app: web
spec:
containers:
- name: web
image: ghcr.io/acme/web:1.4.0
ports:
- containerPort: 8080
Services
Pods are ephemeral and their IPs change. A Service gives a stable virtual IP and DNS name that load-balances across the Pods matching its selector.
| Type | Reachable from | Use |
|---|---|---|
| ClusterIP | Inside the cluster only | Internal service-to-service |
| NodePort | Any node IP on a high port | Basic external access, dev |
| LoadBalancer | Public via cloud LB | Production external services |
apiVersion: v1
kind: Service
metadata:
name: web
spec:
type: ClusterIP
selector:
app: web
ports:
- port: 80 # service port
targetPort: 8080 # container port
Ingress
An Ingress routes external HTTP/HTTPS traffic to Services based on host and path, giving you one entry point with TLS termination instead of one LoadBalancer per service. It requires an ingress controller (nginx, Traefik, or a cloud gateway) running in the cluster.
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: web
annotations:
cert-manager.io/cluster-issuer: letsencrypt
spec:
ingressClassName: nginx
tls:
- hosts: ["app.example.com"]
secretName: web-tls
rules:
- host: app.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service:
name: web
port:
number: 80
ConfigMaps and Secrets
Keep configuration out of images. ConfigMaps hold non-sensitive settings; Secrets hold sensitive values (base64-encoded, and ideally encrypted at rest). Both can be injected as environment variables or mounted as files.
apiVersion: v1
kind: ConfigMap
metadata:
name: app-config
data:
LOG_LEVEL: "info"
FEATURE_FLAG: "true"
---
apiVersion: v1
kind: Secret
metadata:
name: app-secret
type: Opaque
stringData:
DB_PASSWORD: "s3cr3t" # stringData is auto base64-encoded
# inject into a container
envFrom:
- configMapRef:
name: app-config
env:
- name: DB_PASSWORD
valueFrom:
secretKeyRef:
name: app-secret
key: DB_PASSWORD
Namespaces
Namespaces partition a cluster into virtual sub-clusters for isolating teams or environments. They scope names, and combine with resource quotas and RBAC for governance.
kubectl create namespace staging
kubectl get pods -n staging
kubectl config set-context --current --namespace=staging # set default
Essential kubectl
kubectl apply -f deploy.yaml # create/update from manifest
kubectl get pods -o wide # list pods with node/IP
kubectl describe pod web-xyz # events + full spec (debug start)
kubectl logs -f web-xyz # follow logs
kubectl logs web-xyz --previous # logs from a crashed container
kubectl exec -it web-xyz -- sh # shell into a container
kubectl port-forward svc/web 8080:80 # local access to a service
kubectl get events --sort-by=.lastTimestamp
kubectl delete -f deploy.yaml
Rolling Updates and Rollbacks
Deployments update Pods gradually so there is no downtime. maxUnavailable and maxSurge control the pace. If a rollout goes bad, roll back to the previous revision.
kubectl set image deployment/web web=ghcr.io/acme/web:1.5.0
kubectl rollout status deployment/web # watch progress
kubectl rollout history deployment/web # list revisions
kubectl rollout undo deployment/web # revert to previous
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0 # never drop below desired replicas
maxSurge: 1 # add at most one extra pod at a time
Health Probes
Probes let the kubelet know a container's true state. A liveness probe failing restarts the container; a readiness probe failing removes the Pod from Service load balancing until it recovers; a startup probe protects slow-starting apps.
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 10
periodSeconds: 10
readinessProbe:
httpGet:
path: /ready
port: 8080
periodSeconds: 5
Liveness vs readiness
Readiness gates traffic; liveness restarts. Do not point liveness at a dependency like a database, or a transient outage will trigger endless restart loops. Keep liveness checks cheap and self-contained.
Resource Requests and Limits
Requests are what the scheduler reserves for a Pod; limits are the hard ceiling. Exceeding a memory limit causes an OOMKill; exceeding a CPU limit throttles the container. Setting sensible values keeps nodes stable and scheduling predictable.
resources:
requests:
cpu: "250m" # 0.25 vCPU reserved
memory: "256Mi"
limits:
cpu: "500m"
memory: "512Mi"
Horizontal Pod Autoscaling
The HorizontalPodAutoscaler (HPA) adds or removes replicas based on observed metrics like CPU. It requires a metrics server and CPU/memory requests set on the Pods.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
kubectl autoscale deployment web --cpu-percent=70 --min=3 --max=20
kubectl get hpa # see current vs target and replica count
Practice Exercises
- Write a Deployment manifest with 3 replicas of an nginx image, apply it, and confirm that deleting one Pod causes the ReplicaSet to recreate it.
- Expose that Deployment with a ClusterIP Service and reach it from another Pod using its DNS name, then use
port-forwardto hit it locally. - Create a ConfigMap and a Secret, inject both into a container as environment variables, and verify the values with
kubectl exec. - Add liveness and readiness probes to a Deployment, then intentionally break the readiness endpoint and observe the Pod being removed from the Service endpoints.
- Perform a rolling update to a new image tag, watch it with
rollout status, then roll it back withrollout undo. - Set CPU requests and limits, create an HPA targeting 70% CPU, generate load, and observe the replica count scale up and back down.