Failure tolerance starts with a budget
Define acceptable errors, latency and data loss. Availability is the fraction of relevant requests or time meeting the agreed objective. Durability is about retaining data. A system can stay reachable while returning stale or incorrect answers, so your success definition matters.
Redundancy duplicates resources. Fault tolerance is the behavior that continues useful service when resources fail. Three replicas sharing one power supply or one broken deployment are vulnerable to correlated failure. Inspect the entire path for single points of failure, including identity, DNS, secrets, queues and deployment control planes.
Deadlines, timeouts and retries
Propagate an end-to-end deadline. Split its remaining budget across dependencies and queue waiting; avoid letting each nested call wait the whole original duration. A timeout only says the caller stopped waiting, not that the server cancelled its side effect.
Retry transient failures only when the operation is safe to repeat, with a bounded attempt count, exponential backoff and jitter. Avoid retries at every layer: three attempts at each of five layers can amplify one logical operation into 243 deepest-layer attempts. Place retry ownership deliberately.
Circuit breakers and bulkheads
A circuit breaker stops calls when failure evidence crosses its policy. Closed permits calls, open fails fast, and half-open admits limited probes after a waiting period. Scope breakers to sensible failure domains: one tenant's bad input should not disable an entire dependency.
Bulkheads isolate resource pools. Separate image processing from checkout workers so one saturated workload cannot consume all capacity. Bound connection pools and concurrency, reject excess requests with clear retry guidance and preserve capacity for recovery work. Circuit breakers do not replace timeouts or idempotency.
Backpressure, admission and load shedding
An unbounded queue hides overload until memory or latency fails. Measure arrival rate, service rate and queue age. When sustainable service rate is lower than arrivals, more buffering merely delays failure. Slow the producer, reject new work, reduce optional work or add actual processing capacity.
Rate limiting protects a policy boundary; backpressure responds to downstream capacity. Load shedding deliberately rejects lower-priority work. A fallback must be tested under the same failure that activates it: routing all traffic to an undersized backup can turn partial failure into total failure.
Graceful degradation
A product catalog may serve a bounded stale description while recommendations are disabled. Payments should not pretend success because the provider is unavailable. Define fallback freshness, visibility and correctness per feature. Separate control-plane changes from data-plane operation so service can continue through some management outages.
Use canaries, feature flags and rollback plans to limit bad releases. Rollback may be unsafe after an incompatible schema migration; prefer expand/migrate/contract changes. A healthy process metric is not evidence that a rollout preserves business outcomes.
Recovery Point and Recovery Time Objectives
RPO is the tolerated data-loss window; RTO is the targeted restoration duration. These are objectives, not guarantees obtained by purchasing a backup feature. A 15-minute RPO requires more than one nightly backup. RTO includes detection, decision, provisioning, restoration, verification and traffic cutover.
| Strategy | Typical shape | Main trade-off |
|---|---|---|
| Backup and restore | Restore into provisioned capacity | Lower idle cost, longer recovery |
| Pilot light | Essential data/core services remain | More recovery automation required |
| Warm standby | Reduced live capacity in alternate location | Scale-up and failover correctness |
| Active-active | Serve from multiple locations | Conflict, routing and operating complexity |
Replication is not a backup: it can copy accidental deletion or corruption. Use access-isolated, versioned/immutable backups where appropriate, verify keys remain available and practice restores. Include object storage, secrets, configuration, search rebuilds and event history in the inventory.
Worked outage drill
For an orders service, kill the primary DB during a write, then make one zone unavailable. Observe unknown client results, replica lag and failover. After recovery, reconcile payments against orders and verify no double application. Measure actual recovery duration and lost data rather than declaring success because the endpoint returned 200.
Walk through a breaker transition
As an illustrative policy, a client opens its breaker after five qualifying failures and waits thirty seconds before limited probes. While open, it rejects calls to that dependency quickly; healthy unrelated work should continue. Successful probes permit closing, while failing probes reopen the circuit. These numbers are an example to test, not universal production defaults. Decide which errors count, whether state is per instance and whether the fallback is actually safe.
Restore the business, not just a server
A backup is an input to recovery; disaster recovery includes access, capacity, dependencies, verification and traffic cutover. Maintain an incident owner, restoration sequence and communication plan. Cold, warm and hot recovery setups trade ongoing cost against recovery work and time.
Run a drill from a specified failure point: restore data, measure its age, validate critical transactions, reconcile external effects and only then reopen writes. Fault-tolerant placement should cross the failure domains you promise to survive. Multiple replicas help little if the same outage removes their zone, keys or shared storage.
Interview review
For every box in your diagram, ask: what fails, who notices, what retries, what is durable, who takes over and what happens during recovery? Include the failure of the monitoring and coordination systems themselves.
Next: observability and capacity and coordination.