Reliability · 4 / 4
Build a reliable job worker
Your challenge
A worker sends notifications from a queue. It crashes after sending but before acknowledging. Explain what happens and design safe processing.
Try it first. Write down your assumptions and explain your reasoning.
1.Expect redelivery
The queue may deliver the message again after its visibility timeout or lease expires. Acknowledging before sending risks losing work. Acknowledging after sending permits duplicates. The application must choose and document how it handles this boundary.
2.Deduplicate meaningful effects
Use a stable business operation id and persist delivery state. If the provider accepts an idempotency key, pass it through. A local sent flag cannot atomically cover an external provider call; describe this remaining ambiguity rather than claiming exactly once.
3.Bound retries and observe failures
Retry transient errors with exponential backoff and jitter. Send exhausted or permanent failures to a dead-letter path with a reason. Record attempts, age and outcome without logging private message bodies. Size visibility timeouts to processing duration and renew leases for long work. Provide an operator replay process that preserves the original operation id.
Take it one step further
- 1.What makes an error retryable?
- 2.How do you avoid duplicate effects when replaying dead-letter messages?
Self-review
Can you explain each point without looking at the solution?
- Correct acknowledgement boundary
- Stable deduplication identity
- Bounded retries and operational visibility
