A workflow survives the process that starts it
Sending a receipt, charging a customer and provisioning a resource can span seconds to days. Holding an HTTP request or a database transaction for the whole duration is fragile. Store the workflow's state and let workers resume after failures.
A state machine names valid transitions. For a purchase: created -> reserved -> payment_pending -> confirmed, with explicit failed, expired and refund paths. Record why a transition happened, not only its current status. Treat an unknown payment response as pending reconciliation rather than declaring payment failed.
Minimum durable model
| Record | Key fields | Important guarantee |
|---|---|---|
| Workflow | tenant, run ID, state, version | Conditional valid transitions |
| Step attempt | run, step, attempt, lease epoch | One recognized owner for recorded completion |
| Timer | run, purpose, due time | Durable deadline intent |
| Outbox | event ID, destination, status | Publication survives worker restart |
| Inbox | consumer, event ID | Duplicate effects are suppressed atomically |
Keep business state transitions and the outbox in the same local transaction. Workers claim bounded work, use deadlines and renew leases as appropriate. A lease expiration makes another attempt possible; it does not undo the previous worker's external action.
Crash between an effect and its acknowledgment
Worker calls payment provider with stable operation key K
Provider charges once
Worker crashes before storing success
Next attempt queries/retries K
Workflow records the authoritative result
Reusing K matters; generating a fresh key for every attempt can create another charge. If the dependency lacks idempotency or result lookup, document reconciliation and the remaining risk. A workflow engine cannot give a non-idempotent external API magical exactly-once effects.
Cancellation and compensation
Cancellation asks the workflow to stop future work. It can race with a step already completing. Record the request and make each transition check whether cancellation is still meaningful. A paid booking may require a refund rather than a return to never-created state.
A compensation is a business operation with its own authorization, idempotency, failures and retries. Some effects cannot be undone. Preserve compensation_pending rather than hiding a failed refund behind a generic cancelled label. Define a manual-repair owner and an audit trail.
Durable timers and scheduling
Persist a due time and unique purpose, scan or index due work, and treat repeated dispatch as possible. Cron describes when to try; it does not guarantee one successful execution. For calendar schedules specify time zone, daylight-saving skips/repeats, overlap policy and missed-run behavior.
Long sleeps do not belong in a worker process when a durable timer can release the resource. Separate attempt timeout from business expiry. A reservation can expire in ten minutes while each API attempt has a three-second budget.
CQRS without event sourcing
CQRS separates command handling from query handling. Start with different code paths over the same database if that solves the problem. A separate read store adds synchronization and lag; it is justified when the read shape or scale needs it.
Return a write version or operation ID so a UI can show accepted work before the projection catches up. A projection missing a just-created order is not proof that creation failed. Permissions must be checked on reads even if a materialized view was built earlier.
Event sourcing: the history is authoritative
Event sourcing reconstructs business state from accepted events. It is different from saving the latest row and also writing an audit log. Use it when historical reconstruction and domain transitions justify maintaining a permanent replay contract.
For an order stream, ItemAdded, ItemRemoved and OrderConfirmed describe decisions. Enforce an expected stream version when appending: two commands that both read version 8 cannot independently accept contradictory version-9 decisions. After a conflict, reload and reevaluate the business rule.
An event broker distributes events; it is not automatically an event store with per-aggregate queries, expected-version appends and permanent retention. Never assume Kafka retention settings satisfy authoritative history requirements.
Snapshots and projections
A snapshot stores derived aggregate state at a known stream version. Load it and replay later events. Verify schema version and integrity; the event history remains authoritative. Snapshot corruption should be recoverable by replay, though that may be expensive.
A projection keeps an applied-event checkpoint. Updating derived state and advancing that checkpoint must be atomic or idempotently recoverable. Partition checkpoints correctly: one scalar offset cannot describe independent broker partitions. Handle an out-of-order event using the documented sequence contract rather than blindly accepting arrival order.
Replay without re-sending every receipt
Separate pure state reconstruction from side-effect delivery. A replay can rebuild the order list but must not charge every historical order again. Side-effect workers use their own durable deduplication records and explicit replay policy.
Build a new projection alongside the old one, replay retained history, catch up live events, compare outputs and switch readers. Budget the rebuild so it does not overload normal processing. If historical events contain sensitive payloads, deletion and retention requirements may complicate immutable history; design the data boundary before adoption.
Event evolution
Include an event type and schema version. Readers can supply safe defaults for additive changes or use an upcaster to map an old event into the current representation. Never silently reinterpret an old amount stored in paise as rupees.
Test representative historical streams against every new handler version. A new business rule cannot simply reject a formerly valid event during replay. Keep command validation for today's decisions separate from interpreting decisions already accepted.
Exercise
Build a report-generation workflow: accept request, snapshot inputs, queue computation, upload output, publish completion. Draw crashes after every external effect. Add cancellation while upload finishes, a duplicate timer and a projection rebuild. State why a small CRUD service might reasonably choose transactions plus an outbox instead of event sourcing.