Define the contract before the diagram
Design booking for a fixed number of seats. Browsing may be eventually consistent, but accepting two confirmed bookings for one seat is forbidden. A hold expires after a stated interval; payment results can be delayed. The system must reconcile accepted provider charges and show an honest pending state.
Assume an illustrative event has 20,000 seats and 100,000 users arriving in two minutes. Average arrival is about 833 users/second; bursts can be higher. If each user performs five requests, average request load is roughly 4,167/second during that period. The scarce resource is inventory, not a capacity target for selling 100,000 seats.
Minimal architecture
Client -> admission/API -> booking service -> transactional inventory DB
| local transaction
-> outbox -> payment/workflow workers
Payment provider -> verified inbox -> reconciliation -> booking DB
Catalog read model/cache -> browsing only
Start with one transactional ownership boundary for seat allocation. A cache improves browsing, but the displayed availability is not permission to sell. Introduce partitioning only when measured workload and transaction locality justify it.
Data model and access paths
| Table | Keys and fields | Invariant/index |
|---|---|---|
| Seats | event, seat, state, hold ID, expiry, version | Primary key event + seat |
| Holds | owner, hold ID, status, expiry | Index status + expiry for due work |
| Bookings | booking ID, owner, hold, payment status | Unique accepted hold |
| Operations | owner, operation key, request hash, result | Unique owner + operation key |
| Payment attempts | booking, provider operation, state | Unique provider operation |
| Inbox/outbox | event ID, status, next attempt | Unique event identity; due-work index |
A hold for several seats must reserve all requested seats or none. Lock/check them in a consistent order inside the transaction. A separate partial insert per seat can leave a user with half a group unless that outcome is the intended contract.
API contract
POST /holds accepts event and seat IDs with a stable operation key. Respond with accepted hold identity and expiry, or a conflict when any seat is unavailable. Repeating the same key and same payload returns the recorded outcome; a different payload with the same key is rejected.
POST /bookings references an owned valid hold. GET /bookings/id reports pending, confirmed, expired or refund-pending. A private booking is queried using both booking identity and verified owner. Do not let a caller select another account through a request field.
Return bounded retry guidance for admission limits. Queue admission is not a successful reservation and must not be shown as one.
Reserve atomically
For a single-seat teaching example, attempt a conditional update in a transaction:
UPDATE seats
SET state = 'held', hold_id = 'h17', expires_at = CURRENT_TIMESTAMP + INTERVAL '10 minutes',
version = version + 1
WHERE event_id = 'e1' AND seat_id = 's42' AND state = 'available'
RETURNING seat_id;
Zero rows means no reservation. The hold insert, operation result and notification intent belong to the same transaction. Real production code uses bound parameters and validates identities. Expired holds are reclaimed through the authoritative state protocol; a delayed cleanup job cannot justify selling an already-confirmed seat.
Payment is a durable workflow
Create a payment attempt with a stable provider idempotency key, then call outside the inventory transaction. Never hold seat row locks while waiting on a distant provider. A response timeout leaves the attempt unknown: query the provider or wait for a verified callback.
Persist callback event identity before acknowledging durable acceptance. Validate signature, provider account, booking mapping, amount and currency. A valid callback about somebody else's payment is not sufficient authorization to finalize this booking.
Race: payment success versus hold expiry
Both actors update the same authoritative booking/hold state under the chosen transaction protocol. Decide which transition wins based on the business contract. If expiry wins and inventory is reassigned, a late successful payment cannot simply mark the original booking confirmed.
Instead move to refund-pending or a documented alternate-resolution workflow. If finalization wins, cleanup must observe confirmed state and not release its seats. The rule must use durable state, not the order in which two HTTP handlers started.
Crash traces
| Failure point | Visible state | Recovery |
|---|---|---|
| Before hold transaction commits | No accepted hold | Retry same operation identity |
| After commit before HTTP reply | Hold exists, client uncertain | Retrieve recorded operation result |
| Charge accepted before worker records it | Payment pending | Reconcile same provider operation |
| Confirmation committed before email | Booking confirmed | Outbox retries email safely |
| Refund provider unavailable | Refund pending | Retry/reconcile and expose status |
Keep a reconciliation scan for old pending states with an operator escalation deadline. Do not use email delivery as proof of booking success. Do not mark a refund complete until authoritative evidence supports it.
Flash-sale overload
Bound admitted booking concurrency and separate browsing resources from mutations. A waiting room can smooth attempts, but explain queue fairness and token expiry. Bots, refresh storms and retries can make offered load much larger than user count.
Partition inventory by event when reservations normally stay within one event. A popular event remains a hot key domain; spreading random seats across databases can complicate group booking. Prefer measured contention controls and admission before an unnecessary distributed transaction.
Observability and security
Track reservation conflict ratio, lock wait, hold age, unknown payment age, reconciliation backlog and refund-pending age. Page on stuck accepted payments and invariant violations, not every expected sold-out response.
Protect payment credentials, redact provider payloads and audit administrative refunds. Test concurrent last-seat requests, repeated operation keys, forged callbacks, callback reordering and a restart after charge acceptance. Keep private bookings out of shared public caches.
Interview extension
Add a second region. Specify one authoritative owner per event or use a database with the required cross-region transaction guarantees. Draw how a stale-region attempt is fenced. Explain the latency/availability cost and why an eventually consistent availability counter alone cannot prevent overselling.
Continue to durable workflows and regional ownership.