The Warehouse System That Fell a Day Behind
The Warehouse System That Fell a Day Behind
Your order service currently calls the warehouse-allocation service synchronously over HTTP the moment a customer checks out: reserve inventory, pick a fulfillment center, return a confirmation. At normal volume (150 orders/min) this works, but the warehouse-allocation service occasionally takes 8-10 seconds under its own load (it does geo-distance calculations across 40 fulfillment centers), and during Black Friday, checkout traffic spiked to 2,200 orders/min while warehouse-allocation's throughput ceiling stayed around 400/min. The synchronous call chain meant checkout itself started timing out and returning errors to paying customers, even though the order data was perfectly valid ā the bottleneck was entirely downstream. You're asked to redesign the handoff between order placement and warehouse allocation so a slow or temporarily overloaded warehouse system never causes checkout itself to fail, describing exactly what decouples them and what happens to an order if allocation fails for that specific item.
The Warehouse System That Fell a Day Behind
Your order service currently calls the warehouse-allocation service synchronously over HTTP the moment a customer checks out: reserve inventory, pick a fulfillment center, return a confirmation. At normal volume (150 orders/min) this works, but the warehouse-allocation service occasionally takes 8-10 seconds under its own load (it does geo-distance calculations across 40 fulfillment centers), and during Black Friday, checkout traffic spiked to 2,200 orders/min while warehouse-allocation's throughput ceiling stayed around 400/min. The synchronous call chain meant checkout itself started timing out and returning errors to paying customers, even though the order data was perfectly valid ā the bottleneck was entirely downstream. You're asked to redesign the handoff between order placement and warehouse allocation so a slow or temporarily overloaded warehouse system never causes checkout itself to fail, describing exactly what decouples them and what happens to an order if allocation fails for that specific item.
Provide 1ā2 precise sentences for each architectural dimension. Each box guides you on what staff-level interviewers evaluate.
Define SLA targets, hard consistency constraints, and conditions the system must never violate.
Quantify throughput (QPS/RPS), read:write ratios, and peak burst multipliers.
Step-by-step path: client ingress ā API gateway ā queues ā background workers ā persistence.
Database engine, table schema, partition keys (PK/SK), and durability strategy.
What resource hits saturation first under 10x traffic? (CPU, disk IOPS, connection pools, network).
Worker crashes, network partitions, split-brain, poison pill DLQ, retries, and idempotency.
What did you sacrifice in exchange and why? (e.g. eventual consistency vs latency, cost vs redundancy).