The Booking That Existed in Frankfurt but Not in Virginia
The Booking That Existed in Frankfurt but Not in Virginia
Your airline-seat-booking platform runs active-active across three AWS regions (us-east-1, eu-central-1, ap-southeast-1) for low-latency reads and regional failover, with each region's database asynchronously replicating changes to the other two, typically within 200-400ms. A support ticket describes exactly the bug you feared: a customer books the last available seat on a flight via the eu-central-1 (Frankfurt) region, gets a confirmation, then immediately opens the same booking flow on their phone, which happens to route to us-east-1 (Virginia) because of a mobile carrier's DNS behavior -- and Virginia still shows the seat as available, because Frankfurt's write hadn't replicated there yet. A second customer in that 300ms window in Virginia also tries to book the same seat. You're asked to explain precisely why this happened given the async-replication design, and what you'd change so double-booking becomes structurally impossible for the "one seat, one owner" case specifically, without giving up multi-region active-active for everything else.
The Booking That Existed in Frankfurt but Not in Virginia
Your airline-seat-booking platform runs active-active across three AWS regions (us-east-1, eu-central-1, ap-southeast-1) for low-latency reads and regional failover, with each region's database asynchronously replicating changes to the other two, typically within 200-400ms. A support ticket describes exactly the bug you feared: a customer books the last available seat on a flight via the eu-central-1 (Frankfurt) region, gets a confirmation, then immediately opens the same booking flow on their phone, which happens to route to us-east-1 (Virginia) because of a mobile carrier's DNS behavior -- and Virginia still shows the seat as available, because Frankfurt's write hadn't replicated there yet. A second customer in that 300ms window in Virginia also tries to book the same seat. You're asked to explain precisely why this happened given the async-replication design, and what you'd change so double-booking becomes structurally impossible for the "one seat, one owner" case specifically, without giving up multi-region active-active for everything else.
Provide 1ā2 precise sentences for each architectural dimension. Each box guides you on what staff-level interviewers evaluate.
Define SLA targets, hard consistency constraints, and conditions the system must never violate.
Quantify throughput (QPS/RPS), read:write ratios, and peak burst multipliers.
Step-by-step path: client ingress ā API gateway ā queues ā background workers ā persistence.
Database engine, table schema, partition keys (PK/SK), and durability strategy.
What resource hits saturation first under 10x traffic? (CPU, disk IOPS, connection pools, network).
Worker crashes, network partitions, split-brain, poison pill DLQ, retries, and idempotency.
What did you sacrifice in exchange and why? (e.g. eventual consistency vs latency, cost vs redundancy).