The Product Page That Melted Redis
The Product Page That Melted Redis
Your e-commerce product-detail service reads a single row from Postgres per request: price, stock count, and description for one SKU. Traffic is normally 3,000 req/s, spread across 500,000 SKUs, and a single r6g.xlarge Postgres read replica handles it fine at that fan-out. A flash sale drops the price on one SKU and a marketing email goes out to 2 million subscribers; for the next 20 minutes, 80% of traffic ā roughly 12,000 req/s ā hits that one SKU's page. Postgres connections saturate, p99 latency goes from 40ms to 6 seconds, and the on-call engineer pages you asking whether caching this endpoint in front of Postgres would have prevented the outage, and if so, exactly how you'd design the read path, the invalidation path when price or stock changes, and what happens to correctness if two updates land on the same key within the same second.
The Product Page That Melted Redis
Your e-commerce product-detail service reads a single row from Postgres per request: price, stock count, and description for one SKU. Traffic is normally 3,000 req/s, spread across 500,000 SKUs, and a single r6g.xlarge Postgres read replica handles it fine at that fan-out. A flash sale drops the price on one SKU and a marketing email goes out to 2 million subscribers; for the next 20 minutes, 80% of traffic ā roughly 12,000 req/s ā hits that one SKU's page. Postgres connections saturate, p99 latency goes from 40ms to 6 seconds, and the on-call engineer pages you asking whether caching this endpoint in front of Postgres would have prevented the outage, and if so, exactly how you'd design the read path, the invalidation path when price or stock changes, and what happens to correctness if two updates land on the same key within the same second.
Provide 1ā2 precise sentences for each architectural dimension. Each box guides you on what staff-level interviewers evaluate.
Define SLA targets, hard consistency constraints, and conditions the system must never violate.
Quantify throughput (QPS/RPS), read:write ratios, and peak burst multipliers.
Step-by-step path: client ingress ā API gateway ā queues ā background workers ā persistence.
Database engine, table schema, partition keys (PK/SK), and durability strategy.
What resource hits saturation first under 10x traffic? (CPU, disk IOPS, connection pools, network).
Worker crashes, network partitions, split-brain, poison pill DLQ, retries, and idempotency.
What did you sacrifice in exchange and why? (e.g. eventual consistency vs latency, cost vs redundancy).