The Partner Whose Retry Loop Took Down Everyone Else
The Partner Whose Retry Loop Took Down Everyone Else
Your public partner API sits behind 12 stateless gateway instances behind a load balancer, and every partner is contractually capped at 1,000 requests/minute per API key. One integration partner shipped a bug: their client retries every failed request immediately with no backoff, and when your API had a brief 30-second blip, their retry storm hit 40,000 requests/minute. Because the current rate limiter increments an in-memory counter local to whichever gateway instance handles each request, each instance enforced the 1,000 req/min cap only on the requests it happened to see (about 3,333/min each), so across 12 instances the key got up to 12,000 req/min through, 12 times its cap -- and in aggregate that partner consumed enough capacity to push p99 latency for every other partner from 60ms to 900ms for 10 minutes. You're asked to design a rate limiter that actually enforces the 1,000 req/min cap per API key across all 12 instances, and to pick and defend a specific algorithm (token bucket, leaky bucket, fixed window, or sliding window) for this traffic shape.
The Partner Whose Retry Loop Took Down Everyone Else
Your public partner API sits behind 12 stateless gateway instances behind a load balancer, and every partner is contractually capped at 1,000 requests/minute per API key. One integration partner shipped a bug: their client retries every failed request immediately with no backoff, and when your API had a brief 30-second blip, their retry storm hit 40,000 requests/minute. Because the current rate limiter increments an in-memory counter local to whichever gateway instance handles each request, each instance enforced the 1,000 req/min cap only on the requests it happened to see (about 3,333/min each), so across 12 instances the key got up to 12,000 req/min through, 12 times its cap -- and in aggregate that partner consumed enough capacity to push p99 latency for every other partner from 60ms to 900ms for 10 minutes. You're asked to design a rate limiter that actually enforces the 1,000 req/min cap per API key across all 12 instances, and to pick and defend a specific algorithm (token bucket, leaky bucket, fixed window, or sliding window) for this traffic shape.
Provide 1–2 precise sentences for each architectural dimension. Each box guides you on what staff-level interviewers evaluate.
Define SLA targets, hard consistency constraints, and conditions the system must never violate.
Quantify throughput (QPS/RPS), read:write ratios, and peak burst multipliers.
Step-by-step path: client ingress → API gateway → queues → background workers → persistence.
Database engine, table schema, partition keys (PK/SK), and durability strategy.
What resource hits saturation first under 10x traffic? (CPU, disk IOPS, connection pools, network).
Worker crashes, network partitions, split-brain, poison pill DLQ, retries, and idempotency.
What did you sacrifice in exchange and why? (e.g. eventual consistency vs latency, cost vs redundancy).