One Customer, One Shard, One Outage
One Customer, One Shard, One Outage
You run a B2B SaaS HR platform, multi-tenant, currently a single Postgres database holding all 4,000 customer companies. Total data is 8 TB and growing 15% a quarter; write volume is dominated by timesheet and payroll-run inserts, about 6,000 writes/s at peak (month-end payroll runs for the largest customers). One customer ā a 50,000-employee logistics company ā accounts for 22% of all rows and a disproportionate share of writes during their payroll window, and their query latency has degraded from 80ms to over 1.5s p99 as the table has grown, while a hundred small customers with under 50 employees each see no problem at all. Leadership wants to shard the database before the next big customer signs, and wants a plan for the partition key, how many shards to start with, and what happens to the one customer whose data alone could saturate an entire shard.
One Customer, One Shard, One Outage
You run a B2B SaaS HR platform, multi-tenant, currently a single Postgres database holding all 4,000 customer companies. Total data is 8 TB and growing 15% a quarter; write volume is dominated by timesheet and payroll-run inserts, about 6,000 writes/s at peak (month-end payroll runs for the largest customers). One customer ā a 50,000-employee logistics company ā accounts for 22% of all rows and a disproportionate share of writes during their payroll window, and their query latency has degraded from 80ms to over 1.5s p99 as the table has grown, while a hundred small customers with under 50 employees each see no problem at all. Leadership wants to shard the database before the next big customer signs, and wants a plan for the partition key, how many shards to start with, and what happens to the one customer whose data alone could saturate an entire shard.
Provide 1ā2 precise sentences for each architectural dimension. Each box guides you on what staff-level interviewers evaluate.
Define SLA targets, hard consistency constraints, and conditions the system must never violate.
Quantify throughput (QPS/RPS), read:write ratios, and peak burst multipliers.
Step-by-step path: client ingress ā API gateway ā queues ā background workers ā persistence.
Database engine, table schema, partition keys (PK/SK), and durability strategy.
What resource hits saturation first under 10x traffic? (CPU, disk IOPS, connection pools, network).
Worker crashes, network partitions, split-brain, poison pill DLQ, retries, and idempotency.
What did you sacrifice in exchange and why? (e.g. eventual consistency vs latency, cost vs redundancy).