One Customer, One Shard, One Outage
You run a B2B SaaS HR platform, multi-tenant, currently a single Postgres database holding all 4,000 customer companies. Total data is 8 TB and growing 15% a quarter; write volume is dominated by timesheet and payroll-run inserts, about 6,000 writes/s at peak (month-end payroll runs for the largest customers). One customer — a 50,000-employee logistics company — accounts for 22% of all rows and a disproportionate share of writes during their payroll window, and their query latency has degraded from 80ms to over 1.5s p99 as the table has grown, while a hundred small customers with under 50 employees each see no problem at all. Leadership wants to shard the database before the next big customer signs, and wants a plan for the partition key, how many shards to start with, and what happens to the one customer whose data alone could saturate an entire shard.