Skip to main content
BLUEPRINT #05Production Case Studies

Design Figma's 100x PostgreSQL Multi-Tenant Scaling Architecture

Target AWS Architecture:Aurora
Referenced Architecture Primitives (3)
Click any primitive to study its algorithmic deep dive
10-Stage Structure:1. Requirements→2. Sizing→3. Topology→4. Data Model→5. AWS Topology→6. Deep-Dive→7. Failures→8. SRE Playbooks

1. Problem Statement & Scope Clarification

System Mission

Design Figma's collaborative cloud database and multi-tenant scaling infrastructure. Figma's multiplayer canvas operates over real-time WebSockets, persisting complex vector document graphs, version trees, user permissions, and comment threads. As Figma scaled 100Γ—100\times, its single monolithic AWS database experienced severe CPU exhaustion, table lock contention, and replication lag. The platform must scale to hundreds of sharded nodes with zero downtime, zero data corruption, and seamless cross-shard routing.

Functional Requirements

  1. Multi-Tenant Document Persistence: Store hierarchical canvas document trees, node modifications, and organization metadata.
  2. Dynamic Shard Query Routing: Route SQL queries transparently to the appropriate database shard based on org_id or file_key.
  3. Multiplayer Canvas State Sync: Persist real-time CRDT vector operations without blocking on slow OLTP queries.
  4. Live Zero-Downtime Shard Splitting: Split and migrate heavily loaded shards without dropping user WebSocket connections or causing canvas write errors.

Non-Functional Requirements (SLAs & SLOs)

  • High Availability: 99.999%99.999\% uptime for canvas saving and file loading.
  • Low Query Latency: P50<1.0Β msP50 < 1.0\text{ ms}, P99<8.0Β msP99 < 8.0\text{ ms} on sharded queries.
  • Connection Scalability: Support 100,000+100,000+ concurrent application connections via PgBouncer connection multiplexing.

2. Capacity & Scale Estimation (Back-of-the-Envelope Math)

Scale Metrics

  • Active Collaborative Design Files: 50,000,000+50,000,000+.
  • Canvas Vector Nodes per File: 10,000βˆ’1,000,000Β nodes10,000 - 1,000,000\text{ nodes}.
  • Peak Query Throughput: 100,000Β QPS100,000\text{ QPS} against the database tier.
  • Single Limits: A maxed-out db.r6g.16xlarge instance (64 vCPUs, 512 GiB RAM) saturates at β‰ˆ15,000βˆ’20,000Β complexΒ QPS\approx 15,000 - 20,000\text{ complex QPS}.
  • Sharding Requirement: MinimumΒ RequiredΒ Shards=100,000Β QPS10,000Β safeΒ QPS/nodeβ‰ˆ10Β -Β 20Β ShardΒ Clusters\text{Minimum Required Shards} = \frac{100,000\text{ QPS}}{10,000\text{ safe QPS/node}} \approx \mathbf{10\text{ - } 20\text{ Shard Clusters}}

3. High-Level Architecture & Component Mapping

Interactive Architecture Diagram
Synthesizing vector architecture diagram...

Part 2: Production Deep-Dive Locked1 Coin = 24 Hours

Unlock Complete Architecture & Production Runbooks

Your Balance:40 Coins

You have explored the free architectural preview (~48%). Spend 1 Coin to unlock the remaining 5 production deep-dive sections for a full 24 hours.

Sections Included in This 24-Hour Pass:
4. Deep-Dive: Figma's 100x Database Evolution
5. Data Model & Shard Routing Architecture
6. Live Zero-Downtime Shard Migration
7. Comprehensive Trade-off Matrix
8. Real-World Engineering Failure Modes & Post-Mortem Lessons
Keeps page unlocked for exactly 24 hoursSpend coins to fund LLM & compute infrastructure