Design Google Drive File Sync & Storage
1. Problem Statement & Scope
System Mission
Design an enterprise-grade, cross-platform cloud file storage and multi-device synchronization engine capable of managing over 500 Million registered users and billions of files. The architecture must feature Content-Defined Chunking (CDC) using rolling polynomial Rabin Fingerprints for block-level deduplication, differential delta synchronization (transferring only mutated 4MB chunks), optimistic concurrency control for file revision trees, and a sub-500ms real-time mutation fanout pipeline across 250 Million concurrent connected devices.
Synthesizing vector architecture diagram...
Functional Requirements
- Hierarchical File & Directory Namespace: Create, rename, move, copy, and delete files and deeply nested directories with ACID atomicity and POSIX permissions.
- Block-Level Differential Delta Synchronization: Split mutated files using Content-Defined Chunking (CDC, average chunks), compute cryptographic SHA-256 block hashes, and negotiate with the server to upload only novel, uncommitted blocks.
- Global Content Addressable Storage (CAS) Deduplication: Deduplicate identical blocks across all files and users globally, referencing shared binary blocks in S3 via reference counting.
- Real-Time Cross-Device Push Sync: Deliver instant remote file mutation notifications to all active desktop, mobile, and web sessions via persistent WebSocket connections within .
- Revision History & Conflict Resolution: Maintain a 30-day immutable revision tree with optimistic concurrency locking on monotonic revision IDs, automatically generating "Conflicted Copies" during concurrent offline edits.
Non-Functional Requirements (SLAs/SLOs)
- High Availability: availability for file metadata query and sync commit APIs ( downtime/year).
- Data Durability: (11 Nines) zero-byte-loss durability guaranteed by Amazon S3 multi-AZ storage and S3 Object Lock.
- Latency SLOs:
- Delta Commit & Negotiation Latency: (), ().
- Sync Push Fanout Latency: () from commit finalization to peer device WebSocket frame receipt.
- Bandwidth Reduction Efficiency: bandwidth reduction on incremental edits compared to full-file re-uploads.
- CAP / PACELC Classification:
- Metadata & Commit Store: CP / PC-EC (Consistent revision numbers, strict serializability via DynamoDB transactional writes or Aurora PostgreSQL).
- CAS Block Storage: AP / PA-EL (Immutable content-addressed S3 chunks, eventually consistent cache reads).
2. Capacity & Scale Estimation
Traffic Calculations
- User Population: total registered users; Daily Active Users (DAU).
- Active Connected Devices: Average .
- Daily File Modifications / Sync Commits:
- Average .
- Commit API QPS:
- WebSocket Notification Fanout Volume:
- Each commit fans out to an average of other online peer devices owned by the user.
Storage Calculations (3-Year Horizon)
- User Storage Quota & Raw Footprint:
- Free tier allocation ; average usage .
- Deduplication Multiplier (FastCDC Content-Defined Chunking):
- Cross-user block-level deduplication yields an estimated storage footprint reduction.
- Total Unique Chunks Managed:
- Average chunk size .
- Metadata Storage Growth (3 Years):
- .
- Record size (file pointer + 5 chunk hashes + timestamps): .
Network Bandwidth
- Ingress Bandwidth (With vs Without Delta Sync):
- Without Delta Sync (re-uploading the whole edited file; the average edited file is a 40 MB document or spreadsheet, i.e. 10 chunks, since tiny files are rarely the ones being re-saved):
- With Block-Level Delta Sync (only the 1 modified 4MB chunk of those 10 is uploaded per edit): That is the reduction promised in the NFRs; for micro-edits of 100KB inside 100MB documents, CDC achieves . Note that delta sync buys nothing for files smaller than one chunk (): those are always re-uploaded whole, which is why the average is computed over edited files, not all files.
- Egress Bandwidth (Peer Download Path):
Memory & Cache Sizing (80/20 Pareto Working Set)
- Top of active daily files generate of read/sync queries.
- Daily active chunk working set: .
- Global Chunk Hash Cache (ElastiCache Redis):
- Redis Set mapping
sha256 -> exists(32 bytes hash + 16 bytes pointer ):
- Redis Set mapping
- Global Bloom Filter for Chunk Existence (, false positive):
A RedisBloom filter is a single key and therefore lives on a single shard, so one 638 GB filter is impossible; the global filter is partitioned into 256 sub-filters keyed by the first byte of the SHA-256 (
bf:blocks:<xx>, each), which cluster mode spreads across shards. Provision 32 nodes ofcache.r6g.2xlarge( each, total) so that filter growth, replication buffers and fragmentation have headroom; 20 nodes would be full on day one.
Fleet Sizing & Connection Gateway Provisioning
- WebSocket Gateway Fleet:
- Managing persistent WebSocket connections.
- Linux epoll/Netty containers on AWS ECS Fargate: 50,000 idle TCP connections per container (requiring ). Distributed across 3 Availability Zones behind multi-AZ Network Load Balancers (NLB).
3. AWS-First High-Level Architecture
Synthesizing vector architecture diagram...
Follow a file edit from device A to device B. Client A splits the changed file into content-defined chunks and asks the commit API which chunks are new. In the "Metadata Engine & CAS Storage Tier" panel, the sync service checks a Bloom filter of existing chunk hashes (most chunks already exist, even across users), uploads only new chunks to S3 under their SHA-256 hash, and records the new file version in DynamoDB as a list of chunk hashes. The "Real-Time Mutation Fanout Pipeline" panel carries that change from DynamoDB Streams through Kinesis to push workers, which look up device B's connection in Redis and notify it over WebSocket; device B then downloads only the chunks it lacks. In the "Garbage Collection & Archival" panel, chunks whose reference count reaches zero are deleted by a GC worker. Content addressing is the key idea: identical chunks are stored once, and a one-line edit transfers one chunk instead of the whole file.
Data Flow Walkthrough
- Local Chunking & Hashing: When a local file mutates, Client A's OS filesystem watcher activates FastCDC, computing rolling Rabin fingerprints to cut chunks, hashing each chunk with SHA-256.
- Delta Negotiation: Client A invokes
POST /v1/files/sync/commitwith base version and chunk hashes. The Sync Service checks the Redis Bloom filters first: a negative is definitive (the chunk is new, upload it), but a positive is only "probably exists" and must be confirmed with aBatchGetItemon theBLOCK#<sha256>rows before the client is told to skip an upload, because acting on a false positive would commit a manifest that points at a chunk nobody ever stored. For chunks that do exist, the response also carries a proof-of-possession challenge (Section 8.6) so a client cannot attach someone else's chunk by knowing its hash. The service returns presigned S3 PUT URLs for only the genuinely missing chunks. - Immutable CAS Ingestion: Client A uploads only novel chunks directly to
s3://drive-blocks/<sha256_hash>. - Atomic Manifest Finalization: Client A invokes
POST /v1/files/sync/finalize. The Sync Service verifies chunk existence, increments monotonic version (), and commits via DynamoDBTransactWriteItems. - Real-Time Device Fanout: DynamoDB Streams captures the revision commit, routing through Kinesis to Fanout Workers. Workers query the Redis connection registry and push a lightweight
FileUpdatedevent through the WebSocket Gateway to Client B. Client B downloads only the modified 4MB block from S3.
Concrete Step-by-Step Request Walkthrough: Tracing Delta Sync Commit & Fanout
| Step # | Event / Action | Component State | Distributed Transition | Output / Response |
|---|---|---|---|---|
| 1 | User saves 100MB file on Client A; OS watcher triggers FastCDC | Local SQLite stores prior version 14; FastCDC cuts 25 chunks ( each) | FastCDC detects 1 modified chunk; computes 25 SHA-256 hashes () | 24 hashes match local cache; 1 novel hash () identified |
| 2 | Client A sends commit negotiation via POST /v1/files/sync/commit | Sync Service receives base_version=14 and hashes [H1..H25] | Pipelines BF.EXISTS across the Redis Bloom sub-filters; Bloom positives () confirmed via DynamoDB BatchGetItem on BLOCK# rows; is a Bloom negative, so it is definitively missing | Returns 200 OK with missing_chunks: [H17], an S3 presigned PUT URL, and possession challenges for the 24 existing chunks |
| 3 | Client A streams chunk () directly to Amazon S3 | S3 receives binary block at key s3://drive-cas-blocks/H17 | Multi-AZ S3 write committed; MD5 ETag verified by client | S3 returns 200 OK ETag to Client A |
| 4 | Client A invokes POST /v1/files/sync/finalize with proof of upload | Sync Service initiates atomic metadata transaction | DynamoDB executes TransactWriteItems: asserts current_version == 14, writes VER#000015 | Version 15 committed; 200 OK returned to Client A in |
| 5 | DynamoDB Stream emits INSERT event on VER#000015 | Stream captures file ID, user ID, new version, and chunk list | Kinesis partition key hash(user_id) preserves strict mutation ordering | Fanout worker reads Kinesis batch in |
| 6 | Fanout worker queries Redis for active device sessions | Redis returns connection ID for Client B on Gateway Task #42 | Worker publishes IPC notification to WebSocket Gateway Task #42 | Gateway sends WebSocket frame: FileUpdated(file_id, ver=15, [H1..H25]) |
| 7 | Client B receives frame; inspects local SQLite block cache | Local cache contains ; misses | Client B requests only from S3 CAS bucket via CloudFront | Reassembles file locally in ; user sees updated file seamlessly |
4. API Interface Design
1. Delta Sync Commit Negotiation (POST /v1/files/sync/commit)
Negotiates missing chunks with the server prior to data transfer.
httpPOST /v1/files/sync/commit HTTP/1.1 Host: api.drive.aws.internal Authorization: Bearer <jwt_token> Content-Type: application/json { "file_id": "file_88129384bc", "base_version": 14, "file_name": "quarterly_financial_model.xlsx", "parent_folder_id": "folder_root_001", "file_size_bytes": 20971520, "chunk_hashes": [ "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855", "f2ca1bb6c7e907d06dafe4687e579fce76b37e4e93b7605022da52e6ccc26fd2", "a591a6d40bf420404a011733cfb7b190d62c65bf0bcda32b57b277d9ad9f146e", "5e884898da28047151d0e56f8dc6292773603d0d6aabbdd62a11ef721d1542d8", "4b227777d4dd1fc61c6f884f48641d02b4d121d3fd328cb08b5531fcacdabf8a" ] }
Response: 200 OK(Upload Only Required For Missing Chunks)
json{ "file_id": "file_88129384bc", "new_version": 15, "sync_action": "UPLOAD_REQUIRED", "missing_chunks": [ { "chunk_hash": "a591a6d40bf420404a011733cfb7b190d62c65bf0bcda32b57b277d9ad9f146e", "presigned_put_url": "https://s3.us-east-1.amazonaws.com/drive-cas-blocks/a591a6d40bf420404a011733cfb7b190d62c65bf0bcda32b57b277d9ad9f146e?X-Amz-Signature=...", "expires_at": 1773651600 } ] }
2. Finalize Commit (POST /v1/files/sync/finalize)
Atomically promotes revision upon verified upload of missing chunks.
httpPOST /v1/files/sync/finalize HTTP/1.1 Host: api.drive.aws.internal Authorization: Bearer <jwt_token> Content-Type: application/json { "file_id": "file_88129384bc", "base_version": 14, "new_version": 15, "uploaded_chunk_hashes": [ "a591a6d40bf420404a011733cfb7b190d62c65bf0bcda32b57b277d9ad9f146e" ] }
Response: 200 OK
json{ "file_id": "file_88129384bc", "status": "COMMITTED", "version": 15, "committed_at": 1773648500123 }
3. Status Codes & Error Contracts
| HTTP Status | Reason Code | Error Contract Payload | Mitigation / Client Action |
|---|---|---|---|
200 OK | SUCCESS | Entity payload / missing chunk list | Normal completion |
201 Created | FILE_CREATED | New file descriptor metadata | Normal completion |
400 Bad Request | INVALID_CHUNK_HASH | {"error": "SHA-256 hash must be 64 hex characters"} | Client regenerates hash before retry |
404 Not Found | PARENT_FOLDER_NOT_FOUND | {"error": "Parent folder ID does not exist"} | Refresh directory tree from server |
409 Conflict | VERSION_CONFLICT | {"error": "Base version 14 is stale, server at 15"} | Branch into Conflicted Copy |
507 Insufficient Storage | STORAGE_QUOTA_EXCEEDED | {"error": "15GB quota exhausted", "used": 15024} | Alert user to clear space or upgrade (413 is about a single request body being too large, not the account's quota) |
429 Too Many Req | SYNC_RATE_LIMITED | {"error": "Commit frequency exceeded", "retry_after": 2} | Exponential backoff with jitter |
5. Data Models & Storage Architecture
Relational Schema (Amazon Aurora PostgreSQL Alternative)
sqlCREATE TABLE users ( user_id UUID PRIMARY KEY DEFAULT gen_random_uuid(), email VARCHAR(255) UNIQUE NOT NULL, storage_quota_bytes BIGINT NOT NULL DEFAULT 16106127360, -- 15 GB storage_used_bytes BIGINT NOT NULL DEFAULT 0, created_at TIMESTAMPTZ NOT NULL DEFAULT NOW() ); CREATE TABLE files ( file_id UUID PRIMARY KEY DEFAULT gen_random_uuid(), user_id UUID NOT NULL REFERENCES users(user_id) ON DELETE CASCADE, parent_folder_id UUID REFERENCES files(file_id), file_name VARCHAR(255) NOT NULL, is_folder BOOLEAN NOT NULL DEFAULT FALSE, is_deleted BOOLEAN NOT NULL DEFAULT FALSE, current_version INT NOT NULL DEFAULT 1, created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(), updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW(), CONSTRAINT unique_file_in_folder UNIQUE (user_id, parent_folder_id, file_name) ); CREATE TABLE file_versions ( file_id UUID NOT NULL REFERENCES files(file_id) ON DELETE CASCADE, version_number INT NOT NULL, file_size_bytes BIGINT NOT NULL, chunk_manifest JSONB NOT NULL, -- Array of SHA-256 hashes in sequence committed_at TIMESTAMPTZ NOT NULL DEFAULT NOW(), PRIMARY KEY (file_id, version_number) ); CREATE TABLE cas_blocks ( block_sha256 CHAR(64) PRIMARY KEY, size_bytes INT NOT NULL, global_ref_count BIGINT NOT NULL DEFAULT 1, s3_storage_tier VARCHAR(32) NOT NULL DEFAULT 'STANDARD', first_seen_at TIMESTAMPTZ NOT NULL DEFAULT NOW() ); CREATE INDEX idx_files_parent ON files(user_id, parent_folder_id) WHERE is_deleted = FALSE;
DynamoDB Single-Table Schema (GoogleDriveCoreTable)
Partition Key (PK) | Sort Key (SK) | Attributes & Payloads | GSI1-PK / GSI1-SK |
|---|---|---|---|
USER#<user_id> | FILE#<file_id> | name, parent_id, current_ver: 15, is_folder: false, updated_at | FOLDER#<parent_id> / NAME#<name> |
FILE#<file_id> | VER#000015 | size_bytes: 20971520, chunk_hashes: [H1..H5], committed_at: 1773648500 | — |
BLOCK#<sha256> | METADATA | size_bytes: 4194304, global_ref_count: 842, s3_key: "drive-cas-blocks/<sha256>" | — |
USER#<user_id> | DEVICE#<device_id> | conn_id: "WSS_CONN_9918", client_os: "MACOS", ttl: 1773734900 | — |
Unlock Complete Architecture & Production Runbooks
You have explored the free architectural preview (~39%). Spend 1 Coin to unlock the remaining 6 production deep-dive sections for a full 24 hours.