Design a File Sync and Storage Service (Google Drive)
This page is one interview loop in three rounds. All three rounds design the same system. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | A startup's desktop and mobile sync app | A consumer service like Google Drive or Dropbox | The enterprise edition: shared drives, admin controls, compliance |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Volume | 100K users, 200K devices; 1M commits/day (11.6/s, 35/s at peak); about 600 TB stored | 500M users, 100M daily; 250M connected devices; 1B commits/day (11,574/s, 34.7K/s at peak); 7.5 EB logical, about 2.14 EB after dedup | 20M seats in 20,000 organizations; 180M commits/day; 1 EB logical, 625 PB stored |
| Footprint | 1 region, 3 AZs | 1 region, 3 AZs, plus a CDN | 4 regions: a home and a standby region in each of 2 jurisdictions (US, EU) |
| Targets | A change shows up on my other devices within 1 minute; never lose a saved file; 99.9% | Commit P95 < 100 ms; push P95 < 500 ms; ≥ 90% fewer bytes on edits; 99.99% | Tenant isolation; data stays in its jurisdiction; roll a whole drive back to a point in time; survive a region loss |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
Loop Opener: What Is File Sync?
You Already Know One: the Same Folder on Every Device
You save a spreadsheet on your laptop at the office. On the train home you open your phone, and the new version is already there. You edit a slide on the plane with no network; when you land, the laptop quietly sends the change up, and your desktop at home picks it up. It feels like one folder that lives on every device at once.
A file sync service makes that illusion. Each device keeps a real copy of the folder on its own disk. A small program on each device (the sync client) watches for local changes and sends them to the server, and it asks the server for changes made elsewhere and applies them locally. The server is the meeting point: it holds the authoritative list of files and a copy of every file's bytes.
A few words we'll use all page:
| Word | What it means on this page |
|---|---|
| Chunk | A piece of a file, a few megabytes long. Files are uploaded, stored and downloaded as chunks, never as one big blob. |
| Hash | A short fingerprint of some bytes. We use SHA-256: 32 bytes that change completely if even one input bit changes, and no one knows how to find two inputs with the same hash. |
| Content-addressed | Stored under the hash of its own contents. The name is the fingerprint, so a chunk's name tells you exactly which bytes it holds, and the same bytes always get the same name. |
| Manifest | The ordered list of chunk hashes that make up one version of a file. Rebuilding the file means fetching those chunks and joining them in order. |
| Revision | One saved version of a file: a revision number, its manifest, its size, who saved it and when. |
| Namespace | A tree of folders and files with its own change log. In Round 1 it's one user's whole drive; in Round 2 each shared folder gets its own. |
| Cursor | A device's bookmark in a namespace's change log: "I have applied everything up to change number 48,213." |
What Makes It Hard
- Files are big and change a little at a time. A 2 GB video project or a 40 MB spreadsheet gets saved dozens of times a day, and each save changes a few kilobytes. Sending the whole file every time wastes the user's upload link and our bandwidth.
- Devices go offline. A laptop sleeps for a weekend, a phone loses signal in a tunnel. When they come back, they must catch up exactly, without re-reading everything.
- Two people (or two devices) edit the same file. Both were offline, both saved, both come online. We must never silently throw one of those edits away.
- The file tree is metadata that must be exactly right. Renames, moves and deletes must never produce two files with the same name in one folder, a folder inside itself, or a file that points to bytes we deleted.
The Question the Whole Loop Answers
How do we keep every device's copy in agreement, send the fewest bytes, and never silently lose someone's edit?
The answer grows every round:
- Round 1: split the namespace (a database) from the bytes (content-addressed chunks in S3), give every namespace a change log with a cursor, make commits conditional on the base revision, and save a conflicted copy instead of overwriting.
- Round 2: cut chunks by content so an insert only changes nearby chunks, store each chunk once across all users (with proof of possession), push "something changed" nudges over WebSockets, serve immutable chunks from a CDN, and garbage-collect chunks safely.
- Round 3: add organizations: inherited permissions, customer-held keys (which limit dedup), legal holds, data residency, ransomware rollback and audit logs.
How the bytes are stored underneath (replication, erasure coding, durability math) is the subject of the S3-like object storage loop. Here we use Amazon S3 as the byte store and link there instead of re-teaching it.
Round 1 · Mid-level · "Sync for a 100K-User Startup"
~35 min · SDE II (L5) · 1 region, 3 AZs · 100K users, 200K devices · 1M commits/day, 35/s at peak · about 600 TB stored · changes visible within 1 min · 99.9%
R1.1 Establish Design Scope
The interviewer says: "We're a startup building a Dropbox-style app: a folder on your desktop and an app on your phone that stay in sync. Design the backend." Before drawing anything, we ask.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How many users, and how much do they store? | 100,000 users; about 5 GB each on average. | 500 TB of current files, plus version history (R1.7). One region is plenty. |
| What's the largest file? | Up to 10 GB. | A single upload can take an hour on a home connection. It must survive a dropped connection without starting over (step 1.1). |
| Which devices? | A desktop client and a mobile app; about 2 devices per user. | Two very different clients: the desktop runs all day; the phone runs only while the app is open. |
| Sharing between users? | Not yet. | Each user's drive is a private namespace. No cross-user anything this round. |
| Can people edit offline? | Yes. | Two devices can change the same file without seeing each other. We need a conflict rule (step 1.5). |
| Version history? | 30 days, including deleted files. | Old revisions and their chunks stay for 30 days (step 1.6). |
| How fast should a change appear on my other devices? | Within a minute. | Polling every 30 seconds is fast enough this round; push waits for Round 2. |
Out of scope for this round: sharing, cross-user dedup, real-time push, a web editor.
R1.2 Functional Requirements, Derived Step by Step
| Phrase from the problem | Operation |
|---|---|
| "Put a file in the folder" | Upload a new file: send its chunks, then commit it (create revision 1) |
| "Edit the file" | Upload only the chunks that changed, then commit a new revision on top of a named base revision |
| "It shows up on my other devices" | changes(cursor): list everything that changed in my namespace since my bookmark |
| "Open a file on the phone" | Download its chunks by hash and join them |
| "Browse my folders" | List a folder's children |
| "Oops, restore yesterday's version" / "undelete" | Restore an old revision as a new revision; bring a file back from the trash |
Not yet: sharing, dedup across users, real-time push.
R1.3 Non-Functional Requirements: the Questions
We name each quality in words first; the numbers come in R1.7.
- Never lose data. Once the app says "saved to the cloud", the file must survive any single server, disk or AZ failure, and no concurrent edit may silently erase it. This is the correctness bar of the round.
- Bandwidth on slow links. Many users sit behind a 10 Mbps home upload link or a phone connection. A small edit must not cost a big upload.
- Eventual agreement. After everyone stops editing and every device is online, every device must end up with exactly the same tree and the same bytes.
- Availability. If the service is down, the client keeps working offline and catches up later, so we can be a little less strict than a payment system; 99.9% is the target.
R1.4 The API
Chunks travel directly between the client and S3 using pre-signed URLs: an S3 URL that our server signs with its credentials, valid for a short time and for exactly one operation on one key. Our API servers never touch file bytes; they only decide who may upload or download what.
1. Start an upload: "here are my chunks, which do you need?"
httpPOST /v1/uploads HTTP/1.1 Host: api.syncapp.example Authorization: Bearer <device token> Content-Type: application/json { "node_id": "n_8f21", "base_rev": 6, "size": 9437184, "chunks": ["3f9a...c1", "b07e...42", "77d1...9e"] }
The file is 9,437,184 bytes (9 MiB): two full 4 MiB chunks and a last chunk of 1 MiB. The hashes are SHA-256 in hex, shortened here. The server looks up which of this user's chunks it already stores:
json{ "missing": [ { "index": 2, "hash": "77d1...9e", "put_url": "https://syncapp-chunks.s3.us-east-1.amazonaws.com/u/u_1001/77d1...9e?X-Amz-Signature=...", "required_headers": { "x-amz-checksum-sha256": "d9FxT...=" }, "expires_at": "2026-09-28T09:15:00Z" } ] }
2. Upload each missing chunk straight to S3 with a PUT to its URL. The x-amz-checksum-sha256 header is one of the signed headers: the client must send exactly that value, and S3 computes the SHA-256 of the bytes it receives and rejects the upload if they don't match. So a chunk stored under the name 77d1...9e really does contain the bytes whose hash is 77d1...9e; a buggy or malicious client can't put the wrong bytes under a hash.
3. Commit: "make these chunks revision 7, on top of revision 6"
httpPOST /v1/commits HTTP/1.1 Authorization: Bearer <device token> Content-Type: application/json { "commit_id": "c_2b9d0e71", "node_id": "n_8f21", "base_rev": 6, "size": 9437184, "chunks": ["3f9a...c1", "b07e...42", "77d1...9e"], "client_mtime": "2026-09-28T09:02:11Z" }
json{ "node_id": "n_8f21", "rev": 7, "seq": 48214 }
commit_id is a random ID the client makes once per commit and reuses on every retry, so a retried commit can't create two revisions. client_mtime is stored for display only; it never decides anything (R1.9).
4. Catch up: "what changed since my bookmark?"
httpGET /v1/changes?cursor=48213&limit=500 HTTP/1.1 Authorization: Bearer <device token>
json{ "cursor": 48214, "has_more": false, "changes": [ { "seq": 48214, "op": "UPSERT", "node_id": "n_8f21", "parent_id": "n_docs", "name": "q3-plan.docx", "rev": 7, "size": 9437184, "chunks": ["3f9a...c1", "b07e...42", "77d1...9e"] } ] }
5. Download a chunk
httpGET /v1/files/n_8f21/revisions/7/chunks/2 HTTP/1.1 Authorization: Bearer <device token>
The server checks that this user can read that file, then answers 302 Found with a pre-signed S3 GET URL valid for 15 minutes. The client already knows the hash from the manifest, and checks the downloaded bytes against it.
6. Restore and delete
POST /v1/files/n_8f21/restorewith{ "rev": 5, "commit_id": "..." }creates revision 8 whose manifest is revision 5's. No bytes move.DELETE /v1/files/n_8f21?base_rev=7moves the file to the trash (a soft delete, step 1.6).
Status codes
| Code | Meaning |
|---|---|
200 OK | Done. A retried commit whose commit_id was already applied gets 200 with the same body. |
400 Bad Request | A malformed hash or size |
404 Not Found | No such file, or not yours |
409 Conflict | STALE_BASE: the file is past your base_rev, so your edit is a conflict (step 1.5); or MISSING_CHUNKS: some chunk you named isn't stored (upload it and commit again) |
410 Gone | Your cursor is older than the change log we keep: do a full resync (R1.9) |
429 Too Many Requests | A client over its rate limit |
507 Insufficient Storage | The account is over its quota |
Recap
- Bytes go straight between client and S3 by pre-signed URL; S3 checks each chunk's SHA-256.
- Upload is three steps: "which chunks do you need?",
PUTthe missing ones, then a commit that names its base revision and carries acommit_id. - Devices catch up with
changes(cursor);409means conflict or missing chunks;410means resync.
R1.5 Design Evolution: From Whole Files to Chunks, Revisions and a Change Log
Each step is a problem, your turn to think, the answer, and what it costs us.
Step 1.0: The Baseline
The client uploads the whole file to PUT /files/<path> on every save. The server writes it to S3 under the user's path, like u_1001/Documents/q3-plan.docx, overwriting the old object. Every minute, each device lists the user's whole tree and downloads whatever looks different.
It works for a demo. The weaknesses show up as soon as files get big, edits get frequent, and two devices disagree.
Step 1.1: A 2 GB Upload Fails at 90%
The problem: a user drops a 2 GB video into the folder. On a 10 Mbps upload link that's 2 × 8,000 Mb ÷ 10 Mbps = 1,600 seconds, almost 27 minutes. At minute 24 the Wi-Fi hiccups and the single upload dies. The client starts over from zero. What would you do?
Step 1.2: Where Do Files and Folders Live?
The problem: we now have chunks in S3 and manifests. We also have folders, file names, renames, moves, deletes and old versions. The baseline used the path as the S3 key (u_1001/Documents/q3-plan.docx). A user renames Documents to Docs, a folder with 3,000 files.
What would you do?
Step 1.3: Saving a Small Edit Re-uploads the Whole File
The problem: a user edits one cell in a 40 MB spreadsheet (10 chunks) and saves. The client uploads all 10 chunks again. The same user saves 20 times a day. What would you do?
Step 1.4: How Does My Phone Know What Changed?
The problem: the laptop commits revision 7 of q3-plan.docx. The desktop at home and the phone must find out. The user has 1,000 files in 40 folders.
What would you do?
Primitive: Change Data Capture and the Outbox Pattern
Step 1.5: Two Devices Edited the Same File Offline
The problem: Maria's laptop and desktop both have revision 6 of budget.xlsx. On a flight she edits it on the laptop. At home, her partner edits it on the desktop (same account). The desktop comes online first and commits revision 7. The laptop lands and tries to upload its version.
What would you do?
Step 1.6: I Deleted a File and Want It Back
The problem: a user deletes a folder by mistake, or overwrites a report with the wrong one. A day later they want it back. Meanwhile, if we keep every chunk forever, storage grows without end. What would you do?
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | Whole-file uploads under path keys; devices re-list everything | Huge uploads; slow scans |
| 1.1 | 2 GB upload fails at 90% | Fixed 4 MiB chunks, each a separate PUT; a manifest per revision | Chunk bookkeeping |
| 1.2 | Where files and folders live | Namespace in Aurora (nodes with parent IDs); chunks in S3 named by SHA-256; bytes first, metadata second | Two stores, a write order, orphans |
| 1.3 | Small edit re-uploads everything | Upload only chunks whose hash is new | Client hashing; inserts still shift chunks |
| 1.4 | How the phone learns of changes | Per-namespace journal with a cursor; 30 s polling; 30-day log | Mostly empty polls; up to 30 s delay |
| 1.5 | Offline edits on two devices | Commits name a base revision; stale base → conflicted copy | Manual reconciling |
| 1.6 | Undelete and restore | Soft delete, 30-day history, exact ref counts, a locked sweeper | 30 days of retained bytes |
R1.6 Architecture v1
Synthesizing vector architecture diagram...
File bytes never pass through our servers: clients move them to and from S3 with pre-signed URLs. Everything the API does is a small metadata transaction in Aurora.
Tables
sqlCREATE TABLE namespaces ( ns_id BIGINT PRIMARY KEY, -- one per user this round owner_id BIGINT NOT NULL, seq BIGINT NOT NULL DEFAULT 0, -- last journal number used min_seq BIGINT NOT NULL DEFAULT 0, -- oldest journal row still kept used_bytes BIGINT NOT NULL DEFAULT 0 ); CREATE TABLE nodes ( node_id BIGINT PRIMARY KEY, ns_id BIGINT NOT NULL REFERENCES namespaces, parent_id BIGINT, -- NULL for the root folder name TEXT NOT NULL, is_folder BOOLEAN NOT NULL, rev INT NOT NULL DEFAULT 0, -- current revision (files) deleted_at TIMESTAMPTZ, UNIQUE (ns_id, parent_id, name) -- two live files can't share a name (see note) ); CREATE TABLE revisions ( node_id BIGINT NOT NULL, rev INT NOT NULL, size_bytes BIGINT NOT NULL, chunks BYTEA[] NOT NULL, -- 32-byte SHA-256 hashes, in order commit_id UUID NOT NULL UNIQUE, -- idempotency: one revision per commit_id device_id BIGINT NOT NULL, client_mtime TIMESTAMPTZ, -- display only committed_at TIMESTAMPTZ NOT NULL DEFAULT now(), -- server clock PRIMARY KEY (node_id, rev) ); CREATE TABLE journal ( ns_id BIGINT NOT NULL, seq BIGINT NOT NULL, node_id BIGINT NOT NULL, op TEXT NOT NULL, -- UPSERT, MOVE, DELETE, RESTORE rev INT, PRIMARY KEY (ns_id, seq) ); CREATE TABLE chunks ( owner_id BIGINT NOT NULL, hash BYTEA NOT NULL, -- SHA-256 size_bytes INT NOT NULL, ref_count INT NOT NULL, zero_since TIMESTAMPTZ, -- set when ref_count reaches 0 PRIMARY KEY (owner_id, hash) );
A note on the unique name constraint: deleted_at rows would block reusing a name, so in practice the constraint is a partial unique index over rows WHERE deleted_at IS NULL, and PostgreSQL treats NULL parents as distinct, so the root folder's children use the root's node ID as parent_id rather than NULL.
The commit transaction (READ COMMITTED isolation, one short transaction):
UPDATE namespaces SET seq = seq + 1 WHERE ns_id = $ns RETURNING seq— this takes the namespace row lock first, so all commits and sweeps in one namespace run one at a time, and gives us the journal number.- If a revision with this
commit_idalready exists, this is a retry: roll back and return the stored result. SELECT rev FROM nodes WHERE node_id = $n FOR UPDATE; ifrev <> base_rev, roll back and return409 STALE_BASE.- Check that every chunk in the manifest has a
chunksrow, creating rows for the chunks this commit uploaded after aHEADon S3 confirms each exists (S3 reads are strongly consistent after a write). A missing one →409 MISSING_CHUNKS. - Insert the revision, set
nodes.rev = base_rev + 1, insert the journal row with the newseq, incrementref_countfor the manifest's chunks, updateused_bytes. - Commit.
Tracing an edit on the laptop appearing on the desktop
Synthesizing vector architecture diagram...
Only 1 MiB moved in each direction. The desktop found the change on its next 30-second poll, 17 seconds after the commit here.
Tracing a conflict (numbered steps; both devices started from revision 6)
- 18:40:05: the home desktop commits
budget.xlsxwithbase_rev = 6. The transaction findsrev = 6, writes revision 7, journalseq50,112. - 18:55:30: the laptop comes back online and runs
POST /v1/uploadswithbase_rev = 6; the server doesn't check the base here, it only says which chunks are missing. The laptop uploads 2 chunks. - 18:55:33: the laptop's commit with
base_rev = 6findsrev = 7. The transaction rolls back:409 STALE_BASE, current_rev 7. - The laptop renames its local copy to
budget (Maria's conflicted copy 2026-09-28).xlsxand commits it as a new node (revision 1,seq50,113). Its 2 new chunks are already uploaded, so this commit moves no bytes. - It pulls
changes(50111): revision 7 ofbudget.xlsx(seq50,112) and its own new file (seq50,113). It downloads revision 7's changed chunks and writesbudget.xlsx. - The desktop's next poll picks up
seq50,113 and downloads the conflicted copy. Both devices now show both files.
R1.7 Numbers
Targets
| Quality | Target | Why this number |
|---|---|---|
| Availability | 99.9% | 0.1% of a 30.4-day month: 43,776 min × 0.001 ≈ 44 minutes. The client works offline through an outage. |
| Propagation | A change visible on the user's other desktop within 1 minute | The interviewer's number; the chain below shows about 37 s worst case for a small edit. |
| Durability | No committed file lost | S3 is designed for 99.999999999% (11 nines) durability of objects across AZs (AWS's design figure); Aurora keeps six copies of its storage across three AZs. The rest is on us: bytes before metadata, never overwrite, never delete a referenced chunk. |
Traffic
| Item | Math | Result |
|---|---|---|
| Daily active users | we assume half of the 100,000 | 50,000 |
| Commits | 50,000 × 20 saves a day | 1M/day |
| Average rate | 1,000,000 ÷ 86,400 s | 11.6/s |
| Peak | we assume 3× the average in the busiest hour | 35/s |
| New bytes per commit | we assume 3 MB on average: most saved files are small and go whole, big ones send 1–2 changed chunks | 3 MB |
| Upload traffic | 1M × 3 MB = 3 TB/day; 3 × 10¹² × 8 bits ÷ 86,400 s | 278 Mbps average, about 830 Mbps at peak |
| Download traffic | each commit downloaded by the user's one other device | 3 TB/day, 91.2 TB a month |
| Polls | 100,000 desktop clients, 60% online at peak, one poll per 30 s: 60,000 ÷ 30 | 2,000/s |
Storage
| Item | Math | Result |
|---|---|---|
| Current files | 100,000 × 5 GB | 500 TB |
| Version history | 3 TB of new chunk data a day, kept 30 days (replaced and deleted data) | about 90 TB |
| Total in S3 | 500 + 90 ≈ 590; we plan for | 600 TB |
| Chunk objects | we assume an average stored chunk of 2.5 MB (small files make short chunks): 600 TB ÷ 2.5 MB | about 240M objects |
| New chunks a day | 3 TB ÷ 2.5 MB | about 1.2M PUTs |
Metadata (we assume 1,000 files and folders per user)
| Table | Math | Result |
|---|---|---|
nodes | 100M rows × about 400 B | 40 GB |
revisions | 100M current + 30M in history (1M a day × 30) × about 300 B | about 39 GB |
journal | 1M rows a day × 200 B × 30 days | 6 GB |
chunks | 240M rows × about 100 B | 24 GB |
| Total | about 110 GB: one small Aurora cluster |
Propagation chain, worst case for a small edit (every step is dependent, so they add):
| Step | Time |
|---|---|
| Wait for the app to finish writing the file | 2 s |
| Hash, then ask which chunks are missing | 0.3 s |
| Upload one 4 MiB chunk at 10 Mbps | 3.4 s |
| Commit | 0.1 s |
| The other desktop's next poll, at worst | 30 s |
| The poll's round trip | 0.2 s |
| Download 4 MiB at 50 Mbps | 0.7 s |
| Total | about 36.7 s, under 1 minute |
A 2 GB new file takes 27 minutes to upload on a 10 Mbps link, which is simply the user's link; the minute target is for edits.
Monthly cost (us-east-1 on-demand list prices, 730 hours a month, ignoring free tiers; check the AWS Pricing Calculator before quoting):
| Item | Math | Monthly |
|---|---|---|
| S3 Standard storage | 600 TB: 50 TB × $0.023 + 450 TB × $0.022 + 100 TB × $0.021 per GB | ≈ $13,150 |
| S3 requests | 36.5M PUTs × $0.005/1,000 ≈ $182; about 73M GETs and HEADs × $0.0004/1,000 ≈ $29 | ≈ $210 |
| Data transfer out | 91.2 TB: 10 TB × $0.09 + 40 TB × $0.085 + 41.2 TB × $0.07 per GB | ≈ $7,184 |
| Aurora PostgreSQL | writer and reader, db.r7g.large at about $0.28/h × 2 × 730 ≈ $409; 110 GB × $0.10 ≈ $11; I/O ≈ $100 | ≈ $520 |
| Fargate (ARM) | 4 API tasks × (1 vCPU + 2 GB ≈ $0.0395/h) × 730 ≈ $115; scheduled jobs ≈ $10 | ≈ $125 |
| ALB | $0.0225 × 730 ≈ $16, plus 20 LCU (60,000 open connections ÷ 3,000 per LCU) × $0.008 × 730 ≈ $117 | ≈ $133 |
| CloudWatch, logs, secrets | estimate | ≈ $100 |
| Total | ≈ $21.4K/month |
About $0.21 per user a month. Storage is 61% of the bill and data transfer out is 34%; the servers are almost free. Every later round keeps that shape.
R1.8 Trade-Offs
Fixed-size vs content-defined chunks
| Fixed 4 MiB chunks (chosen for now) | Content-defined chunks (Round 2) | |
|---|---|---|
| Where cuts go | Every 4,194,304 bytes | Where a rolling hash of the content matches a pattern |
| Edit in place, append | Only the touched chunk changes | Same |
| Insert or delete near the start | Every later chunk changes | Only the chunks near the edit change |
| Client work | Hash only | A rolling hash over every byte, then hash |
| Known in practice | Dropbox documents a 4 MiB block for its content_hash (its API hashes each 4 MiB block and then the list of block hashes) | Backup and dedup tools; the research line from LBFS (Rabin fingerprints) to FastCDC (Gear hash) |
For a startup, fixed chunks are simple and good enough; the insert problem becomes worth solving when bandwidth and storage are big bills (Round 2).
Polling vs push
| Polling every 30 s (chosen) | Long polling | WebSocket push | |
|---|---|---|---|
| Delay | Up to 30 s | About a round trip | About a round trip |
| Server work | 2,000 mostly empty requests a second | Held requests; must wake them on a commit | A connection per device; must route nudges |
| Complexity | Trivial | Moderate | A connection fleet and a routing layer |
| When it fits | "Within a minute", 100K users | Middle ground; Dropbox's public API offers a long-poll call for exactly this | Half a second at hundreds of millions of devices (Round 2) |
Relational vs key-value namespace
| Aurora PostgreSQL (chosen) | DynamoDB | |
|---|---|---|
| Commit | One transaction with row locks: namespace, node, revision, journal, ref counts | A TransactWriteItems of up to 100 items; ref counts for a 2,385-chunk file don't fit |
| Listing and moves | SQL on parent_id; cycle checks in one transaction | A GSI per parent; checks written by hand |
| Scale ceiling | One writer; storage up to 256 TiB per cluster on recent versions | Horizontal, no practical table size limit |
| At 35 commits/s | Plenty | Also fine, more work to build |
At this size the relational database wins on simplicity. Round 2's 35,000 commits a second and hundreds of terabytes of metadata change the answer.
R1.9 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| Commit after a partial chunk upload (the client crashed or its URLs expired) | A commit names a chunk that isn't in S3 | The commit checks every new chunk with a HEAD before writing: 409 MISSING_CHUNKS with the list; the client uploads them and commits again with the same commit_id. Nothing points at missing bytes. |
| A device offline for weeks | Its cursor is older than min_seq (30 days of journal) | 410 Gone. The client does a full resync: it lists the whole tree with the current seq, compares each file's hash list with its local files, downloads what differs, and commits local-only changes with their old base revisions (stale ones become conflicted copies). Then it keeps the new cursor. Slow once, but correct. |
| Client clock skew | A laptop whose clock says 2031 or 2019 | We never use mtime to decide anything. Revision numbers and journal seqs come from the server's transaction; committed_at uses the server clock; client_mtime is display only. A skewed clock can't make an old version win. |
| A lost commit response | The client retries a commit that already succeeded | The retry finds its commit_id already applied (step 2 of the transaction) and gets the same 200, not a second revision and not a false conflict. |
| The Aurora writer fails | About 30 seconds of commit errors while the reader is promoted | Clients retry with exponential backoff and jitter; they're offline-capable anyway. No committed transaction is lost: Aurora's storage is shared by writer and reader. |
| A pre-signed URL expires mid-upload (15-minute URLs, a slow link) | S3 answers 403 | The client asks POST /v1/uploads again for fresh URLs; chunks already uploaded show up as "have" and are skipped. |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | Bytes first, metadata second; conditional commits on a base revision; idempotent commits by commit_id; Aurora and S3 across three AZs; 30-day version history as the backup users can reach themselves REL 4 · REL 9 · REL 10 |
| Performance Efficiency | Bytes go straight between devices and S3; only new chunks move; a change log instead of folder scans PERF 3 |
| Security | Short-lived pre-signed URLs, one key and one operation each; S3 verifies each chunk's SHA-256; every download authorized against the namespace; all API and S3 traffic over TLS SEC 3 · SEC 9 |
| Cost Optimization | About $21.4K a month; unchanged chunks never re-sent; deleted data freed after 30 days plus a 7-day grace COST 4 · COST 6 |
| Operational Excellence | Skipped this round: basic alarms on commit errors, 409 rates and sweeper lag. |
| Sustainability | Skipped this round: a handful of small ARM tasks; the storage itself is the footprint. |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Splits the namespace (a transactional database) from the bytes (content-addressed chunks in object storage), and states the write order: bytes first.
- Uploads in chunks with per-chunk retry, and sends only chunks whose hash is new.
- Uses a change log with a cursor instead of rescanning, and compacts it with a resync path.
- Makes every commit conditional on a base revision, and turns conflicts into conflicted copies rather than merges or last-writer-wins.
- Keeps version history with soft deletes, and frees chunks only when nothing references them.
- Never trusts client clocks.
Follow-up questions
-
"Why not let the API servers receive the file bytes and write them to S3?" Answer: every byte would cross our servers twice (in and out), our fleet would need to scale with bandwidth, not requests, and a 10 GB upload would tie up a server for an hour. Pre-signed URLs let S3 carry the bytes while we keep control: each URL names one key, one operation, a checksum and an expiry.
-
"Two chunks from different files of the same user have the same hash. Is that a problem?" Answer: it's a feature: they're the same bytes, stored once under
u/<user>/<hash>, and theref_countcounts both references. SHA-256 collisions between different contents are not a practical concern; nobody has ever produced one. -
"A user's disk fills up locally and the client can't write a downloaded file. What happens to the cursor?" Answer: the cursor only moves after a change is applied locally. The client stops at the change it couldn't apply, reports it, and retries later; the server doesn't care how far behind a device is, up to the 30-day log.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Store files under their path in S3" | Renames and moves become mass copies; S3 has no folders. |
| "Last writer wins, by modification time" | Silently deletes someone's edit, and client clocks are wrong often. |
| "Merge the two versions' chunks" | Byte-splicing a ZIP-based .xlsx or .docx produces a corrupt file. |
| "Write the revision, then upload the chunks" | A crash in between leaves a revision pointing at bytes that never arrived. |
| "Delete a file's chunks when the file is deleted" | Chunks are shared between revisions and files; undo becomes impossible. |
Round 2 · Senior · "100M Users, 250M Devices, 1B Changes a Day"
~40 min · Senior SDE (L6) · 1 region, 3 AZs, plus a CDN · 500M users, 100M daily, 250M connected devices · 1B commits/day, 34.7K/s at peak · 2.14 EB stored after dedup · commit P95 < 100 ms · push P95 < 500 ms · 99.99%
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "We built sync for a 100,000-user startup: about 1 million commits a day, 35 a second at peak, 600 TB in S3. The namespace (folders, files, revisions, manifests) lives in Aurora PostgreSQL; the bytes are fixed 4 MiB chunks in S3, named by their SHA-256 hash under each user's prefix and uploaded straight from the client with pre-signed URLs that make S3 check the checksum. Bytes go first, metadata second. A commit is one transaction: take the namespace lock, check the file is still at the base revision, write the revision, a journal row with the next sequence number, and exact reference counts. A stale base gets
409and the client saves a conflicted copy; nothing is ever merged. Devices pollchanges(cursor)every 30 seconds; the journal keeps 30 days and older cursors do a full resync. Deletes are soft, history is 30 days, and a sweeper under the same namespace lock frees chunks whose count stayed at zero for 7 days. About $21.4K a month. Open costs: an insert shifts every fixed chunk, the same file is stored once per user, polling doesn't scale to push speed, and one Aurora writer won't hold the next round."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: metadata in one transactional database, bytes in S3, and a journal every device reads with its cursor.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Big upload fails | Fixed 4 MiB chunks, separate PUTs, a manifest | Chunk bookkeeping |
| 1.2 | Where files live | Namespace in Aurora; content-addressed chunks in S3; bytes first | Two stores, orphans |
| 1.3 | Small edit re-uploads all | Upload only new hashes | Inserts shift chunks |
| 1.4 | Learning of changes | Journal + cursor, 30 s polling, 30-day log | Delay, empty polls |
| 1.5 | Offline edits collide | Base revision check → conflicted copy | Manual reconciling |
| 1.6 | Undelete | Soft delete, history, exact ref counts, locked sweeper | Retained bytes |
Open costs: fixed chunks; per-user copies of identical files; polling; one relational writer.
R2.1 The Scope Raise
Interviewer: "We're now a consumer service the size of Google Drive or Dropbox: 500 million accounts, 100 million people using it every day, 250 million devices connected at once. A billion saves a day. Edits have to show up on your other devices in half a second. Millions of people store the same installers and email attachments, and storage is our biggest bill; it has to come down. Folders can be shared with thousands of people. Last month a company shared a 10 GB video with 10,000 employees at 9 a.m. and our origin fell over."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How many devices are online at once, and how often do users save? | Plan for 250M connected devices (100M daily users × 2.5); about 10 commits per daily user per day. | 1B commits a day, 11,574/s average, 34.7K/s at peak (R2.6); a connection fleet for 250M sockets (step 2.3). |
| What does "half a second" cover? | From the commit being acknowledged to the other device receiving the notification, P95. | Push, not polling; we budget the chain in R2.6. |
| How much do identical files matter? | We estimate cross-user dedup at about 3.5× (an assumption the storage team wants us to use). | 7.5 EB of logical data becomes about 2.14 EB stored, if we store each chunk once for everyone (step 2.2). |
| What's a typical edit? | Mostly documents and spreadsheets tens of MB in size; small edits, often in the middle of the file. | Inserts in the middle break fixed chunks; content-defined chunking (step 2.1). |
| How big do shared folders get? | Up to 10,000 members; a few folders hold millions of files. | Shared folders become their own namespaces; moves must not rewrite millions of rows (step 2.5); nudges fan out to thousands of devices. |
| The 10 GB video at 9 a.m.? | 10,000 people, same office network, same 15 minutes. | Serve immutable chunks from a CDN with request collapsing (step 2.4). |
| Any change to durability or history? | Same promises: never lose a committed file, 30 days of history. | Garbage collection must stay safe while chunks are shared by millions of files (step 2.6). |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Users | 100K, 200K devices | 500M accounts, 100M daily, 250M connected devices |
| Commits | 1M/day, 35/s peak | 1B/day, 11,574/s average, 34.7K/s peak |
| Stored | 600 TB | 7.5 EB logical, about 2.14 EB after dedup |
| Chunking | Fixed 4 MiB | Content-defined, 1 MB min / 4 MB average / 8 MB max |
| Dedup | Per user | Across all users, with proof of possession |
| Propagation | Polling every 30 s | Push nudges, P95 < 500 ms |
| Sharing | None | Shared folders up to 10,000 members |
| Targets | 99.9%; within 1 min | Commit P95 < 100 ms; push P95 < 500 ms; ≥ 90% bandwidth saved on edits; 99.99% |
The "Not yet" list from R1.2 comes back: sharing, cross-user dedup and real-time push are all in scope now.
A note on the availability target: we promise 99.99% (4.4 minutes a month), not "five nines". 99.999% allows 5.26 minutes a year, less than one bad deploy.
R2.2 What Breaks in the Round 1 Design
| Round 1 choice | How it fails at the new scope |
|---|---|
| Fixed 4 MiB chunks | Insert one sentence near the start of a 40 MB document: every later chunk boundary shifts by those bytes, all 10 hashes change, and we re-upload all 40 MB. The "90% saved" promise holds only for edits in place. |
| Polling every 30 s | 250M devices ÷ 30 s = 8.3M requests a second, nearly all answering "nothing new", and still up to 30 s late against a 0.5 s target. |
| Per-user chunk keys | The same 200 MB installer is stored once per user who has it. The 3.5× dedup the storage plan counts on never happens. But dedup across users needs a global "do we have this chunk?" lookup at 35,000 commits a second, and it opens a privacy hole (step 2.2). |
| Direct S3 downloads | 10,000 employees fetching the same 2,500 chunks at once means thousands of GETs a second on the same S3 keys (S3 supports about 5,500 GETs a second per prefix before it scales out) and 100 TB of egress at S3's internet prices. |
| One Aurora writer, paths computed on read | 35,000 commit transactions a second and hundreds of terabytes of metadata exceed one writer (and one cluster's 256 TiB storage limit). Moving a folder with a million files must stay a small operation. |
| Exact ref counts in the commit transaction | A popular installer chunk is referenced by millions of files: incrementing one counter item on every commit makes it the hottest item in the system, and a 2,500-chunk file can't update 2,500 counters inside one transaction. Orphans and deletes need a different collector (step 2.6). |
The order we fix it in: chunking first (2.1), because it decides what a chunk is; then dedup across users (2.2), which changes chunk names; then push (2.3) and the CDN (2.4), which move the bytes and the news; then moves at scale (2.5); and last the garbage collector (2.6), which must be safe against everything before it.
R2.3 New Requirements and API Additions
Negotiation now returns three kinds of answers per chunk
httpPOST /v1/uploads HTTP/1.1 Authorization: Bearer <device token> Content-Type: application/json { "ns_id": "ns_7Qm", "node_id": "n_8f21", "base_rev": 14, "new_chunks": [ { "hash": "e19c...5a", "size": 3981204 }, { "hash": "c41e...07", "size": 4012331 }, { "hash": "a0d3...b8", "size": 2210876, "from": { "node_id": "n_3c10", "rev": 2 } } ] }
The client sends only hashes that are not in its base revision's manifest: it already knows the server has those. from says "this chunk is also in another file of mine, n_3c10 revision 2"; the server checks that revision lists that hash and that the user can read it.
json{ "session": "us_4c1e77", "upload": [ { "hash": "e19c...5a", "put_url": "https://drive-chunks.s3.us-east-1.amazonaws.com/e19c/e19c...5a.1?X-Amz-Signature=...", "required_headers": { "x-amz-checksum-sha256": "4ZxQ...=" }, "expires_at": "2026-09-28T09:15:00Z" } ], "prove": [ { "hash": "c41e...07", "challenge": { "user": "u_1001", "hash": "c41e...07", "day": "2026-09-28", "ranges": [[1048576, 1024], [2621440, 1024], [3801088, 1024]], "expires_at": "2026-09-28T09:15:00Z", "sig": "..." } } ], "have": [ "a0d3...b8" ] }
- upload: nobody has this chunk.
.1at the end of the key is the chunk's generation (step 2.6). - prove: another user has it, and it's popular enough to skip the upload; prove you hold the bytes (step 2.2). Each range is
[offset, length], all inside the chunk's 4,012,331 bytes. - have: you already own it (it's in a file you can read).
Commit carries the full manifest and the proofs
json{ "commit_id": "c_91e0aa3f", "session": "us_4c1e77", "ns_id": "ns_7Qm", "node_id": "n_8f21", "base_rev": 14, "size": 41890112, "manifest": ["7a11...0c", "e19c...5a", "c41e...07", "a0d3...b8", "..."], "proofs": [ { "hash": "c41e...07", "challenge": { "...": "echoed as received" }, "proof": "5be0...9d" } ] }
proof is HMAC-SHA256, keyed with the challenge's sig (a value the client can't compute itself), over the bytes of the three ranges joined in order. Response: { "rev": 15, "seq": 912004411 }.
The commit checks ownership of every hash. The negotiation stored an upload_sessions item (us_4c1e77: user, node, base revision, the hashes granted for upload, the hashes sent a challenge, the from references it accepted, and an expiry our servers check with their own clock). At commit, each hash in manifest must be one of: (a) in the base revision's manifest; (b) granted for upload in this session (and now verified in S3); (c) covered by a proof verified in this commit; (d) in a from revision the user can read. Any other hash gets 403 NOT_OWNED. Without this check a client could skip the negotiation and commit a manifest of hashes it merely knows.
The push channel (one WebSocket per device, wss://push.drive.example/v1/connect)
json{ "type": "hello", "device_id": "d_77", "cursors": { "ns_7Qm": 912004410, "ns_team9": 1022 } }
json{ "type": "nudge", "ns": "ns_7Qm", "seq": 912004411, "pull_within_ms": 0 }
A nudge says only "namespace ns_7Qm has reached seq 912,004,411". It carries no file data; the device then calls GET /v1/changes?cursor=ns_7Qm:912004410,ns_team9:1022, which returns changes for each namespace it mounts. pull_within_ms lets the server spread a big fan-out (step 2.4): the device waits a random time up to that value before pulling.
Shared folders and membership
POST /v1/shared-folderswith{ "node_id": "n_proj" }turns a folder into its own namespace (a background job, step 2.5).POST /v1/namespaces/ns_team9/memberswith{ "user_id": "u_2044", "role": "editor" }; the member's root gets a mount node pointing atns_team9, and their devices addns_team9to their cursors.
Moves
httpPOST /v1/nodes/n_proj/move HTTP/1.1 Content-Type: application/json { "commit_id": "c_5d2e19b0", "new_parent_id": "n_archive", "new_name": "Project X" }
200 for a move inside one namespace; 409 CYCLE if n_archive is inside n_proj; 202 Accepted with a job ID for a move into a different namespace (step 2.5).
Recap
- Clients ask only about chunks new to the file; the server answers upload, prove or have.
- Proofs are HMACs over server-chosen byte ranges; commits carry the full manifest, and every hash must be owned (base revision, this session's grants or proofs, or a readable
from), else403 NOT_OWNED. - One WebSocket per device carries tiny nudges; the device pulls changes per namespace.
- Shared folders are namespaces with members; moves are small unless they cross namespaces.
R2.4 Design Evolution: Content-Defined Chunks, One Copy of Everything, and Push
Step 2.1: Inserting One Sentence Near the Start Re-uploads the Whole File
The problem: a user adds a sentence at the top of a 40 MB report. With fixed 4 MiB chunks, every byte after the sentence moves, so every chunk's contents and hash change. We upload 40 MB for a 100-byte edit, and yesterday's 10 chunks can't be shared with today's. What would you do?
Step 2.2: Millions of Users Upload the Same Installer
The problem: 3 million users have the same 200 MB installer in their drives, and email attachments repeat the same way. With per-user chunk keys we store 600 TB for one file. The storage team's plan assumes 3.5× savings from storing identical chunks once, across all users. What would you do?
Primitive: Bloom Filters and Counting Filters · Drill: The Web Crawler That Forgot Its History. If the filter were used and filled past its design size, its false-positive rate would climb (at 15%, 15% of truly new chunks would cost an extra index read, and a rebuild or resizing of the sub-filter would be due). And a Valkey set of all 511 billion SHA-256 hashes at about 80 bytes per member with overhead would need about 41 TB of RAM against the filter's 639 GB: exact membership at 64 times the memory, which is why an exact index belongs in DynamoDB, not in a cache.
Step 2.3: Polling 250 Million Devices
The problem: Round 1's desktops poll every 30 seconds. At 250M devices that's 8.3M requests a second, and the target is now 0.5 seconds, not 30. What would you do?
Primitives: WebSocket, SSE and Long Polling · Change Data Capture and the Outbox Pattern · Drills: The Gateway Restart That DDOSed the Chat Fleet · The Dual-Write That Broke Search Consistency. Why the stream and not a second write, or a poller: a service that writes DynamoDB and then publishes to a queue can crash between the two and silently drop the nudge; and an application thread polling a "to publish" table every 500 ms adds a query load that never stops, a 500 ms latency floor, and an index on unpublished rows, while the table's own change stream is ordered per item, kept 24 hours, and read by Lambda at no request charge.
Step 2.4: 10,000 People Download the Same 10 GB at 9 a.m.
The problem: a company shares a 10 GB town-hall recording with 10,000 employees. At 9:00 their clients all get the nudge and start downloading: 10,000 × 10 GB = 100 TB, the same 2,500 chunks for everyone, within minutes. What would you do?
Primitive: Distributed Cache Patterns and Eviction · Drill: The Image Service That Made Every Origin Server Sweat. Content-addressed names are the answer to its "how long until viewers stop seeing the old image?": zero, because a new version is a new URL; and the CDN-versus-bigger-origin trade-off is the egress price, request collapsing and latency argument above.
Step 2.5: Moving a Folder With 1 Million Files
The problem: a user moves Projects/ (1M files in 40,000 folders) into Archive/. If files stored their full path, we'd rewrite a million rows. And two devices move folders at the same time: one moves A into C (which is inside B), while the other moves B into A.
What would you do?
Primitive: Database Isolation Levels, ACID and Concurrency Anomalies · Drill: The On-Call Doctor Anomaly. Why not run everything at strict serializable isolation instead? In a relational database, serializable transactions abort more often under contention and have to be retried, and at 35,000 commits a second we'd pay that on every commit to protect the rare move; we pay for the extra checks only on moves.
Step 2.6: Chunks Nobody References Any More
The problem: about 4 PB of new chunk data arrives every day, and in a steady state about as much stops being referenced: revisions age out of the 30-day history, files leave the trash, uploads are abandoned before their commit. Round 1's exact reference counts, updated inside each commit, don't survive this round (R2.2). But deleting a chunk that someone is about to reference loses their file. What would you do?
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | An insert re-uploads the file | FastCDC (Gear hash), 1 / 4 / 8 MB | Variable sizes; client CPU |
| 2.2 | Everyone stores the same installer | Global chunk index; server-side dedup; proofs of possession; a random threshold before skipping uploads | Proof reads; extra uploads below the threshold |
| 2.3 | Polling 250M devices | WebSocket gateways, registry, fast path + stream backstop, coalesced nudges | The largest compute fleet |
| 2.4 | 10,000 downloads at 9 a.m. | CloudFront, immutable year-long caching, Origin Shield, jittered starts, LAN sync | CDN bill; caches we can't take back |
| 2.5 | Moving 1M files | Parent IDs; cycle check as condition checks in the move transaction | Depth cap 64; cross-namespace jobs |
| 2.6 | Unreferenced chunks | Mark, doom, resurrect, second mark, sweep; generations; versioned bucket | Weeks of delay before space is freed |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
The write path: metadata commits in DynamoDB, bytes straight into S3; nudges leave by the fast path and again through the table's own stream.
Synthesizing vector architecture diagram...
The read path: the manifest comes from the API and is never cached; chunks come from the nearest cache that has them, and at most about once from S3.
Tables (DynamoDB, on-demand)
| Table | Key | Main attributes |
|---|---|---|
namespaces | ns_id | seq (last journal number), min_seq, kind (user root or shared folder), owner |
nodes | ns_id + node_id | parent_id, name, is_folder, rev, size, deleted_at; a GSI on ns_id#parent_id + name lists a folder |
names | ns_id#parent_id + lowercased name | node_id; written with attribute_not_exists in the same transaction, so two files can't share a name in a folder |
revisions | node_id + rev (Number) | manifest (32-byte hashes; manifests over 64 KB are stored as a content-addressed S3 object and referenced by hash), size, commit_id, device_id, committed_at |
journal | ns_id + seq (Number) | node_id, op, rev, parent_id, name, the manifest or its reference: everything a device needs, in one Query |
commits | commit_id | ns_id, node_id, rev, seq: the idempotency record, written with attribute_not_exists in the commit transaction; a TTL clears it after 7 days, eventually |
chunks | hash | state, gen, size, created_at, doomed_at, uploaders (count up to T) |
members | ns_id + user_id | role; a GSI by user_id lists what a user mounts |
upload_sessions | session_id | user, node, base revision, hashes granted, hashes challenged, accepted from references, expires_at (checked with our clock; a TTL clears it later) |
chunk_uploaders | hash + user_id | one item per distinct uploader below T, written with attribute_not_exists |
The seq and rev sort keys are Numbers, so they sort as numbers.
The commit transaction. First the handler checks ownership: every hash in the manifest must be in the base revision, granted or proven in this upload session (upload_sessions), or in a readable from revision; any other hash is 403 NOT_OWNED. Then it reads namespaces.seq (say s), then runs one TransactWriteItems: update the namespace to s + 1 if seq = s; put journal (ns, s + 1) if it doesn't exist; update the node to rev = base_rev + 1 if rev = base_rev; put the revision; put commits (commit_id) if it doesn't exist. If the commits condition is what failed, it's a retry: read the stored result, return it, and re-send the nudge. If rev failed: 409 STALE_BASE. If seq failed: another commit in this namespace won the race; read seq again and retry. Commits to one namespace are serialized through its head item, which caps a single namespace at a few hundred commits a second: plenty for a person, and a known limit for the busiest shared folders.
Tracing an edit with one new chunk
Synthesizing vector architecture diagram...
One chunk of about 4 MB moved each way for a 40 MB file: 90% saved. The stream backstop sends the same nudge a moment later; the desktop sees seq already applied and does nothing.
Tracing the shared-file herd (numbered steps)
- 08:59:40: the town-hall video (10 GB, 2,500 chunks after CDC) is committed to the shared folder
ns_allhands(10,000 members). One journal row,seq1,204. - The router reads the members from its cache: 10,000 users, about 25,000 devices, 18,000 of them online. It sends 18,000 nudges with
pull_within_ms = 300000. - Each client waits a random 0–300 s, then pulls
changes, gets the manifest and signed URLs, and fetches chunks with 4 parallel downloads. So about 60 new clients start each second instead of 18,000 at once. - The first requests for each chunk miss at their edge, then at the regional edge cache, then at Origin Shield, which fetches from S3 once per chunk: about 2,500
GETs on S3 in total. Everything after is served from cache. - On the office LAN, clients that already hold chunks serve them to their neighbors; the internet link carries far less than 100 TB.
Tracing a folder move (numbered steps)
- The desktop moves
Projects(n_proj) intoArchive(n_arch), both inns_7Qm. - The API walks
n_arch's ancestors:n_arch→n_root.n_projisn't among them. - It reads
seq= 912,004,500 and runs one transaction: updaten_proj.parent_id = n_archifrevmatches;ConditionCheckn_arch.parent_id = n_root; delete the oldnamesitem and put the new one if absent; bumpseqto 912,004,501 if it's still 912,004,500; put the journal rowMOVE; put thecommitsrecord. Seven items. - Other devices pull one
MOVEchange and rename one directory locally. None of the 1M files inside was touched.
Losing an AZ. The API, routers and gateways run in three AZs; DynamoDB, S3 and Lambda are regional. Devices on the lost AZ's gateways reconnect to the others with jittered backoff and resume tickets (R2.6 sizes the survivors).
R2.6 Numbers and Cost
Traffic
| Item | Math | Result |
|---|---|---|
| Connected devices | 100M daily users × 2.5 devices | 250M (a planning peak) |
| Commits | 100M × 10 a day | 1B/day |
| Average rate | 10⁹ ÷ 86,400 s | 11,574/s |
| Peak | we assume 3× in the busiest hour | 34,722/s |
| Push frames | each commit nudges 1.5 other devices on average: 11,574 × 1.5 | 17,361/s, 52,083/s at peak (shared folders add more, coalesced to one per device per namespace per second) |
| Upload without deltas | the average edited file is 40 MB (10 chunks): 10⁹ × 40 MB × 8 ÷ 86,400 s | 3.70 Tbps |
| Upload with CDC deltas | about one new 4 MB chunk per commit: 10⁹ × 4 MB × 8 ÷ 86,400 s | 370 Gbps (1.11 Tbps at peak): 90% saved |
| Downloads | 370 Gbps × 1.5 devices | 556 Gbps: 6.0 PB a day, 182.4 PB a month |
| API requests | about 5 per commit (negotiate, commit, 1.5 change pulls, 1.5 URL batches) | 5B/day, 58K/s average |
Delta sync saves nothing on a file smaller than one chunk: it is always sent whole. That's why the 40 MB average is taken over edited files, which skew large.
Storage and dedup
| Item | Math | Result |
|---|---|---|
| Logical data | 500M users × 15 GB average (history included) | 7.5 EB |
| Stored, after dedup | 7.5 EB ÷ 3.5 (an assumption) | 2.14 EB |
| Unique chunks | 2.143 × 10¹⁸ B ÷ 4,194,304 B (4 MiB) average | about 511B |
| New chunk data | 1B new chunks a day × 4 MB | 4 PB/day; about as much is deleted a day in a steady state |
Metadata (we assume 1,000 files and folders per account)
| Table | Math | Result |
|---|---|---|
nodes | 500B items × about 300 B | 150 TB |
revisions, current | 500B × about 250 B (big manifests live in S3) | 125 TB |
revisions, 30-day history | 30 × 1B × 500 B | 15 TB |
journal, 30 days | 30 × 1B × 200 B | 6 TB |
chunks | 511B × about 100 B | 51 TB |
names | 500B × about 100 B | about 50 TB |
nodes GSI (by parent and name) | 500B × about 150 B | about 75 TB |
| Total | about 475 TB |
If we kept every revision for 3 years instead of 30 days, revisions alone would be 1.095 trillion × 500 B = 547.5 TB; the 30-day history is what keeps all tables together near 475 TB.
DynamoDB writes per commit. Transactional writes cost 2 units per item (up to 1 KB): namespace, journal, node, revision and commit record = 5 × 2 = 10; plus the new chunk's PENDING and LIVE updates, 2 more: 12 write units per commit. At peak that's 34,722 × 12 ≈ 417,000 write units a second, about 69,000 on each of the busiest tables: above the default on-demand limit of 40,000 per table, so we raise those quotas (they're adjustable) and pre-warm the tables before launch.
Gateways, with per-AZ rounding. We assume one 4 vCPU / 16 GB Fargate task holds 50,000 idle WebSockets (a figure to load-test). 250M ÷ 50,000 = 5,000 tasks; spread over 3 AZs that's 1,666.7 per AZ, rounded up to 1,667 per AZ, 5,001 tasks. If an AZ is lost, 3,334 tasks carry 250M ÷ 3,334 ≈ 75,000 connections each for the minutes Fargate needs to add tasks. That's 16 GB ÷ 75,000 ≈ 213 KB of memory per connection, an emergency ceiling we must load-test, not a normal operating point.
Latency budgets (P95; dependent steps add, parallel steps count once, as their maximum)
| Negotiate | ms |
|---|---|
| ALB and WAF | 3 |
| Token check (local signature) | 1 |
BatchGetItem on chunks | 10 |
| Sign URLs and challenges | 1 |
| Total | 15 |
| Commit | ms |
|---|---|
| ALB and WAF | 3 |
| Token check | 1 |
In parallel: strongly consistent chunk reads 10, HEAD of new chunks 20, proof range GETs 30 | 30 |
Chunk state updates (PENDING to LIVE) | 10 |
Read namespaces.seq | 5 |
TransactWriteItems | 25 |
| Total | 74, under 100 |
| Push, from commit acknowledged to frame received | ms |
|---|---|
| Commit service to router | 2 |
| Members cache and registry lookups | 2 |
| Router to gateway | 2 |
| Gateway to device over the internet (an assumption for P95) | 80 |
| Total | 86, under 500 |
The stream backstop is slower: DynamoDB Streams' own delay (AWS publishes no figure, so we measure it) plus up to 250 ms, because Lambda polls each stream shard 4 times a second, plus the same 84 ms of routing and delivery. It only decides the latency when the fast path was lost.
Availability. 99.99% of a 30.4-day month is 43,776 min × 0.0001 ≈ 4.4 minutes.
Monthly cost (us-east-1 list prices, decimal GB, 730 hours; at this size real contracts are negotiated, so read these as proportions):
| Item | Math | Monthly |
|---|---|---|
| S3 storage, Intelligent-Tiering | 2.143B GB. Assumed mix: 10% frequent ($0.021 at this volume), 20% infrequent ($0.0125), 70% archive instant access ($0.004) = $0.0074/GB ≈ $15.86M; monitoring 511B objects × $0.0025/1,000 ≈ $1.28M | ≈ $17.14M |
| (for comparison) S3 Standard | 50 TB × $0.023 + 450 TB × $0.022 + the rest × $0.021 | (≈ $45.0M) |
| Non-current versions (GC safety net) | 28 PB × $0.0074 | ≈ $0.21M |
| S3 requests | 30.4B PUTs × $0.005/1,000 ≈ $152K; origin GETs (80% of 1.5B daily downloads miss the cache for personal chunks) 36.5B × $0.0004/1,000 ≈ $14.6K; HEADs ≈ $12.2K; proof ranges (0.3 claimed chunks per commit × 3) ≈ $10.9K | ≈ $0.19M |
| CloudFront egress | 182.4 PB: 1 TB free, then $0.085, $0.080, $0.060, $0.040, $0.030 and $0.025 through 5 PB, $0.020 for the remaining 177,376 TB | ≈ $3.69M |
| CloudFront requests | 45.6B × $0.01/10,000 ≈ $45.6K; Origin Shield 36.5B × $0.0075/10,000 ≈ $27.4K | ≈ $0.07M |
| DynamoDB | writes 1B × 12 × 30.4 = 364.8B × $0.625/M ≈ $228K; reads about 10 per commit ≈ $38K; storage 475 TB × $0.25 ≈ $119K; point-in-time recovery × $0.20 ≈ $95K | ≈ $0.48M |
| Gateways | 5,001 tasks × (4 vCPU × $0.03238 + 16 GB × $0.00356)/h × 730 h = 5,001 × $136.13 | ≈ $0.68M |
| NLBs (TCP) | 250M connections ÷ 100,000 per unit = 2,500 units × $0.006 × 730 ≈ $10,950, plus about 50 NLBs × $16 | ≈ $0.01M |
| Sync API, routers | about 100 tasks of 2 vCPU / 4 GB on average × $57.67 | ≈ $0.01M |
| ALB and AWS WAF | ALB ≈ $10K; WAF 5B requests a day × 30.4 × $0.60/M ≈ $91K | ≈ $0.10M |
| Valkey, Lambda, GC scans | 6 cache.r7g.xlarge ≈ $1.5K; stream Lambdas ≈ $3K; weekly scans (the revisions table about $2.1K a scan, plus S3 GETs of large manifests: about 1% of 530B revisions = 5.3B × $0.0004/1,000 ≈ $2.1K a scan, ≈ $9.2K a month) and batch jobs ≈ $29K | ≈ $0.03M |
| CloudWatch, logs, VPC endpoints, cross-AZ | estimates | ≈ $0.08M |
| Total | ≈ $22.7M/month |
About $0.23 per daily user a month, or $0.045 per account. Storage is 76% of the bill and CDN egress 16%; everything else together is under 8%. Two choices dominate the total: dedup (without it, storage would be 3.5 times larger) and tiering (S3 Standard would add about $28M). At this scale the next big lever is to stop renting object storage and build our own, which is what the S3-like object storage loop designs.
R2.7 Trade-Offs
Chunk size
| Average chunk | Dedup and delta | Chunks to track (2.14 EB) | Our view |
|---|---|---|---|
| 1 MB | Best: small edits send 1 MB; more matches across files | about 2.1 trillion: 4× the index, the manifests and the S3 PUTs | Too much metadata |
| 4 MB (chosen) | A small edit sends about 4 MB | about 511B | Balanced |
| 16 MB | Worse: small edits send 16 MB; fewer matches | about 128B | Too coarse for documents |
CDC vs fixed chunks
| Fixed | CDC (chosen) | |
|---|---|---|
| Insert near the start | Everything after it changes | 1–2 chunks change |
| Chunk sizes | Uniform | 1–8 MB |
| Client CPU | Hash only | A rolling hash too (cheap) |
| Cross-user dedup | Only if files align on 4 MiB boundaries | Finds shared content at any offset |
Cross-user dedup vs privacy
| Per-user only | Server-side, always upload | Skip uploads after T copies, with proofs (chosen) | |
|---|---|---|---|
| Storage saved | Within each user | Full (3.5×) | Full (3.5×) |
| Upload bandwidth saved on shared content | No | No | Yes, for popular chunks |
| Existence leak ("does anyone have this?") | None | None | Only "at least T copies exist", for popular chunks |
| Works with per-user encryption keys | Yes | No: different keys, different ciphertexts | No |
The last row is Round 3's problem: dedup across users and keys held per customer don't mix.
Push vs poll, and whose gateway
| Poll every 30 s | Our WebSocket fleet (chosen) | API Gateway WebSockets | |
|---|---|---|---|
| Delay | Up to 30 s | About 90 ms | About the same |
| Monthly cost at 250M devices | WAF alone ≈ $13.1M, plus servers for 8.3M requests/s | ≈ $0.69M (gateways + NLBs) | ≈ $2.74M in connection-minutes, before messages |
| Limits | – | Our own load tests | 2-hour maximum connection, 10-minute idle timeout |
R2.8 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| Concurrent edits to one file (two people in a shared folder) | The second commit's base is stale | 409 STALE_BASE → a conflicted copy. We never merge bytes of formats we don't understand; a byte-level merge of two .xlsx versions is a corrupt ZIP. |
| Laptops sleep and wake | Half-open sockets (a sleeping laptop sends no close); then thousands of reconnects at 9 a.m. | Gateways ping every 60 s and drop a connection after missed pongs, removing its registry entry. On wake, the client sees a jump between its wall clock and its monotonic clock, drops the old socket at once, reconnects with jitter and a resume ticket, and only after its gateway has registered it in the registry does it pull changes with its saved cursors, so no nudge can fall between the pull and the registration: a catch-up, not a rescan. |
| A move cycle | Two concurrent moves | The second fails its condition checks and gets 409 CYCLE (step 2.5). |
| Orphan chunks (an upload never committed) | PENDING items with no commit | The sweep's scan finds PENDING items older than 7 days, sets them DELETED (conditionally) and deletes that generation's object. |
| A stale "have" hint (a chunk deleted since a cache or index read said it existed) | A commit names a chunk that isn't LIVE | Every commit re-reads chunk states strongly consistently; DELETED → 409 MISSING_CHUNKS, the client re-uploads under the next generation. A "maybe present" answer from any cache or filter is never trusted without that read. |
| The fast path loses a nudge | A device doesn't hear about a commit | The stream backstop sends it; the 10-minute pull catches anything else. |
| The stream consumer lags | Iterator age grows | Only the backstop is late; alarm on iterator age; Lambda retries a failing batch until it succeeds (records stay 24 hours), so we alarm long before that. |
| A hot shared namespace | Many commits contend on one head item | Conflicts retry with jitter; above a few hundred commits a second the namespace is rate-limited (429) and the client queues locally. |
| An AZ is lost | A third of the sockets drop | Reconnects with jitter; survivors run at 75,000 connections each until Fargate adds tasks. |
Primitive: Circuit Breaker, Bulkhead and Fault Tolerance
R2.9 Production Gotchas
| Gotcha | Why it hurts | What we do |
|---|---|---|
| Chunks in database BLOBs | Bloats the database, its backups and its buffer cache; costs many times what object storage does | Metadata in DynamoDB, bytes in S3 |
Trusting client mtime | Skewed clocks make old versions win | Server-assigned revisions and seqs; mtime for display |
| A notification per chunk | A 1 GB file = 250 nudges; devices pull half-written state | One nudge per commit, after the transaction, coalesced per namespace |
| Polling at scale | 8.3M empty requests a second | Push nudges; a slow safety-net pull |
| Auto-merging binary files | Byte-spliced ZIP containers are corrupt | Conflicted copies; merge only formats the server understands |
| "Skip the upload if the hash exists" | Anyone who knows a hash gets the file (Dropship) | Proofs of possession, and a random threshold before skipping |
| Deleting a chunk when its count hits 0 | Races a new reference | Doom, resurrect, second mark, then delete by generation |
R2.10 Pillar Check
| Pillar | What Round 2 covers |
|---|---|
| Reliability | Conditional commits with idempotency records; retries that re-send the nudge; a stream backstop; safe GC with generations and a versioned bucket; gateways sized per AZ REL 4 · REL 10 · REL 11 |
| Performance Efficiency | CDC deltas save 90% of upload bytes; commit P95 about 74 ms; push about 86 ms; immutable chunks cached at the edge PERF 3 · PERF 4 |
| Security | Proofs of possession against Dropship; a threshold against the existence side channel; S3-verified checksums against poisoning; logical-size quotas; signed short-lived URLs SEC 3 · SEC 8 |
| Cost Optimization | About $22.7M a month, 76% storage; dedup and Intelligent-Tiering are worth tens of millions a month; a Bloom filter rejected by arithmetic; CloudFront over S3 egress COST 5 · COST 8 |
| Operational Excellence | Alarms on commit P95, push lag, conflict rate, stream iterator age, GC backlog and quarantine size OPS 8 |
| Sustainability | Light this round: every byte stored once, deltas instead of whole files, cold data in cheaper tiers SUS 4 |
R2.11 Round 2 Rubric and Follow-Ups
What a strong senior (L6) answer shows
- Explains why fixed chunks break on inserts, and describes content-defined chunking correctly (a rolling hash picks cut points; Gear-based FastCDC is the modern choice, Rabin the older one), with min / average / max sizes.
- Dedups across users and immediately raises proof of possession and the existence side channel.
- Does the arithmetic before adding a Bloom filter or a cache, and knows a filter's "maybe" must be confirmed.
- Pushes nudges and pulls data; keeps a stream backstop, re-sends on duplicate retries, and handles reconnect storms.
- Serves immutable chunks from a CDN and notices the customer's own link is the real limit.
- Makes moves one update, and sees write skew in concurrent cycle checks.
- Builds a GC that can't delete a chunk a commit is about to use.
Follow-up questions
-
"Why not merge concurrent edits to text files automatically?" Answer: a three-way merge of plain text is possible when both edits touch different lines, but the sync client can't know a file is plain text in a format where merging is safe (a
.csvread by a program may break on a merged line), and a wrong merge is silent data corruption. Conflicted copies are loud and safe. Real-time co-editing belongs in an editor that understands the document, not in file sync. -
"A device has been offline for 45 days and has local edits." Answer: its cursor is below
min_seq:410. It resyncs: lists the tree at the currentseq, compares hashes, downloads what changed, and commits its local edits with their original base revisions, so any that went stale become conflicted copies. Nothing it wrote offline is lost. -
"How would you know dedup is really 3.5×?" Answer: it's measured, not assumed: logical bytes over stored bytes, per day of uploads and in total, from the quota counters and S3 Storage Lens. If it's lower, storage cost rises in proportion; the plan should show the cost at 2× and 5× too.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Smaller fixed chunks fix inserts" | Every later chunk still shifts. |
| "FastCDC uses Rabin fingerprints" | FastCDC rolls a Gear hash; Rabin CDC is the older approach. |
| "If the hash exists, don't ask for the bytes" | Hands any file to anyone who knows its hashes. |
| "A Bloom filter in front of the index" (without the math) | Here it costs more than three times the reads it saves. |
| "Check for cycles, then move" | Two moves can each pass and together make a cycle. |
| "Reference counts from the change stream" | At-least-once delivery double-counts. |
Round 3 · Architect · "Enterprise: Sharing, Keys, Laws, Ransomware"
~45 min · Principal (L7) · 4 regions: a home and a standby in the US and in the EU · 20M seats in 20,000 organizations · 180M commits/day · 1 EB logical, 625 PB stored · tenant isolation · residency · point-in-time rollback · survive a region loss
R3.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 3. If you're starting here, it's everything you need from Rounds 1 and 2.
Round 2 in 60 seconds. "We run a consumer sync service for 500 million accounts from one region: 250 million connected devices, a billion commits a day, 34,700 a second at peak. Clients cut files with FastCDC, a Gear rolling hash with 1, 4 and 8 MB bounds, so an insert changes only nearby chunks and an edit uploads about 4 MB of a 40 MB file. Chunks are stored once for everyone, 7.5 EB becomes about 2.14 EB, under a global chunk index in DynamoDB; skipping an upload needs a proof of possession, and only after a random number of users have uploaded the chunk, so the answer doesn't reveal private files. Metadata is DynamoDB: a commit is one transaction on the namespace head, the node, the revision, the journal and an idempotency record, conditional on the base revision; stale bases become conflicted copies. Devices hold WebSockets to 5,001 gateway tasks and get tiny nudges, sent right after the commit and again from the journal's change stream; they pull changes by cursor. Chunks are served from CloudFront with year-long caching and Origin Shield. Moves are one update with cycle checks as condition checks. GC is mark, doom, resurrect, second mark, sweep, with a generation in every key and a versioned bucket as a safety net. About $22.7 million a month, three quarters of it storage. Open costs: no organizations or permissions beyond folder membership, our keys only, one region, and nothing between a ransomware infection and every synced device."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: metadata transactions, bytes stored once in S3, nudges pushed, chunks pulled through the CDN.
Round 2 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | Inserts shift chunks | FastCDC, 1 / 4 / 8 MB | Variable sizes |
| 2.2 | Same installer everywhere | Global index, proofs, random threshold | Proof reads, some extra uploads |
| 2.3 | Polling 250M devices | WebSocket gateways, fast path + stream backstop | The biggest fleet |
| 2.4 | 9 a.m. herd | CloudFront, Origin Shield, jitter, LAN sync | CDN bill |
| 2.5 | Moving 1M files | Parent IDs, cycle checks in the transaction | Depth cap |
| 2.6 | Orphan chunks | Mark, doom, resurrect, sweep, generations | Delayed reclaim |
Open costs: permissions; customer keys; legal holds; residency; ransomware; audit.
R3.1 The Scope Raise
Interviewer: "We're launching the enterprise edition. Customers are organizations, some with 400,000 employees and millions of shared files in drives that are twelve folders deep. Their security teams want to hold their own encryption keys. Their legal teams put holds on employees' files during lawsuits. European customers' data must stay in the EU. Last quarter a customer's laptop got ransomware, encrypted 50,000 files, and our sync faithfully spread the damage to every device. And admins need an audit log of every access."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How big is the enterprise edition? | About 20M seats in 20,000 organizations; the largest has 400,000 seats. Plan 50 GB per seat, including version history. | 1 EB of logical data; per-tenant sizing, and one very large tenant to plan for (R3.6). |
| How do permissions work? | Shared drives owned by the organization; roles (viewer, commenter, editor, manager); folders inherit from their parents, and some folders are restricted to a smaller group. Link sharing inside the organization. | A permission check on every read and write, walking up to 12 levels (step 3.1). |
| What does "hold their own keys" mean? | They want to revoke our access to their data at any moment, and see every use of their key in their own logs. | Per-tenant keys in their control; no dedup across tenants (step 3.2). |
| What does a legal hold require? | Nothing in scope may be permanently deleted until the hold is released, whatever users or retention rules do. | Holds become a separate layer that every purge and the GC must respect (step 3.3). |
| Where must EU data stay? | In the EU: files, metadata, backups and logs. | Two jurisdictions, each with a home and a standby region (step 3.4). |
| What happened with the ransomware? | 50,000 files rewritten in 20 minutes from one laptop; every device synced the encrypted versions. | Detect mass rewrites, stop that device, and roll the drive back to a point in time (step 3.5). |
| What must the audit log hold? | Who did what to which file, when and from where; kept for the contract period; exportable by the customer. | An append-only per-tenant audit stream (step 3.6). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Customers | Consumers | 20,000 organizations, 20M seats |
| Stored | 2.14 EB after global dedup | 1 EB logical → 625 PB, dedup only within a tenant |
| Commits | 1B/day | 180M/day (9 per seat per day) |
| Permissions | Folder membership | Roles, inheritance, restricted folders, org links |
| Keys | Ours | Per tenant, optionally held by the customer |
| Deletion | Trash, 30-day history, GC | + legal holds; 90-day history |
| Regions | 1 | 4: US home + standby, EU home + standby |
| Targets | Commit P95 < 100 ms; 99.99% | + tenant isolation; residency; drive rollback; survive a region (RPO: seconds for metadata, 15 min for chunks) |
R3.2 What Breaks in the Round 2 Design
| Round 2 choice | What breaks at the new scope |
|---|---|
| Permission = "you're a member of this namespace" | Enterprise permissions differ folder by folder, 12 levels deep, with restricted subfolders. Walking 12 parents with a database read each, on every read, is 12 dependent round trips. |
| Global dedup under our key | A chunk shared by two tenants can't be encrypted under both tenants' keys. If one revokes their key, the other's files must still open. |
| Delete means delete (after the GC) | A user deleting a held file, the 90-day history purge and the GC would all destroy evidence a court ordered kept. |
| One region | EU data would sit in a US region; a regional outage stops every tenant. |
| Sync applies every valid commit everywhere | A ransomware process produces perfectly valid commits. We spread them to every device within a second. |
| Application logs | Unstructured, mixed across tenants, deletable by us, not exportable per customer. |
R3.3 New Requirements and API Additions
Permissions on a folder
httpPUT /v1/nodes/n_board/permissions HTTP/1.1 Authorization: Bearer <admin or manager token> Content-Type: application/json { "base_acl_ver": 7, "inherit": false, "entries": [ { "principal": "group:board@acme.example", "role": "editor" }, { "principal": "user:cfo@acme.example", "role": "manager" } ], "link": null }
inherit: false makes n_board a restricted folder: permissions from its ancestors stop here. base_acl_ver makes the change conditional, like a commit: two admins editing the same folder's permissions can't overwrite each other.
| Role | Can |
|---|---|
| Viewer | List, open, download |
| Commenter | Viewer + comment |
| Editor | Commenter + create, edit, move and delete inside |
| Manager | Editor + change permissions, including restricting a folder |
Tenant key configuration
json{ "mode": "customer_key", "kms_key_arn": "arn:aws:kms:eu-central-1:444455556666:key/mrk-1a2b3c4d5e6f47a8b9c0d1e2f3a4b5c6", "replica_key_arn": "arn:aws:kms:eu-west-1:444455556666:key/mrk-1a2b3c4d5e6f47a8b9c0d1e2f3a4b5c6" }
The key lives in the customer's AWS account (444455556666) as a multi-Region key (the mrk- prefix), with a replica in the standby region. Their key policy grants our service roles GenerateDataKey and Decrypt. The replica in eu-west-1 is a separate key resource with its own state and its own policy, so cutting us off means disabling (or removing our grant from) both keys; our revocation check reads the key state in both regions. Multi-Region keys can't be created in an external key store, so tenants who keep keys outside AWS register a separate key in each region.
Legal holds
httpPOST /v1/admin/holds HTTP/1.1 Content-Type: application/json { "hold_id": "h_2291", "matter": "Case 2026-114", "scope": { "users": ["u_4410"], "shared_drives": ["ns_legal7"] } }
DELETE /v1/admin/holds/h_2291 releases it. Holds have no expiry: only a release ends one.
Point-in-time restore for a drive
json{ "scope": { "ns": "ns_fin" }, "to": "2026-09-28T08:55:00Z", "only_changes_by": { "device_id": "d_77" }, "dry_run": true }
POST /v1/admin/restores with dry_run: true returns counts (files to restore, files also changed by others since then, deletes to undo); without it, it starts the job.
An audit event
json{ "event_id": "ev_01J8Z6Q4M9", "tenant": "t_acme", "time": "2026-09-28T09:14:02.118Z", "actor": "user:bob@acme.example", "device": "d_4410", "ip": "203.0.113.24", "action": "download_url_issued", "node": "n_deck", "rev": 12, "decision": "allow", "via": "group:board@acme.example on n_board" }
Customers read theirs with GET /v1/admin/audit?from=...&to=...&actor=..., or receive daily files in their own S3 bucket.
Recap
- Roles with inheritance and restricted folders; permission changes are conditional on a version.
- Keys per tenant, optionally in the customer's own account, multi-Region within a jurisdiction.
- Holds have no expiry; restores are dry-run first; audit events are structured, per tenant, with the reason for each decision.
R3.4 Design Evolution: Organizations, Keys, Law and Attackers
Step 3.1: Who Can See This File?
The problem: Bob opens deck.pptx in Acme Shared / Finance / 2026 / Q3 / Board / ..., twelve folders deep. Permissions come from the shared drive, from Finance (the finance group can view) and from Board, a restricted folder that only the board group can see. This check runs on every open, download, list and commit: tens of thousands a second.
What would you do?
Step 3.2: Customers Want Their Own Keys
The problem: Acme's security team wants chunks encrypted under a key in their AWS account, so they can cut us off. But in Round 2 a popular chunk is stored once and shared by every tenant who has it. Whose key encrypts it? What would you do?
Step 3.3: A Legal Hold on an Employee's Files
The problem: Acme's legal team places a hold on Dana's files and the Legal-7 shared drive. Next week Dana deletes a folder and empties the trash, the 90-day history purge runs, and the GC sweeps unreferenced chunks.
What would you do?
Step 3.4: EU Data Stays in the EU
The problem: our European customers' contracts require their files, metadata, backups and logs to stay in the EU. Today everything lives in one US region. And a regional outage takes every tenant down together. What would you do?
Primitive: Cloud Disaster Recovery and Multi-Region Active-Active · Drill: The Booking That Existed in Frankfurt but Not in Virginia. Its first question is exactly the active-active collision above (two regions accepting writes to the same record during replication lag), which one owner region per tenant prevents; its second (why not replicate synchronously to every region) is answered by the latency of a cross-region round trip on every commit, by a commit failing whenever any region is down, and by multi-Region strong consistency's three-region, no-transaction limits.
Step 3.5: Ransomware Encrypted and Synced 50,000 Files
The problem: at 09:02 ransomware on Tom's laptop starts rewriting files in the Finance shared drive with encrypted versions, about 40 a second. Each is a valid commit from an authenticated device. Within a second each change reaches every member's devices, which faithfully download the encrypted versions.
What would you do?
Step 3.6: Every Access Must Be Auditable
The problem: Acme's security team asks: "Who downloaded the board deck in the last 90 days, from where, and why were they allowed?" Our application logs are text, mixed across tenants, sampled, and kept for 14 days. What would you do?
Primitive: Change Data Capture and the Outbox Pattern
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | Who can see this file? | In-memory tree per tenant, fed from the stream; revoke waits for every copy; fail closed | A stateful service; slower revokes |
| 3.2 | Customers' own keys | SSE-KMS per tenant, Bucket Keys, dedup inside a tenant, encrypted names, 1-day CDN cache | About $4.0M a month of lost dedup |
| 3.3 | Legal holds | A holds layer every deletion path checks; Object Lock legal hold as a second layer | Storage kept for the hold's life |
| 3.4 | EU data in the EU | Home + standby per jurisdiction; tenant map with owner epochs; failover from the standby | A second stack and copy |
| 3.5 | Ransomware | Detector on the journal; device holds; metadata-only rollback | 90-day history |
| 3.6 | Audit | Per-tenant events, batched through Firehose, Object Lock compliance mode | Volume |
R3.5 Global Architecture
Synthesizing vector architecture diagram...
Content never crosses the jurisdiction line; only the tenant map (IDs, regions, epochs) is global.
Synthesizing vector architecture diagram...
Each regional stack adds four things to Round 2: a permission service fed from the stream, override tables (holds and device holds), a detector, and an audit pipeline.
Tracing a shared-drive read with inherited permissions
Synthesizing vector architecture diagram...
The check never reads Finance's entries: the restricted Board folder cut inheritance, so the finance group can't see the deck even though it can see Finance.
Tracing a ransomware rollback (numbered steps)
- 09:02:10: Tom's laptop (
d_77) starts rewriting files inns_fin, about 40 a second. - 09:07:10: the detector's 5-minute window closes with 12,000 rewrites (40/s × 300 s), 96% of sampled chunks random-looking and mismatched. It writes a device hold for
d_77. By then, 12,000 files have been rewritten and pushed. - 09:07:11:
d_77's next commit gets423 Locked. The admin is paged. - 09:15: the admin runs a dry run to 09:02:05, limited to
d_77: 12,000 files to restore, 14 also edited by others since, 0 deletes. - 09:16: the job commits 12,000 restores in 600 transactions of 20 and puts the pre-attack versions of the 14 shared-edit files next to their current ones as copies. Two seconds of writes at 300 transactions a second.
- 09:16 onward: nudges go out; devices pull the rollback and fetch old chunks. The encrypted revisions remain in history for 90 days, then are purged normally.
(The detector here stopped the device after 12,000 files, not the interviewer's 50,000: a lower threshold stops it sooner but raises false positives. Either way the rollback is the same metadata operation.)
Tracing a legal hold (numbered steps)
- Legal places
h_2291on useru_4410(Dana) and shared drivens_legal7. The API writes the hold and an audit event. - A job marks Dana's owned files' revisions and all of
ns_legal7's asheld_by: [h_2291]and starts an S3 Batch Operations job that sets Object Lock legal hold on the chunk versions they reference. - Dana deletes
Contracts/and empties the trash. Her view no longer shows it; the purge job checks theholdstable itself (not onlyheld_by, which a new revision may not carry yet), findsh_2291covers Dana, and keeps the revisions. - The weekly mark finds those revisions, so their chunks stay
LIVE. Even a buggy sweep couldn't destroy them: its delete would only add a delete marker, and no one, lifecycle included, can remove a version under legal hold. - Months later, legal releases the hold. A job clears
held_byforh_2291, removes Object Lock legal holds from chunk versions no other hold references, and the normal purge then deletes what Dana deleted.
Deleting data really means a sequence of windows. When a user permanently deletes a file (with no hold), its bytes disappear from our systems only after: the trash (30 days, unless emptied), the 90-day version history, up to about 21 days of GC quarantine and marks, 7 days as a non-current S3 version, and the same in the standby region (which needs delete-marker replication turned on and its own lifecycle rule, or the standby copy would outlive the original); for US tenants, CDN edge caches can keep a chunk for up to their cache lifetime (1 day for customer-key tenants, otherwise up to a year) unless we invalidate it, although without a valid signed URL no one can fetch it; its name and metadata stay in DynamoDB point-in-time backups for up to 35 days, and audit events that mention it stay for the contract period. Anything we can restore from, we must count. When a customer deletes their KMS key (after a waiting period of 7 to 30 days), everything encrypted under it, in every copy and backup, becomes unreadable at once: that is the fastest real deletion we offer. Local copies on devices are outside all of this.
R3.6 Numbers and Cost
Tenants and data per jurisdiction (assumptions from the scope raise: 50 GB per seat including 90-day history, 1.6× dedup inside tenants, 9 commits per seat per day)
| US | EU | Total | |
|---|---|---|---|
| Seats | 13M (65%) | 7M (35%) | 20M |
| Logical data | 13M × 50 GB = 650 PB | 350 PB | 1 EB |
| Stored, ÷ 1.6 | 406.25 PB | 218.75 PB | 625 PB |
| Commits | 13M × 9 = 117M/day, 1,354/s | 63M/day, 729/s | 180M/day, 2,083/s; 6,250/s at a 3× peak |
| New chunk data (4 MB per commit) | 468 TB/day | 252 TB/day | 720 TB/day |
| Connected devices (2 per seat, 60% online) | 15.6M | 8.4M | 24M |
| Gateway tasks (50,000 each, per-AZ rounding) | 312 = 104 per AZ | 168 = 56 per AZ | 480 |
With global dedup (3.5×) the same 1 EB would be 286 PB. Tenant keys cost us 339 PB of storage.
Permission service memory. The largest tenant needs about 11.5 GB per copy (step 3.1). Across all tenants we assume about 100 folders per seat: 2B folders × 200 B = 400 GB, plus explicit entries, about 1 TB per copy, 3 TB with three copies, sharded by tenant over both jurisdictions.
KMS requests (per commit: 1 GenerateDataKey for the new chunk; Decrypts for downloads: in the US 2, since a 1-day CDN cache misses about half of 4 downloads, in the EU all 4, since EU chunks are served from S3 directly; 1 Decrypt and 1 GenerateDataKey for replication to the standby)
| Without Bucket Keys | Quota (default) | |
|---|---|---|
| us-east-1 | 117M × 4 = 468M/day = 5,417/s, 16,250/s at peak | 100,000/s |
| us-west-2 | 117M × 1 = 1,354/s | 100,000/s |
| eu-central-1 | 63M × 6 = 378M/day = 4,375/s, 13,125/s at peak | 20,000/s: 66% at peak |
| eu-west-1 | 63M × 1 = 729/s | 100,000/s |
| Cost | (468M + 117M + 378M + 63M) = 1,026M/day × 30.4 × $0.03/10,000 ≈ $94K a month |
With Bucket Keys, KMS requests fall by up to 99%, to under $1K a month, and eu-central-1's quota stops mattering. For tenants who opt out of Bucket Keys, their share counts against eu-central-1's 20,000 a second, which we watch and raise before it matters.
Audit volume. 20M seats × 60% active × 200 events a day = 2.4B events, about 1.2 TB a day raw, about 150 GB compressed; a year is about 55 TB, under Object Lock.
Failover timing (dependent steps add):
| Step | Time |
|---|---|
| Detection: the standby's canaries (every minute) fail twice; Route 53's 30-second alarm may come sooner but is best-effort | 2 min |
| Alarm, page, engineer online | 1 min |
| Confirm it's the region (canaries, several services, AWS Health) | up to 5 min |
| Flip the ARC routing control and the tenant map, from the standby | a few seconds |
| DNS TTL (60 s) and client reconnects with jitter (up to 60 s) | 2 min |
| Total | about 10 minutes of write downtime; reads of cached data continue on devices throughout |
RPO: metadata, usually about a second; chunks, 15 minutes for 99.99% of objects under Replication Time Control, and in practice recovered from the committing devices.
Monthly cost (list prices, decimal GB, 730 hours; us-east-1 and us-west-2 at us-east-1 prices; EU S3 Intelligent-Tiering at Frankfurt's prices, $0.0225 / $0.0135 / $0.005 per GB for the frequent, infrequent and archive instant tiers, which with the same 10/20/70 mix is $0.00845/GB, 14% above us-east-1; Glacier Instant Retrieval in eu-west-1 at $0.004; other EU lines at us-east-1 prices; check each region's prices in the calculator)
| Line | US | EU |
|---|---|---|
| S3 Intelligent-Tiering, home (Round 2's $0.0074/GB blend) + monitoring | 406.25M GB × $0.0074 ≈ $3,006K + 96.9B objects × $0.0025/1,000 ≈ $242K = $3,248K | 218.75M GB × $0.00845 ≈ $1,848K + $130K = $1,979K |
| Standby copy in S3 Glacier Instant Retrieval ($0.004/GB; 90-day minimum, 128 KB minimum object size, $0.03/GB to read back) | $1,625K | $875K |
Replication: transfer $0.02/GB, RTC $0.015/GB, replica PUTs $0.02/1,000 | 14.2 PB a month: $285K + $213K + 3.56B × $0.02/1,000 ≈ $71K = $569K | 7.66 PB: $153K + $115K + $38K = $306K |
| Other S3 requests | $28K | $15K |
| Downloads (4 × 4 MB per commit): CloudFront in the US; straight from S3 in the EU (residency) | 56.9 PB through CloudFront, tiered ≈ $1,177K + requests $14K = $1,191K | 30.6 PB of S3 data transfer out: 10 TB × $0.09 + 40 TB × $0.085 + 100 TB × $0.07 + the rest × $0.05 ≈ $1,536K; GETs ≈ $3K; regional caching proxies ≈ $30K = $1,569K |
| DynamoDB, home and replica | writes 117M × 12 × 30.4 × 2 × $0.625/M ≈ $53K; 40 TB × 2 × $0.25 + backups ≈ $28K; reads ≈ $10K = $91K | ≈ $48K |
| Compute: gateways, 20% warm standby, API, permission service, detector | $42K + $8K + $40K = $90K | ≈ $50K |
| KMS: a key per tenant and its replica ($1 each; keys held in customers' own accounts are billed to them, so this is an upper bound), requests with Bucket Keys | 26,000 keys + requests = $27K | 14,000 keys = $15K |
| Audit (Firehose, dynamic partitioning, storage) | $3K | $2K |
| WAF, load balancers, Valkey, logs, CloudWatch | estimate $100K | estimate $60K |
| Region pair total | ≈ $6.97M | ≈ $4.92M |
| Enterprise edition total | ≈ $11.9M/month |
About $0.59 per seat a month. The EU pair costs 70% as much as the US pair for 54% as many seats, because residency takes away the CDN's cheaper egress. The two lines architects argue about: tenant-scoped dedup costs about $4.0M a month (220.5 PB extra in the US × $0.0074 ≈ $1.63M and 118.75 PB in the EU × $0.00845 ≈ $1.00M at home, plus the standby copies, 339 PB × $0.004 ≈ $1.36M), 34% of the bill, and the standby copy plus replication costs about $3.4M a month, the price of surviving a region inside each jurisdiction. Everything else (compute, KMS, audit, metadata) is under 5%.
R3.7 Trade-Offs
Dedup vs tenant keys
| Global dedup, our key | Per-tenant dedup, tenant keys (chosen for enterprise) | Client-side encryption | |
|---|---|---|---|
| Storage for 1 EB | 286 PB | 625 PB | Up to 1 EB (no dedup at all) |
| Customer can revoke | No | Yes, and sees every use (fewer with Bucket Keys) | Yes; we never hold the key |
| Server features (preview, search, scanning) | Yes | Yes | No |
| Existence side channel across tenants | Controlled by the threshold | None across tenants | None |
The consumer service keeps global dedup (Round 2); the enterprise edition pays for isolation.
Precomputed permissions vs check on read
| Effective ACL stored on every node | Walk the database per check | In-memory tree per tenant (chosen) | |
|---|---|---|---|
| Check cost | 1 read | Up to 12 dependent reads (≈ 60 ms) | Microseconds |
| Change cost | Rewrite the whole subtree (millions of items) | 1 write | 1 write + stream apply |
| Revocation | After the rewrite finishes | Immediate | Immediate once every copy acknowledges (≤ 2 s) |
| Risk | A half-finished rewrite | Latency | A stateful service to rebuild and shard |
Retention vs cost. Each day of version history keeps about one day of new chunk data: 720 TB. At home Intelligent-Tiering plus the standby copy, about $0.0118/GB a month (the US and EU blends weighted by their new data):
| History | Kept | Monthly |
|---|---|---|
| 30 days | 21.6 PB | ≈ $0.25M |
| 90 days (chosen) | 64.8 PB | ≈ $0.76M |
| 365 days | 263 PB | ≈ $3.1M |
90 days is our choice because ransomware and accidental deletes are often found weeks later; tenants can buy longer.
Closing the loop. Round 1 answered "how do devices agree?" with a journal, a cursor and conditional commits. Round 2 kept exactly those and changed what a chunk is, who stores it and how the news travels. Round 3 added organizations, keys, laws and attackers, and still every change in the system, a restore, a rollback, a move, a revoked permission, is a commit on a base revision, recorded in a journal that every device and every index reads in order. One ordered log per namespace, conditional commits on top of immutable chunks is the invariant that survived all three rounds.
R3.8 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| A KMS outage or a revoked key | Decrypt fails; S3 returns errors for that tenant's objects | Reads that need KMS fail closed: files already on devices keep working, a US tenant's cached CDN copies expire within a day, uploads queue on clients. We never fall back to an unencrypted path. For a revoked key, that's the intended result; we check both regions' key state, invalidate a US tenant's CDN paths and tell its admins. With Bucket Keys, S3 may keep using a cached bucket-level key for a short while, so revocation can take effect slightly later. |
| A permission cache stale after a revoke | A copy of the permission service is behind the stream | The revoke waits for every copy to acknowledge the new acl_ver; a copy that can't within 2 s stops serving that tenant until it catches up. Download URLs expire in 5 minutes. |
| A region outage | The standby's canaries fail against a home region | After confirmation, the ARC routing control and the tenant map flip together from the standby, about 10 minutes in all; until then the standby forwards writes home; revisions with unreplicated chunks are marked pending and re-uploaded by their devices; a reconciliation report after failback. |
| A GC bug near held data | The sweep tries to delete chunks that held revisions reference | Three layers stop it: held revisions keep their chunks marked; the sweep's delete only adds a delete marker, and the held version can't be removed by lifecycle or by anyone else; other deleted versions survive 7 days as non-current versions. Before deleting, the sweeper HEADs each candidate and checks x-amz-object-lock-legal-hold; any candidate with the hold ON is skipped and pages us, because it means a bug. |
| A ransomware false positive | A user's legitimate bulk encryption triggers a device hold | The device's commits stop; nothing is rolled back without an admin. The admin clears the hold, and the tenant's threshold can be tuned. |
| eu-central-1 KMS throttling | ThrottlingException from KMS during a peak | Bucket Keys on by default; tenants who opt out are budgeted; the client retries with jitter; we raise the quota ahead of growth. |
R3.9 Runbook and Incident Response
Golden signals, per tenant and per region OPS 8 · REL 6
| Signal | Alarm | Severity | First action |
|---|---|---|---|
| Commit P95 | > 100 ms for 10 min | P2 | DynamoDB throttles or conflicts on hot namespaces; S3 HEAD latency |
| Push lag (commit to nudge sent) | P95 > 500 ms for 10 min; stream iterator age > 30 s | P2 | Router health; backstop Lambda errors |
| Conflicted-copy rate | above 2× the same hour last week | P3 | A client release that sends stale bases? |
| Dedup hit rate | chunks per commit found existing, down by a third | P3 | A client chunking change (parameters must never change silently) |
| GC backlog | DOOMED count or quarantine age above plan; the mark didn't finish | P3 | Scale the batch jobs; never shorten the quarantine to catch up |
Sweep candidates under legal hold (the sweeper's HEAD shows x-amz-object-lock-legal-hold: ON) | any | P1 | Pause GC and purges; find the path that ignored the holds table |
| CDN hit ratio | drops by a third | P3 | Cache policy changed? Query strings in the cache key? |
| Mass-change alerts | any device hold created | P2 | Confirm with the tenant's admin; prepare the dry-run restore |
| KMS throttles | any sustained ThrottlingException | P2 | Which tenants? Bucket Keys off? Request a quota increase |
| Replication backlog | S3 replication latency > 15 min, or global table lag > 5 s | P2 | The failover RPO is growing |
Ransomware rollback procedure SEC 10
- Confirm the device hold is in place (the detector writes it; command 2 writes it by hand).
- With the tenant's admin, choose the rollback time: the detector's first suspicious commit minus a few seconds.
- Run the restore as a dry run; review files also edited by others since (they'll get copies, not overwrites).
- Run it. Watch the rollback's transactions and the devices' pull traffic.
- The device is cleaned or reimaged by the customer; only then does the admin clear its hold.
GC pause procedure (a suspected GC bug) REL 9
- Set the pause flag (command 3); mark, doom and sweep jobs check it before each batch.
- Count what the last sweep deleted; non-current versions from the last 7 days can be restored by removing their delete markers.
- Fix, re-run a mark in dry-run mode, compare, and only then clear the flag.
Regional failover procedure REL 13
- Confirm it's the region: the standby's canaries, several services and AWS Health agree; Route 53's health-check status (command 7) is a hint only, since its API may be unavailable when us-east-1 is impaired.
- Check replica state and chunk replication lag (commands 4 and 5) to estimate what may be pending.
- From the standby region, flip the ARC routing control and each affected tenant's owner with a new epoch (command 8). Confirm commits succeed in the standby.
- Start the "pending chunks" job: it asks committing devices to re-upload missing chunks.
- Fail back later, tenant by tenant at quiet hours, after replication catches up; run the reconciliation report.
Go deeper: CLI playbook
Plain commands an on-call engineer runs one at a time. Replace names, times and IDs with real ones.
text# 1. Alarms firing for the sync service in a region aws cloudwatch describe-alarms --region eu-central-1 --state-value ALARM --alarm-name-prefix drive- # 2. Put a device hold by hand aws dynamodb put-item --region eu-central-1 --table-name device_holds --item '{"device_id":{"S":"d_77"},"tenant_id":{"S":"t_acme"},"reason":{"S":"mass encryption"},"created_at":{"S":"2026-09-28T09:07:10Z"}}' # 3. Pause garbage collection aws ssm put-parameter --region eu-central-1 --name /drive/gc/paused --value true --type String --overwrite # 4. Replica status of the nodes table, seen from the standby aws dynamodb describe-table --region eu-west-1 --table-name nodes --query "Table.Replicas" # 5. Chunk replication latency over the last hour aws cloudwatch get-metric-statistics --region eu-central-1 --namespace AWS/S3 --metric-name ReplicationLatency --dimensions Name=SourceBucket,Value=drive-chunks-euc1 Name=DestinationBucket,Value=drive-chunks-euw1 Name=RuleId,Value=all-chunks --start-time 2026-09-28T08:00:00Z --end-time 2026-09-28T09:00:00Z --period 300 --statistics Maximum # 6. Legal hold on one chunk version by hand aws s3api put-object-legal-hold --bucket drive-chunks-euc1 --key t/t_acme/9c1f/9c1f...e2.1 --legal-hold Status=ON # 7. Health check status for a home region's endpoint (a hint only: may fail when us-east-1 is impaired) aws route53 get-health-check-status --health-check-id 0a1b2c3d-4e5f-6789-abcd-ef0123456789 # 8. Flip a tenant to the standby with a new owner epoch, run in the standby region aws dynamodb update-item --region eu-west-1 --table-name tenant-map --key '{"tenant_id":{"S":"t_acme"}}' --update-expression "SET owner_region = :r, owner_epoch = owner_epoch + :one" --condition-expression "owner_region = :old" --expression-attribute-values '{":r":{"S":"eu-west-1"},":one":{"N":"1"},":old":{"S":"eu-central-1"}}' # 9. State of a tenant's key aws kms describe-key --region eu-central-1 --key-id arn:aws:kms:eu-central-1:444455556666:key/mrk-1a2b3c4d5e6f47a8b9c0d1e2f3a4b5c6 --query "KeyMetadata.KeyState" # 10. Drop a US tenant's cached chunks from the CDN after a key revocation (EU tenants aren't on the CDN) aws cloudfront create-invalidation --distribution-id E2DRIVEEXAMPLE --paths "/t/t_globex/*"
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | A standby per jurisdiction with replicated metadata, chunks and keys; failover run from the standby; owner epochs; pending-chunk recovery from devices; drive rollback from history REL 9 · REL 10 · REL 13 |
| Security | Permission checks that fail closed; per-tenant keys the customer can revoke; encrypted names; a ransomware detector with device holds; tamper-proof audit logs SEC 3 · SEC 4 · SEC 8 · SEC 10 |
| Performance Efficiency | Permission checks in memory instead of 12 reads; tenants served from their own jurisdiction PERF 1 · PERF 4 |
| Cost Optimization | About $11.9M a month, $0.59 per seat; the cost of isolation ($4.0M), of region survival ($3.4M) and of EU residency without a CDN shown line by line; the standby copy in a cheaper class; retention priced by days COST 5 · COST 8 |
| Operational Excellence | Golden signals per tenant; runbooks for rollback, GC pause and failover; residency and hold requirements written into the design OPS 1 · OPS 10 |
| Sustainability | Data kept in the users' own jurisdiction; history and holds sized by policy, not "forever"; the standby copy in a cold class SUS 1 · SUS 4 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Keeps permission checks fast without rewriting subtrees, and makes revocation fail closed.
- Explains exactly what customer-held keys promise (revocable, visible) and what they cost (dedup across tenants), and why convergent encryption isn't the answer.
- Treats legal holds and device holds as override layers that every deletion or commit path consults, with an independent second layer.
- Designs residency as home and standby within a jurisdiction, one writer per tenant, and failover driven from the standby; names the split-brain window and how it's reconciled.
- Rolls back ransomware through metadata, not bytes, because chunks are immutable and history is kept.
- States deletion honestly, as a list of windows including backups, caches and devices.
Follow-up questions
-
"A customer revokes their key in the middle of the day. What exactly happens?" Answer: once the customer has disabled both the primary key and its replica in the standby region (each has its own state; disabling only one leaves the other region readable), S3 can no longer decrypt their objects: downloads fail, uploads fail. That's usually quick, but with Bucket Keys S3 may keep using a cached bucket-level key for a short while, so it isn't instant. For a US tenant, CDN copies live up to a day unless we invalidate them, which we do. Devices keep their local files. Metadata fields we encrypted with the tenant's data key become unreadable once our 5-minute cache of the data key expires. If they re-enable the key, everything works again; if they delete it after the waiting period, the data is gone for good.
-
"Why not detect ransomware on the device instead?" Answer: we should, too, but the device is the thing that's compromised, and it can be told to lie. The server sees every commit whatever the device claims, so the server-side detector and device holds are the backstop that can't be switched off by the attacker.
-
"A 400,000-seat tenant wants its own isolated stack." Answer: the design already partitions by tenant everywhere (keys, chunk prefixes, permission shards, audit prefixes), so a dedicated cell is the same stack with one tenant in its tenant map entry. It costs a fixed overhead per cell, which is why it's a paid option, not the default.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Store every file's effective permissions" | One membership change rewrites millions of items. |
| "Convergent encryption keeps dedup and privacy" | Equal ciphertexts leak equal contents; guessable files can be confirmed. |
| "Copy held files to a vault" | A snapshot misses later changes and doubles storage. |
| "Active-active for every tenant" | Last writer wins and local transactions let two regions both commit. |
| "Restore from the infected laptop" | Its copies are the encrypted ones. |
| "Deleted means gone" | Trash, history, quarantine, versions, replicas and backups each keep a copy for a while. |
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scoping: file sizes, devices, offline edits, history, "within a minute" | Restate Round 1 in 60 seconds | Restate Round 2 in 60 seconds |
| 5–15 min | Requirements and API: upload negotiation, commit with base revision and commit_id, changes(cursor) | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Steps 1.0–1.6: chunks → namespace vs bytes → only new hashes → journal and cursor → conflicted copies → soft delete and ref counts | Steps 2.1–2.6: CDC with the worked example → global dedup, proofs and the threshold → push with a backstop → CDN herd → moves and write skew → safe GC | Steps 3.1–3.6: permission service → tenant keys → legal holds → residency and failover → ransomware rollback → audit |
| 40–50 min | Numbers, cost, fixed vs CDC, poll vs push | Storage and dedup, Bloom arithmetic, write units, gateways per AZ, budgets, cost | Per-jurisdiction sizing, KMS rates, cost of isolation and survival, retention |
| 50–60 min | Failures and pillar check | Failures, gotchas, pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint.
The Two Sentences That Matter Most
- Opening any round: "The file tree is metadata that must be exactly right, so every change is a conditional commit on a base revision, logged in an ordered journal that devices follow with a cursor; the bytes are immutable chunks named by their hash, uploaded first and never overwritten, and when two edits collide we keep both as a conflicted copy."
- When scale arrives: "I'll cut chunks by content so edits move one chunk, store each chunk once with proofs before anyone skips an upload, push tiny nudges and let devices pull, serve chunks from a CDN because they never change, and delete a chunk only after it has been unreferenced across two marks."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "What if two devices edit the same file offline?" (REL 4) | The second commit's base is stale, so it gets 409 and is saved as a conflicted copy; nothing is merged or lost. | 1 | Step 1.5 |
| "What if the GC deletes a chunk someone just referenced?" (REL 9) | It can't: commits resurrect doomed chunks, deletion waits for a second mark, keys carry generations, and deletes leave recoverable versions. | 2 | Step 2.6 | |
| "What if a region fails?" (REL 13) | Each tenant fails over to its standby in the same jurisdiction, with DNS and ownership flipped together from the standby itself, in about 10 minutes. | 3 | Step 3.4 | |
| Performance | "How do you avoid re-uploading big files?" (PERF 3) | Content-defined chunks: an edit changes one or two chunks, so we send about 4 MB of a 40 MB file. | 2 | Step 2.1 |
| "10,000 people download the same file." (PERF 4) | Immutable chunks cached at the edge with Origin Shield, jittered starts, and LAN sync for the office link. | 2 | Step 2.4 | |
| Security | "Can someone get a file by knowing its hash?" (SEC 3) | No: skipping an upload needs a proof of possession, and only after a random number of users uploaded the chunk. | 2 | Step 2.2 |
| "Can customers hold their own keys?" (SEC 8) | Yes, per-tenant SSE-KMS keys in their account; revocable and visible, at the price of dedup across tenants. | 3 | Step 3.2 | |
| "How would you notice ransomware?" (SEC 4) | A detector on the journal stream flags mass rewrites of random-looking data and holds the device; rollback is metadata-only. | 3 | Step 3.5 | |
| Cost | "What does it cost?" (COST 5) | About $21.4K, $22.7M and $11.9M a month; storage dominates every round, so dedup and tiering are the levers. | 1–3 | R1.7, R2.6, R3.6 |
| "Why not a Bloom filter in front of the chunk index?" (COST 5) | About $19.4K a month of cache to save about $5.7K of reads; the arithmetic says no. | 2 | Step 2.2 | |
| Operations | "How do you know sync is healthy?" (OPS 8) | Golden signals per tenant and region: commit P95, push lag, conflicted copies, dedup hit rate, GC backlog, replication lag. | 3 | R3.9 |
| Sustainability | "Is this wasteful?" (SUS 4) | Each byte is stored once per scope, edits move deltas, cold data sits in cheaper tiers, and history is sized by policy. | 2–3 | R2.10, R3.7 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| Chunking | Fixed chunks, only new hashes uploaded | Content-defined (Gear-based FastCDC), sizes justified | Chunk keys scoped per tenant |
| Dedup | Within one user | Across users, proofs of possession, side channel handled | Within a tenant only, and why |
| Metadata | One transactional database, namespace lock, journal | DynamoDB transactions per namespace; moves by parent ID with cycle checks | Encrypted names; home and standby per jurisdiction |
| Change propagation | Polling a cursor | Push nudges, stream backstop, re-sent on retries | Permission and ransomware services fed from the same streams |
| Conflicts | Conflicted copies, never mtime | Conflicted copies at scale; no binary merges | Rollbacks that keep other users' later edits |
| Deletion | Soft delete, history, exact ref counts | Mark, doom, resurrect, sweep, generations | Holds as a layer, Object Lock as a second, deletion windows stated honestly |
| Evolving under new scope | Builds from whole-file uploads step by step | Opens with "what breaks", fixes chunking first | Adds organizations, keys, law and attackers without breaking conditional commits on an ordered log |