Design a Mobile Chat Client Architecture
This page is one interview loop in three rounds. All three rounds design the same system. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
In this loop the phone is part of the system. It has its own database, a radio that wants to sleep, an operating system that freezes and kills apps, and (in Round 3) the only copy of the encryption keys. Every round designs both halves: the app on the device and the backend it talks to. How the server orders messages and fans them out to group members is the subject of the chat and instant messaging loop; here we link to it instead of teaching it again. The local database basics (WAL mode, background schedulers, schema migrations) are taught in the offline-first news feed loop.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | A marketplace app's buyer–seller chat that works offline | A consumer messenger: receipts, push, photos and videos, battery | End-to-end encrypted, several devices per user, restorable on a new phone |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Traffic | 1M DAU; 20M messages/day (≈ 810/s at peak); ~100K open sockets | 30M DAU; 500M messages/day (≈ 20K/s at peak); 10M open sockets | 100M DAU; 1.7B messages/day (≈ 69K/s at peak); ≈ 7.5 stored envelopes per message |
| On the device | Recent chats in SQLite; an outbox | + a push handler, an upload manager, a receipt batcher; < 1.5% battery a day for messaging | + a crypto engine, keys in secure hardware, an encrypted database, encrypted backups |
| Targets | A sent message appears at once; none lost or doubled | Send-to-display < 100 ms on good networks; reconnect < 300 ms P99 on good networks | Keys never leave devices unprotected; a new phone restores its history |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
Loop Opener: What Makes Mobile Chat Hard?
You Already Use One: a Walkie-Talkie That Keeps Falling Asleep
Imagine chatting over a walkie-talkie that switches itself off every few seconds to save its battery. When it's off, it can't hear anything. When you pick it up, it takes a moment to come back, and it may have moved to a different channel. Yet the people you talk to expect every word to arrive, in order, and nothing said twice.
A phone is that walkie-talkie. Its radio sleeps to save battery, the operating system freezes apps you aren't looking at and kills them when it needs memory, and you walk from Wi-Fi to 5G in the middle of a sentence. A chat app still has to feel instant.
| You do | The app does |
|---|---|
| Type "Is the bike still for sale?" and tap send with no signal | Shows the message at once with a clock icon, saves it on the phone, and sends it when the network returns |
| Put the phone in your pocket | Closes its connection so the radio can sleep; the OS's push service will wake it |
| Get a reply while the app is closed | A notification appears; opening the app shows the reply already in the chat |
| Walk out of the house onto 5G | Notices the network change, reconnects in a fraction of a second, and fetches anything missed |
What Makes It Hard
- Keeping a connection costs battery; dropping it costs latency. A connection held open in the background wakes the radio for every keep-alive packet. A closed connection means the next message has to find another way in.
- The OS is in charge. iOS and Android decide when a backgrounded app may run, for how long, and whether a "wake up" push is delivered at all.
- Messages must never be lost or shown twice, even when the app is killed halfway through sending one, or a retry races the original.
- The device may be the only place the data exists. An unsent message in the outbox, and (in Round 3) the keys that decrypt everything, live only on the phone.
The Question the Whole Loop Answers
How do we make chat feel live on a device that keeps going to sleep, losing its network and getting killed?
The answer gets sharper every round:
- Round 1: the local database is the truth for the screen. A message is written locally first and sent from an outbox; the network only fills and drains that database.
- Round 2: a live socket only while the app is in the foreground, OS push notifications for everything else, media on a separate resumable path, and receipts as cursors, all inside a battery budget.
- Round 3: the keys live on the devices, not on our servers. Every device is a separate encrypted endpoint, the server carries sealed envelopes, and restoring a new phone becomes a key-management problem.
Round 1 · Mid-level · "A Chat Screen That Works Offline"
~35 min · SDE II (L5) · 1 region, 3 AZs · 1M DAU · 20M messages/day, ≈ 810/s peak · ~100K open sockets · 99.9%
R1.1 Establish Design Scope
The interviewer says: "Our second-hand marketplace has an in-app chat between buyers and sellers. Users complain that messages take seconds to appear, disappear when they're on the train, and sometimes show up twice. Design the chat on the phone, and what it needs from the backend." We ask before we draw, and we say what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| One-to-one, or groups too? | One-to-one: a buyer and a seller. | Each conversation has two members; one recipient per message. |
| Can users send while offline? | Yes. Queue it and send it later. | We need a durable outbox on the phone (step 1.1). |
| How much history lives on the phone? | Recent chats, enough to read on the train. | A local database with a size limit; older history is fetched from the server on scroll. |
| Photos or videos? | Not yet. | Text frames only. Media gets its own path in Round 2. |
| Receipts ("delivered", "read")? | Just "sent" for now. | One tick means "the server has it". Delivered and read come in Round 2. |
| Encryption? | TLS to our servers is enough for now. | The server can read messages. End-to-end encryption is Round 3. |
| How many users? | About 1M daily users. | Small numbers; we derive them in R1.7. |
Out of scope for this round:
- Receipts beyond "sent".
- Push notifications when the app is in the background.
- Media.
- End-to-end encryption and multiple devices.
The interviewer will widen this scope later. Write your out-of-scope list where you can see it: in a multi-round loop, it all comes back.
R1.2 Functional Requirements, Derived Step by Step
We read the problem one phrase at a time and turn each phrase into a requirement:
| Phrase from the problem | Requirement |
|---|---|
| "Messages take seconds to appear" | Send optimistically: the sender's own message appears the instant they tap send, before the server replies. |
| "Disappear on the train" | Durable offline send: a message written offline survives the app being killed, and goes out when the network returns. |
| "Show up twice" | Exactly once in effect: retries never create a second copy. |
| "Show the conversation" | Local history: the conversation list and each chat render from storage on the phone. |
| "Receive a reply" | Live receive while the app is open, and catch-up of everything missed when the app opens. |
Not yet: delivered and read receipts, push notifications, media, end-to-end encryption.
R1.3 Non-Functional Requirements: the Questions
Numbers come in R1.7. For now, the questions and why each matters:
- Instant local display. How long from tapping send to the message on screen? It can't depend on the network at all, so it must come from a local write.
- No lost messages. If the app is killed between "tap" and "sent", is the message still there on the next launch?
- No duplicates. If the network drops after the server stored a message but before the phone heard back, the phone will retry. Does the other person see it twice?
- Order. Do both people see the conversation in the same order, even if their phones' clocks disagree?
- Battery. How long do we keep a connection open, and what does it cost the phone when the app is in the background?
R1.4 The API
The app speaks two protocols: a WebSocket for live traffic while the app is open, and plain HTTPS for catch-up and history.
A WebSocket is a connection that starts as an HTTP request and then "upgrades" into a two-way channel over the same TCP connection. After the upgrade, either side can send a message (a frame) at any time.
The backend protocol is the one in the chat loop's Round 1 API: the app gets a short-lived connect ticket over HTTPS, then opens the socket with it. The frames that matter to the phone are three:
SEND_MESSAGE (phone → server), JSON in Round 1:
| Field | Type | Example | Why |
|---|---|---|---|
type | string | "SEND_MESSAGE" | Frame kind |
client_msg_id | UUID string | "9b1deb4d-3b7d-4bad-9bdd-2b0d7b3dcb6d" | Generated on the phone before the first attempt; the same on every retry (step 1.2) |
conversation_id | string | "c_7Hk2" | Which chat |
body.text | string | "Is the bike still for sale?" | The message |
ACK (server → phone), sent only after the message is stored:
| Field | Type | Example | Why |
|---|---|---|---|
client_msg_id | UUID string | "9b1deb4d-…" | Lets the phone find its pending row |
server_msg_id | string | "718293847561029384" | The server's name for the message |
seq | integer | 10482 | The per-conversation sequence number: the message's place in the order |
sent_at | ISO-8601 string | "2026-09-26T18:04:11.135Z" | Server time, for display |
RECEIVE_MESSAGE (server → phone): conversation_id, server_msg_id, client_msg_id (the sender's), seq, sender_id, body, sent_at.
Catch-up and history over HTTPS:
httpGET /v1/conversations/c_7Hk2/messages?after_seq=10477&limit=50 HTTP/1.1 Host: chat.example-market.com Authorization: Bearer <access_token>
httpHTTP/1.1 200 OK Content-Type: application/json { "messages": [ { "seq": 10478, "server_msg_id": "718293847561029300", "client_msg_id": "1f0c…", "sender_id": "u_802", "body": { "text": "Yes, still available" }, "sent_at": "2026-09-26T18:01:02.410Z" } ], "has_more": false }
after_seq reads forward from where the phone is (catch-up); before_seq reads backward (scrolling up into older history). And one call to learn which conversations changed:
httpGET /v1/inbox?changed_since=inb_8Kq2xW HTTP/1.1 Authorization: Bearer <access_token>
It returns each changed conversation with its newest seq and a new sync_token to send next time. The token is opaque: the server builds it from its own change positions, not from a millisecond timestamp, because two changes written at nearly the same moment can become visible out of order and a pure time cursor would skip the later-committed one. (If a timestamp is used, the server re-reads with an overlap margin of a few seconds, and the phone deduplicates.)
| Status | When |
|---|---|
401 Unauthorized | Missing or expired token or ticket |
403 Forbidden | The caller isn't a member of the conversation |
429 Too Many Requests | Over the send or read limit (with Retry-After) |
The server never trusts identity from a frame: sender_id is taken from the authenticated connection, and every send and read checks that the caller is a member of conversation_id.
Recap
- A WebSocket while the app is open; HTTPS for catch-up and history.
SEND_MESSAGEcarries a phone-generatedclient_msg_id;ACKmaps it to the server'sseq.- Catch-up by a per-conversation cursor,
after_seq, plus an inbox call for "what changed".
Let's build it, starting with the simplest thing that works.
R1.5 Design Evolution: From "Send and Wait" to a Local-First Chat
Every step below follows the same pattern: a problem, your turn to think, the answer, and what the answer costs us. The cost is always the next problem.
Step 1.0: The Baseline
The chat screen calls POST /v1/messages when the user taps send, waits for the response, and then shows the message. To see new messages, it asks the server again when the screen opens.
Synthesizing vector architecture diagram...
Everything the screen shows comes from a network call. If the call is slow or fails, the screen has nothing.
What's good about it: it's simple, and the server is the only copy of anything. What it costs us: on a weak signal the user stares at a spinner after every send; with no signal the message is simply gone; and the screen is blank on the train.
Step 1.1: The Message Only Appears After the Server Replies
The problem: on a train with one bar of signal, the round trip takes two seconds, so each message appears two seconds after the user taps send. With no signal, the request fails and the text the user typed is lost. What would you do?
Primitive: Change Data Capture & the Outbox Pattern · Primitive: Write-Ahead Log & LSM Trees
Step 1.2: A Retry Sent the Message Twice
The problem: the phone sends "Can you do $80?". The server stores it, but the ACK is lost when the train enters a tunnel. The outbox worker times out and sends it again. The seller now sees "Can you do $80?" twice.
What would you do?
Primitive: Distributed Unique ID Generators
Step 1.3: Replies Arrive, but Slowly
The problem: while both people have the chat open, a reply only shows up when the screen re-fetches. We could poll every 3 seconds. What should the phone hold open, and when? What would you do?
Primitive: WebSocket, SSE & Long Polling
Step 1.4: I Opened the App After a Day: What Did I Miss?
The problem: the buyer opens the app after a day away. Three sellers replied, one of them twice. The chat list must show the right conversations and unread counts, and the first chat they open must be complete. What would you do?
Step 1.5: Messages Show Out of Order
The problem: the buyer's phone clock is 4 minutes fast. They reply to the seller's "Yes, still available" and their reply appears above it. On the seller's phone, the order is right. And a message sent while offline shows between two older ones. What would you do?
Primitive: Distributed Unique ID Generators
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | Send over HTTP; show on reply | Spinners; lost messages offline |
| 1.1 | Appears only after the reply | Optimistic insert into SQLite (WAL, one writer); SENDING rows are the outbox | Must handle retries and failures |
| 1.2 | Retry sent it twice | Client message ID created before the first attempt; server dedup; one in flight per conversation | An idempotency record per message |
| 1.3 | Replies are slow | WebSocket in the foreground only; closed in the background | Connection management; a gap after every background period |
| 1.4 | What did I miss? | Inbox token + per-conversation after_seq cursor, one transaction per page; outbox reconciled against catch-up | An extra request per changed chat |
| 1.5 | Wrong order | Sort by server seq; pending last; gap fetch; server times and a measured clock offset | A pending message may move once |
R1.6 Architecture v1
Now the concepts get names on both sides.
Synthesizing vector architecture diagram...
On the phone, every arrow into the database goes through one writer, and the screen only ever reads. The outbox worker sends through the socket manager, and runs catch-up over HTTPS before it drains. On the server, a message is stored with its sequence number before it is acknowledged; the backend details are in the chat loop's Architecture v1.
The local schema (SQLite, SQL shown as a data model):
sqlCREATE TABLE conversation ( conversation_id TEXT PRIMARY KEY, peer_user_id TEXT NOT NULL, title TEXT, newest_seq INTEGER NOT NULL DEFAULT 0, -- newest seq the server has told us about synced_seq INTEGER NOT NULL DEFAULT 0, -- we hold every message up to here last_read_seq INTEGER NOT NULL DEFAULT 0, -- unread = newest_seq - last_read_seq; raised on own sends last_activity_ms INTEGER NOT NULL -- server time, for the chat list order ); CREATE TABLE message ( local_id INTEGER PRIMARY KEY, -- order the user wrote pending messages conversation_id TEXT NOT NULL REFERENCES conversation(conversation_id), client_msg_id TEXT NOT NULL, -- created before the first send attempt server_msg_id TEXT UNIQUE, -- NULL until acknowledged seq INTEGER, -- NULL until acknowledged sender_id TEXT NOT NULL, body TEXT NOT NULL, state TEXT NOT NULL CHECK (state IN ('SENDING','SENT','FAILED','RECEIVED')), created_local_ms INTEGER NOT NULL, -- phone clock: display of pending rows only sent_at_ms INTEGER, -- server clock attempts INTEGER NOT NULL DEFAULT 0, next_attempt_ms INTEGER ); CREATE UNIQUE INDEX idx_message_order ON message(conversation_id, seq); CREATE UNIQUE INDEX idx_message_client ON message(conversation_id, sender_id, client_msg_id); CREATE INDEX idx_outbox ON message(conversation_id, local_id) WHERE state = 'SENDING'; CREATE TABLE sync_state ( id INTEGER PRIMARY KEY CHECK (id = 1), inbox_token TEXT, -- changed_since for GET /v1/inbox clock_offset_ms INTEGER NOT NULL DEFAULT 0 -- server time minus phone time );
A unique index in SQLite allows many rows with a NULL seq, which is exactly what pending messages need. The partial index makes "next pending message in this conversation" a cheap lookup.
The message's life on the phone:
Synthesizing vector architecture diagram...
A message leaves SENDING only when the server has proved it has the message: an ACK, or the message appearing in catch-up. Retry keeps the same client_msg_id, so a retry can never make a second copy.
Trace 1: sent offline, then reconnect
Synthesizing vector architecture diagram...
Catch-up runs before the outbox drains, so a message the server already had would have been matched instead of resent.
Trace 2: a lost ACK and a retry
Synthesizing vector architecture diagram...
The server stores the message once; the phone gets the same answer twice.
R1.7 Numbers
Traffic (assumptions: each daily user sends 20 messages a day; the evening peak is 3.5× the daily average; 10% of daily users have the app open at peak)
| Quantity | Math | Value |
|---|---|---|
| Messages per day | 1M × 20 | 20M |
| Messages/s, average | 20,000,000 ÷ 86,400 | ≈ 231 |
| Messages/s, peak | 231 × 3.5 | ≈ 810 |
| Open sockets at peak | 1M × 10% | 100K (foreground apps only) |
| Heartbeats/s at peak | 100,000 ÷ 30 s | ≈ 3,333 |
| Registry refreshes/s (the gateway refreshes a user's entry on each heartbeat) | same | ≈ 3,333: a small fraction of one cache node |
| App opens with catch-up (assume 10 per user per day) | 10M ÷ 86,400 × 3.5 | ≈ 405/s at peak |
On the phone (a row is about 350 bytes: a ~60-byte text, three 36-character IDs, a numeric server ID, a few integers, SQLite's record overhead, and two index entries)
| Quantity | Math | Value |
|---|---|---|
| Per message | ≈ 240 B row + ≈ 110 B of index entries | ≈ 350 B |
| A typical user (5 active chats × 100 messages) | 500 × 350 B | ≈ 175 KB |
| Our cap: newest 1,000 messages in each of the 50 most recent chats | 50,000 × 350 B | ≈ 17.5 MB |
| WAL file between checkpoints | 1,000 pages × 4 KB (SQLite's default checkpoint threshold) | ≈ 4 MB |
Older messages are deleted locally in a background job and fetched again with before_seq when the user scrolls that far. Pending (SENDING, FAILED) rows are never evicted.
Monthly cost. The backend is the chat loop's Round 1 system at exactly this traffic (1M DAU, 20M messages a day, 100K sockets), so we reuse its derivation, minus the push notifications we don't send yet (us-east-1 list prices, 730 hours a month; rounded):
| Item | Math (from the chat loop's R1.7) | Monthly |
|---|---|---|
| DynamoDB writes | 20M × 12 units × 30.4 ≈ 7.3B × $0.625 per million | ≈ $4,560 |
| DynamoDB reads | ≈ 65M units/day × 30.4 × $0.125 per million (includes 10M catch-ups a day) | ≈ $250 |
| DynamoDB storage, a year in | ~4.4 TB × $0.25 per GB | ≈ $1,100 |
Gateways, 9 × c7g.large | 9 × $0.0725 × 730 | ≈ $480 |
Message service, 3 × c7g.large | 3 × $0.0725 × 730 | ≈ $160 |
| ElastiCache (Valkey) registry, 2 nodes | 2 × ~$0.175 × 730 | ≈ $256 |
| NLB | fixed fee + a few capacity units | ≈ $35 |
| Data transfer, CloudWatch, misc. | ≈ $420 | |
| Total | ≈ $7,300/month |
The phone adds almost nothing to the server's bill: catch-up is a few small reads per app open. What the phone costs is paid on the phone: storage (a few MB at most) and battery (a socket only in the foreground).
R1.8 Trade-Offs
How the app hears about new messages in the foreground.
| Polling | Server-Sent Events (SSE) over HTTP/2 | WebSocket (our choice) | |
|---|---|---|---|
| Direction | Client asks | Server → client only | Both ways on one connection |
| Sending a message | A request | A separate request | A frame on the same connection |
| Latency | Half the interval on average | Instant | Instant |
| Battery | A radio wake-up per poll | One connection | One connection |
| State on the server | None | A long-lived response per client | A long-lived socket per client |
SSE over HTTP/2 is a real option: it rides ordinary HTTP, multiplexes with other requests, and reconnects on its own. But chat sends constantly in both directions (messages, and in Round 2 receipts and typing), and with SSE every send is a separate HTTP request with its own headers, while a WebSocket frame costs a few bytes of header. Both need a stateful connection tier on the server, so SSE doesn't save us that. We choose WebSockets and keep HTTPS polling as a fallback for networks that block them.
Where the local data lives.
| Option | Good at | Why not (or why) |
|---|---|---|
| SQLite (chosen), through Room on Android and GRDB or Core Data on iOS | Queries, indexes, transactions, observed queries | The standard answer; one writer at a time, which our writer queue respects |
| A key-value store or plain files | Simple blobs | No queries ("pending rows in this chat, in order"), no multi-row transactions |
| Realm or another object database | Convenient object model | A second storage engine to learn and migrate; fine, but gives us nothing SQLite doesn't |
| Memory only | Fast | Gone when the OS kills the app, which is exactly when the outbox matters |
R1.9 Failure Modes
| Failure | What the user would see | How the design responds |
|---|---|---|
| App killed with messages in the outbox | Nothing: the messages are still there with clock icons | Rows are durable before they're shown. On the next launch, catch-up runs, reconciles by client_msg_id, and the outbox sends the rest. In Round 1 "next launch" means the next time the user opens the app; Round 2 adds background work. |
| Socket drops mid-send | The clock icon stays a little longer | The worker resends the same frame after reconnecting; the server's idempotency record (or catch-up) ensures one copy. |
| Server restarts a gateway | A brief reconnect | The socket manager reconnects with jittered backoff and catches up by cursor; no message is lost because the server log is the truth. |
| Corrupted local database (a bad flash sector, a crash during a file operation outside SQLite) | The app fails to open the chat | On open, the app runs a quick integrity check after a crash. If the database is damaged, it first tries to read the SENDING rows (the only data that exists nowhere else), then deletes the file, recreates the schema, re-queues those rows, and resyncs recent chats from the server. If even the pending rows can't be read, they're lost; we tell the user which chat had unsent messages. |
| Phone clock very wrong | Nothing in ordering | Order comes from seq; display times from the server. (A clock wrong by days breaks TLS certificate checks; the app shows "check your date and time".) |
| The phone storage is full | The send can't be written | The write fails before the screen shows anything, so we never show a message we didn't save. The app asks the user to free space; the eviction job trims old history first. |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | Durable local write before display; retries with the same client ID; catch-up before draining the outbox, so reconciliation doesn't depend on an expiring record; backoff with jitter REL 4 · REL 5 |
| Performance Efficiency | The screen reads local data only; WAL so reads never wait on the writer; catch-up downloads only what's new PERF 3 |
| Security | TLS on every connection; identity from the authenticated socket, never from the frame; membership checked on every send and read SEC 2 · SEC 9 |
| Cost Optimization | About $7,300/month, derived; the phone adds almost nothing to the server bill COST 5 |
| Operational Excellence | Light this round: client metrics for send-to-ack time and outbox age OPS 8 |
| Sustainability | Light this round: no socket in the background, so no background radio wake-ups SUS 3 |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Makes the local database the source of what the screen shows, and writes before showing.
- Explains why the outbox survives the app being killed, and that there's one writer.
- Generates the message ID on the phone before the first attempt, and explains why text or time windows can't deduplicate.
- Opens a socket only in the foreground, and knows the OS suspends backgrounded apps.
- Catches up by cursor in transactions, and reconciles pending messages against catch-up.
- Orders by the server's sequence, not the phone's clock.
Follow-up questions
-
"Why not let the server assign the message ID and have the phone wait for it?" Answer: then the phone can't name a message until the server has answered, which is exactly the answer that gets lost. If the first attempt is stored but the reply is lost, a retry has nothing in common with the first attempt, and the server stores a second copy. The ID has to exist before the first attempt.
-
"The user edits a pending message before it's sent. What happens?" Answer: if the row is still
SENDINGand not in flight, the writer updates its body in place; the next attempt carries the new text with the sameclient_msg_id. If it's in flight, we treat the edit as a separate "edit" message that follows it, because we can't know whether the server already stored the old text. -
"How does the app delete old messages without deleting something important?" Answer: a background job deletes
SENTandRECEIVEDrows beyond the cap per chat, oldest first, in small transactions through the same writer. It never touchesSENDINGorFAILEDrows, and it never movessynced_seq: the cursor says what the phone has received, and older messages can be fetched again withbefore_seq.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Show a spinner until the server replies" | The user waits on every message, and an offline message is lost when the app closes. |
| "Deduplicate by text within a few seconds" | People send the same text twice, and retries can come much later. |
| "Keep the socket open in the background" | The OS suspends or restricts the app, and every heartbeat wakes the radio. |
| "Sort by the phone's timestamp" | Phone clocks disagree; an offline message sorts into the past. |
| "Keep the outbox in memory" | The OS kills backgrounded apps; the outbox is the one thing only the phone has. |
Round 2 · Senior · "30M Daily Users, Receipts, Media and Battery"
~40 min · Senior SDE (L6) · 1 region, 3 AZs · 30M DAU · 10M open sockets at peak · 500M messages/day, ≈ 20K/s peak · 99.99% · send-to-display < 100 ms and reconnect < 300 ms P99 on good networks · < 1.5% battery a day for messaging
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "We built buyer–seller chat for a marketplace: 1M daily users, 20M messages a day, about 810 a second at peak, 100K open sockets. The rule on the phone is that the screen reads only from a local SQLite database in WAL mode, and every write goes through one serial writer. Tapping send writes the message as
SENDINGwith a client message ID generated before the first attempt, and shows it at once; those rows are the outbox. An outbox worker sends them over a WebSocket that is open only while the app is in the foreground, one in flight per conversation, retrying with the same ID; the server stores once and answers retries with the original sequence number. When the app opens, it catches up per conversation from a cursor, one transaction per page, and reconciles pending rows against catch-up by client ID before sending anything. Messages sort by the server's sequence, pending ones last, and times come from the server. The backend is the chat loop's Round 1 system, about $7,300 a month. Open costs: nothing reaches a phone whose app is in the background, there are no receipts or media, and we haven't counted the battery."
Architecture v1, compact
Synthesizing vector architecture diagram...
The screen reads, one writer writes, the outbox drains through a foreground socket after catch-up.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Appears only after the reply | Optimistic local write; SENDING rows are the outbox | Retries |
| 1.2 | Retry sent it twice | Client message ID; server dedup; one in flight per chat | An idempotency record |
| 1.3 | Slow replies | Foreground-only WebSocket | A gap after each background period |
| 1.4 | What did I miss? | Inbox token + per-chat cursor; reconcile outbox against catch-up | A request per changed chat |
| 1.5 | Wrong order | Server seq; pending last; server times | A pending message may move once |
Open costs: no delivery while the app is backgrounded or closed; no delivered or read receipts; no media; no numbers for battery or data use.
R2.1 The Scope Raise
Interviewer: "We're now a consumer messenger: 30 million people use it daily, and at the evening peak 10 million have it open. People expect a message to arrive when their phone is in their pocket, and to see sent, delivered and read ticks. They send photos and videos up to a few hundred MB, often on a bad connection. They walk from Wi-Fi to cellular mid-conversation. And our top support complaint is battery drain."
We ask back, and say what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| What's the mix of one-to-one and group messages? | About 85% one-to-one; 15% go to small groups averaging 8 members (at most 256). | A message reaches recipients on average. Receipts multiply by recipients (step 2.5). Group fan-out itself is the chat loop's job. |
| Must messages arrive when the app is closed, even force-quit? | Closed, yes. Force-quit is the user's choice, but a notification should still appear. | The socket can't be the background path; OS push is (step 2.1). On iOS only a visible alert survives a force-quit (step 2.2). |
| How big is media? | 10% of messages carry an attachment; 500 KB on average, videos up to 500 MB. | Media leaves the socket for a resumable upload path (step 2.4). |
| What do receipts look like in groups? | "Delivered" and "read" ticks in one-to-one; "read by 5" in groups. | Receipts become cursors, not per-message events (step 2.5). |
| What's the battery budget? | Messaging must use under 1.5% of battery a day for a normal user. | We budget radio wake-ups explicitly (step 2.3). |
| What happens when the network changes? | It should feel seamless: under 300 ms to be connected again. | Detect the change and reconnect proactively; the target holds only on good networks (step 2.3). |
| Several devices per user? | Not yet: one phone per account. | Still one socket and one push token per user. Multi-device is Round 3. |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Users | 1M DAU | 30M DAU |
| Open sockets at peak | ~100K | 10M |
| Messages | 20M/day; ≈ 810/s peak | 500M/day; ≈ 5,787/s average, ≈ 20K/s peak |
| Recipients per message | 1 | 1.9 (small groups) |
| Receipts | "Sent" only | Sent, delivered, read |
| App states served | Foreground only | Foreground, background, killed |
| Media | None | Photos and videos up to 500 MB, resumable |
| Device budget | A few MB of storage | < 1.5% battery a day; bounded data use |
| Availability | 99.9% | 99.99% |
The "Not yet" list from R1.2 is now mandatory: receipts, push and media.
R2.2 What Breaks in the Round 1 Design
| Round 1 choice | What breaks at the new scope |
|---|---|
| Socket only in the foreground, nothing else | A message to a phone in a pocket waits until the user happens to open the app. |
| (The obvious fix) keep the socket open in the background | The OS suspends or restricts the app anyway, and each keep-alive wakes the radio: we'd blow the battery budget many times over (step 2.3). |
| Nothing wakes a closed app | Force-quit or OS-killed apps receive nothing at all. |
| One text frame type over the socket | A 200 MB video sent as frames would block every other message behind it for minutes, and restart from zero on every drop. |
| A receipt per message | 500M messages × 1.9 recipients × 2 receipts is 1.9 billion receipt frames a day, and a flood on the sender's phone in groups. |
| Reconnect when a heartbeat is missed | After walking out of Wi-Fi range, the dead socket isn't noticed for up to two heartbeat intervals: minutes of silence. |
A state column with SENDING/SENT only | No place for delivered and read, and no rules for which update wins when they race. |
R2.3 New Requirements and API Additions
The socket speaks binary now. Round 1's frames were JSON. From Round 2 the socket carries compact binary frames: each field has a number and a type, written as a tag byte, then the value (a Protocol Buffers-style encoding). R2.6 does the byte count and R2.7 the trade-off. Here are the frames as field tables; the "bytes" column is our estimate for a typical frame.
Connecting and resuming in the upgrade request. Under the WebSocket protocol (RFC 6455) the phone can't send frames until the server's 101 Switching Protocols arrives, so anything the server needs for a resume goes into the upgrade request itself:
httpGET /v1/ws?inbox_token=inb_8Kq2xW&app_version=8.14.0&platform=ios&network=cellular&caps=1f HTTP/1.1 Host: chat.example.com Authorization: Bearer <access_token> Upgrade: websocket Connection: Upgrade Sec-WebSocket-Version: 13 Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
| Parameter | Why |
|---|---|
Authorization header | Unlike a browser, a native app can set headers on a WebSocket request, so we skip Round 1's separate ticket call and its round trips on every reconnect. The gateway checks the token's signature in memory. |
inbox_token | Resume: "what changed since" (opaque, R1.4) |
app_version, caps | Which frames and features this app understands (for example "watermark receipts v2") |
platform | iOS or Android: push behaviour differs |
network | Wi-Fi or cellular: the heartbeat interval depends on it |
The server answers 101 and, as its first frame in the same flight, WELCOME: the recommended heartbeat interval and the conversations that changed since inbox_token. There is no separate hello frame, so resuming costs no extra round trip.
RECEIVE_MESSAGE (server → phone), the frame that dominates traffic:
| # | Field | Type | Bytes on the wire |
|---|---|---|---|
| 1 | type | enum | 2 |
| 2 | conversation_id | 16 raw bytes | 18 |
| 3 | client_msg_id | 16 raw bytes | 18 |
| 4 | server_msg_id | fixed 64-bit | 9 |
| 5 | seq (20418 needs 3 varint bytes) | varint | 4 |
| 6 | sender_id | 16 raw bytes | 18 |
| 7 | sent_at_ms | varint | 7 |
| 8 | text | string, ~60 bytes | 62 |
| 9 | media | nested, optional | 0 here |
| Total: ≈ 138 B of payload, plus a 4-byte WebSocket header (a payload over 125 bytes needs the 2-byte extended length) | ≈ 142, call it ≈ 140 |
Heartbeats:
| Frame | Fields | Why |
|---|---|---|
PING (phone → server) | client_time_ms | Keeps NAT mappings alive; proves the socket works |
PONG (server → phone) | server_time_ms, recommended_interval_s | Clock offset for display; lets us tune intervals per network remotely |
Receipts are cursors, one frame per conversation (step 2.5):
| Frame | Fields | Meaning |
|---|---|---|
DELIVERED (recipient → server) | conversation_id, up_to_seq | "This phone has stored every message up to here" |
READ (recipient → server) | conversation_id, up_to_seq | "The user has seen every message up to here" |
RECEIPT_UPDATE (server → sender) | conversation_id, delivered_up_to, read_up_to (one-to-one) or seq, read_count, member_count (groups) | Moves the ticks on the sender's phone |
The server checks that the caller is a member of the conversation, and clamps up_to_seq to the member's own range: from the seq at which they joined to the smaller of the conversation's newest seq and the seq at which they left (if they did). A DELIVERED cursor is also capped at the highest seq the server has actually sent or served to that device. A buggy or hostile client can't mark messages it hasn't received.
Push payloads. On iOS we send through APNs (Apple Push Notification service); on Android through FCM (Firebase Cloud Messaging). Both limit the payload to 4 KB.
An APNs alert (headers apns-push-type: alert, apns-priority: 10, apns-collapse-id: c_9Jq1):
json{ "aps": { "alert": { "title": "Priya", "body": "See you at 7!" }, "mutable-content": 1, "thread-id": "c_9Jq1", "sound": "default" }, "conv": "c_9Jq1", "up_to_seq": 20418 }
An FCM data message (HTTP v1 API; data values must be strings):
json{ "message": { "token": "<device registration token>", "android": { "priority": "HIGH", "ttl": "86400s" }, "data": { "conv": "c_9Jq1", "up_to_seq": "20418", "sender": "Priya", "text": "See you at 7!" } } }
Resumable media upload:
httpPOST /v1/media/uploads HTTP/1.1 Authorization: Bearer <access_token> Content-Type: application/json { "conversation_id": "c_9Jq1", "mime": "video/mp4", "size_bytes": 209715200, "sha256": "9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08" }
httpHTTP/1.1 201 Created Content-Type: application/json { "media_id": "m_5Qz8", "part_size_bytes": 8388608, "part_count": 25, "part_urls": [ { "part": 1, "url": "https://chat-media-prod.s3.us-east-1.amazonaws.com/…&X-Amz-Signature=…" } ], "urls_expire_in_s": 3600 }
Then: PUT each part to its pre-signed URL; POST /v1/media/uploads/m_5Qz8/part-urls with { "parts": [7, 8, 9] } to get fresh URLs for parts not yet uploaded; and POST /v1/media/uploads/m_5Qz8/complete with { "parts": [ { "part": 1, "etag": "…" }, … ] }. Every call checks that the caller created m_5Qz8. A file of 8 MiB or less is a single pre-signed PUT.
R2.4 Design Evolution: Push, Radios, Media and Receipts
Step 2.1: Messages Don't Arrive When the App Is in the Background
The problem: Bob's phone is in his pocket, the app backgrounded. Priya sends him a message. The server stores it, looks for Bob's socket, and finds none: we closed it when the app left the screen (step 1.3). Bob sees nothing until he opens the app. What would you do?
Step 2.2: Silent Pushes Get Dropped
The problem: on Android, a few percent of data messages never lead to a notification: the handler was killed mid-write (low memory), or FCM deprioritized the message to normal priority and Doze held it. (A force-stopped app is a different case: FCM delivers nothing to it, not even notification messages, so it is reached only when the user opens it.) Separately, the registry sometimes says Bob is online, the server writes the message to his socket, and Bob's phone has already walked into a lift: the socket is dead but nobody knows yet. What would you do?
Step 2.3: Heartbeats Drain the Battery; Long Silence Kills the Connection
The problem: a teammate set a 10-second heartbeat to detect dead sockets quickly. Battery complaints jumped. Another tried 5 minutes; on some cellular networks, sockets started dying silently: the phone and the server both thought they were connected, and messages vanished into them until the deadline caught them. And when users walk out of the house, the app is deaf for a minute. What would you do?
Step 2.4: A 200 MB Video Over the Socket Stalls the Chat
The problem: a user sends a 200 MB video on a train. Sent as socket frames, every message queued behind it waits minutes (one TCP connection delivers bytes in order: head-of-line blocking), the gateway buffers hundreds of MB, and when the connection drops at 180 MB, everything starts again. Meanwhile the user switched apps, and iOS suspended us. What would you do?
Loop: Design Google Drive (file sync and storage) · Loop: Design S3-Like Object Storage
Step 2.5: Read Receipts Flood the Network in Groups
The problem: a first version sent a DELIVERED and a READ frame for every message, from every recipient. Priya's message to an 8-person group produces 14 receipt frames to her phone; a user scrolling through 200 unread messages fires 200 READ frames in a second.
What would you do?
Step 2.6: The Message State Is Confusing
The problem: the app stores is_sent, is_delivered, is_read as three booleans. A READ watermark arrives before the DELIVERED one (they travel separately), and the message shows two grey ticks after it was shown blue. A retry of a failed message leaves is_sent = true from an old attempt.
What would you do?
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | Nothing arrives in the background | OS push as the background path: APNs alerts with a Notification Service Extension; FCM high-priority data messages; socket only in the foreground | Someone else's rules; best-effort delivery |
| 2.2 | Silent pushes get dropped; dead sockets | Delivery acknowledged by the phone; deadlines in a sorted set; escalation to OS-shown notifications; collapse by conversation (APNs collapse ID, Android tag), no FCM collapse keys | More visible notifications; ≈ 13 s worst case |
| 2.3 | Heartbeats drain; NATs drop | Adaptive intervals (illustrative 60 s / 180 s) tuned by a server override; network-change callbacks; make before break; resume in 3 round trips; battery budget ≈ 0.3% | Tuning; conditional latency targets |
| 2.4 | 200 MB video on the socket | Direct-to-S3 multipart uploads with pre-signed URLs, run by OS background transfers; references checked for ownership; orphan cleanup with conditional states; 30-day media expiry | A media service and second protocol |
| 2.5 | Receipt floods | Watermark cursors, batched and debounced; forward-only; group summaries | Receipts lag up to 3 s; summaries |
| 2.6 | Confusing message state | One ordered state per message, moved forward only by acks and cursors | – |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
The phone has three ways in: the socket in the foreground, OS push in the background, and HTTPS for catch-up and media. All three write through the same single writer. On the server, the chat core is the chat loop's Round 2 pipeline; what's new here is the deadline set in ElastiCache, push workers talking to APNs and FCM directly, and a media path that never touches the gateways.
The pieces added in this round:
- Push workers (ECS) hold HTTP/2 connections to APNs (token-based auth) and use FCM's HTTP v1 API. They read from a queue fed by the delivery workers and the deadline sweeper, coalesce per recipient and conversation, and delete device tokens that APNs or FCM report as no longer valid. Every push payload is built from the stored message, so a replayed or retried push carries the same
up_to_seqand the same collapse ID. - The deadline set in the routing ElastiCache cluster, and a sweeper in the delivery service.
- The media service (ECS behind the ALB), the S3 media bucket (lifecycle: abort incomplete multipart uploads after 2 days; expire objects after 30 days; blocked public access, reachable only through CloudFront with origin access control), and CloudFront with signed URLs.
- Consumers of the ingest and delivery logs follow the usual rules for stream consumers: a record that keeps failing is retried a bounded number of times, then parked in a dead-letter topic with an alarm, so one poison record can't stall a partition; and anything parked is repaired from the stored message, not by replaying the raw log.
The server-side media record (DynamoDB):
| Attribute | Example | Notes |
|---|---|---|
PK / SK | MEDIA#m_5Qz8 / META | One item per upload |
owner | u_401 | Checked on every media call and on attach |
state | UPLOADING, COMPLETE, DELETING | Conditional transitions only |
size_bytes, sha256, mime | Size checked at complete; SHA-256 checked per part (a multipart upload's checksum is a composite of its parts, not the file's hash) | |
attached_to | { "c_9Jq1#4e0c1a7e…" } (conversation + client message ID) | Added with "only if state = COMPLETE" before the message is stored |
created_at | server time | For the 24-hour orphan rule; our clock, not DynamoDB's (DynamoDB has no server-time function for conditions) |
Trace 1: a backgrounded delivery with escalation (Android)
Synthesizing vector architecture diagram...
The phone's own DELIVERED cursor is the only thing that clears the deadline.
Trace 2: a network handoff (Wi-Fi to 5G, RTT 50 ms)
Synthesizing vector architecture diagram...
Three round trips from "the OS told us" to "resumed". The old socket's registry entry is overwritten by the new connection, and its cleanup deletes the entry only if it still names the old connection.
Trace 3: a large media send
Synthesizing vector architecture diagram...
The socket carries one small frame for a 200 MB video, and nothing is re-uploaded after the tunnel.
R2.6 Numbers and Cost
Traffic (the Round 2 figures, derived; peak is 3.5× the average)
| Quantity | Math | Value |
|---|---|---|
| Messages/s, average | 500,000,000 ÷ 86,400 | ≈ 5,787 |
| Messages/s, peak | 5,787 × 3.5 | ≈ 20,255, call it 20K |
| Recipient deliveries/day | 500M × 1.9 | 950M (≈ 38.5K/s at peak) |
| Receipt frames, naive | 950M × 2 | 1.9B/day (≈ 77K/s at peak) |
| Receipt frames, watermarks (step 2.5) | 1.9B ÷ 3 | ≈ 633M/day (≈ 25.7K/s at peak) |
A correction to a common receipt figure. A common shortcut counts 2 receipts per message (1B a day, ≈ 40.5K/s at peak). That is right only when every message has one recipient. With groups, receipts are per recipient: 1.9B a day before watermarks.
Frame sizes: our byte budget, estimates.
| Frame | JSON | Binary | Where the difference comes from |
|---|---|---|---|
RECEIVE_MESSAGE with a 60-byte text | ≈ 360 B | ≈ 140 B (R2.3) | Three 36-character UUID strings become 16 raw bytes each; field names become one-byte tags; the timestamp becomes a varint |
| Daily message frames (500M up + 950M down = 1.45B) | ≈ 522 GB | ≈ 203 GB | ≈ 61% smaller |
A second correction: the often-quoted "≈ 800 B JSON vs ≈ 160 B binary, 400 vs 80 GB a day" describes an end-to-end encrypted frame (a ratchet public key, an authentication tag, ciphertext) and counts only the upload direction. Round 2 has no end-to-end encryption, so its frame is smaller; Round 3 re-derives the encrypted one.
Heartbeats: blend the rates, not the intervals. Assume half the open sockets are on Wi-Fi (60 s) and half on cellular (180 s). The average ping rate is per second, an effective interval of 90 s, not the 120 s you get by averaging the intervals.
| Quantity | Math | Value |
|---|---|---|
| Pings/s at peak (10M sockets) | 10,000,000 × 0.0111 | ≈ 111,000 (and as many pongs) |
| Bytes per ping on the wire | ~11 B payload + WebSocket header and mask + TLS record + TCP/IP headers | ≈ 100 B |
| Heartbeat bandwidth at peak | 111,000 × 100 B | ≈ 11.1 MB/s ≈ 89 Mbps each way |
| Per day (assume 6M sockets open on average) | 6M × 0.0111 × 86,400 × 100 B × 2 directions | ≈ 1.15 TB |
Heartbeats move about 5.7 times the bytes of all binary message frames (1.15 TB vs 203 GB a day). The interval matters more than the encoding.
Gateways: round per AZ, and survive an AZ. A naive 200 tasks (50,000 sockets each) holds exactly 10M, with no headroom, and "67 per AZ" is 201.
| Quantity | Math | Value |
|---|---|---|
| Two AZs must hold 10M | 10,000,000 ÷ 2 ÷ 50,000 | 100 per AZ |
| Plus 10% for deploys and uneven spread | 100 × 1.1 | 110 per AZ, 330 tasks (c7g.xlarge) |
| Normal load | 10M ÷ 330 | ≈ 30,300 sockets per task |
| After losing an AZ | 10M ÷ 220 | ≈ 45,500 per task |
50,000 sockets per task is an assumption to load-test (TLS buffers per connection decide it).
The routing cache (ElastiCache for Valkey; operations per second at peak):
| Operation | Math | Per second |
|---|---|---|
| Registry refreshes (gateways refresh each live socket's entry every 60 s, 150 s expiry, batched) | 10M ÷ 60 | ≈ 166.7K |
| Registry lookups | 38.5K deliveries | ≈ 38.5K |
| Deadline adds and removes (live deliveries ≈ 45% of 38.5K ≈ 17.3K; Android data pushes ≈ 7.3K; × 2) | (17.3K + 7.3K) × 2 | ≈ 49K |
| Delivered cursors | ≈ receipt frames | ≈ 25.7K |
| Connects and disconnects (15 app opens per user a day: 450M ÷ 86,400 × 3.5 ≈ 18.2K, × 2) | ≈ 36.5K | |
| Total | ≈ 316K |
At a planning figure of 80K operations a second per shard (to load-test), 4 shards would run at 99% of that figure with no headroom, so we use 5 shards (≈ 63K each at peak), each a primary and a replica in another AZ: 10 × cache.r7g.xlarge.
Push volume. Assume 55% of recipient deliveries find the app not in the foreground, and per-conversation coalescing keeps 57% of those: 300M pushes a day, ≈ 3,472/s on average and ≈ 12.2K/s at peak. With 60% of phones on Android, FCM sees ≈ 7.3K/s at peak, ≈ 437K a minute against FCM's default quota of 600,000 messages a minute per project: inside it, but at 73%, so we alarm at 80% and request an increase before growth. (FCM also limits a single Android device to 240 messages a minute and 5,000 an hour; our coalescing stays far below that.) APNs publishes no fixed per-app rate; we watch its throttling responses.
DynamoDB writes, counted honestly (write units per message, average):
| Write | Units |
|---|---|
| Message item (≈ 600 B, one conditional put) | 1 |
| Inbox bump: 1 base + 2 index writes, about one bump per message after the chat loop's coalescing | 3 |
| Read cursors: 633M ÷ 2 ≈ 317M a day ≈ 0.63 per message × (1 base + 1 index write) | 1.27 |
| Delivered cursors, flushed from the cache at most every 10 s per member: ≈ 0.3 per message × 1 | 0.3 |
| Media records: 10% of messages × (create + complete + attach) | 0.3 |
| Total | ≈ 5.87 |
The peak is far above DynamoDB's default quotas (40,000 write units a second per table; 80,000 per account for provisioned tables), so we request increases before launch.
Reads. 450M app opens a day × ≈ 4 units (an inbox query and a page or two of catch-up) = 1.8B, plus history scrolling ≈ 0.3B: ≈ 2.1B a day ≈ 24.3K units/s on average. For an AZ loss, 3.33M phones reconnect through 220 surviving gateways admitting 300 a second each (66,000/s, about 50 s); at ≈ 0.8 read units per reconnect that's ≈ 53K extra read units a second, which we provision as a floor during the 4 evening peak hours with scheduled scaling.
Media (step 2.4; 50M files a day, 500 KB average):
| Quantity | Math | Value |
|---|---|---|
| Upload volume | 50M × 500 KB | 25 TB/day ≈ 2.31 Gbps on average (free into S3) |
| Stored (30-day expiry) | 25 TB × 30 | 750 TB |
| Downloads | 50M × 1.9 recipients × 500 KB | 47.5 TB/day ≈ 1,444 TB/month through CloudFront |
| Upload requests | 50M × 1.04 (multipart parts for large videos) × 30.4 | ≈ 1.58B PUTs a month |
Monthly cost (us-east-1 list prices, 730 hours and 30.4 days a month; rounded)
| Item | Math | Monthly |
|---|---|---|
| DynamoDB writes (provisioned, auto scaling at 70%) | 34.0K ÷ 0.7 ≈ 48.5K units × $0.00065 × 730 | ≈ $23.0K |
| DynamoDB reads (70%) + evening storm floor | 34.7K × $0.00013 × 730 ≈ $3.3K; 53K × $0.00013 × 4 h × 30.4 ≈ $0.8K | ≈ $4.1K |
| DynamoDB storage | messages 90 days hot: 500M × 600 B × 90 = 27 TB, + ~1 TB of other items; 28,000 GB × $0.25 | ≈ $7.0K |
Gateways, 330 × c7g.xlarge | 330 × $0.145 × 730 | ≈ $34.9K |
| NLB | ≈ 60 capacity units (active connections and bytes are about equal) × $0.006 × 730 + fixed fee | ≈ $0.3K |
| ALB (catch-up, media API) | ≈ 200 capacity units, driven by new connections | ≈ $1.2K |
ElastiCache (Valkey), 10 × cache.r7g.xlarge | 10 × ~$0.350 × 730 | ≈ $2.6K |
| MSK for the chat core | assume 6 × kafka.m7g.large at ≈ $0.20/h (list price, an assumption) + storage + cross-AZ client traffic | ≈ $1.6K |
| Workers (sequencers, delivery, push, media) | assume 30 × c7g.xlarge | ≈ $3.2K |
| Data transfer out (chat) | ≈ 27.8 MB/s on average (catch-up responses 15.6, pongs 6.7 (≈ 17.6 TB a month, ≈ $1.2K at the $0.07 tier), pushes to APNs/FCM 3.5, message frames 1.5, receipts 0.4) ≈ 73 TB: 10 TB × $0.09 + 40 TB × $0.085 + 23 TB × $0.07 | ≈ $5.9K |
| Cross-AZ traffic inside the region | registry refreshes ≈ 100K/s on average × ≈ 150 B ≈ 15 MB/s ≈ 39 TB a month, about ⅔ of it crossing AZs ≈ 26 TB; gateway ↔ chat core ↔ cache traffic ≈ 34 TB; ≈ 60 TB × $0.02 per GB ($0.01 each way) | ≈ $1.2K |
| S3 media storage | 50 TB × $0.023 + 450 TB × $0.022 + 250 TB × $0.021 (per GB) | ≈ $16.3K |
| S3 upload requests | 1.58B × $0.005 per 1,000 | ≈ $7.9K |
| CloudFront data out | 10 TB × $0.085 + 40 × $0.080 + 100 × $0.060 + 350 × $0.040 + 524 × $0.030 + 420 × $0.025 (per GB) | ≈ $50.3K |
| CloudFront requests + S3 GETs on misses | 2.89B × $0.01 per 10,000 ≈ $2.9K; assume 80% miss: 2.31B × $0.0004 per 1,000 ≈ $0.9K | ≈ $3.8K |
| Push to APNs and FCM | sent directly: no per-message fee | $0 |
| CloudWatch, SQS, misc. | ≈ $5.0K | |
| Total | ≈ $168K/month |
About 0.56 cents per daily user a month. Media is 47% of the bill ($78.3K of ≈ $168.3K), and most of that is CloudFront delivering photos and videos. Gateways are 21%, DynamoDB 20%. Two things people expect to be expensive aren't: heartbeats (their bytes cost a few hundred dollars) and binary-vs-JSON (the saving is about 319 GB a day, under $1K a month). What they cost is paid on the phone, in battery.
Sending pushes through Amazon SNS instead would add about 300\text{M} \times 30.4 \times \1.00 per million (publish plus delivery) ≈ **\9.1K a month**.
R2.7 Trade-Offs
Socket or push, by app state.
| App state | Path | Why |
|---|---|---|
| Foreground | WebSocket | Tens of milliseconds; the radio is already awake for the user |
| Just backgrounded (seconds) | Close the socket; push from now on | iOS will suspend us anyway; closing on our terms avoids half-dead sockets |
| Background or killed | Push (iOS alert; Android high-priority data message) | The OS's shared connection costs us nothing until something arrives |
| Force-quit (iOS) | Alert push, shown by the system | Background pushes aren't delivered to force-quit apps |
| Force-stopped (Android) | Nothing until the user opens the app | FCM can't reach a force-stopped app |
| Muted chat | No alert; best-effort hint at most | Don't wake the phone for something the user asked not to hear |
Heartbeat interval vs battery vs detection time.
| Interval | Radio wake-ups in an idle hour | Dead socket noticed after (2 missed) | Risk |
|---|---|---|---|
| 10 s | 360 | 20 s | Battery; this is the "hog" |
| 60 s (Wi-Fi start) | 60 | 2 min | Fine on Wi-Fi, where the radio is cheap |
| 180 s (cellular start) | 20 | 6 min | A carrier NAT with a shorter timeout silently kills sockets |
| 10 min | 6 | 20 min | Beyond the NLB's 350 s default and many NATs |
We don't pick detection time by heartbeat alone: network-change callbacks, send timeouts and the server's delivery deadline all shorten it.
Binary vs JSON frames.
| Binary, Protocol Buffers-style (chosen for the socket) | JSON | |
|---|---|---|
| Size | ≈ 140 B per message frame | ≈ 360 B |
| Parsing on old phones | Fast, few allocations | Slower |
| Evolution across app versions | Numbered fields; old apps skip unknown fields | Unknown keys ignored too, but no schema to check |
| Debugging | Needs a decoder | Readable |
| What it saves us | ≈ 319 GB a day, under $1K a month | – |
The money is small. We choose binary for the socket because it is typed, versioned by field numbers, and cheap to parse; REST APIs stay JSON.
Push through SNS vs directly. SNS mobile push hides the provider protocols and token handling for ≈ $9.1K a month at our volume. We send directly: at 300M pushes a day the saving pays for the code, and we control collapse IDs, priorities and push types per message. A smaller team should use SNS (it supports these headers too).
Media retention. 30 days on the server keeps storage at 750 TB (≈ $16.3K a month); keeping everything would grow by 760 TB a month, forever. The price is "no longer available" on old media for users who lost their phone, which Round 3's backups address.
R2.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| Silent-push throttling (Low Power Mode, Doze, app standby buckets) | Hints arrive late or never | Nothing depends on them: messages go as alerts or high-priority data messages, and step 2.2's deadline escalates to a notification the OS shows by itself. |
| Radio thrashing (a client release with a short heartbeat) | Battery complaints; OS energy warnings for the app | The server's recommended_interval_s override fixes it without a release; MetricKit and Android vitals by app version tell us which build. |
| NAT drops on one carrier | Sockets on that carrier die silently; deadlines fire more often | Per-carrier deadline-escalation rate is a metric; we lower that carrier's interval in the override table. |
| Orphaned media (cancelled or abandoned uploads) | Storage growth; staging files on phones | S3 aborts incomplete multipart uploads after 2 days; the cleanup job deletes unattached media with conditional state changes; the phone deletes old staging files daily. |
| Receipt stampede in a group (200 members open an announcement at once) | A burst of READ frames | Watermarks are one per member; the server coalesces the sender's summary to one every 2 seconds. |
| Reconnect storm (an AZ fails; 3.33M sockets drop) | A wave of TLS handshakes, token checks and resume reads | Apps reconnect with full jitter (a random wait between 0 and a doubling cap), gateways admit at most 300 new connections a second each and close the rest at TCP accept before any TLS work, TLS session resumption makes the admitted ones cheaper, and resume reads only what changed. Immediate reconnects would line up millions of handshakes and token checks in the same second, and every failure would retry in lockstep. |
| APNs or FCM outage | Pushes fail or queue | Push workers back off per provider; the queue holds work, coalesced per conversation, so a backlog shrinks as it waits; messages are stored, so opening the app always shows them. |
| An old app version gets a new frame type | The old app can't show it | The connect request carries the app version and capabilities; the server sends only frames the app declared, and a "please update" placeholder message for content it can't render. |
Drill: WebSocket reconnect thundering herd (the storm is answered in the "Reconnect storm" row above; WebSockets vs SSE over HTTP/2 in R1.8)
R2.9 Production Gotchas
| Gotcha | Symptom | Cause | Fix |
|---|---|---|---|
| Plaintext messages in local storage | Chats readable from an unencrypted device backup or a jailbroken phone | The database file sits in normal app storage and is included in OS backups | This round: the OS's file protection (iOS Data Protection; Android's file-based encryption), and exclude the database from OS backups. Round 3 encrypts the database itself. |
| Persistent socket in the background | Battery complaints; the socket dies anyway | Fighting the OS | Close on background; push is the background path |
| Media streamed over the socket | Chat stalls behind a video; restarts from zero | Head-of-line blocking; no resume | Direct-to-S3 multipart uploads with pre-signed URLs, in OS background sessions |
| Per-message receipts | Receipt frames outnumber messages several times | A frame per message per recipient | Watermark cursors, batched |
| FCM collapse keys for chat | Notifications delayed for minutes in busy chats | Collapsible messages are throttled (burst of 20, then one every 3 minutes per device) | Non-collapsible messages; coalesce on our side; replace by notification tag |
| High-priority FCM with no notification | Android messages start arriving late | FCM deprioritizes apps whose high-priority messages don't show anything | High priority only when a notification follows |
| VoIP pushes for messages | The app is killed; VoIP pushes stop | PushKit requires reporting a call | Alerts with a Notification Service Extension |
R2.10 Pillar Check
| Pillar | What Round 2 adds |
|---|---|
| Reliability | Delivery proven by the phone's own cursor; deadlines and escalation to OS-shown notifications; AZ-sized gateway headroom and admission control; quotas (DynamoDB, FCM) raised ahead of need REL 1 · REL 5 · REL 10 |
| Performance Efficiency | Foreground socket for tens-of-milliseconds delivery; resume in three round trips; media off the socket and served by a CDN; latency targets stated with their network conditions PERF 4 · PERF 5 |
| Security | Bearer token on the upgrade, checked in memory; membership and ownership checks on every receipt, media call and media reference; pre-signed URLs scoped to one part and one hour; private bucket behind CloudFront signed URLs SEC 3 · SEC 9 |
| Cost Optimization | ≈ $168K/month, derived; media named as half the bill; direct APNs/FCM instead of SNS; 30-day media expiry; incomplete uploads aborted COST 5 · COST 6 · COST 8 |
| Operational Excellence | Client and server metrics per app version, platform and carrier; a server-side override for heartbeat intervals; alarms with first actions OPS 4 · OPS 8 |
| Sustainability | The battery budget: no background socket, OS push instead, heartbeats only when idle, batched receipts, coalesced pushes (≈ 0.3% of a battery a day) SUS 2 · SUS 3 |
Alarms and first actions
| Signal | Alarm | First action |
|---|---|---|
| Send-to-ACK P99, by network type (client-reported) | > 2× baseline for 10 min | Server stage timings, or one carrier or region of users? |
| Deadline escalations per distinct recipient device | > 3% of deliveries for 15 min | Which platform, carrier or app version? Heartbeat override? |
| Push send failures by provider | > 2% for 10 min | Provider status; invalid-token cleanup; back off |
| FCM quota use | > 80% of the per-minute quota | Request an increase; check coalescing |
| Reconnect time P99 and detection time, by network | > 2× baseline | A gateway problem, or an app release? |
| Incomplete multipart uploads (S3 Storage Lens) | growing week over week | Client upload bug? Lifecycle rule intact? |
| Battery: background wake-ups per user (MetricKit, Android vitals) | > 2× the last release | Roll back the release's interval or push behaviour |
R2.11 Round 2 Rubric and Follow-Ups
What a senior (L6) answer adds over L5
- Uses the OS push service as the background path and knows each platform's rules: iOS alerts vs background pushes, Android high priority vs deprioritization, force-quit and force-stop.
- Defines "delivered" as the phone's acknowledgement, and escalates on a deadline rather than sending more pushes.
- Explains why sockets die (NAT timeouts, the NLB idle timeout), treats interval values as tunable assumptions, and reacts to network-change callbacks.
- Does the battery arithmetic in wake-ups, and shows that a background socket alone would break the budget.
- Moves media to resumable direct-to-S3 uploads run by the OS, and cleans up orphans without racing attaches.
- Turns receipts into forward-only cursors and counts them per recipient.
- Re-checks carried numbers: receipts per recipient, heartbeat rates blended correctly, gateway headroom per AZ, and finds that media dominates the bill.
Follow-up questions
-
"Typing indicators?" Answer: foreground only, over the socket, never stored and never pushed. At most one every few seconds per user, sent only to members who have the chat open, and the indicator expires on the receiving side after about 5 seconds unless refreshed.
-
"The user sends a photo with no signal. What does the recipient see, and when?" Answer: the sender sees the photo at once, from the local file, marked sending. The upload manager queues the upload; the message waits in the outbox, behind its upload, because a media reference is only valid once the upload is
COMPLETE. When the network returns, the OS runs the upload, then the outbox sends the message. The recipient sees the message with its inline thumbnail immediately, and the full image when their phone downloads it. -
"Why not have the Notification Service Extension fetch everything the user missed?" Answer: it has about 30 seconds and a tight memory limit, and it runs once per notification. It should do one small, bounded job (write the message in the payload, or fetch one), and leave a full catch-up to the app when the user opens it.
Interview gotchas from this round
| Gotcha | Why it's wrong |
|---|---|
| "Silent pushes deliver messages on iOS" | They're throttled, may be dropped, and aren't delivered to force-quit apps. Use alerts. |
| "Delivered means APNs accepted it" | Only the phone's own acknowledgement proves delivery. |
| "180 seconds matches carrier NAT timeouts" | Carrier timeouts vary; the value is a starting point tuned per network. |
| "1B receipts a day" | That's per message; receipts are per recipient: 1.9B before watermarks. |
| "Average the heartbeat intervals" | Average the rates: 60 s and 180 s blend to 90 s, not 120 s. |
| "200 gateways for 10M sockets" | Exactly full; an AZ loss strands a third of users. |
Round 3 · Architect · "End-to-End Encrypted, Multi-Device, and Restorable"
~45 min · Principal (L7) · priced as one home region, 3 AZs (geography as in the chat loop's Round 3) · 100M DAU, ~2.2 active devices each · 1.7B messages/day, ≈ 69K/s peak · ≈ 7.5 encrypted envelopes per message · 40M open sockets at peak · 99.99%
R3.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 3. If you're starting here, it's everything you need from Rounds 1 and 2.
Round 2 in 60 seconds. "We run a messenger for 30 million daily users: 10 million open sockets at peak, 500 million messages a day, about 20,000 a second at peak, 1.9 recipients per message. On the phone, the screen reads a local SQLite database through one writer; sends go through an outbox with client-generated IDs. The socket is open only in the foreground, with adaptive heartbeats (60 s Wi-Fi, 180 s cellular as tunable starting points) and reconnects on network-change callbacks in three round trips. In the background, OS push is the path: iOS alerts with a Notification Service Extension, Android high-priority data messages that always end in a notification. Delivery counts only when the phone acknowledges it; a deadline escalates to a notification the OS shows by itself, collapsed per conversation. Media goes straight to S3 in resumable 8 MiB parts through the OS's background transfers, and the message carries only a checked reference. Receipts are forward-only cursors, and each message has one ordered delivery state. The battery cost is about 0.3% a day. About $168K a month, almost half of it media delivery. Open costs: the server can read every message, a user has exactly one device, a lost or replaced phone loses its history, and a lost phone's database is protected only by the OS."
Architecture v2, compact
Synthesizing vector architecture diagram...
Socket in the foreground, push in the background, media on its own path, everything landing in one local database.
Steps so far
| Step | Problem | Component |
|---|---|---|
| 1.1–1.2 | Instant send; no duplicates | Local write first; outbox; client message ID |
| 1.3–1.5 | Live receive; catch-up; order | Foreground socket; cursors; server seq |
| 2.1–2.2 | Background delivery | OS push; phone-acknowledged delivery; deadline escalation; collapse per conversation |
| 2.3 | Battery and handoffs | Adaptive heartbeats; network callbacks; 3-round-trip resume |
| 2.4 | Big media | Resumable direct-to-S3 uploads in OS background sessions |
| 2.5–2.6 | Receipts; state | Watermark cursors; one forward-only state |
Open costs: we can read every message; one device per user; no restore; keys and data protected only by the OS.
R3.1 The Scope Raise
Interviewer: "We're growing to 100 million daily users, and we're making every conversation end-to-end encrypted: we must be unable to read messages, even if ordered to or breached. People use a phone, a tablet and a desktop, all in sync. A new phone must restore history. During outages, messages will arrive out of order and late. A lost phone must not expose chats. And notifications must still show who wrote and what, without our servers seeing the text."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| What exactly is end-to-end encrypted? | Message content, media, and read receipts. The server may know who talks to whom and when. | Keys live only on devices; the server routes sealed envelopes; read receipts become encrypted control messages (step 3.1). Metadata stays visible, and we say so. |
| How many devices per user? | A phone plus up to four linked devices; about 2.2 active on average. | Every device is its own encryption endpoint; every message fans out per device (step 3.3). |
| Restore from what? | From an encrypted backup the user turns on, or from the old phone if they still have it. | Backups encrypted on the device with a key we can't read; recovery becomes a key-custody decision (step 3.4). |
| How late can messages arrive? | Minutes to hours during outages, and out of order. | Keys for skipped messages must be kept, bounded (step 3.2). |
| What does "lost phone must not expose chats" mean? | A thief with the phone, or with a copy of its files, can't read history; the user can cut the phone off from another device. | Database encryption keyed from secure hardware; device revocation (step 3.5). |
| Which regions? | Assume the chat loop's home-region design for geography. Focus on the devices and keys. | We price one home region and link the chat loop for regions and residency. |
| Message volume? | About 17 messages per daily user per day, as today. | 1.7B messages a day (R3.6). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Users | 30M DAU | 100M DAU |
| Messages | 500M/day; ≈ 20K/s peak | 1.7B/day; ≈ 19,676/s average, ≈ 69K/s peak |
| Devices | 1 per user | ~2.2 active per user; up to 5 |
| What the server can read | Everything | Metadata only: routing, sizes, timing |
| Server copies of messages | A 90-day log per conversation | A mailbox per device, deleted when the device acknowledges |
| History on a new phone | From the server | From an encrypted backup or the old phone |
| Lost phone | OS protection | Encrypted database; keys from secure hardware; remote revocation |
| Notification text | Written by the server | Decrypted on the device |
R3.2 What Breaks in the Round 2 Design
| Round 2 choice | What breaks at the new scope |
|---|---|
| The server stores readable messages and builds previews | Violates end-to-end encryption outright. |
| One device, one socket, one push token per user | A tablet and a desktop need their own copies, each encrypted for them. |
| A per-conversation log on the server that any device can catch up from | Under end-to-end encryption, the log would be ciphertext only one device can decrypt; a newly added device can't read it. |
| History lives on the server | A new phone has nowhere to get history from, unless the user made an encrypted backup. |
| Keys (such as they were) handled by our servers | Our servers must never hold a key that decrypts messages. |
| Database protected only by the OS | A copy of the file (a backup, a forensic extraction) is readable wherever the OS protection doesn't apply. |
| Server-generated notification text | The server can't see the text. |
R3.3 New Requirements and API Additions
Device registration and keys. Each device creates its own key pairs and uploads only the public halves:
httpPOST /v1/devices HTTP/1.1 Authorization: Bearer <access_token> Content-Type: application/json { "device_name": "Priya's iPad", "identity_key": "base64…", "signed_prekey": { "id": 17, "public_key": "base64…", "signature": "base64…" }, "signed_pq_prekey": { "id": 4, "public_key": "base64…", "signature": "base64…" }, "one_time_prekeys": [ { "id": 5001, "public_key": "base64…" } ], "one_time_pq_prekeys": [ { "id": 901, "public_key": "base64…", "signature": "base64…" } ], "device_list_update": { "version": 8, "devices": ["d_1", "d_2", "d_3"], "signature": "base64…" } }
The device list update is signed by an existing device of the same account, which approved the new one (the user scanned a QR code on it). The server checks the signature and the version before storing the list, and never accepts a device list for another user.
Fetching keys to start a session (one bundle per device of the target user; each fetched one-time prekey is removed so it's used once). Because every fetch consumes keys, bundle fetches are rate-limited per requesting account and device (a few dozen new users an hour, with bursts for contact syncs), so nobody can drain a victim's one-time prekeys on purpose. Only the device that owns a key set can upload or replace its prekeys; the server checks the authenticated device, not a field in the body:
httpGET /v1/keys/u_802/devices HTTP/1.1 Authorization: Bearer <access_token>
httpHTTP/1.1 200 OK Content-Type: application/json { "device_list": { "version": 8, "devices": ["d_1", "d_2", "d_3"], "signature": "base64…" }, "bundles": [ { "device_id": "d_1", "identity_key": "…", "signed_prekey": { … }, "signed_pq_prekey": { … }, "one_time_prekey": { "id": 5001, … }, "one_time_pq_prekey": { "id": 901, … } } ] }
A sealed send: one message, one envelope per device (field table; bytes are estimates for a short text):
| Field | Type | Bytes | Visible to the server? |
|---|---|---|---|
client_msg_id | 16 raw bytes | 18 | Yes |
conversation_id | 16 raw bytes | 18 | Yes |
envelopes[] | repeated | per device | – |
↳ device_id | varint | 2 | Yes |
↳ type | enum: prekey message (first in a session) or normal | 2 | Yes |
↳ ratchet_public_key | 32 bytes | 34 | Yes, but meaningless without private keys |
↳ counter, previous_counter | varints | 6 | Yes |
↳ ciphertext | bytes (plaintext padded to a multiple of 160 B) | 162 | No |
↳ auth_tag | bytes | ≈ 18 | – |
Each envelope as delivered, with the routing header the server adds (seq, server_msg_id, sender user and device, sent_at), is about 300 bytes in binary. The same thing in JSON, with UUID strings and the sealed part base64-encoded, is about 590 bytes. (The often-quoted "800 B vs 160 B" counts this kind of frame, but without padding the ciphertext; our estimate pads, which hides message length at the cost of bytes.)
If the envelope set doesn't match the recipient's current devices, the server refuses the whole send:
httpHTTP/1.1 409 Conflict Content-Type: application/json { "error": "DEVICE_MISMATCH", "user_id": "u_802", "device_list_version": 9, "missing_devices": ["d_4"], "extra_devices": ["d_2"] }
The sender's device fetches bundles for d_4, drops d_2, re-encrypts and retries with the same client_msg_id.
Encrypted backups: POST /v1/backups/snapshots returns pre-signed URLs for an encrypted snapshot; POST /v1/backups/increments for a daily increment; GET /v1/backups/latest lists what exists. The server stores opaque bytes per account and checks that the caller owns the account. If the user chose a PIN-protected backup key, the key vault has two calls: PUT /v1/vault/entry (store the wrapped backup key with a PIN-derived verifier) and POST /v1/vault/recover (guess-limited).
Revoking a device: DELETE /v1/devices/d_2 from another signed-in device, carrying a device list update without d_2, signed by that device. The server verifies that d_2 belongs to the caller's account before doing anything.
Safety numbers are computed on the devices, and must cover every device identity key of both users (so adding a device changes the number), unless the account uses one identity key shared by its linked devices, as Signal does. The server only serves the public keys.
R3.4 Design Evolution: Keys, Devices, Backups and Lost Phones
Step 3.1: We Must Not Be Able to Read Messages
The problem: today TLS protects messages between the phone and our servers, and our servers see plaintext. The requirement is that only the devices at both ends can read a message, and that a breach of our servers, or an order to hand over messages, yields nothing readable. What would you do?
Step 3.2: Messages Arrived Out of Order and Some Won't Decrypt
The problem: during a network outage, Bob's phone receives Alice's messages 7, 8 and 11 now, and 9 and 10 an hour later. With a ratchet, each message's key comes from the one before it. Separately, some users report "couldn't decrypt this message" after a session state was lost or corrupted (a crash, a bug, an old file restored from an OS backup). What would you do?
Step 3.3: Phone, Tablet, Desktop
The problem: Bob reads chats on his phone, his tablet and his desktop. Alice's message must reach all three, and Bob's own messages sent from the desktop must show up on his phone. What would you do?
Step 3.4: Restore History on a New Phone
The problem: Priya drops her phone in a lake and buys a new one. Our servers hold no readable history, and her mailbox only holds messages not yet delivered. Her tablet has some history. What does her new phone show? What would you do?
Step 3.5: A Lost Phone
The problem: Priya's other phone was stolen, not drowned. The thief has the device; maybe they copied its files. Its database holds years of history and every session key. How much can they read, and how does Priya cut it off? What would you do?
Primitive: OAuth2, OIDC & Distributed Token Authentication · Drill: JWT revocation token blacklist
Step 3.6: Notification Previews Without Us Reading the Text
The problem: in Round 2 the server wrote "Priya: See you at 7!" into the push. Now it only has ciphertext. Users still expect the sender and the text on the lock screen. What would you do?
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | We must not read messages | Signal Protocol: X3DH/PQXDH from prekey bundles; Double Ratchet per message; opaque envelopes; receipts as encrypted control messages | Server search, previews, moderation, server history |
| 3.2 | Out of order; won't decrypt | Skipped keys (≤ 1,000 skips, ≤ 2,000 stored, ≤ 30 days); processed-envelope IDs in the same transaction; one writer across app and extension; decryption-error repair | Stored keys; a repair protocol |
| 3.3 | Several devices | Signed device lists; per-device sessions and mailboxes; own devices included; sender keys for groups | ≈ 5.38 envelopes per message |
| 3.4 | Restore on a new phone | Device-to-device transfer; encrypted backups; recovery key or PIN + guess-limited vault; restored phone is a new device; no sessions in backups | History lost with the key; a vault to run |
| 3.5 | Lost phone | SQLCipher with a random key in hardware-backed storage; available after first unlock; revocation: refresh token, in-memory list, mailbox and push token deleted | Lock-state trade-off; key code on two platforms |
| 3.6 | Previews without plaintext | Per-device push carrying the envelope; decrypt in the extension or handler; 20 s (iOS) / ≈ 5 s (Android, longer work to an expedited job) internal deadlines; "New message" fallback | A second process in the crypto path; more pushes |
R3.5 Global Architecture
Synthesizing vector architecture diagram...
Keys are created and used only on the left. The right side stores public keys, routes ciphertext, and keeps backups and media it can't read. The only secret the server side guards is the vault's wrapped backup keys, and those are useless without the user's PIN and the enclave's guess limit.
The server's tables (DynamoDB):
| Item | PK | SK | Attributes |
|---|---|---|---|
| Device list | USER#<user> | DEVLIST | version, devices, signature (by an existing device) |
| Device | USER#<user> | DEVICE#<device> | identity_key, signed_prekey, signed_pq_prekey, push_token, app_version, capabilities, revoked |
| One-time prekey | OTPK#<device> | <key_id> (Number) | public_key; taken with a conditional delete, so two senders never get the same one |
| Mailbox envelope | MBOX#<device> | MSG#<conversation>#<seq, 12 digits>#<sender device> for messages; CTRL#<server id> for receipts, typing and sender keys | ciphertext, sender, sent_at, expires_at (30 days, written with our clock); ttl for cleanup only |
| Conversation meta | CONV#<conv> | META | last_seq (never expires), members |
Envelopes are deleted when their device acknowledges them. Envelopes for a device that never comes back carry a 30-day ttl attribute for cleanup. DynamoDB deletes expired items only eventually (typically within a few days), so readers also skip any envelope whose expires_at is in the past by our servers' clocks.
Trace 1: the first message to a new contact (Alice → Bob, who has 3 devices)
Synthesizing vector architecture diagram...
One network trip to the key directory per new contact, then normal sends.
Trace 2: a restore on a new phone
Synthesizing vector architecture diagram...
History comes from the backup; sessions are always new.
R3.6 Numbers and Cost
Traffic (100M DAU × 17 messages; peak 3.5× average)
| Quantity | Math | Value |
|---|---|---|
| Messages/day | 100M × 17 | 1.7B |
| Messages/s | 1.7B ÷ 86,400; × 3.5 | ≈ 19,676 average; ≈ 68,866 peak |
| Message envelopes per message | step 3.3 | 5.38 |
| Read-receipt envelopes per message | read watermarks 1.9 ÷ 3 ≈ 0.63 per message, each to the sender's 2.2 devices and the reader's other 1.2: 0.63 × 3.4 | ≈ 2.15 |
| Envelopes stored per message | 5.38 + 2.15 | ≈ 7.53 |
| Envelopes/day | 1.7B × 7.53 | ≈ 12.8B (≈ 148K/s average, ≈ 519K/s peak) |
| Open sockets at peak | assume a third of daily users' phones and tablets in the foreground (33.3M, Round 2's ratio) + 6.7M desktops | ≈ 40M |
Gateways (Round 2's rule: two AZs hold everything, plus 10%): per AZ, × 1.1 = 440 per AZ, 1,320 tasks.
DynamoDB writes per message. Every envelope is written once and deleted once when acknowledged (1 unit each, as envelopes are about 400 bytes with attributes):
| Write | Units |
|---|---|
| Envelopes: 7.53 × (put + delete) | 15.06 |
| Delivered cursors (coalesced) | 0.3 |
| Media records | 0.3 |
| Total | ≈ 15.66 (Round 2: 5.87; there are no server inbox bumps any more, but there are 7.5 times as many items) |
Each mailbox is its own partition key, so the per-partition limit (1,000 write units a second) only matters for a single device, far above what one device receives. The table and account quotas must be raised far ahead of launch.
Key directory. Assume each daily user starts sessions with 0.5 new contacts' users a day (new contacts, new devices, reinstalls): devices = 110M bundle fetches a day (≈ 1,270/s). Each takes one classical and one post-quantum one-time prekey (conditional deletes; the post-quantum key is about 1.6 KB, so 2 units), and the owner later uploads replacements: ≈ 6 write units per fetch, ≈ 660M a day ≈ 7.6K units/s on average. Storage: 220M devices × (50 classical one-time keys × ~100 B + 50 post-quantum ones × ~1.7 KB) ≈ 19.8 TB, most of it post-quantum keys.
Device-list checks. Every send is checked against the recipients' current device lists, cached by version in ElastiCache: the recipients' lists and the sender's own (for its other devices), ≈ 68,866 × (1.9 + 1) ≈ 200K lookups/s at peak.
The routing cache (per second at peak): registry refreshes 40M ÷ 60 ≈ 667K; device-list lookups ≈ 200K; delivery deadlines ≈ (259K live envelope deliveries + 50.8K Android pushes) × 2 (added and removed) ≈ 620K; connects and disconnects ≈ 122K (15 opens per user a day); cursors ≈ 60K. ≈ 1.67M operations a second → at 80K per shard that is 21 shards with no headroom, so we run 24 shards (≈ 70K each at peak), 48 × cache.r7g.xlarge.
Pushes. Phones and tablets receive 5.38 × (1.6 ÷ 2.2) ≈ 3.91 message envelopes per message; with Round 2's 55% not in the foreground and 57% kept after coalescing: 2.09B pushes a day (≈ 24.2K/s average, ≈ 84.7K/s peak). FCM's share (60%) is ≈ 50.8K/s ≈ 3.05M a minute: five times FCM's default project quota of 600,000 a minute. We must have the increase approved before launch. Through SNS these would cost ≈ $63.5K a month; we send directly.
Media (10% of messages, 500 KB average, encrypted on the device with a random per-file key that travels inside the message):
| Quantity | Math | Value |
|---|---|---|
| Uploads | 170M files × 500 KB | 85 TB/day |
| Stored, 30 days | 85 TB × 30 | 2,550 TB |
| Downloads per file | 1.9 recipients × 1.6 devices that download + 0.5 of the sender's other devices | 3.54 |
| Downloads | 170M × 3.54 × 500 KB × 30.4 | ≈ 9,147 TB/month |
Encrypted media can't be recompressed, thumbnailed or deduplicated by us: the sender's device makes the thumbnail and puts it inside the message.
Backups (assume 60% of daily users turn them on; text-only backups average 15 MB encrypted; media backup is out of scope):
| Quantity | Math | Value |
|---|---|---|
| Snapshots | 60M × 15 MB | 900 TB in S3 Standard-IA, replaced monthly |
| Increments | 40% of backup users a day × 200 KB = 4.8 TB/day; ≈ 15 days' worth live on average | ≈ 72 TB in S3 Standard |
| Requests | 60M snapshot PUTs + 24M × 30.4 ≈ 730M increment PUTs a month |
Monthly cost (us-east-1 list prices; rounded)
| Item | Math | Monthly |
|---|---|---|
| DynamoDB writes: envelopes | 308K ÷ 0.7 ≈ 440K units × $0.00065 × 730 | ≈ $208.9K |
| DynamoDB writes: key directory | 7.6K ÷ 0.7 ≈ 10.9K × $0.00065 × 730 | ≈ $5.2K |
| DynamoDB reads (mailbox fetches, bundles) + AZ-loss floor | ≈ 21.5K units on average ÷ 0.7 ≈ 30.7K × $0.00013 × 730 ≈ $2.9K; 13.3M reconnects ÷ (880 × 300/s) ≈ 50 s, floor ≈ 211K units for 4 evening hours ≈ $3.3K | ≈ $6.2K |
| DynamoDB storage | pending envelopes ≈ 10 TB + prekeys 19.8 TB + ≈ 1 TB other ≈ 31 TB × $0.25 per GB | ≈ $7.8K |
Gateways, 1,320 × c7g.xlarge | 1,320 × $0.145 × 730 | ≈ $139.7K |
| NLB | bytes dominate: ≈ 720 capacity units × $0.006 × 730 | ≈ $3.2K |
| ALB (keys, media, backups, catch-up) | ≈ 690 capacity units from new connections | ≈ $4.1K |
ElastiCache (Valkey), 48 × cache.r7g.xlarge | 48 × ~$0.350 × 730 | ≈ $12.3K |
| MSK | assumed, scaled with envelope bytes; confirm | ≈ $8.0K |
| Workers (router, delivery, push, media, key directory) | assume 100 × c7g.xlarge | ≈ $10.6K |
| Data transfer out (envelopes ≈ 44 MB/s, pongs ≈ 30, pushes ≈ 24, other ≈ 5: ≈ 104 MB/s ≈ 272 TB) | 10 TB × $0.09 + 40 × $0.085 + 100 × $0.07 + 122 × $0.05 (per GB) | ≈ $17.4K |
| Cross-AZ traffic | ≈ 236 TB × $0.02 per GB | ≈ $4.7K |
| S3 media storage | 50 TB × $0.023 + 450 × $0.022 + 2,050 × $0.021 (per GB) | ≈ $54.1K |
| S3 media PUTs | 170M × 1.04 × 30.4 ≈ 5.38B × $0.005 per 1,000 | ≈ $26.9K |
| CloudFront data out | first 5,024 TB across the tiers ≈ $139.8K + 4,123 TB × $0.020 ≈ $82.5K | ≈ $222.2K |
| CloudFront requests + S3 GETs on misses | 18.3B × $0.01 per 10,000 ≈ $18.3K; 80% miss × $0.0004 per 1,000 ≈ $5.9K | ≈ $24.2K |
| Backups | 900 TB × $0.0125 ≈ $11.25K + 72 TB × $0.023 ≈ $1.7K + PUTs ≈ $4.3K + restores ≈ $1.1K | ≈ $18.3K |
| Key vault (a few enclave instances; KMS) | ≈ $0.5K | |
| CloudWatch, SQS, misc. | ≈ $20.0K | |
| Total | ≈ $794K/month |
About 0.79 cents per daily user a month. Media is 41% of the bill ($327.4K), envelope writes 26%, gateways 18%. The server can't read anything, yet it does more work than before: every message becomes 7.5 stored and deleted envelopes.
Levers, in order:
- Reserved capacity for the steady part of the provisioned DynamoDB capacity (a one-year commitment; this table is single-Region, so it qualifies): a large cut on the $209K write line at list prices.
- Media auto-download policy: don't auto-download video on tablets and desktops. Cutting downloads from 3.54 to 2.5 per file saves roughly a quarter of the CloudFront line.
- Mailbox cursors instead of deletes (R3.7): about $100K of writes, traded for about $38–44K of storage, a race to handle, and ciphertext kept longer.
- Fewer post-quantum one-time prekeys per device: most of the key directory's storage.
R3.7 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| E2E vs server features | E2E for everything, including media and read receipts | Server search, link previews, server-written notifications, content moderation, server history. Metadata stays visible; hiding the sender from the server too ("sealed sender") is a further option with its own abuse-control costs. |
| Backup recovery vs security | Recovery key by default; PIN + guess-limited vault as an opt-in | With the key only: history lost with the key. With the vault: we operate hardware whose failure or compromise matters, and must prevent counter rollback. |
| Per-device fan-out vs group sender keys vs MLS | Pairwise sessions for one-to-one; sender keys for groups | Pairwise for a group of 8 would be 16.6 encryptions and ≈ 5 KB of upload per message from the sender's phone, instead of one. Sender keys make removals expensive: every remaining member rotates. MLS (the chat loop's choice) makes group changes cost roughly the log of the group size, but needs every change applied in one agreed order. The server's per-device mailbox writes are the same in all three. |
| Delete on ack vs cursor + TTL mailboxes | Delete each envelope when acknowledged | 7.53 extra writes per message (≈ $100K a month). The alternative: never delete, advance a per-device cursor, and let a 30-day ttl clean up: 12.8B × 400 B × 30 days ≈ 154 TB more storage (≈ $38K), or about 175 TB (≈ $44K) allowing for TTL deletion lag, envelopes kept longer, and a real race: many senders write one mailbox, so a time-ordered cursor can skip an envelope whose write finished late. It would need a re-read overlap window and writers that re-stamp slow writes. |
| SQLCipher vs OS file protection alone | Both | SQLCipher costs some CPU per page and a key to manage, and it protects copies of the file. Neither protects against code running inside our app on an unlocked phone. |
| Key available after first unlock vs only while unlocked | After first unlock | Notifications decrypt while the phone is locked; a stolen phone that is on and was unlocked once depends on the OS. |
Closing the loop. The opening question was: how do we make chat feel live on a device that keeps going to sleep, losing its network and getting killed? The answer is now:
- Live: the local database is the truth for the screen; a message is written before it's shown, and the network only fills and drains the database.
- Asleep: a socket only in the foreground; the OS's push connection in the background; delivery counts only when the device says so, and deadlines escalate to notifications the OS shows by itself.
- Losing its network: an outbox with IDs made before the first attempt, catch-up by cursor, resumable uploads the OS runs, and reconnects triggered by the OS's network callbacks.
- Killed, lost or replaced: everything important is durable on the device, encrypted with a key that never leaves it; a new device is a new endpoint with new sessions, restored from a backup whose key the user holds.
R3.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| Ratchet desync after a restore (someone restored an OS backup that included old state) | Decryption errors from one contact | Sessions are never in our backups, and the database is excluded from OS backups. If it happens anyway, the decryption-error control message makes the sender start a new session and re-encrypt from its local copy. |
| Key-directory outage | New sessions can't start; new devices can't be reached | Existing sessions keep working: their keys are on the devices. Sends to new contacts wait in the outbox, and the app says "waiting to connect securely". The directory is multi-AZ, and its bundles are served from a replica cache for reads. |
| Device-list change attack (a compromised server adds its own device to Bob's list) | Alice's phone would encrypt to an attacker's device | The list must be signed by one of Bob's existing devices; Alice's phone checks the signature and shows "Bob added a new device" in the chat. Safety numbers let users compare keys in person; a public, auditable log of device keys (key transparency) makes silent changes detectable at scale. |
| Notification extension timeouts | "New message" instead of the preview | Our internal deadline (20 s on iOS; about 5 s on Android, where longer work moves to an expedited WorkManager job) shows the fallback; the app decrypts on open. We alarm on the fallback rate per app version. |
| One-time prekeys run out (many new sessions to one device while it's offline) | Sessions start without a one-time key | The bundle falls back to the signed prekey (and a last-resort post-quantum key), which is still secure but gives the first message weaker forward secrecy. Devices replenish when they come online; we alarm on devices with fewer than 10 left. |
| Old app version in a device list (a tablet that hasn't updated) | It doesn't support the newest protocol or a new content type | Each device advertises capabilities in the directory; senders use what every recipient device supports, and old apps show "update to view" for unknown content types. After a published cutoff date, devices below the minimum version are dropped from device lists until they update. |
| Key vault unavailable | PIN recovery fails for a while | Recovery-key restores and device-to-device transfer still work; vault calls are retried, and a failed call doesn't count as a guess. |
| Push provider outage | No notifications on one platform | Messages wait in mailboxes; desktops and foreground apps still get them over sockets; push workers queue and coalesce (R3.9). |
R3.9 Runbook and Incident Response
| Signal | Source | Alarm | First action |
|---|---|---|---|
| Send success (ACK within 10 s) | Client telemetry | < 99% for 10 min | Server stage timings; 409 mismatch rate; a client release? |
Time to DELIVERED P95 | Server | > 2× baseline | Push delays? Deadline escalations? Mailbox write throttling? |
Push acknowledged rate (pushes followed by DELIVERED within 60 s) | Server | < 90% for 15 min | Provider status by platform; token failures; FCM quota |
| Reconnect time P99, by network | Client | > 2× baseline | Gateway load; NLB health; app version |
| Decrypt failures, counted as distinct (sender device, recipient device) pairs | Client | > 3× baseline for 15 min | Decrypt-failure procedure below |
| Upload failures | Client + S3 | > 5% for 15 min | Pre-signed URL expiry? Checksum mismatches? A client release? |
| Notification fallback rate ("New message" shown) | Client | > 5% of notifications | Extension timeouts by OS and app version |
| Devices low on one-time prekeys | Key directory | > 1% of active devices below 10 | Replenishment bug in a release? |
Counting decrypt failures by distinct device pairs matters: one broken session retrying a hundred times is one problem, not a hundred, and a spike in pairs is what a real bug looks like.
Procedure: a push-provider outage REL 11 · OPS 10
- Confirm it's the provider. Send failures by provider and error code (CLI 1), and the provider's status page. If only one of our push workers' connections fails, recycle the workers instead.
- Don't retry harder. Push workers already back off per provider; raising concurrency adds to the provider's load. Check that the queue is growing and coalescing (its size should be bounded by users × conversations, not messages).
- Tell users what still works. Foreground apps and desktops still receive over sockets. Mailboxes hold everything.
- When the provider recovers, the queue drains coalesced: one push per conversation with the newest
up_to_seq, never a burst of stale notifications. Watch FCM quota use during the drain. - Afterwards: compare time-to-
DELIVEREDduring the outage with the alarm, and review whether escalations behaved.
Procedure: a decrypt-failure spike SEC 10 · OPS 10
- Slice it. By app version, platform, message type (prekey or normal), and whether the recipient restored or reinstalled recently (CLI 2).
- One app version? Halt its rollout; if it's live, disable the new feature by remote flag. A client bug that corrupts session state can't be fixed server-side, but the decryption-error repair recovers each session as the sender's device retries.
- Prekey messages only? Check the key directory: bundles served with a wrong signature, a one-time key served twice (the conditional delete must make that impossible), or a stale device list cache. Invalidate the device-list cache (versions make stale entries detectable).
- Treat it as a possible attack if failures cluster on key changes for particular users: check device-list change rates and signatures, and involve the security team.
- Afterwards: a post-incident review; add the failing sequence to the client's protocol test suite.
Go deeper: CLI playbook
Plain commands an on-call engineer runs, one at a time. Replace names and times with real ones.
text# 1. Push send failures by provider (custom metric from the push workers) aws cloudwatch get-metric-statistics --namespace Chat/Push --metric-name SendFailures --dimensions Name=Provider,Value=FCM --statistics Sum --period 60 --start-time 2026-09-28T18:00:00Z --end-time 2026-09-28T19:00:00Z # 2. Decrypt failures by app version (custom metric from client telemetry) aws cloudwatch get-metric-statistics --namespace Chat/Client --metric-name DecryptFailurePairs --dimensions Name=AppVersion,Value=8.14.0 --statistics Sum --period 300 --start-time 2026-09-28T18:00:00Z --end-time 2026-09-28T19:00:00Z # 3. Mailbox table write throttling aws cloudwatch get-metric-statistics --namespace AWS/DynamoDB --metric-name WriteThrottleEvents --dimensions Name=TableName,Value=Mailboxes --statistics Sum --period 60 --start-time 2026-09-28T18:00:00Z --end-time 2026-09-28T19:00:00Z # 4. Scale push workers aws ecs update-service --cluster chat --service push-workers --desired-count 60 # 5. Confirm the media bucket's lifecycle rules (abort incomplete uploads, 30-day expiry) aws s3api get-bucket-lifecycle-configuration --bucket chat-media-prod
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | Out-of-order and lost-session handling on devices; key-directory outages don't break existing sessions; encrypted backups and device-to-device restore; revocation stops queueing for dead devices REL 9 · REL 11 |
| Performance Efficiency | Sender keys keep group sends to one encryption; previews decrypted in ~1–2 s inside the extension's limits; device lists checked from a versioned cache PERF 1 · PERF 3 |
| Security | End-to-end encryption with the Signal Protocol; per-device identities and signed device lists; SQLCipher with a device-bound key; backups we can't read; revocation with short tokens and an in-memory list SEC 2 · SEC 8 · SEC 9 · SEC 10 |
| Cost Optimization | ≈ $794K/month, derived; media and per-device envelope writes named as the drivers; reserved capacity, download policy and the mailbox-cursor option as levers COST 5 · COST 7 · COST 8 |
| Operational Excellence | Decrypt failures counted by distinct device pairs; push-outage and decrypt-spike procedures; client telemetry by app version OPS 8 · OPS 10 |
| Sustainability | Envelopes deleted on acknowledgement; media expires after 30 days; backups only on Wi-Fi while charging; no auto-download of video on secondary devices SUS 3 · SUS 4 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Distinguishes "encrypted" from "end-to-end", and explains session setup and the Double Ratchet at the level of what each key does, without hand-waving.
- Bounds skipped keys and explains why both "drop" and "keep everything" are wrong; handles duplicate delivery paths and cross-process state.
- Treats every device as an endpoint, verifies device lists, and counts the per-device fan-out.
- Makes backup recovery an explicit key-custody decision, keeps sessions out of backups, and names the history that can be lost.
- Knows what secure hardware can and can't hold, and the lock-state trade-off notifications force.
- Prices the system and sees that the server does more work, not less, after it stops reading messages.
Follow-up questions
-
"Can we detect spam if we can't read messages?" Answer: partly. Metadata signals (new accounts sending to many strangers, fan-out rates, reports) work without content, and a reported message arrives decrypted from the reporter's device. What we can't do is scan every message; that's the product decision end-to-end encryption makes.
-
"A user has 5 devices and joins a 256-member group. What happens?" Answer: each of their devices needs the group's sender keys from every member device, and every member device needs theirs: one sender-key message from each of their 5 devices to up to member devices, over pairwise sessions, plus fetching bundles for any devices they haven't talked to. That's a burst of thousands of small envelopes on a join, which is why very large groups favour MLS.
-
"Why not store message history on the server, encrypted with a key all the user's devices share?" Answer: that's a legitimate design (some apps do it): a per-user history key, distributed to each new device by an existing one. It gives new devices history without a backup. The costs: one key that decrypts everything a user ever had, held on every device, and server-side storage of all history. We chose device-held history with backups; a product that needs seamless history on new devices might choose the other.
Interview gotchas from this round
| Gotcha | Why it's wrong |
|---|---|
| "TLS plus encryption at rest is end-to-end" | We hold the keys, so we can read everything. |
| "Store the Signal keys in the Secure Enclave" | The Secure Enclave works with P-256 keys; Curve25519 keys live in the encrypted database and in memory while used. |
| "Back up the sessions too" | Restored ratchet state desyncs and reuses keys. A restored phone is a new device. |
| "Keep all skipped keys forever" | Undoes forward secrecy and lets a malicious sender exhaust the phone. |
| "One key shared across a user's devices" | One compromise exposes all devices, and revocation means re-keying everyone. |
| "The lock screen protects a lost phone" | Files can be copied; only encryption with a device-bound key protects them. |
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scoping: offline send? history on the phone? receipts? encryption? | Restate the Round 1 design in 60 seconds | Restate the Round 2 design in 60 seconds |
| 5–15 min | Requirements + API: frames, client message ID, cursors | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Steps 1.0–1.5: local write → outbox → client ID → foreground socket → catch-up → server order | Steps 2.1–2.6: OS push; deadlines and escalation; heartbeats and handoffs; resumable media; watermarks; state machine | Steps 3.1–3.6: Signal Protocol; skipped keys; devices; backups; lost phone; encrypted previews |
| 40–50 min | Numbers: traffic, local storage per chat, sockets, cost | Numbers: receipts per recipient, frame bytes, heartbeat rates, pushes, media, cost | Numbers: envelopes per message, key directory, backups, cost |
| 50–60 min | Failures + pillar check | Failures, battery budget, pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint.
The Two Sentences That Matter Most
- Opening a round: "Before I design: what must work offline, what happens when the app is in the background or killed, and who is allowed to read the messages?"
- When the scope is raised: "Here's what breaks, and I'll fix it in this order: anything that can lose or duplicate a message, then anything that drains the battery or depends on the OS behaving, then cost."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "What if the app is killed mid-send?" (REL 4) | The message was written to the local outbox before it was shown, with an ID made before the first attempt, so the next launch sends it once. | 1 | Steps 1.1, 1.2 |
| "How do you know a message was delivered?" (REL 5) | Only when the device acknowledges it; a deadline escalates to a notification the OS shows by itself. | 2 | Step 2.2 | |
| "What happens to history on a new phone?" (REL 9) | It comes from another device or an encrypted backup whose key the user holds; sessions are always new. | 3 | Step 3.4 | |
| Performance | "How fast is reconnecting after a network change?" (PERF 4) | Three round trips from the OS's callback, about 170 ms on a good network; we measure detection separately. | 2 | Step 2.3 |
| "Why does the chat screen never spin?" (PERF 3) | It only reads the local database; the network fills it. | 1 | Step 1.1 | |
| Security | "Can you read users' messages?" (SEC 9) | No: devices encrypt with the Signal Protocol and we route sealed envelopes; we see metadata only. | 3 | Step 3.1 |
| "What if a phone is stolen?" (SEC 8) | The database is encrypted with a device-bound key, and another device revokes it: tokens, sockets, mailbox and push token. | 3 | Step 3.5 | |
| "How do you stop a fake device being added?" (SEC 2) | Device lists must be signed by an existing device, and every change shows a notice. | 3 | Step 3.3, R3.8 | |
| Cost | "Where does the money go?" (COST 5) | About $7.3K, $168K and $794K a month; media delivery is the biggest line, then per-device envelope writes and gateways. | 1–3 | R1.7, R2.6, R3.6 |
| "Why not SNS for push?" (COST 11) | At 300M pushes a day it costs ≈ $9.1K a month and hides the headers we tune; a small team should still use it. | 2 | R2.7 | |
| Operations | "How do you know the clients are healthy?" (OPS 8) | Client telemetry by app version and network: send success, reconnect time, decrypt failures by distinct device pair, battery wake-ups. | 2–3 | R2.10, R3.9 |
| Sustainability | "Why does your app not drain the battery?" (SUS 3) | No socket in the background; the OS's shared push connection instead; heartbeats only when idle; about 0.3% a day. | 2 | Step 2.3 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| The device as a system | Local database as the screen's truth; one writer; outbox survives kills. | App states and OS rules on both platforms; background transfers; a battery budget in wake-ups. | Secure hardware, encrypted database, cross-process session state, and what a lost device exposes. |
| Exactly once | Client message ID before the first attempt; reconciliation against catch-up. | Delivery proven by the device; forward-only cursors; pushes that replace instead of stack. | Processed-envelope IDs in the decrypt transaction; resend from the sender's copy on decryption errors. |
| Connectivity | Foreground socket; catch-up by cursor. | Push as the background path; deadlines and escalation; adaptive heartbeats; network callbacks. | Per-device mailboxes and pushes; revocation that closes sockets. |
| Security | TLS; identity from the connection. | Ownership checks on every reference; pre-signed, scoped URLs. | End-to-end encryption, signed device lists, backups we can't read, key custody as a product decision. |
| Numbers | Messages, local storage per chat, sockets, a cost. | Re-checks carried figures (receipts per recipient, heartbeat rate blending, headroom per AZ); finds media dominates. | Envelopes per message, key directory, backups; sees the server's work grow under E2E. |
| Evolving under new scope | Builds from the baseline one problem at a time. | Opens with what the OS will and won't allow. | Changes what the server is allowed to know, and says what users can lose. |