Design an Offline-First Mobile News Feed
This page is one interview loop in three rounds. All three rounds design the same system. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
In this loop the phone is part of the system. It has its own database, a flaky network, a battery, and an operating system that can kill the app at any moment. Every round designs both halves: the app on the device and the backend it talks to. How the server builds each user's feed (fanout, timelines, ranking) is the subject of the news feed loop; here we link to it instead of teaching it again.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | A news app that opens instantly and shows the last feed with no signal | A social app where people like, comment and post while offline | iOS, Android and web; dozens of app versions in the field; frequent releases |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Traffic | 1M DAU; 10M refreshes/day (≈ 116/s, ≈ 347/s at peak) | 50M DAU; 1B delta syncs/day (≈ 11,574/s, ≈ 35K/s at peak); 250M offline actions/day | 200M installs, 80M DAU; 1.6B syncs/day (≈ 55.6K/s at peak); 30 app versions in use |
| On the device | ~100 MB of cache; no disk or network work on the main thread | + an outbox of offline actions; ≤ ~4 KB per sync; < 1.5% battery a day for sync | + schema migrations that keep queued actions; encrypted data; safe on a lost phone |
| Targets | Cached content on screen fast after launch; 99.9% API | Every offline action applied exactly once in effect; 99.99% sync API | A bad feature switched off in minutes; crash-free users ≥ 99.5% per version |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
Loop Opener: What Is an Offline-First App?
You Already Use One: a Notebook You Can Read on the Plane
Think of a notebook where a friend copies the latest news for you every morning. On the plane you can't call anyone, but you can still read the notebook, and you can scribble notes in the margin. When you land, you hand the notes over and get the pages you missed.
An offline-first app works the same way. It keeps its own copy of your feed on the phone, so it opens instantly and still works in a tunnel. The network is only used to catch up: download what changed, upload what you did.
| You do | The app does |
|---|---|
| Open the app in a subway tunnel | Shows the feed it saved last time, straight from the phone's storage |
| Pull to refresh with one bar of signal | Downloads only what changed since the last refresh, in the background |
| Like a post with no signal | Shows the like at once, writes it down locally, and sends it when the network returns |
| Close the app, and the OS kills it | Nothing is lost: everything important was already saved on the phone |
What Makes It Hard
- The network drops constantly, often in the middle of a request, and we don't know whether the server got it.
- The OS kills apps at any moment to free memory, and it decides when (and whether) the app may run in the background.
- Batteries and data plans are precious. An app that wakes the radio every minute gets noticed, uninstalled, and restricted by the OS.
- Offline actions must not be lost or applied twice. A like sent three times because of three retries must count once.
- You can't update every phone at once. Last year's app version is still out there, talking to today's server.
The Question the Whole Loop Answers
How do we make the app instant and fully usable offline, and still keep it in agreement with the server?
The answer gets sharper every round:
- Round 1: the screen reads from a database on the phone, never from the network. The network only fills that database.
- Round 2: actions made offline go into a durable queue on the phone and reach the server exactly once in effect; refreshes download only what changed, deletions included.
- Round 3: the hard part is time. Dozens of app versions, database upgrades on millions of phones, and releases we can't recall, so the server has to stay compatible and able to switch features off remotely.
Round 1 · Mid-level · "A Feed That Opens Instantly, Even Offline"
~35 min · SDE II (L5) · 1 region, 3 AZs · 1M DAU · ~347 refreshes/s peak · 99.9% API · ~100 MB on the device
R1.1 Establish Design Scope
The interviewer says: "Our news app shows a spinner every time it opens, and a blank screen on the subway. Design the feed so it opens instantly and works offline." We ask before we draw.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Must people be able to read offline? | Yes. The last feed they saw, with images. | The app needs its own copy of the feed on the phone, not just a copy of the last network response in memory. |
| Can they act offline (like, comment, save)? | Not yet. Reading only. | No offline write path this round. Round 2 adds it. |
| How much should we keep? | About the last 200 items. | A few sessions of reading. It fixes our storage budget (R1.7). |
| Images too? | Yes, cached, so an offline feed isn't a wall of grey boxes. | An image cache on disk with a size limit, and images sized for the phone's screen (step 1.4). |
| Which platforms? | iOS and Android. | Two apps with the same design. The web comes in Round 3. |
| How fresh? | Refresh when the app opens, and on pull to refresh. | No background sync yet. Opening shows the saved feed at once, then refreshes quietly. |
| How many users? | About 1M daily users, each opening the app about 10 times a day. | We derive the traffic in R1.7. |
Out of scope for this round: offline actions, syncing only the changes, deletions reaching the phone, and background refresh.
The interviewer will widen this scope later. Write your out-of-scope list where you can see it: in a multi-round loop, some of it comes back.
R1.2 Functional Requirements, Derived Step by Step
| Phrase from the problem | Requirement |
|---|---|
| "Opens instantly" | On launch, show the saved feed before any network call finishes |
| "Refresh" | On open and on pull to refresh, fetch the newest page and merge it into the saved feed |
| "Keep scrolling" | Page to older items with a cursor, saving them too |
| "Works on the subway" | Everything above (text and images) works with no network, from what was saved |
Not yet: actions while offline, and a refresh that sends only what changed. The interviewer will bring them back.
R1.3 Non-Functional Requirements: the Questions
Numbers come in R1.7. For now, the questions and why each matters:
- Time to content. What do we measure, from when? A cold start (the app's process isn't running) includes the OS launching the process, which we don't control and which takes hundreds of milliseconds or more on a mid-range phone. So we measure two things separately: time to first frame (the platform's launch metric), and time to cached content: from the first frame to the saved feed on screen. That second number is ours to make small.
- Smooth scrolling. A display refreshing at 60 Hz draws a frame every ms; at 120 Hz, every ms. Any work on the main thread (the one thread that draws the UI and handles touches) that takes longer than that drops a frame, and the user sees a stutter.
- Storage. A phone is shared with every other app and with the user's photos. How much may we take, and what happens when the phone is nearly full?
- Data usage. Many users pay per gigabyte. How many bytes does one day of use cost them?
R1.4 The API
One endpoint for the feed, and images from a CDN. Every call carries the app's normal login token.
1. The newest page (a refresh).
httpGET /v1/feed?limit=20 HTTP/1.1 Host: api.example-news.com Authorization: Bearer <token> Accept-Encoding: gzip If-None-Match: "h-8f2c91"
httpHTTP/1.1 200 OK Content-Type: application/json Content-Encoding: gzip ETag: "h-8f2d07" { "items": [ { "post_id": "97301141913751557", "sort_key": 23198400000, "title": "City opens a new bike bridge", "summary": "The 400 m bridge links the two riverbanks...", "author": { "id": "401", "name": "Metro Desk" }, "image": { "id": "7Qz2", "width": 1600, "height": 900, "url_template": "https://img.example-news.com/i/7Qz2/w{width}.{format}", "widths": [360, 720, 1080], "formats": ["avif", "webp", "jpg"] } } ], "next_cursor": "eyJ2IjoxLCJzIjoyMzE5ODM5OTg3NiwicCI6Ijk3MzAxMTQxMzk2OTY0MzIwIn0.k3Jd9w", "server_time": "2026-09-28T08:00:00.120Z" }
ETagandIf-None-Match. An ETag is a version label for a response. The app sends back the last one it saw; if the newest page hasn't changed, the server answers304 Not Modifiedwith no body, a few hundred bytes instead of a full page.sort_keyis the server's order for the feed (here, the post's time in milliseconds from the ID). The app sorts by it and never by its own clock.url_templatelets the app pick the image width that fits its screen (step 1.4).
2. Older items.
httpGET /v1/feed?limit=20&cursor=eyJ2IjoxLCJzIjoyMzE5ODM5OTg3NiwicCI6Ijk3MzAxMTQxMzk2OTY0MzIwIn0.k3Jd9w HTTP/1.1 Host: api.example-news.com Authorization: Bearer <token>
The cursor is opaque: a signed token saying "continue after this item". The app stores it and sends it back, and never builds or edits one. The server side of this cursor (why offsets repeat items, and how the signature works) is step 1.3 of the news feed loop.
| Status | When | What the app does |
|---|---|---|
200 OK | A page | Save it, then the screen updates from the database |
304 Not Modified | Nothing new at the top | Nothing; update "last checked" |
400 Bad Request | A cursor fails its signature check | Drop the cursor; page again from the top |
401 Unauthorized | Token expired | Refresh the token, retry once |
429 / 503 | Overload or outage | Keep showing the saved feed; retry later with backoff |
Recap
- Two calls: the newest page (with an ETag) and older pages (with a cursor).
- Images come from a CDN, in sizes the app chooses.
- The saved feed must show before any of these calls finishes.
Let's build it, starting with the simplest thing that works.
R1.5 Design Evolution: From "Call the API and Render" to a Local Database
Every step below follows the same pattern: a problem, your turn to think, the answer, and what the answer costs us. The cost is always the next problem.
Step 1.0: The Baseline
The screen calls the API when it opens, parses the JSON, keeps the list in memory and draws it. Images are downloaded as each row scrolls into view.
Synthesizing vector architecture diagram...
The screen owns the data, and the data lives only as long as the screen does.
What's good about it: it is simple, and what you see is always what the server just said.
What it costs us: on a slow network the user stares at a spinner; with no network they get an error screen; and every launch downloads everything again.
Step 1.1: The App Shows a Spinner on Every Launch
The problem: users open the app on the train. For three to ten seconds they see a spinner, and in a tunnel they see "No connection". They read these same stories an hour ago. What would you do?
Step 1.2: Scrolling Stutters
The problem: the saved feed now shows instantly, but scrolling stutters every time a refresh lands, and some Android users get "App isn't responding" dialogs on old phones. What would you do?
Step 1.3: Reads Block Writes During a Refresh
The problem: a refresh writes 20 items while the screen's query is reading. With SQLite's default setup, the writer waits for readers and readers wait for the writer, so the list freezes for a moment, and saving 20 items one by one takes far longer than it should. What would you do?
Primitive: Write-Ahead Log & LSM-Trees
Step 1.4: Images Reload Every Time and Eat the User's Data
The problem: the text now loads from the phone, but every image is downloaded again on each launch, at full 1600 px size, even to show a 360 px thumbnail. Users on metered plans complain about data use, and offline the images are grey boxes. What would you do?
Primitive: Distributed Cache Patterns & Eviction · Drill: CDN edge image delivery
Step 1.5: The Cache Fills the Phone
The problem: a heavy reader scrolls deep every day. After a month the app uses 1.2 GB, the phone warns "Storage almost full", and the app store's storage screen names us. What would you do?
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | Screen calls the API and keeps the list in memory | Spinners; blank screen offline; everything re-downloaded |
| 1.1 | Spinner every launch | SQLite as the single source of truth; the screen observes a query; the network only writes | Two copies of the feed; a schema on every phone |
| 1.2 | Scrolling stutters | No disk or network work on the main thread; off-thread decode and downsampling | Threading discipline |
| 1.3 | Reads block writes | WAL mode; one transaction per refresh; automatic checkpoints | WAL maintenance; queries must close promptly |
| 1.4 | Images reload and eat data | CDN width and format variants; immutable URLs; memory + disk LRU cache | An eviction policy; variant generation |
| 1.5 | The cache fills the phone | 200 items; 90 MB of images; cache files where the OS may clear them | Old items gone offline |
Two costs stay open for Round 2: the phone can't act offline, and a refresh downloads the whole newest page even when one item changed.
R1.6 Architecture v1
Now the concepts get AWS names and platform names.
Synthesizing vector architecture diagram...
Follow the arrows into the screen: there is only one, from the database. The refresh worker is the only thing that talks to the feed API, and it only writes to the database. Images are the one thing the screen asks for by URL, and they come from the phone's cache first, then CloudFront.
The pieces:
- Feed API on ECS (Fargate) behind an Application Load Balancer in three AZs. It reads each user's feed from the feed service the news feed loop designs (a precomputed timeline plus a batch read of the posts), and computes the ETag.
- S3 and CloudFront for images. An upload triggers a small job that writes the three widths in each format.
- On the phone: SQLite in WAL mode, an observed query, a refresh worker on a background thread, and an image cache.
The local schema (SQL; the same on both platforms):
sqlCREATE TABLE feed_item ( post_id TEXT PRIMARY KEY, sort_key INTEGER NOT NULL, -- server's order; never the phone's clock title TEXT NOT NULL, summary TEXT, author_id TEXT NOT NULL, author_name TEXT NOT NULL, image_id TEXT, -- builds URLs from the template image_w INTEGER, image_h INTEGER, saved_at_ms INTEGER NOT NULL -- phone's clock; only for "last updated" on screen ); CREATE INDEX feed_item_order ON feed_item (sort_key DESC, post_id); CREATE TABLE feed_meta ( -- one row per key key TEXT PRIMARY KEY, -- 'head_etag', 'older_cursor', 'last_server_time' value TEXT NOT NULL );
The screen's query is SELECT ... FROM feed_item ORDER BY sort_key DESC, post_id DESC LIMIT 200, served by the index.
Trace 1: a cold launch in a tunnel.
Synthesizing vector architecture diagram...
Time to cached content is measured from the first frame to the rows on screen, and it doesn't depend on the network at all.
Trace 2: a refresh with new items.
Synthesizing vector architecture diagram...
If nothing changed, the API answers 304 and no row is written at all.
R1.7 Numbers
Traffic (assumptions: 10 refreshes per user per day; older-page requests are few and folded into that number; peak is 3× the average)
| Quantity | Math | Value |
|---|---|---|
| Refreshes per day | 1M × 10 | 10M |
| Refreshes/s, average | 10,000,000 ÷ 86,400 | ≈ 116 |
| Refreshes/s, peak | 115.7 × 3 | ≈ 347 |
Payloads (assumptions: a full page of 20 items is about 60 KB on the wire after gzip, including author objects and image metadata; 20% of refreshes find nothing new and get a 304; about 5 new images per refresh on average, counting the 304s, at a ~60 KB phone-sized variant)
| Quantity | Math | Value |
|---|---|---|
| API bytes per day | 10M × 80% × 60 KB | 480 GB |
| API bytes per month | 480 GB × 30.4 | ≈ 14.6 TB |
| API egress at peak | 347 × 60 KB ≈ 20.8 MB/s | ≈ 167 Mbps |
| Image bytes per day | 10M × 5 × 60 KB | 3 TB |
| Image bytes per month | 3 TB × 30.4 | ≈ 91.2 TB |
| Image requests per month | 10M × 5 × 30.4 | ≈ 1.52B |
Per user, per day: MB, about 106 MB a month. Without the disk cache, every refresh would reload all 20 images: MB a day for images alone.
On the phone
| Store | Math | Value |
|---|---|---|
| Feed rows | 200 items × ~3 KB (text, author, image metadata, index) | ≈ 0.6 MB |
| WAL file | up to 1,000 pages × 4 KB (SQLite's default; Android's framework default is smaller), checkpointed | ≤ ~4 MB in normal use |
| Image disk cache | cap | 90 MB |
| Total | < 100 MB |
Monthly cost (us-east-1 on-demand list prices; CloudFront at United States and Europe rates; rounded)
| Item | Math | Monthly |
|---|---|---|
| CloudFront image data out | 10 TB × $0.085 + 40 TB × $0.080 + 41.2 TB × $0.060 | ≈ $6,520 |
| CloudFront image requests | 1.52B HTTPS requests × $0.01 per 10,000 | ≈ $1,520 |
| API data out through the ALB | 10 TB × $0.09 + 4.6 TB × $0.085 (EC2 data-transfer-out tiers) | ≈ $1,290 |
| Feed API, 3 Fargate tasks (Graviton, 1 vCPU, 2 GB) | 3 × ($0.03238 + 2 × $0.00356)/h × 730 h | ≈ $90 |
| DynamoDB post reads | only changed pages are rebuilt: 8M × 20 items × 0.5 RRU = 80M RRU/day ≈ 2.43B × $0.125 per million | ≈ $300 |
| ALB | processed bytes dominate: 14.6 TB a month ≈ 20 GB an hour ≈ 20 LCU × $0.008 × 730, + $0.0225 × 730 | ≈ $133 |
| S3 images and variant generation | originals + 3 widths × 3 formats, a year in | ≈ $230 |
| CloudWatch, misc. | ≈ $200 | |
| Total | ≈ $10.3K/month |
Why 3 API tasks: one task handles about 250 refreshes a second (an assumption to confirm by load test), so the 347/s peak needs 2; losing an AZ must leave 2, so we run 1 per AZ, 3 in all.
The feed service's own cost (timelines and fanout) belongs to the server loop; this table is what the offline-first app adds or resizes.
Say the headline: delivery is almost the whole bill, and the phone's cache is what keeps it small. Without the disk cache, images would be about 365 TB a month (≈ $18.6K) plus 6.1B requests (≈ $6.1K), about $24.7K instead of $8K. The cache also saves the user about 9 MB a day.
R1.8 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| A database vs a file or HTTP cache on the phone | SQLite: query by order, update one item, merge pages, all in transactions | A schema to migrate. An HTTP cache can replay the last response, but can't merge pages or update one item, and it's evicted by rules we don't control. |
| SQLite vs a key-value store | SQLite for the feed; a key-value store (DataStore, UserDefaults) only for small settings | Key-value stores are great for a handful of values; for "the newest 200 by sort_key, then page older" we'd rebuild indexes and transactions by hand. |
| How much to cache | 200 items, 90 MB of images | Items past 200 aren't there offline. More costs the user storage for items they rarely reread. |
| Refresh on open vs in the background | On open and pull to refresh only | The first seconds after opening show slightly old items. Background refresh arrives in Round 2, with the OS's limits. |
R1.9 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| A refresh interrupted mid-write (app killed, battery dies) | Nothing | One transaction per refresh: SQLite rolls the unfinished one back on next open. The old feed stays whole. |
| A corrupted database file (rare; a failing flash chip, a bug) | A query returns SQLITE_CORRUPT, or a startup PRAGMA quick_check fails | In Round 1 the database holds only a copy of server data, so we delete it and fetch a fresh page. Round 2 adds data that exists only on the phone, and this gets harder. |
| Low storage | Writes fail with SQLITE_FULL; the OS clears cache directories | Shrink the image cache first, stop fetching off-screen images, keep the text. The database is small and trimmed to 200 rows. |
| The API is down | Refreshes fail | The app shows the saved feed with a "last updated" note, and retries with backoff. That's the point of the design. |
| A bad image variant (a broken encoder output) | Broken images for one format | Fall back to the next format in the list; purge the variant path in CloudFront; since URLs are immutable, re-encode under a new image ID. |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | The app works with no network, from the local database; a refresh is one transaction, so a crash never leaves half a feed; the API runs in three AZs REL 11 · REL 10 |
| Performance Efficiency | The screen reads local data only; no disk or network work on the main thread; WAL so reads and a write don't wait on each other; images sized for the screen PERF 3 · PERF 4 |
| Security | Light this round: the database and images live in the app's private storage, encrypted at rest by the OS; tokens on every call SEC 8 |
| Cost Optimization | About $10.3K/month; delivery is most of it, and the disk cache and ETags cut it by roughly two thirds COST 8 |
| Operational Excellence | Skipped this round. |
| Sustainability | Bytes the user doesn't need are never sent: 304s, right-sized AVIF and WebP variants, a disk cache instead of re-downloads SUS 3 · SUS 4 |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Makes the local database the only source for the screen, and says why an in-memory list isn't enough (the OS kills apps).
- Defines what "instant" measures: time to cached content after the first frame, not the OS's process launch.
- Keeps every disk and network call off the main thread, and knows the frame budgets (16.7 ms at 60 Hz, 8.33 ms at 120 Hz).
- Knows what WAL mode changes (readers and the writer don't block each other; still one writer) and batches writes in a transaction.
- Serves images as CDN variants with immutable URLs and caches them on disk with a size cap.
- Derives the traffic, bytes per user and the monthly cost, and sees that delivery dominates.
Follow-up questions
-
"Why not use the HTTP cache instead of a database?" Answer: it works for images, and we use it for them. For the feed, an HTTP cache stores whole responses keyed by URL. It can't answer "the newest 200 items across the last five refreshes", can't change one item in place, and its eviction isn't ours to control. The feed needs queries and transactions, which is what a database is for.
-
"The phone's clock is wrong by a day. Does anything break?" Answer: ordering doesn't: we sort by the server's
sort_key. "Updated 5 minutes ago" would be wrong if computed from the phone's clock, so we compute it fromserver_time: the app stores the difference between the server's time and its own clock at each refresh, and uses it for display only. -
"How would you know if time to content got worse in the field?" Answer: the app records first frame to cached content, and first frame itself, and reports percentiles per app version (Round 3 builds this telemetry). Both platforms also collect launch times (Android vitals, iOS MetricKit), which catch regressions we didn't instrument.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Keep the feed in memory" | The OS kills background apps; the next launch starts empty. |
| "The cold start is under 16 ms" | A cold start includes the OS launching the process. Measure time to cached content after the first frame separately. |
| "WAL means many writers at once" | WAL lets readers and one writer proceed together. Writers still take turns. |
| "Download full images and scale on the phone" | The user pays for bytes that are thrown away; use CDN variants. |
| "Put the database in the cache directory" | The OS may delete it when space is low, taking offline reading with it. |
Round 2 · Senior · "Likes and Posts Offline, Synced Safely at 50M DAU"
~40 min · Senior SDE (L6) · 1 region, 3 AZs · 50M DAU · ~35K syncs/s and ~8.7K offline actions/s at peak · 99.99% sync API · ≤ ~4 KB per sync · < 1.5% battery a day
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "We built a news app for 1M daily users that opens instantly and reads offline. The rule is: the screen reads from a SQLite database on the phone, never from the network. The screen observes a query; a refresh worker on a background thread fetches the newest page and writes it in one transaction, and the screen updates because the table changed. SQLite runs in WAL mode, so the screen's reads and the refresh's write don't wait on each other, though writes still take turns. No disk or network work on the main thread: frames are 8.33 ms at 120 Hz. Images come from CloudFront as fixed width and format variants with immutable URLs, into a 90 MB disk cache. We keep the newest 200 items. About 347 refreshes a second at peak and about $10.3K a month, almost all of it delivery. Two costs are open: the phone can't act offline, and every refresh downloads a whole page even if one item changed."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: one path into the screen, from the local database.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Spinner every launch | SQLite as the single source of truth for the screen | Two copies of the feed |
| 1.2 | Scrolling stutters | Nothing slow on the main thread | Threading discipline |
| 1.3 | Reads block writes | WAL mode; one transaction per refresh | WAL maintenance |
| 1.4 | Images reload | CDN variants; immutable URLs; disk LRU cache | Eviction policy |
| 1.5 | The cache fills the phone | 200 items, 90 MB of images | Old items gone offline |
Open costs: no offline actions; full-page refreshes.
R2.1 The Scope Raise
Interviewer: "We're a social app now: 50 million daily users. People like, comment and post on the subway, and they expect to see it happen at once. Refreshes must send only what changed, because our users are on expensive data plans. The OS kills our app in the middle of uploads, phone clocks are all over the place, and sometimes the server has to refuse an action. And the battery team says sync may use at most 1.5% of the battery a day."
A scope raise is not the end of scoping. We ask back, and say what each answer changes.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| Which actions must work offline, and must they show at once? | Like, unlike, comment, post (with a photo). Yes, instantly, with no spinner. | An optimistic local write plus a durable queue on the phone (step 2.1). |
| How much traffic? | 50M DAU; about 20 syncs and 5 actions per user per day; around 200M installs in total. | 1B syncs and 250M actions a day (R2.6). |
| May a refresh download the whole page? | No. Only what changed, ideally a few KB. | A change log per user and a server-assigned cursor (step 2.3). |
| What happens when the OS kills the app mid-upload? | Nothing may be lost, and nothing may be applied twice. | Idempotency keys created on the phone, and a queue item removed only after the server acknowledges (step 2.2). |
| Can we trust the phone's clock? | No. Some are days off; some users change them on purpose. | No device timestamps in any cursor or ordering decision (step 2.3). |
| Can the server refuse an action? | Yes: the post was deleted, the user was blocked, the comment breaks the rules. | Error classes, a quarantine, and rolling back what the screen showed (step 2.5). |
| Must deleted posts vanish from phones? | Yes, including phones that were offline when it happened. | Tombstones in the change stream, with a retention window (step 2.4). |
| What's the budget for background sync? | Under 1.5% of battery a day, and respect the user's data plan. | OS-scheduled background work, prefetch by network type, and few wake-ups (step 2.6). |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Users | 1M DAU | 50M DAU (≈ 200M installs) |
| Reads | 10M refreshes/day, a full page each | 1B syncs/day ≈ 11,574/s, ≈ 35K/s at peak; ≤ ~4 KB each |
| Writes from phones | None | 250M actions/day ≈ 2,894/s, ≈ 8.7K/s at peak, many made offline |
| Offline | Read only | Read and write; nothing lost, nothing applied twice |
| Deletes | Not propagated | Reach every phone, including ones that were offline |
| On the phone | ~100 MB | 500 items, 150 MB of images, a bounded queue; < 1.5% battery/day for sync |
| Availability | 99.9% | 99.99% for the sync API (4.4 min/month) |
R2.2 What Breaks in the Round 1 Design
| Round 1 choice | What breaks at the new scope |
|---|---|
| Every refresh downloads the newest page | 1B × 60 KB = 60 TB a day of API egress, and a page that says nothing about items that changed or were deleted further down. |
| No write path | Tapping "like" in a tunnel either fails or spins. |
| A request per action, retried on failure | A like whose response was lost is sent again; without a key, the server counts it twice. |
| Anything ordered by the phone's clock | A phone set two hours ahead asks for "changes since two hours from now" and silently skips them. |
| "Deleted posts drop out of the next page" | A post deleted further down, or while the phone was off, stays on the phone forever. |
| Refresh only on open | Fine, but anything more (background refresh, prefetching) must now live inside the OS's rules and the battery budget. |
R2.3 New Requirements and API Additions
1. Delta sync. The phone sends its sync cursor (an opaque token for "how far I've read my change log") and the IDs of the posts it's showing, so their counters can be refreshed in the same call.
httpPOST /v1/sync HTTP/1.1 Host: api.example-social.com Authorization: Bearer <token> Content-Type: application/json Accept-Encoding: gzip { "cursor": "c1.ep_Hx7.450912.Qm8aZ1", "limit": 100, "visible_post_ids": ["97301141913751557", "97301141861222400"] }
httpHTTP/1.1 200 OK Content-Type: application/json Content-Encoding: gzip { "entries": [ { "seq": 450913, "op": "ADD", "version": 3, "post": { "post_id": "97301143011622913", "sort_key": 23198431000, "author": { "id": "802", "name": "Priya" }, "text": "Sunset from the ferry", "image": { "id": "m_91xQ", "width": 1080, "height": 1350 }, "like_count": 12, "comment_count": 2, "viewer_liked": false } }, { "seq": 450914, "op": "UPDATE", "version": 5, "post": { "post_id": "97301141861222400", "text": "Edited: the meetup moved to 7pm" } }, { "seq": 450915, "op": "DELETE", "version": 8, "post_id": "97301141396964320" } ], "counters": { "97301141913751557": { "like_count": 131, "comment_count": 9, "status": "ACTIVE" } }, "cursor": "c1.ep_Hx7.450915.pT3xLw", "has_more": false, "server_time": "2026-09-28T08:00:00.120Z" }
A DELETE entry is a tombstone: a record that says "this was deleted", so a phone can remove its copy. If the phone's cursor is older than the log still holds, the answer is 410 Gone with "error": "CURSOR_EXPIRED", and the phone rebuilds its feed from the head:
httpGET /v1/sync/head?limit=20 HTTP/1.1 Authorization: Bearer <token>
It returns the newest 20 posts (about 60 KB) and a fresh cursor; older posts are fetched as the user scrolls. The shape is the same as above.
2. Batch mutations. Queued actions go up in batches, each with an idempotency key created on the phone (client_mutation_id, a UUIDv7):
httpPOST /v1/mutations:batch HTTP/1.1 Host: api.example-social.com Authorization: Bearer <token> Content-Type: application/json { "mutations": [ { "client_mutation_id": "01927e3a-7f12-7000-8000-123456789abc", "type": "LIKE", "post_id": "97301141913751557" }, { "client_mutation_id": "01927e3a-7f13-7000-8000-123456789abd", "type": "COMMENT", "post_id": "97301141913751557", "payload": { "text": "Great shot!" } } ] }
httpHTTP/1.1 200 OK Content-Type: application/json { "results": [ { "client_mutation_id": "01927e3a-7f12-7000-8000-123456789abc", "status": "APPLIED" }, { "client_mutation_id": "01927e3a-7f13-7000-8000-123456789abd", "status": "APPLIED", "comment_id": "97301143552201984" } ], "server_time": "2026-09-28T08:00:01.015Z" }
The user is always the one in the token; the payload never says who is acting. The batch is processed in order, and each mutation gets its own result:
| Result | Meaning | What the phone does |
|---|---|---|
APPLIED | Done now | Remove from the queue |
DUPLICATE | Already done by an earlier attempt; same result returned | Remove from the queue |
RETRY (or a whole-batch 429, 5xx, timeout, network error) | Transient | Keep it; back off with jitter; honour Retry-After |
REJECTED with a code (POST_DELETED, FORBIDDEN, INVALID, VERSION_CONFLICT) | Permanent: retrying can't help | Move to quarantine; roll back what the screen showed; tell the user |
Whole-batch 401 | Token expired | Refresh the token and resend; never drop the queue |
R2.4 Design Evolution: Offline Writes, Delta Sync and Battery
Step 2.1: Liking a Post With No Signal
The problem: a user in a tunnel taps "like". Today the app calls the API, which fails. Product wants the heart to fill instantly, and the like to reach the server whenever the network allows, even if the app is killed in between. What would you do?
Primitive: Change Data Capture & the Outbox Pattern
Step 2.2: The App Was Killed While Sending, and the Retry Liked Twice
The problem: the sync worker sends a batch with a comment. The server stores it, but the response is lost when the train enters a tunnel, and then the OS kills the app. On the next launch the worker sends the batch again, and the post now shows the same comment twice. On another phone, a foreground sync and a background job both drained the queue at the same moment and sent everything twice. What would you do?
Primitive: Distributed Unique ID Generators
Step 2.3: Refreshing Downloads the Whole Feed
The problem: each of 1B daily syncs downloads a 60 KB page, 60 TB a day, though on average only two or three items changed. The team proposes GET /v1/changes?since=<timestamp of my last sync>.
What would you do? And what exactly is wrong with "since a timestamp"?
Primitive: Change Data Capture & the Outbox Pattern · Primitive: Message Queues vs Event Streams · Primitive: Distributed Locks & Leases · Drill: CDC outbox dual-write drift
Step 2.4: A Deleted Post Still Shows on Phones
The problem: an author deletes a post. Phones that sync later still show it: a delta that only carries new and changed posts has nothing to say about a post that no longer exists. A phone that was off for three weeks is even worse. What would you do?
Step 2.5: The Server Rejected an Action
The problem: a user comments offline on a post that was deleted while they were in the tunnel. The server answers 404. The worker retries forever with backoff, and the twenty likes queued behind that comment never leave the phone. The comment still shows as posted.
What would you do?
Primitive: Two-Phase Commit & Saga Orchestration
Step 2.6: Background Sync Drains Batteries
The problem: to keep the feed fresh, a teammate added a background timer that syncs every minute and downloads full-size images for everything new. Android marks the app for excessive wake-ups, iOS stops giving it background time, reviews say "battery hog", and some users hit their data cap. What would you do?
Loop: Design a Notification System
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | Liking offline | Optimistic update: outbox row in one local transaction; screen = server state + pending overlay; outbox media folder; pre-signed URL at send time | Something to undo when refused |
| 2.2 | Killed mid-send; double likes | UUIDv7 key born with the action; remove only after ACK; durable server guards (the like itself, a dedup item with no expiry); side effects from the change stream; one claimed drainer per phone | Dedup storage; a claim protocol |
| 2.3 | Full-page refreshes | Per-user FeedSyncLog with dense seq; cursor advanced in the same local transaction; one fenced writer per user fed from the fanout stream; META.last_seq never expires; head call reads last_seq first | ~2.3B log entries a day |
| 2.4 | Deleted posts linger | DELETE tombstones; versions on every entry; 14-day local tombstones; gap detection by dense numbers; 410 CURSOR_EXPIRED → head rebuild keeping the outbox; visible-post status backstop | 14 days of log; rebuilds after long absences |
| 2.5 | Rejected actions block the queue | Transient / auth / permanent classes; quarantine + rollback + notice; dependencies; full jitter; queue cap | Rollback UX |
| 2.6 | Background sync drains batteries | WorkManager / BGTaskScheduler with constraints; coalesced wake-ups; prefetch by network and battery; capped silent pushes as hints | Background freshness is the OS's call |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
Two loops meet here. Writes go phone → outbox → mutation API → social table, and every consequence (fanout, the followers' logs, counters, notifications) flows from that table's change stream. Reads go log → sync API → phone, where they land in the database the screen observes. Nothing on the server writes the log except the log writers, and nothing on the phone writes the screen except the database.
Local schema (SQLite)
sqlCREATE TABLE post ( -- server-confirmed state only post_id TEXT PRIMARY KEY, sort_key INTEGER NOT NULL, version INTEGER NOT NULL, -- apply an entry only if its version is higher author_id TEXT NOT NULL, author_name TEXT NOT NULL, text TEXT, image_id TEXT, like_count INTEGER NOT NULL DEFAULT 0, comment_count INTEGER NOT NULL DEFAULT 0, viewer_liked INTEGER NOT NULL DEFAULT 0 ); CREATE INDEX post_order ON post (sort_key DESC, post_id); CREATE TABLE pending_mutation ( -- the outbox local_seq INTEGER PRIMARY KEY AUTOINCREMENT, -- send order client_mutation_id TEXT NOT NULL UNIQUE, -- UUIDv7, the idempotency key type TEXT NOT NULL, -- LIKE, UNLIKE, COMMENT, CREATE_POST, EDIT_POST target_id TEXT NOT NULL, -- post ID, or a local ID for an offline post depends_on TEXT, -- client_mutation_id this one needs first payload_json TEXT NOT NULL, payload_version INTEGER NOT NULL, -- lets a newer app read an older row (Round 3) state TEXT NOT NULL CHECK (state IN ('QUEUED', 'IN_FLIGHT')), claim TEXT, -- run_id of the drainer that claimed it attempts INTEGER NOT NULL DEFAULT 0, next_attempt_at_ms INTEGER NOT NULL DEFAULT 0 -- local scheduling only ); CREATE TABLE quarantined_mutation ( client_mutation_id TEXT PRIMARY KEY, type TEXT NOT NULL, payload_json TEXT NOT NULL, error_code TEXT NOT NULL, user_notified INTEGER NOT NULL DEFAULT 0 ); CREATE TABLE tombstone (post_id TEXT PRIMARY KEY, version INTEGER NOT NULL, received_at_ms INTEGER NOT NULL); CREATE TABLE sync_state (key TEXT PRIMARY KEY, value TEXT NOT NULL); -- 'cursor', 'server_clock_offset_ms'
Server tables (DynamoDB)
| Table | Item | PK | SK | Attributes |
|---|---|---|---|---|
FeedSyncLog | Entry | USER#<user_id>#E#<epoch> | seq (Number, ≥ 1) | op, post_id, version, sort_key, src (the Kinesis sequence number of the record that produced it), expires_at (TTL, 14 days) |
FeedSyncLog | High-water mark (one per epoch) | USER#<user_id>#E#<epoch> | 0 | last_seq; no TTL |
FeedSyncLog | Current epoch | USER#<user_id> | 0 | epoch (which log the user writes to now); no TTL |
| Social table | Like | POST#<post_id> | LIKE#<user_id> | state (LIKED or UNLIKED), last_cmid_by_device (a map: device → last applied client_mutation_id), updated_at |
| Social table | Comment | POST#<post_id> | C#<comment_id> | author_id, text, status |
| Social table | Dedup guard | USER#<user_id> | CMID#<client_mutation_id> | result_id (the comment or post ID); no TTL |
seq is a Number, so it sorts numerically (as a string, "10" would sort before "9"). Counters (like_count, comment_count) are summed from the social table's change stream into ElastiCache and written back to the post item every few seconds, instead of being incremented on the post item by every like, which would turn a viral post into one hot item. DynamoDB Streams orders records per item only: each item's records arrive in order on one shard, but the likes of one post are different items, and when DynamoDB splits a hot partition, a viral post's likes can spread across shards. So the aggregator stores, on the post item, the sequence number of the last record it applied per (post, stream shard), and applies a batch from a shard only if the batch's first record is newer than that shard's stored number, so a replayed batch after a crash isn't counted twice.
Trace 1: a like made offline, synced later.
Synthesizing vector architecture diagram...
If the response is lost, the row stays; the retry's conditional put fails and the answer is DUPLICATE, which removes the row the same way.
Trace 2: a rejected comment, rolled back.
Synthesizing vector architecture diagram...
Trace 3: a delta sync with a tombstone.
Synthesizing vector architecture diagram...
One local transaction applies the entries and moves the cursor, so a crash leaves either the old state and old cursor, or the new state and new cursor.
R2.6 Numbers and Cost
Traffic (the interviewer's figures, derived; peak is 3× the average)
| Quantity | Math | Value |
|---|---|---|
| Syncs per day | 50M × 20 | 1B |
| Syncs/s, average | 1,000,000,000 ÷ 86,400 | ≈ 11,574 |
| Syncs/s, peak | 11,574 × 3 | ≈ 34,722, call it 35K |
| Actions per day | 50M × 5 | 250M |
| Actions/s, average | 250,000,000 ÷ 86,400 | ≈ 2,894 |
| Actions/s, peak | 2,894 × 3 | ≈ 8,681, call it 8.7K |
| Mutation requests (assume 2.5 actions per batch) | 8,681 ÷ 2.5 at peak | ≈ 3.5K/s |
Bytes on the wire
| Flow | Math | Value |
|---|---|---|
| Delta egress at peak | 35,000 × 4 KB = 140 MB/s | ≈ 1.12 Gbps |
| The same with full pages | 35,000 × 60 KB = 2.1 GB/s | ≈ 16.8 Gbps |
| Delta egress per month | 1B × 4 KB × 30.4 | ≈ 121.6 TB |
| Full pages per month | 1B × 60 KB × 30.4 | ≈ 1.82 PB |
| Action ingress at peak | 8,700 × 1 KB = 8.7 MB/s | ≈ 69.6 Mbps |
| All API egress per month | 121.6 TB deltas + ~1.5 TB mutation responses + ~1.8 TB head rebuilds (1M a day × 60 KB) | ≈ 125 TB |
The 4 KB is our planning figure for an average sync, the NFR's ceiling: about 2.3 new log entries (below), each hydrated to about 1 KB, plus counters for the posts on screen and headers, after gzip. Many syncs return nothing and are far smaller.
Synthesizing vector architecture diagram...
Syncing only what changed cuts peak egress by 15 times, and it's the user's data plan that saves the most.
The change log (assumptions: 0.2 posts per user per day; 200 followers per post on average; 2% of posts later edited or deleted; every action also appends one SELF entry to its own author's log, for their other devices)
| Quantity | Math | Value |
|---|---|---|
| Posts per day | 50M × 0.2 | 10M |
| Feed entries | 10M × 200 | 2.0B |
| Edit and delete entries | 10M × 2% × 200 | 40M |
| Own-action entries | 250M | 250M |
| Log entries per day | ≈ 2.29B | |
| Entries/s, average and peak | 2.29B ÷ 86,400; × 3 | ≈ 26.5K; ≈ 79.5K |
| New entries per sync, average | 2.29B ÷ 1B | ≈ 2.3 |
| Stored for 14 days | 2.29B × 14 × ~250 B (item, keys, overhead) | ≈ 32B items ≈ 8 TB |
On the phone
| Store | Math | Value |
|---|---|---|
| Posts | 500 × ~2 KB | ≈ 1 MB |
| Outbox | ≤ 1,000 rows × ~1 KB | ≤ 1 MB |
| Tombstones | deletes seen in 14 days, a few thousand × ~60 B | < 0.5 MB |
| WAL | at most about 4 MB between checkpoints at SQLite's default (less with Android's framework defaults); journal_size_limit trims it back after one | ≤ ~4 MB normally |
| Image cache | cap | 150 MB |
| Total | ≈ 160 MB, plus photos waiting in the outbox |
Cellular data per user per day: KB of sync, plus about 1.5 MB of images (assumption: ~20 new images viewed, mostly thumbnails on cellular and feed-width on Wi-Fi, plus ~20% prefetched and never viewed, plus avatars). About 1.6 MB a day, 49 MB a month.
Fleet sizing (per-task rates are assumptions to confirm by load test; every fleet must carry the peak with one AZ lost, and we round per AZ)
| Fleet | Per task | Peak need | With one AZ lost | Tasks |
|---|---|---|---|---|
| Sync API (2 vCPU, 4 GB) | 500 syncs/s | 35,000 ÷ 500 = 70 | 2 AZs carry 70 → 35 per AZ | 105 |
| Mutation API (2 vCPU, 4 GB) | 250 batches/s | 3,500 ÷ 250 = 14 | 7 per AZ | 21 |
| Log writers (1 vCPU, 2 GB) | 4,000 entries/s | 79,500 ÷ 4,000 ≈ 20 | 10 per AZ | 30 |
Quotas. The log's peak is about 79.5K entries/s × 2 writes (entry + META) ≈ 159K write units a second, far above DynamoDB's default of 40,000 per table (and 80,000 per account for provisioned tables), so we request the increases before launch. The Kinesis stream needs 80 shards at peak (1,000 records/s each); we run 100 to leave headroom for skew between shards. Background pushes, at a few per device per day (say 150M a day, about 5.2K/s at peak), sit under FCM's default 600,000 messages a minute (10K/s) per project, but close enough to alarm on.
Monthly cost (us-east-1 list prices; CloudFront at United States and Europe rates; rounded)
| Item | Math | Monthly |
|---|---|---|
| CloudFront image data | 50M × 1.5 MB × 30.4 ≈ 2,280 TB: $39.8K for the first 1,024 TB (tiers $0.085 → $0.030) + 1,256 TB × $0.025 | ≈ $71.2K |
| CloudFront image requests | 50M × ~30 a day × 30.4 ≈ 45.6B × $0.01 per 10,000 | ≈ $45.6K |
| API data out through the ALB | 125 TB: 10 × $90 + 40 × $85 + 75 × $70 per TB | ≈ $9.6K |
| ALB | new connections dominate: (11.6K syncs + 1.16K mutation batches)/s on average ÷ 25 per LCU ≈ 509 LCU × $0.008 × 730 | ≈ $3.0K |
| ECS on Fargate (Graviton) | 126 × $0.079/h + 30 × $0.0395/h, × 730 | ≈ $8.1K |
FeedSyncLog writes, provisioned | 53K WCU average ÷ 70% target use ≈ 75.7K WCU × $0.00065 × 730 | ≈ $35.9K |
FeedSyncLog reads, provisioned | ~1.6 strong reads per sync (the Query, plus META when it's empty, plus a rare second Query): 18.5K RCU ÷ 70% × $0.00013 × 730 | ≈ $2.5K |
| Sync hydration reads | the ~2.3 new posts per sync in one eventually consistent BatchGetItem, ~1 RRU per sync: 30.4B RRU × $0.125 per million | ≈ $3.8K |
| Social table writes, on demand | 250M × ~1.9 WRU (assume 70% are likes at 1 WRU; a comment or post with its dedup item in a transaction is 4) × 30.4 ≈ 14.4B × $0.625 per million | ≈ $9.0K |
| DynamoDB storage and counters | log 8 TB + dedup items ~5.5 TB a year in (75M a day × ~200 B with DynamoDB's per-item overhead), × $0.25/GB; counter write-backs ~$1K | ≈ $4.4K |
Kinesis sync-log-events | 100 shards × $0.015 × 730 + 69.6B PUT units × $0.014 per million | ≈ $2.1K |
| ElastiCache for Valkey, counters | 6 × cache.r7g.xlarge × $0.3496 × 730 | ≈ $1.5K |
| Cross-AZ traffic to the cache | ~30 TB a month of counter replies, two thirds cross-AZ × $0.02/GB | ≈ $0.4K |
| CloudWatch, WAF, variants, push senders, misc. | ≈ $10K | |
| Total | ≈ $207K/month |
Three notes on that table:
- Why provisioned capacity for the log. At on-demand prices its writes would be 69.6\text{B} \times 2 \times \0.625 \text{ per million} \approx $87\text{K}$. The load is steady and predictable (it follows posting), which is what provisioned capacity with auto scaling is for, and it cuts the line by about 60%.
- Why the API goes straight to the ALB, not through CloudFront. CloudFront's data price is a little lower, but every API call would also pay its request fee: 1B syncs + 100M batches a day ≈ 33.4B requests a month ≈ $33K, more than all the API's bytes cost.
- What the deltas saved. Full pages would be 1.82 PB a month of API egress, about $95K instead of $9.6K, and about 1.1 MB a day more on every user's data plan (20 × 60 KB instead of 20 × 4 KB).
The feed service itself (timelines, fanout, ranking) and media storage are priced in the server loop; this is what the offline-first sync adds. About $0.004 per daily user per month, and images are more than half of it.
R2.7 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| Optimistic vs pessimistic UI | Optimistic, with an overlay and quarantine | The screen sometimes shows something that's later taken back; we design the notice. Pessimistic is right where a false "done" is costly (payments), not for a like. |
| Cursor vs timestamp sync | A dense per-user seq assigned by one fenced writer | A second write path and ~2.3B entries a day. Timestamps (device or server) skip late commits. |
| Tombstone window length | 14 days | Longer: more log storage (~0.57 TB per extra day). Shorter: more 410s and 60 KB head rebuilds for people back from a holiday. We tune it from the distribution of time between syncs. |
| Managed sync (AppSync) vs a custom sync API | Custom | AppSync still documents Delta Sync on versioned DynamoDB data sources, but its cursor is a time (lastSync) checked against a delta table's TTL, it's built for syncing whole models, not per-user feed logs, and the client that used it, Amplify DataStore, isn't supported in Amplify Gen 2 (Gen 1 is in maintenance mode, critical fixes only, and reaches end of life on May 1, 2027). We'd build the phone side ourselves either way. |
| Live updates in the foreground: timer vs SSE vs WebSocket | A sync every 5 minutes while the feed is on screen, plus on open and pull to refresh | Up to a few minutes' lag in the foreground. SSE (server-sent events, one-way from server to client over HTTP) would be enough for "new posts available" nudges; a WebSocket (a full-duplex connection where both sides send at will) is what chat needs, because the client sends and receives continuously. A feed has minute-level freshness, and a held connection per active phone costs server memory and keeps the radio awake. |
R2.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| WAL file ballooning (checkpoint starvation) | The -wal file grows from ~4 MB to hundreds of MB on some phones; queries slow down | A checkpoint can't copy past, or reset the log under, a reader still using an old snapshot. The usual cause is a read transaction left open: a query cursor never closed, or a long export. Fixes: close every query promptly (scoped reads); set journal_size_limit so the file is truncated after a checkpoint; run PRAGMA wal_checkpoint(TRUNCATE) in the idle, charging maintenance task; report WAL size in telemetry. |
| Clock skew | Wrong "5 min ago" labels | Nothing sync-related uses the phone's clock. Display times use the offset from server_time. |
| A poison mutation | One action refused forever | Classified as permanent, quarantined, rolled back, noticed; the queue keeps moving (step 2.5). |
| Process death mid-flight | Unknown whether the server applied a batch | Rows stay until acknowledged; the resend gets DUPLICATE; rows left IN_FLIGHT by a dead run_id are reset at startup (step 2.2). |
| Durability on power loss | The last action before a phone died is missing | With synchronous = NORMAL, a WAL database survives an app crash, but a power loss can roll back the last commits. We run the writer connection with synchronous = FULL: one extra flush per write transaction, which is cheap at a few user actions a minute. |
| Reconnect storm: one phone | A train in and out of tunnels: 30 network flaps a minute | Wait for the network to stay up for 3 s before sending; full-jitter backoff; one coalesced sync per wake-up. |
| Reconnect storm: the fleet | After a 20-minute backend outage, millions of phones sync at once; each new connection costs a TLS handshake and an auth check | Phones retry with full jitter and honour Retry-After. The sync API sheds early: an overloaded task answers 503 with a random Retry-After before it verifies the token or reads DynamoDB, so shedding costs almost nothing. TLS session resumption cuts repeat-handshake cost. Each task's admission limit is the fleet's safe rate divided by the current task count, recomputed on every scale event, so the sum never exceeds what the log table is provisioned for. See R2.7 for why we don't hold WebSockets open. |
| Tombstone resurrection after a long absence | A deleted post reappears | The version rule plus 14-day local tombstones reject late UPDATEs; anyone away longer gets 410 and a head rebuild; the visible-post status check removes stragglers (step 2.4). |
| A stale log writer (lost its shard, still running) | Two workers write one user's log | The "only if SEQ#n doesn't exist" condition fails for the stale one's write, and its checkpoint is refused; it stops. |
| Log table throttling (auto scaling lags a spike) | Log writers get throttled | Writers back off; records wait in the Kinesis stream (24-hour default retention, raised to 7 days), so new posts reach phones late, never lost. Syncs keep serving what's there. |
Drill: WebSocket reconnect thundering herd (the fleet-level storm above, and the SSE vs WebSocket row in R2.7)
R2.9 Production Gotchas
| Gotcha | Symptom | Cause | Fix |
|---|---|---|---|
| UI bound to network callbacks | Blank lists after rotation or process death; a newer local change overwritten by an older response | The screen shows responses instead of the database | Network writes to SQLite only; the screen observes SQLite |
| Optimistic state written into the server row | The heart flickers off when a sync arrives before the like is sent | Sync overwrites the column the optimistic write changed | Keep server state and pending actions apart; the view combines them |
| Main-thread database or JSON work | Jank; ANRs on old phones | "It's only a small query" | Background dispatchers; StrictMode in debug; hang reports in the field |
| Non-idempotent replays | Double comments and like counts | No key, or a key forgotten after 24 hours | Key born with the action; durable server guards |
| Aggressive background wake-ups | Battery complaints; OS restrictions | Timers and polling | OS schedulers, coalescing, capped hints |
| Pre-signed URL requested at compose time | Offline photo posts fail with 403 after the tunnel | The URL expired while queued | Request it when sending |
| Evicting a post that a queued action references | A queued comment whose post vanished from the screen | Eviction didn't look at the outbox | Eviction runs in a write transaction and skips posts referenced by pending_mutation, quarantine or an open screen |
R2.10 Pillar Check
| Pillar | What Round 2 adds |
|---|---|
| Reliability | Idempotent actions with durable guards; nothing removed from the outbox before an acknowledgement; error classes and quarantine; full-jitter retries; fleet shedding before expensive work; AZ-loss sizing for every fleet; quotas raised ahead REL 4 · REL 5 · REL 1 |
| Performance Efficiency | ≤ ~4 KB syncs; the screen never waits on the network; counters for visible posts only; log reads are one strongly consistent Query PERF 3 |
| Security | The acting user always comes from the token; keys namespaced by user; permission and ownership checks on every referenced post and media_id; signed cursors bound to a log epoch SEC 3 |
| Cost Optimization | ≈ $207K/month, derived; deltas instead of full pages save ~$85K a month; provisioned capacity for the steady log saves ~$51K; API not routed through CloudFront to avoid ~$33K of request fees COST 7 · COST 8 |
| Operational Excellence | Light this round: server alarms on log-writer lag (Kinesis iterator age), log throttling, 410 rate and quarantine rate; client metrics arrive in Round 3 OPS 8 |
| Sustainability | OS-scheduled, coalesced background work; prefetch only on Wi-Fi with battery to spare; fewer bytes per sync on every phone SUS 2 · SUS 3 |
Alarms and first actions
| Signal | Alarm | First action |
|---|---|---|
| Sync API P99 latency | > 300 ms for 5 min | Per-stage timings: log Query, hydration, counters |
| Log writer lag | Kinesis iterator age > 60 s for 5 min | Check log table throttling; scale writers |
410 CURSOR_EXPIRED rate | > 3× the weekly baseline | A log retention or META bug, or a writer that restarted numbering |
| Quarantine rate | > 0.1% of distinct users in 15 min | A server validation change, or a client bug in one version |
Mutation DUPLICATE rate | Sudden jump | A client resending acknowledged rows (claim bug), or lost responses at the edge |
R2.11 Round 2 Rubric and Follow-Ups
What a senior (L6) answer adds over L5
- Writes offline actions to a local outbox in the same transaction as the optimistic change, and keeps server state separate from pending changes.
- Makes retries safe with keys born on the phone and guards on the server that don't expire, and checks who the caller is and what they may touch.
- Rejects timestamp cursors for the right reason (late commits, not just skewed clocks), and designs how the sequence is assigned: one fenced writer, dense numbers, a high-water mark that survives expiry.
- Carries deletes as tombstones with versions, bounds the window, and detects a phone that fell out of it without a clock.
- Separates transient from permanent errors, and designs the rollback the user sees.
- Lives inside the OS's background rules, and checks the battery budget with rough numbers.
- Derives traffic, bytes and cost, and finds that the log and image delivery dominate.
Follow-up questions
-
"A user writes 10 comments offline, then reconnects on a slow network. What do they see?" Answer: all 10 show at once with a "Sending" badge, in the order written. The engine sends them in batches in
local_seqorder; each result removes its row and the server's version of the comment arrives through the log, replacing the local one byclient_mutation_idin one transaction, so nothing jumps. A slow network only means the badges clear slowly. -
"The same user has a phone and a tablet. They like a post on the phone and unlike it on the tablet while both are offline." Answer: each device's queue is sent when it reconnects, and the server applies them in the order they arrive: last write wins for that user's like, which is acceptable for a like (only the user writes it). Each applied action appends a
SELFentry to the user's own log, so both devices converge on the final state at their next sync. For something that matters more, like an edit,EDIT_POSTcarries the version it was based on, and a stale one getsVERSION_CONFLICTand goes to quarantine with the user's text. -
"Why not use CRDTs so phones never conflict?" Answer: a CRDT (a data type whose copies merge the same way whatever order updates arrive in) earns its complexity when many people edit the same object at once, like a shared document. Here, likes are per-user state, comments are append-only, and posts have one author. A server order with version checks is simpler and enough. The file sync loop discusses true multi-writer conflicts.
Interview gotchas from this round
| Gotcha | Why it's wrong |
|---|---|
| "Sync since the last timestamp" | Device clocks lie, and even server timestamps skip writes that commit late. |
| "Take the next number from a counter, then write the entry" | Two writers can commit out of order; a reader skips the lower number forever. |
| "Keep idempotency keys for 24 hours" | A phone can hold an unacknowledged action for weeks. |
| "Retry everything with backoff" | A permanent error blocks the queue forever. |
| "Rely on DynamoDB TTL to delete old log entries on time" | TTL deletes eventually, typically within days after expiry; design so lingering entries are harmless. |
| "Silent pushes keep the feed live" | Apple says two or three an hour at most, delivery isn't guaranteed, and force-quit apps don't get them. |
Round 3 · Architect · "Many App Versions, Platforms and a Risky Rollout"
~45 min · Principal (L7) · 1 region, 3 AZs (global placement is the server loop's Round 3) · 200M installs, 80M DAU · ~55.6K syncs/s peak · 30 app versions · crash-free users ≥ 99.5% per version · a bad feature off in minutes
R3.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 3. If you're starting here, it's everything you need from Rounds 1 and 2.
Round 2 in 60 seconds. "We run an offline-first social feed for 50 million daily users: a billion syncs a day, about 35,000 a second at peak, and 250 million actions, many made offline. On the phone, SQLite is the only source for the screen. A tap writes an outbox row in one transaction, with a UUIDv7 key born with the action; the screen shows server state plus pending actions, so a sync can't erase a pending like. One claimed drainer sends batches; rows leave the outbox only when the server says applied or duplicate. The server's guards never expire: the like item itself, or a dedup item written in the same transaction as a comment. Permanent errors go to a quarantine and the screen rolls back with a notice. Reads are delta syncs: each user has a change log in DynamoDB with dense sequence numbers, written by one fenced writer per user from the fanout stream, with a high-water mark that never expires. The phone advances its cursor in the same transaction that applies entries. Deletes are versioned tombstones kept 14 days; a phone away longer gets 410 and rebuilds from the head, keeping its outbox. Background work is scheduled by WorkManager and BGTaskScheduler; prefetch depends on network and battery. About 4 KB a sync, 1.12 Gbps of delta egress at peak, about $207K a month, over half of it images. Open costs: one protocol version, a schema we can't change safely, releases we can't pull back, plaintext data on lost phones, and no idea what happens on real devices."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: actions go up through the outbox, changes come down through a per-user log, and the phone's database is where they meet.
Steps so far
| Step | Problem | Component |
|---|---|---|
| 1.1–1.3 | Spinners, jank, lock waits | SQLite as the screen's only source; nothing slow on the main thread; WAL |
| 1.4–1.5 | Images and storage | CDN variants, immutable URLs, disk LRU; budgets |
| 2.1 | Acting offline | Optimistic overlay + outbox in one transaction |
| 2.2 | Duplicates | Keys born on the phone; durable server guards; one claimed drainer |
| 2.3 | Full refreshes | Per-user log, dense seq, fenced single writer, cursor in the same transaction |
| 2.4 | Deletes | Versioned tombstones; 14-day window; 410 → head rebuild |
| 2.5 | Rejections | Error classes, quarantine, rollback |
| 2.6 | Battery | OS schedulers, coalescing, adaptive prefetch |
Open costs: one protocol version; destructive local upgrades; no way to stop a bad release; plaintext on lost phones; no telemetry.
R3.1 The Scope Raise
Interviewer: "We're on iOS, Android and the web now, with 200 million installs. We ship every week, but a third of our users haven't updated in months, and we count 30 app versions still in use. Next quarter's release changes the local database a lot, and people will have offline actions queued when they upgrade. Last month a bad release crashed on some Android phones for four days while we waited for the store. Families share tablets, phones get lost, and we learn about problems from one-star reviews."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How many users, and on what? | 200M installs; 80M daily users: 70M in the apps, 10M on the web. | 1.6B syncs a day (R3.6); a third platform with different storage. |
| How old is the oldest version still in use? | About two years. 30 versions each have at least 0.1% of users. | The server must speak several protocol versions, and we need a policy for retiring them (step 3.1). |
| Can the database change destroy anything? | Not queued actions or drafts. Cached posts can be re-downloaded. | Migrate what exists only on the phone; rebuild what the server has (step 3.2). |
| How fast must a bad release be stopped? | In minutes, on phones that already have it. | Store rollouts can't do that; server-controlled flags and kill switches can (step 3.3). |
| What must a lost or shared phone not reveal? | The feed, drafts and messages, and nobody should stay signed in on a lost phone. | Encrypted local data, per-account files, token revocation and wipe on sign-out (step 3.4). |
| What do we want to see from devices? | Crashes, sync health, queue depth, time to content, per version, without collecting personal content. | A sampled, privacy-aware telemetry pipeline (step 3.5). |
| Must web and apps behave the same? | Yes: same conflicts, same rollbacks. | One sync specification and conformance tests; platform storage underneath (step 3.6). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Users | 50M DAU | 200M installs; 80M DAU (70M apps, 10M web) |
| Syncs | 1B/day; 35K/s peak | 1.6B/day ≈ 18,519/s; ≈ 55.6K/s at peak |
| Actions | 250M/day | 400M/day ≈ 4,630/s; ≈ 13.9K/s at peak |
| Platforms | iOS, Android | + web |
| Versions in the field | Assumed one | 30 app versions; 3 sync protocol versions |
| Local schema | Fixed | Changes with releases; queued actions and drafts survive |
| Stopping a bad release | An app-store update (days) | A kill switch (minutes) |
| Device data | OS defaults | Encrypted, per account, wiped on sign-out and revocation |
| Visibility | Server metrics | + client telemetry per version |
R3.2 What Breaks in the Round 2 Design
| Round 2 choice | What breaks at the new scope |
|---|---|
| One sync protocol | A new entry type (say, a poll) sent to a two-year-old app crashes its parser, or silently drops data. |
| Schema changes by "drop and recreate" | The outbox and drafts live in the same database. Dropping it loses offline actions on millions of phones. |
| Fixes ship through the stores | Google Play's halt stops new installs of a staged release; phones that have it keep it. Apple's phased release covers automatic updates only; anyone who updates by hand gets the new version at once. A fix takes a build, a review and days of update adoption. |
| OS-default storage, one database per app | On a shared tablet, the second account sees the first account's feed and drafts. A lost phone keeps a valid session. |
| No client telemetry | A crash on one Android model, a migration failing on old phones, or a queue that never drains is invisible until reviews pile up. |
| Two native apps | The web app reimplements sync from memory, and its conflict rules drift. |
R3.3 New Requirements and API Additions
1. Every request says who is talking.
httpPOST /v1/sync HTTP/1.1 Host: api.example-social.com Authorization: Bearer <token> X-App-Version: 5.14.2 X-Platform: android X-Sync-Protocol: 2 Content-Type: application/json { "cursor": "c1.ep_Hx7.450912.Qm8aZ1", "capabilities": ["op.edit", "media.avif", "post.poll"] }
The response carries a client policy and the current flags version:
json{ "entries": [], "cursor": "c1.ep_Hx7.450912.Qm8aZ1", "client_policy": { "min_supported": "5.2.0", "recommended": "5.14.0", "flush_only_until": null }, "flags_version": 43 }
A version below min_supported gets 400 with "error": "CLIENT_VERSION_UNSUPPORTED" on reads, and the app shows its upgrade screen (step 3.1).
2. Flags and kill switches, served as static files from CloudFront. A small latest.json names the current version; each version's document has its own path and never changes:
httpGET /flags/v43/android.json HTTP/1.1 Host: cfg.example-social.com
json{ "version": 43, "kill": { "poll_posts": true }, "flags": { "poll_posts": { "default": false, "rollout_percent": 20, "min_app": "5.14.0" }, "avif_images": { "default": true } }, "client_policy": { "min_supported": "5.2.0", "recommended": "5.14.0" }, "telemetry": { "session_sample_percent": 10 } }
3. Telemetry, in batches:
httpPOST /v1/telemetry:batch HTTP/1.1 Authorization: Bearer <token> Content-Type: application/json { "install_id": "i_4f9c2e", "app_version": "5.14.2", "platform": "android", "device_class": "mid-2023", "day": "2026-09-28", "health": { "sessions": 6, "crashes": 0, "syncs_ok": 19, "syncs_failed": 2, "queue_depth_max": 3, "queue_age_max_s": 540, "quarantined": 0, "ttc_ms_p50": 38, "ttc_ms_p95": 120, "wal_bytes_max": 4194304 } }
install_id is a random ID made at install time, not an advertising ID and not the user ID.
R3.4 Design Evolution: Versions, Migrations, Rollouts, Devices and Platforms
Step 3.1: 30 App Versions Talk to Our API
The problem: product wants polls in the feed. The new server sends op: "ADD" with "post_type": "POLL" and new fields. Version 5.3, still used by 4% of people, crashes on the unknown type; version 4.9 ignores it and advances its cursor, so those users will never see those posts even after they update.
What would you do?
Step 3.2: A New Release Changes the Local Database
The problem: release 5.15 restructures the local tables. The quick way is "if the schema version changed, delete the database and resync". But millions of phones have queued likes, comments and photo posts in the outbox, and some have drafts. What would you do?
Step 3.3: A Release Is Crashing on Some Phones
The problem: release 5.15's new poll screen crashes on one family of Android phones. The team halts the Play rollout, but the 20% who already have 5.15 keep crashing, and some crash at launch. A fixed build needs review and days to reach people. What would you do?
Step 3.4: A Lost Phone Exposes the Feed and Drafts
The problem: a user loses their phone on a train. Their session is still valid, the feed and drafts are on it, and their outbox might send a half-written post. On a family tablet, a child opens the app and sees a parent's feed. What would you do?
Primitive: OAuth2, OIDC & Distributed Token Authentication · Drill: JWT revocation token blacklist
Step 3.5: We Can't See What's Happening on Devices
The problem: reviews say "my likes disappear" and "the app is slow since the update". Server dashboards are green. We don't know which versions, devices or networks are affected. What would you do?
Step 3.6: Web and Mobile Behave Differently
The problem: the web team wrote its own sync. On the web, a rejected comment silently disappears instead of going to quarantine, two open tabs send the same queued like twice, and a user who opens the site once a week on Safari finds their offline drafts gone. What would you do?
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | 30 versions | Additive protocol; declared capabilities with fallbacks; encoders per protocol version; min_supported that still lets the outbox flush | Encoders and fallbacks for every change |
| 3.2 | Local schema changes | Rebuild cache tables, migrate local-only tables; user_version steps in transactions; payload_version decoders; fixture matrix; rescue the queue first | A growing test matrix |
| 3.3 | Crashing release | Features ship dark; staged binaries; flags with a separate kill layer; AppConfig for authoring, CloudFront for delivery; versioned documents; crash-loop safe mode | Flag debt |
| 3.4 | Lost or shared phone | Data Protection class that allows background sync; Keychain / Keystore tokens; no cloud backup of the database; per-account files; wipe on sign-out; 1-hour tokens + denylist | Key handling; complex sign-out |
| 3.5 | No visibility | Crash service + daily health summary + sampled traces → Firehose → S3 → Athena; distinct-install thresholds; privacy rules | A pipeline and a privacy review |
| 3.6 | Platforms drift | One spec, golden-trace conformance suite, shared core; Web Locks; persistent storage requests | Cross-team coordination |
R3.5 Global Architecture
Synthesizing vector architecture diagram...
Three clients share one sync core and one specification. The sync gateway turns one internal model into each protocol version. Two new paths serve every version: flags flow from AppConfig to static documents on CloudFront, and telemetry flows from phones to Firehose, S3 and Athena. The feed backend below the log (fanout, timelines, ranking) is unchanged from Round 2 and the server loop.
Trace 1: an old app version syncs a poll.
Synthesizing vector architecture diagram...
When this user updates, the new app refetches posts it stored as fallbacks and shows the poll.
Trace 2: a kill switch.
Synthesizing vector architecture diagram...
Trace 3: a local migration with queued actions.
Synthesizing vector architecture diagram...
Every path keeps the outbox if it can be read at all, and every path ends with caches rebuilt from the server.
R3.6 Numbers and Cost
Traffic (80M DAU; 20 syncs and 5 actions per user per day as in Round 2; peak 3× average)
| Quantity | Math | Value |
|---|---|---|
| Syncs per day | 80M × 20 | 1.6B |
| Syncs/s, average / peak | 1.6B ÷ 86,400; × 3 | ≈ 18,519; ≈ 55.6K |
| Actions per day | 80M × 5 | 400M |
| Actions/s, average / peak | 400M ÷ 86,400; × 3 | ≈ 4,630; ≈ 13.9K |
| Log entries per day | 2.29B × 80 ÷ 50 | ≈ 3.66B (≈ 42.4K/s, ≈ 127K/s at peak) |
| Delta egress per month | 1.6B × 4 KB × 30.4 | ≈ 194.6 TB (plus mutation replies and head rebuilds: ≈ 200 TB) |
Sync traffic across versions (planning split by protocol: v3 72%, v2 23%, v1 5%). Every version reads the same log; only the encoding differs. We assume encoding for v1 and v2 costs about 10% more CPU per sync (translation and fallbacks), which adds to the gateway fleet.
Write rates we must size (heartbeat-like traffic that grows with installs, not with engagement)
| Flow | Assumption | Rate |
|---|---|---|
| Health summaries | 1 per active install per day | 80M/day ≈ 926/s average |
| Sampled session traces | 80M × 5 sessions × 10% | 40M/day ≈ 463/s average |
| Telemetry uploads, peak | (80M + 40M) ÷ 86,400 × 3 | ≈ 4.2K/s |
| Telemetry bytes, peak | 80M × 2 KB + 40M × 20 KB = 960 GB/day ≈ 11.1 MB/s; × 3 | ≈ 33 MB/s |
| Flag document fetches | ~3 per DAU per day (a new version, or a cold start with a stale copy) | 240M/day ≈ 2.8K/s average |
| Token refreshes (1-hour access tokens) | ~3 per DAU per day, each rotating the refresh token in the session table, with a short grace window for the previous token | 240M/day ≈ 2.8K/s average, 8.3K/s at peak |
The telemetry peak of about 33 MB/s is over Firehose's default Direct PUT quota of 5 MiB/s per stream (with 2,000 requests/s and 500,000 records/s) in us-east-1. We request an increase before launch (or spread across several streams), and the ingest service batches up to 500 records per PutRecordBatch call.
Fleet sizing (same per-task rates as Round 2, AZ-loss sizing, rounded per AZ)
| Fleet | Peak need | Per AZ, with one AZ lost | Tasks |
|---|---|---|---|
| Sync gateway | 55.6K × 1.028 ÷ 500 ≈ 114.3 | 58 | 174 |
| Mutation API | 13.9K ÷ 2.5 ≈ 5.6K batches/s ÷ 250 ≈ 22.4 | 12 | 36 |
| Log writers | 127K ÷ 4,000 ≈ 31.8 | 16 | 48 |
| Telemetry ingest (1 vCPU) | 4.2K ÷ 1,000 ≈ 4.2 | 3 | 9 |
Why not let phones read AppConfig directly? At list price AppConfig charges about $0.0008 each time a client receives configuration. 240\text{M} \times \0.0008 = $192\text{K} a day, about \5.8M a month. The same documents from CloudFront cost about $7.6K.
Monthly cost (us-east-1 list prices; CloudFront at United States and Europe rates; rounded)
| Item | Math | Monthly |
|---|---|---|
| CloudFront image data | 80M × 1.5 MB × 30.4 ≈ 3,648 TB: $39.8K for the first 1,024 TB + 2,624 TB × $0.025 | ≈ $105.4K |
| CloudFront image requests | 80M × 30 × 30.4 ≈ 73B × $0.01 per 10,000 | ≈ $73.0K |
| Flags documents | 7.3B requests × $0.01 per 10,000 + ~11 TB at the marginal $0.025/GB | ≈ $7.6K |
| API data out through the ALB | 200 TB: 10 × $90 + 40 × $85 + 100 × $70 + 50 × $50 per TB | ≈ $13.8K |
| ALB | new connections: (18.5K syncs + 1.85K mutation batches + 1.39K telemetry uploads)/s on average ÷ 25 ≈ 870 LCU × $0.008 × 730 | ≈ $5.1K |
| ECS on Fargate (Graviton) | 210 × $0.079/h + 57 × $0.0395/h, × 730 | ≈ $13.8K |
FeedSyncLog, provisioned | writes 84.8K WCU ÷ 70% × $0.00065 × 730 ≈ $57.5K; reads 29.6K RCU ÷ 70% × $0.00013 × 730 ≈ $4.0K | ≈ $61.5K |
| Sync hydration reads | ~1 RRU per sync: 1.6B × 30.4 ≈ 48.6B RRU × $0.125 per million | ≈ $6.1K |
| Social table writes, on demand | 400M × 1.9 WRU × 30.4 ≈ 23.1B × $0.625 per million | ≈ $14.4K |
| Session table (token refreshes) | 240M × 30.4 ≈ 7.3B WRU × $0.625 per million | ≈ $4.6K |
| DynamoDB storage and counters | log 12.8 TB + dedup ~8.8 TB a year in (120M a day × ~200 B), × $0.25/GB, + counters ~$1.6K | ≈ $7.0K |
Kinesis sync-log-events | 160 shards (128 at peak, plus headroom for skew) × $0.015 × 730 + 111.3B PUT units × $0.014 per million | ≈ $3.3K |
| ElastiCache for Valkey, counters, + cross-AZ | 9 × cache.r7g.xlarge × $0.3496 × 730, + ~$0.6K | ≈ $2.9K |
| Telemetry | Firehose: 80M records billed at 5 KB + 800 GB/day of 20 KB records ≈ 36.5 TB × $0.029/GB ≈ $1.1K; S3 (90 days of Parquet) ≈ $0.4K; Parquet conversion ≈ 36.5 TB × $0.018/GB ≈ $0.66K; Athena ≈ $0.5K | ≈ $2.7K |
| CloudWatch, WAF, variants, push senders, misc. | ≈ $15K | |
| Total | ≈ $336K/month |
About $0.004 per daily user per month, the same as Round 2: the new work (telemetry, flags, extra protocol versions) is small next to delivery. Images are still more than half the bill. The regional price matters too: these are United States and Europe CloudFront rates, and viewers in Asia or South America cost more per GB. CloudFront's cheaper price classes limit which edge locations serve our content; they don't limit where users are, so a user in Asia is then served from a farther edge, slower, not refused.
Firehose bills each record in 5 KB increments, so the 2 KB health summaries are billed as 5 KB. At this size it doesn't matter ($0.35K a month for all of them); it would if we sent one tiny record per event.
R3.7 Trade-Offs
| Choice | We chose | What we give up |
|---|---|---|
| Supporting old versions vs forcing upgrades | Support two years of apps with fallbacks; min_supported as the last resort, still letting the outbox flush | Encoders, fallbacks and a test matrix; a slower path to removing old code |
| Client-side vs server-driven logic | Server-driven where it changes often (flags, kill layer, client policy, fallback text); client-side where it must work offline (outbox, rollback, rendering) | More server round trips for configuration; server-driven UI beyond this is harder to test and to keep offline |
| Telemetry detail vs privacy | Daily aggregates from everyone, 10% detailed sampling, no content, random install IDs | Rare problems in unsampled sessions show up later; we can raise sampling by flag during an incident |
| Flags from AppConfig directly vs static documents | AppConfig for authoring; CloudFront for delivery | A publisher to run; up to about 5.6 minutes to reach a phone in use, instead of a live connection |
| A shared core vs native clients | One spec and conformance suite always; a shared core where it pays | The shared core must fit three lifecycles; native teams lose some freedom |
Closing the loop. The opening question was: how do we make the app instant and fully usable offline, and still keep it in agreement with the server? The answer is now:
- Instant: the screen reads only the phone's database; the network never sits between a tap and the screen (Round 1).
- Fully usable offline: actions go into an outbox in the same transaction as the change the user sees, and the server's guards make every retry safe (Round 2).
- In agreement with the server: changes come down through a per-user log with numbers the server assigns, deletes included, and a phone that falls behind rebuilds instead of guessing (Round 2).
- And it stays that way over time: every version speaks a compatible protocol, local upgrades protect what only the phone has, and features can be turned off without an app-store release (Round 3).
R3.8 Failure Modes
| Failure | What you'd see | How the design responds |
|---|---|---|
| A migration bug corrupts local data | Migration failures for one source version in health summaries; users report missing drafts | The step rolls back; rescue copies the outbox to a fresh database and flushes it; caches rebuild from the head. We kill the feature that needed the new schema if it's the cause, and fix the migration in the next build. |
| The flags service is unreachable | Publisher errors; latest.json can't be refreshed | Phones keep their last known document; compiled defaults for new flags are off. Flags documents are served with stale-if-error, so CloudFront keeps serving the last copy if S3 errors. Only kill switches are delayed, so we page on it. |
| A telemetry flood (a bug makes one version upload every minute) | Ingest traffic jumps; Firehose throttles | Ingest drops, never blocks: telemetry is best-effort and shares nothing with sync. Lower the sample rate by flag; the daily summary has a per-install cap enforced at ingest. |
| An API change breaks an old version | Sync errors for one app version only | Contract tests replay recorded requests from every supported version against the new server in CI; server deploys canary with error rates split by X-App-Version, and roll back automatically on a jump. |
| A kill switch crashes old versions | Crashes right after a flags change | Flags documents are additive and old apps ignore unknown keys (tested); a flags change is rolled out like code: to 1% of installs first, except for kills. |
| A crash at launch before flags load | A crash loop in one version | The crash-loop guard starts in safe mode and fetches latest.json before enabling features (step 3.3). |
| A release fills phones' storage (a cache limit removed by mistake) | Storage and WAL size in health summaries climb for one version | Kill the feature, or lower the cache budget by flag (budgets are flags too). |
R3.9 Runbook and Incident Response
Golden signals, per app version and platform OPS 8 · OPS 4
| Signal | Alarm | Severity | First action |
|---|---|---|---|
| Crash-free users (distinct installs) | New version < 99.5%, or 0.2 points below the previous version on the same rollout day, with ≥ 10,000 installs | P1 | Halt the store rollout; find the crashing feature; kill it |
| Sync success rate | < 99% of syncs per version for 15 min | P2 | Split by error code: 410s, 5xxs, parse failures on the phone |
| Outbox depth / age (p99 per version) | Age p99 > 12 h | P2 | A drainer bug, or a server rejection loop; check quarantine rate |
| Quarantine rate | > 0.1% of distinct users in 15 min | P2 | A server validation change or a client payload bug |
| Time to cached content | p95 > 250 ms for a version | P3 | A slow migration, WAL growth, or main-thread work in that build |
| Rollback rate (optimistic changes undone) | 2× baseline | P3 | Server rejections spiking for one action type |
| Battery and data | Android vitals excessive wake locks; MetricKit energy by version; background syncs per install per day | P3 | A background schedule or prefetch rule changed |
Log writer lag, log throttling, 410 rate | as in R2.10 | P2 | as in R2.10 |
Staged-rollout halt procedure OPS 6 · REL 8
- Stop the spread. Pause the phased release in App Store Connect; halt the staged rollout in the Play Console. (Neither removes the version from phones that have it.)
- Find the feature. Crash reports by version and device model; the top crash's stack names the screen or flag.
- Kill it. Set the flag in the kill layer in AppConfig and deploy (CLI 2). Watch
flags_versionreach the sync fleet (CLI 3), then crash-free users recover over the next minutes as phones sync. - If there's no flag for it, raise
min_supportedonly as a last resort, and only withflush_only_untilset, so queued actions still drain. - Fix forward. Ship a new build with a higher version number; resume the rollout from 1%.
- Record it: who, what, when; add a flag for anything that shipped without one.
Go deeper: CLI playbook
Plain commands an on-call engineer runs, one at a time. Replace the IDs with real ones.
text# 1. Crash-free users by app version, yesterday (health summaries in Athena) aws athena start-query-execution --work-group telemetry --query-string "SELECT app_version, 1 - CAST(approx_distinct(IF(crashes > 0, install_id)) AS double) / approx_distinct(install_id) AS crash_free_users FROM health_daily WHERE day = '2026-09-27' GROUP BY app_version ORDER BY crash_free_users" # 2. Kill switch: deploy the new kill-layer version at once aws appconfig start-deployment --application-id a1b2c3d --environment-id e4f5g6h --configuration-profile-id k7l8m9n --configuration-version 44 --deployment-strategy-id AppConfig.AllAtOnce # 3. Which flags version is published? aws s3 cp s3://example-flags/flags/latest.json - # 4. Log writer lag aws cloudwatch get-metric-statistics --namespace AWS/Kinesis --metric-name GetRecords.IteratorAgeMilliseconds --dimensions Name=StreamName,Value=sync-log-events --statistics Maximum --period 60 --start-time 2026-09-28T09:00:00Z --end-time 2026-09-28T10:00:00Z # 5. Log table write throttling aws cloudwatch get-metric-statistics --namespace AWS/DynamoDB --metric-name WriteThrottleEvents --dimensions Name=TableName,Value=FeedSyncLog --statistics Sum --period 60 --start-time 2026-09-28T09:00:00Z --end-time 2026-09-28T10:00:00Z # 6. Scale the sync gateway aws ecs update-service --cluster sync --service sync-gateway --desired-count 220
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | A versioned, additive protocol with fallbacks; migrations that protect local-only data and rescue the queue; kill switches with safe defaults and a crash-loop guard; changes rolled out in stages REL 3 · REL 8 · REL 1 |
| Performance Efficiency | One log read serves every protocol version; flags as static documents at the edge; time to cached content tracked per version PERF 4 · PERF 5 |
| Security | Data Protection class chosen for background sync; tokens in Keychain and Keystore; no cloud backup of the database; per-account files; wipe on sign-out; 1-hour tokens with a revocation denylist; telemetry with no content SEC 8 · SEC 2 · SEC 7 |
| Cost Optimization | ≈ $336K/month, derived; flags from CloudFront instead of about $5.8M of direct AppConfig reads; provisioned log capacity; telemetry aggregated per day so record rounding doesn't matter COST 5 · COST 8 |
| Operational Excellence | Per-version golden signals counted by distinct installs; a staged-rollout halt procedure; kill switches in minutes; contract tests for every supported version OPS 4 · OPS 6 · OPS 10 |
| Sustainability | Battery and data use measured per version (Android vitals, MetricKit, background syncs per install); budgets and prefetch rules adjustable by flag; telemetry sampled and sent with existing syncs SUS 2 · SUS 3 · SUS 4 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Treats time as a dimension of the system: many versions at once, and designs the protocol, fallbacks and retirement policy for it.
- Splits local data into "copy of the server" (rebuild) and "only on this phone" (migrate, rescue first), and tests every upgrade path.
- Knows what app-store rollouts can't do, and puts a kill layer, safe defaults and a crash-loop guard in place before they're needed.
- Chooses encryption classes by what background work must do, and designs revocation and wipe for lost and shared devices, including backups and reinstalls.
- Builds telemetry that counts people, not events, and respects privacy by design.
- Keeps three platforms honest with one spec and a conformance suite.
- Prices the new paths and catches the traps (AppConfig per-receive pricing, Firehose's default quota and record rounding).
Follow-up questions
-
"The backend becomes multi-region (see the server loop's Round 3). What happens to the per-user log and cursors in a failover?" Answer: the log stays single-writer per user: only the user's home region writes it. If the home region is lost and the partner takes over that user's writes, it must not continue the old numbers, because the last entries may not have replicated. So the partner starts a new log epoch for the moved users; phones presenting a cursor from the old epoch get
410and rebuild from the head, keeping their outboxes. The epoch is part of the log's key (USER#<id>#E#<epoch>, with oneMETAper epoch), so late replication of old-epoch entries from the returning region lands in the old log and can't touch the new one. Actions whose guards replicated before the failure stay safe. For the last seconds of lag, a resent comment can be applied twice, so a reconciliation job flags comments sharing aclient_mutation_idonce the region returns. Failover is detected by health checks evaluated outside the failed region (Route 53's health checkers, not calls into the failed region's APIs), and the DNS move is gated: users are routed to the partner only after it has been promoted to write their logs, and the partner needs capacity for the rebuild storm (millions of 60 KB head calls). Where users' data may live is a residency decision made before any of this: a user may only fail over to a region their data is allowed to be in. -
"Why not force everyone onto the latest version every month?" Answer: the store can't make people update; a forced-upgrade screen turns into uninstalls for users short of storage or on old OS versions that the new build doesn't support. We force only below
min_supported, with data (share under 0.5% for 30 days) or a security reason, and even then let the outbox flush. -
"A user says their likes keep disappearing. How do you debug it?" Answer: from their
install_id(they can share it from a debug screen), check the health summary: quarantine count and rollback rate. If rejections are the cause, the server logs byclient_mutation_idshow the error code. If there are no rejections, check for a sync overwriting server state in one version (the overlay bug from R2.9) or a drainer resending acknowledged rows (aDUPLICATEspike). The version split usually points at the build.
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scoping: offline reading? actions? how much to keep? images? freshness? | Restate the Round 1 design in 60 seconds | Restate the Round 2 design in 60 seconds |
| 5–15 min | Requirements; what "instant" measures; API with ETag and cursor | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Steps 1.0–1.5: local database as the only source, main-thread rule, WAL, CDN variants and disk cache, budgets | Steps 2.1–2.6: outbox + overlay, idempotency, per-user log and cursor, tombstones, error classes, battery | Steps 3.1–3.6: protocol versions, migrations, flags and kill layer, lost phones, telemetry, one spec |
| 40–50 min | Numbers: traffic, bytes per user, device storage, cost | Numbers: syncs, bytes, log entries, fleet, cost | Numbers: heartbeat-like rates, quotas, fleets, cost |
| 50–60 min | Failures + pillar check | Failures + pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint.
The Two Sentences That Matter Most
- Opening a round: "Before I design: must it work offline, for reading only or for actions too, and how much may we keep on the phone?"
- When the scope is raised: "Here's what breaks, and I'll fix it in this order: anything that can lose or duplicate a user's action, then anything that can show them something wrong, then battery, data and cost."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "What happens with no network?" (REL 11) | The screen reads the phone's database, so reading works; actions queue in an outbox and are sent later. | 1–2 | Steps 1.1, 2.1 |
| "What if a request is sent twice?" (REL 4) | Each action carries a key born on the phone, and the server's guard for it never expires. | 2 | Step 2.2 | |
| "What if a million phones reconnect at once?" (REL 5) | Full-jitter retries with Retry-After, and tasks shed before any expensive work. | 2 | R2.8 | |
| "How do you change the API with old apps in the field?" (REL 3) | Additive changes, declared capabilities with fallbacks, and a minimum version that still lets queued actions drain. | 3 | Step 3.1 | |
| Performance | "Why does the app feel instant?" (PERF 3) | Nothing on screen waits for the network, and nothing slow runs on the main thread. | 1 | Steps 1.1–1.3 |
| "How do you keep syncs small?" (PERF 3) | A per-user change log read from a server-assigned cursor, about 4 KB a sync. | 2 | Step 2.3 | |
| Security | "What if a phone is lost?" (SEC 8) | Encrypted per-account data, tokens in the Keychain or Keystore, revocation with a denylist, and a wipe on the next contact. | 3 | Step 3.4 |
| "Can a user act on something they shouldn't?" (SEC 3) | The acting user comes from the token, keys are namespaced by user, and every referenced post and media ID is checked. | 2 | Step 2.2 | |
| Cost | "Where does the money go?" (COST 8) | About $10.3K, $207K and $336K a month; image delivery is more than half from Round 2 on, and the phone's cache is what keeps it down. | 1–3 | R1.7, R2.6, R3.6 |
| "Why not use the managed service for everything?" (COST 5) | We priced it: AppConfig per-receive fees would be about $5.8M a month at our scale, so it authors flags and CloudFront delivers them. | 3 | Step 3.3, R3.6 | |
| Operations | "How do you stop a bad release?" (OPS 6) | Halt the store rollout, then switch the feature off with the kill layer; phones in use pick it up within minutes. | 3 | Step 3.3, R3.9 |
| "How do you know the app is healthy on real phones?" (OPS 4) | Daily health summaries and sampled traces per version, with thresholds counted by distinct installs. | 3 | Step 3.5 | |
| Sustainability | "How do you protect battery and data?" (SUS 2) | OS-scheduled, coalesced background work, prefetch only on Wi-Fi with battery to spare, and deltas instead of pages. | 2 | Step 2.6 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| Local data | The local database is the screen's only source; WAL; transactions | Server state and pending actions kept apart; a durable outbox with one claimed drainer | Local-only data migrated and rescued; caches rebuilt; per-account, encrypted, excluded from backup |
| Sync | Head page with ETag; cursor for older pages | A per-user log with dense, server-assigned numbers, a fenced writer and a cursor advanced in the same transaction; versioned tombstones | One log, several protocol encoders; capabilities and fallbacks; epochs when the writer changes |
| Writes | None | Idempotency keys born on the phone; durable server guards; error classes, quarantine and rollback | Old payload versions still accepted; flush allowed below the minimum version |
| Device limits | Main-thread rule; frame budgets; storage budgets | OS schedulers; battery and data arithmetic; silent-push limits | Budgets per version and device class, adjustable by flag; measured in the field |
| Numbers | Traffic, bytes per user, device storage, cost | Syncs, egress, log entries, fleets with AZ loss, quotas, cost | Heartbeat-like write rates, service quotas, pricing traps, cost |
| Well-Architected trade-offs | Storage for speed and offline use | Background freshness for battery; log storage for small syncs | Old-version support for reach; telemetry detail for privacy |
| Evolving under new scope | Builds from the baseline, one problem at a time | Opens with what can lose or duplicate actions | Designs for the versions and phones it can't update |