Design a Video Streaming Platform (YouTube)
This page is one interview loop in three rounds. All three rounds design the same system. Each round opens with the interviewer raising the scope, and the design from the round before has to evolve to meet it.
| Round 1: Mid-level | Round 2: Senior | Round 3: Architect | |
|---|---|---|---|
| Story | A company's internal training-video site | A public video platform like YouTube | Live events, global delivery economics, copyright and moderation |
| Level (Amazon) | SDE II (L5) | Senior SDE (L6) | Principal (L7) |
| Traffic | 500 uploads a day; 100K views a day; about 1,500 people watching at the busiest hour | 4.32M uploads a day (50/s, 150/s peak); 5B views a day (57.9K/s, 173.6K/s peak) | The same, plus a live final with 10M people watching at once |
| Data | About 11 TB of uploads and 11 TB of renditions a month | 6.48 PB raw and 8.16 PB transcoded a day; 469 PB delivered a day (43.4 Tbps average) | Several processing regions; three CDNs plus caches inside ISPs; storage tiers by popularity |
| Footprint | 1 region + CloudFront | 1 processing region; CloudFront worldwide | 3 processing regions; multi-CDN |
| Targets | Playback starts in < 2 s; watchable within an hour; 99.9% | Time to first frame P95 < 800 ms; rebuffering < 0.5% of sessions; 720p ready within half the video's length | Live latency about 4 s (low-latency mode) or 10 s (standard); cost per view-hour down 30% |
| Reading time | ~35 min | ~40 min | ~45 min |
You can start at any round. Rounds 2 and 3 open with a "Where we left off" summary that catches you up.
Loop Opener: What Is a Video Platform?
A TV Station Where Anyone Can Broadcast
Imagine a TV station with one unusual rule: anyone can walk in with a tape and put it on the air. The station doesn't broadcast on one channel at one quality. Every viewer tunes in whenever they like, on a phone in a train tunnel, a laptop on office Wi-Fi or a 4K TV, and each one gets a picture that fits their screen and their connection.
That is a video platform. People upload a video once, and many people watch it, on many devices, over fast and slow networks. YouTube, Vimeo and every company's internal video portal work this way.
A few words we'll use all page:
| Word | What it means on this page |
|---|---|
| Codec | The compression method for the pictures: H.264 (also called AVC), H.265 (HEVC), VP9 or AV1. Newer codecs need fewer bits for the same quality but more compute to encode. |
| Container | The file format that holds encoded video and audio: MP4, fragmented MP4 (fMP4) or MPEG-TS. A codec is how the pictures are squeezed; a container is the box they travel in. |
| Bitrate | How many bits a second the encoded stream uses, for example 2.5 Mbps. It sets both the file size and the network speed needed to play it. |
| Rendition | One encoded version of a video: a resolution, a bitrate and a codec, such as "720p, 2.5 Mbps, H.264". |
| Ladder | The set of renditions we produce for one video. |
| Keyframe and GOP | A keyframe (an I-frame; an IDR frame is a keyframe that nothing after it refers back past) can be decoded on its own. Other frames only store changes. A GOP (group of pictures) is a keyframe plus the frames up to the next one. You can only start decoding at a keyframe. |
| Segment | A few seconds of one rendition, fetched as one HTTP request. |
| Manifest (playlist) | A small text file that lists the renditions and their segments. HLS calls it a playlist (.m3u8); DASH calls it an MPD. |
| ABR | Adaptive bitrate streaming: the player picks a rendition for each segment based on its measured bandwidth and how much video it has buffered. |
| CDN | A content delivery network: many cache servers near viewers that serve the segments so our own servers don't have to. |
| Transcoding | Decoding the upload and re-encoding it into each rendition of the ladder. |
What Makes It Hard
- Raw video is enormous. A 10-minute phone video is hundreds of megabytes; a camera file can be tens of gigabytes. Moving it through our servers is the first mistake.
- One upload becomes many files. Every video is converted into several sizes and formats, and each conversion is real compute. Transcoding is the platform's biggest compute bill.
- Delivery is the biggest bill of all. Every view sends tens of megabytes. Multiply by billions of views a day and the bandwidth bill dwarfs everything else. That fact drives every decision in Rounds 2 and 3.
- Viewers are impatient. People leave if playback takes more than a couple of seconds to start or stalls in the middle.
The Question the Whole Loop Answers
How do we turn any upload into smooth playback on any device and network, at a cost we can afford?
The answer grows every round:
- Round 1: upload straight to storage, transcode into a small ladder with a managed service, cut it into segments with a manifest, and serve it through a CDN with signed cookies.
- Round 2: transcode in parallel chunks on a Spot GPU fleet, shield the origin from viral bursts, count views without a hot row, tier the petabytes, and add DRM. Delivery turns out to be nearly 90% of the bill.
- Round 3: add live streaming, spread delivery across several CDNs, catch copyrighted and harmful uploads, enforce country rules, and cut the cost of every view-hour by 30%.
Round 1 · Mid-level · "Training Videos for One Company"
~35 min · SDE II (L5) · 1 region + CloudFront · 500 uploads/day · 100K views/day · playback starts < 2 s · watchable within an hour · 99.9%
R1.1 Establish Design Scope
The interviewer says: "Our company wants an internal video site. Teams record trainings and all-hands talks, upload them, and employees watch them. Design it." Before drawing anything, we ask.
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How long are the videos? | Up to 1 hour; most are about 10 minutes. | Files up to a few GB, so uploads must survive a dropped connection (multipart, step 1.1). |
| Which devices? | Browsers on laptops, and phones. | We need formats every browser and phone can play: H.264 in HLS works everywhere (step 1.3). |
| How many quality levels? | A few. | A ladder of four renditions (step 1.2). |
| Private or public? | Private to employees. | Every request must prove the viewer is an employee, including segment requests (step 1.6). |
| Live? | Not yet. | Files only. Round 3 adds live. |
| How soon after upload must a video be watchable? | Within an hour. | Transcoding can be a background job with a queue; no rush-hour tricks. |
| How much traffic? | About 100,000 employees. Around 500 uploads and 100,000 views a day. | Small compute, but the delivery bytes add up fast (R1.7). |
Out of scope for this round: resumable uploads over flaky mobile networks, view counts, DRM, and anything public.
R1.2 Functional Requirements, Derived Step by Step
| Phrase from the problem | Operation |
|---|---|
| "Teams upload recordings" | createVideo(title, size) returns where to upload; completeUpload(id) finishes it |
| "Employees watch them" | getPlayback(id) returns a manifest URL and the cookies that let the player fetch segments |
| "Is it ready?" | getVideo(id) returns metadata and status (UPLOADING, PROCESSING, READY, FAILED) |
| "Find the all-hands from March" | listVideos(channel) and search(title) |
Not yet: resumable uploads for huge files over bad networks, adaptive bitrate at global scale, view counts.
R1.3 Non-Functional Requirements: the Questions
We name each quality first; the numbers come in R1.7.
- Playback start time. How long from pressing play to the first frame? Viewers notice anything over about 2 seconds.
- Smoothness. Playback must not stall on a hotel Wi-Fi or a phone on 4G. That needs more than one quality level.
- Durability. A recorded all-hands can't be recorded again. The original upload must not be lost.
- Cost. Compute for transcoding, storage that grows every day, and delivery bytes. We'll find out which one dominates.
R1.4 The API
Create a video and get upload URLs
httpPOST /v1/videos HTTP/1.1 Host: video.corp.example Authorization: Bearer <SSO access token> Content-Type: application/json { "title": "Q3 all-hands", "channel_id": "ch_eng", "file_name": "allhands.mp4", "size_bytes": 4500000000 }
httpHTTP/1.1 201 Created Content-Type: application/json { "video_id": "v_8f3k2m", "upload": { "type": "multipart", "upload_id": "mpu_2b91", "part_size_bytes": 67108864, "parts": [ { "part_number": 1, "url": "https://corp-video-raw.s3.us-east-1.amazonaws.com/raw/v_8f3k2m/source?partNumber=1&uploadId=mpu_2b91&X-Amz-Signature=..." }, { "part_number": 2, "url": "https://corp-video-raw.s3.us-east-1.amazonaws.com/raw/v_8f3k2m/source?partNumber=2&uploadId=mpu_2b91&X-Amz-Signature=..." } ], "urls_expire_at": "2026-10-05T15:00:00Z" } }
- Files up to 100 MB get one pre-signed
PUTURL. Bigger files get a multipart upload: the file is sent in parts (here 64 MiB each), each part retried on its own. A 4.5 GB file is ⌈4,500,000,000 ÷ 67,108,864⌉ = 68 parts. - A pre-signed URL is a normal S3 URL with a signature in the query string. It lets the browser upload one object (or one part) directly to S3 for a limited time, without AWS credentials.
Finish the upload
httpPOST /v1/videos/v_8f3k2m/complete HTTP/1.1 Authorization: Bearer <SSO access token> Content-Type: application/json { "upload_id": "mpu_2b91", "parts": [ { "part_number": 1, "etag": "\"a54f...\"" }, { "part_number": 2, "etag": "\"77c1...\"" } ] }
httpHTTP/1.1 202 Accepted Content-Type: application/json { "video_id": "v_8f3k2m", "status": "PROCESSING", "status_url": "/v1/videos/v_8f3k2m" }
Check status and metadata
httpGET /v1/videos/v_8f3k2m HTTP/1.1 Authorization: Bearer <SSO access token>
json{ "video_id": "v_8f3k2m", "title": "Q3 all-hands", "status": "READY", "duration_s": 3540, "renditions": ["1080p", "720p", "480p", "360p"], "created_at": "2026-10-05T14:02:11Z" }
Start playback
httpGET /v1/videos/v_8f3k2m/playback HTTP/1.1 Authorization: Bearer <SSO access token>
httpHTTP/1.1 200 OK Set-Cookie: CloudFront-Policy=eyJTdGF0ZW1lbnQi...; Domain=media.corp.example; Path=/; Secure; HttpOnly Set-Cookie: CloudFront-Signature=Qm9i...; Domain=media.corp.example; Path=/; Secure; HttpOnly Set-Cookie: CloudFront-Key-Pair-Id=K2JCJMDEHXQW5F; Domain=media.corp.example; Path=/; Secure; HttpOnly Content-Type: application/json { "manifest_url": "https://media.corp.example/v/v_8f3k2m/master.m3u8", "cookies_expire_at": "2026-10-05T15:10:00Z" }
Status codes
| Code | Meaning |
|---|---|
200 OK / 201 Created | Done / video record and upload URLs created |
202 Accepted | Upload complete; processing started |
400 Bad Request | File over 5 GB, a missing part, or a part list that doesn't match what S3 has |
401 Unauthorized / 403 Forbidden | No valid SSO token / not allowed to upload to this channel |
404 Not Found | No such video |
409 Conflict | playback for a video that isn't READY yet; the body says its status |
Recap
- Uploads go straight to S3 with pre-signed URLs; big files in parts.
- Playback returns a manifest URL plus cookies that unlock the video's segments.
- Four statuses; a video is only playable when
READY.
R1.5 Design Evolution: From One File to a Streamed Ladder
Each step is a problem, your turn to think, the answer, and what it costs us.
Step 1.0: The Baseline
A web app on two EC2 instances. The browser uploads the file to the app, which writes it to S3. To play, the browser gets an <video src="..."> pointing at the app, which streams the original file back from S3.
Synthesizing vector architecture diagram...
Every byte, in and out, passes through our app servers, and every viewer gets the same file the camera produced.
It works for a demo. Two problems show up in the first week.
Step 1.1: Uploads Saturate Our Servers
The problem: a team uploads ten 4 GB recordings after an offsite. The app servers' network links and memory fill up, uploads time out at 90%, and every other request slows down. What would you do?
Step 1.2: A 4 GB Camera File Won't Play on Phones
The problem: a 1-hour recording from a camera is 4.5 GB at about 10 Mbps. On a phone over 4G it takes forever to start, stalls constantly, and some phones can't decode its format at all. What would you do?
Step 1.3: Playback Stutters on Slow Networks
The problem: we now have four renditions. Which one does the player download? A laptop on hotel Wi-Fi starts at 1080p, the connection drops to 2 Mbps, and the video freezes every 20 seconds. What would you do?
Step 1.4: Viewers Far From Our Region Buffer
The problem: the London and Singapore offices complain. Every segment request travels to us-east-1 and back: from Singapore that's a round trip of more than 200 ms before the first byte, and the long path loses packets, which throttles throughput.
What would you do?
Primitive: Distributed Cache Patterns and Eviction
Step 1.5: "Is My Video Ready Yet?"
The problem: after uploading, people refresh the page over and over. Some click play too early and get errors. When a job fails, nobody finds out. What would you do?
Step 1.6: Anyone With the Link Can Watch
The problem: a manager pastes a segment URL into a public chat channel. It plays for anyone. These are internal videos. What would you do?
Round 1 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.0 | (baseline) | App servers proxy uploads and the original file | Choked servers; one file for all devices |
| 1.1 | Uploads saturate servers | Pre-signed multipart uploads straight to S3 | Short-lived URLs; abandoned parts |
| 1.2 | Originals don't play on phones | MediaConvert ladder of 4 H.264 renditions | Transcoding cost and delay |
| 1.3 | Stutters on slow networks | HLS: 4-second segments, manifests, ABR | Many small files |
| 1.4 | Far viewers buffer | CloudFront with OAC; immutable segments | CDN is the biggest bill |
| 1.5 | "Is it ready?" | DynamoDB state machine, conditional transitions, events | Two small functions |
| 1.6 | Anyone with the link | CloudFront signed cookies, 1-hour expiry | Key management |
R1.6 Architecture v1
Synthesizing vector architecture diagram...
Bytes never touch our code: uploads go to the raw bucket and viewers read through CloudFront. Our code only moves the status along and hands out permissions.
S3 layout
| Bucket | Key | What it is | Lifecycle |
|---|---|---|---|
corp-video-raw | raw/{video_id}/source | The original upload | Glacier Instant Retrieval after 30 days; abort incomplete multipart uploads after 7 days |
corp-video-out | v/{video_id}/master.m3u8 | Master manifest | Kept |
corp-video-out | v/{video_id}/{rendition}/index.m3u8 | Media playlist | Kept |
corp-video-out | v/{video_id}/{rendition}/init.mp4, seg_00000.m4s … | Init segment and 4-second fragments | Kept |
The Videos table
| Attribute | Example | Notes |
|---|---|---|
video_id (partition key) | v_8f3k2m | Random, so items spread evenly |
status | READY | Changed only by conditional writes |
channel_id, title, owner | ch_eng, Q3 all-hands, u_1042 | Search and listing |
duration_s, renditions | 3540, [1080p, 720p, 480p, 360p] | Written when the job completes |
job_id | 1712345678901-abc123 | The MediaConvert job |
created_at | 2026-10-05T14:02:11Z | Sort key of a channel_id index for "newest in this channel" |
Title search uses a small OpenSearch domain fed from the table's DynamoDB stream (the table's own change log) by a Lambda function. We never write to both from the API: a crash between the two writes would leave them different forever, while the stream carries every committed change.
Tracing an upload: the all-hands, 1 hour, 4.5 GB
Synthesizing vector architecture diagram...
At 20 Mbps of office upload bandwidth, 4.5 GB takes 4.5 × 8,000 ÷ 20 = 1,800 s = 30 minutes; the transcode then takes minutes to tens of minutes (MediaConvert gives no fixed time, so we measure it). Both fit "watchable within an hour" only if the upload link is that fast, which is why the 1-hour target is measured from upload complete.
Tracing playback start
Synthesizing vector architecture diagram...
The first segment at 480p is 4 s × (1.2 + 0.128) Mbps = 5.31 Mb ≈ 664 KB. R1.7 adds up the whole path.
R1.7 Numbers
Targets
| Quality | Target | Why this number |
|---|---|---|
| Playback start | P95 < 2 s from pressing play | The budget below gives about 1.6 s on a cache miss |
| Watchable | Within 1 hour of upload complete | From the scope |
| Availability | 99.9% for playback | 0.1% of a 30.4-day month: 43,776 min × 0.001 ≈ 44 minutes |
| Durability | Originals and renditions in S3 | S3 is designed for 99.999999999% (11 nines) durability of objects; that's AWS's design figure for S3, not something we add |
Traffic and bytes
| Item | Math | Result |
|---|---|---|
| Uploads | given | 500 a day |
| Average source | 10 min at about 10 Mbps: 600 s × 10 Mbps ÷ 8 | 750 MB |
| Raw uploads a day | 500 × 750 MB | 375 GB |
| Raw a month | 375 GB × 30.4 | ≈ 11.4 TB |
| Ladder bitrate | (5 + 2.5 + 1.2 + 0.6) Mbps video + 4 × 0.128 Mbps audio | 9.812 Mbps |
| Renditions per video | 600 s × 9.812 Mbps ÷ 8 | ≈ 736 MB |
| Renditions a month | 500 × 736 MB × 30.4 | ≈ 11.2 TB |
| Transcoding a day | 500 × 10 min = 5,000 min = 83.3 video-hours; × 4 renditions | 333 rendition-hours |
| Views | given | 100,000 a day |
| Bytes per view | we assume 6 minutes watched at 2.5 Mbps on average: 360 s × 2.5 Mbps ÷ 8 | 112.5 MB |
| Delivered a day | 100,000 × 112.5 MB | 11.25 TB |
| Delivered a month | 11.25 TB × 30.4 | ≈ 342 TB |
| Busiest hour | we assume 15% of views: 15,000 starts × 6 min ÷ 60 min | 1,500 watching at once, ≈ 3.75 Gbps |
Notice the ratio: we store about 11 TB of renditions a month but send about 342 TB. Every view re-sends the bytes.
Playback start budget, P95, cache miss, a remote office on a 10 Mbps link (every step waits for the one before, so they add):
| Step | Time |
|---|---|
GET playback to the API (SSO token check, sign cookies) | 150 ms |
| TCP and TLS to the CloudFront edge | 100 ms |
master.m3u8, edge miss to S3 in us-east-1 | 250 ms |
480p/index.m3u8 and init.mp4 (fetched in parallel), edge miss | 250 ms |
| First segment, 664 KB at 10 Mbps = 531 ms, plus a 200 ms miss | 731 ms |
| Decode and render | 100 ms |
| Total | 1,581 ms |
On a cache hit the three misses shrink to about 30 ms each and the total is about 1.0 s.
Monthly cost (us-east-1 list prices; CloudFront per-area tiered prices; 30.4-day month)
| Item | Math | Monthly |
|---|---|---|
| CloudFront data out | 342 TB: we assume 70% of viewers in the US and 30% in Europe. Each area's tiers: first 1 TB free, next 9 TB $0.085/GB, next 40 TB $0.080, next 100 TB $0.060, next 350 TB $0.040; the 1 TB free tier is per account, so it applies once. US 239 TB ≈ $13,541; Europe 103 TB ≈ $7,206 | ≈ $20,747 |
| CloudFront requests | (90 segments + 2 playlists) × 100K views × 30.4 ≈ 280M; first 10M free; US $0.0100 and Europe $0.0120 per 10,000: 270M ÷ 10,000 × (0.7 × 0.0100 + 0.3 × 0.0120) | ≈ $286 |
| MediaConvert (Basic tier) | 500 × 10 min × 30.4 = 152,000 output-minutes per rendition. Normalized minutes: HD (720p and 1080p) count 2×, SD count 1×, so 6 per source minute: 912,000. First 100,000 at $0.0075, the next 812,000 at $0.0053 | ≈ $5,054 |
| S3 storage, month 12 | Raw: last 30 days in Standard, 11.4 TB × $0.023 ≈ $262; 11 older months in Glacier Instant Retrieval, 125 TB × $0.004 ≈ $502. Renditions: 134 TB in Standard, first 50 TB at $0.023 + 84 TB at $0.022 ≈ $3,005 | ≈ $3,769 |
| API, Lambda, DynamoDB | A few million small requests | ≈ $30 |
| OpenSearch | 2 × t3.small.search × $0.036/h × 730 h | ≈ $53 |
| Logs, metrics, Secrets Manager | an estimate | ≈ $50 |
| Total, month 12 | ≈ $29,990 |
In month 1, storage is only about 22.6 TB × $0.023 ≈ $520, so the total is about $26,735. Either way delivery is about 70% of the bill, even for an internal site. MediaConvert, the thing people worry about, is 17%.
R1.8 Trade-Offs
Managed transcoding vs our own fleet
| MediaConvert (chosen) | Our own FFmpeg workers on EC2 | |
|---|---|---|
| Cost at 333 rendition-hours a day | ≈ $5K a month | Probably a few hundred dollars of compute |
| Engineering | Job templates and events | Queue, autoscaling, encoder tuning, patching, on-call |
| Formats and features | HLS, DASH, captions, many codecs | Whatever we build and test |
| Scale limit | Service quotas per account; cost grows with minutes | Only our fleet |
At 500 videos a day, the difference is one engineer's time for a month a year. That's COST 11 (the cost of effort): the managed service wins. Round 2 runs the same comparison at 4.32M videos a day and gets the opposite answer.
HLS vs DASH
| HLS (chosen) | MPEG-DASH | |
|---|---|---|
| Apple devices | Native in Safari and iOS | Needs a JavaScript player on macOS; not native on iPhone |
| Other browsers | Through a small JavaScript player (Media Source Extensions) | Same |
| Segments | fMP4 (CMAF) or MPEG-TS | fMP4 (CMAF) |
With CMAF (the common fMP4 segment format both use), the same segment files can serve both: only the manifests differ. We start with HLS alone and can add a DASH manifest later over the same segments.
How many renditions?
| Ladder | Upside | Downside |
|---|---|---|
| 2 (720p, 360p) | Half the transcoding and storage | A big quality jump when the network dips |
| 4 (chosen) | Smooth steps from 0.6 to 5 Mbps | 4× the transcoding of one rendition |
| 6+ with 4K | Great on big screens | Nobody watches a slide deck in 4K |
R1.9 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| A MediaConvert job fails | Status stuck in PROCESSING; job ERROR event | The job-done Lambda resubmits once for errors marked retryable. A second failure, or an input error ("unsupported codec"), sets FAILED with the reason, and the owner sees it and can re-upload. An alarm fires on any video in PROCESSING for more than 2 hours, in case an event was lost. |
| An upload is abandoned halfway | Parts sit in S3, invisible and billed | A lifecycle rule aborts incomplete multipart uploads 7 days after they start (lifecycle runs asynchronously, so "7 days" is when they become eligible). The API also treats a UPLOADING video older than 24 hours as expired and hides it. |
| The first viewers of a new video | Cache misses at every edge | Expected; the budget in R1.7 already assumes a miss. What we avoid is caching an error: the API never returns a manifest URL before READY, and the distribution's error-caching TTL for 403 and 404 is set to a few seconds, so a mistaken early request can't pin an error in the cache. |
| An S3 event is delivered twice | Two "object created" events | The conditional UPLOADED → PROCESSING write and the stored job_id make the second one a no-op (MediaConvert's client request token only de-dupes within one minute). |
| The signing key leaks | Anyone could mint cookies | Add a new key to the key group, switch the API to it, remove the old key; cookies signed with the old key stop working once the key-group change has finished deploying (minutes). |
R1.10 Pillar Check
| Pillar | What Round 1 covers |
|---|---|
| Reliability | Serverless API and managed transcoding across AZs; idempotent state transitions; retry once, then fail visibly REL 11 |
| Performance Efficiency | CloudFront edges near every office; 4-second segments and ABR; P95 start about 1.6 s PERF 4 |
| Security | SSO for the API; private bucket behind OAC; short-lived signed cookies; pre-signed upload URLs scoped to one key; HTTPS for every request SEC 3 · SEC 9 |
| Cost Optimization | Managed transcoding because the effort is worth more than the savings; raw files to Glacier Instant Retrieval after 30 days; about $30K a month, 70% delivery COST 11 · COST 8 |
| Operational Excellence | Skipped this round: managed services; one alarm on stuck PROCESSING videos. |
| Sustainability | Skipped this round: renditions never exceed the source resolution; old raw files move to colder storage. |
R1.11 Round 1 Rubric and Follow-Ups
What a strong mid-level (L5) answer shows
- Sends uploads straight to S3 with pre-signed multipart URLs and explains why proxying is wrong.
- Explains codec vs container and why we transcode into a ladder.
- Explains ABR: aligned keyframes, segments, master and media playlists, and why the player chooses.
- Puts a CDN in front with a private origin and different cache rules for segments and manifests.
- Tracks status with conditional writes and makes retries and notifications idempotent.
- Protects private video with signed cookies, and says why not signed URLs.
- Finds that delivery, not transcoding, dominates the bill.
Follow-up questions
-
"Why not have the browser play the MP4 file directly with a
<video>tag?" Answer: one file means one bitrate for every network and device. It can't adapt mid-video, and a 1080p file stalls on a weak connection. Segments plus manifests let the player switch every 4 seconds. -
"Why not trigger transcoding from the
completeAPI call instead of the S3 event?" Answer: the API call and the object's existence can disagree: a client could callcompletetwice, or crash before calling it after S3 assembled the object. The S3 event fires exactly when the object exists; the conditional status write makes duplicates harmless. -
"An employee leaves the company. Can they still watch?" Answer: for at most an hour, until their cookies expire; the API refuses to renew them because SSO does. If that hour matters, shorten the cookie lifetime; the player renews in the background.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Proxy uploads through the app" | Servers hold connections for minutes copying bytes; a failure restarts the file. |
| "Serve the original file" | Wrong bitrate, codec and size for most viewers; no adaptation. |
| "One medium quality for everyone" | Networks change during a video. |
| "S3 is fast, skip the CDN" | Distance, not S3, is the latency; and every viewer re-fetches the same bytes. |
| "Poll S3 for the manifest" | Existence isn't a status; failures are invisible. |
| "Unguessable URLs are private enough" | URLs leak and never expire. |
| "Transcoding is the big cost" | Delivery is about 70% of the bill. |
Round 2 · Senior · "Millions of Uploads, Billions of Views"
~40 min · Senior SDE (L6) · 1 processing region, CloudFront worldwide · 4.32M uploads/day · 5B views/day · time to first frame P95 < 800 ms · rebuffering < 0.5% of sessions · 720p ready within 0.5× the video's length · 99.99% playback
R2.0 Where We Left Off
This is what the candidate says aloud in the first 60 seconds of Round 2. If you're starting here, it's everything you need from Round 1.
Round 1 in 60 seconds. "We built an internal video site: 500 uploads and 100,000 views a day. Browsers upload straight to S3 with pre-signed multipart URLs, so no bytes pass through our servers. An S3 event starts a MediaConvert job that transcodes each upload into four H.264 renditions, 360p to 1080p, cut into 4-second segments with aligned keyframes, with HLS master and media playlists. Players use adaptive bitrate: they start at 480p and switch rendition every segment as bandwidth changes. CloudFront serves everything from a private bucket; segments are cached for a year because they never change, manifests for 5 minutes. A DynamoDB state machine with conditional writes tracks each video from
UPLOADINGtoREADY, and signed cookies with a 1-hour expiry keep the videos private. About $30K a month, and 70% of it is delivery, not transcoding. Open costs: one managed transcoder billed per minute, no resumable uploads for bad networks, no view counts, no DRM, and nothing public."
Architecture v1, compact
Synthesizing vector architecture diagram...
Round 1 in one picture: bytes go around our code, not through it.
Round 1 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 1.1 | Uploads saturate servers | Pre-signed multipart to S3 | Abandoned parts |
| 1.2 | Originals don't play on phones | MediaConvert, 4-rung H.264 ladder | Per-minute transcoding bill |
| 1.3 | Stutters | HLS segments, ABR | Many small files |
| 1.4 | Far viewers buffer | CloudFront, OAC | Delivery is the biggest bill |
| 1.5 | Is it ready? | Conditional state machine | – |
| 1.6 | Link sharing | Signed cookies | Key management |
Open costs: a per-minute transcoder, fragile uploads on mobile, no counts, no DRM, one audience.
R2.1 The Scope Raise
Interviewer: "Now it's public, like YouTube. People upload from phones everywhere. Some videos go viral and get millions of viewers in an hour. We need view counts, and some partner content must be protected by DRM."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How much is uploaded? | More than 500 hours of video a minute. (That's the figure YouTube has cited publicly since 2019.) | 720,000 hours a day. We assume 10 minutes per video: 4.32M uploads a day (R2.6). |
| How big can an upload be? | Match YouTube. Its help page has listed 256 GB or 12 hours, whichever is less; YouTube has said it's testing larger uploads with some creators. | Resumable sessions; part size must adapt so 256 GB fits in S3's 10,000 parts (R2.3). |
| How many views? | 5 billion a day. | 57.9K playback starts a second on average; delivery bytes dominate (step 2.4). |
| How fast must a video be watchable? | 720p within half the video's length; the full ladder later. | Parallel transcoding with a fast lane for the lower rungs (step 2.1). |
| Viral videos? | Millions of viewers in the first hour, starting within seconds of publishing. | Protect the origin from simultaneous misses (step 2.3). |
| View counts? | Accurate enough, visible within a couple of minutes, never double-counted by our own retries. | A buffered stream pipeline, not a counter row (step 2.5). |
| Protected content? | Partner films need DRM on top of access control. | License servers and encrypted segments (step 2.7). |
Scope change
| Round 1 | Round 2 | |
|---|---|---|
| Uploads | 500 a day, office networks | 4.32M a day, phones on flaky networks, up to 256 GB |
| Views | 100K a day, employees | 5B a day, the whole internet |
| Processing target | Within an hour | 720p within 0.5× duration |
| Playback target | Start < 2 s | First frame P95 < 800 ms; rebuffering < 0.5% of sessions |
| New features | – | View counts, DRM, several codecs |
| Availability | 99.9% | 99.99% for playback (about 4.4 minutes a month) |
The "Not yet" list from R1.2 comes back: resumable uploads, ABR at global scale and view counts are all in scope now.
R2.2 What Breaks in the Round 1 Design
| Round 1 piece | What breaks at the new scale |
|---|---|
| MediaConvert per minute | 720,000 hours a day is 1.31 billion source minutes a month. Even counting only the five rungs below 4K (7.5 normalized minutes per source minute) at the lowest listed Basic-tier rate of $0.0038, that's 9.85 billion × $0.0038 ≈ $37M a month. |
| One job per video | One GPU (at the throughput we assume in step 2.2) needs 6 ÷ 8 = 45 minutes for a 1-hour video's full ladder. The 720p target for that video is 30 minutes, and a crash at 95% throws away 43 minutes of work. |
| Uploads over phones | A 1.5 GB upload over a 20 Mbps uplink takes 10 minutes. Phones switch networks, sleep and lose signal; a 1-hour URL expiry and no way to ask "which parts arrived?" means starting over. |
| One origin behind the CDN | A viral video's first segments miss in every CloudFront cache region at the same moment, and each one asks S3. |
| A counter | A viral video at 100,000 views a second would need 100,000 writes a second on one DynamoDB item; one partition takes at most 1,000 write units a second. |
| Keep everything in S3 Standard | We add 14.64 PB a day, about 445 PB a month. At $0.021/GB, every month adds about $9.35M to the monthly bill, forever. |
| Abandoned uploads | If 5% of upload bytes are abandoned, that's 6.48 PB × 5 ÷ 95 ≈ 341 TB a day of invisible parts, about 10.4 PB a month. |
R2.3 New Requirements and API Additions
Resumable upload sessions
httpPOST /v1/uploads HTTP/1.1 Authorization: Bearer <access token> Content-Type: application/json { "file_name": "trip.mov", "size_bytes": 1500000000, "content_type": "video/quicktime", "title": "Iceland in 10 minutes" }
httpHTTP/1.1 201 Created Content-Type: application/json { "upload_id": "up_71c9", "video_id": "v_Q2x9", "part_size_bytes": 16777216, "part_count": 90, "checksum_algorithm": "CRC32C", "session_expires_at": "2026-10-12T14:00:00Z" }
- Part size adapts to the file. S3 allows at most 10,000 parts, each 5 MiB to 5 GiB. We use 16 MiB, or more when the file needs it: ⌈size ÷ 10,000⌉ rounded up to a whole MiB. 1.5 GB ÷ 16 MiB = 89.4, so 90 parts. A 256 GB file needs 25.6 MB per part, so 25 MiB (26,214,400 bytes): 9,766 parts.
- URLs in batches.
GET /v1/uploads/up_71c9/part-urls?from=1&count=20returns 20 pre-signed part URLs valid for 1 hour. The phone asks for more as it goes, so a URL is never older than an hour. (A pre-signed URL signed with temporary credentials also stops working when those credentials expire, another reason to hand them out in small batches.) - Per-part checksums. Each part is sent with a CRC32C checksum that S3 verifies on arrival; a part damaged in transit is rejected and re-sent.
- Resume.
GET /v1/uploads/up_71c9asks S3 which parts it has (ListParts) and returns{ "parts_received": [1, 2, …, 57] }. The phone sends only the rest. - Complete.
POST /v1/uploads/up_71c9/completeassembles the object. The session expires after 7 days; the API checkssession_expires_atitself, and an S3 lifecycle rule cleans up the parts.
View events
httpPOST /v1/views/events HTTP/1.1 Content-Type: application/json { "events": [ { "view_id": "vw_5e1b0c7a", "video_id": "v_Q2x9", "type": "qualified", "position_s": 31.2, "watch_s": 30.0, "client_ts": "2026-10-05T14:03:02Z" } ] }
httpHTTP/1.1 202 Accepted
- The player creates one random
view_idwhen playback starts. It sendsstart,qualified(once the viewer has watched long enough to count as a view, a threshold we choose) andend: three small requests per view. 202means "queued", not "counted".
Playback by device capability
httpPOST /v1/videos/v_Q2x9/playback HTTP/1.1 Content-Type: application/json { "codecs": ["avc1", "hvc1"], "drm": ["fairplay"], "screen_height": 1170 }
json{ "manifest_url": "https://media.example.com/v/3f/v_Q2x9/r1/h264/master.m3u8", "license_url": null, "cookies_expire_at": "2026-10-05T14:33:00Z" }
The API picks the manifest for the codecs the device can decode (H.264 everywhere; HEVC for the 4K rung on devices that have it; AV1 comes in Round 3) and, for protected titles, the DRM system the device supports.
DRM licenses (described). For protected titles, the player's DRM module sends a license request to our license service with its playback token. The service checks the viewer's rights and returns the content key, encrypted so only that device's DRM module can use it (step 2.7).
R2.4 Design Evolution: From One Job to a Parallel Pipeline
Step 2.1: Transcoding a 1-Hour Video Takes Hours
The problem: one worker transcodes a whole upload, rung after rung. A 1-hour 4K upload needs about 45 minutes of GPU time for the full ladder, and 720p isn't out until it's that rung's turn. At 50 uploads a second, one job per video also means a crash near the end wastes most of the work. What would you do?
Primitive: Message Queues vs Event Streams
Step 2.2: The GPU Fleet Costs a Fortune
The problem: the demand is 720,000 hours of video a day, each hour becoming 6 rendition-hours. At on-demand prices a fleet that size costs about $22M a month. What would you do?
Primitives: Circuit Breaker, Bulkhead and Fault Tolerance Patterns (the fast and full lanes are bulkheads: a 4K backlog can never delay 720p) · Drill: Message Queue for an Order Pipeline. A worker that crashes halfway through a task never deletes the message, so it reappears after the visibility timeout and another worker redoes the whole task (outputs are idempotent); a task that fails 5 times goes to a dead-letter queue and marks the video FAILED. We use a queue rather than calling workers directly because workers come and go with Spot, uploads arrive in bursts of 3×, and the two lanes need different priorities: the queue absorbs all three.
Step 2.3: A Viral Video's First Minute Floods the Origin
The problem: a creator with 50 million subscribers publishes. Within 2 seconds, hundreds of thousands of players around the world ask for master.m3u8 and the first segments. None of them are cached anywhere yet.
What would you do?
Drill: CDN Edge Image Delivery (the same two questions, for images)
Step 2.4: Egress Is Most of the Bill
The problem: the first month's bill arrives. Nearly 90% of it is CloudFront data transfer. What would you do?
Step 2.5: A Counter Row per Video Melts on Viral Videos
The problem: each view does UpdateItem ADD views 1 on the video's DynamoDB item. A viral video gets 100,000 views a second.
What would you do?
Primitive: Database Sharding and Partition Keys
Step 2.6: Storage Grows by Petabytes a Day
The problem: 6.48 PB of raw uploads and 8.16 PB of renditions arrive every day. In S3 Standard, a year of it (5.34 EB × $0.021) costs about $112M a month and rising. What would you do?
Step 2.7: Partner Content Must Be Protected
The problem: a film studio will license its catalog only if "the files can't be downloaded and shared". Our signed cookies already stop strangers. What would you do?
Round 2 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | Hours to transcode | Keyframe-aligned 40 s chunks, fast and full lanes, whole-file audio, Step Functions + SQS | Exact joins; moving parts |
| 2.2 | GPU fleet cost | 22,500 GPUs average; 80% Spot, SQS visibility leases, idempotent outputs, re-sent callbacks | Reruns; Spot scarcity |
| 2.3 | Viral bursts | Origin Shield, immutable versioned paths, prefetch | Shield request fees |
| 2.4 | Egress | ABR start-low, screen-size caps, 98.5% hit ratio | Extra codecs later |
| 2.5 | Hot counter | Kinesis → Flink dedupe → conditional last_window totals | Counts lag 1–2 min |
| 2.6 | Petabytes a day | Standard → Glacier IR → Deep Archive for raw; Intelligent-Tiering for single-file renditions | Slow cold access |
| 2.7 | Protected content | CENC cbcs, Widevine, FairPlay, PlayReady licenses | License path |
R2.5 Architecture v2
Synthesizing vector architecture diagram...
Four independent paths: uploads land in S3, processing is a queue-driven fleet, delivery never touches our servers except for playback authorization, and views flow through a stream into conditional writes.
Data layout
| Store | Key | Holds |
|---|---|---|
Videos table | pk = VIDEO#{id}, sk = META | owner, title, status (UPLOADING → PROCESSING → PLAYABLE → READY / FAILED), current version, renditions, visibility |
Videos table | pk = VIDEO#{id}, sk = VIEWS | total, last_window; never expires |
Jobs table | pk = JOB#{id}#{version} | chunk total per lane, set of done chunks per lane, task tokens |
| S3 raw | {h2}/{video_id}/source | original; tiered as in step 2.6 |
| S3 renditions | {h2}/{video_id}/r{n}/{codec}/{rung}.mp4, .../{rung}.m3u8, {h2}/{video_id}/master.m3u8 | one CMAF file per rendition; immutable except the master |
| S3 scratch | scratch/{video_id}/{lane}/{chunk}/{rung}.mp4 | chunk outputs; lifecycle expiry after 1 day |
Search is fed the same way as Round 1, from the Videos table's DynamoDB stream, now into a larger OpenSearch cluster.
Trace: upload to 720p ready, a 10-minute 1.5 GB phone video
Synthesizing vector architecture diagram...
The upload itself took about 10 minutes at 20 Mbps (1.5 GB × 8 ÷ 20 Mbps = 600 s); processing starts only when it's complete. Full-lane tasks (≈ 32 s each) run in parallel, so under normal load READY follows a couple of minutes after PLAYABLE.
Trace: a DRM playback
Synthesizing vector architecture diagram...
The license call happens in parallel with fetching the first segment, so it adds about one round trip to the start, not a full sequence.
R2.6 Numbers and Cost
Traffic
| Item | Math | Result |
|---|---|---|
| Uploaded hours a day | 500 h/min × 60 × 24 | 720,000 |
| Uploads a day | 720,000 × 60 ÷ 10 min average | 4.32M |
| Uploads a second | 4.32M ÷ 86,400 | 50 average, 150 at 3× peak |
| Views a second | 5 × 10⁹ ÷ 86,400 | 57,870 average, 173,610 peak |
| View events | 3 per view: 15 × 10⁹ a day ÷ 86,400 | 173.6K/s average, 520.8K/s peak |
| View-hours a month | 5 × 10⁹ × 5 min ÷ 60 × 30.4 | 12.67 billion |
Storage
| Item | Math | Result |
|---|---|---|
| Raw per video | 20 Mbps source × 600 s ÷ 8 | 1.5 GB |
| Raw a day | 4.32M × 1.5 GB | 6.48 PB |
| Ladder bitrate | 15 (4K) + 5.5 (1080p) + 2.5 + 1.2 + 0.6 + 0.3 Mbps + 0.128 audio | 25.2 Mbps |
| Renditions per video | 600 s × 25.2 Mbps ÷ 8 | 1.89 GB |
| Renditions a day | 4.32M × 1.89 GB | 8.16 PB |
| Both, per year | (6.48 + 8.16) PB × 365 | ≈ 5.34 EB |
Bandwidth
| Item | Math | Result |
|---|---|---|
| Upload ingress | 50/s × 1.5 GB = 75 GB/s | 600 Gbps average, 1.8 Tbps peak |
| Delivery | step 2.4 | 43.4 Tbps average, 130.2 Tbps peak |
| From origin at a 98.5% hit ratio | 43.4 Tbps × 0.015 | ≈ 651 Gbps average, 1.95 Tbps peak |
Memory
| Item | Math | Result |
|---|---|---|
| Metadata cache for watch pages | we assume the top 20% of views touch 50M videos; 50M × 4 KB | 200 GB: 6 shards of cache.r7g.2xlarge (52.8 GiB each, 200 GB ÷ 6 ≈ 33 GB used), each with a replica |
| Flink dedupe state | 5 × 10⁹ view IDs a day × about 50 B | ≈ 250 GB, on local disk (RocksDB state), not RAM |
| Flink window state | we assume 2M distinct videos per minute × 64 B | ≈ 128 MB |
Time to first frame, P95 < 800 ms. The playback call and the TLS connection to CloudFront happen while the watch page loads, so they're off the critical path. After "play" (dependent steps add; parallel fetches count once):
| Step | P95 |
|---|---|
| New connection to the edge if the page's one was closed (2 round trips at 60 ms) | 120 ms |
master.m3u8 (edge hit) | 70 ms |
| Video and audio playlists, in parallel | 70 ms |
| Init segments, in parallel | 70 ms |
| First 360p segment, 300 KB at an assumed 8 Mbps P95 start throughput (300 ms) plus a round trip; audio segment in parallel | 370 ms |
| Decode and render | 50 ms |
| Total | 750 ms |
That only holds on a cache hit, which is why prefetching (step 2.3) and the hit ratio matter: a miss adds 100–200 ms, and the long tail of rarely watched videos lives in the slowest 5%.
Quotas to raise before launch (current defaults). CloudFront: 150 Gbps and 250,000 requests a second per distribution; we need 43.4 Tbps and about 9 million requests a second on average (775 billion a day ÷ 86,400), so this is a conversation with AWS, not a form. EC2: GPU-family vCPU quotas. Step Functions: we fit the default 5,000 state transitions a second. SQS: 120,000 in flight per standard queue.
Monthly cost (us-east-1 list prices; CloudFront per area and tiered; end of year one for storage)
We assume where viewers are: 30% US/Mexico/Canada, 25% Europe, 15% India, 15% the Hong Kong/Singapore/Southeast Asia group, 10% South America and 5% Japan. Each area's volume is priced on its own tiers (an assumption about how the tiers aggregate), and at petabyte volumes nearly all of it sits in the "over 5 PB" tier.
| Item | Math | Monthly |
|---|---|---|
| CloudFront data out | 14.25 EB split by area. "Over 5 PB" rates: US and Europe $0.020/GB, India $0.072, Asia group $0.060, South America $0.040, Japan $0.060, with the first ~5 PB of each area at the higher tiers. US 4.28 EB ≈ $85.5M; Europe 3.56 EB ≈ $71.3M; India 2.14 EB ≈ $153.9M; Asia 2.14 EB ≈ $128.3M; South America 1.43 EB ≈ $57.1M; Japan 0.71 EB ≈ $42.8M | ≈ $539.0M |
| CloudFront requests | per view: 75 video + 75 audio segments (5 min ÷ 4 s) + 5 playlists = 155; × 5B × 30.4 = 23.56 trillion; blended $0.0124 per 10,000 (US $0.0100, most areas $0.0120, South America $0.0220) | ≈ $29.2M |
| Origin Shield | 1.5% of requests miss the edges; we assume 90% of those pass through the Shield: 318 billion × $0.0085 per 10,000 | ≈ $0.27M |
S3 GETs | we assume the Shield hits 60%, so 40% of its 318 billion requests reach S3: 127 billion × $0.0004 per 1,000 | ≈ $0.05M |
| GPU fleet | step 2.2 | ≈ $11.5M |
| Raw storage | Standard 45.4 PB × $0.021 ≈ $0.95M; Glacier IR 583 PB × $0.004 ≈ $2.33M; Deep Archive 1,737 PB × $0.00099 ≈ $1.72M | ≈ $5.0M |
| Rendition storage | 2,978 PB after a year. We assume the last 30 days plus 10% of older bytes stay frequent: 518 PB × $0.021 ≈ $10.9M; 90% of days 30–90 in the infrequent tier: 441 PB × $0.0125 ≈ $5.5M; 90% of older days in archive instant access: 2,020 PB × $0.004 ≈ $8.1M; monitoring 23.7 billion objects × $0.0025 per 1,000 ≈ $0.06M | ≈ $24.5M |
S3 PUTs | about 200 per video (90 upload parts, 90 chunk outputs, packaged files): 26.3 billion × $0.005 per 1,000 | ≈ $0.13M |
| Transfer Acceleration | we assume 25% of upload bytes use it: 49.2 PB × $0.04/GB | ≈ $1.97M |
| Orchestration | Step Functions 1.31 billion transitions × $0.025 per 1,000 ≈ $32.8K; SQS 19.7 billion requests × $0.40 per million ≈ $7.9K; jobs table ≈ $4.9K | ≈ $0.05M |
| View pipeline | ALB: 173.6K new connections a second ÷ 25 per LCU = 6,944 LCUs × $0.008 × 730 ≈ $40.6K; collectors 35 ARM Fargate tasks ≈ $2.0K; Kinesis 200 shards ≈ $2.3K; Flink 48 KPUs (+1 for the application) ≈ $4.2K; DynamoDB 33.3K conditional writes a second (2M videos a minute) ≈ $54.7K on demand | ≈ $0.10M |
| Metadata and APIs | cache nodes (12 × $0.70/h) ≈ $6.1K; API Fargate ≈ $14K; API load balancer ≈ $40K; DynamoDB reads ≈ $3K; search ≈ $20K (an estimate) | ≈ $0.08M |
| Total | ≈ $611.9M |
Delivery (data out plus requests) is $568M, 93% of the bill. The GPU fleet everyone worries about is 1.9%. Per view-hour: $611.9M ÷ 12.67 billion ≈ $0.048.
Three things the table teaches:
- Segment count is money. Requests alone are $29M a month, driven by 4-second segments with separate audio. Round 3 revisits segment length.
- Where the bytes don't go. GPU workers read the raw file from S3 through a VPC gateway endpoint (free). Through a NAT gateway, 6.48 PB a day at $0.045/GB would add about $292K a day. And nothing on the hot path crosses AZs between our own instances: S3 is regional, so reading it from any AZ costs nothing extra, while instance-to-instance traffic across AZs is billed $0.01/GB in each direction.
- Upload acceleration isn't free. Transfer Acceleration costs $0.04/GB on top of free S3 ingress; that's $2M a month even at 25% use. Round 3 moves uploads to regional buckets instead.
R2.7 Trade-Offs
MediaConvert vs our own fleet at this scale
| MediaConvert | Own GPU fleet (chosen) | |
|---|---|---|
| Monthly cost | ≥ $37M even without 4K, at the lowest listed Basic rate | ≈ $11.5M including 4K |
| Control | Its ladder features and its queueing | Chunking, lanes, per-title ladders, custom codecs |
| Effort | Low | A team to run the fleet and tune encoders |
| When it wins | Round 1 scale, spiky or rare workloads, broadcast features | Steady, huge volume |
We keep MediaConvert for unusual inputs our fleet rejects (exotic professional codecs), so the fleet stays simple.
Chunk size
| Chunk | Tasks per 10-min video | Effect |
|---|---|---|
| 4 s (one GOP) | 150 per lane (900 at one task per rung) | Fastest turnaround; 3.9 billion tasks a day at one task per rung, each paying per-task overhead |
| 40 s (chosen) | 15 per lane | Turnaround in seconds; overhead small |
| 5 min | 2 per lane | Little parallelism for short videos; a Spot loss wastes minutes |
Segment length (what players fetch; separate from chunks)
| 2 s | 4 s (chosen) | 6 s | |
|---|---|---|---|
| Start and switch speed | Fastest | Good | Slower to adapt |
| Requests per 5-min view | 300 | 150 (+ playlists) | 100 |
| Compression | Worse (a keyframe every 2 s) | Good | Best |
Apple's HLS authoring guidance recommends 6-second segments. We chose 4 s for faster adaptation on mobile networks; Round 3 moves VOD to 6 s once request costs matter more.
Codec ladder
| Codec | Plays on | Bits for equal quality | Encode cost |
|---|---|---|---|
| H.264 | Everything | Baseline | Lowest; hardware everywhere |
| HEVC | Apple devices, most TVs, browsers only where the hardware supports it | Fewer (often 30–40% less, content-dependent) | Higher |
| AV1 | Newer phones, TVs and browsers | Fewer still (often 30–50% less than H.264) | Highest; hardware encoders only on recent GPUs (the A10G decodes AV1 but can't encode it) |
Every video gets H.264; HEVC only for the 4K rung; AV1 only for popular videos (Round 3).
Exact vs approximate counts
| Exact, transactional | Approximate, windowed (chosen) | |
|---|---|---|
| Write rate on a viral video | 100K/s on one item: impossible | 1 a minute |
| Freshness | Instant | 1–2 minutes |
| Correct under retries | Needs idempotency keys per view in the transaction | Dedupe in Flink plus conditional last_window |
| Fraud | Counted, then must be subtracted | Recomputed from S3 events |
R2.8 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| Chunk join artifacts | A visible glitch or a stall every 40 s in one rung | Closed GOPs with an IDR at every boundary; frame-count timestamps; the packager's settings check re-queues a mismatched chunk; an automated check decodes each joined file across every boundary before READY. |
| Audio/video drift | Lips out of sync late in long videos | Audio is one continuous encode; video timestamps come from frame counts on a constant frame rate, never from summed rounded durations. |
| A Spot reclaim wave | Hundreds of instances get 2-minute notices at once | Workers finish their ≤ 32 s task or release it; lost tasks reappear after the 120 s visibility timeout. The fast lane's on-demand floor keeps 720p flowing. The autoscaler shifts to other instance types and AZs. |
| Origin thundering herd | S3 503 Slow Down on new, popular videos | Origin Shield collapses requests; hashed prefixes spread load; prefetch warms phones. Alarm on S3 5xx for the renditions bucket. |
| Hot counter | Throttling on a video's item, or one hot Kinesis shard | The item gets one conditional write per window, whatever the view rate. Kinesis is partitioned by the random view_id, so a viral video's events spread over all 200 shards; Flink pre-sums per worker before the per-video total. Alarm on the busiest shard's write rate, not only the stream's total. |
| Zombie multipart uploads | Storage grows with no visible objects | Lifecycle rule AbortIncompleteMultipartUpload after 7 days; an S3 Storage Lens alarm on incomplete-upload bytes, in case the rule is missing on a new bucket. |
| The search indexer stops for more than 24 hours | Search misses new videos | A DynamoDB stream keeps records for 24 hours, and a consumer can start from the oldest record, the newest, or a sequence number it saved, but never from a timestamp; after a gap longer than 24 hours only the oldest (TRIM_HORIZON) or the newest (LATEST) are useful. So after a longer outage we first restart the consumer at LATEST, then load a fresh table export to S3 (which requires point-in-time recovery on the table) into the index. The two overlap, which is harmless because each index write carries the item's version and an older version never replaces a newer one. |
| The license service is down | Protected titles won't start | It runs in three AZs behind its own load balancer; players cache licenses for the session. Public videos are unaffected because they don't use it. |
R2.9 Production Gotchas
| Gotcha | Why it hurts | What we do |
|---|---|---|
| Monolithic transcoding | Turnaround grows with length; a crash near the end loses everything | 40 s chunks in two lanes. |
Long cache times on master.m3u8 | The master changes when the top rungs finish (and on takedowns); a 24-hour cache pins the old one | max-age=60 on masters; immutable, versioned paths for everything else. |
| Audio muxed into every rendition | The same audio is stored and delivered 6 times over; an ABR switch can make audio jump | One audio rendition bound to all video rungs with #EXT-X-MEDIA:TYPE=AUDIO. |
| Proxying uploads through app servers | Servers hold multi-GB streams; slow clients exhaust connections | Pre-signed part URLs; the API never sees the bytes. |
| One object per segment at this scale | Trillions of objects; Intelligent-Tiering monitoring and PUT fees in the millions | One CMAF file per rendition, byte-range playlists. |
| S3 through a NAT gateway | $0.045/GB on 6.48 PB a day | VPC gateway endpoint for S3. |
| Default CloudFront quotas | 150 Gbps per distribution by default | Raised with AWS before launch. |
R2.10 Pillar Check
| Pillar | What Round 2 covers |
|---|---|
| Reliability | Idempotent chunk tasks with SQS leases; re-sent callbacks; Spot interruption handling; fast lane on an on-demand floor; quotas raised ahead REL 1 · REL 5 · REL 11 |
| Performance Efficiency | Parallel chunked transcoding (720p in about 39 s for a 10-minute video); Origin Shield and a 98.5% hit ratio; start-low ABR; first frame P95 about 750 ms PERF 2 · PERF 4 |
| Security | Signed cookies for private videos; DRM for licensed titles; private origin behind OAC; TLS everywhere SEC 3 · SEC 8 · SEC 9 |
| Cost Optimization | Own fleet vs $37M+ of MediaConvert; 80% Spot; lifecycle tiers; one object per rendition; delivery identified as 93% of the bill COST 7 · COST 8 · COST 4 |
| Operational Excellence | Alarms on queue age per lane, stuck videos, S3 5xx, hit ratio; a rebuild path for the search index OPS 8 |
| Sustainability | Light this round: renditions never above the source; cold data in archive tiers; screen-size caps send fewer bytes SUS 4 |
R2.11 Round 2 Rubric and Follow-Ups
What a strong senior (L6) answer shows
- Splits transcoding in time at keyframes, explains closed GOPs, frame-count boundaries and whole-file audio, and chooses a chunk size with reasons.
- Sizes the GPU fleet from rendition-hours, labels the throughput assumption, and uses Spot with leases that don't depend on a database clock.
- Makes retries safe everywhere: idempotent outputs, conditional completion, and callbacks re-sent when a retry finds the work already recorded.
- Explains how misses multiply through a CDN and what Origin Shield actually collapses.
- Does the egress math and concludes that delivery, not compute, is the bill.
- Counts views without a hot row and without double-counting replays.
- Tiers storage by access pattern and notices the per-object costs.
- Separates access control (signed cookies) from copy protection (DRM).
Follow-up questions
-
"Why not transcode on the phone before uploading?" Answer: phones could shrink the upload, but we'd lose the original quality forever, battery and heat would suffer, and results would vary by device. Some apps do a light re-encode of very large files; the ladder is still made on our side from the best source we have.
-
"A popular creator uploads at the daily peak. Is their 720p still fast?" Answer: yes. The fast lane is about 7% of the GPU work and is always polled first, with an on-demand floor, so it keeps up even at 3× the average upload rate. The 1080p and 4K rungs may lag by hours at the peak; the master playlist gains them when they're ready.
-
"Why doesn't the count use exactly-once transactions?" Answer: it doesn't need them. Flink's checkpointed state makes the dedupe exact within its 24 hours, and the conditional
last_windowmakes the write idempotent per window. Anything left over (late events, fraud) is fixed by recomputing from the stored events.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "A faster machine" | Doesn't parallelize and still loses everything on a crash. |
| "GPUs = video-hours ÷ throughput" | Forgets the 6 rungs: 6× too small. |
| "Decode once halves the fleet" | Encoding is the bottleneck; decode-once saves decode and I/O. |
| "Scale the origin for viral videos" | The load is duplicates; collapse them instead. |
| "Increment a counter per view" | 100× over the partition limit, and replays double-count. |
| "Signed URLs protect partner films" | They stop strangers, not subscribers who save the files. |
| "Transcoding is the big cost" | Delivery is 93%; the GPU fleet is 2%. |
Round 3 · Architect · "Live, Global, and Paying for Every Byte"
~45 min · Principal (L7) · 3 processing regions, 3 CDNs plus ISP caches · VOD as in Round 2, plus a live final with 10M viewers at once · live latency about 4 s (low-latency mode) or 10 s (standard) · cost per view-hour down 30%
R3.0 Where We Left Off
Round 2 in 60 seconds. "We run a public platform from one processing region: 500 hours uploaded a minute, 4.32M uploads and 5 billion views a day. Phones upload through resumable sessions with adaptive part sizes. Each upload is probed and planned as 40-second chunks on keyframe boundaries; a fast lane makes 720p and below within about 39 s for a 10-minute video, and a full lane adds 1080p and 4K. About 22,500 GPUs on average, 80% Spot, with SQS visibility timeouts as leases, idempotent outputs and callbacks re-sent on retry. Renditions are single CMAF files with byte-range playlists in S3 Intelligent-Tiering; raw files go from Standard to Glacier Instant Retrieval to Deep Archive. CloudFront with Origin Shield serves 43.4 Tbps on average at a 98.5% hit ratio; players start low and climb, capped by screen size. Views flow through Kinesis and Flink into one conditional write per video per minute. DRM protects partner titles. About $612M a month at list prices, $0.048 per view-hour, and 93% of it is delivery. Open costs: no live, one CDN, one processing region, nothing catches copyrighted or harmful uploads, no country rules, and the bill."
Architecture v2, compact
Synthesizing vector architecture diagram...
Round 2 in one picture: a queue-driven transcoding fleet, one CDN with a shield, a stream for counts.
Round 2 step summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 2.1 | Hours to transcode | 40 s chunks, fast and full lanes | Exact joins |
| 2.2 | GPU cost | 80% Spot, SQS leases | Reruns |
| 2.3 | Viral bursts | Origin Shield, versioned paths | Shield fees |
| 2.4 | Egress | Start-low ABR, screen caps, hit ratio | – |
| 2.5 | Hot counter | Kinesis → Flink → conditional totals | 1–2 min lag |
| 2.6 | Petabytes a day | Tiers; one file per rendition | Slow cold reads |
| 2.7 | Protected titles | DRM with three systems | License path |
Open costs: VOD only, one CDN, one region, no rights or moderation, no country rules, $0.048 per view-hour.
R3.1 The Scope Raise
Interviewer: "We're adding live: creators stream every day, and next summer we have the rights to a football final. Rights holders want their content caught when people re-upload it. Harmful content must come down fast, and some countries require specific videos to be blocked there. And the CFO wants the cost per view-hour down 30%."
| We ask | Interviewer answers | What it changes in the design |
|---|---|---|
| How live is live? | Sports fans need a few seconds, so spoilers don't arrive by text first. Most creators are fine with 10 seconds; chat is a separate system. | Two latency modes (step 3.1). The chat loop covers chat. |
| The biggest event? | The final: 10M people watching at once, for about 3 hours. | 10M × 4 Mbps = 40 Tbps on top of the evening VOD peak (step 3.2). |
| Creator streams? | About 50,000 at once at the busiest time. | A live transcoding fleet, and fingerprinting of live streams too. |
| What do rights holders want? | To upload reference files and choose, per territory: block, monetize, or just track. | Audio and video fingerprinting against references (step 3.3). |
| How fast must harmful content come down? | New viewers must be blocked within seconds of a decision, and the video gone from caches within minutes. | A moderation pipeline and a takedown path through every cache layer (step 3.4). |
| Country rules? | A court or regulator in one country can require a video be unavailable there only. | Per-country availability as a separate override layer (step 3.5). |
| Failures? | A whole AWS Region can be lost without stopping uploads. | Several processing regions (R3.5, R3.8). |
Scope change
| Round 2 | Round 3 | |
|---|---|---|
| Content | VOD | VOD + live (creators and a 10M-viewer final) |
| Delivery | CloudFront | 3 CDNs + caches inside ISPs, steered per session |
| Processing | 1 region | 3 regions: us-east-1, eu-west-1, ap-south-1 |
| Rules | None | Copyright matching, moderation, per-country availability |
| Cost | $0.048 per view-hour | ≤ $0.034 (−30%) |
R3.2 What Breaks in the Round 2 Design
| Round 2 piece | What breaks |
|---|---|
| A VOD-only pipeline | Files are processed after they're complete. A live stream never completes, and 40-second chunks alone are 40 s of latency. |
| A single CDN | The final adds 40 Tbps on top of an evening VOD peak of about 95 Tbps (after this round's savings). One vendor's outage or capacity shortfall is then a global outage, and we have no price competition. |
| No fingerprinting | Re-encoding changes every byte, so a file hash never matches a re-upload. Trimming, cropping or changing the pitch defeats any exact comparison. |
| Manual moderation | 4.32M uploads a day at 1 minute of human viewing each is 72,000 hours a day, about 9,000 reviewers on 8-hour shifts, and still too slow. |
| One global catalog | A video is either available or not. No way to block it in one country, and no single place where overrides live. |
| One processing region | A Region outage stops all uploads, and far-away uploaders pay for Transfer Acceleration. |
R3.3 New Requirements and API Additions
Create a live stream
httpPOST /v1/live/streams HTTP/1.1 Authorization: Bearer <access token> Content-Type: application/json { "title": "Cup final", "latency_mode": "low", "redundant_ingest": true }
json{ "stream_id": "ls_77ab", "ingest": [ { "protocol": "srt", "url": "srt://ingest-a.live.example.com:9000", "passphrase_ref": "sec_91" }, { "protocol": "srt", "url": "srt://ingest-b.live.example.com:9000", "passphrase_ref": "sec_92" } ], "playback_url": "https://live.example.com/l/ls_77ab/master.m3u8" }
- RTMP (Real-Time Messaging Protocol, over TCP; RTMPS with TLS) is what most streaming software sends. SRT (Secure Reliable Transport, over UDP) resends lost packets within a latency window we choose and encrypts the stream; it holds up better over long, lossy links. We accept both; the final uses SRT to two independent ingest points.
latency_modeislow(about 4 s) orstandard(about 10 s). R3.6 shows why not everyone getslow.
Rights-holder references
httpPOST /v1/rights/references HTTP/1.1 Authorization: Bearer <partner token> Content-Type: application/json { "asset_id": "ref_ucl_final_feed", "upload_id": "up_9f2", "match": ["audio", "video"], "policy": [ { "territories": ["DE", "AT"], "action": "block" }, { "territories": ["*"], "action": "monetize" } ] }
Moderation and legal decisions
httpPOST /v1/decisions HTTP/1.1 Authorization: Bearer <reviewer token> Content-Type: application/json { "video_id": "v_Q2x9", "action": "block", "scope": { "countries": ["TR"] }, "source": "legal", "reason_code": "court_order", "reference": "case 2026/1142" }
httpHTTP/1.1 201 Created Content-Type: application/json { "decision_id": "dec_01J9ZK4T7M", "video_id": "v_Q2x9", "effective": "immediately" }
A decision is never edited. Lifting it is a new decision that names it: { "action": "lift", "lifts": "dec_01J9ZK4T7M" }.
Availability
GET /v1/videos/v_Q2x9/availability?country=TR returns { "available": false, "reason": "legal" }, and playback there returns 451 Unavailable For Legal Reasons.
R3.4 Design Evolution: Live, Rules and Money
Step 3.1: Live Streams
The problem: the VOD pipeline starts when a file is complete. A live stream is never complete, and fans want to see a goal within a few seconds of the stadium. What would you do?
Step 3.2: 10M Viewers for a Final
The problem: during the final, the CDNs must carry about 106 Tbps (R3.6): the evening VOD peak plus 40 Tbps of live. If our one CDN has a bad hour, 10M people see a spinner at the same moment. What would you do?
Step 3.3: Catch Copyrighted Uploads
The problem: a broadcaster's match highlights are re-uploaded hundreds of times a day, cropped, mirrored, with the audio sped up 5%. What would you do?
Step 3.4: Harmful Content Must Come Down Fast
The problem: a violent video spreads. Reviewers decide to remove it. It's cached at hundreds of edge locations in three CDNs and inside ISPs, playing in thousands of sessions, and listed in search and recommendations. What would you do?
Step 3.5: Available in Some Countries, Not Others
The problem: a court in one country orders a video unavailable there. It must stay available everywhere else. Next week the creator edits the title, the video is re-encoded for AV1, and a Region fails over. What would you do?
Step 3.6: Cut Cost per View-Hour by 30%
The problem: Round 2 costs about $612M a month at list prices, $0.048 per view-hour. The CFO wants $0.034. What would you do?
Round 3 Step Summary
| Step | Problem | Component | What it costs us |
|---|---|---|---|
| 3.1 | Live streams | SRT/RTMP ingest, redundant transcoders, LL-HLS parts and 2 s segments | Always-on compute; requests |
| 3.2 | 10M viewers | 3 CDNs, per-session steering with keep-alive documents, player-side failover, event reservations | Three contracts to keep identical |
| 3.3 | Copyright | Audio and video fingerprints, aligned-offset matching, per-territory policies | False positives |
| 3.4 | Harmful content | Classifiers, exposure-ranked review, takedown through every layer | Staffing; errors both ways |
| 3.5 | Country rules | Append-only restrictions layer, checked at playback and every edge | Geo accuracy |
| 3.6 | −30% per view-hour | AV1 for popular, per-title, earned 4K, L4 GPUs, 6 s segments, ISP caches | Pipeline complexity |
R3.5 Global Architecture
VOD across three Regions
Synthesizing vector architecture diagram...
Each upload is processed in the Region nearest the uploader, its home. Popular videos are copied to a partner Region; restrictions and metadata are global tables.
- Home Region per video. The upload API sends each upload to the nearest healthy Region; that Region processes and stores it. Its paths start with the Region (
/use1/3f/v_Q2x9/...), and the CDN's cache behavior for each prefix points at that Region's bucket, with Origin Shield in the same Region. - Writes have one home. A video's metadata is written only in its home Region, so the global table's last-writer-wins rule decides nothing in normal operation. Restrictions are append-only. View counts are per Region: each Region's Flink writes its own counter item (
VIEWS#use1,VIEWS#euw1, …) with its ownlast_window, and a video's total is the sum of its items, so Regions never write the same item. - Popular videos get a second copy. When a video passes 1,000 views a day (about the top 5% of uploads), a job copies its renditions to a partner Region. The CDN's origin group for that prefix lists both, and fails over per request on errors or timeouts from the home bucket: no central automation.
Live
The diagram in step 3.1: two ingest points, two transcoding pipelines in two AZs, one packager origin with Origin Shield, three CDNs, and an archive into the VOD pipeline.
Delivery and policy
Synthesizing vector architecture diagram...
The playback API decides who may watch and where from; the CDNs enforce the same token, country check and blocklist; player measurements steer the next sessions.
Trace: the final starts
Synthesizing vector architecture diagram...
10M joins spread over 15 minutes is 10,000,000 ÷ 900 s ≈ 11,100 playback requests a second, on top of normal VOD traffic. The edges, not our APIs, carry the heavy load.
Trace: a copyrighted upload
Synthesizing vector architecture diagram...
Matching finishes alongside the fast lane, so the restriction exists before the video is public.
Trace: an urgent takedown
Synthesizing vector architecture diagram...
New sessions stop in seconds. Open sessions of a globally blocked video stop when the invalidation completes (usually seconds), because the origin no longer has the files; for any block, they stop within 10 minutes, when their cookies can't be renewed.
R3.6 Numbers and Cost
Live: the final
| Item | Math | Result |
|---|---|---|
| Live bandwidth | 10M × 4 Mbps average (we assume the mix of TVs and phones averages 4 Mbps on the ladder above) | 40 Tbps |
| Evening VOD peak after this round's savings | 130.2 Tbps × 0.727 | 94.7 Tbps |
| VOD on CDNs | 94.7 × (1 − 0.30 ISP caches) | 66.3 Tbps |
| Total on CDNs | 66.3 + 40 | 106.3 Tbps |
| Split | 50% CloudFront / 30% CDN B / 20% CDN C | 53.2 / 31.9 / 21.3 Tbps |
| Requests | 3M low-latency × 4/s + 7M standard × 1/s (we assume 30% choose low latency) | 19M a second |
| Bytes for 3 hours | 40 Tbps × 10,800 s ÷ 8 | 54 PB |
| Delivery cost | 54 PB at about $0.020/GB (the over-5 PB US and Europe rate, and we assume similar contracted rates at B and C) | ≈ $1.08M per final |
| Request cost | 19M × 10,800 s = 205 billion × $0.0124 per 10,000 | ≈ $0.25M per final |
| On IVS instead | 10M × 3 h = 30M viewer-hours × $0.096 (Full HD, largest North America tier) | ≈ $2.88M, before input charges |
Per viewer-hour, our CDN delivery at 4 Mbps is 1.8 GB × $0.020 = $0.036, against IVS's $0.096 to $0.144: about 2.7 to 4 times more. That's why creator live is built, not bought.
Live: creators
| Item | Math | Result |
|---|---|---|
| Streams at the peak | given | 50,000 |
| Transcoded (more than 50 viewers) | 20% (assumption) | 10,000 |
| GPUs at the peak | 10,000 ÷ 4 ladders per L4 (assumption) | 2,500 |
| Monthly | average half the peak: 1,250 × $0.8048 (g6.xlarge) × 730 h, on demand | ≈ $0.73M |
Fingerprinting and moderation: ≈ $97K and ≈ $9K a month (steps 3.3 and 3.4).
Storage per tier (end of year one, as in Round 2)
| Data | Tier | Round 2 | Round 3 |
|---|---|---|---|
| Raw, days 0–7 | S3 Standard | 45.4 PB, $0.95M | same |
| Raw, days 7–97 | Glacier Instant Retrieval | 583 PB, $2.33M | same (the source for earned 4K and AV1) |
| Raw, older | Glacier Deep Archive | 1,737 PB, $1.72M | same |
| Renditions, H.264/HEVC | Intelligent-Tiering | 2,978 PB, $24.5M | 2,978 × (1 − 0.595 × 0.9) ≈ 1,384 PB, ≈ $11.4M |
| Renditions, AV1 | Intelligent-Tiering | – | 5% of uploads × about 0.6 of a ladder's bytes ≈ 3% of 2,978 PB = 89 PB; popular, so in the frequent tier: 89 PB × $0.021 ≈ $1.87M |
| Popular copies in a partner Region | Intelligent-Tiering | – | Storage of the second copy: about 5% of rendition bytes ≈ 74 PB, popular so frequent tier: × $0.021 ≈ $1.55M. Replication: 5% × 8.16 PB a day = 408 TB × $0.02/GB ≈ $8.2K a day, ≈ $0.25M a month (transfer out of ap-south-1 costs more than $0.02/GB, so this is a floor) |
(0.595 is the 4K rung's share of the ladder's bytes: 15 ÷ 25.2 Mbps. 0.74 is its share of the GPU work, by pixels.)
VOD cost, before and after (monthly, list prices)
| Line | Round 2 | Round 3 without ISP caches | Round 3 with ISP caches |
|---|---|---|---|
| CDN data out | $539.0M | × 0.727 = $391.9M | 0.727 × (0.7 × $539.0M + 0.3 × 14.25 EB × $0.005/GB) = $289.9M |
| CDN requests | $29.2M | × 2/3 = $19.5M | × 2/3 × 0.7 = $13.6M |
| GPUs | $11.5M | $11.5M × 0.334 × 0.45 = $1.73M, + AV1 $11.5M × 0.45 × 0.25 = $1.29M: $3.0M | $3.0M |
| Raw storage | $5.0M | $5.0M | $5.0M |
| Rendition storage | $24.5M | $11.4M + AV1 $1.87M ≈ $13.3M | $13.3M |
| Transfer Acceleration | $1.97M | $0.39M | $0.39M |
| Everything else (Shield, S3 requests, orchestration, views, APIs) | 0.270 + 0.051 + 0.131 + 0.046 + 0.104 + 0.084 ≈ $0.69M | + fingerprinting $0.10M, moderation $0.01M, popular replication $0.25M and second-copy storage $1.55M, steering $0.05M, playlist edge function $0.10M ≈ $2.74M | $2.74M |
| Total | $611.9M | $435.9M (−28.8%) | $328.0M (−46.4%) |
| Per view-hour (12.67 billion) | $0.0483 | $0.0344 | $0.0259 |
Live is on top: about $0.73M a month for creator transcoding, plus about $1.3M of delivery and requests per final.
R3.7 Trade-Offs
Live latency vs stability
| Low latency (4.4 s) | Standard (10.4 s) | Real-time (under 1 s, WebRTC) | |
|---|---|---|---|
| Buffer against a bad network | 1.5 s | 6 s | Almost none |
| Requests per viewer a second | 4 | 1 | – (a persistent connection) |
| Caching | CDN-cacheable parts | CDN-cacheable segments | Not cacheable; servers relay every stream |
| Used for | Sports, auctions, anything with spoilers | Most creators | Guest calls on a stream, not the audience |
Multi-CDN vs single
| Single CDN | Three CDNs (chosen) | |
|---|---|---|
| A vendor outage | Global outage | Players switch in two failed fetches; caps keep the rest inside capacity |
| Price | One negotiation | Competition and commitments across three |
| Complexity | One token scheme, one rule set | Three implementations of tokens, country checks and blocklists, kept identical by tests |
| Cache efficiency | One cache for everything | Each CDN caches its share; misses consolidate through CloudFront's Origin Shield |
Fingerprint sensitivity
| Minimum match | Catches | Risk |
|---|---|---|
| 3 s | Short clips, samples | Many false positives: common sounds and music |
| 10 s (chosen default) | Most re-uploads | Short excerpts slip through |
| 30 s | Only long copies | Highlights clips slip through |
Rights holders can ask for shorter thresholds on specific references, after review.
Build vs buy is the table in step 3.6.
What changed from Round 1. Round 1 answered "how do we get a video from a camera to a screen?" with one managed job and one CDN. Round 3 answers "how do we do that for the world, live, under law and on budget?" with three Regions, three CDNs, a live path, and rules that sit beside the catalog. The core never changed: bytes go around our code, not through it; every file is cut into segments any cache can serve; and the player, not the server, adapts to the network.
R3.8 Failure Modes
| Trigger | What you'd see | How the design responds |
|---|---|---|
| A live encoder or ingest fails during the final | One pipeline's input or output stops | Two venue encoders feed two ingest points and two transcoding pipelines in two AZs, time-aligned. The packager switches to the other input; viewers see nothing, or at worst one short gap. Alarm on input loss on either pipeline, so we're back to two before a second failure. |
| A CDN has an outage | Errors and stalls on one CDN, in one country or everywhere | Players switch after two failed fetches, within seconds, without any central decision. The steering document drops the CDN's weight within a minute; if the others would overflow, it sets the bitrate caps (step 3.2). |
| A fingerprint false-positive storm | One reference suddenly claims thousands of uploads an hour | A circuit breaker per reference: if it matches far above its normal rate (say, 100× its hourly baseline), its block policy is switched to track until a human reviews the reference. Claims it made meanwhile are re-evaluated, not left standing. |
| A regional processing outage | Uploads to that Region fail; its videos' origin errors | Probes running in the other Regions detect it, and flip an Application Recovery Controller routing control (its data plane spans five Regions, so it doesn't depend on the failed one) that removes the Region from the upload API's choices. Uploads in progress there must restart in another Region (a multipart upload can't move between Regions). Its popular videos keep playing from their partner Region through per-request origin failover. Its long-tail videos play only from cache until it returns. Jobs in flight wait and resume. |
| Restriction propagation stalls | A blocked video still plays through one CDN | The 10-minute reconciliation repairs the blocklist; the canary alarms first. The playback API still refuses new sessions and cookie renewals, so the exposure is limited to open sessions for at most the 10-minute cookie life. |
| A takedown is wrong | A lawful video is blocked | The creator appeals; a reviewer issues a lift decision naming the block; the same stream removes it from every blocklist. Nothing was deleted in the meantime. |
R3.9 Runbook and Incident Response
| Signal | Alarm | Severity | First action |
|---|---|---|---|
| Time to first frame, P95, by CDN and country (player-reported) | > 800 ms for 10 min | P2 | Hit ratio and errors on that CDN; steering weights |
| Rebuffering, share of sessions, by CDN | > 0.5% for 10 min; > 1.5% | P2 / P1 | CDN failover procedure below |
| Edge hit ratio | < 98% for 15 min | P2 | New cache-key parameters? A purge? Origin Shield health |
| Error rate by CDN | 5xx > 1% for 5 min | P1 | CDN failover procedure |
| Transcode queue age, fast lane / full lane | oldest > 60 s / > 2 h | P1 / P2 | Stuck-queue procedure below |
| Live latency (player-reported) and ingest health | > 8 s in low-latency mode; any pipeline input loss | P1 | Check both pipelines; the packager's active input |
| Blocklist canary | a known-blocked video plays anywhere | P1 | Force reconciliation; check the enforcement consumer |
| Fingerprint matches per reference | > 100× hourly baseline | P2 | The circuit breaker should have tripped; review the reference |
CDN failover procedure REL 5
- Confirm it's the CDN: errors on one CDN only, across many videos, while the others are healthy.
- Players are already switching. Check that the other CDNs' error rates stay flat as they absorb traffic.
- Set the failing CDN's weight to 0 in the steering document and deploy it. Sessions starting from now avoid it.
- If the others are above 80% of their reserved capacity, set the bitrate caps (720p for live, above-720p VOD capped). Remove the caps when load falls below 60%.
- Restore the CDN's weight in steps of 10% an hour once its own status and our canaries are clean.
Stuck transcode queue procedure REL 7
- Is it the fast lane? That's P1: videos aren't becoming watchable.
- Is the fleet running? Compare instances in service with desired capacity. If Spot capacity is short, raise the on-demand share of the fast-lane group.
- Are tasks failing? Check the dead-letter queue: one bad input can't block others, but a bad encoder release can fail everything. Roll back the release, then move the dead-letter messages back.
- Are workflows waiting on callbacks that never came? List running executions older than an hour; for each, read the job item. If all chunks are recorded, re-send the callback (the same rule as step 2.2).
Go deeper: CLI playbook
Plain commands an on-call engineer runs one at a time. Replace names, IDs and ARNs with real ones.
text# 1. Oldest message in the fast lane (a CloudWatch metric, not a queue attribute) aws cloudwatch get-metric-statistics --region us-east-1 --namespace AWS/SQS --metric-name ApproximateAgeOfOldestMessage --dimensions Name=QueueName,Value=transcode-fast --start-time 2026-10-05T09:00:00Z --end-time 2026-10-05T09:30:00Z --period 60 --statistics Maximum # 2. Messages waiting and in flight aws sqs get-queue-attributes --region us-east-1 --queue-url https://sqs.us-east-1.amazonaws.com/123456789012/transcode-fast --attribute-names ApproximateNumberOfMessages ApproximateNumberOfMessagesNotVisible # 3. Fast-lane fleet and an emergency scale-up aws autoscaling describe-auto-scaling-groups --region us-east-1 --auto-scaling-group-names gpu-fast-lane aws autoscaling set-desired-capacity --region us-east-1 --auto-scaling-group-name gpu-fast-lane --desired-capacity 6000 # 4. Move dead-letter messages back to their source queue after a fix aws sqs start-message-move-task --region us-east-1 --source-arn arn:aws:sqs:us-east-1:123456789012:transcode-fast-dlq # 5. Workflows still running aws stepfunctions list-executions --region us-east-1 --state-machine-arn arn:aws:states:us-east-1:123456789012:stateMachine:vod-pipeline --status-filter RUNNING --max-items 50 # 6. CloudFront error rate (CloudFront metrics live in us-east-1 with Region=Global) aws cloudwatch get-metric-statistics --region us-east-1 --namespace AWS/CloudFront --metric-name 5xxErrorRate --dimensions Name=DistributionId,Value=E2ABC123EXAMPLE Name=Region,Value=Global --start-time 2026-10-05T09:00:00Z --end-time 2026-10-05T09:30:00Z --period 60 --statistics Average # 7. Invalidate everything tagged with a video, then track it aws cloudfront create-invalidation --distribution-id E2ABC123EXAMPLE --paths "#vid:v_Q2x9" aws cloudfront get-invalidation --distribution-id E2ABC123EXAMPLE --id I2J0I21PCUYOIK # 8. Add a video to the CloudFront emergency blocklist (needs the store's current ETag) aws cloudfront-keyvaluestore describe-key-value-store --kvs-arn arn:aws:cloudfront::123456789012:key-value-store/0f1e2d3c-4b5a-6978-8a9b-0c1d2e3f4a5b aws cloudfront-keyvaluestore put-key --kvs-arn arn:aws:cloudfront::123456789012:key-value-store/0f1e2d3c-4b5a-6978-8a9b-0c1d2e3f4a5b --key v_Q2x9 --value block --if-match ETVPDKIKX0DER # 9. Live channel state for the final aws medialive describe-channel --region eu-west-1 --channel-id 1234567 # 10. Deploy the steering document (weights, caps) aws appconfig start-deployment --region us-east-1 --application-id abc1234 --environment-id def5678 --deployment-strategy-id AppConfig.AllAtOnce --configuration-profile-id ghi9012 --configuration-version 43
R3.10 Pillar Check
| Pillar | What Round 3 adds |
|---|---|
| Reliability | Redundant live paths; player-side CDN failover; three processing Regions with failover driven from outside the failed one; per-request origin failover for popular videos REL 10 · REL 13 |
| Performance Efficiency | LL-HLS at about 4.4 s glass to glass; steering by measured time to first frame; AV1 and per-title ladders PERF 1 · PERF 4 |
| Security | An append-only restrictions layer that no other writer touches; the same token, country and blocklist checks at every CDN; a takedown path through every cache layer; deletion claims that include backup windows SEC 10 |
| Cost Optimization | Cost per view-hour from $0.048 to $0.034 without ISP caches and $0.026 with them; build vs buy per workload; IVS priced against our own delivery COST 5 · COST 6 · COST 8 |
| Operational Excellence | Golden signals by CDN and country; CDN-failover and stuck-queue procedures; canaries for the blocklist; event rehearsals OPS 8 · OPS 10 |
| Sustainability | Expensive encodes only for videos people watch; 27% fewer bytes per view; newer GPUs doing more per watt; cold data in archive tiers SUS 3 · SUS 5 |
R3.11 Round 3 Rubric and Follow-Ups
What an architect (L7) answer adds over L6
- Builds live as a continuous pipeline, adds up the glass-to-glass latency, and offers modes instead of one answer.
- Treats CDNs as interchangeable suppliers: steering from real measurements, player-side failover, keep-alive steering documents, and identical rules at every one.
- Proves the design survives losing its largest CDN during the biggest event, with numbers.
- Knows why hashes don't find copies, and makes matching robust and reversible.
- Puts availability rules in a separate, append-only layer that survives every update, restore and failover, and takes it down through every cache.
- Is honest about what "deleted" means while backups exist.
- Cuts cost with levers whose assumptions are labeled, finds that software alone lands at −29%, and knows what closes the gap.
Follow-up questions
-
"Why not make every stream low latency?" Answer: 4 requests a second per viewer instead of 1, and a 1.5 s buffer instead of 6 s, so more stalls on weak networks. For the final, with 30% in low latency, that's 19M requests a second instead of 10M. Low latency is worth it where spoilers matter, not for a cooking stream.
-
"A country orders a video blocked, and a week later the creator re-uploads it as a new video." Answer: the restriction is on the old video ID, so the new one isn't blocked by it. The legal team can ask for the ruling to cover copies; then we add the blocked video's fingerprints as a reference with a
blockpolicy for that country, and the matcher catches re-uploads before they publish. -
"Why do other CDNs pull through CloudFront's Origin Shield? Isn't that a dependency on one vendor?" Answer: it is, and we accept it for the common case because it collapses origin requests across all CDNs into one. The other CDNs keep a direct origin path, behind a secret header, for when CloudFront is the problem.
Interview gotchas from this round's wrong answers
| Gotcha | Why it's wrong |
|---|---|
| "Upload a file every minute for live" | Minutes of latency and a seam every minute. |
| "Our CDN is big enough" | One vendor's bad hour becomes everyone's; no price leverage. |
| "Hash the files to find copies" | Any re-encode changes every byte. |
| "Delete the file to take it down" | Caches, open sessions and other CDNs still serve it. |
"A blocked_countries field on the video" | The next update or restore can overwrite it. |
| "Deleted means gone" | Not until backups and old versions expire. |
| "Lower quality to cut cost" | View-hours are the product; cut bytes per quality instead. |
Loop Closer: Interview Strategy for All Three Rounds
How to Run Each 60-Minute Round
| Time | Round 1 | Round 2 | Round 3 |
|---|---|---|---|
| 0–5 min | Scoping: length, devices, private or public, how fast ready | Restate Round 1 in 60 seconds | Restate Round 2 in 60 seconds |
| 5–15 min | Requirements and API (pre-signed uploads, status, playback cookies) | Scope raise → what breaks | Scope raise → what breaks |
| 15–40 min | Steps 1.0–1.6: proxy → direct upload → ladder → HLS and ABR → CDN → state machine → signed cookies | Steps 2.1–2.7: chunked DAG → Spot fleet → Origin Shield → egress math → view counts → storage tiers → DRM | Steps 3.1–3.6: live → multi-CDN → fingerprints → moderation → country rules → −30% |
| 40–50 min | Start-time budget, bytes, cost (delivery 70%) | Fleet sizing, egress, the $612M table, trade-offs | Live and CDN math, cost per view-hour before and after, trade-offs |
| 50–60 min | Failures and pillar check | Failures, gotchas, pillar check | Failures, runbook, pillar check |
For how to spend a single 45-minute round, see the 45-minute interview blueprint. For blob storage internals, see the S3 object storage loop; for live chat beside a stream, the chat loop.
The Two Sentences That Matter Most
- Opening any round: "Video is a bytes-and-money problem: I'll keep every byte out of our servers, turn each upload into a ladder of segmented renditions the player can switch between, and let CDNs serve the bytes, because delivery will be most of the bill."
- When scale arrives: "I'll transcode in parallel keyframe-aligned chunks on Spot GPUs with idempotent tasks, protect the origin with a shield, count views through a stream with conditional writes, and then attack cost per view-hour where it actually lives: bytes per view and who carries them."
Well-Architected Review Sheet
Interviewers rarely ask "which pillar is this?". They ask the pillar's question in plain words. Rehearse one sentence per row.
| Pillar | Question you'll hear | One-sentence answer | Round | Backed by |
|---|---|---|---|---|
| Reliability | "What if a transcoding worker dies mid-task?" (REL 11) | Its SQS lease expires within 120 s and another worker redoes the task; outputs are idempotent and completion is a conditional write. | 2 | Step 2.2 |
| "What if a Region fails?" (REL 13) | Probes in other Regions flip a routing control; uploads go elsewhere; popular videos play from their partner Region per request. | 3 | R3.5, R3.8 | |
| "What if your CDN fails during the final?" (REL 10) | Players switch CDNs after two failed fetches; bitrate caps keep the other two inside capacity. | 3 | Step 3.2 | |
| Performance | "How fast does playback start?" (PERF 4) | About 750 ms P95 on a cache hit: start low, prefetch, edge hits. | 1–2 | R1.7, R2.6 |
| "How long until an upload is watchable?" (PERF 2) | 720p in about 39 s for a 10-minute video, because chunks run in parallel and the fast lane goes first. | 2 | Step 2.1 | |
| "How live is live?" (PERF 1) | 4.4 s in low-latency mode, 10.4 s standard, with every delay counted. | 3 | Step 3.1 | |
| Security | "How do you keep private videos private?" (SEC 3) | Short-lived signed cookies over a private origin; DRM when authorized viewers mustn't keep copies. | 1–2 | Steps 1.6, 2.7 |
| "How do you block a video in one country?" (SEC 10) | An append-only restrictions layer checked at playback and at every CDN's edge, which nothing else writes. | 3 | Step 3.5 | |
| Cost | "Where does the money go?" (COST 8) | Delivery: 70% in Round 1, 93% at scale. | 1–2 | R1.7, R2.6 |
| "Managed or your own transcoding?" (COST 5) | Managed at 500 videos a day; our own Spot GPU fleet at 4.32M, where MediaConvert would be $37M+ a month. | 1–2 | R1.8, R2.7 | |
| "How do you cut cost per view-hour 30%?" (COST 6) | Earned 4K and AV1, per-title ladders, newer GPUs, longer segments: −29%; ISP caches or better contracts close the gap. | 3 | Step 3.6 | |
| Operations | "How do you know playback is healthy?" (OPS 8) | Player-reported first-frame time and rebuffering by CDN and country, plus queue age per lane. | 2–3 | R3.9 |
| Sustainability | "How do you avoid waste?" (SUS 3) | Expensive encodes only for watched videos, fewer bytes per view, and cold data in archive tiers. | 2–3 | Steps 2.6, 3.6 |
Rubric Across Levels
| Dimension | L5 (Round 1) | L6 (Round 2) | L7 (Round 3) |
|---|---|---|---|
| Upload | Pre-signed multipart, straight to S3 | Resumable sessions, adaptive part size, checksums | Regional upload buckets, Region failover from outside |
| Transcoding | Managed job, a small H.264 ladder | Chunked DAG, fast and full lanes, Spot fleet sized from rendition-hours | Earned 4K, AV1 for popular videos, per-title ladders, live |
| Delivery | HLS, ABR, CloudFront, signed cookies | Origin Shield, versioned paths, egress math, DRM | Three CDNs steered per session, ISP caches, survive losing one |
| Correctness | Conditional status transitions; re-sent notifications | Idempotent tasks, re-sent callbacks, windowed conditional counts | Append-only overrides that survive every update, restore and fallback |
| Cost | Finds delivery is 70% | Prices every line; delivery 93% | Cost per view-hour −29% to −46% with labeled levers |
| Evolving under new scope | Builds from one proxied file | Opens with the byte and fleet math | Adds live and rules without breaking the VOD core |