Design S3-Like Distributed Object Storage
1. Problem Statement & Scope Clarification
System Mission
Design a planetary-scale, multi-tenant distributed Object Storage Service (equivalent to Amazon Simple Storage Service / AWS S3) engineered to ingest, persist, and serve exabytes of unstructured binary objects (images, videos, documents, database backups, analytical datasets) with (11 9s) annual durability, sub- read latency, and strong read-after-write consistency for PUT and DELETE operations.
Functional Requirements
- Object CRUD Operations: Clients upload (
PUT), read (GET), delete (DELETE), and inspect headers (HEAD) for immutable objects identified by a unique(bucket_name, object_key)coordinate. - Multipart Parallel Upload: Objects larger than (and up to ) can be split into arbitrary chunks ( per part) uploaded concurrently in random order and assembled atomically.
- Prefix-Based Listing & Delimiters: Hierarchical folder simulation allowing clients to list objects under a given bucket prefix (e.g.,
GET /photos/2026/09/?delimiter=/) with cursor-based pagination. - Presigned URLs & Fine-Grained Access Control: Generate cryptographically signed, time-bounded URLs allowing untrusted web/mobile clients to stream direct uploads and downloads without passing through application backend servers.
- Intelligent Lifecycle Tiering & Versioning: Automatic object versioning (
version_id) and background policy transition from High-Throughput NVMe/EBS Standard storage to lower-cost Infrequent Access (S3-IA) and archival cold tiers (S3 Glacier / Glacier Deep Archive).
Non-Functional Requirements (SLAs & SLOs)
- Durability: (11 9s) annual durability per object via Reed-Solomon Erasure Coding ( or ) striped across independent failure domains (racks and availability zones).
- Availability: ("four nines") uptime SLA for the data and metadata planes across multi-AZ deployments.
- Latency (P99):
- Metadata Operations (
HEAD, SmallGET): . - Streaming Ingress (
PUTchunks): Time-to-First-Byte (TTFB) , sustaining link saturation. - Streaming Egress (
GETstreaming): TTFB , continuous line-rate throughput.
- Metadata Operations (
- Consistency Model: Strong Read-After-Write Consistency for
PUTandDELETErequests of both new and overwritten objects. Any subsequent read instantly reflects the latest committed byte stream and metadata. - Throughput & Scale: Support total objects, initial raw storage capacity (scaling to Exabytes), and peak ingress/egress network bandwidth exceeding .
2. Capacity & Scale Estimation (Back-of-the-Envelope Math)
Ingestion Scale & Traffic Projections
- Total Objects Stored: ().
- Average Object Size:
- Small objects / thumbnails ( of count): .
- Medium documents / images ( of count): .
- Large media files / backups ( of count): .
- Weighted Average Object Size:
- Total Raw Data Footprint: (For a baseline initial cluster footprint, we model a active tier with live objects).
Request Throughput (QPS)
- Read/Write Ratio: read-heavy workload.
- Average Write QPS:
- Peak Write QPS ( peak):
- Average Read QPS:
- Peak Read QPS ( peak):
Network Bandwidth & Egress Sizing
- Assuming average streaming read size is for online traffic (larger payloads fetched via byte-range streams):
- Peak Ingress Bandwidth:
Erasure Coding Storage & Rebuild Bandwidth Mathematics
To deliver durability without the cost penalty of triple replication, we deploy Reed-Solomon Erasure Coding ( data chunks, parity chunks):
- Storage Overhead:
- For a raw active working set:
- Durability Guarantee: The system can sustain the complete, simultaneous destruction of any 4 storage nodes/racks/AZs without a single bit of data loss.
- Reconstruction Bandwidth Sizing: When a storage node holding a NVMe drive fails, the cluster must read surviving chunks to reconstruct the 1 lost chunk for every stripe: To rebuild within an MTTR (Mean Time to Repair) window of (): AWS multi-chassis VPC network fabrics easily accommodate this rebuild overhead without degrading client traffic.
Why Reed-Solomon Across 3 Availability Zones? Slicing an object into 8 data chunks and 4 parity chunks yields a storage overhead factor of only (compared to for 3-way replication). Striping exactly 4 chunks into each of 3 distinct AWS Availability Zones ensures the cluster can withstand the total, catastrophic loss of an entire Availability Zone (4 chunks) with zero data loss and uninterrupted read reconstruction using the surviving 8 chunks.
Metadata Footprint Estimation
- Metadata Record Size: Bucket name (), object key (), version (), size (), ETag/MD5 (), timestamp (), storage class (), chunk manifest pointers () .
- With B-Tree and primary key index overhead .
- For : This dataset is distributed across an Amazon Aurora PostgreSQL Multi-Master / Multi-AZ Cluster combined with an Amazon DynamoDB Single-Table Store.
3. AWS-First High-Level Architecture
The architecture decouples the Control & Metadata Plane (handling authorization, key indexes, bucket ACLs, and multipart manifests) from the Data Plane (responsible for chunking, erasure coding, streaming I/O, and raw block storage on NVMe clusters).
Synthesizing vector architecture diagram...
Control Plane vs. Data Plane Decoupling
By completely isolating metadata operations (bucket ACLs, version catalogues, chunk location tables in Aurora/Redis) from raw binary payload streaming (NVMe data nodes), the system sustains line-rate multi-gigabit throughput while guaranteeing single-digit millisecond latency for HEAD and prefix listing queries.
Data Flow Walkthrough
- Request Ingress & Signature Validation: Requests arrive via Route 53 Anycast DNS and NLB. The API Gateway validates the request using AWS Signature Version 4 (SigV4) against an in-memory AWS IAM/KMS caching layer.
- Metadata Authorization & Resolution:
- For
PUT: The gateway queries the Metadata Plane to verify bucket existence, write permissions, and bucket quota. - For
GET: The gateway queries Aurora/Redis to retrieve the object's chunk manifest, version ID, and the physical network addresses of the 12 storage nodes holding the stripes.
- For
- Data Streaming & Erasure Coding:
- On
PUT: The client payload streams directly into the Data Router. As data arrives, the Erasure Coding Engine slices the stream into equal data chunks, computes parity chunks via Galois Field SIMD instructions, and concurrently writes the 12 chunks to 12 distinct storage nodes distributed across 3 Availability Zones. - Once all 8 data chunks and at least 2 parity chunks acknowledge persistence (a durable quorum of ), the Metadata Plane updates the object status to
ACTIVE, and an HTTP 200/201 with the object ETag is returned to the client.
- On
- Zero-Copy Streaming Read:
- On
GET: The gateway connects concurrently to the first 8 fastest responding storage nodes holding data chunks. Chunks stream directly to the client over an HTTP/2 response stream. If 1 or 2 nodes are slow or offline, parity chunks are read in parallel and decoded on the fly with zero client-perceived disruption.
- On
Concrete Step-by-Step Request Walkthrough: Tracing an Object Upload (PUT)
| Step # | Event / Action | Component State | Distributed Transition | Output / Response |
|---|---|---|---|---|
| 1 | Client issues PUT /media/intro.mp4( payload with SigV4 header) | Gateway validates HMAC-SHA256 signature; checks bucket ACL in Redis cache | Gateway generates unique 128-bit object_id;acquires metadata staging lock | Gateway ready for body stream; HTTP 100 Continue emitted if requested |
| 2 | Gateway streams payload into Erasure Coding buffer | Slices into 8 Data Chunks ( each: ) | Vandermonde matrix multiplies over to generate 4 Parity Chunks () | 12 Chunks ready in memory ( total, each) |
| 3 | Parallel storage write across 3 Availability Zones | Storage Router selects 12 nodes across AZ-a, AZ-b, AZ-c | Issues concurrent HTTP/2 chunk writes with SHA-256 chunk trailers | Nodes persist chunks to NVMe append-only block containers ( container_42.blob) |
| 4 | Quorum confirmation | 12 nodes acknowledge receipt; disk fdatasync verified | Requires at least durable node confirmations | Chunk manifest constructed with node UUIDs and byte offsets |
| 5 | Atomic Metadata Commit | Aurora PostgreSQL transaction commitsobject_metadata and chunk_manifest | Atomically transitions state from PENDINGto COMMITTED; updates GSI index | Commits in ; Redis cache invalidated |
| 6 | Client Completion Acknowledgment | Gateway closes client connection stream | Emits HTTP 200 OK with MD5/SHA256 ETag and version tracking token | Client receives ETag: "9b10e43f0563..."and x-amz-version-id: "v-88319" |
4. API Interface Design & Wire Protocol
The service exposes an S3-compatible REST API over HTTP/1.1 and HTTP/2 with strict AWS Signature Version 4 authentication.
1. Object Upload (PUT)
httpPUT /my-media-bucket/videos/2026/presentation.mp4 HTTP/1.1 Host: s3.us-east-1.amazonaws.com Authorization: AWS4-HMAC-SHA256 Credential=AKIAIOSFODNN7EXAMPLE/20260916/us-east-1/s3/aws4_request, SignedHeaders=content-length;content-type;host;x-amz-content-sha256;x-amz-date, Signature=fe5f80f779dad5328e... Content-Type: video/mp4 Content-Length: 67108864 x-amz-date: 20260916T120000Z x-amz-storage-class: STANDARD x-amz-checksum-sha256: 47DEQpj8HBSa+/TImW+5JCeuQeRkm5NMpJWZG3hSuFU= <binary video payload stream (64 MB)>
Response:
httpHTTP/1.1 200 OK x-amz-id-2: LriByRT6OQIqZ0yKq3h8dQe9n1F2c... x-amz-request-id: 4A41C16489469FB2 Date: Wed, 16 Sep 2026 12:00:01 GMT ETag: "c8b417e8ef28258b538053c9e99a8123" x-amz-version-id: "9mQkPzY_1n0J12Vw" x-amz-server-side-encryption: AES256 Content-Length: 0
2. Initiate Multipart Upload (POST ?uploads)
httpPOST /my-media-bucket/backups/database_dump.tar.gz?uploads HTTP/1.1 Host: s3.us-east-1.amazonaws.com Authorization: AWS4-HMAC-SHA256 Credential=... Content-Type: application/gzip x-amz-date: 20260916T120500Z Response: 200 OK Content-Type: application/xml <?xml version="1.0" encoding="UTF-8"?> <InitiateMultipartUploadResult xmlns="http://s3.amazonaws.com/doc/2006-03-01/"> <Bucket>my-media-bucket</Bucket> <Key>backups/database_dump.tar.gz</Key> <UploadId>upload_id_7a8b9c0d1e2f3a4b</UploadId> </InitiateMultipartUploadResult>
3. Upload Part (PUT ?partNumber=N&uploadId=X)
httpPUT /my-media-bucket/backups/database_dump.tar.gz?partNumber=1&uploadId=upload_id_7a8b9c0d1e2f3a4b HTTP/1.1 Host: s3.us-east-1.amazonaws.com Content-Length: 104857600 Content-MD5: Q2hlY2sgSW50ZWdyaXR5IQ== <binary part 1 payload (100 MB)> Response: 200 OK ETag: "b10a8db164e0754105b7a99be72e3fe5"
4. Complete Multipart Upload (POST ?uploadId=X)
httpPOST /my-media-bucket/backups/database_dump.tar.gz?uploadId=upload_id_7a8b9c0d1e2f3a4b HTTP/1.1 Host: s3.us-east-1.amazonaws.com Content-Type: application/xml <CompleteMultipartUpload> <Part> <PartNumber>1</PartNumber> <ETag>"b10a8db164e0754105b7a99be72e3fe5"</ETag> </Part> <Part> <PartNumber>2</PartNumber> <ETag>"1b2cf535f27731c974343645a3985328"</ETag> </Part> </CompleteMultipartUpload> Response: 200 OK <CompleteMultipartUploadResult> <Location>https://my-media-bucket.s3.amazonaws.com/backups/database_dump.tar.gz</Location> <Bucket>my-media-bucket</Bucket> <Key>backups/database_dump.tar.gz</Key> <ETag>"3a72810f92b0e9d6b2c8a14e1011e40b-2"</ETag> <VersionId>4vJk90Wqz18LkmQ</VersionId> </CompleteMultipartUploadResult>
5. Data Models & Storage Architecture
Relational Schema: Amazon Aurora PostgreSQL (Control & Object Catalog)
The metadata tier manages buckets, object hierarchies, and multipart tracking.
sql-- 1. Buckets Table CREATE TABLE storage_buckets ( bucket_id UUID PRIMARY KEY DEFAULT gen_random_uuid(), bucket_name VARCHAR(63) UNIQUE NOT NULL, owner_account_id VARCHAR(32) NOT NULL, region VARCHAR(32) NOT NULL, versioning_status VARCHAR(16) DEFAULT 'DISABLED' CHECK (versioning_status IN ('DISABLED', 'ENABLED', 'SUSPENDED')), encryption_algorithm VARCHAR(32) DEFAULT 'AES256', kms_master_key_id VARCHAR(128), created_at TIMESTAMPTZ DEFAULT NOW(), storage_quota_bytes BIGINT DEFAULT 109951162777600 -- 100 TB Default ); CREATE INDEX idx_buckets_owner ON storage_buckets(owner_account_id); -- 2. Object Metadata Table (Strong Read-After-Write Consistency Catalog) CREATE TABLE object_metadata ( object_uuid UUID DEFAULT gen_random_uuid(), bucket_id UUID NOT NULL REFERENCES storage_buckets(bucket_id) ON DELETE CASCADE, object_key VARCHAR(1024) NOT NULL, version_id VARCHAR(64) NOT NULL, is_latest BOOLEAN NOT NULL DEFAULT TRUE, is_delete_marker BOOLEAN NOT NULL DEFAULT FALSE, object_size_bytes BIGINT NOT NULL, etag VARCHAR(64) NOT NULL, content_type VARCHAR(128) DEFAULT 'application/octet-stream', storage_class VARCHAR(32) DEFAULT 'STANDARD' CHECK (storage_class IN ('STANDARD', 'INTELLIGENT_TIERING', 'STANDARD_IA', 'GLACIER', 'DEEP_ARCHIVE')), lifecycle_expiry_ts TIMESTAMPTZ, created_at TIMESTAMPTZ DEFAULT NOW(), manifest_id UUID NOT NULL, PRIMARY KEY (bucket_id, object_key, version_id) ); -- Fast range scanning for hierarchical listing with delimiters (e.g. photos/2026/%) CREATE INDEX idx_objects_prefix_listing ON object_metadata (bucket_id, object_key text_pattern_ops, is_latest) WHERE is_delete_marker = FALSE; -- 3. Multipart Upload Sessions Table CREATE TABLE multipart_uploads ( upload_id VARCHAR(64) PRIMARY KEY, bucket_id UUID NOT NULL REFERENCES storage_buckets(bucket_id), object_key VARCHAR(1024) NOT NULL, storage_class VARCHAR(32) DEFAULT 'STANDARD', initiator_identity VARCHAR(64) NOT NULL, status VARCHAR(16) DEFAULT 'INITIATED' CHECK (status IN ('INITIATED', 'COMPLETING', 'ABORTED', 'COMPLETED')), created_at TIMESTAMPTZ DEFAULT NOW(), abort_rule_deadline TIMESTAMPTZ NOT NULL DEFAULT (NOW() + INTERVAL '7 days') ); -- 4. Multipart Parts Table CREATE TABLE multipart_parts ( upload_id VARCHAR(64) NOT NULL REFERENCES multipart_uploads(upload_id) ON DELETE CASCADE, part_number INT NOT NULL CHECK (part_number BETWEEN 1 AND 10000), part_size_bytes BIGINT NOT NULL, etag VARCHAR(64) NOT NULL, manifest_id UUID NOT NULL, uploaded_at TIMESTAMPTZ DEFAULT NOW(), PRIMARY KEY (upload_id, part_number) );
DynamoDB Single-Table Design: BlobChunkManifestTable
While Aurora handles SQL prefix listing and ACID operations on buckets, raw chunk placements are persisted in DynamoDB for ultra-low latency () key-value lookups by the data router during streaming:
PK (Partition Key) | SK (Sort Key) | Attributes | Description |
|---|---|---|---|
MANIFEST#<manifest_id> | STRIPE#0 | k_data: 8, m_parity: 4, stripe_size_bytes: 67108864, chunk_size_bytes: 8388608 | Stripe parameters |
MANIFEST#<manifest_id> | CHUNK#01 | node_id: "sn-az1-04", container_file: "blk_9921.blob", offset: 4194304, size: 8388608, sha256: "8f1a..." | Data chunk 1 |
MANIFEST#<manifest_id> | CHUNK#09 | node_id: "sn-az3-01", container_file: "blk_1042.blob", offset: 1048576, size: 8388608, sha256: "2d4c..." | Parity chunk 1 |
Unlock Complete Architecture & Production Runbooks
You have explored the free architectural preview (~40%). Spend 1 Coin to unlock the remaining 6 production deep-dive sections for a full 24 hours.