Skip to main content
BLUEPRINT #04Location & Geospatial

Design a Real-Time Ride-Sharing Dispatch Service (Uber/Lyft)

Target AWS Architecture:DynamoDBS3AuroraMemoryDB
10-Stage Structure:1. Requirements→2. Sizing→3. Topology→4. Data Model→5. AWS Topology→6. Deep-Dive→7. Failures→8. SRE Playbooks

1. Problem Statement & Scope

System Mission

Design a real-time, highly available, low-latency ride-matching and dispatch platform (similar to Uber or Lyft) capable of tracking millions of active drivers sending continuous GPS coordinates, matching rider requests with the nearest available drivers in under 2 seconds, calculating dynamic surge pricing, and handling atomic ride acceptances under high concurrency.

Functional Requirements

  1. Real-Time Driver Location Ingestion: Ingest live GPS coordinates (lat, lng, bearing, status) from 1 Million active drivers every 4 seconds.
  2. Ride Request & Driver Discovery: Riders request a ride specifying pickup and dropoff coordinates; system discovers top NN nearby available drivers within a 5Β km5\text{ km} radius in <500Β ms< 500\text{ ms}.
  3. Atomic Ride Dispatch & Match Acceptance: Dispatch ride offers sequentially or in batched rings to drivers; ensure exactly one driver accepts the ride without race conditions.
  4. Dynamic Surge Pricing: Calculate real-time supply vs. demand multiplier per spatial grid cell (e.g. 1.5Γ—1.5\times surge).
  5. Trip State Machine: Track trip lifecycle states (REQUESTED, MATCHED, ARRIVING, IN_TRIP, COMPLETED, CANCELLED).

Non-Functional Requirements (SLAs/SLOs)

  • High Availability: 99.999%99.999\% uptime for location ingestion and dispatch pipelines.
  • Ultra-Low Latency: Location Ingestion Latency <50Β ms< 50\text{ ms}; Driver Matching Latency <1Β s< 1\text{ s}.
  • Consistency: Strong consistency on driver acceptance (atomic lock) to prevent double-booking.
  • Scalability: Support 1M1\text{M} active drivers and 10M10\text{M} daily completed trips.

2. Capacity & Scale Estimation

Traffic Calculations

  • Active Driver Count: 1 Million active drivers (10610^6).
  • GPS Ingestion Rate: Each driver emits a location ping every 4 seconds (0.25Β updates/sec/driver0.25\text{ updates/sec/driver}).
  • Global Ingestion : LocationΒ WriteΒ QPS=1,000,0004=250,000Β QPS\text{Location Write QPS} = \frac{1,000,000}{4} = \mathbf{250,000\text{ QPS}}
    • Peak Ingestion (2Γ—2\times rush hour): 500,000Β QPS\mathbf{500,000\text{ QPS}}.
  • Rider Demand & Search :
    • 10Β Million10\text{ Million} ride requests per day.
    • Active rider search : β‰ˆ2,000Β searchΒ QPS\approx 2,000\text{ search QPS} (Peak 5,000Β QPS5,000\text{ QPS}).

Storage & Bandwidth Estimation

  • Location Payload Size:
    • driver_id (16 Bytes), lat (8 Bytes), lng (8 Bytes), bearing (4 Bytes), timestamp (8 Bytes), status (4 Bytes) β‰ˆ48Β Bytes\approx \mathbf{48\text{ Bytes}} (with JSON/Protobuf envelope β‰ˆ100Β Bytes\approx 100\text{ Bytes}).
  • Ingestion Network Bandwidth: 250,000Β writes/secΓ—100Β Bytes=25Β MB/s=200Β Mbps250,000\text{ writes/sec} \times 100\text{ Bytes} = 25\text{ MB/s} = \mathbf{200\text{ Mbps}}
  • Active Driver Ephemeral Location Cache ( Geospatial / In-Memory MemoryDB): 1,000,000Β driversΓ—200Β Bytesβ‰ˆ200Β MBΒ RAM1,000,000\text{ drivers} \times 200\text{ Bytes} \approx \mathbf{200\text{ MB RAM}} (Easily stored in memory with sub-millisecond query latency).
  • Historical Location Audit Log (30-Day Retention in ): 25Β MB/sΓ—86,400Β s/dayΓ—30Β daysβ‰ˆ64.8Β TB/month25\text{ MB/s} \times 86,400\text{ s/day} \times 30\text{ days} \approx \mathbf{64.8\text{ TB/month}}

3. AWS-First High-Level Architecture

Interactive Architecture Diagram
Synthesizing vector architecture diagram...

Part 2: Production Deep-Dive Locked1 Coin = 24 Hours

Unlock Complete Architecture & Production Runbooks

Your Balance:40 Coins

You have explored the free architectural preview (~42%). Spend 1 Coin to unlock the remaining 7 production deep-dive sections for a full 24 hours.

Sections Included in This 24-Hour Pass:
4. API Interface Design
5. Data Models & Storage Architecture
6. Component Deep Dives & Workflows
7. Architectural Trade-Off Matrix & Primitive Links
8. Critical Edge Cases & Distributed Failure Modes
9. Production Pitfalls & Anti-Patterns (The "Gotchas")
10. Production Runbook & Operational Best Practices
Keeps page unlocked for exactly 24 hoursSpend coins to fund LLM & compute infrastructure