System design explainer

Design Uber

The interview question behind every ride-hail: a rider taps a button and a driver appears. How do you match riders to drivers in seconds, price it fairly when demand spikes, and track it all live?

The takeaway, up front

Uber is a real-time geospatial matching problem. The core loop: drivers stream their location into a geo index every few seconds; rider requests are collected for 2–4 seconds and matched as a batch assignment problem; price is a control loop that balances local supply against demand. Nail those three — geo index, batch matching, surge control — and the rest is plumbing.

1 · Requirements

What are we actually building?

The analogy: Uber is a stock exchange where the goods are car seats and the traders move. Riders post bids (ride requests), drivers post asks (availability), and the exchange clears the market every few seconds — with the price floating to keep both sides showing up.

Functional

  • Rider requests a ride: pickup, dropoff, upfront fare
  • Match a driver within seconds, show live ETA
  • Real-time trip tracking for rider and driver
  • Surge pricing when demand exceeds supply
  • Driver app: go online/offline, accept or decline, navigation
  • Payments, receipts, ratings after dropoff

Non-functional

  • Match latency: p99 under 2 seconds from tap to driver assigned
  • Location ingestion: tens of thousands of driver pings per second
  • Availability: 99.99% for matching — a city can't go dark
  • ETA accuracy: within ~20% or riders stop trusting the app
  • State safety: a ride is never double-assigned or lost mid-trip

🚫 Common misconception

"Matching is just nearest-driver-wins, instantly." It isn't. Production systems hold requests for a few seconds and solve a batch assignment — because greedy matching lets two riders grab the same driver and produces globally worse pairings. The batch window is the difference between a demo and a real marketplace. Nearest-driver is the starting point you mention, then improve on.

2 · Back-of-the-envelope

Capacity math

AssumptionValue
Rides per day (one large metro)500,000
Ride-request QPS500k ÷ 86,400 ≈ 6 avg · peak 5× ≈ 30 QPS — tiny!
Active drivers50,000
Location updates50k drivers × 1 ping / 5s = 10,000 QPS writes — this dominates
Trip record~2 KB → 500k × 2 KB ≈ 1 GB/day per metro
Geo index memory50k drivers × ~200 bytes ≈ 10 MB — the index is small; the write rate is the challenge
Matching computeBatch every 3s, ~100 requests/batch, assignment over ~1k nearby drivers — milliseconds on one box

The line to say out loud: "Ride requests are only ~30 QPS at peak — the hard scaling problem is 10k QPS of location writes and answering 'who's near this point' in milliseconds."

Go deeper: why location writes dominate everything

Every design decision in the serving path bends around the location stream: the geo index must absorb 10k writes/sec with sub-10ms reads, driver positions go stale in seconds so caching is nearly useless, and a hot downtown cell can receive a wildly disproportionate share of both writes and queries. That's why Uber built dedicated geospatial indexing (originally geohash grids, later S2 cells) instead of leaning on a general database.

3 · Architecture

The system, end to end

flowchart TB
    RA["Rider app"]
    DA["Driver app"]
    GW["API gateway"]
    LOC["Location ingest
Kafka stream"] GEO["Geo index
Redis GEO / S2 cells"] DISP["Dispatch service
batch matching engine"] PRICE["Pricing service
surge control loop"] ETA["ETA service
routing engine"] TRIP["Trip store
Postgres + PostGIS"] NOTIF["Notifications
push + SMS"] RA --> GW DA --> GW DA --> LOC LOC --> GEO GW --> DISP DISP --> GEO DISP --> PRICE DISP --> ETA DISP --> TRIP DISP --> NOTIF PRICE -. "fare quote" .-> RA ETA -. "live tracking" .-> RA

Notice the two speeds: the location stream (firehose, seconds-old data is fine) and the dispatch path (must decide in ~2s). They only meet inside the geo index.

4 · Component deep-dives

A ride, step by step

sequenceDiagram
    autonumber
    participant R as Rider app
    participant D as Dispatch
    participant G as Geo index
    participant Dr as Driver app
    participant P as Pricing
    R->>D: POST /v1/rides (pickup, dropoff)
    D->>D: hold in 3s batch window
    D->>G: drivers near pickup
    G-->>D: candidate set
    D->>D: solve assignment (not greedy)
    D->>P: fare quote with surge
    P-->>D: $24.50 at 1.4x
    D->>Dr: offer ride (15s to accept)
    Dr-->>D: accept
    D-->>R: 202 matched, driver 4 min away
    Dr->>D: location stream (pickup to dropoff)
    D-->>R: live tracking over websocket

The surge control loop

Surge isn't price gouging — it's a thermostat. Too many waiting riders? Raise the price: some riders wait, and drivers from quiet zones drive toward the heat. Too many idle drivers? Let it decay. The loop runs every ~30 seconds per neighborhood cell.

flowchart LR
    A["Measure per cell, every 30s:
waiting riders vs free drivers"] --> B{"demand > supply?"} B -->|"yes"| C["Raise multiplier
+0.25 steps, cap ~3x"] B -->|"no"| D["Decay toward 1.0x"] C --> E["Higher fares: some riders wait,
drivers reposition toward demand"] D --> E E --> A
Go deeper: geohash vs S2 cells

A geohash turns a lat/lng into a string where shared prefixes mean nearby — great for "find drivers near me" as a prefix scan. But geohash cells have awkward edge cases at boundaries and wildly varying cell sizes by latitude. S2 (Google's spherical geometry library) projects the globe onto a cube and numbers cells along a Hilbert curve — uniform-ish cells, no edge discontinuities, and clean parent/child hierarchies for zooming out when no driver is nearby. Either is defensible in an interview; S2 is the "I've read the engineering blogs" answer.

Go deeper: the assignment algorithm

With a batch of riders R and candidate drivers D, you want to minimize total pickup ETA — the classic assignment problem, solvable optimally by the Hungarian algorithm in O(n³). At city scale that's too slow per batch, so production systems use greedy-with-lookahead, local search, or batched Hungarian on small partitions. Name Hungarian, then say you'd approximate it — that sentence alone signals senior-level judgment.

5 · API + data model

The contract

POST /v1/rides
{"pickup": {"lat": 37.77, "lng": -122.41}, "dropoff": {"lat": 37.79, "lng": -122.39}}
→ 202 Accepted
{"ride_id": "r_8f2k", "driver": "d_114", "eta_min": 4, "fare": {"amount": 24.50, "surge": 1.4}}

GET /v1/rides/r_8f2k          # ride state machine: requested → matched → pickup → trip → done
POST /v1/drivers/location     # driver heartbeat: {lat, lng, status}
GET /v1/rides/r_8f2k/track    # websocket: live driver position

Data model: Driver {id, cell, lat, lng, status, updated_at} in the geo index (ephemeral — a driver who stops heartbeating vanishes); Ride {id, rider_id, driver_id, state, fare, events[]} in Postgres as the durable record. The ride state machine is the source of truth — the map is just a view.

6 · Trade-offs

What you give up, on purpose

DecisionWhyCost
3-second batch windowGlobally better matches, no double-assignmentRiders wait ~1.5s longer on average for a match
Surge pricingBalances the marketplace in minutesRider anger — needs transparent UI and caps
ETA from routing engine, not straight lineTrustworthy ETAsCPU-heavy; precompute per metro, cache aggressively
Ephemeral driver presenceGeo index stays small and freshMust handle drivers blinking in and out (tunnels, dead zones)
Upfront fareRider certaintyCompany absorbs traffic/detour risk — needs good ETA models
7 · Failure modes

What breaks, and what saves you

8 · What I'd actually build

Opinionated, concrete, shippable

Dispatch: Go service, 3s batch loop, greedy assignment with one lookahead pass. Geo index: Redis with GEO commands, S2 cells as the sharding key, 5s driver heartbeats over a Kafka topic. ETA: Valhalla/OSRM routing engine per metro, precomputed. Trips: Postgres + PostGIS as the durable record. Surge: per-cell control loop every 30s, capped at 3×, every change logged for audit. One team, one quarter, and it handles the math above.

9 · Interview tips

How to run the room

  1. Lead with the geo index. "Drivers stream locations into a geospatial index" in minute two shows you see the real problem.
  2. Say the batch-matching sentence early, then defend it when challenged — it's the highest-signal moment in the interview.
  3. Draw surge as a control loop, not a pricing table. Interviewers remember the thermostat analogy.
  4. Quantify: 30 QPS of requests vs 10k QPS of location writes. The contrast is the whole scaling story.
  5. Keep a safety/fraud bullet in your pocket (fake GPS, collusion) — one sentence shows marketplace maturity.
  6. End with what you'd measure: match latency, ETA error, driver utilization, surge frequency.
10 · Interactive widget

Matching simulator — add riders, watch dispatch work

A living city grid. Green drivers cruise; tap Add rider (or tap the map) and watch the dispatcher pair the nearest free driver, draw the pickup route, and compute an ETA. Hit Rush hour to flood demand and watch surge pricing kick in.

available driver en route to pickup on trip waiting rider
Available drivers: – Waiting riders: – Surge: 1.0× Avg pickup ETA: – Trips completed: 0

Simplified on purpose: real dispatch batches requests for a few seconds and solves an assignment problem instead of pure nearest-driver — but the pieces you see here (geo lookup, ETA math, surge reacting to the waiting/free ratio) are the real ones.

v2026.10.03-01