Hands On Kafka

Hands On Kafka

Uber-Lite: Architecting High-Scale Geo-Spatial Matchmaking Systems

Lesson 46 — The DriverProcess Logic: Upsert into a Composite-Key RocksDB Store

Jul 19, 2026
∙ Paid

The Naive Approach — And Why It Dies at 10k Drivers/sec

A typical engineer models driver location like this:

// Naive: KTable keyed by driverId
store.put(driverId, new DriverLocation(lat, lng));

This compiles. It runs in dev. It collapses in production for two independent reasons.

Reason 1: O(N) matching. A rider requests a pickup. Your matching engine needs all drivers within ~1 km. With a driverId-keyed store, there is no spatial index. You iterate every entry and compute distance. At 50,000 active drivers, that is 50,000 deserialization + Haversine calls per rider request. At 200 rider requests/sec, that is 10M operations/sec — on a single StreamThread — before it has processed any driver events. The process-latency-max metric will flatline your Grafana dashboard around 8,000 drivers.

Reason 2: Ghost drivers. A driver moves from cell A to cell B. You put(driverId, newLocation). The old record is overwritten. Fine for a simple KV lookup. But once you switch to the correct composite key design (explained below), the old composite key (cellA, driverId) is never deleted. Within 90 seconds of a 3,000-event/sec stream, you accumulate tens of thousands of stale entries — ghost drivers that appear to be at locations they left long ago. Your matching engine dispatches riders to phantoms.

Both failures are architectural. You cannot patch them with tuning.


The Uber-Lite Architecture — Inverted Spatial Index in RocksDB

User's avatar

Continue reading this post for free, courtesy of Kafka.

Or purchase a paid subscription.
© 2026 SystemDR · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture