URL Shortener System Design: Encoding, Storage, and Redirects
We design a URL shortener from the write path up: base62 and base64 encoding, random IDs versus counters, and where collision handling actually lives. Then we trace a redirect — the 301 versus 308 decision, what the browser and the cache do with each, and why the analytics write is the real bottleneck. The scale story is in the asymmetry: one write to the database, hundreds of redirects per short code, so the cache hit rate is the entire game.
Topics covered:
- Base62 encoding and reversible IDs
- Counter-based vs random ID generation
- 301 vs 302 vs 308 redirect semantics
- Cache-first read path and stale redirects
- Click analytics and the write path
Related articles
Health Checks and Observability
Active vs passive health checks, circuit breaker interplay, and the four concrete reasons a dead node still gets production traffic.
Capacity Planning: The Math
QPS times latency equals concurrency, utilization targets, and headroom: the arithmetic behind fleet sizing that survives peak traffic.
Backpressure and Flow Control
Bounded queues, rejection vs blocking, TCP-style windowing, and why an unbounded queue turns a slow downstream into an OOM crash.
More in System Design
Load Balancers, Explained
Layer 4 vs layer 7, health checks, stickiness, and the failure modes behind the most trusted box in your architecture diagram.
WatchSearch System Design: Inverted Indexes, Tokenization, and Ranking
How a search system actually works — inverted indexes, tokenization, relevance ranking, and the pipeline between a keystroke and ranked results.
DetailsNotification System Design: Fanout, Retries, and Delivery Channels
Designing a notification system — provider abstraction, fanout, retry policies, and how push, email, and SMS channels fail differently.
DetailsChat System Design: WebSockets, Presence, and Message Ordering
Designing a chat system — WebSocket connections, message ordering, presence, and what happens to undelivered messages when a client goes offline.
DetailsRate Limiter Design: Token Bucket, Sliding Window, and Distributed Limits
How rate limiters actually work — token bucket, fixed window, sliding window, and what breaks when the limiter spans multiple machines.
DetailsMessage Queue Design: Topics, Consumer Groups, and Delivery Guarantees
How message queues actually work — topics, partitions, consumer groups, exactly-once semantics, and what guarantees the broker really gives.
DetailsDatabase Replication Topologies Visualized: Leader, Multi-Leader, and Quorum
How database replication topologies actually work — single leader, multi leader, quorum, and what happens to reads and writes when a node fails.
DetailsCaching Strategies: Cache Aside, Write Through, and Cache Invalidation
How caching strategies actually work in production systems — cache aside, read through, write through, write back, and where stale reads come from.
DetailsLoad Balancer Algorithms: Round Robin, Least Connections, and Hashing
How load balancer algorithms actually distribute requests — round robin, least connections, IP hashing, and the trade-offs each one makes.
DetailsDepth, delivered weekly
One technical dispatch a week — articles and episode notes before they go public.
One technical dispatch per week. No noise.