The Runtime Theory
System DesignIn production

Rate Limiter Design: Token Bucket, Sliding Window, and Distributed Limits

Recording in progress
#rate-limiting#distributed-systems#api

The inner mechanics of rate limiting: we implement token bucket, fixed window, sliding window log, and sliding window counter on a single machine, then add a second machine and watch the classic failure — a client at the edge of its limit can slip through the gap between two limiter nodes. We cover where the limit state lives, how Redis changes the arithmetic, and the counter-intuitive trade-off that local counters and approximate global limits are the shape most real APIs ship with.

Topics covered:

  • Token bucket: burst capacity and refill rate
  • Fixed window and the boundary burst
  • Sliding window log vs counter memory cost
  • Distributed limits and the Redis race
  • Returning 429 correctly: headers and retry-after

Related articles

More in System Design

09:03
system design

Load Balancers, Explained

Layer 4 vs layer 7, health checks, stickiness, and the failure modes behind the most trusted box in your architecture diagram.

Watch
In production
system design

Search System Design: Inverted Indexes, Tokenization, and Ranking

How a search system actually works — inverted indexes, tokenization, relevance ranking, and the pipeline between a keystroke and ranked results.

Details
In production
system design

Notification System Design: Fanout, Retries, and Delivery Channels

Designing a notification system — provider abstraction, fanout, retry policies, and how push, email, and SMS channels fail differently.

Details
In production
system design

Chat System Design: WebSockets, Presence, and Message Ordering

Designing a chat system — WebSocket connections, message ordering, presence, and what happens to undelivered messages when a client goes offline.

Details
In production
system design

URL Shortener System Design: Encoding, Storage, and Redirects

Designing a URL shortener — base62 encoding, ID generation, redirect caching, and how the write and read paths differ in scale.

Details
In production
system design

Message Queue Design: Topics, Consumer Groups, and Delivery Guarantees

How message queues actually work — topics, partitions, consumer groups, exactly-once semantics, and what guarantees the broker really gives.

Details
In production
system design

Database Replication Topologies Visualized: Leader, Multi-Leader, and Quorum

How database replication topologies actually work — single leader, multi leader, quorum, and what happens to reads and writes when a node fails.

Details
In production
system design

Caching Strategies: Cache Aside, Write Through, and Cache Invalidation

How caching strategies actually work in production systems — cache aside, read through, write through, write back, and where stale reads come from.

Details
In production
system design

Load Balancer Algorithms: Round Robin, Least Connections, and Hashing

How load balancer algorithms actually distribute requests — round robin, least connections, IP hashing, and the trade-offs each one makes.

Details

Depth, delivered weekly

One technical dispatch a week — articles and episode notes before they go public.

One technical dispatch per week. No noise.