Chat System Design: WebSockets, Presence, and Message Ordering
We design a chat service around the connection, not the message: how a WebSocket stays alive across proxies, why the gateway holds the only copy of the socket, and what ordering guarantees survive when two gateways are in the room. We trace a message from one phone through gateway, queue, and delivery, then kill the recipient's connection mid-flight and show where the message waits. Presence and read receipts are the same problem — a state that must be eventually consistent — so we build them once.
Topics covered:
- WebSocket lifecycle and gateway topology
- Ordering per room with multiple gateways
- Offline message storage and sync
- Presence: heartbeat, timeout, and false offlines
- Scaling: fanout to one user vs fanout to one room
Related articles
Leader-Follower vs P2P Architectures
Quorum write cost, gossip convergence time, and when peer-to-peer wins — CDNs, blockchains, and systems that can live without ordering.
A Load Balancer Is Not a Magic Box
Layer 4 vs layer 7, hashing vs health checks, stickiness, and the failure modes hiding behind the most over-trusted building block in web infrastructure.
Health Checks and Observability
Active vs passive health checks, circuit breaker interplay, and the four concrete reasons a dead node still gets production traffic.
More in System Design
Load Balancers, Explained
Layer 4 vs layer 7, health checks, stickiness, and the failure modes behind the most trusted box in your architecture diagram.
WatchSearch System Design: Inverted Indexes, Tokenization, and Ranking
How a search system actually works — inverted indexes, tokenization, relevance ranking, and the pipeline between a keystroke and ranked results.
DetailsNotification System Design: Fanout, Retries, and Delivery Channels
Designing a notification system — provider abstraction, fanout, retry policies, and how push, email, and SMS channels fail differently.
DetailsRate Limiter Design: Token Bucket, Sliding Window, and Distributed Limits
How rate limiters actually work — token bucket, fixed window, sliding window, and what breaks when the limiter spans multiple machines.
DetailsURL Shortener System Design: Encoding, Storage, and Redirects
Designing a URL shortener — base62 encoding, ID generation, redirect caching, and how the write and read paths differ in scale.
DetailsMessage Queue Design: Topics, Consumer Groups, and Delivery Guarantees
How message queues actually work — topics, partitions, consumer groups, exactly-once semantics, and what guarantees the broker really gives.
DetailsDatabase Replication Topologies Visualized: Leader, Multi-Leader, and Quorum
How database replication topologies actually work — single leader, multi leader, quorum, and what happens to reads and writes when a node fails.
DetailsCaching Strategies: Cache Aside, Write Through, and Cache Invalidation
How caching strategies actually work in production systems — cache aside, read through, write through, write back, and where stale reads come from.
DetailsLoad Balancer Algorithms: Round Robin, Least Connections, and Hashing
How load balancer algorithms actually distribute requests — round robin, least connections, IP hashing, and the trade-offs each one makes.
DetailsDepth, delivered weekly
One technical dispatch a week — articles and episode notes before they go public.
One technical dispatch per week. No noise.