The Runtime Theory
System DesignIn production

Notification System Design: Fanout, Retries, and Delivery Channels

Recording in progress
#notifications#async#architecture

A notification system is a distributed system where the failure is guaranteed — every delivery channel fails differently and on its own schedule. We design the pipeline from event ingestion to provider fanout, then trace a notification through queue, channel adapter, and provider call, showing where state is lost and where retries create duplicates. The three channels — push, email, SMS — get separate failure profiles and separate retry policies, and we cover the dedupe key that keeps a retried batch from spamming a user.

Topics covered:

  • Event ingestion and template resolution
  • Channel providers and the abstraction layer
  • Fanout to multiple users and channel priority
  • Retry with exponential backoff and dead-lettering
  • Idempotency keys and duplicate suppression

Related articles

More in System Design

09:03
system design

Load Balancers, Explained

Layer 4 vs layer 7, health checks, stickiness, and the failure modes behind the most trusted box in your architecture diagram.

Watch
In production
system design

Search System Design: Inverted Indexes, Tokenization, and Ranking

How a search system actually works — inverted indexes, tokenization, relevance ranking, and the pipeline between a keystroke and ranked results.

Details
In production
system design

Chat System Design: WebSockets, Presence, and Message Ordering

Designing a chat system — WebSocket connections, message ordering, presence, and what happens to undelivered messages when a client goes offline.

Details
In production
system design

Rate Limiter Design: Token Bucket, Sliding Window, and Distributed Limits

How rate limiters actually work — token bucket, fixed window, sliding window, and what breaks when the limiter spans multiple machines.

Details
In production
system design

URL Shortener System Design: Encoding, Storage, and Redirects

Designing a URL shortener — base62 encoding, ID generation, redirect caching, and how the write and read paths differ in scale.

Details
In production
system design

Message Queue Design: Topics, Consumer Groups, and Delivery Guarantees

How message queues actually work — topics, partitions, consumer groups, exactly-once semantics, and what guarantees the broker really gives.

Details
In production
system design

Database Replication Topologies Visualized: Leader, Multi-Leader, and Quorum

How database replication topologies actually work — single leader, multi leader, quorum, and what happens to reads and writes when a node fails.

Details
In production
system design

Caching Strategies: Cache Aside, Write Through, and Cache Invalidation

How caching strategies actually work in production systems — cache aside, read through, write through, write back, and where stale reads come from.

Details
In production
system design

Load Balancer Algorithms: Round Robin, Least Connections, and Hashing

How load balancer algorithms actually distribute requests — round robin, least connections, IP hashing, and the trade-offs each one makes.

Details

Depth, delivered weekly

One technical dispatch a week — articles and episode notes before they go public.

One technical dispatch per week. No noise.