Data Partitioning Explained
We place rows into partitions and watch two strategies divide the data: range partitioning draws contiguous keys to the same node, hash partitioning scatters them. Both trade locality against balance, and we measure the cost of each decision with real workloads.
Topics covered:
- Partition by range: sorted keys, efficient scans, hotspot risk
- Partition by hash: uniform distribution, scattered reads
- Secondary indexes: the fan-out problem in sharded systems
- Rebalancing: how partitions move when nodes join or leave
- Hotspots: why one key can saturate a partition — celebrities, counters
- Production patterns: DynamoDB partition keys and Cassandra composite keys
Related articles
Eventual Consistency Is a Spectrum
Strong, causal, read-your-writes — the consistency zoo and when each model is the right engineering choice.
Why Distributed Systems Are Hard
The CAP theorem's real meaning, partial failures, and the eight fallacies — why networked code is a different discipline.
Consensus Algorithms Aren't About Agreement
Raft, leader election, and split-brain — what consensus actually solves and what it doesn't.
More in Distributed Systems
Raft Consensus Explained
Leader election, log replication, and safety — the consensus algorithm that powers etcd, CockroachDB, and TiKV, explained from first principles.
WatchDistributed Locks and Leases
What a lock actually protects in a distributed system — lease-based locks with expiry, fencing tokens against stale holders, and why client crashes are the hard case.
DetailsGossip Protocols
Membership, failure detection, and state propagation — how nodes exchange information through random peer conversations so the whole cluster converges without a coordinator.
DetailsChaos Engineering Basics
Fault injection with a controlled blast radius — killing nodes, dropping packets, and inducing latency to verify that recovery paths actually work, not just that they exist.
DetailsDistributed Tracing Explained
Trace context propagation, span trees, and sampling — how one request's work is reconstructed across services using trace IDs, span IDs, and parent-child relationships.
DetailsLeader Election in Practice
How real systems elect leaders — ZooKeeper's Zab, etcd's Raft, and lease-based locks — plus fencing tokens and why a stale leader must be fenced before it writes.
DetailsConsistent Hashing Visualized
Watch keys land on a hash ring — what consistent hashing actually does when a cache node dies, and why ring position and virtual nodes determine how many keys move.
DetailsDistributed Transactions Explained
Two-phase commit, prepare and commit phases, and the coordinator failure window — how databases coordinate atomic writes across machines and what happens when a participant crashes.
DetailsRaft Consensus Visualized
See Raft's term clock, randomized leader election, and log replication in motion — what actually happens in etcd, CockroachDB, and TiKV when a server fails or the network splits.
DetailsDepth, delivered weekly
One technical dispatch a week — articles and episode notes before they go public.
One technical dispatch per week. No noise.