Distributed Transactions Explained
A distributed transaction spans multiple machines, so atomicity stops being a log-local property. We walk through two-phase commit as the protocol machines actually run: the coordinator writes a prepare record, participants flush and vote, then everyone waits for the commit decision.
Topics covered:
- Why a single-machine atomicity model breaks across shards
- Two-phase commit: prepare phase, vote, commit phase
- The coordinator's write-ahead log and why the decision must be durable
- The blocking window: what happens when the coordinator crashes
- 3PC and why it trades blocking for liveness
- Commit in practice: Spanner's Paxos groups under a 2PC layer
- When distributed transactions are worth the coordination cost
Related articles
Two-Phase Commit and Why It Fails
How two-phase commit coordinates atomic transactions across machines, why a crashed coordinator blocks every participant, and why Paxos and Raft now carry the decision.
Eventual Consistency Is a Spectrum
Strong, causal, read-your-writes — the consistency zoo and when each model is the right engineering choice.
Why Distributed Systems Are Hard
The CAP theorem's real meaning, partial failures, and the eight fallacies — why networked code is a different discipline.
More in Distributed Systems
Raft Consensus Explained
Leader election, log replication, and safety — the consensus algorithm that powers etcd, CockroachDB, and TiKV, explained from first principles.
WatchData Partitioning Explained
Range, hash, and hybrid partitioning — where each key's row actually lives, how partitions are balanced, and what happens when one partition becomes a hotspot.
DetailsDistributed Locks and Leases
What a lock actually protects in a distributed system — lease-based locks with expiry, fencing tokens against stale holders, and why client crashes are the hard case.
DetailsGossip Protocols
Membership, failure detection, and state propagation — how nodes exchange information through random peer conversations so the whole cluster converges without a coordinator.
DetailsChaos Engineering Basics
Fault injection with a controlled blast radius — killing nodes, dropping packets, and inducing latency to verify that recovery paths actually work, not just that they exist.
DetailsDistributed Tracing Explained
Trace context propagation, span trees, and sampling — how one request's work is reconstructed across services using trace IDs, span IDs, and parent-child relationships.
DetailsLeader Election in Practice
How real systems elect leaders — ZooKeeper's Zab, etcd's Raft, and lease-based locks — plus fencing tokens and why a stale leader must be fenced before it writes.
DetailsConsistent Hashing Visualized
Watch keys land on a hash ring — what consistent hashing actually does when a cache node dies, and why ring position and virtual nodes determine how many keys move.
DetailsRaft Consensus Visualized
See Raft's term clock, randomized leader election, and log replication in motion — what actually happens in etcd, CockroachDB, and TiKV when a server fails or the network splits.
DetailsDepth, delivered weekly
One technical dispatch a week — articles and episode notes before they go public.
One technical dispatch per week. No noise.