The Runtime Theory
Distributed SystemsIn production

Chaos Engineering Basics

Recording in progress
#chaos-engineering#fault-injection#reliability

Chaos engineering is not random destruction; it is a controlled experiment with a hypothesis. We run three faults against a live cluster — a node kill, a network partition, and artificial latency — and watch what the system's recovery paths actually do.

Topics covered:

  • The experiment format: steady state, hypothesis, injection, blast radius
  • A node kill: how the cluster detects it and re-replicates data
  • A network partition: where quorum loss freezes writes and why
  • Latency injection: how retries and timeouts interact under load
  • Blast radius: why faults run in a small, isolated scope first
  • How Chaos Monkey, Litmus, and Gremlin schedule and execute faults
  • Steady-state verification: comparing metrics before and after

Related articles

More in Distributed Systems

08:50
distributed systems

Raft Consensus Explained

Leader election, log replication, and safety — the consensus algorithm that powers etcd, CockroachDB, and TiKV, explained from first principles.

Watch
In production
distributed systems

Data Partitioning Explained

Range, hash, and hybrid partitioning — where each key's row actually lives, how partitions are balanced, and what happens when one partition becomes a hotspot.

Details
In production
distributed systems

Distributed Locks and Leases

What a lock actually protects in a distributed system — lease-based locks with expiry, fencing tokens against stale holders, and why client crashes are the hard case.

Details
In production
distributed systems

Gossip Protocols

Membership, failure detection, and state propagation — how nodes exchange information through random peer conversations so the whole cluster converges without a coordinator.

Details
In production
distributed systems

Distributed Tracing Explained

Trace context propagation, span trees, and sampling — how one request's work is reconstructed across services using trace IDs, span IDs, and parent-child relationships.

Details
In production
distributed systems

Leader Election in Practice

How real systems elect leaders — ZooKeeper's Zab, etcd's Raft, and lease-based locks — plus fencing tokens and why a stale leader must be fenced before it writes.

Details
In production
distributed systems

Consistent Hashing Visualized

Watch keys land on a hash ring — what consistent hashing actually does when a cache node dies, and why ring position and virtual nodes determine how many keys move.

Details
In production
distributed systems

Distributed Transactions Explained

Two-phase commit, prepare and commit phases, and the coordinator failure window — how databases coordinate atomic writes across machines and what happens when a participant crashes.

Details
In production
distributed systems

Raft Consensus Visualized

See Raft's term clock, randomized leader election, and log replication in motion — what actually happens in etcd, CockroachDB, and TiKV when a server fails or the network splits.

Details

Depth, delivered weekly

One technical dispatch a week — articles and episode notes before they go public.

One technical dispatch per week. No noise.