/lab/failure-lab
Failure Lab
SYSTEMS / SIMULATION
Break a system. Watch the reasoning happen.
HYPOTHESIS
The instinct that makes a request pipeline resilient — backpressure, circuit breaking, load shedding, recovery — is more convincing demonstrated than described, if the propagation is actually computed rather than scripted per scenario.
HOW TO OPERATE IT
A five-stage pipeline (Traffic → Checkout → Payment → Orders → Fulfillment) with five live parameters (traffic load, payment failure rate, database pressure, network latency, service availability). Every tick recomputes queue depth, latency, and throughput per stage from the tick before it — nothing about a cascade is pre-authored. Four optional incident scenarios script realistic parameter changes over time; the visitor can still intervene freely during any of them.
WHAT WAS DIFFICULT
Deriving the live event stream and the post-incident outcome summary (time to recovery, primary failure, key decision) directly from simulation state transitions, so every log line and every outcome field is provably computed rather than a plausible-sounding canned string.
WHAT IT TAUGHT ME
The shape of this pipeline — and the instinct for where backpressure and circuit breaking need to sit — comes directly from debugging real production Shopify checkout and payment flows, not from reading about distributed systems theory.
TECHNOLOGY
- Deterministic tick-based simulation
- finite-state circuit breaker
- computed event log
REAL SKILLS THIS DRAWS ON
- Distributed-systems reasoning
- backpressure & circuit-breaking design
- incident response