Skip to content

/lab/failure-lab

02

Failure Lab

SYSTEMS / SIMULATION

Break a system. Watch the reasoning happen.

HYPOTHESIS

The instinct that makes a request pipeline resilient — backpressure, circuit breaking, load shedding, recovery — is more convincing demonstrated than described, if the propagation is actually computed rather than scripted per scenario.

HOW TO OPERATE IT

A five-stage pipeline (Traffic → Checkout → Payment → Orders → Fulfillment) with five live parameters (traffic load, payment failure rate, database pressure, network latency, service availability). Every tick recomputes queue depth, latency, and throughput per stage from the tick before it — nothing about a cascade is pre-authored. Four optional incident scenarios script realistic parameter changes over time; the visitor can still intervene freely during any of them.

WHAT WAS DIFFICULT

Deriving the live event stream and the post-incident outcome summary (time to recovery, primary failure, key decision) directly from simulation state transitions, so every log line and every outcome field is provably computed rather than a plausible-sounding canned string.

WHAT IT TAUGHT ME

The shape of this pipeline — and the instinct for where backpressure and circuit breaking need to sit — comes directly from debugging real production Shopify checkout and payment flows, not from reading about distributed systems theory.

TECHNOLOGY

  • Deterministic tick-based simulation
  • finite-state circuit breaker
  • computed event log

REAL SKILLS THIS DRAWS ON

  • Distributed-systems reasoning
  • backpressure & circuit-breaking design
  • incident response
BUILD SOMETHING