It's 2AM. Checkout is down.
You have 90 minutes.
A high-stakes, live timed flight simulator outage for SREs and platform engineers. Real Kubernetes cluster, live cascading failure under traffic load, automated 7-step evaluation. Zero multiple choice.
90-Minute Operational Run of Show
Engineered like an authentic on-call pager rotation. No slides, no tutorials โ pure hands-on diagnostic mastery.
Pager Briefing & Topology
Incident Commander announces Sev-1 alert. Review microservices architecture, initial P99 latency alerts, and customer impact metrics.
Live Outage Triage
Terminal sandboxes provisioned. Run kubectl, isolate poisoned pods, fix connection saturation, and execute 7-step remediation under live loadgen traffic.
Post-Mortem & Leaderboard
Automated After-Action Review (AAR) analysis. Root cause post-mortem breakdown, MTTR metrics reveal, and leaderboard ranking.
Team Readiness Debrief
SRE capability benchmarking. Direct comparison readout, team on-call drills synthesis, and next simulation preview.
Rules of Engagement
Every engineer receives an independent, disposable Linux/k3d container sandbox provisioned in seconds. Your commands and state do not affect other participants.
A background traffic generator relentlessly fires HTTP requests at the gateway. A fix is only scored once the cluster sustains at least 30 seconds at 0% error rate.
Graded strictly on infrastructure state via ./grade.workstation.sh. Checks triage, isolation, pool unclogging, root-cause patch, and scale-up.
Register for 2AM Live
Reserve your live workstation sandbox. Space is capped to ensure cluster performance.
