VIBEINFRA COCKPIT // INCIDENT LABS
XP: 4,850 XPAVAILABLE INCIDENTS: 28
ON-CALL INCIDENT SIMULATOR

VibeInfra Cockpit

Choose an incident to practice. Investigate live telemetry, inspect container socket logs, modify Kubernetes manifests, and restore degraded production systems with automated verification.

First incident is on us

See how you investigate under pressure.

1. Read the incident brief2. Enter a disposable cluster3. Restore the SLO

No credit card · No local setup · 0 VibeCredits

Start Free Kubernetes Incident
P1 CriticalKubernetes
Free Trial Sandbox

Incident #22 — Kubernetes CrashLoopBackOff

The production Ingress gateway returns HTTP 503. Workloads in the auth namespace fail to stay up. Use live kubectl diagnostics to locate the missing secret and restore HTTP 200 OK.

Cluster Telemetry● 503 Outage
Ingress Status:HTTP 503 Service Unavailable
Target SLO:HTTP 200 OK Sustained
⏱️ 45 mins · Advanced · 140 XPLaunch Simulator (Free) →
P1 OutageObservability
Free Trial Sandbox

Incident #35 — Dashboards Green, Customers Angry

Checkout fails for real customers, yet all synthetic dashboards report 100% Green due to an unrevealed prober payload blindspot. Cross-examine independent metric streams.

Telemetry Blindspot● Silent Failure
Synthetic Probe:100% (False Green)
Real Checkout QPS:0% Completed
⏱️ 45 mins · Expert · 240 XPLaunch Simulator (Free) →

Choose your mission path

Start Here

Incident #22 — Kubernetes Microservices CrashLoopBackOff
P1 CriticalFREE

Incident #22 — Kubernetes Microservices CrashLoopBackOff

The production Ingress gateway returns HTTP 503 Service Unavailable. Workloads in the auth namespace are crashing with OOMKill/CrashLoop. Use live kubectl diagnostics to resolve.

Kubernetes (K8s)Traefik IngressSecrets & ConfigMaps
Advanced·45 mins
Launch Free
Incident #35 — Dashboards Green, Customers Angry
P1 OutageFREE

Incident #35 — Dashboards Green, Customers Angry

Support tickets pile up: checkout is failing for real customers, yet all synthetic probers read 100% Green due to an unrevealed payload blindspot. Cross-examine metric streams.

Go MicroservicesSynthetic ProberPrometheus Metrics
Expert·45 mins
Launch Free
Incident #23 — API Gateway 5xx Errors Under Load
P1 IncidentFREE

Incident #23 — API Gateway 5xx Errors Under Load

Customers see intermittent 502/503/504 errors from the API under moderate load. The backend reports healthy — the failure is in the upstream connection pool of Nginx.

NginxReverse ProxyGo BackendKeepalive
Intermediate·30 mins
Launch Free

Kubernetes Firefights

Incident #22 — Kubernetes Microservices CrashLoopBackOff
P1 CriticalFREE

Incident #22 — Kubernetes Microservices CrashLoopBackOff

The production Ingress gateway returns HTTP 503 Service Unavailable. Workloads in the auth namespace are crashing with OOMKill/CrashLoop. Use live kubectl diagnostics to resolve.

Kubernetes (K8s)Traefik IngressSecrets & ConfigMaps
Advanced·45 mins
Launch Free
Incident #21 — Payment Gateway Outage Triage
P1 Outage🪙 30 Credits

Incident #21 — Payment Gateway Outage Triage

E-commerce checkout transactions are failing with timeout cascades across payment pods. Isolate faulty sidecars, clear socket leaks, and restore payment processing.

KubernetesgRPCSidecarsEnvoy Proxy
Advanced·45 mins
Start Incident (30 Cr)
Incident #32 — Go Memory Leak & OOMKill Rolling Restarts
P1 Incident🪙 30 Credits

Incident #32 — Go Memory Leak & OOMKill Rolling Restarts

Pod replicas are sequentially killed every 12 minutes by the Linux cgroup OOM killer during peak load. Analyze pprof heap profiles and tune memory limits.

KubernetesGo pprofcgroups v2Memory Profiling
Advanced·45 mins
Start Incident (30 Cr)

Linux & Kernel

Incident #24 — Background Worker Queue Meltdown
P1 Incident🪙 30 Credits

Incident #24 — Background Worker Queue Meltdown

Checkouts succeed and payments are captured, but confirmation emails and order pipelines aren't processing. Investigate asynchronous queue backpressure and lock contention.

GoPostgreSQLAsync Worker QueueRedis
Advanced·45 mins
Start Incident (30 Cr)
Incident #25 — Product Image Uploads Failing (S3)
P2 Incident🪙 30 Credits

Incident #25 — Product Image Uploads Failing (S3)

The upload API returns 200 OK to every request, but photos never appear on the storefront. Investigate AWS-shaped S3 bucket policies and SSM key path mismatches.

GoAWS S3IAM PolicySSM Parameter Store
Intermediate·30 mins
Start Incident (30 Cr)
Incident #28 — Linux Kernel nf_conntrack Table Exhaustion
P1 Incident🪙 30 Credits

Incident #28 — Linux Kernel nf_conntrack Table Exhaustion

Incoming TCP SYN packets are silently dropped before reaching Nginx. App logs report 0% CPU, but customer requests hang for 30s. Tune kernel Netfilter ring buffers and sysctl.

Linux KernelNetfiltersysctlTCP Sockets
Advanced·45 mins
Start Incident (30 Cr)

Docker Incidents

Incident #26 — Deploy Pipeline Backlog & Runner Starvation
P1 Incident🪙 30 Credits

Incident #26 — Deploy Pipeline Backlog & Runner Starvation

PRs merge to main and CI reports green, but production containers never update. Investigate asynchronous deploy queue locks and runner resource exhaustion.

DockerCI/CD RunnersPostgreSQLDeploy Lock
Intermediate·35 mins
Start Incident (30 Cr)
Incident #31 — Docker /var/lib/docker Inode & Disk Saturation
P2 Incident🪙 30 Credits

Incident #31 — Docker /var/lib/docker Inode & Disk Saturation

Containers fail to spawn with 'no space left on device', even though df -h shows 40% disk space free. Uncover inode exhaustion from abandoned build caches.

Docker Engineoverlay2Linux InodesStorage Quota
Intermediate·30 mins
Start Incident (30 Cr)
P1 Incident🪙 30 Credits

Incident #63 — The Page That Landed Mid-Deploy

A routine release is queued — ship it. While you're heads-down on your ticket, a real Postgres connection-pool leak on an unrelated service self-triggers, unannounced, and every second it goes unnoticed is real, permanent damage.

GoPostgreSQLDeploy PipelineLive-Interrupt Drill
Expert·45 mins
Start Incident (30 Cr)

Networking Incidents

Incident #23 — API Gateway 5xx Errors Under Load
P1 IncidentFREE

Incident #23 — API Gateway 5xx Errors Under Load

Customers see intermittent 502/503/504 errors from the API under moderate load. The backend reports healthy — the failure is in the upstream connection pool of Nginx.

NginxReverse ProxyGo BackendKeepalive
Intermediate·30 mins
Launch Free
Incident #27 — Flash Sale Traffic Surge: Load Balancer Partition
P1 Outage🪙 30 Credits

Incident #27 — Flash Sale Traffic Surge: Load Balancer Partition

Traffic surges to 2,000+ RPS. The autoscaler scales backend instances from 2 to 10, but 90% of traffic still lands on one overloaded instance. Fix hash ring imbalance.

NginxLoad BalancingConsistent HashingGo Backend
Expert·60 mins
Start Incident (30 Cr)
Incident #30 — Ingress TLS Certificate Expiration Outage
P1 Incident🪙 30 Credits

Incident #30 — Ingress TLS Certificate Expiration Outage

Browser TLS warnings halt 100% of consumer traffic. The cert-manager webhook failed silently during secret renewal. Re-issue and hot-reload TLS certs without gateway downtime.

TLS / SSLCert-ManagerNginx IngressOpenSSL
Beginner·25 mins
Start Incident (30 Cr)

Observability & Data Incidents

Incident #35 — Dashboards Green, Customers Angry
P1 OutageFREE

Incident #35 — Dashboards Green, Customers Angry

Support tickets pile up: checkout is failing for real customers, yet all synthetic probers read 100% Green due to an unrevealed payload blindspot. Cross-examine metric streams.

Go MicroservicesSynthetic ProberPrometheus Metrics
Expert·45 mins
Launch Free
Incident #59 — Leaked Cloud API Key & Security Compromise Triage
P1 Security🪙 30 Credits

Incident #59 — Leaked Cloud API Key & Security Compromise Triage

A production master API key was accidentally committed to public Git history. Revoke access, quarantine compromised resources, and rotate credentials with zero downtime.

Git HistorySecret RotationSecurity TriageAudit Logs
Intermediate·30 mins
Start Incident (30 Cr)
Incident #62 — Prometheus PromQL Silence & Dropped P0 Alert
P1 Incident🪙 30 Credits

Incident #62 — Prometheus PromQL Silence & Dropped P0 Alert

Production was down for 40 minutes before on-call was alerted because a missing group_left() in a PromQL alert rule caused vector matching to return empty sets.

PrometheusPromQLAlertmanagerMetric Vector Matching
Advanced·40 mins
Start Incident (30 Cr)