About

About VibeInfra

VibeInfra is an incident-readiness platform for infrastructure engineers. We build production-realistic outage simulations that run in live disposable sandboxes, so engineers meet a failure mode for the first time in practice rather than at 2am.

What we do

Most infrastructure training explains a failure and then asks you to recognise it in a multiple-choice question. That teaches vocabulary, not diagnosis. The engineer who has read about CrashLoopBackOff and the engineer who has traced one back to a readiness probe on the wrong port are not the same engineer, and only one of them is useful at 2am.

So every VibeInfra lab is a real outage. We provision a multi-container environment, inject a genuine fault, and hand you logs, metrics, a topology view and a terminal attached to a system that is actually broken. You diagnose it, you fix it, and an automated evaluator checks the system recovered — it grades the state you left behind, not an answer you selected. Then you get an after-action report.

What we cover

The catalog is organised as a skill ladder across nine domains: database and caching, networking and edge, async systems and streaming, Linux kernel and OS, cloud infrastructure, CI/CD and deploy, security, observability, and Kubernetes. Each incident carries a difficulty tier from Beginner to Expert and an honest duration.

The technologies are the ones that actually page people: Linux, Docker, Kubernetes, Nginx, PostgreSQL, Kafka, Prometheus and AWS. The full live catalog is at /labs/, and machine-readable at api.vibeinfra.id/api/v1/courses.

How we build

  • Production realism over toy examples. If a scenario could not happen in a real system, it does not teach anything transferable.
  • Fast feedback. Grading runs against the live environment, so you find out whether your fix worked in seconds, not on submission.
  • Learn by doing. No lab is passed by reading. Every one of them requires touching the system.
  • Incremental progression. The ladder exists so nobody's first incident is unwinnable and nobody's tenth is trivial.

Open by default

The incident catalog is a public, unauthenticated index — it is meant to be read, crawled and cited. There is a documented REST API, an OpenAPI 3.1 description, an MCP server for AI agents, SDKs for JavaScript and Python, and agent-facing instructions at /agents.md. We also publish open-source infrastructure tooling through our GitHub organization.

There is a free tier with foundation labs, course chapters, open-source repositories and technical documentation. Advanced flight simulators and verified skill certificates are on the paid tier; see pricing.

Reach us

Email [email protected], or see the contact page for the right channel. How we handle personal data is set out in our privacy policy.