Difficulty Distribution
Skills You'll Practice (15)
All 5 Scenarios
Click any scenario to practice it. No account required.
Production Database Down
It's a Monday at 9:03 AM. You get paged: the production database is unreachable. Your monitoring shows 100% error rate on the checkout service. Your on-call engineer is stuck on a call. 50,000 users are currently online and can't complete purchases. Walk me through your incident response.
Design a CI/CD Pipeline
Your company has 5 developers pushing to GitHub multiple times daily. Currently deploys are manual — someone SSHes in and runs a script. This takes 2 hours every time. Design a CI/CD pipeline that: runs tests automatically, deploys to staging on every PR, deploys to production on merge to main, with a rollback mechanism. What tools would you use and why?
Kubernetes Pods Keep Crashing
Your Kubernetes production cluster has 3 pods that keep OOMKilled every 15-20 minutes. The app is a Python Flask API. Memory limits are set to 512MB. The pods restart, run fine for a bit, then crash again. Walk me through your debugging approach.
Terraform State File Corrupted
Your team's Terraform state file (stored in S3) got corrupted after a failed apply. You can't run terraform plan without errors. You have 3 environments (dev, staging, prod) all managed from this state. What do you do to recover without destroying infrastructure?
Deployment Rollback Strategy
You're rolling out a new microservice to production. It's a payment processing service handling $2M/day. The deploy needs zero downtime. What deployment strategy do you use? What's your rollback plan? How do you validate success?