Define clear Major Incident criteria
ITILCreate a severity matrix based on impact and urgency. Document it, train on it, and review it quarterly.
Best Practices
Proven practices from ITIL and Agile methodologies, organized by when they apply in the incident lifecycle.
Create a severity matrix based on impact and urgency. Document it, train on it, and review it quarterly.
Ensure 24/7 coverage with clear escalation paths. Use tools like PagerDuty or Opsgenie to manage rotations and avoid fatigue.
Practice major incidents before they happen. Inject failures into pre-production environments and run the full process.
Runbooks rot fast. Assign ownership and review them monthly. Store them where they're easy to find (wiki, runbook repo).
The Incident Manager should coordinate, not troubleshoot. When the MIM gets pulled into technical detail, coordination suffers.
Use 15–30 minute sprints. If a hypothesis doesn't pan out within the timebox, pivot. Prevents tunnel vision.
Silence breeds anxiety. Even "we're still investigating" is better than no update. Set regular communication cadences.
If a recent change caused the incident, roll it back first. Restore service, then figure out what went wrong.
Don't wait until you're desperate. Bring in senior engineers or management before things get worse. There's no penalty for escalating.
Focus on what the system allowed to happen, not who made a mistake. Blame kills learning. Psychological safety drives improvement.
PIR action items must be first-class work items with owners, priorities, and deadlines. Put them in your sprint backlog.
Share post-incident reports across the organization. Other teams will learn from your incidents and may spot similar risks in their own systems.
Track MTTD, MTTA, MTTR, and recurrence rate. Use these metrics to identify systemic issues and justify investment in reliability.
Assess where your organization stands and identify the next steps for improvement.
No formal process. Incidents are handled ad-hoc by whoever is available. No post-incident learning.
Basic process is documented. Roles exist but aren't always followed. PIRs happen sometimes.
Process is consistently followed. Metrics are tracked. PIRs produce real improvements. Tooling supports the process.
The organization anticipates and prevents incidents. Chaos engineering, game days, and continuous improvement are embedded in culture.