Process Framework
Major Incident Lifecycle The five phases of managing a major incident, from first alert to post-incident improvement. Click each phase to explore objectives, best practices, and ITIL/Agile perspectives.
Detection Triage Investigation Resolution Post-Incident
1. Detection & Alerting Identify that a major incident is occurring through monitoring, user reports, or automated alerts.
0–5 minKey objectives
Automated monitoring triggers alert based on thresholds Service desk or NOC receives initial report Preliminary impact assessment is performed Major Incident criteria are evaluated Best practices
Define clear thresholds for auto-detection (e.g., >5% error rate, latency >2s) Ensure monitoring covers all critical services and dependencies Use a single pane of glass for alert aggregation (PagerDuty, Opsgenie, etc.) Automate initial correlation so you avoid alert storms ITIL Perspective
ITIL calls this "Event Management" feeding into "Incident Management." The key is reducing Mean Time to Detect (MTTD).
Agile Perspective
Agile teams often embed monitoring into their CI/CD pipeline, and "you build it, you run it" accelerates detection.
2. Triage & Declaration Classify severity, declare the major incident, and mobilize the response team.
5–15 min3. Investigation & Diagnosis Systematically investigate root cause while working toward a fix or workaround.
15 min – 2 hr4. Resolution & Recovery Implement the fix or workaround to restore service, then verify stability.
Varies5. Post-Incident Review Conduct a blameless review to learn from the incident and drive systemic improvements.
Within 48 hr