MIM Guide

Process Framework

Major Incident Lifecycle

The five phases of managing a major incident, from first alert to post-incident improvement. Click each phase to explore objectives, best practices, and ITIL/Agile perspectives.

Key objectives

  • Automated monitoring triggers alert based on thresholds
  • Service desk or NOC receives initial report
  • Preliminary impact assessment is performed
  • Major Incident criteria are evaluated

Best practices

  • Define clear thresholds for auto-detection (e.g., >5% error rate, latency >2s)
  • Ensure monitoring covers all critical services and dependencies
  • Use a single pane of glass for alert aggregation (PagerDuty, Opsgenie, etc.)
  • Automate initial correlation so you avoid alert storms

ITIL Perspective

ITIL calls this "Event Management" feeding into "Incident Management." The key is reducing Mean Time to Detect (MTTD).

Agile Perspective

Agile teams often embed monitoring into their CI/CD pipeline, and "you build it, you run it" accelerates detection.