MIM Guide

Best Practices

ITIL & Agile Best Practices

Proven practices from ITIL and Agile methodologies, organized by when they apply in the incident lifecycle.

Preparation

Define clear Major Incident criteria

ITIL

Create a severity matrix based on impact and urgency. Document it, train on it, and review it quarterly.

Maintain an on-call rotation

Agile

Ensure 24/7 coverage with clear escalation paths. Use tools like PagerDuty or Opsgenie to manage rotations and avoid fatigue.

Run regular incident drills (Game Days)

Both

Practice major incidents before they happen. Inject failures into pre-production environments and run the full process.

Keep runbooks up-to-date

ITIL

Runbooks rot fast. Assign ownership and review them monthly. Store them where they're easy to find (wiki, runbook repo).

During the Incident

Separate the MIM from the technical work

ITIL

The Incident Manager should coordinate, not troubleshoot. When the MIM gets pulled into technical detail, coordination suffers.

Timebox investigation efforts

Agile

Use 15–30 minute sprints. If a hypothesis doesn't pan out within the timebox, pivot. Prevents tunnel vision.

Communicate early and often

Both

Silence breeds anxiety. Even "we're still investigating" is better than no update. Set regular communication cadences.

Prefer rollback over forward-fix

Agile

If a recent change caused the incident, roll it back first. Restore service, then figure out what went wrong.

Escalate early, not late

ITIL

Don't wait until you're desperate. Bring in senior engineers or management before things get worse. There's no penalty for escalating.

After the Incident

Blameless post-incident reviews

Agile

Focus on what the system allowed to happen, not who made a mistake. Blame kills learning. Psychological safety drives improvement.

Track remediation items as real work

Agile

PIR action items must be first-class work items with owners, priorities, and deadlines. Put them in your sprint backlog.

Publish PIRs broadly

Both

Share post-incident reports across the organization. Other teams will learn from your incidents and may spot similar risks in their own systems.

Measure and improve over time

ITIL

Track MTTD, MTTA, MTTR, and recurrence rate. Use these metrics to identify systemic issues and justify investment in reliability.

Maturity Model

Incident Management Maturity

Assess where your organization stands and identify the next steps for improvement.

1

Level 1: Reactive

No formal process. Incidents are handled ad-hoc by whoever is available. No post-incident learning.

  • No defined severity criteria
  • No on-call rotation
  • No formal communication process
  • No PIRs conducted
2

Level 2: Defined

Basic process is documented. Roles exist but aren't always followed. PIRs happen sometimes.

  • Severity matrix exists
  • On-call rotation in place
  • Bridge call process defined
  • PIRs for SEV-1 only
3

Level 3: Managed

Process is consistently followed. Metrics are tracked. PIRs produce real improvements. Tooling supports the process.

  • MTTR tracked and improving
  • PIRs for all major incidents
  • Remediation items tracked to completion
  • Regular process reviews
4

Level 4: Proactive

The organization anticipates and prevents incidents. Chaos engineering, game days, and continuous improvement are embedded in culture.

  • Regular game days / chaos engineering
  • Proactive monitoring and alerting
  • Cross-team incident learning
  • Automation reduces MTTR consistently