MIM Guide

Advanced MIM

Leading Incidents When It Gets Hard

Process gets you started. Judgement gets you through. These are the situational skills that separate experienced major incident managers from people following a checklist.

When nothing is going to plan

Running Incidents in the Fog of War

Incidents rarely fail because of tech. They fail because no one takes control of uncertainty. Your job is to make decisions and move forward even when the picture is incomplete.

Decision-making with incomplete data

You will never have the full picture. Accept it. Make a decision based on the best available information and communicate it clearly. A wrong decision acted upon quickly is often better than the right decision 20 minutes too late.

Avoiding paralysis when tech teams disagree

When SMEs contradict each other, don't wait for them to reach consensus organically, because it won't happen. Acknowledge both views, ask each team what they need to prove or disprove their hypothesis in the next 10 minutes, and set a time-box. If still no consensus, you make the call.

Knowing when to force direction vs. facilitate

In the first 15 minutes, facilitate: gather data, set structure, assign owners. After that, if progress stalls, switch to direction mode. The bridge needs a commander, not a chairperson.

Recognising and killing false leads early

Watch for teams chasing rabbit holes. If someone has been investigating the same thread for 20+ minutes with no movement, intervene: "What would it take to confirm or eliminate this theory in the next 5 minutes?" If they can't answer, it's a false lead.

Stakeholder management under pressure

Handling Difficult Stakeholders Without Losing the Bridge

You're managing three conversations at once: with the technical team, with business stakeholders, and with yourself. Each requires a different register.

The "Just Fix It" Exec

They don't want detail, they want confidence. Give them a clear status, a timeline (even if estimated), and a single point of contact. "We have our best people on it, our current ETA is X, and I'll update you in 30 minutes." That's it. Don't over-explain.

The "Why Wasn't This Prevented?" Interrogator

Acknowledge the frustration, defer the analysis. "That's exactly the right question, and we will address it in the PIR. Right now, my focus is resolution." Never get drawn into a blame conversation during a live incident. It helps no one and distracts everyone.

The "Update Every 2 Minutes" Stakeholder

Set a cadence and hold to it. "I will send updates every 15 minutes unless there's a material change. Unscheduled updates from me mean things are moving." This trains stakeholders, reduces noise, and protects your bandwidth.

The art of saying "no" without saying "no"

Replace "no" with "here's what we're doing instead." When a stakeholder demands something unrealistic, redirect: "I understand the urgency. What we're doing right now is X because that gives us the fastest path to recovery. Here's what I need from you to make that happen."

The pitfalls that kill incidents

Top Ways Major Incidents Go Off the Rails

These aren't hypotheticals. Every experienced MIM has lived through most of these. Recognise them early and call them out.

Too many voices, no clear command

When everyone leads, no one does. If you're on the bridge and there's no clear MIM, volunteer or nominate one immediately. Ambiguous ownership is the fastest way to lose time.

Engineers problem-solving in silence

Silence on the bridge is a red flag. Quiet engineers are either deep in the problem or completely lost. Regular check-ins ("What's your current status?") are not micromanagement, they're your primary sensing mechanism.

Overloading comms channels with noise

Constant updates, multiple threads, people pasting logs into Slack all drown the signal. Keep your incident channel focused. Side channels for technical deep-dives are fine; keep the main thread clean.

Premature conclusions

"We've seen this before, it's the same issue as last Tuesday." Maybe. But acting on that assumption without verification has extended more incidents than it's solved. Validate before you commit.

Lack of clear ownership per workstream

If two people own the same thing, no one owns it. Explicitly assign: "You own database, you own network, I own comms." Repeat it out loud so it's acknowledged.

Status updates that are too vague or too technical

"Teams are working on it" tells stakeholders nothing. "The SQL log files confirm the replica is behind by 40 minutes" tells them too much. Aim for: impact + action + ETA.

Communication as your primary tool

Your Words Are the Incident

People judge the severity and progress of an incident more by your communication than the actual outage. Confidence is contagious, and so is panic.

The three-part update structure

Every stakeholder update should answer: What we know / What we're doing / What we need. This structure prevents waffle, forces clarity, and signals competence even when the situation is messy.

Building confidence with no progress to report

"We don't yet know the root cause, but we have eliminated X and Y and are now focused on Z" is infinitely more reassuring than "still investigating." Show that you're moving, even when you're not moving forward.

Controlling tone to prevent escalation panic

Calm, measured tone in your updates sets the emotional register for the whole incident. If you sound panicked in writing, executives assume the situation is out of control. Choose words that convey urgency without alarm: "We're actively progressing" not "We're scrambling."

Cutting filler and waffle

"As per my previous update, the teams are continuing to look into the ongoing issue with the service which is experiencing degradation..." That is noise. Replace with: "Payment service still down. Engineering focused on DB replication lag. ETA 45 mins." Every word should earn its place.

Herding technical teams without authority

Leading Without Being the Smartest Person in the Room

Your job isn't to understand every system. It is to understand how to move people effectively. Technical authority earns respect; facilitation skills earn results.

Guiding SMEs without micromanaging

Give experts the space to work but set clear time-boxes and check-in points. "I'll come back to you in 15 minutes, so let me know if you need anything before then." This respects their expertise while maintaining momentum.

Challenging engineers constructively

When a team has been pursuing a line of investigation for too long, ask questions rather than issue directives: "Help me understand what evidence would confirm or disprove this theory." It gets them thinking critically without triggering defensiveness.

Keeping momentum when teams go down rabbit holes

Name it early. "I think we might be getting into the weeds here, so let's time-box this 10 minutes and reconvene." Giving permission to pause and re-evaluate is often all people need.

Knowing when to escalate or swap resources

Escalating isn't failure, it's judgement. If a team has been stuck for more than 30 minutes, either bring in additional expertise or explicitly ask: "Do you have everything you need, or should we pull in additional resource?"

Balancing speed vs. stability

Fix Fast or Fix Right?

The pressure to restore service quickly is real, but poorly considered fixes can turn a 1-hour incident into a 6-hour one. This is where your judgement is tested.

Understanding the risk trade-off

Every proposed fix during a live incident should be assessed on two axes: likelihood of success and blast radius if it fails. A high-risk change on an already degraded system needs a much higher confidence bar than a reversible, low-impact one.

When to approve a risky change

If the service is completely down, a risky change is more justifiable than if it's partially degraded. Ask: "What's our fallback if this doesn't work?" If there's no clear answer, the change isn't ready.

The danger of "quick fixes"

A temporary workaround that masks the underlying problem can delay detection of a more serious issue. Approve them, but always flag: "This is a temporary mitigation. We need to confirm root cause before declaring this resolved."

Business pressure vs. technical reality

Executives will sometimes push for risky actions because they're under pressure. Your role is to be honest about the risk, not compliant. "I understand the urgency. Proceeding with this change without testing carries a significant risk of extending the outage. My recommendation is X."

The human side of incidents

Keeping People Effective When Stress Takes Over

You're managing people, not just systems. Stress degrades decision-making, communication, and creativity. Recognising this in others, and in yourself, is a core MIM skill.

Recognising burnout mid-incident

Engineers who've been on a bridge for 3+ hours make worse decisions. Watch for slow responses, repetitive suggestions, or unusual quietness. Where possible, rotate people. Where not, explicitly ask: "Do you need 5 minutes?" Permission to pause often resets performance.

Managing dominant or emotional personalities

The loudest voice isn't always right. When someone is dominating or becoming emotional, name it calmly: "I hear you, so let's make sure we're capturing everything. For now, let's focus on what's in front of us." Redirect, don't suppress.

Keeping quieter engineers engaged

Some of your best contributors will say nothing unless asked. Directly address them: "I haven't heard from the network team in a while, what's your read on this?" Creating explicit space for quieter voices often surfaces insights that would otherwise be lost.

You as the emotional stabiliser

Whether you like it or not, your energy sets the emotional tone. If you're calm, people believe the situation is manageable. If you're frantic, it amplifies everyone's anxiety. This isn't about being emotionless, it's about being deliberate in how you show up.

Post-incident truth vs. blame culture

Running PIRs That Actually Improve Things

A PIR that only produces surface-level fixes is worse than no PIR, because it creates the illusion of improvement. A great PIR is uncomfortable, honest, and actionable.

Creating psychological safety

Before any PIR, set the tone: "This is a blameless review. We're here to improve the system, not evaluate individuals." People will only be honest if they believe honesty is safe.

Extracting real lessons vs. surface-level fixes

"We'll add more monitoring" appears on almost every PIR. Push deeper: "Why wasn't this caught by existing monitoring? What assumption were we making that turned out to be wrong?" Real lessons change process or architecture, not just tooling.

Turning insights into actionable change

Every action item needs an owner, a due date, and a review mechanism. "Team to investigate X" is not an action item. "Jane to complete runbook for DB failover by end of sprint, reviewed at next ops meeting" is.

Handling defensive teams

When teams become defensive, it's usually because they feel blamed. Redirect to the system: "I'm not asking why you made that decision. I'm asking what information was available to you at the time, and what would need to change for a different outcome." It shifts the frame entirely.

The illusion of control

You're Not in Control, You're Creating Control

The best MIMs don't try to control the incident. They create the conditions in which control becomes possible.

Accepting chaos as the default

Incidents are chaotic by nature. Accepting this stops you from wasting energy trying to eliminate uncertainty and lets you focus on reducing it incrementally. Your goal isn't certainty, it is enough clarity to take the next step.

Creating structure through cadence, roles, and clarity

Structure is created through rhythm: regular check-ins, clear role assignments, consistent update formats. These signals tell everyone that someone is in command, even when the situation is unclear.

Why overconfidence is dangerous

"We know what this is" is one of the most dangerous phrases in incident management. Premature certainty closes off investigative paths and creates false confidence in stakeholders. Stay curious. Stay open.

The MIM as signal, not noise

Your presence should reduce confusion, not add to it. Every time you speak on the bridge, you should be adding structure, clarity, or direction. If you're just thinking out loud, do it off the bridge.