Summer SaleSee pricing

Incident updates people actually read

During an outage, the incident channel fills with questions the last update should have answered. Here is a communication structure that keeps stakeholders informed and off the bridge.

Updated 14 July 20263 min read

Every major incident has two workstreams. There's the technical one, where people are actually fixing the thing. And there's the communication one, where everyone else decides how much to panic.

Most teams staff the first and improvise the second. The result is familiar: engineers get pulled off the fix to answer "any update?" pings, a director hears about the outage from a customer instead of from IT, and the incident channel becomes an archaeology site where the current status is buried under forty messages. The outage was 90 minutes. The reputational damage came from the silence, not the downtime.

Good incident communication isn't a talent. It's a structure and a rhythm, and both are learnable.

The rhythm: updates on a clock, not on progress

The single most effective change you can make is committing to a cadence. Every 30 minutes for a severity 1, every hour for a severity 2, adjusted to your world, but fixed and stated.

The reason is psychological, not informational. What stakeholders can't tolerate isn't bad news, it's not knowing when they'll hear next. "Next update at 14:30" buys you thirty minutes of silence in which nobody pings the bridge. And the discipline holds even when there's nothing new: "no change since the last update, still working the same theory, next update at 15:00" is a perfectly good update. It proves the process is alive, which is exactly what the reader is checking.

The structure: four lines, always the same

Every update answers four things, in the same order every time.

What's happening: the impact in plain terms, from the user's side. "Email delivery is delayed up to 40 minutes for all customers", not "the MTA cluster is degraded."

What we know: the current best understanding, honestly labelled. "We believe this is linked to this morning's storage change" is fine. Guessing dressed as certainty is how you end up walking back statements later, and walked-back statements are what people remember.

What we're doing: the active workstream, one or two lines. This is where trust is built, because it shows motion without requiring the reader to understand the technology.

What's next: the next update time, and if you have one, a realistic expectation. Resist the pressure to promise a fix time you don't have. "We expect to know within the hour whether the rollback resolves it" is honest and useful.

Four lines. When the structure never changes, readers stop reading defensively and start scanning, and a scanning reader doesn't ping the bridge.

One incident, three audiences

The classic failure is writing one update and sending it everywhere. The bridge needs system names and hypotheses. The executive needs impact, scope, and trajectory in business terms. The customer needs what they'll experience and what, if anything, they should do, in language that doesn't assume they know what a failover is.

Same facts, three renderings. This used to be the expensive part, which is why it rarely happened under pressure. It's also the part AI is genuinely good at: give an assistant the technical bridge update and have it produce the executive and customer versions against fixed rules for tone, length, and what never gets mentioned externally. The incident lead reviews, sends, done. Our incident pack in the Opstimio library includes the prompt and templates wired this way, one truth in, three audiences out, because rewriting under pressure at 2 a.m. is precisely when unforced errors happen.

Words that cause trouble

A few small rules prevent most communication damage. Don't name vendors or blame in live updates; there's a review for that later, with cooler heads. Don't say "resolved" until it's verified from the user's side; "we're seeing recovery and are monitoring" survives a relapse, "resolved" doesn't. Don't use "intermittent" as a synonym for "we don't know"; readers can tell. And numbers beat adjectives everywhere: "affecting roughly 200 users at the Utrecht site" informs, "widespread issues" alarms.

The quiet payoff

Teams that fix incident communication notice something odd within a quarter: their incidents feel smaller, even when the technical facts are the same. Stakeholders who trust the updates stop escalating around them. Engineers work uninterrupted. The post-incident review starts from a clean timeline instead of channel archaeology.

Write the four-line structure down, set the cadence per severity, and put one person on comms who is explicitly not also fixing the thing. That's the whole intervention. The next incident will still be stressful. It just won't be lonely at the top of the org chart, wondering.

Free, no account needed

Put this into practice today

Reading is the easy part. Start with a free tool: grab the sample pack of ready-to-use prompts, or take the two-minute baseline to see where your operation should start.

Ready-to-use tools for this

2 in the library

Locked previews. The article teaches the approach; these are the ready-made tools that do the work.