Skip to content
Glowo
Esc
navigateopen⌘Jpreview
On this page

Incidents

Tell users what is happening, keep them updated, and leave a record when it is over.

An incident is the human narrative over the top of the machine data. Monitors say a component is red; an incident says what happened, who it affects, and what you are doing.

Statuses

Incidents move through four states:

Status Say this when
INVESTIGATING You know something is wrong, not yet why
IDENTIFIED You know the cause
MONITORING A fix is deployed, you are watching it hold
RESOLVED It is over

Move through them honestly. Jumping to IDENTIFIED before you know, and then walking it back, costs more trust than saying “investigating” for longer.

Severity

Severity Meaning
NONE Informational
MINOR Limited impact, most users unaffected
MAJOR Significant impact
CRITICAL Service is unusable for most or all users

Visibility

Visibility Behaviour
DRAFT Not public. Write it before you publish it.
PUBLISHED Live on the status page
SCHEDULED Publishes automatically at scheduledFor

SCHEDULED is for planned announcements. For planned downtime specifically, a maintenance window is the better tool.

Updates

An incident is a timeline. Each update adds an entry with its own text and can change the status.

Post an update when something changes, and post one even when nothing has — “still investigating, next update in 30 minutes” is far better than silence, because silence reads as “nobody is working on this”.

Updates can be edited after posting, which is recorded.

Auto-incidents

With autoIncident enabled, a status page opens an incident automatically when a monitor goes down, so your page reflects reality even at 3am before a human arrives.

A database constraint guarantees only one open auto-incident per monitor and page, so a flapping monitor cannot fill your page with duplicates.

Treat the auto-incident as a placeholder and add a human update as soon as you can. “Our monitoring detected an issue” tells users less than one sentence of explanation.

Resolving, reopening, archiving

  • Resolve when it is genuinely over.
  • Reopen if it comes back — better than opening a second incident, because the timeline stays in one place.
  • Archive to remove an incident from the default view while keeping it. An incident must be archived before it can be deleted.

Both resolve and reopen support an undo, for the moment you resolve the wrong incident.

Writing well under pressure

  • Lead with impact. “Sign-in is failing for all users” before “we are seeing elevated error rates in the auth service”.
  • Say who is affected, and say when only some people are.
  • Give a next-update time and meet it.
  • Avoid blame and internal jargon. Nobody outside needs the name of the service that failed.
  • Do not promise a cause you have not confirmed.

Afterwards, consider a postmortem.

Was this page helpful?