# Incidents & escalation

> **For agents:** `list_incidents(org="acme", status="triggered")` is the list that matters during an outage — nobody has picked those up yet. `get_incident(org="acme", id="01J…")` returns the full timeline: every escalation step, who was paged on which channel, who acknowledged. Writes — `trigger_incident`, `acknowledge_incident`, `resolve_incident`, and the policy/service tools — need `alert:write`.

Incidents are opened by an uptime monitor whose `serviceId` is set (see [Uptime](https://docs.bugwatch.io/product/uptime.md)), by `trigger_incident`, or by the REST route below.

An **incident** is one firing of one problem. An **escalation policy** decides who hears about it and how insistently. A **service** binds the two together.

## The chain

A policy is an ordered list of rules:

| Field | Meaning |
|---|---|
| `position` | Rule `0` fires the moment the incident opens — not after a delay. |
| `delaySeconds` | How long to wait before moving to the **next** rule. |
| `urgency` | `high` pages; `low` notifies quietly and never rings a phone. |
| `targets` | An on-call `schedule` (resolved to whoever is on call right now) or a `user` directly. |

When the chain runs out and nobody has acknowledged, the policy can `repeatCount` the whole thing again after `repeatDelaySeconds`. When that is exhausted too, the incident timeline records it — the difference between "nobody was paged" and "everybody was paged and nobody answered" is one worth being able to see afterwards.

**Set `fallbackUserId`.** When a rule targets a schedule that covers nobody, the fallback is what stops the page evaporating. It is the single most common way an incident goes unanswered, and it is why [coverage gaps](https://docs.bugwatch.io/product/on-call.md) are drawn in red.

## What acknowledging does

Acknowledging stops the escalation **immediately** — every pending step, including the notification ladder already queued for the person who acked.

If the service sets `ackTimeoutSeconds` and nobody resolves the incident in that time, it **re-escalates from the top**. This is deliberate. Acking is how you say "I have this"; without the timer it becomes how you make the noise stop, and an incident that was acknowledged and then forgotten is indistinguishable from one nobody ever saw.

`autoResolveSeconds` on the service closes an incident that nothing resolved — useful for sources that never send a recovery signal.

**Resolved is final.** Acknowledging or resolving an incident that is already resolved, from a stale tab, a second responder or an agent working from an old list, answers `{"status": "resolved"}` and changes nothing: not the record, not who resolved it, not the timeline, not the status page. A second acknowledge on an acknowledged incident keeps the first responder.

## Deduplication

An incident is identified by `(service, dedupKey)` while it is open. Triggering with the same key joins the existing incident and appends to its timeline instead of opening a second one and paging the rotation twice. Once an incident is **resolved**, the same key opens a fresh one — a partial unique index in the control database enforces exactly that.

## One timer

All of the above lives in one Durable Object per incident, holding one action queue and one alarm set to the nearest action. That is what makes acknowledging exact: the ack cancels pending escalations in the same transition that sets the status, so there is no window in which a cancelled page still goes out.

Durable Object alarms are at-least-once, so every notification carries a deterministic id (`incident:rule:user:channel:attempt`). A replayed alarm produces the same ids and delivery collapses them. Paging someone twice is survivable; paging them twelve times because an alarm retried is not.

## Timeline

`get_incident` returns every event: `triggered`, `escalated` (with the rule and the users it resolved to), `notified` (with the channel, whether it came from the fallback, and **whether it was actually delivered**), `acknowledged`, `unacknowledged` (an ack that timed out), `repeated`, `exhausted`, `resolved`. A `no_policy` entry means the service has no escalation policy attached — the misconfiguration that otherwise looks exactly like a delivery failure.

The timeline is kept for **90 days**, the same as monitor history and for the same reason: it is a timeline, not a ledger. A retro reaches back weeks, not months, and when the blow-by-blow is swept the incident itself keeps the summary the history list reads — when it opened, who acknowledged, when it resolved.

A `notified` entry with `delivered: false` means the chain reached that person's turn and had nowhere to send. The step is still recorded — losing it would lose why the incident escalated past them — but it is not a page, and the dashboard says so in red rather than showing it as one. `undeliveredReason` names which kind of nowhere:

| Reason | What it means | What fixes it |
|---|---|---|
| `no_channels` | The service has no notification channels configured, and the responder has no registered device | Add a channel to the service, or register a device |
| `channels_missing` | The service names channels, and **none of them still exist** | Attach a channel that exists — the ids it holds are stale |
| `no_channel_for_medium` | The service has channels, just none of this rung's kind | Add a channel of that kind, or change the ladder |

`channels_missing` is its own reason rather than being folded into the last one: a service whose only channel was deleted used to report "none of this kind", which sends the responder to add a channel they already had.

## In the dashboard

**Incidents** in the sidebar (or `G` `N`) is the list: what is unacknowledged first, then everything open, then history. The hero counts what nobody has picked up and how long the oldest of those has been waiting; **median ack** is measured over the incidents on screen.

Acknowledge and Resolve are on each open row and on the incident page. If the escalation refuses the transition the dashboard says so rather than showing a tick — an ack the DO did not accept means the chain is still paging people.

The incident page is its timeline, written as sentences: who was paged, on which channel, whether the page came from a schedule or from the policy's fallback, when somebody answered. Two states that look identical in a status field are kept apart there: *nobody was paged* (`no_policy` — the service has no policy; or a step that says **"Would have paged Rae"** because there was nowhere to send it) and *everybody was paged and nobody answered* (`exhausted` after steps that really went out). Below the list, **Services** and **Escalation policies** show what would happen the next time something fires: a service with no policy is called out in red, and so is a policy with no fallback, because an uncovered rotation under it pages nobody.

People are shown by name. Services, policies and their steps are created over the API or with the agent tools (`create_service`, `create_escalation_policy`) — the dashboard reads them, and an editor is a later slice.

## REST routes

| Route | Scope |
|---|---|
| `GET/POST /v1/orgs/{org}/escalation-policies` | `org:read` / `alert:write` |
| `GET/PATCH/DELETE /v1/orgs/{org}/escalation-policies/{id}` | `org:read` / `alert:write` |
| `GET/POST /v1/orgs/{org}/services` | `org:read` / `alert:write` |
| `PATCH/DELETE /v1/orgs/{org}/services/{id}` | `alert:write` |
| `GET/POST /v1/orgs/{org}/incidents` | `org:read` / `alert:write` |
| `GET /v1/orgs/{org}/incidents/{id}` | `org:read` |
| `POST /v1/orgs/{org}/incidents/{id}/acknowledge` | `alert:write` |
| `POST /v1/orgs/{org}/incidents/{id}/resolve` | `alert:write` |

`PATCH` with `rules` replaces the whole chain — positions and targets only mean anything as an ordered set.

## Where a page goes

An escalation step aimed at a person notifies **the service's channels and that person's own phones**. The phones are a `mobile_push` channel, created when they register a device ([Mobile notifications](https://docs.bugwatch.io/product/mobile-notifications.md)) — so a responder who has the app is paged even by a service that was never wired to Slack, and a new handset is reachable tonight with no change to any policy.

The message on the service's shared channels still names who the step was for: those channels are read by everybody, and an unnamed page is an anonymous one.

## The ladder

An escalation step is not one notification — it is a **ladder** of rungs, each aimed at one medium:

| Urgency | Rungs |
|---|---|
| `high` | the person's **push** at once, and the service's **Slack**, **PagerDuty** and **webhook** channels at once; **email** after two minutes |
| `low` | the service's **email** and **Slack**, at once. Nothing here pages. |

Each rung goes **only** to channels of its own medium: `push` means that person's phones and nothing else, and every other rung means the service's channels of that kind. A rung whose medium the service has not configured is recorded as `delivered: false` with `undeliveredReason: "no_channel_for_medium"` rather than quietly going out over whatever else is to hand.

Email is the laggard rung because it is the "you still have not seen this" medium — and acknowledging inside two minutes cancels it before it is sent, along with every other pending step.

**Each person can set their own ladder** (`ladder` in [notification preferences](https://docs.bugwatch.io/product/mobile-notifications.md)): which media reach them, in what order, with what delays. One thing is not theirs to change — a high-urgency ladder's first rung is always immediate. A ladder that starts at +5 minutes is a page that arrives five minutes late, every time, silently, and that is the one setting a person must not be able to configure into their own pager.

## Not yet
- **Issue alert routing.** Alert rules still notify their own channels directly; routing them onto services (as monitors now are) is the next slice.
- **PagerDuty Events API v2 ingest**, event rules, maintenance windows and dependency suppression.
- **Editing services and policies in the dashboard.** They are read-only there; use the API or the agent tools to create and change them.

## What else changed (incident correlation)

`GET /v1/orgs/{org}/incidents/{id}/context` — or `get_incident_context` — answers the first question a responder asks after being paged: *what else moved around this?*

It compares the window before the incident opened with the window since, and lists what rose:

- **Error rate** for the service's project, if it has one.
- **Probe p95 latency** for the monitor that triggered it. Buckets where nothing passed are excluded rather than read as 0 ms — a failed check has no latency to report.
- **Deploys** in the two hours before it opened, with the commit.

Three rules shape it, and they are worth stating because each is a thing it deliberately does *not* do:

1. **It never delays or replaces the page.** The correlation runs when somebody asks for it, never on the paging path. A query between a monitor going down and a phone ringing is how a pager stops being trusted.
2. **It cannot act.** No acknowledge, no resolve, no suppress — there is no mutation in the endpoint at all.
3. **It says when it could not look.** The response carries a `checked` object, so "nothing stood out" is distinguishable from "we had nothing to check with". A reader who suspects a check was skipped will redo it by hand.

Findings are ranked with measured signals ahead of deploys — a metric that moved is evidence, a deploy that happened is a coincidence until somebody checks it — and a deploy is never called a strong lead on timing alone. Only increases are reported: a metric that *fell* during an outage is usually traffic draining away, which is a symptom rather than a lead.

Correlation is available for the first 72 hours of an incident. Past that it is a retro, not a triage.

## In the dashboard

An incident's page carries an **Incident agent** panel with all four drafts, and it offers only the ones that can say something true right now:

| While it is open | Once it is resolved |
|---|---|
| **Catch me up** — the running brief | — |
| **What else changed** — correlation | **What else changed** |
| **Customer update** — the status-page draft | **Customer update** |
| — | **Retro draft** |

A retro is not offered on an open incident: drafted from half a timeline, it is missing the half people most want to argue about. "Catch me up" is not offered once it is resolved, because there is nobody to catch up.

Every one is **copy, never publish**. These are drafts, and sending a customer-facing or written-down artifact stays a person's decision — the same rule that keeps `NON_TOOL_ROUTES` human-initiated.

## The running brief

`GET /v1/orgs/{org}/incidents/{id}/brief` — or `get_incident_brief` — is the only one of these written for an incident that is **still happening**.

It answers the question a person actually asks in the ten seconds after their phone rings: *is somebody already on this, or is it me?* So it is ordered for that reader rather than chronologically — chronological order is how the important thing ends up in the middle.

1. **What is wrong with the response**, above everything. A service with no escalation policy means nothing is paging anybody and this will not escalate on its own. A source that fired again while the incident was open means whatever was done has not held.
2. **Who has it** — acknowledged and by whom, or paged and unanswered, or nobody reached; plus when the chain escalates next and to whom.
3. **What has already been tried**, so the reader does not repeat it. Never blank: "nothing yet" is stated, because an empty list reads as *not loaded* and the reader waits for something that is not coming.
4. **The history**, newest first, if they want it.

The distinction it works hardest to preserve is the same one the retro makes, and it changes what the reader should do:

| State | What it means | What the reader does |
|---|---|---|
| `unacknowledged` | Pages were delivered; nobody has answered | Somebody has been told. You may still be second |
| `not_reaching_anyone` | Pages went out and **none** were delivered | Nobody has been told. Assume it is yours |
| `nobody_paged` | No policy, or nothing to send to | The chain will never produce a responder |

`not_reaching_anyone` and `nobody_paged` outrank `unacknowledged` deliberately: waiting is a reasonable plan only when somebody is going to arrive.

Read-only, and off the paging path by construction like the rest of the incident agent.

## Retro drafts

`GET /v1/orgs/{org}/incidents/{id}/retro` — or `draft_incident_retro` — assembles a retro from the incident's own timeline.

It measures what a retro written the next morning gets wrong, because those facts were recorded at the time and memory is not:

- **How long each stage took** — detection to acknowledgement, acknowledgement to resolution. "Never acknowledged" and "still open" are stated, not left blank.
- **What the response itself got wrong**, above the timeline rather than buried under it: a page that was never delivered and why, a chain that fell through to its fallback, a source that fired again while the incident was open, a service with no escalation policy, and pages that were delivered and ignored.

The last distinction matters: *nobody acknowledged* and *nothing was ever delivered* are different failures with different fixes, and the draft never reports the first when the second is true — saying somebody ignored a page that never arrived sends the retro after the wrong people.

**It does not write the analysis.** The draft ends with questions to fill in, and says so in the document. An autofilled conclusion is one nobody argues with, and the arguing is the point of a retro.

Read-only: it records nothing and changes nothing about the incident.

## Status-page update drafts

`GET /v1/orgs/{org}/incidents/{id}/status-draft` — or `draft_status_update` — drafts the update your customers read.

It is the retro's public sibling and it is written to a different standard, because the audience is different. A retro is read by the people who were there and can correct it. A status update is quoted back, screenshotted, and cannot be unsaid.

**The stage is derived, not chosen.** `investigating` until somebody acknowledges, `identified` after, `resolved` when the incident is. Only `monitoring` can be asked for, and only once the incident is acknowledged — an operator who can pick the stage is an operator who can pick `resolved` during an outage.

**It refuses an all-clear the data does not support.** A `resolved` draft is downgraded to `monitoring` while any affected component is still reporting `down` or `degraded` on the published page — checked against what the public can currently see, not against the incident record, because the incident closing and the monitors recovering are two different facts. A component that is merely `unknown` does not hold a genuine all-clear hostage.

**What it will not say:**

| Never | Why |
|---|---|
| A cause | Not even a likely one, and not even when correlation found a strong lead. Correlation is for the responder; a cause published in the first ten minutes is a cause retracted later |
| Anything internal | No service, monitor or incident ids, hostnames, stack frames, customer names or responders. Components are named in the words your status page already uses, because that is what a reader recognises |
| Nothing about the next update | Every unresolved draft promises one. A status update with no stated next update turns every reader into a support ticket asking when the next one is |

Each draft carries an **`omissions`** list — what it deliberately left blank, in words. That is not decoration: an operator who cannot see *why* a draft is thin will fill the gap with a guess, which is the thing this exists to avoid. "No cause is stated because none is confirmed" turns an omission into a decision.

Read-only. Publishing an update is a human action; nothing here writes to a status page.
