# Status pages

> **For agents:** `list_status_pages(org="acme")` for the pages, `get_status_page(org="acme", id="01J…")` for one with its components. Writes — `create_status_page`, `update_status_page`, `delete_status_page` — need `alert:write`. Maintenance: `list_maintenance`, `schedule_maintenance`, `cancel_maintenance`. Incidents: `list_status_incidents` for what the page is showing, `post_status_update` to publish an update about one (`alert:write`, audited, and public immediately). Custom domains: `list_status_page_domains`, `add_status_page_domain`, `verify_status_page_domain`, `remove_status_page_domain`. Example: `create_status_page(org="acme", slug="acme", name="Acme Status", published=true, components=[{name:"Checkout", sourceKind:"monitor", sourceId:"01J…"}])`.

A status page is what your customers read when something is wrong. It shows the health of the monitors and services you choose, under names you pick.



## Subscribers

People outside your company can ask to hear when a page changes status — by email, or by webhook for a team that wants it in their own system.

**Email is double opt-in.** Anyone can type anyone's address into a public form, so the row is inert until the person holding that inbox clicks the confirmation link. Until then nothing is sent to them, ever. Submitting the form twice creates one row and one confirmation, which is also what stops the form being used to mail-bomb an address.

**A webhook proves it asked.** On subscribe, the URL receives a `POST` of `{"type": "bugwatch.status.subscription", "page": "<slug>", "challenge": "spch_…"}` and must answer `2xx` with the challenge, either as the whole body or as `{"challenge": "spch_…"}`, within 5 seconds. Otherwise the subscription is refused with `400` and nothing is stored. Without this step anybody could aim a page's notices at any URL. A webhook must also be `https://`, because announcing an outage over cleartext announces it to the network.

**Limits.** A page takes at most 100 webhook subscribers and 1,000 unconfirmed email subscriptions a day. One address is sent at most 3 confirmation emails an hour, across every page. Past those limits the form keeps answering the same flat `ok`, but nothing is stored or sent. The form is also rate-limited per IP.

**Delivery is per subscriber.** If some subscribers fail (a webhook down, a mail provider hiccup), the retry goes only to them. Nobody gets the same notice twice because somebody else's endpoint was broken.

Every notice carries a one-click unsubscribe. It is a capability — no account, no login — because a subscription somebody cannot leave is one they report as spam, and a status page in spam folders is a status page nobody receives.

### What is and is not an email

The page republishes on **every** health change: a monitor that flaps twice a minute republishes twice a minute. A subscriber list wired naively to that is a mailing list nobody stays on. So a notice goes out only when the page's **overall status changes**, and three rules follow from it:

- **A component moving while the summary holds is not an email.** A second service degrading during an outage is not a second announcement.
- **Nothing is sent on the first publish.** A page has no previous status the first time anybody looks, and "operational" arriving out of nowhere is not news.
- **`unknown` is never an incident.** A page whose monitors have not reported yet, or whose only component is paused, is not down — and mailing everybody that it is would be the status page raising a false alarm about itself. Moving *out* of `unknown` is not a recovery either; we simply started knowing.

### Where this lives, and why not on the status page

Subscribing, confirming and unsubscribing are handled by the **API**, not by the status Worker.

That Worker binds exactly one thing — the KV namespace holding pre-built snapshots — so the page stays up during precisely the outage it is reporting. Putting a form that writes to the database on it would give it a dependency that can be down at the worst moment. The page links out instead: the subscription surface may be unavailable while the page itself is not, which is the right way round.

## The uptime strip

Each monitor-backed component carries **ninety days** of per-day history, oldest first, with the uptime figure for the window beside it.

Four things about it are deliberate, and each is a way the strip could flatter the service:

- **A day with no events is not a green day.** `monitor_events` stores state *changes*, so a three-day outage is two rows. The strip carries the state forward across days that recorded nothing, rather than counting rows per day and painting two of the three green.
- **Before the monitor existed is `unknown`, never `operational`.** A page whose first render shows ninety green days is making a claim about a service nobody has watched yet. The same rule the component status already follows for a paused monitor.
- **Ninety days is the whole window, because that is what is retained.** `monitor_events` is swept at 90 days, so asking for a longer history would not produce more of it — it would produce confident history derived from an absence.
- **A day is ranked by its worst state, not its average.** Twenty-three good hours and one outage is a down day; the percentage in the tooltip says 95.83%, and the cell is red.

A **service**-backed component has no strip. Its status comes from open incidents, which are not a per-day signal, and an empty strip would read as ninety unknown days about something the strip does not apply to.

Days are UTC. A status page has no reader's timezone to use, and picking the organization's would make two readers disagree about which day an outage fell on.

## Why it stays up

The page a customer loads is served by a **separate Worker with one binding**: a KV namespace holding a pre-built snapshot. No D1, no Durable Objects, no service binding back to the API. It renders what it finds and never computes a status.

That is the whole design, and everything else follows from it. A status page that has to ask the control plane how things are going goes dark at exactly the moment somebody is looking for it — which is the moment it exists for. So the snapshot is built on **write**, not on read:

- editing the page rebuilds it,
- a monitor changing state rebuilds every published page that shows that monitor,
- an incident opening or resolving rebuilds every published page that shows its service.

A snapshot the Worker cannot parse returns **503**, not a green page. On a status page, failing loudly beats failing reassuringly.

## Components

A component is a **public name for something internal**. Readers see `Checkout`; they never see a monitor id, and no id appears in the served page.

| Field | Meaning |
|---|---|
| `name` | What readers see, e.g. `Checkout` |
| `description` | Optional second line under the name |
| `sourceKind` | `monitor` or `service` |
| `sourceId` | The monitor or service whose health this component reports |

Components are an **ordered list** and the order is the display order, so `update_status_page` replaces the whole list rather than patching entries. A source must belong to the same organization as the page.

A page can carry up to 40 components. Past that it has stopped being a status page and become a dashboard.

## What each colour means

| Shown | When |
|---|---|
| Operational | The monitor is `up`, or the service has no open incident |
| Degraded | The monitor is `degraded`, or the service's open incidents are all `sev3`/`sev4` |
| Major outage | The monitor is `down`, or the service has an open `sev1`/`sev2` incident |
| Maintenance | A maintenance window covers a component that is otherwise fine |
| Unknown | The monitor is paused, disabled, or has not reported yet |

Two of those rows are deliberate refusals:

- **A paused or disabled monitor is `Unknown`, not `Operational`.** A probe nobody is running proves nothing, however green it was when somebody switched it off.
- **Maintenance never masks a real fault.** A component that is genuinely down during a window is shown as down. Saying "scheduled maintenance" over an unplanned outage is how a status page stops being believed.

The page's own status is its **worst** component, and a page with no components reads `Unknown` rather than a green light nobody earned.

## Incidents and updates

Incidents on anything the page shows are listed above the components, along with everything published about each one.

`POST /v1/orgs/{org}/status-pages/{id}/incidents/{incidentId}/updates` (tool: `post_status_update`) publishes one: a stage and what you want readers to know, in your words. It is public the moment it is written — the page rebuilds on the write rather than waiting for the next health change.

| Stage | What it tells a reader |
|---|---|
| `investigating` | We know something is wrong and are looking |
| `identified` | We know what it is |
| `monitoring` | A fix is out and we are watching it |
| `resolved` | It is over |

**The stage you publish is what the page shows.** Without an update the page derives one — `investigating` until somebody acknowledges the page, then `identified` — but that derived value answers a question about *us*, not about the work. An incident whose fix is already deployed is still `acknowledged` internally, and `monitoring` has no internal equivalent at all.

Two things override that, in opposite directions:

- **A closed incident always reads `resolved`**, whatever was published last. Closure is a fact about the system rather than a claim about what happened, and a page still saying "investigating" under a closed incident is worse than one that says `resolved` with nobody's words under it.
- **Publishing `resolved` does not close anything.** If the incident is still open — reopened, or the all-clear went out early — the page keeps showing the real stage. An update is a description, not a control.

Updates are separate from an incident's internal timeline. That timeline records who was paged, who acknowledged, and what the escalation chain did; none of it reaches a reader, and publishing is always something a person or an agent chose to do.

`list_status_incidents` shows what the page is showing right now, updates included. It reads the page's own assembly rather than a separate query, so it cannot disagree with what readers see.

**Resolved incidents stay on the page for 7 days**, with their updates, and stop colouring any component the moment they close. "Was there an outage this morning?" is most of what a status page is asked after the fact, and it is the one question the page can answer with nobody awake — dropping a resolved incident on the spot also meant the all-clear could never be read.

## Scheduled maintenance

`POST /v1/orgs/{org}/status-pages/{id}/maintenance` (tool: `schedule_maintenance`) puts a window on the page: a title, an optional body, `startsAt` and `endsAt` in epoch ms, and the component ids it covers. **Omit the components and it covers the whole page**, which is what "we are doing scheduled maintenance" usually means. `list_maintenance` shows them, past ones included; `cancel_maintenance` removes one and the page stops showing it at once.

### A window also stops the paging

An open window **suppresses pages** for anything on the components it covers — the monitors those components point at, and the monitors that escalate through a covered service. Nobody is woken for work you scheduled.

What it does *not* do is change what is true:

- The monitor still goes `down`, so consensus and recovery keep working.
- The transition still lands in the monitor's timeline, marked as suppressed and naming the window. A withheld page is visible afterwards rather than looking like one that went missing.
- The status page still shows the component as down if it is genuinely down. Maintenance never masks a real fault, here as everywhere else on the page.

Three rules decide the edges, and each exists because the obvious alternative is worse:

- **A monitor that is still down when the window closes pages then**, saying how long the page was withheld. Otherwise scheduling a window over a monitor would silence it permanently — the monitor never transitions again, so nothing would ever look — which is worse than having no windows at all.
- **Recovery is silent only if the outage was.** If the DOWN was withheld, nobody is waiting for an all-clear and sending one is noise. If people *were* paged — the outage started before the window did — they get the recovery even mid-window, because the alternative is a page left open forever.
- **A window that cannot be read suppresses nothing.** Every failure here resolves toward paging. Silence is the expensive direction, and no corrupt row or database hiccup is allowed to buy it.

Cancelling a window (`cancel_maintenance`) restores paging immediately.

A component id that is not on the page is refused, not ignored — a window that silently covers less than you asked for is one you find out about from a reader.

**The window is decided when the page is read, not when it is published.** That follows from the one thing this page exists to guarantee: it is served by a Worker that reads a single pre-built snapshot and no database. If a window starting at 02:00 needed the control plane to rebuild the page at 02:00, the status page would depend on exactly the system it is supposed to outlive. So the *schedule* travels inside the snapshot, and the reader's clock resolves it — the window opens and closes on time even if the API has been down for a week, and cancelling one takes effect on the next read.

Two consequences worth knowing:

- A window is visible on the page from the moment you create it, as **upcoming**, and changes no component's colour until it starts. That is deliberate: "there is planned work on Saturday" is part of what a status page is for.
- **Maintenance never masks a real fault** — the rule above applies here too. A component that is genuinely down during its own maintenance window is shown as down.

## Your own domain

A page is served at `status.bugwatch.io/{slug}` until you point a hostname of your own at it. After that, `https://status.acme.com/` **is** the page.

Add the domain, add two or three DNS records, and wait. Nothing else happens — the certificate is ordered, issued and renewed for you, and the domain goes live on its own once your records resolve. It is re-checked every few minutes, so you do not have to come back and press anything.

```bash
curl -X POST https://api.bugwatch.io/v1/orgs/acme/status-pages/01J.../domains \
  -H "authorization: Bearer $BUGWATCH_TOKEN" \
  -H "content-type: application/json" \
  -d '{"hostname": "status.acme.com"}'
```

The response carries the records to add:

| Record | Why |
|---|---|
| `CNAME status.acme.com` → your fallback origin | Sends the traffic to us |
| `TXT _cf-custom-hostname.status.acme.com` | Proves you control the name |
| `TXT _acme-challenge.status.acme.com` | Lets the certificate authority issue for it |

Validation is by TXT record rather than by HTTP. An HTTP check would need the hostname to already point at us before the certificate exists — and a status page is exactly the thing people aim at us while their own origin is down.

### The three states

| State | What it means |
|---|---|
| `pending` | Waiting on DNS, or on the certificate. Normal for the first few minutes |
| `active` | Serving. The hostname resolves here **and** the certificate is live |
| `failed` | Something needs you: the hostname was blocked or moved, or a step timed out. `lastError` carries the certificate authority's own words |

`active` deliberately requires both halves. A hostname that resolves to us before its certificate exists shows every visitor a TLS warning, which reads as your status page being broken — so it is not published until both are ready. The same rule runs backwards: if a live domain later fails or its certificate expires, it stops being served rather than answering over a certificate that no longer validates.

### What a custom domain does not do

**It serves one page.** `status.acme.com/some-other-slug` is a 404, including the JSON view. Without that rule, pointing a hostname at us would quietly turn it into a public mirror of every status page we host, under your brand.

**It does not add a dependency.** The Worker serving your page still reads pre-built snapshots from one key-value store and talks to no database — a custom domain is one extra lookup in that same store, written ahead of time. This is the whole reason the page stays up during an outage, and it is not traded away for a nicer URL.

**Hostnames are claimed once.** A hostname belongs to one page across the whole service. No wildcards: `*.acme.com` would map every subdomain you have, including ones you add later and never think about, onto a public page.

A page may hold up to five domains, which is enough for a rename with an overlap.

Removing a domain stops it serving immediately and releases the certificate.

## Publishing

A page is created unpublished. Nothing is served until `published` is true, so a half-built page cannot be found by guessing its slug.

Turning `published` off **deletes the snapshot** — the public URL stops answering, it does not merely stop updating. Renaming the slug retires the old URL in the same way, so a rename is a rename rather than a fork.

Slugs are globally unique across Bugwatch because they are public URLs; a taken slug returns `409` without saying who has it.

## The public views

| Path | Returns |
|---|---|
| `/{slug}` | The page: one self-contained HTML document, no scripts, no fonts, no external requests |
| `/{slug}.json` | The same snapshot as JSON, for anything polling it |
| `/` | On a custom domain, that domain's page |
| `/index.json` | On a custom domain, that domain's snapshot as JSON |

Both are cached for 15 seconds. A status page is polled hard during an outage, and a cache that holds for a few seconds is what keeps it answering.

## REST

| Method | Path | Scope |
|---|---|---|
| `GET` | `/v1/orgs/{org}/status-pages` | `org:read` |
| `POST` | `/v1/orgs/{org}/status-pages` | `alert:write` |
| `GET` | `/v1/orgs/{org}/status-pages/{id}` | `org:read` |
| `PATCH` | `/v1/orgs/{org}/status-pages/{id}` | `alert:write` |
| `DELETE` | `/v1/orgs/{org}/status-pages/{id}` | `alert:write` |
| `GET` | `/v1/orgs/{org}/status-pages/{id}/domains` | `org:read` |
| `POST` | `/v1/orgs/{org}/status-pages/{id}/domains` | `alert:write` |
| `POST` | `/v1/orgs/{org}/status-pages/{id}/domains/{domainId}/verify` | `alert:write` |
| `DELETE` | `/v1/orgs/{org}/status-pages/{id}/domains/{domainId}` | `alert:write` |

```bash
curl -X POST https://api.bugwatch.io/v1/orgs/acme/status-pages \
  -H "authorization: Bearer $BUGWATCH_TOKEN" \
  -H "content-type: application/json" \
  -d '{
    "slug": "acme",
    "name": "Acme Status",
    "headline": "How Acme is doing today",
    "published": true,
    "components": [
      { "name": "Website",  "sourceKind": "monitor", "sourceId": "01J..." },
      { "name": "Payments", "sourceKind": "service", "sourceId": "01J..." }
    ]
  }'
```

## See also

- [Uptime monitors](https://docs.bugwatch.io/product/uptime.md) — what the `monitor` components read from
- [Incidents & escalation](https://docs.bugwatch.io/product/incidents.md) — what the `service` components read from
