bugwatch docs

Status pages

For agents: list_status_pages(org="acme") for the pages, get_status_page(org="acme", id="01J…") for one with its components. Writes — create_status_page, update_status_page, delete_status_page — need alert:write. Maintenance: list_maintenance, schedule_maintenance, cancel_maintenance. Incidents: list_status_incidents for what the page is showing, post_status_update to publish an update about one (alert:write, audited, and public immediately). Custom domains: list_status_page_domains, add_status_page_domain, verify_status_page_domain, remove_status_page_domain. Example: create_status_page(org="acme", slug="acme", name="Acme Status", published=true, components=[{name:"Checkout", sourceKind:"monitor", sourceId:"01J…"}]).

A status page is what your customers read when something is wrong. It shows the health of the monitors and services you choose, under names you pick.

Subscribers

People outside your company can ask to hear when a page changes status — by email, or by webhook for a team that wants it in their own system.

Email is double opt-in. Anyone can type anyone's address into a public form, so the row is inert until the person holding that inbox clicks the confirmation link. Until then nothing is sent to them, ever. Submitting the form twice creates one row and one confirmation, which is also what stops the form being used to mail-bomb an address.

A webhook proves it asked. On subscribe, the URL receives a POST of {"type": "bugwatch.status.subscription", "page": "<slug>", "challenge": "spch_…"} and must answer 2xx with the challenge, either as the whole body or as {"challenge": "spch_…"}, within 5 seconds. Otherwise the subscription is refused with 400 and nothing is stored. Without this step anybody could aim a page's notices at any URL. A webhook must also be https://, because announcing an outage over cleartext announces it to the network.

Limits. A page takes at most 100 webhook subscribers and 1,000 unconfirmed email subscriptions a day. One address is sent at most 3 confirmation emails an hour, across every page. Past those limits the form keeps answering the same flat ok, but nothing is stored or sent. The form is also rate-limited per IP.

Delivery is per subscriber. If some subscribers fail (a webhook down, a mail provider hiccup), the retry goes only to them. Nobody gets the same notice twice because somebody else's endpoint was broken.

Every notice carries a one-click unsubscribe. It is a capability — no account, no login — because a subscription somebody cannot leave is one they report as spam, and a status page in spam folders is a status page nobody receives.

What is and is not an email

The page republishes on every health change: a monitor that flaps twice a minute republishes twice a minute. A subscriber list wired naively to that is a mailing list nobody stays on. So a notice goes out only when the page's overall status changes, and three rules follow from it:

  • A component moving while the summary holds is not an email. A second service degrading during an outage is not a second announcement.
  • Nothing is sent on the first publish. A page has no previous status the first time anybody looks, and "operational" arriving out of nowhere is not news.
  • unknown is never an incident. A page whose monitors have not reported yet, or whose only component is paused, is not down — and mailing everybody that it is would be the status page raising a false alarm about itself. Moving out of unknown is not a recovery either; we simply started knowing.

Where this lives, and why not on the status page

Subscribing, confirming and unsubscribing are handled by the API, not by the status Worker.

That Worker binds exactly one thing — the KV namespace holding pre-built snapshots — so the page stays up during precisely the outage it is reporting. Putting a form that writes to the database on it would give it a dependency that can be down at the worst moment. The page links out instead: the subscription surface may be unavailable while the page itself is not, which is the right way round.

The uptime strip

Each monitor-backed component carries ninety days of per-day history, oldest first, with the uptime figure for the window beside it.

Four things about it are deliberate, and each is a way the strip could flatter the service:

  • A day with no events is not a green day. monitor_events stores state changes, so a three-day outage is two rows. The strip carries the state forward across days that recorded nothing, rather than counting rows per day and painting two of the three green.
  • Before the monitor existed is unknown, never operational. A page whose first render shows ninety green days is making a claim about a service nobody has watched yet. The same rule the component status already follows for a paused monitor.
  • Ninety days is the whole window, because that is what is retained. monitor_events is swept at 90 days, so asking for a longer history would not produce more of it — it would produce confident history derived from an absence.
  • A day is ranked by its worst state, not its average. Twenty-three good hours and one outage is a down day; the percentage in the tooltip says 95.83%, and the cell is red.

A service-backed component has no strip. Its status comes from open incidents, which are not a per-day signal, and an empty strip would read as ninety unknown days about something the strip does not apply to.

Days are UTC. A status page has no reader's timezone to use, and picking the organization's would make two readers disagree about which day an outage fell on.

Why it stays up

The page a customer loads is served by a separate Worker with one binding: a KV namespace holding a pre-built snapshot. No D1, no Durable Objects, no service binding back to the API. It renders what it finds and never computes a status.

That is the whole design, and everything else follows from it. A status page that has to ask the control plane how things are going goes dark at exactly the moment somebody is looking for it — which is the moment it exists for. So the snapshot is built on write, not on read:

  • editing the page rebuilds it,
  • a monitor changing state rebuilds every published page that shows that monitor,
  • an incident opening or resolving rebuilds every published page that shows its service.

A snapshot the Worker cannot parse returns 503, not a green page. On a status page, failing loudly beats failing reassuringly.

Components

A component is a public name for something internal. Readers see Checkout; they never see a monitor id, and no id appears in the served page.

Field Meaning
name What readers see, e.g. Checkout
description Optional second line under the name
sourceKind monitor or service
sourceId The monitor or service whose health this component reports

Components are an ordered list and the order is the display order, so update_status_page replaces the whole list rather than patching entries. A source must belong to the same organization as the page.

A page can carry up to 40 components. Past that it has stopped being a status page and become a dashboard.

What each colour means

Shown When
Operational The monitor is up, or the service has no open incident
Degraded The monitor is degraded, or the service's open incidents are all sev3/sev4
Major outage The monitor is down, or the service has an open sev1/sev2 incident
Maintenance A maintenance window covers a component that is otherwise fine
Unknown The monitor is paused, disabled, or has not reported yet

Two of those rows are deliberate refusals:

  • A paused or disabled monitor is Unknown, not Operational. A probe nobody is running proves nothing, however green it was when somebody switched it off.
  • Maintenance never masks a real fault. A component that is genuinely down during a window is shown as down. Saying "scheduled maintenance" over an unplanned outage is how a status page stops being believed.

The page's own status is its worst component, and a page with no components reads Unknown rather than a green light nobody earned.

Incidents and updates

Incidents on anything the page shows are listed above the components, along with everything published about each one.

POST /v1/orgs/{org}/status-pages/{id}/incidents/{incidentId}/updates (tool: post_status_update) publishes one: a stage and what you want readers to know, in your words. It is public the moment it is written — the page rebuilds on the write rather than waiting for the next health change.

Stage What it tells a reader
investigating We know something is wrong and are looking
identified We know what it is
monitoring A fix is out and we are watching it
resolved It is over

The stage you publish is what the page shows. Without an update the page derives one — investigating until somebody acknowledges the page, then identified — but that derived value answers a question about us, not about the work. An incident whose fix is already deployed is still acknowledged internally, and monitoring has no internal equivalent at all.

Two things override that, in opposite directions:

  • A closed incident always reads resolved, whatever was published last. Closure is a fact about the system rather than a claim about what happened, and a page still saying "investigating" under a closed incident is worse than one that says resolved with nobody's words under it.
  • Publishing resolved does not close anything. If the incident is still open — reopened, or the all-clear went out early — the page keeps showing the real stage. An update is a description, not a control.

Updates are separate from an incident's internal timeline. That timeline records who was paged, who acknowledged, and what the escalation chain did; none of it reaches a reader, and publishing is always something a person or an agent chose to do.

list_status_incidents shows what the page is showing right now, updates included. It reads the page's own assembly rather than a separate query, so it cannot disagree with what readers see.

Resolved incidents stay on the page for 7 days, with their updates, and stop colouring any component the moment they close. "Was there an outage this morning?" is most of what a status page is asked after the fact, and it is the one question the page can answer with nobody awake — dropping a resolved incident on the spot also meant the all-clear could never be read.

Scheduled maintenance

POST /v1/orgs/{org}/status-pages/{id}/maintenance (tool: schedule_maintenance) puts a window on the page: a title, an optional body, startsAt and endsAt in epoch ms, and the component ids it covers. Omit the components and it covers the whole page, which is what "we are doing scheduled maintenance" usually means. list_maintenance shows them, past ones included; cancel_maintenance removes one and the page stops showing it at once.

A window also stops the paging

An open window suppresses pages for anything on the components it covers — the monitors those components point at, and the monitors that escalate through a covered service. Nobody is woken for work you scheduled.

What it does not do is change what is true:

  • The monitor still goes down, so consensus and recovery keep working.
  • The transition still lands in the monitor's timeline, marked as suppressed and naming the window. A withheld page is visible afterwards rather than looking like one that went missing.
  • The status page still shows the component as down if it is genuinely down. Maintenance never masks a real fault, here as everywhere else on the page.

Three rules decide the edges, and each exists because the obvious alternative is worse:

  • A monitor that is still down when the window closes pages then, saying how long the page was withheld. Otherwise scheduling a window over a monitor would silence it permanently — the monitor never transitions again, so nothing would ever look — which is worse than having no windows at all.
  • Recovery is silent only if the outage was. If the DOWN was withheld, nobody is waiting for an all-clear and sending one is noise. If people were paged — the outage started before the window did — they get the recovery even mid-window, because the alternative is a page left open forever.
  • A window that cannot be read suppresses nothing. Every failure here resolves toward paging. Silence is the expensive direction, and no corrupt row or database hiccup is allowed to buy it.

Cancelling a window (cancel_maintenance) restores paging immediately.

A component id that is not on the page is refused, not ignored — a window that silently covers less than you asked for is one you find out about from a reader.

The window is decided when the page is read, not when it is published. That follows from the one thing this page exists to guarantee: it is served by a Worker that reads a single pre-built snapshot and no database. If a window starting at 02:00 needed the control plane to rebuild the page at 02:00, the status page would depend on exactly the system it is supposed to outlive. So the schedule travels inside the snapshot, and the reader's clock resolves it — the window opens and closes on time even if the API has been down for a week, and cancelling one takes effect on the next read.

Two consequences worth knowing:

  • A window is visible on the page from the moment you create it, as upcoming, and changes no component's colour until it starts. That is deliberate: "there is planned work on Saturday" is part of what a status page is for.
  • Maintenance never masks a real fault — the rule above applies here too. A component that is genuinely down during its own maintenance window is shown as down.

Your own domain

A page is served at status.bugwatch.io/{slug} until you point a hostname of your own at it. After that, https://status.acme.com/ is the page.

Add the domain, add two or three DNS records, and wait. Nothing else happens — the certificate is ordered, issued and renewed for you, and the domain goes live on its own once your records resolve. It is re-checked every few minutes, so you do not have to come back and press anything.

curl -X POST https://api.bugwatch.io/v1/orgs/acme/status-pages/01J.../domains \
  -H "authorization: Bearer $BUGWATCH_TOKEN" \
  -H "content-type: application/json" \
  -d '{"hostname": "status.acme.com"}'

The response carries the records to add:

Record Why
CNAME status.acme.com → your fallback origin Sends the traffic to us
TXT _cf-custom-hostname.status.acme.com Proves you control the name
TXT _acme-challenge.status.acme.com Lets the certificate authority issue for it

Validation is by TXT record rather than by HTTP. An HTTP check would need the hostname to already point at us before the certificate exists — and a status page is exactly the thing people aim at us while their own origin is down.

The three states

State What it means
pending Waiting on DNS, or on the certificate. Normal for the first few minutes
active Serving. The hostname resolves here and the certificate is live
failed Something needs you: the hostname was blocked or moved, or a step timed out. lastError carries the certificate authority's own words

active deliberately requires both halves. A hostname that resolves to us before its certificate exists shows every visitor a TLS warning, which reads as your status page being broken — so it is not published until both are ready. The same rule runs backwards: if a live domain later fails or its certificate expires, it stops being served rather than answering over a certificate that no longer validates.

What a custom domain does not do

It serves one page. status.acme.com/some-other-slug is a 404, including the JSON view. Without that rule, pointing a hostname at us would quietly turn it into a public mirror of every status page we host, under your brand.

It does not add a dependency. The Worker serving your page still reads pre-built snapshots from one key-value store and talks to no database — a custom domain is one extra lookup in that same store, written ahead of time. This is the whole reason the page stays up during an outage, and it is not traded away for a nicer URL.

Hostnames are claimed once. A hostname belongs to one page across the whole service. No wildcards: *.acme.com would map every subdomain you have, including ones you add later and never think about, onto a public page.

A page may hold up to five domains, which is enough for a rename with an overlap.

Removing a domain stops it serving immediately and releases the certificate.

Publishing

A page is created unpublished. Nothing is served until published is true, so a half-built page cannot be found by guessing its slug.

Turning published off deletes the snapshot — the public URL stops answering, it does not merely stop updating. Renaming the slug retires the old URL in the same way, so a rename is a rename rather than a fork.

Slugs are globally unique across Bugwatch because they are public URLs; a taken slug returns 409 without saying who has it.

The public views

Path Returns
/{slug} The page: one self-contained HTML document, no scripts, no fonts, no external requests
/{slug}.json The same snapshot as JSON, for anything polling it
/ On a custom domain, that domain's page
/index.json On a custom domain, that domain's snapshot as JSON

Both are cached for 15 seconds. A status page is polled hard during an outage, and a cache that holds for a few seconds is what keeps it answering.

REST

Method Path Scope
GET /v1/orgs/{org}/status-pages org:read
POST /v1/orgs/{org}/status-pages alert:write
GET /v1/orgs/{org}/status-pages/{id} org:read
PATCH /v1/orgs/{org}/status-pages/{id} alert:write
DELETE /v1/orgs/{org}/status-pages/{id} alert:write
GET /v1/orgs/{org}/status-pages/{id}/domains org:read
POST /v1/orgs/{org}/status-pages/{id}/domains alert:write
POST /v1/orgs/{org}/status-pages/{id}/domains/{domainId}/verify alert:write
DELETE /v1/orgs/{org}/status-pages/{id}/domains/{domainId} alert:write
curl -X POST https://api.bugwatch.io/v1/orgs/acme/status-pages \
  -H "authorization: Bearer $BUGWATCH_TOKEN" \
  -H "content-type: application/json" \
  -d '{
    "slug": "acme",
    "name": "Acme Status",
    "headline": "How Acme is doing today",
    "published": true,
    "components": [
      { "name": "Website",  "sourceKind": "monitor", "sourceId": "01J..." },
      { "name": "Payments", "sourceKind": "service", "sourceId": "01J..." }
    ]
  }'

See also