bugwatch docs

Alerts

For agents: list_channels(org="acme"), create_channel(org="acme", name="#alerts", config={kind:"slack_webhook", url:"https://hooks.slack.com/services/…"}), list_alert_rules(org="acme", project="web"), create_alert_rule(…), update_alert_rule(…), delete_alert_rule(…). Writes need alert:write; channel secrets are never returned, only a masked target.

Channels

A channel is an organization-level destination. Eight kinds:

kind Config Notes
slack_webhook {"url": "https://hooks.slack.com/…"} Incoming webhook; URL must start with https://hooks.slack.com/
discord {"url": "https://discord.com/api/webhooks/…"} Channel webhook; posts an embed whose title links to the issue or incident
teams {"url": "https://…​.logic.azure.com/…"} A Workflows webhook (Power Automate), posting an Adaptive Card. Not the retired Office 365 connector format
webhook {"url": "https://…"} Generic JSON POST, HMAC-signed (below)
email {"to": "oncall@example.com"} Needs the transactional email provider configured on the notify Worker; otherwise the delivery is recorded as skipped
pagerduty {"routingKey": "…"} Events API v2 routing key, 16–64 chars
voice {"to": "+15551234567"} A phone call that reads the alert aloud. E.164 only. Needs Twilio configured on the notify Worker; otherwise the delivery is recorded as skipped
mobile_push {"userId": "…"} One person's registered phones and browsers. Not creatable here — it appears when that person registers a device, and only they can remove it by revoking their devices (Mobile notifications)

Configs are encrypted at rest with AES-GCM (enc:v1: prefix) when the CHANNEL_ENC_KEY secret is set; GET /v1/orgs/{org}/channels reports encryptionEnabled and, per channel, encrypted plus a masked target such as hooks.slack.com/…/Ab3F or ••••7f2c. The full secret is never readable through the API, and audit rows reduce a channel config to its kind.

Chat channels, and why their URLs are host-checked

Slack, Discord and Teams all work the same way: the URL is the credential, so anyone holding it can post into that channel. Each is checked against the host its provider actually serves webhooks from — hooks.slack.com, discord.com/api/webhooks/, and a .logic.azure.com Workflows URL for Teams. A channel that posted to any URL you named would be a request-forgery surface pointed at whatever the notify Worker can reach, which is why the check exists rather than as a typo-catcher.

If your URL is legitimate and refused — a self-hosted relay, a provider not listed here, a Teams tenant on a host these patterns miss — use the generic webhook kind. It accepts any https:// URL and posts the alert as signed JSON. That escape hatch is the reason the other kinds can afford to be strict.

Teams uses Adaptive Cards through a Workflows webhook, deliberately. Most examples online still show the MessageCard format from Office 365 connectors; Microsoft has retired those, so a channel built on them would ship with a known expiry date.

Discord escapes markdown in the description and not in the title, because Discord renders one and not the other. An exception message containing backticks or underscores keeps the characters that identify it instead of turning into formatting.

Voice: the channel that can wake somebody

A voice channel places a real phone call. It is the only destination here that breaks through Do Not Disturb, Focus and silent mode without an app installed — which is why it exists and why it is worth configuring before any rotation goes live.

The call says what broke, twice, with a pause between: somebody woken by a ringing phone does not take in the first sentence, and a pager you have to call back to find out what happened is a worse pager. It never reads the URL aloud — text-to-speech turns a link into syllables nobody can write down.

Acknowledging by keypad. When the call is about an escalated incident, it offers "press 4 to acknowledge". Pressing 4 acknowledges the incident exactly as the dashboard button does, and the call says so before hanging up.

Three things it will not do, each deliberate:

  • It does not offer the key when there is nothing to acknowledge. A monitor-down notification with no incident behind it is announce-only. Offering a key that does nothing teaches people it does not work, and the next time it matters they will not press it.
  • It never claims an acknowledgement that did not happen. If the incident refuses the transition, the call says the page is still chasing somebody rather than thanking you.
  • It says so out loud when no key was pressed — "this page remains unacknowledged" — so hanging up on a call is never mistaken for handling it.

The acknowledgement URL is a capability: it carries the incident id and an HMAC of it, so it cannot be forged or reused for a different incident, and it only ever travels to the telephony provider over TLS. This is the same model as a heartbeat ping URL.

Configuration lives on the notify Worker (TWILIO_ACCOUNT_SID, TWILIO_AUTH_TOKEN, TWILIO_FROM_NUMBER) plus a shared VOICE_ACK_SECRET that the uptime Worker verifies. With the ack secret unset the call is still placed — announce-only. A page that says what broke is worth far more than no page at all.

Email and voice channels ask their destination first

An email channel can point at any address, so it sends nothing until the address agrees. Creating one emails the address a one-time link ("an organization wants to send alerts here"), and the channel is listed as verified: false (awaiting confirmation) until the link is followed. An org can have at most 10 email channels waiting at once. Without this, a channel was a way to send mail from the platform's domain to anyone.

Slack, Discord, Teams, generic webhooks and PagerDuty reach only endpoints the organization itself controls, and a phone's mobile channel is created by its owner, so these are verified when created. A deployment with no email provider cannot ask an address anything, so there an email channel is trusted as before. A voice channel asks by phone. Creating one places a single call to the number: who is calling, the channel's name, and "to accept these calls, press 1. To refuse, hang up." Only pressing 1 verifies the channel. Silence, any other key, or hanging up is a no, and no alert ever calls a number that has not said yes. Consent calls are limited to 2 a day per number, across every organization, and 5 a day per organization, so deleting and re-adding a channel cannot ring somebody repeatedly. The key press lands on the uptime Worker, so this needs VOICE_ACK_ORIGIN on notify. Without it a voice channel is trusted as before, the same rule email follows without a mail provider.

Channels that existed before these checks keep working.

Deleting a channel

Channel ids are stored as plain lists on the things that page through them — services, monitors and alert-rule actions — so deleting a channel detaches it from all of them in the same operation. The response reports how many rows were touched as detachedFrom.

That matters because a reference left dangling is worse than no reference at all: the service still looks configured, and it pages nobody. Detaching makes the configuration tell the truth, but it can leave something with no way to page — check anything the response counted.

References are also validated when they are written. A service, monitor or escalation policy naming a channel, policy, project, schedule or user that does not exist in the organization is a 400, not a row that silently reaches nobody.

Generic webhooks receive the alert signal as JSON (kind, orgId, projectId, issueId, shortId, title, culprit, level, environment, release, url) with headers x-bugwatch-timestamp and x-bugwatch-signature: v1=<hex HMAC-SHA256 of "<timestamp>.<body>">.

Each webhook channel has its own signing secret (whsec_…). It is returned once, in the response that creates the channel (and shown once in the dashboard), and never again. POST /v1/orgs/{org}/channels/{id}/rotate-secret (rotate_channel_secret) issues a new one; the old one stops verifying at the next delivery. A receiver that checks the signature therefore knows the alert came from its channel. A shared secret could not tell one customer's channel from another's, so anyone able to verify could also forge for everybody else.

Channels created before per-channel secrets are listed with legacySigning: true and are signed with the deployment's WEBHOOK_SIGNING_SECRET until rotated. Rotate them.

Rules

Rules belong to a project (GET/POST /v1/orgs/{org}/projects/{project}/alert-rules, PATCH/DELETE …/alert-rules/{id}):

{
  "name": "Production errors → Slack + PagerDuty",
  "environment": "production",
  "conditions": [{ "type": "new_issue" }, { "type": "regression" }, { "type": "frequency", "count": 100, "windowMinutes": 10 }],
  "filters": [{ "type": "min_level", "value": "error" }, { "type": "release", "value": "2.*" }],
  "actions": [{ "type": "channel", "channelId": "01J…" }, { "type": "service", "serviceId": "01M…" }],
  "frequencyS": 1800,
  "enabled": true
}
  • Conditions (1–10, any matches): new_issue, regression, or frequency with count (1–1,000,000) events within windowMinutes (1–1440). Frequency thresholds are evaluated from the per-issue minute counters kept by the project coordinator, so they fire even for issues that are not new.

  • Filters (0–10, all must match): environment, min_level (debug < info < warning < error < fatal), release (glob-style pattern). A rule-level environment is a shorthand for the same filter.

  • Actions (1–20), of two kinds, both scoped to the organization and refused at write time if not:

    • channel — posts. A Slack message, a webhook, an email, a PagerDuty event.
    • service — escalates. Opens an incident on that service, so the error is chased through the service's escalation policy to whoever is on call. This is the same path a DOWN monitor has taken since uptime shipped, reached from the error side: before it existed, a monitor could page a person and an error could not.

    A rule may carry both, and usually should. The post is how the team hears; the page is how somebody is made to answer.

  • frequencyS: minimum seconds between notifications per rule per issue, 60–604,800 (default 1800).

  • PATCH replaces the whole definition (same fields as create); toggling enabled alone still requires the full body.

Escalating an error onto a service

A service action opens an incident keyed on (service, issue). That key is the whole design:

  • an issue that keeps matching its rule joins the incident it already opened rather than starting a new one every throttle window — an incident is about the thing that is broken, not about each time we noticed;
  • once that incident is resolved, the next match opens a fresh one, because reopening something a person closed is worse than opening something new;
  • a queue redelivery joins too, so the at-least-once delivery of the alert queue cannot wake two people for one error.

Severity is sev3, deliberately. An error somebody wrote a paging rule for is not automatically the worst thing happening; severity is the responder's to raise, and a system that claims the top one every time has stopped saying anything.

Nothing about this suppresses the rule's channel actions. If the escalation cannot be started — the incident service is unreachable, or a deploy has no incident binding — the channel posts still go out and the failure is logged explicitly. Somebody hears either way; what is lost is that nobody is being chased, and that distinction is the reason the rule was written with a service in it.

There is no automatic resolution. A monitor recovers and its incident resolves; an error has no equivalent signal, so an issue-driven incident ends the way the service says — its autoResolveSeconds, or a person resolving it. That is stated rather than left to be discovered at the point somebody wonders why the incident is still open.

When your own quota is the problem

Two signals are organization-level and carry no project: org.quota_warning at 80% of the period's event allowance, and org.quota_exceeded when errors and transactions actually start being dropped.

They go to every announcement channel in the organization — slack_webhook, discord, teams, webhook and email — are never throttled, and escalate nothing. Project alert rules cannot apply to a signal with no project, and there is no sensible way to opt out of being told that your own error monitoring has stopped.

They never reach voice, pagerduty or mobile_push. Those three exist to make a human answer: a call rings a phone, a PagerDuty event opens an incident at the other end, and a mobile channel is aimed at one person's own handsets. A quota threshold is news, and news does not ring a phone — the post is how the team hears, the page is how somebody is made to answer, and a pager that goes off about a bill is a pager slightly less likely to be answered at 3am about an outage.

If nobody in the organization has an announcement channel, the notice reaches nobody. That is deliberate: the alternative is a phone call.

Why both, rather than just the one that matters: at 100% the events are already gone and the notice is an apology. At 80% it is a chance to act.

A metered organization gets the warning too. Nothing stops for it — overage is billed instead — so 80% means the bill is about to grow rather than that data is about to be lost. Suppressing it would leave the only customers who never hear about their usage being the ones paying for it.

This matters more than it looks, because the failure it replaces was invisible from both ends. The SDK reads X-Sentry-Rate-Limits and backs off correctly and quietly — no failed requests, nothing in your logs. And the flag was set server-side and forgotten. The first anybody learned was during an incident, looking for an exception that was never recorded.

Delivery semantics

  • No rules configured → every channel in the organization is notified on new issues and regressions. Alerts-on-by-default beats a silent first week; add a rule to narrow.
  • Signals: issue.new, issue.regressed, issue.frequency. Each carries the dashboard deep link.
  • Throttle: at most one notification per rule + issue per frequencyS, tracked in KV. The throttle is marked only after delivery rows are durably claimed, so a crash cannot silence an alert that never went out. frequencyS is the only thing that suppresses a repeat notification — if a rule allows two firings in an hour, you get two.
  • Exactly-once-ish: every (signal, channel) attempt gets a deterministic delivery row keyed by the signal's occurrence — the event that raised the issue, the monitor's state transition, the rule's threshold crossing. A queue redelivery of the same firing reuses that row and skips channels already marked sent, so nothing double-posts; a genuinely new firing gets its own row and notifies, even if it lands seconds after the last one. Transient failures (5xx, 429) retry via the queue; 4xx is recorded as failed and not retried; unconfigured email is skipped.
  • Delivery runs in a separate notify Worker so a dashboard or API incident cannot stop paging.
  • Retention: delivery rows are kept for 30 days and swept nightly — long enough to reconstruct who was notified during an incident retro, not a permanent ledger.

Ignored issues

ignore_issue mutes an issue: it stops matching new_issue/regression signals until reopened. Frequency rules still see its counters; ignore is a status, not a filter.