Mobile notifications
For agents:
list_devices(org="acme")shows what would be paged andrevoke_device(org="acme", id="…")stops one;get_notification_prefs/update_notification_prefsread and write your rules. Registering a device and confirming a delivery receipt are the app's own jobs and have no tool — an agent confirming on a handset's behalf would mark a silent device fresh, which is exactly the state a receipt exists to detect.
The pager is the product. This page is what happens between an escalation deciding to page you and a phone ringing.
In the dashboard
Settings → Devices lists every phone that would be paged: what it is, when it was added, and when it last confirmed a delivery. A device that has never confirmed one is called out, because it is on its way to being ignored for pages and that is worth knowing before a shift rather than during one.
Handsets register themselves from the app, so there is nothing to add for them. What the page is for is the two things a handset cannot do for itself: see the whole set, and revoke one you no longer have — and one thing you do here directly, Enable notifications in this browser.
Settings → Notification rules is the other half: what reaches you, how loudly, and when it waits until morning. Every category is listed with the insistence it earns — including Page, which is shown as always on rather than as a switch that would do nothing, because somebody who cannot see the page row does not learn that it is exempt and goes on believing quiet hours will silence it.
Two details on that page are deliberate. A mute is described by the time it ends, never by the duration it was set for — that has been running down since it was chosen. And your escalation ladder is shown as the default until you set one of your own: rendering the default into editable fields would mean that changing one rung silently freezes today's default for the rest, so the day the default improves, the people who once opened the page are the ones who never get it.
Devices belong to a person
A device is one handset, registered by the app with its own push credentials. Devices are strictly per-user: an org admin can see that somebody has devices (seat management) but never their tokens, and never edits their notification rules — quiet hours you did not set are quiet hours that lose you a page.
| Field | Meaning |
|---|---|
platform |
ios, android, or web for a browser |
name |
"Rae's iPhone", "Chrome on macOS" — shown when revoking, so you can tell two devices apart |
hasPushToken / hasVoipToken |
Whether credentials are held. The tokens themselves are never returned by any route: a listing is for recognising and revoking a device, not for reading its secrets back out |
lastReceiptAt |
The last delivery this device confirmed (below) |
A push token identifies one physical device, so re-registering an existing token moves it — a reinstall or a restore to a new phone updates the row rather than creating a second target that pages the same handset twice. Revoking clears the credentials with the row, because a revoked device that still holds a live-looking token is one careless query away from being paged again.
Escalation policies name people, never devices. A handset registered this morning is paged tonight with no change to any policy.
Registering your first device in an organization — handset or browser — also creates your mobile channel — a mobile_push notification channel whose config is just your user id. That is what lets a page reach your phones through the same delivery machinery every other channel uses: claim rows, redelivery resume, the skip-what-already-sent rule, the 30-day delivery audit. An escalation step aimed at you notifies the service's channels and your mobile channel, so a responder with the app is paged even by a service that was never wired to Slack.
The channel is created by you registering a device and never through the channels API, and an admin cannot delete it — both halves of one rule: nobody else changes where your pages go. It holds no secret; the push tokens live in devices and never enter a channel config.
Browser notifications
A browser is a device in the same registry, not a separate kind of thing. Settings → Devices → Enable notifications in this browser registers the one you are sitting at, and from then on it is paged exactly as a handset is: the same escalation policies, the same notification rules, the same quiet hours, the same test page.
It is worth doing first, because it is the only paging channel that needs nothing procured. Voice needs a telephony account, iOS critical alerts need an entitlement Apple grants at its discretion, Android needs a Firebase project, email needs a transactional provider. Browser push needs a keypair your deployment generates for itself. On a new deployment it is the difference between a pager and a mailing list.
What it is not: a browser notification does not break through Do Not Disturb, and a closed laptop receives nothing until it wakes. It is a good desk pager and a poor night pager. For a rotation that covers the night, pair it with voice or the mobile app.
Three things about how it behaves are deliberate:
- Permission is asked when you click the button, never on page load. A prompt nobody expects is the reliable way to get "Block" forever — and once blocked, the site cannot ask again; you have to allow it in your browser's own site settings.
- A page stays on screen until you dismiss it, and a repeat page re-alerts rather than silently replacing the last one. An escalation moving to the next person must not look identical to nothing happening.
- Clicking a notification opens the dashboard, not the incident. Push payloads travel through a third-party push service, so they carry ids and nothing that identifies a person or a request — there is no link in them to follow, and inventing one would send you to the wrong place rather than to a slightly less specific right one.
The alert text itself is encrypted end-to-end to your browser (RFC 8291): the push service that carries it can see that a message went to you and nothing about what it said.
A browser confirms deliveries too. When you enable notifications, this browser is issued a receipt token (stored only as a hash on our side, and in this browser's own storage). Each time the service worker puts a notification on screen, it posts that token back, so a browser goes stale and is badged exactly like a handset when its confirmations stop: permission revoked in site settings, the service worker unregistered, or the subscription silently rotated. The receipt needs no session; the token alone identifies the device, and it can do nothing except mark that one browser as having received a push. Re-enabling notifications issues a new token; revoking the device retires it. A browser registered before receipts existed (2026-10-02) confirms nothing until you turn its notifications off and on again; until then it is still paged, because "never confirmed" counts as fresh.
If the button is not there, this deployment has no VAPID keypair set — see the deploy runbook.
Categories, and the insistence each earns
Too loud and people mute the app; too quiet and a page is missed. Each category has a default, and only the first is non-negotiable.
| Category | Default insistence | Breaks Do Not Disturb |
|---|---|---|
| Page — an escalation step targeting you | Critical alert + call UI | Always |
| Incident update — someone acked, it escalated past you, it resolved | Normal, updates the thread in place | No |
| Handoff — your shift starts or ends | Time-sensitive | No |
| Coverage gap — a rotation you own is uncovered in the next 24 h | Time-sensitive | No |
| Monitor state — a monitor went down or recovered and did not page you | Bundled | No |
| Issue alert — new issue, regression, or spike matching a rule | Bundled | No |
| Agent run — a fix-agent run opened a PR or gave up | Normal | No |
| Account — quota, credits, payment | Normal | No |
A monitor going down that did not page you is a monitor state notification, not a page. Calling it a page is how a pager stops meaning anything.
A page is never suppressed
This is enforced in the server-side sender, not in a client. Not by a muted category, not by quiet hours, not by a preference toggle, not by a page: false in a preferences payload — that request is refused with a 400. Someone who silenced notifications at 11pm did not mean "do not wake me when production is down", and a client bug must not be able to disable paging.
Everything else obeys you, in order: category toggle, then mutes, then quiet hours.
- Quiet hours are wall clock in your own IANA zone, and they wrap midnight. The window stays 22:00–08:00 local on both sides of a DST change — quiet hours that drift by an hour twice a year are quiet hours nobody trusts. A held notification comes back with the instant the hold ends, so the morning summary lands on the boundary. A malformed window delivers rather than holds.
- Mutes always expire. There is no indefinite switch to send: the schema requires an
until, and a mute longer than 30 days is refused. An invisible permanent mute is how pages get missed months later by someone who forgot they set it. - Per-device page routing. Choose which of your devices take pages, so a shared tablet does not ring at midnight. Choosing none means all of them, which is the right default.
- Your own ladder.
laddersets which media an escalation reaches you through, in what order, with what delays —[{channel: "push", afterSeconds: 0}, {channel: "slack", afterSeconds: 300}]says "my phone now, the team's room in five minutes if I have not answered". Only media the product can actually deliver are accepted (push,slack,email,webhook,pagerduty); at most 8 rungs, none more than an hour out. A high-urgency ladder's first rung is always forced to immediate — see Incidents.
Not becoming noise
- One thread per incident. Every update to an incident replaces its notification in place. You never scroll five notifications about one outage.
- Chatty categories bundle on a rolling five-minute window: "3 monitors degraded" is one useful notification where three are three chances to learn the app is noise.
- Stale devices stop counting. A device that has not confirmed a delivery in a week is probably gone and is not paged — unless every device looks stale, in which case they are all paged. "Probably gone" is not a reason to page nobody.
- No devices is not a delivery. When there is nothing to page, the escalation is told so and falls through to its other channels. Reporting a page that never left would be a lie.
What actually goes over the wire
The payload shapes are pinned with golden tests so the apps are built against a fixed contract.
| Transport | Used for | Shape |
|---|---|---|
| APNS alert | Every iOS notification | interruption-level per insistence; a page also carries a critical sound at full volume, which is the only thing that sounds through iOS Focus (it needs Apple's critical-alert entitlement). thread-id and apns-collapse-id are the incident, so an update replaces the page in place. |
| APNS VoIP (PushKit) | Pages only | Carries no aps dictionary — iOS rejects a VoIP push that has one — so the call screen's caller label travels in the payload. This is what makes a page present as an incoming call, answerable without unlocking. |
| FCM HTTP v1 | Every Android notification | Data-only and HIGH priority for a page. A notification block is drawn by the system without waking the app, which means no full-screen intent and no page. collapse_key is the incident. |
Nothing that identifies a person goes to a push provider. The payload carries ids, a severity, a title and one short line — never an event body, a stack frame, a request, or a user's address. Push payloads pass through Apple's and Google's infrastructure and their logs; the allowed keys are a checked list rather than a review habit.
A page expires after five minutes and everything else after an hour. A page delivered an hour late rings for an incident somebody else already took, which is worse than one never delivered.
Send me a test page
A pager nobody has tested is a pager nobody trusts. Settings → Devices has one button, and POST /v1/orgs/{org}/test-page is the same thing over REST.
It is a real page down the real path: the same queue, the same notify Worker, the same mobile channel, the same critical-alert and PushKit payloads at the same insistence. A test that took any shortcut would only prove the shortcut works.
Two things keep it honest:
- It goes only to your own phones. Never the organization's channels, never the service's Slack. A drill that wakes the team is a drill nobody runs twice.
- The payload says it is a test (
test: "1"), so the app can label it and nobody starts an incident response over it.
Each press is its own page — two in a row are two pages, never one deduped against the other. With no registered device the request is refused and says so, rather than reporting a page that reached nothing.
Send one before every shift you are on call for. It is the only way to find out that a phone stopped ringing while it still mattered that it did.
Canary devices
Everything else in the product exists to tell somebody that something is broken. The canary is what notices when that is what is broken.
A canary is an ordinary registered handset flagged in the device registry (devices.is_canary). Every five minutes it receives a synthetic page down the ordinary path — the alerts queue, the notify Worker, the mobile channel, the real critical-alert and PushKit payloads — and confirms it with a delivery receipt like any other device. A watchdog with its own shortcut to the phone would be testing the shortcut.
When a canary stops confirming for three rounds, the paging path is presumed broken. Three rather than one: a single push can be late for reasons that are nobody's emergency, and a watchdog that cries at the first one is a watchdog somebody mutes.
Something outside has to watch the watchdog. Both halves of the canary — sending the synthetic page and noticing that rounds were missed — run on the same scheduled Worker. If that Worker stops, no page goes out and nothing notices, so the failure alarm never fires either: silence would look exactly like health. CANARY_HEARTBEAT_URL is pinged after every healthy round so a dead-man's-switch elsewhere alarms when it stops arriving. It is deliberately never pinged while a canary is silent — otherwise the external switch would report fine at the moment the alarm reports broken, and the external one is the one being watched.
The alarm does not go through Bugwatch. The thing that would carry it is the thing under suspicion, so a missed canary is reported over a plain webhook (CANARY_ALERT_URL) to something that is not us. It carries device ids, platforms and timestamps — never anything about a person.
A deployment with no canary configured stays silent. An alarm on every install that never wanted the watchdog is noise, and noise is how a real alarm gets ignored.
Delivery receipts
The provider is not trusted to deliver. The app confirms each push it actually received (POST /v1/orgs/{org}/devices/{id}/receipt), and a browser's service worker confirms each notification it shows (POST /devices/receipt with its receipt token). That timestamp is what the staleness rule above reads. A device whose receipts stop is going stale; a device the provider reports as gone — APNS 410/BadDeviceToken, FCM UNREGISTERED — is revoked immediately rather than a week later, because until it is, it is a delivery target that silently absorbs pages.
One dead handset never holds up the others: a permanent failure on one device revokes that device and the send continues, while a 5xx or a 429 is retried with the escalation step still open.
REST routes
| Route | Scope |
|---|---|
GET /v1/orgs/{org}/devices |
org:read, session only |
POST /v1/orgs/{org}/devices |
org:read, session only. Returns channelId — your mobile channel |
DELETE /v1/orgs/{org}/devices/{id} |
org:read, session only |
POST /v1/orgs/{org}/devices/{id}/receipt |
org:read, session only |
POST /devices/receipt {"token": "rcpt_…"} |
the browser's receipt token; no session |
POST /v1/orgs/{org}/test-page |
org:read, session only |
GET /v1/orgs/{org}/notification-prefs |
org:read, session only |
PUT /v1/orgs/{org}/notification-prefs |
org:read, session only |
"Session only" is deliberate: a device and a set of notification rules belong to a person, and an API token — which may outlive whoever created it — is the wrong thing to hang either on. These routes answer 400 to a bearer caller rather than guessing whose phone it is.
Chat
Everything the app does beyond the page and on-call awareness is a conversation over the same tool registry the MCP server exposes — so a tool added for the web is answerable on the phone the same day, and the two surfaces cannot drift.
POST /v1/orgs/{org}/chat streams the answer as server-sent events. Reads run freely: anything the caller's own scopes already allow, enforced by the route each tool wraps rather than by a second list. Every fact comes from a tool call — a companion that invents an all-clear is a liability.
Writes are never performed by the model. Proposing an action emits a confirm event carrying the exact tool, its arguments, and a sentence saying what will happen; the turn then ends. Nothing changes until the person taps it and the client calls back with confirmed, which runs that one tool through the same REST handler the web and MCP use. One card at a time.
| Allowed from a phone | Never from a phone |
|---|---|
acknowledge_incident, resolve_incident, create_override, resolve_issue, ignore_issue, unresolve_issue, pause_monitor, resume_monitor |
Minting tokens, connecting repositories, checkout and credit purchases, member changes, deleting anything, editing alert rules or escalation policies |
Everything on the left is reversible, scoped to one subject, and something a person does during an incident. Everything on the right is human-initiated on the web by policy, and a phone at 3am is the worst place to relax that.
Chat is metered against the same agent credits as the fix agent. Running out degrades chat and nothing else — the paging path never calls a model, never reads a balance, and cannot be reached from here.
Not yet
- The apps themselves. Signed binaries, the critical-alert entitlement and store listings need Apple and Google developer accounts and real devices. Everything on this page is the server half, built and tested without them.
- A chat client. The endpoint is live; the app that talks to it is U4b.
- Voice and SMS. There is no voice provider; SMS is U7 and ships last. The ladder deliberately schedules neither — a rung for a medium nobody can deliver is a duplicate under another name.