bugwatch docs

On-call schedules

For agents: who_is_on_call(org="acme", id="01J…") answers the only question that matters during an incident; get_schedule_coverage(org="acme", id="01J…") shows the next seven days and counts uncovered windows; list_schedules / get_schedule describe what exists. Writes — create_schedule, update_schedule, delete_schedule, create_override, delete_override — need alert:write. Example: who_is_on_call(org="acme", id="01J…", at=1772953200000).

A schedule says who is on call at any instant. It is a stack of rotation layers plus overrides, evaluated in the schedule's own timezone.

The model

Piece What it is
Schedule A name, an IANA timezone, and a stack of layers.
Layer A rotation: an ordered list of users, a rotation type, a handoff time, and the dates it is active.
Restriction A wall-clock window a layer is limited to — "this layer only covers 09:00–17:00", or "only Saturdays".
Override One person, one window, beating every layer. Illness, holiday, a handover.

Layers stack: the highest-position layer active at time T wins. Overrides beat every layer.

Rotation types

  • daily — hands off every day at handoffLocalTime.
  • weekly — hands off on handoffWeekday (0 = Sunday) at handoffLocalTime.
  • custom — hands off every turnLengthSeconds. A turn that is a whole number of days advances in wall-clock time like the others; a shorter turn (12 hours, say) has no "same time tomorrow" to preserve and advances by real elapsed seconds.

The first user of a layer is on call from the layer's startDate until its first handoff, even when that is a partial turn. Starting a rotation on a Monday at 09:00 with a Monday 09:00 handoff puts the first person on call immediately, not the second.

Timezones and DST — the part that quietly goes wrong

Two defensible semantics exist for what happens when the clocks change:

  • fixed duration — turns are exactly n × 86 400 seconds, and the handoff time drifts by an hour twice a year.
  • fixed wall clock — the handoff is always Monday 09:00 local, and two turns a year are 23 or 25 hours long.

Bugwatch implements fixed wall clock. People say "handoff is Monday at 9" and mean it. Rotation boundaries advance by incrementing wall-clock time in the schedule's IANA zone, never by adding seconds to a timestamp.

That is also why timeZone is an IANA zone name (America/New_York) and never a UTC offset (-05:00). Offsets change twice a year; zones do not.

Wall-clock times create two cases that need a stated policy, and both are decided by the engine rather than left to whoever asks:

Case What happens Policy
Nonexistent local time Spring forward: 02:00–03:00 never occurs Moves to the first instant that does exist — the transition itself
Ambiguous local time Fall back: 01:00–02:00 happens twice Takes the first occurrence

A restriction of "09:00 for 8 hours" therefore ends at 17:00 local on every day of the year, including the 23- and 25-hour ones — the real elapsed time is what varies, not the wall clock.

Gaps are an answer, not an error

When no layer covers an instant and no override applies, the schedule resolves to nobody, and both who_is_on_call and get_schedule_coverage say so explicitly (userId: null, source: "gap").

This is deliberate. Most "nobody got paged" incidents are not a delivery failure — they are an uncovered window that nobody noticed, usually created by a restriction that does not tile the day or a layer whose endDate passed. get_schedule_coverage returns a gaps count for exactly this reason; check it after every schedule change.

Coverage

GET /v1/orgs/{org}/schedules/{id}/coverage?from=…&to=… returns contiguous, non-overlapping segments over the window (default: the next 7 days, maximum 90). Every instant belongs to exactly one segment — intervals are half-open, [startsAt, endsAt) — so the boundary where one person hands over to the next is unambiguous.

Notification rules

Which notifications reach you, how insistently, and on which device is decided by a per-user rule set: a category taxonomy, quiet hours in your timezone, time-boxed mutes, and per-device page routing. One rule is not configurable and is enforced in the sender rather than in any client: a page is never suppressed. Someone who silenced notifications at 11pm did not mean "do not wake me when production is down", and a client bug must not be able to disable paging.

The rules, the device registry and the push payload contract are all on Mobile notifications.

The coverage index

Answering "who is on call" from source means reassembling the schedule from its tables and replaying the rotation from its start date. That is fine for a dashboard read. It is the wrong thing to do on the paging path, where the answer is needed while somebody's service is down.

Each schedule therefore has a Durable Object holding a rolling 30-day coverage index, built by the same buildCoverage every other surface uses — a cached answer and a computed one cannot disagree. An escalation step resolves its targets from that index instead of from the database.

It is a cache, and it behaves like one:

  • An instant outside the built window, or an index older than a day, is rebuilt before it is answered. A stale index paging the wrong person is worse than a slow one paging the right person.
  • Every write to a schedule, layer or override invalidates it; the next read rebuilds it, and a nightly alarm rolls the window forward.
  • If the index cannot answer — unreachable, unbuilt, deleted — the escalation resolves from the database instead. A cache miss is never read as "nobody is on call", because that would page the policy's fallback and record a gap that does not exist.

The dashboard and the REST API still resolve from source, so what you see there is the schedule itself rather than a cache of it.

In the dashboard

People appear by name (falling back to their email, and to a short id for anyone who has since left the organization — a rotation that named them still has to say who was on call). Overrides are added by picking a member, not by pasting a user id.

On-call in the sidebar lists every schedule with who is on call right now, and counts the schedules where that answer is nobody.

Opening one shows a week of coverage as a row per local day. Each row is scaled to its own length, so the 23- and 25-hour days around a transition are drawn as they actually are rather than stretched to a nominal 24. Uncovered windows are hatched in red and counted above the calendar; the legend keeps one colour per person across the whole week. Every time on the page — the week range, the tooltips, the override windows — is rendered in the schedule's timezone, not the viewer's, so two people in different countries reading the same calendar see the same thing.

Overrides can be added and removed from the same page. You pick the window in your own local time, because a cover is an agreement about instants rather than about anybody's wall clock.

REST routes

Route Scope
GET /v1/orgs/{org}/schedules org:read
POST /v1/orgs/{org}/schedules alert:write
GET /v1/orgs/{org}/schedules/{id} org:read
PATCH /v1/orgs/{org}/schedules/{id} alert:write
DELETE /v1/orgs/{org}/schedules/{id} alert:write
GET /v1/orgs/{org}/schedules/{id}/on-call org:read
GET /v1/orgs/{org}/schedules/{id}/coverage org:read
POST /v1/orgs/{org}/schedules/{id}/overrides alert:write
DELETE /v1/orgs/{org}/schedules/{id}/overrides/{overrideId} alert:write

PATCH with layers replaces the whole stack. Positions, rotation order and restrictions only make sense as a set, and editing one layer of a stack in isolation is how coverage gaps get introduced by accident.

Not yet

Escalation policies, incidents and paging are the next phase — a schedule answers who, and nothing yet routes a page to them. The precomputed coverage cache that will serve the paging hot path lands with it, since that is the thing which reads it.