# On-call schedules

> **For agents:** `who_is_on_call(org="acme", id="01J…")` answers the only question that matters during an incident; `get_schedule_coverage(org="acme", id="01J…")` shows the next seven days and counts uncovered windows; `list_schedules` / `get_schedule` describe what exists. Writes — `create_schedule`, `update_schedule`, `delete_schedule`, `create_override`, `delete_override` — need `alert:write`. Example: `who_is_on_call(org="acme", id="01J…", at=1772953200000)`.

A schedule says who is on call at any instant. It is a stack of **rotation layers** plus **overrides**, evaluated in the schedule's own timezone.

## The model

| Piece | What it is |
|---|---|
| **Schedule** | A name, an IANA timezone, and a stack of layers. |
| **Layer** | A rotation: an ordered list of users, a rotation type, a handoff time, and the dates it is active. |
| **Restriction** | A wall-clock window a layer is limited to — "this layer only covers 09:00–17:00", or "only Saturdays". |
| **Override** | One person, one window, beating every layer. Illness, holiday, a handover. |

Layers stack: **the highest-position layer active at time T wins**. Overrides beat every layer.

## Rotation types

- **daily** — hands off every day at `handoffLocalTime`.
- **weekly** — hands off on `handoffWeekday` (0 = Sunday) at `handoffLocalTime`.
- **custom** — hands off every `turnLengthSeconds`. A turn that is a whole number of days advances in wall-clock time like the others; a shorter turn (12 hours, say) has no "same time tomorrow" to preserve and advances by real elapsed seconds.

The first user of a layer is on call from the layer's `startDate` until its first handoff, even when that is a partial turn. Starting a rotation on a Monday at 09:00 with a Monday 09:00 handoff puts the *first* person on call immediately, not the second.

## Timezones and DST — the part that quietly goes wrong

Two defensible semantics exist for what happens when the clocks change:

- **fixed duration** — turns are exactly *n* × 86 400 seconds, and the handoff time drifts by an hour twice a year.
- **fixed wall clock** — the handoff is always Monday 09:00 local, and two turns a year are 23 or 25 hours long.

**Bugwatch implements fixed wall clock.** People say "handoff is Monday at 9" and mean it. Rotation boundaries advance by incrementing wall-clock time in the schedule's IANA zone, never by adding seconds to a timestamp.

That is also why `timeZone` is an IANA zone name (`America/New_York`) and never a UTC offset (`-05:00`). Offsets change twice a year; zones do not.

Wall-clock times create two cases that need a stated policy, and both are decided by the engine rather than left to whoever asks:

| Case | What happens | Policy |
|---|---|---|
| **Nonexistent** local time | Spring forward: 02:00–03:00 never occurs | Moves to the first instant that does exist — the transition itself |
| **Ambiguous** local time | Fall back: 01:00–02:00 happens twice | Takes the **first** occurrence |

A restriction of "09:00 for 8 hours" therefore ends at 17:00 local on every day of the year, including the 23- and 25-hour ones — the real elapsed time is what varies, not the wall clock.

## Gaps are an answer, not an error

When no layer covers an instant and no override applies, the schedule resolves to **nobody**, and both `who_is_on_call` and `get_schedule_coverage` say so explicitly (`userId: null`, `source: "gap"`).

This is deliberate. Most "nobody got paged" incidents are not a delivery failure — they are an uncovered window that nobody noticed, usually created by a restriction that does not tile the day or a layer whose `endDate` passed. `get_schedule_coverage` returns a `gaps` count for exactly this reason; check it after every schedule change.

## Coverage

`GET /v1/orgs/{org}/schedules/{id}/coverage?from=…&to=…` returns contiguous, non-overlapping segments over the window (default: the next 7 days, maximum 90). Every instant belongs to exactly one segment — intervals are half-open, `[startsAt, endsAt)` — so the boundary where one person hands over to the next is unambiguous.

## Notification rules

Which notifications reach you, how insistently, and on which device is decided by a per-user rule set: a category taxonomy, quiet hours in **your** timezone, time-boxed mutes, and per-device page routing. One rule is not configurable and is enforced in the sender rather than in any client: **a page is never suppressed**. Someone who silenced notifications at 11pm did not mean "do not wake me when production is down", and a client bug must not be able to disable paging.

The rules, the device registry and the push payload contract are all on [Mobile notifications](https://docs.bugwatch.io/product/mobile-notifications.md).

## The coverage index

Answering "who is on call" from source means reassembling the schedule from its tables and replaying the rotation from its start date. That is fine for a dashboard read. It is the wrong thing to do on the **paging** path, where the answer is needed while somebody's service is down.

Each schedule therefore has a Durable Object holding a rolling 30-day coverage index, built by the same `buildCoverage` every other surface uses — a cached answer and a computed one cannot disagree. An escalation step resolves its targets from that index instead of from the database.

It is a cache, and it behaves like one:

- An instant outside the built window, or an index older than a day, is rebuilt before it is answered. A stale index paging the wrong person is worse than a slow one paging the right person.
- Every write to a schedule, layer or override invalidates it; the next read rebuilds it, and a nightly alarm rolls the window forward.
- If the index cannot answer — unreachable, unbuilt, deleted — the escalation resolves from the database instead. A cache miss is never read as "nobody is on call", because that would page the policy's fallback and record a gap that does not exist.

The dashboard and the REST API still resolve from source, so what you see there is the schedule itself rather than a cache of it.

## In the dashboard

People appear by name (falling back to their email, and to a short id for anyone who has since left the organization — a rotation that named them still has to say who was on call). Overrides are added by picking a member, not by pasting a user id.


**On-call** in the sidebar lists every schedule with who is on call *right now*, and counts the schedules where that answer is nobody.

Opening one shows a week of coverage as a row per local day. Each row is scaled to its own length, so the 23- and 25-hour days around a transition are drawn as they actually are rather than stretched to a nominal 24. Uncovered windows are hatched in red and counted above the calendar; the legend keeps one colour per person across the whole week. Every time on the page — the week range, the tooltips, the override windows — is rendered in the **schedule's** timezone, not the viewer's, so two people in different countries reading the same calendar see the same thing.

Overrides can be added and removed from the same page. You pick the window in your own local time, because a cover is an agreement about instants rather than about anybody's wall clock.

## REST routes

| Route | Scope |
|---|---|
| `GET /v1/orgs/{org}/schedules` | `org:read` |
| `POST /v1/orgs/{org}/schedules` | `alert:write` |
| `GET /v1/orgs/{org}/schedules/{id}` | `org:read` |
| `PATCH /v1/orgs/{org}/schedules/{id}` | `alert:write` |
| `DELETE /v1/orgs/{org}/schedules/{id}` | `alert:write` |
| `GET /v1/orgs/{org}/schedules/{id}/on-call` | `org:read` |
| `GET /v1/orgs/{org}/schedules/{id}/coverage` | `org:read` |
| `POST /v1/orgs/{org}/schedules/{id}/overrides` | `alert:write` |
| `DELETE /v1/orgs/{org}/schedules/{id}/overrides/{overrideId}` | `alert:write` |

`PATCH` with `layers` **replaces the whole stack**. Positions, rotation order and restrictions only make sense as a set, and editing one layer of a stack in isolation is how coverage gaps get introduced by accident.

## Not yet

Escalation policies, incidents and paging are the next phase — a schedule answers *who*, and nothing yet routes a page to them. The precomputed coverage cache that will serve the paging hot path lands with it, since that is the thing which reads it.
