# Uptime monitors

> **For agents:** `list_monitors(org="acme")` for the fleet, `get_monitor(org="acme", id="01J…")` for one, `get_monitor_history(org="acme", id="01J…", since=24, interval=5)` for uptime and latency, `list_monitor_events(org="acme", id="01J…")` for the incident timeline. Writes — `create_monitor`, `update_monitor`, `pause_monitor`, `resume_monitor`, `snooze_monitor`, `unsnooze_monitor`, `delete_monitor` — need `alert:write`. Example: `create_monitor(org="acme", name="Checkout API", check={type:"http", url:"https://api.example.com/health", keyword:"ok"}, intervalSeconds=60, regions=["wnam","weur","apac"])`.

A monitor is a check plus a schedule plus the regions that run it. Monitors are **organization-level**, not per-project: they page notification channels, the same channels alert rules use ([Alerts](https://docs.bugwatch.io/product/alerts.md)).

## Check types

| Type | What it asserts | Key fields |
|---|---|---|
| `http` | A request returns a status inside `expectedStatus` (default `[200, 399]`), optionally contains `keyword`, does not contain `keywordAbsent`, and answers within `maxLatencyMs` | `url`, `method`, `headers`, `body`, `expectedStatus`, `keyword`, `keywordAbsent`, `followRedirects`, `timeoutMs`, `maxLatencyMs`, `auth` (`basic` or `bearer`) |
| `tcp` | A TCP connection to `host:port` is accepted within `timeoutMs` | `host`, `port`, `timeoutMs` |
| `tls` | A TLS handshake with `host:port` completes within `timeoutMs` | `host`, `port` (default 443), `minDaysUntilExpiry`, `timeoutMs` |
| `dns` | A DNS-over-HTTPS query for `name`/`recordType` returns `NOERROR` with at least one answer, and — when `expected` is set — an answer equal to it | `name`, `recordType` (`A`, `AAAA`, `CNAME`, `MX`, `TXT`, `NS`), `expected`, `resolver` (an `https://` DNS-over-HTTPS JSON endpoint; default Cloudflare), `timeoutMs` |
| `heartbeat` | Your job pinged the monitor's URL within `expectedEverySeconds + graceSeconds` | `expectedEverySeconds`, `graceSeconds` |

HTTP credentials (`auth`, any header that looks like `authorization`, `cookie`, `token` or `key`, and a password in the URL) are sealed with AES-GCM in the database and returned masked as `••••••`. Only the uptime Worker opens them, to run the check. Re-send the real value to change it; sending the mask back is not a way to read it.

## What Bugwatch cannot check

Probes run inside Cloudflare Workers, which have outbound `fetch` and TCP sockets and nothing lower. So, plainly:

- **No ICMP ping.** There are no raw sockets, so there is no `ping` check and no packet-loss metric. A TCP check against a port the host actually serves is the closest equivalent, and it is a better signal anyway.
- **Bodies are read up to 1 MB.** A `keyword` check looks at the first megabyte of the response and stops reading there, so put the keyword near the top of a large page.
- **No UDP**, so no DNS-over-UDP (the `dns` check uses DNS-over-HTTPS against a resolver you choose), no NTP, no syslog, no QUIC-specific probing.
- **No traceroute or MTR**, for the same reason. When a check fails, Bugwatch tells you *which regions* failed and why; it cannot tell you which hop.
- **TLS certificate expiry is not enforced yet.** The `tls` check completes a real handshake — an expired, self-signed, or hostname-mismatched certificate fails it, because the handshake fails. But Workers do not expose the peer certificate, so Bugwatch cannot read the *notAfter* date and cannot warn you 14 days ahead. `minDaysUntilExpiry` is accepted and stored so nothing breaks when the runtime gains that API; until then it does nothing, and the check result says `handshake ok; expiry not inspectable on Workers`.

## Regions

Probes run from eight Cloudflare regions, addressed by location hint:

`wnam` (US West) · `enam` (US East) · `weur` (Europe West) · `eeur` (Europe East) · `apac` (Asia Pacific) · `oc` (Oceania) · `afr` (Africa) · `sam` (South America)

Location hints are a best effort, not a guarantee — Cloudflare places the probe near the hint, not at a fixed address. Pick regions your users are actually in; three is a good default (`["wnam", "weur", "apac"]`).

Each region's schedule is offset by a **stable jitter of ±10 % of the interval**, derived from a hash of (monitor id, region). Two consequences worth knowing: a monitor on a 60-second interval is probed somewhere in a 12-second spread rather than by eight simultaneous requests, so you do not see a synthetic traffic spike every minute; and the offset does not move between checks, so the interval between two checks *from the same region* stays even.

## Consensus: what pages and what does not

Every region keeps a small ring of recent pass/fail results. Two thresholds and a quorum turn those into one monitor state:

- **`failureThreshold`** (default 2) — consecutive failures before *that region* counts as failing.
- **`recoveryThreshold`** (default 2) — consecutive passes before that region counts as passing again.
- **`quorum`** — how many regions must agree. Defaults to `ceil(regions / 2)`: 2 of 3, 3 of 5, 4 of 8.

| State | When | Pages? |
|---|---|---|
| `pending` | Created, no result yet | no |
| `up` | Quorum of regions passing | no |
| `degraded` | At least one region failing, fewer than quorum | **no** |
| `down` | Quorum of regions failing | **yes** |
| `paused` | You paused it | no |

`degraded` is the whole point of the model. One POP with a bad path to your origin, or one region hitting a cold cache, produces a red region and nothing else: the dashboard shows it, the timeline records it, nobody's phone rings. Recovery is symmetric and deliberately conservative — from `down`, a quorum of regions must *pass* before the monitor returns to `up`, so a single region flapping green cannot close an incident.

Raise `failureThreshold` for an endpoint that is legitimately slow at times; raise `quorum` for something you only care about when it is broken everywhere; lower `quorum` to 1 for a monitor where any regional failure is real (a CDN edge, say).

## Notify, or escalate

A monitor can do one of two things when it goes down, and they are not the same promise.

**Notify channels** (the default) sends one message per state change to the channels you pick. Nothing chases anybody: if the person who sees the Slack message is asleep, the outage waits.

**Escalate through a service** attaches the monitor to a service, so a DOWN opens an [incident](https://docs.bugwatch.io/product/incidents.md) and that service's escalation policy pages whoever is on call, repeating and escalating until somebody acknowledges. Set it in the monitor editor ("When it goes down") or with `serviceId` on `create_monitor` / `update_monitor`.

With a service attached:

- The monitor's own channels are **not** notified on a DOWN — being paged twice for one outage trains people to ignore the page that matters. They remain as a fallback and are used only if the escalation cannot be started at all.
- A flapping monitor rejoins the incident that is already open (`dedupKey` is the monitor), so a service that fails three times in ten minutes pages once and keeps one timeline.
- Recovery resolves the incident and says why: the timeline reads *"Resolved: the monitor recovered"*, never as though a person had judged it fixed.

## Dependencies: page once, not seven times

A monitor can declare the monitors whose failure would explain its own:

```
update_monitor(org="acme", id="01J…", dependsOn=["01J…database", "01J…gateway"])
```

While **any declared upstream is `down`**, this monitor records its own DOWN — state, timeline, status page — and **withholds the page**. A database going down should ring once, not once for every service behind it.

- **One level deep, no transitive resolution.** If A depends on B and B depends on C, a C outage does not suppress A. This is deliberate: a graph walk on the paging path is somewhere for silence to hide.
- **Two monitors that depend on each other suppress neither.** Both would otherwise withhold and nobody would be paged — silence produced purely by configuration. The reciprocal edge is ignored, so a cycle degrades to both paging: noisy and safe rather than quiet and wrong.
- **A monitor still down when its upstream recovers pages then**, saying how long the page was withheld. Otherwise declaring a dependency would be a way to silence a monitor for good.
- **A dependency on a deleted monitor is refused**, because nothing would ever clear it.
- **Failure pages.** If the dependency cannot be read at all, the monitor pages. Silence is never the fallback.

`dependsOn` is returned by `get_monitor` and on create/update. It is **omitted** from `list_monitors` rather than returned empty — `[]` would read as "none declared", and loading edges for every row is an N+1 on the paging path.

## Paging

`down` enqueues a `monitor.down` signal to the monitor's `channelIds`, delivered by the same notify Worker, with the same retry and de-duplication behaviour, as issue alerts. De-duplication is per **state transition**, so a monitor that flaps — down, recovers, down again — pages for each outage rather than once per hour. Monitor pages are never throttled; only the consensus model decides whether something pages. The recovery transition (`down` → `up`) sends `monitor.up` to the same channels so the thread closes itself. Nothing else pages: `degraded`, `pending`, and pausing are silent by design.

A monitor with no `channelIds` still records state and history — it just never notifies. That is a reasonable configuration for a monitor you are still tuning.

## Snooze: quiet the paging, not the monitor

When a monitor is down and somebody is already on it, **snooze** it for 15 minutes, an hour, 4 hours or 24 hours: **Snooze** on the monitor's page, `POST /v1/orgs/{org}/monitors/{id}/snooze` with `{"minutes": 60}`, or `snooze_monitor`. The monitor keeps checking, keeps its state and history, and **your public status page still shows what is really happening**. Only the DOWN page is held back.

Three rules make it safe:

- **Always time-boxed.** Those four durations are the only ones accepted; there is no "until I turn it back on". A permanent mute is how a page gets missed months later by somebody who forgot they set it.
- **It never swallows an outage.** If the monitor is still down when the snooze ends, it is paged at its next check, once, saying it is still down after the suppression ended. Lift it early (`DELETE …/snooze`, `unsnooze_monitor`) and the same happens straight away.
- **Always visible.** The monitor's page says "Paging snoozed until …" with an **Unsnooze** button, the monitor list marks it *snoozed*, the API reports `snoozedUntil` and `snoozedBy` while a snooze is in effect, and the timeline records who set or lifted it, and until when.

**Snooze is not pause.** Pausing stops the checks, and the status page shows the component as Unknown. Snoozing keeps everything running and only quiets the page. **And it is not acknowledge:** an incident that is already open keeps escalating until somebody acknowledges it. Snooze is about the *next* page, and acknowledging is how you answer the current one.

## Heartbeat monitors (cron jobs)

A heartbeat monitor inverts the direction: Bugwatch never calls out, your job calls in. Create one with `expectedEverySeconds` (how often the job runs) and `graceSeconds` (how late it may be before that counts as a miss), then put its ping URL at the *end* of the job, so it only fires when the work actually finished:

```sh
# nightly backup, runs at 02:00
pg_dump … | gzip > /backups/$(date +%F).sql.gz
curl -fsS https://uptime.bugwatch.io/ping/<token>
```

```
0 2 * * *  /usr/local/bin/backup.sh && curl -fsS https://uptime.bugwatch.io/ping/<token>
```

- The URL is returned as `heartbeatUrl` when you create the monitor (and by `get_monitor`); the token is the credential, so treat it like one — anyone who has it can mark your job healthy.
- `GET` and `POST` both work. `-fsS` makes curl silent on success and loud on failure, so a broken ping does not pass silently in your job's logs.
- The monitor goes `down` when `now > lastPingAt + expectedEverySeconds + graceSeconds`. Set the grace to cover normal variance in run time — for a 1-hour job, 300 seconds is a sane start. A job that has never pinged stays `pending`, not `down`; the deadline clock starts at the first ping.
- Heartbeat monitors have no regions and no latency; the deadline watcher reports as the pseudo-region `internal`.

## Plan limits

| Plan | Minimum interval | Monitors |
|---|---|---|
| Free | 5 minutes | 3 |
| Starter | 1 minute | 10 |
| Team | 30 seconds | 50 |
| Business | 30 seconds | 200 |
| Scale | 10 seconds | 500 |

The minimum interval bounds how often we probe *out* from our regions, so it applies to `http`, `tcp`, `tls` and `dns` checks. **It does not apply to heartbeats**: there the interval is how often *your* job calls in, which costs us nothing outbound, so an every-minute cron can say so on any plan (the floor there is 10 seconds, the schema minimum).

Creating or updating a monitor below the applicable minimum returns `400 {"error": "interval below plan minimum", "minIntervalSeconds": …, "plan": …}`; exceeding the monitor count returns `402 {"error": …, "limit": …, "plan": …}`. Checks do not consume your event allowance — monitors are counted, not metered ([Billing & quotas](https://docs.bugwatch.io/account/billing-and-quotas.md)).

## REST routes and tools

All org-level, under `/v1/orgs/{org}/monitors`; reads need `org:read`, writes need `alert:write`.

| Method | Path | Tool |
|---|---|---|
| `GET` | `/v1/orgs/{org}/monitors` | `list_monitors` |
| `POST` | `/v1/orgs/{org}/monitors` | `create_monitor` |
| `GET` | `/v1/orgs/{org}/monitors/{id}` | `get_monitor` |
| `PATCH` | `/v1/orgs/{org}/monitors/{id}` | `update_monitor` |
| `DELETE` | `/v1/orgs/{org}/monitors/{id}` | `delete_monitor` |
| `POST` | `/v1/orgs/{org}/monitors/{id}/pause` | `pause_monitor` |
| `POST` | `/v1/orgs/{org}/monitors/{id}/snooze` | `snooze_monitor` |
| `DELETE` | `/v1/orgs/{org}/monitors/{id}/snooze` | `unsnooze_monitor` |
| `POST` | `/v1/orgs/{org}/monitors/{id}/resume` | `resume_monitor` |
| `GET` | `/v1/orgs/{org}/monitors/{id}/history?since=24&interval=5` | `get_monitor_history` |
| `GET` | `/v1/orgs/{org}/monitors/{id}/events` | `list_monitor_events` |

`lastLatencyMs` on a monitor is the **last passing check's** response time, and `null` when there isn't one — a failing check or a heartbeat, both of which have no response time to report (a refused connection "fails" in single-digit milliseconds; that is how long we waited, not how fast the service is). The dashboard renders `null` as `—`.

`history` returns `series` (per bucket: `checks`, `passed`, `p50`, `p95`) and `regions` (the same per region over the whole window). `checks` and `passed` count every check; **`p50` and `p95` are computed over passing checks only**, so a failed check's 0 ms does not drag the percentile down while a monitor is down. A bucket or region where nothing passed has `passed = 0` and no meaningful latency — the dashboard shows those as `—` rather than as 0 ms. `interval` is minutes, one of 1, 5, 15, 60, 1440. Like every analytics-backed response these are **sampling-weighted estimates** — accurate for uptime percentage and latency percentiles, not a ledger of individual checks.

`events` returns the state-transition timeline: `fromState`, `toState`, `regionsFailing` of `regionsTotal`, a short `detail` (`status 503`, `timeout`, `keyword missing`, `last ping 4210s ago`), and `at`.

The dashboard page for a monitor is `https://app.bugwatch.io/o/{org}/monitors/{id}` — the same link every tool returns.
