Uptime monitors
For agents:
list_monitors(org="acme")for the fleet,get_monitor(org="acme", id="01J…")for one,get_monitor_history(org="acme", id="01J…", since=24, interval=5)for uptime and latency,list_monitor_events(org="acme", id="01J…")for the incident timeline. Writes —create_monitor,update_monitor,pause_monitor,resume_monitor,snooze_monitor,unsnooze_monitor,delete_monitor— needalert:write. Example:create_monitor(org="acme", name="Checkout API", check={type:"http", url:"https://api.example.com/health", keyword:"ok"}, intervalSeconds=60, regions=["wnam","weur","apac"]).
A monitor is a check plus a schedule plus the regions that run it. Monitors are organization-level, not per-project: they page notification channels, the same channels alert rules use (Alerts).
Check types
| Type | What it asserts | Key fields |
|---|---|---|
http |
A request returns a status inside expectedStatus (default [200, 399]), optionally contains keyword, does not contain keywordAbsent, and answers within maxLatencyMs |
url, method, headers, body, expectedStatus, keyword, keywordAbsent, followRedirects, timeoutMs, maxLatencyMs, auth (basic or bearer) |
tcp |
A TCP connection to host:port is accepted within timeoutMs |
host, port, timeoutMs |
tls |
A TLS handshake with host:port completes within timeoutMs |
host, port (default 443), minDaysUntilExpiry, timeoutMs |
dns |
A DNS-over-HTTPS query for name/recordType returns NOERROR with at least one answer, and — when expected is set — an answer equal to it |
name, recordType (A, AAAA, CNAME, MX, TXT, NS), expected, resolver (an https:// DNS-over-HTTPS JSON endpoint; default Cloudflare), timeoutMs |
heartbeat |
Your job pinged the monitor's URL within expectedEverySeconds + graceSeconds |
expectedEverySeconds, graceSeconds |
HTTP credentials (auth, any header that looks like authorization, cookie, token or key, and a password in the URL) are sealed with AES-GCM in the database and returned masked as ••••••. Only the uptime Worker opens them, to run the check. Re-send the real value to change it; sending the mask back is not a way to read it.
What Bugwatch cannot check
Probes run inside Cloudflare Workers, which have outbound fetch and TCP sockets and nothing lower. So, plainly:
- No ICMP ping. There are no raw sockets, so there is no
pingcheck and no packet-loss metric. A TCP check against a port the host actually serves is the closest equivalent, and it is a better signal anyway. - Bodies are read up to 1 MB. A
keywordcheck looks at the first megabyte of the response and stops reading there, so put the keyword near the top of a large page. - No UDP, so no DNS-over-UDP (the
dnscheck uses DNS-over-HTTPS against a resolver you choose), no NTP, no syslog, no QUIC-specific probing. - No traceroute or MTR, for the same reason. When a check fails, Bugwatch tells you which regions failed and why; it cannot tell you which hop.
- TLS certificate expiry is not enforced yet. The
tlscheck completes a real handshake — an expired, self-signed, or hostname-mismatched certificate fails it, because the handshake fails. But Workers do not expose the peer certificate, so Bugwatch cannot read the notAfter date and cannot warn you 14 days ahead.minDaysUntilExpiryis accepted and stored so nothing breaks when the runtime gains that API; until then it does nothing, and the check result sayshandshake ok; expiry not inspectable on Workers.
Regions
Probes run from eight Cloudflare regions, addressed by location hint:
wnam (US West) · enam (US East) · weur (Europe West) · eeur (Europe East) · apac (Asia Pacific) · oc (Oceania) · afr (Africa) · sam (South America)
Location hints are a best effort, not a guarantee — Cloudflare places the probe near the hint, not at a fixed address. Pick regions your users are actually in; three is a good default (["wnam", "weur", "apac"]).
Each region's schedule is offset by a stable jitter of ±10 % of the interval, derived from a hash of (monitor id, region). Two consequences worth knowing: a monitor on a 60-second interval is probed somewhere in a 12-second spread rather than by eight simultaneous requests, so you do not see a synthetic traffic spike every minute; and the offset does not move between checks, so the interval between two checks from the same region stays even.
Consensus: what pages and what does not
Every region keeps a small ring of recent pass/fail results. Two thresholds and a quorum turn those into one monitor state:
failureThreshold(default 2) — consecutive failures before that region counts as failing.recoveryThreshold(default 2) — consecutive passes before that region counts as passing again.quorum— how many regions must agree. Defaults toceil(regions / 2): 2 of 3, 3 of 5, 4 of 8.
| State | When | Pages? |
|---|---|---|
pending |
Created, no result yet | no |
up |
Quorum of regions passing | no |
degraded |
At least one region failing, fewer than quorum | no |
down |
Quorum of regions failing | yes |
paused |
You paused it | no |
degraded is the whole point of the model. One POP with a bad path to your origin, or one region hitting a cold cache, produces a red region and nothing else: the dashboard shows it, the timeline records it, nobody's phone rings. Recovery is symmetric and deliberately conservative — from down, a quorum of regions must pass before the monitor returns to up, so a single region flapping green cannot close an incident.
Raise failureThreshold for an endpoint that is legitimately slow at times; raise quorum for something you only care about when it is broken everywhere; lower quorum to 1 for a monitor where any regional failure is real (a CDN edge, say).
Notify, or escalate
A monitor can do one of two things when it goes down, and they are not the same promise.
Notify channels (the default) sends one message per state change to the channels you pick. Nothing chases anybody: if the person who sees the Slack message is asleep, the outage waits.
Escalate through a service attaches the monitor to a service, so a DOWN opens an incident and that service's escalation policy pages whoever is on call, repeating and escalating until somebody acknowledges. Set it in the monitor editor ("When it goes down") or with serviceId on create_monitor / update_monitor.
With a service attached:
- The monitor's own channels are not notified on a DOWN — being paged twice for one outage trains people to ignore the page that matters. They remain as a fallback and are used only if the escalation cannot be started at all.
- A flapping monitor rejoins the incident that is already open (
dedupKeyis the monitor), so a service that fails three times in ten minutes pages once and keeps one timeline. - Recovery resolves the incident and says why: the timeline reads "Resolved: the monitor recovered", never as though a person had judged it fixed.
Dependencies: page once, not seven times
A monitor can declare the monitors whose failure would explain its own:
update_monitor(org="acme", id="01J…", dependsOn=["01J…database", "01J…gateway"])
While any declared upstream is down, this monitor records its own DOWN — state, timeline, status page — and withholds the page. A database going down should ring once, not once for every service behind it.
- One level deep, no transitive resolution. If A depends on B and B depends on C, a C outage does not suppress A. This is deliberate: a graph walk on the paging path is somewhere for silence to hide.
- Two monitors that depend on each other suppress neither. Both would otherwise withhold and nobody would be paged — silence produced purely by configuration. The reciprocal edge is ignored, so a cycle degrades to both paging: noisy and safe rather than quiet and wrong.
- A monitor still down when its upstream recovers pages then, saying how long the page was withheld. Otherwise declaring a dependency would be a way to silence a monitor for good.
- A dependency on a deleted monitor is refused, because nothing would ever clear it.
- Failure pages. If the dependency cannot be read at all, the monitor pages. Silence is never the fallback.
dependsOn is returned by get_monitor and on create/update. It is omitted from list_monitors rather than returned empty — [] would read as "none declared", and loading edges for every row is an N+1 on the paging path.
Paging
down enqueues a monitor.down signal to the monitor's channelIds, delivered by the same notify Worker, with the same retry and de-duplication behaviour, as issue alerts. De-duplication is per state transition, so a monitor that flaps — down, recovers, down again — pages for each outage rather than once per hour. Monitor pages are never throttled; only the consensus model decides whether something pages. The recovery transition (down → up) sends monitor.up to the same channels so the thread closes itself. Nothing else pages: degraded, pending, and pausing are silent by design.
A monitor with no channelIds still records state and history — it just never notifies. That is a reasonable configuration for a monitor you are still tuning.
Snooze: quiet the paging, not the monitor
When a monitor is down and somebody is already on it, snooze it for 15 minutes, an hour, 4 hours or 24 hours: Snooze on the monitor's page, POST /v1/orgs/{org}/monitors/{id}/snooze with {"minutes": 60}, or snooze_monitor. The monitor keeps checking, keeps its state and history, and your public status page still shows what is really happening. Only the DOWN page is held back.
Three rules make it safe:
- Always time-boxed. Those four durations are the only ones accepted; there is no "until I turn it back on". A permanent mute is how a page gets missed months later by somebody who forgot they set it.
- It never swallows an outage. If the monitor is still down when the snooze ends, it is paged at its next check, once, saying it is still down after the suppression ended. Lift it early (
DELETE …/snooze,unsnooze_monitor) and the same happens straight away. - Always visible. The monitor's page says "Paging snoozed until …" with an Unsnooze button, the monitor list marks it snoozed, the API reports
snoozedUntilandsnoozedBywhile a snooze is in effect, and the timeline records who set or lifted it, and until when.
Snooze is not pause. Pausing stops the checks, and the status page shows the component as Unknown. Snoozing keeps everything running and only quiets the page. And it is not acknowledge: an incident that is already open keeps escalating until somebody acknowledges it. Snooze is about the next page, and acknowledging is how you answer the current one.
Heartbeat monitors (cron jobs)
A heartbeat monitor inverts the direction: Bugwatch never calls out, your job calls in. Create one with expectedEverySeconds (how often the job runs) and graceSeconds (how late it may be before that counts as a miss), then put its ping URL at the end of the job, so it only fires when the work actually finished:
# nightly backup, runs at 02:00
pg_dump … | gzip > /backups/$(date +%F).sql.gz
curl -fsS https://uptime.bugwatch.io/ping/<token>
0 2 * * * /usr/local/bin/backup.sh && curl -fsS https://uptime.bugwatch.io/ping/<token>
- The URL is returned as
heartbeatUrlwhen you create the monitor (and byget_monitor); the token is the credential, so treat it like one — anyone who has it can mark your job healthy. GETandPOSTboth work.-fsSmakes curl silent on success and loud on failure, so a broken ping does not pass silently in your job's logs.- The monitor goes
downwhennow > lastPingAt + expectedEverySeconds + graceSeconds. Set the grace to cover normal variance in run time — for a 1-hour job, 300 seconds is a sane start. A job that has never pinged stayspending, notdown; the deadline clock starts at the first ping. - Heartbeat monitors have no regions and no latency; the deadline watcher reports as the pseudo-region
internal.
Plan limits
| Plan | Minimum interval | Monitors |
|---|---|---|
| Free | 5 minutes | 3 |
| Starter | 1 minute | 10 |
| Team | 30 seconds | 50 |
| Business | 30 seconds | 200 |
| Scale | 10 seconds | 500 |
The minimum interval bounds how often we probe out from our regions, so it applies to http, tcp, tls and dns checks. It does not apply to heartbeats: there the interval is how often your job calls in, which costs us nothing outbound, so an every-minute cron can say so on any plan (the floor there is 10 seconds, the schema minimum).
Creating or updating a monitor below the applicable minimum returns 400 {"error": "interval below plan minimum", "minIntervalSeconds": …, "plan": …}; exceeding the monitor count returns 402 {"error": …, "limit": …, "plan": …}. Checks do not consume your event allowance — monitors are counted, not metered (Billing & quotas).
REST routes and tools
All org-level, under /v1/orgs/{org}/monitors; reads need org:read, writes need alert:write.
| Method | Path | Tool |
|---|---|---|
GET |
/v1/orgs/{org}/monitors |
list_monitors |
POST |
/v1/orgs/{org}/monitors |
create_monitor |
GET |
/v1/orgs/{org}/monitors/{id} |
get_monitor |
PATCH |
/v1/orgs/{org}/monitors/{id} |
update_monitor |
DELETE |
/v1/orgs/{org}/monitors/{id} |
delete_monitor |
POST |
/v1/orgs/{org}/monitors/{id}/pause |
pause_monitor |
POST |
/v1/orgs/{org}/monitors/{id}/snooze |
snooze_monitor |
DELETE |
/v1/orgs/{org}/monitors/{id}/snooze |
unsnooze_monitor |
POST |
/v1/orgs/{org}/monitors/{id}/resume |
resume_monitor |
GET |
/v1/orgs/{org}/monitors/{id}/history?since=24&interval=5 |
get_monitor_history |
GET |
/v1/orgs/{org}/monitors/{id}/events |
list_monitor_events |
lastLatencyMs on a monitor is the last passing check's response time, and null when there isn't one — a failing check or a heartbeat, both of which have no response time to report (a refused connection "fails" in single-digit milliseconds; that is how long we waited, not how fast the service is). The dashboard renders null as —.
history returns series (per bucket: checks, passed, p50, p95) and regions (the same per region over the whole window). checks and passed count every check; p50 and p95 are computed over passing checks only, so a failed check's 0 ms does not drag the percentile down while a monitor is down. A bucket or region where nothing passed has passed = 0 and no meaningful latency — the dashboard shows those as — rather than as 0 ms. interval is minutes, one of 1, 5, 15, 60, 1440. Like every analytics-backed response these are sampling-weighted estimates — accurate for uptime percentage and latency percentiles, not a ledger of individual checks.
events returns the state-transition timeline: fromState, toState, regionsFailing of regionsTotal, a short detail (status 503, timeout, keyword missing, last ping 4210s ago), and at.
The dashboard page for a monitor is https://app.bugwatch.io/o/{org}/monitors/{id} — the same link every tool returns.