bugwatch docs

Migrating from Opsgenie or PagerDuty

For agents: diff_imported_schedule(org="acme", id="01J…", provider="opsgenie"|"pagerduty", source={…}, userMap={"ada@acme.test": "usr_ada"}). Read-only, org:read. Returns import warnings first, then the coverage differences.

Moving a rotation is when an on-call programme is most likely to break, and it breaks quietly: somebody is paged at 3am for six months, a cutover moves a handoff by an hour, and nobody finds out until an incident lands in the hour nobody covers.

So the migration path is a diff you sign off on, not an import button. Nothing here activates a rotation.

Point your PagerDuty integrations here

The slowest part of leaving PagerDuty is rarely the schedules — it is the dozens of things already sending to events.pagerduty.com/v2/enqueue: Datadog monitors, Grafana, cron wrappers, a Terraform module somebody wrote years ago.

Bugwatch accepts PagerDuty Events API v2 at its own endpoint, in exactly PD's shape, so each of those changes a URL and a routing key rather than being rewritten.

  1. POST /v1/orgs/{org}/services/{id}/inbound-key mints a routing key for a service. It is returned once — it is a credential, and whoever holds it can open an incident on that service. Re-minting replaces the old key; integrations still pointed at it start getting 400 routing_key not found, which is the visible failure a rotation should have. Signed-in users only, never an API token.
  2. In each integration, change the events URL to POST {uptime origin}/v2/enqueue and the routing key to the one you just minted.

Nothing else changes. event_action of trigger, acknowledge and resolve all behave as they do in PD, dedup_key addresses the same incident across all three, and the response body is the {"status": "success", "message": "Event processed", "dedup_key": "…"} your tooling already parses.

PD payload.severity Bugwatch
critical sev1 — pages
error sev2 — pages
warning sev3
info sev4

Three behaviours are worth knowing because they are deliberate:

  • An acknowledge or resolve for an incident that does not exist succeeds. PD does the same. An integration that acknowledges before its trigger has landed is not in error, and failing it would turn ordinary retry ordering into paging noise.
  • An unknown routing key is 400, not 401. Also PD's behaviour, and it matters: an integration retrying a 401 forever behaves very differently from one that gives up on a 400.
  • A dedup_key is scoped to the service, so two teams that both named their alert disk-full never share an incident.

images, links and client_url are accepted and discarded — they are display sugar for PD's own interface, and storing them without showing them would be worse than saying so.

How it works

Build the schedule in Bugwatch, then compare it against the Opsgenie one:

curl -X POST https://api.bugwatch.io/v1/orgs/acme/schedules/01J.../import-diff \
  -H "authorization: Bearer $BUGWATCH_TOKEN" \
  -H "content-type: application/json" \
  -d '{
    "provider": "opsgenie",
    "source": { "id": "sch-og", "timezone": "Europe/London", "rotations": [ ... ] },
    "userMap": { "ada@acme.test": "usr_ada", "ben@acme.test": "usr_ben" },
    "days": 30
  }'

source is the provider's schedule payload, as their schedules API returns it. provider is stated rather than sniffed: the two shapes are similar enough that a guess would sometimes be wrong, and being wrong means diffing against a schedule that was never read correctly — which reads as a real disagreement rather than as an error. You supply it; we do not hold a token for your Opsgenie account. The endpoint stores nothing and changes nothing — it cannot activate a rotation, which is the surest way to honour the sign-off requirement.

The comparison runs both schedules through the same resolution engine and compares who is on call at every instant, not how the rotations are written. Two rotations expressed completely differently and putting the same person on call every minute are not a difference; it is the minutes that matter.

Read the warnings first

import.warnings names everything the mapping could not carry across. Read it before the deltas — a difference caused by a dropped participant is not a scheduling disagreement, and reading them the other way round leads you to fix the wrong thing.

Code What it means
non_user_participant A team or escalation was in the rotation order. It cannot be imported, so the rotation is short by one and everybody after it lands on a different day
unsupported_rotation_type A rotation type with no equivalent. It covers nothing until rebuilt by hand
empty_rotation Nothing importable was left in it
restriction_spans_days An Opsgenie window running e.g. Monday→Friday. Ours repeat, so it was imported as the start day only
unsupported_restriction A window wrapping past midnight. Splitting it would be guesswork about intent
schedule_disabled The Opsgenie schedule is disabled; the coverage shown is what it would be
overrides_not_fetched No overrides were supplied for an Opsgenie schedule. They come from a separate endpoint, so the export may be missing every swapped shift
override_scoped_to_rotations An Opsgenie override that applied to some rotations only. Ours cover the whole schedule, so it was imported wider than it was written
non_user_override An override assigned to a team rather than a person. Not imported — inventing somebody is worse
unusable_override An override with no parseable window or user

An invalid timezone is refused outright rather than warned about: every handoff would land at the wrong instant.

PagerDuty specifics

Three differences between their model and ours are places a handoff can move silently, so each is handled explicitly:

  • Weekday numbering. Theirs is ISO-8601 (1 = Monday … 7 = Sunday); ours is 0 = Sunday. Restricted windows are renumbered on import.
  • The rotation anchor. Turns are computed from rotation_virtual_start, not start — start is only when the layer became active, and anchoring on it puts everybody on the wrong turn.
  • Layer order. PagerDuty lists layers newest-first with no precedence field; ours is "higher position wins", so the order is inverted on import.

Overrides, from both

An override is somebody's swapped weekend, and losing it silently is how a migration pages the person who arranged not to be paged. Both providers' overrides are imported — but Opsgenie's arrive differently, and that difference is yours to handle:

PagerDuty returns overrides inside the schedule payload, so they come across with everything else.

Opsgenie serves them from a separate endpoint — GET /v2/schedules/{id}/overrides. Fetch it and pass the array as overrides alongside the schedule. If you do not, the import warns overrides_not_fetched rather than assuming there were none: "we looked and there are none" and "we did not look" are opposite facts for whoever signs off. Pass overridesFetched: true with an empty array to say you looked.

One Opsgenie override cannot be carried across faithfully. Theirs can apply to specific rotations; ours cover the whole schedule. Such an override is imported anyway — dropping it would page the person it excuses — and warned as override_scoped_to_rotations, because it now excuses them from every rotation in that window. The warning is what makes the resulting coverage delta explicable instead of mysterious.

A payload with a missing or unreadable timezone is refused outright rather than warned about — a missing one is the worse case, because it would otherwise be read in the runtime's zone and move every handoff without failing.

What a difference is

kind Meaning
only_theirs They have cover and we have none. The dangerous direction — this is time that goes unpaged after cutover
different_user Both sides name somebody, and it is not the same person
only_ours We cover a window they leave uncovered. Worth seeing, rarely harmful

A source user with no userMap entry is a difference with unmappedUser: true, never an assumed match. Two ids that look alike are not evidence of the same person, and an import signed off on an unmatched roster is an import nobody checked.

The number to sign off on

unexplainedCount reaching zero is the goal. Alongside it, uncoveredByUsMs is the milliseconds where the source has cover and we do not.

boundaryToleranceMs exists because two systems rarely agree to the millisecond on a handoff — differences shorter than it move to suppressed rather than deltas. Two rules keep it honest: suppressed differences are still returned in full, so the setting is auditable, and it never reduces uncoveredByUsMs. Time nobody would be paged for is not a rounding question.

In the dashboard

On-call → a schedule → Compare with Opsgenie. Paste the payload, match the people it finds, and read the report:

  • Would go unpaged — hours the Opsgenie schedule covers and this one does not. The only figure here that means somebody does not get woken up.
  • Differences — reaching zero is the sign-off.

There is no "apply" button, and that is deliberate: the endpoint behind this screen cannot activate a rotation.

Escalation policies

Schedules answer who is on call. Policies answer what happens when they do not pick up, and both providers model one the way we do — ordered rungs, a delay, a set of targets. The mapping is mostly faithful, so what matters is the short list of places it is not. Each is warned rather than guessed at.

Code What it means
unsupported_target A team, service or webhook target. We page people and rotations, so that rung is short by one
empty_rule Nothing importable was left on the rung — it pages nobody
no_rules The policy has no rungs at all
urgency_assumed Neither provider carries urgency on the policy — PagerDuty puts it on the service, Opsgenie on the alert. Every rung imports as high, which is the safe direction, but it is an assumption
delay_unit_unsupported An Opsgenie delay in hours or days. Refused rather than converted — wrong by 60× on a paging delay is worse than a blank you fill in
notify_type_narrowed An Opsgenie notifyType that selected who inside a rotation. Imported as the whole rotation, which pages more people than it did
repeat_not_carried The policy repeats with no wait interval, which would re-page immediately

Two model differences are handled by name rather than left to chance:

  • PagerDuty's num_loops counts the first run; our repeatCount counts the passes after it. Off by one here is a chain that pages everybody one extra time on every incident.
  • Opsgenie's delay is the wait before a rung fires; ours is the wait after it. Imported unshifted, the entire chain runs a rung out of time — the first page late, the last one early.

No fallback is ever invented. Our uncovered-rotation fallback has no equivalent in either provider, and guessing one would page somebody who never agreed to it.

Services and their integrations

A schedule says who is on call. A policy says what happens when they do not pick up. A service is what an alert arrives at, and its integrations are the routing keys that carry it there — so this import produces exactly what the repoint checklist below consumes.

The shape maps cleanly. What does not is that both providers let a service exist in states we have no equivalent for, and each one imported quietly produces a service that looks configured and pages nobody:

Code What it means
no_escalation_policy PagerDuty: no policy. Opsgenie: no team, which is where their escalation lives. Either way, alerts arrive, an incident opens, and nobody is told
service_disabled Disabled in PagerDuty. Imported anyway — leaving it out makes the repoint checklist look complete while an integration nobody accounted for still points at the old system
no_integrations Nothing routes to it, so nothing will after the cutover either
integration_disabled Disabled in Opsgenie, and kept off the repoint list: repointing it here would quietly turn back on something somebody switched off there
unsupported_integration A cloud connector or vendor inbound with no routing key we can issue. It must be rebuilt, not repointed
timeout_not_carried PagerDuty's auto_resolve_timeout: 0 means their account default, not immediately. Ours would resolve the incident on the spot, so it is imported as never and left for you

Opsgenie has no per-service auto-resolve or acknowledgement timeout, and none is invented: guessing an account default from a payload that does not contain it would close incidents nobody asked to close.

The routing-key repoint

Everything up to here is reversible and invisible: importing a schedule, diffing coverage, drafting a policy. Repointing an integration's routing key is the moment alerts stop arriving at the old system and start arriving here — done by a person, in somebody else's console, one integration at a time, usually at the end of a long day.

So the checklist is an order, and the order is the point:

  1. Nothing moves before sign-off. Not before coverage is signed off, and not while the diff still shows time nobody here would be paged for. A cutover onto a rotation with a hole pages nobody, and nobody finds out until it is somebody's night.
  2. One integration, then a verified page, then the rest. Not a test page from our side — an alert that actually came through the repointed integration and reached a phone. A checklist that repoints twelve and then checks is one that finds out about a mistake twelve times.
  3. The old system keeps receiving until then. Running both for an evening is noise; running neither is an outage nobody is told about.
  4. Decommissioning is last, because it is the only irreversible step. Until it is done, a mistake anywhere above costs duplicate pages rather than silence.

Two kinds of integration are kept off the list entirely:

Why
Its service was never mapped Repointing it is not "nothing happens" — alerts arrive here and reach nobody, while the old system has stopped getting them
Email integrations They carry a generated address, not a routing key. That needs a new address and a change at the sender, which is a different job with a different owner

What is still yours to do

The import is complete for both providers — schedules, rotations, overrides, escalation policies, services and integrations — and the repoint checklist covers the cutover (#63).

What no import can do for you is the part that needs a person: reading the warnings, mapping their users onto yours, signing off the coverage diff, and making the repoint changes in your provider's console. Nothing here activates a rotation or moves a routing key on its own, by design.