# Migrating from Opsgenie or PagerDuty

> **For agents:** `diff_imported_schedule(org="acme", id="01J…", provider="opsgenie"|"pagerduty", source={…}, userMap={"ada@acme.test": "usr_ada"})`. Read-only, `org:read`. Returns import warnings first, then the coverage differences.

Moving a rotation is when an on-call programme is most likely to break, and it breaks quietly: somebody is paged at 3am for six months, a cutover moves a handoff by an hour, and nobody finds out until an incident lands in the hour nobody covers.

So the migration path is a **diff you sign off on**, not an import button. Nothing here activates a rotation.


## Point your PagerDuty integrations here

The slowest part of leaving PagerDuty is rarely the schedules — it is the dozens of things already sending to `events.pagerduty.com/v2/enqueue`: Datadog monitors, Grafana, cron wrappers, a Terraform module somebody wrote years ago.

Bugwatch accepts **PagerDuty Events API v2** at its own endpoint, in exactly PD's shape, so each of those changes a URL and a routing key rather than being rewritten.

1. `POST /v1/orgs/{org}/services/{id}/inbound-key` mints a routing key for a service. It is returned **once** — it is a credential, and whoever holds it can open an incident on that service. Re-minting replaces the old key; integrations still pointed at it start getting `400 routing_key not found`, which is the visible failure a rotation should have. Signed-in users only, never an API token.
2. In each integration, change the events URL to `POST {uptime origin}/v2/enqueue` and the routing key to the one you just minted.

Nothing else changes. `event_action` of `trigger`, `acknowledge` and `resolve` all behave as they do in PD, `dedup_key` addresses the same incident across all three, and the response body is the `{"status": "success", "message": "Event processed", "dedup_key": "…"}` your tooling already parses.

| PD `payload.severity` | Bugwatch |
|---|---|
| `critical` | `sev1` — pages |
| `error` | `sev2` — pages |
| `warning` | `sev3` |
| `info` | `sev4` |

Three behaviours are worth knowing because they are deliberate:

- **An `acknowledge` or `resolve` for an incident that does not exist succeeds.** PD does the same. An integration that acknowledges before its trigger has landed is not in error, and failing it would turn ordinary retry ordering into paging noise.
- **An unknown routing key is `400`, not `401`.** Also PD's behaviour, and it matters: an integration retrying a `401` forever behaves very differently from one that gives up on a `400`.
- **A `dedup_key` is scoped to the service**, so two teams that both named their alert `disk-full` never share an incident.

`images`, `links` and `client_url` are accepted and discarded — they are display sugar for PD's own interface, and storing them without showing them would be worse than saying so.

## How it works

Build the schedule in Bugwatch, then compare it against the Opsgenie one:

```bash
curl -X POST https://api.bugwatch.io/v1/orgs/acme/schedules/01J.../import-diff \
  -H "authorization: Bearer $BUGWATCH_TOKEN" \
  -H "content-type: application/json" \
  -d '{
    "provider": "opsgenie",
    "source": { "id": "sch-og", "timezone": "Europe/London", "rotations": [ ... ] },
    "userMap": { "ada@acme.test": "usr_ada", "ben@acme.test": "usr_ben" },
    "days": 30
  }'
```

`source` is the provider's schedule payload, as their schedules API returns it. `provider` is stated rather than sniffed: the two shapes are similar enough that a guess would sometimes be wrong, and being wrong means diffing against a schedule that was never read correctly — which reads as a real disagreement rather than as an error. **You supply it; we do not hold a token for your Opsgenie account.** The endpoint stores nothing and changes nothing — it cannot activate a rotation, which is the surest way to honour the sign-off requirement.

The comparison runs both schedules through the same resolution engine and compares **who is on call at every instant**, not how the rotations are written. Two rotations expressed completely differently and putting the same person on call every minute are not a difference; it is the minutes that matter.

## Read the warnings first

`import.warnings` names everything the mapping could not carry across. Read it before the deltas — a difference caused by a dropped participant is not a scheduling disagreement, and reading them the other way round leads you to fix the wrong thing.

| Code | What it means |
|---|---|
| `non_user_participant` | A team or escalation was in the rotation order. It cannot be imported, so the rotation is **short by one** and everybody after it lands on a different day |
| `unsupported_rotation_type` | A rotation type with no equivalent. It covers nothing until rebuilt by hand |
| `empty_rotation` | Nothing importable was left in it |
| `restriction_spans_days` | An Opsgenie window running e.g. Monday→Friday. Ours repeat, so it was imported as the start day only |
| `unsupported_restriction` | A window wrapping past midnight. Splitting it would be guesswork about intent |
| `schedule_disabled` | The Opsgenie schedule is disabled; the coverage shown is what it *would* be |
| `overrides_not_fetched` | No overrides were supplied for an Opsgenie schedule. They come from a **separate endpoint**, so the export may be missing every swapped shift |
| `override_scoped_to_rotations` | An Opsgenie override that applied to some rotations only. Ours cover the whole schedule, so it was imported **wider** than it was written |
| `non_user_override` | An override assigned to a team rather than a person. Not imported — inventing somebody is worse |
| `unusable_override` | An override with no parseable window or user |

An invalid timezone is refused outright rather than warned about: every handoff would land at the wrong instant.

### PagerDuty specifics

Three differences between their model and ours are places a handoff can move silently, so each is handled explicitly:

- **Weekday numbering.** Theirs is ISO-8601 (1 = Monday … 7 = Sunday); ours is 0 = Sunday. Restricted windows are renumbered on import.
- **The rotation anchor.** Turns are computed from `rotation_virtual_start`, not `start` — `start` is only when the layer became active, and anchoring on it puts everybody on the wrong turn.
- **Layer order.** PagerDuty lists layers newest-first with no precedence field; ours is "higher position wins", so the order is inverted on import.

### Overrides, from both

An override is somebody's swapped weekend, and losing it silently is how a migration pages the person who arranged not to be paged. Both providers' overrides are imported — but Opsgenie's arrive differently, and that difference is yours to handle:

**PagerDuty** returns overrides inside the schedule payload, so they come across with everything else.

**Opsgenie serves them from a separate endpoint** — `GET /v2/schedules/{id}/overrides`. Fetch it and pass the array as `overrides` alongside the schedule. If you do not, the import warns `overrides_not_fetched` rather than assuming there were none: *"we looked and there are none"* and *"we did not look"* are opposite facts for whoever signs off. Pass `overridesFetched: true` with an empty array to say you looked.

One Opsgenie override cannot be carried across faithfully. Theirs can apply to **specific rotations**; ours cover the whole schedule. Such an override is imported anyway — dropping it would page the person it excuses — and warned as `override_scoped_to_rotations`, because it now excuses them from every rotation in that window. The warning is what makes the resulting coverage delta explicable instead of mysterious.

A payload with a missing or unreadable timezone is refused outright rather than warned about — a missing one is the worse case, because it would otherwise be read in the runtime's zone and move every handoff without failing.

## What a difference is

| `kind` | Meaning |
|---|---|
| `only_theirs` | **They have cover and we have none.** The dangerous direction — this is time that goes unpaged after cutover |
| `different_user` | Both sides name somebody, and it is not the same person |
| `only_ours` | We cover a window they leave uncovered. Worth seeing, rarely harmful |

A source user with no `userMap` entry is a difference with `unmappedUser: true`, never an assumed match. Two ids that look alike are not evidence of the same person, and an import signed off on an unmatched roster is an import nobody checked.

## The number to sign off on

`unexplainedCount` reaching **zero** is the goal. Alongside it, `uncoveredByUsMs` is the milliseconds where the source has cover and we do not.

`boundaryToleranceMs` exists because two systems rarely agree to the millisecond on a handoff — differences shorter than it move to `suppressed` rather than `deltas`. Two rules keep it honest: suppressed differences are still returned in full, so the setting is auditable, and it **never reduces `uncoveredByUsMs`**. Time nobody would be paged for is not a rounding question.

## In the dashboard

On-call → a schedule → **Compare with Opsgenie**. Paste the payload, match the people it finds, and read the report:

- **Would go unpaged** — hours the Opsgenie schedule covers and this one does not. The only figure here that means somebody does not get woken up.
- **Differences** — reaching zero is the sign-off.

There is no "apply" button, and that is deliberate: the endpoint behind this screen cannot activate a rotation.

## Escalation policies

Schedules answer *who is on call*. Policies answer *what happens when they do not pick up*, and both providers model one the way we do — ordered rungs, a delay, a set of targets. The mapping is mostly faithful, so what matters is the short list of places it is not. Each is warned rather than guessed at.

| Code | What it means |
|---|---|
| `unsupported_target` | A team, service or webhook target. We page people and rotations, so that rung is **short by one** |
| `empty_rule` | Nothing importable was left on the rung — it pages nobody |
| `no_rules` | The policy has no rungs at all |
| `urgency_assumed` | Neither provider carries urgency on the *policy* — PagerDuty puts it on the service, Opsgenie on the alert. Every rung imports as **high**, which is the safe direction, but it is an assumption |
| `delay_unit_unsupported` | An Opsgenie delay in hours or days. Refused rather than converted — wrong by 60× on a paging delay is worse than a blank you fill in |
| `notify_type_narrowed` | An Opsgenie `notifyType` that selected *who inside* a rotation. Imported as the whole rotation, which **pages more people than it did** |
| `repeat_not_carried` | The policy repeats with no wait interval, which would re-page immediately |

Two model differences are handled by name rather than left to chance:

- **PagerDuty's `num_loops` counts the first run**; our `repeatCount` counts the passes *after* it. Off by one here is a chain that pages everybody one extra time on every incident.
- **Opsgenie's `delay` is the wait *before* a rung fires**; ours is the wait *after* it. Imported unshifted, the entire chain runs a rung out of time — the first page late, the last one early.

**No fallback is ever invented.** Our uncovered-rotation fallback has no equivalent in either provider, and guessing one would page somebody who never agreed to it.

## Services and their integrations

A schedule says who is on call. A policy says what happens when they do not pick up. A **service** is what an alert arrives at, and its integrations are the routing keys that carry it there — so this import produces exactly what the repoint checklist below consumes.

The shape maps cleanly. What does not is that both providers let a service exist in states we have no equivalent for, and each one imported quietly produces a service that looks configured and pages nobody:

| Code | What it means |
|---|---|
| `no_escalation_policy` | PagerDuty: no policy. Opsgenie: no team, which is where their escalation lives. Either way, alerts arrive, an incident opens, and **nobody is told** |
| `service_disabled` | Disabled in PagerDuty. Imported anyway — leaving it out makes the repoint checklist look complete while an integration nobody accounted for still points at the old system |
| `no_integrations` | Nothing routes to it, so nothing will after the cutover either |
| `integration_disabled` | Disabled in Opsgenie, and kept **off** the repoint list: repointing it here would quietly turn back on something somebody switched off there |
| `unsupported_integration` | A cloud connector or vendor inbound with no routing key we can issue. It must be **rebuilt**, not repointed |
| `timeout_not_carried` | PagerDuty's `auto_resolve_timeout: 0` means *their account default*, not *immediately*. Ours would resolve the incident on the spot, so it is imported as never and left for you |

Opsgenie has no per-service auto-resolve or acknowledgement timeout, and none is invented: guessing an account default from a payload that does not contain it would close incidents nobody asked to close.

## The routing-key repoint

Everything up to here is reversible and invisible: importing a schedule, diffing coverage, drafting a policy. Repointing an integration's routing key is the moment alerts stop arriving at the old system and start arriving here — done by a person, in somebody else's console, one integration at a time, usually at the end of a long day.

So the checklist is an **order**, and the order is the point:

1. **Nothing moves before sign-off.** Not before coverage is signed off, and not while the diff still shows time nobody here would be paged for. A cutover onto a rotation with a hole pages nobody, and nobody finds out until it is somebody's night.
2. **One integration, then a verified page, then the rest.** Not a test page from our side — an alert that actually came through the repointed integration and reached a phone. A checklist that repoints twelve and then checks is one that finds out about a mistake twelve times.
3. **The old system keeps receiving until then.** Running both for an evening is noise; running neither is an outage nobody is told about.
4. **Decommissioning is last**, because it is the only irreversible step. Until it is done, a mistake anywhere above costs duplicate pages rather than silence.

Two kinds of integration are kept **off** the list entirely:

| | Why |
|---|---|
| Its service was never mapped | Repointing it is not "nothing happens" — alerts arrive here and reach nobody, while the old system has stopped getting them |
| Email integrations | They carry a generated address, not a routing key. That needs a new address and a change at the sender, which is a different job with a different owner |

## What is still yours to do

The import is complete for both providers — schedules, rotations, overrides, escalation policies, services and integrations — and the repoint checklist covers the cutover ([#63](https://github.com/abdallahk/apm/issues/63)).

What no import can do for you is the part that needs a person: reading the warnings, mapping their users onto yours, signing off the coverage diff, and making the repoint changes in your provider's console. Nothing here activates a rotation or moves a routing key on its own, by design.
