bugwatch docs

Performance

For agents: query_performance(org="acme", project="web", since=24, limit=50) for the per-transaction table; get_stats(org="acme", project="web", dataset="transactions", since=24, interval=60) for volume over time; get_web_vitals(org="acme", project="web", since=24) for p75 Core Web Vitals per transaction. All are estimates — say "≈" when you quote them.

What is collected

Sentry SDKs send transactions (tracesSampleRate controls how many) — one root operation such as an HTTP request or a page load — with spans for the work inside it. Bugwatch counts each transaction as one event on the meter; spans are free. The processor stores the full transaction JSON in R2 and writes one analytics point per transaction plus up to 20 of its slowest spans (by exclusive time) to the spans dataset. Transaction names come from the SDK; use the router integrations so they are route templates (GET /orders/:id) rather than raw URLs.

Transaction overview

GET /v1/orgs/{org}/projects/{project}/transactions?since=24&limit=50 (event:read):

Field Meaning
transaction Name
throughput Estimated transaction count in the window
p50, p75, p95, p99 Duration percentiles in ms
failures Transactions whose status is not ok (or unset)

since is a window in hours, 1–2160 (90 days); limit 1–200. The response includes cached: true when it was served from the 60-second query cache.

The rows are a top-N; the summary is not. transactions is the limit busiest routes by throughput, so the response also carries totals — throughput, percentiles and failures across the whole project in the same window — and truncated, which is true when there are more routes than were returned. Summing the rows to get a project number is wrong in a way that hides exactly what you are looking for: the truncation drops the quietest routes, and a route served five hundred times that fails every single time contributes nothing at all to a failure rate computed that way (issue #185). The dashboard leads with totals and says when the table beneath it is a slice.

Example (compact MCP rendering):

GET /orders/:id: 18432 req, p50 41ms, p95 220ms, p99 610ms, 12 failures
POST /checkout: 2210 req, p50 130ms, p95 890ms, p99 2100ms, 31 failures
https://app.bugwatch.io/o/acme/p/web/performance

Stats time series

GET /v1/orgs/{org}/projects/{project}/stats (event:read):

  • dataset — errors (default), transactions, or sessions.
  • interval — bucket size in minutes: 1, 5, 15, 60 (default), or 1440.
  • since — hours, 1–2160.
  • issue — errors only: restrict to one issue by its internal id (the ULID from the issue record, not the short id).

Returns {"series": [{"t": "2026-09-02T10:00:00Z", "events": 1234}, …], "cached": false}. Timestamps are ISO-8601 UTC.

Sampling, honestly

Analytics Engine samples under load. Every query Bugwatch runs is sampling-aware — counts are SUM(_sample_interval), percentiles are quantileWeighted(…, _sample_interval) — so estimates are unbiased, but they are estimates. Exact numbers exist only where they matter for money and issue state: the billing meter and issue counters come from Durable Objects and the control database, never from the analytics index. Charts therefore can disagree slightly with the usage page; the usage page is right.

Performance queries are cached for 60 seconds per distinct query and are billed per query upstream, which is why the API does not expose arbitrary SQL — filters are typed parameters only, and the query builder always pins the project.

Release health

Session-based crash-free rates live with releases: Releases.

Web vitals

GET /v1/orgs/{org}/projects/{project}/vitals (tool: get_web_vitals) returns the p75 of LCP, FCP, CLS, INP and TTFB for each transaction over the window, sampling-aware like every other Analytics Engine read.

Each vital comes back as { "p75", "samples" }, or as null when nothing measured it. That distinction is the point of the endpoint, not a detail:

  • A backend endpoint has no Core Web Vitals at all — nothing in a server transaction produces an LCP. Reporting 0 for it would put a perfect score on the transaction we know least about.
  • samples is how many events actually carried that vital, which is usually fewer than the transaction's throughput. INP and LCP are not reported by every browser, so a low samples against a high throughput means the number describes a subset of your traffic.

In the dashboard, Performance shows one tile per vital with the worst route for that metric named underneath, rated against Google's published thresholds (green / amber / red). Deliberately not a project-wide average: a throughput-weighted mean of per-route p75s is not a p75, and this is a number teams optimise against. A vital nothing measured shows an em-dash and no rating dot at all — a green dot would say we checked and it was fine.

If no browser transaction reported vitals in the range, the section says so in those words rather than showing an empty table, because "no measurements" and "all good" are not the same finding.

CLS is returned on its natural 0–1 scale. It is the one vital where 0 is the best possible score rather than a missing value, so it is measured across every transaction that reported paint timing rather than only those with a non-zero shift — otherwise every perfectly stable page would drop out and the score would read worse than reality.

GET /v1/orgs/acme/projects/web/vitals?since=24&limit=50
{ "vitals": [ { "transaction": "/checkout",
                "lcp": { "p75": 2410, "samples": 8800 },
                "cls": { "p75": 0.04, "samples": 9100 },
                "inp": { "p75": 180, "samples": 6200 },
                "fcp": { "p75": 1200, "samples": 9100 },
                "ttfb": { "p75": 320, "samples": 9100 } } ], "cached": true }

One transaction in detail

GET /v1/orgs/{org}/projects/{project}/transaction?name=… (tool: get_transaction) answers the follow-up question the table raises: this route is slow — where is the time going?

The name is a query parameter, not a path segment, because a transaction name is whatever your SDK sent and routinely contains slashes (GET /api/users/:id). A query parameter carries it exactly as sent, with no percent-encoding for a proxy to decode differently.

It returns four things:

  • summary — throughput, p50/p75/p95/p99 and failure rate over the window.
  • series — the same latency as a time series, so you can see whether a bad p95 is a step change or a slow drift.
  • spans — where the time actually goes, ranked by total exclusive time rather than by the slowest single call. A query run 400 times at 8 ms owns more of the route than one 900 ms call, and only the total says so.
  • linkedIssues — the error issues whose events share a trace with this route's recent transactions, most events first, each with its shortId, title, status and sampling-weighted event count. linkedTraces is how many traces were looked through (the most recent ones, up to 200), so an empty list is read against a denominator: "no errors in 180 traces" is a finding, and "no trace ids at all" is a different one. At high volume these are traces from a sample, as everything in Analytics Engine is.

found is false when nothing matched in the window. Analytics Engine returns a row of zeros for an aggregate that matched nothing, so without that flag a route with no traffic would read as a route with a perfect 0 ms p50 — say "no traffic", not "fast".

Span rows are absent when the SDK sampled the transaction out; spans are sent for sampled transactions only, so an empty spans list with a healthy summary means the traces were not kept, not that the route does no work.

In the dashboard

Performance → any transaction name opens this view: the summary tiles, the p95 series behind them, and the span table with each span's share of the window's exclusive time — so a row says how much of the route it owns, not merely how long it took. A share is shown as — rather than 0% when there is no exclusive time to attribute, because 0% is a precise-looking number for an absence.

The spans shown are the top ones by total exclusive time, so the shares are shares of that set and the page says so. Below them, Errors in these traces lists the linked issues, each opening the issue, with the number of traces looked through in the header. A route with no trace ids says so instead of reading as one that never errors.