# Filters & scrubbing

> **For agents:** `get_project(org="acme", project="web")` shows `storeIp` and `filters`; `update_project(org="acme", project="web", filters={dropCrawlers:true, messageDenyList:["ResizeObserver loop"]})` changes them (`project:write`). Filtered events appear in `get_usage` as `filtered` outcomes with a reason, never as issues.

## Inbound filters

Filters drop events cheaply and record a `filtered` usage outcome with a reason. They never count against your quota and never create issues. Configure them per project via `PATCH /v1/orgs/{org}/projects/{project}` `{"filters": {…}}` (merged with the existing settings):

| Setting | Default | Stage | Drops when |
|---|---|---|---|
| `dropCrawlers` | on | request | User-Agent matches known bots, headless browsers, uptime checkers, SEO crawlers (reason `crawler`) |
| `dropLegacyBrowsers` | off | request | IE ≤ 11 or the old Android stock browser (`legacy-browser`) |
| `allowedDomains` | `[]` | request | The `Origin`/`Referer` host is not one of the listed domains (`*.example.com` allowed). Requests without either header — server SDKs — always pass (`disallowed-domain`) |
| `dropLocalhost` | off | event | `request.url` host or `server_name` is `localhost`, `127.x`, or `::1` (`localhost`) |
| `dropBrowserExtensions` | on | event | The *throwing* frame's URL is a `chrome-extension://`, `moz-extension://`, `safari-extension://`… script (`browser-extension`) |
| `messageDenyList` | `[]` | event | Any of up to 50 case-insensitive patterns ([RE2 syntax](#deny-list-patterns-use-re2-syntax)) matches the message, log entry, or an exception value (`message-deny-list`). Matched against the first 4096 characters |

Request-stage filters run in the ingest Worker before the payload is parsed; event-stage filters run in the processor. Both read their settings from the edge config cache, so a change takes up to about a minute to apply everywhere — and a change that *enables* a filter applies sooner than that, because an absent config is re-checked every few seconds rather than every minute.

If that cache cannot be read at all, filters **fail open**: the event is stored rather than dropped. By the time the processor sees an event its body is already durable and the event is already counted, so the alternative would be discarding a real error because a settings lookup came back empty.

## PII scrubbing

Scrubbing runs in the processor **before an event is written to the record you read**, on every error event and transaction, and cannot be turned off. The request as it arrived is held first in a 7-day replay buffer (`raw/`), so that a failure in processing never loses an event. It is read only to process or reprocess an event; the dashboard and API never return its contents. Scrub in your SDK as well (`beforeSend`) for anything that must never be stored at all.

Scrubbing is key-based plus pattern-based:

**Sensitive keys** (case-insensitive, anywhere in `request`, `extra`, `tags`, `contexts`, `breadcrumbs`, `user`): `password`, `passwd`, `pwd`, `secret`, `token`, `api_key`, `apikey`, `access_token`, `auth`, `authorization`, `credentials`, `cookie`, `set-cookie`, `card*` (prefix), `ssn`, `stripetoken`, `mysql_pwd`. Values are replaced with `[Filtered]`.

Matching is **per key part**, so compound names resolve the way you would expect: `x-auth-token`, `refresh_token`, `client_secret`, `accessToken` and `x-api-key` are all filtered, as are plurals (`cookies`, `tokens`) and camelCase. It is deliberately *not* substring matching on the whole key — that would take `author` with it, and this product uses `author` for the suspect commit in fix context.

Two consequences worth knowing:

- `request.cookies`, the dedicated field in the Sentry request interface, is filtered as a whole — not only the `Cookie` header beside it.
- The `card*` wildcard over-matches (`cardinality` is filtered along with `cardholder`). The trade is one-sided and deliberate: masking a metric costs a number nobody reads twice, and leaking a card number costs a customer.

**Deeply nested data** stops being walked past 32 levels and is replaced with `[Filtered: nesting too deep]` rather than passed through. Nothing that deep is readable in the UI, and a scrubber that quietly gives up is worse than none at all.

**Query strings** are scrubbed by parameter name in both `request.query_string` and `request.url`, since the sensitive name is data there, not an object key.

**Patterns** (applied to every string, including messages, exception values, and frames): 13–19 digit sequences that pass the Luhn check (card numbers; epoch-millisecond timestamps do not) and `-----BEGIN … PRIVATE KEY-----` blocks become `[Filtered]`. Messages and stack traces are otherwise left intact — masking whole traces would destroy the product.

**Client IP**: `user.ip_address` is **deleted unless the project has `storeIp: true`**. The flag is off by default and is the only opt-in in the scrubber. Only that field is treated as an address: an IP your SDK sends elsewhere, such as an `X-Forwarded-For` request header or a tag, is kept.

SDK-side scrubbing (`beforeSend`, `send_default_pii=False`) still applies and runs first; Bugwatch's pass is a backstop.

## No PII in analytics

The analytics index (Analytics Engine) receives only ids, low-cardinality dimensions (level, environment, release, platform, country), and `user_hash` — the first 16 hex characters of `sha256("<project_id>:<user id, email, or username>")`. Raw identities, emails, messages, and bodies never enter it, which matters because that store is append-only with no per-row deletion. Event bodies live in R2 under the project's prefix and expire with the project's retention tier ([Security & privacy](https://docs.bugwatch.io/account/security-and-privacy.md)).

## Deny-list patterns use RE2 syntax

Deny-list patterns are regular expressions run on a **linear-time engine** (RE2 semantics): matching takes time proportional to the length of the message, whatever the pattern. No pattern can stall event processing, so none is refused for its shape. `^(\w+\s?)+$`, which takes seconds on a JavaScript engine against a 26-character message that fails to match, takes about a millisecond here.

The syntax is the familiar one — literals, `.`, classes like `[a-z]` and `\d\w\s`, anchors `^ $ \b`, groups, alternation `|`, and repeats `* + ? {n,m}` — with two features left out because they need backtracking: **lookaround** (`(?=…)`, `(?!…)`, `(?<=…)`, `(?<!…)`) and **backreferences** (`\1`). A pattern that uses them is refused when you save it, with an error saying why. Matching is case-insensitive; `(?-i)` inside a pattern turns that off for the rest of that pattern only.

Patterns are matched against the first 4096 characters of each candidate string, which is far more than any error message needs.
