Documentation
The rules an incident is graded by
Every number the incident engine judges by, read live from /v1/thresholds. A board that grades incidents by numbers nobody can see is asking to be trusted rather than checked. Every change to one of these is dated in the changelog, with its commit and what it moved.
build a13f65a · these are the rules core is running right now, not a copy of them
Engine rules
| Rule | Value | Why it is that number |
|---|---|---|
Outage: failing checks, in either windowoutage_error_rate | 20% | Not 100%. A venue failing one request in four is unusable for trading, and waiting for total failure means reporting the incident after the desks have already noticed. Read in both the 60-second and the 5-minute window since 12 Sep 2026: before that only the short window could conclude an outage, so a venue failing by timing out — which delivers fewer samples per minute, often too few for the short rule — was graded degraded at 100% failure. See /changelog. |
Degraded: failing checks over 5 minutesdegraded_error_rate | 5% | The quiet window. A 4% error rate is real but is not worth waking anyone at 3am over a 60-second sample of twelve requests. |
Outage: how long the failure must hold before the grade is publishedsustained_outage_ms | 120000ms | A rate says whether a component is failing, not how long it has been failing, and the top grade is a claim about both. In the month to 12 Sep 2026 up24 published 190 outages, 26 of them shorter than a minute and the shortest 15.2 seconds — each one a page for a fault that was over before the alert could be read. Until the failure has held this long the same evidence is published one grade down: the incident, its permalink and its alert at the degraded floor all exist from the first bad window, and only the word waits. Two minutes is where the data is empty — the longest sub-minute outage ran 50.0 seconds and the next one up ran 230.0 — so no real event sits near the line. On the way up only: an incident that reaches outage keeps the grade afterwards. |
Degraded: p95 against the venue’s own 7-day baselinelatency_multiple | 4× | Relative, never absolute. 200ms is healthy for a venue on another continent and a serious regression for one two racks away, so one absolute number would be simultaneously too tight and too loose. The baseline is the median of the last 7 days’ p95s — a median so that one bad day does not raise the bar against detecting the next one. |
Degraded: absolute floor under the latency rulelatency_floor_ms | 500ms | Stops the ratio firing on an endpoint that is fast in absolute terms: 4× a 30ms baseline is 120ms, which nobody would call an incident. Raised from 250ms on 28 Aug 2026, before the first 90-day window filled — see /changelog. |
Outage: share of a 5-minute window a connected feed was silentstall_outage_rate | 50% | A stalled feed is degraded, not down, until the silence is most of the window. One pause straddling a window boundary fails two of them, which is 33% at six windows a minute and an "outage" for a book that went quiet for twelve seconds. Socket-level failures keep the request rules above: a venue refusing the connection is down whatever the feed would say. |
Degraded: sequence gaps tolerated over 5 minutesseq_gap_limit | 3 | Not zero. A single dropped book update during a reconnect is normal; a feed dropping updates repeatedly is a data-quality fault, and that is what the count catches. |
Unknown: measurements needed in the short window (60 seconds at a 5 s cadence)min_short_samples | 4 | A REST endpoint produces 12 samples a minute; this floor is a third of that. Below it and below the long floor as well, the component is `unknown` before any rate is read — absence of evidence is never operational. Clearing it is not enough to be called healthy: that takes the degraded floor below. |
Unknown: measurements needed in the long window (five minutes at a 5 s cadence)min_long_samples | 10 | The same rule over the long window, tolerating a probe that missed a few cycles. It is also the smallest long window that may call an outage without resolving 5%: ten checks can resolve 20%, and a component failing by timeout delivers fewer checks exactly when it is worst — fifteen in the five-minute window of a 5-second endpoint, all failing, used to read operational. |
Days of history the latency baseline is drawn frombaseline_days | 7 days | The median of those days’ p95s. Until a venue has one, the latency rule is skipped entirely rather than guessed at. |
How long a better state must hold before it is publishedrecovery_ms | 120000ms | Recovery is the direction where being wrong is expensive: an outage that flickers green for one window and back to red produces two incidents for one event. Getting worse publishes immediately — the windows above already smoothed it. This is the hold on top of recovery, not the whole of it: a component reads better only once its long window’s failure rate has fallen back under the line, so a component polled slower than every five seconds carries a fault for up to its long window after it recovered — 30 minutes for an RPC block read polled every 60 seconds, 6 hours for the heavier methods polled every 720 seconds. |
How long evidence must be missing before a component is `unknown`unknown_after_ms | 120000ms | Longer than a probe deploy takes, so a rolling restart does not blank the board, and short enough that a probe that died stops being reported as healthy within a couple of minutes. |
Quiet hold on a component that keeps re-openingflap_quiet_ms | 900000ms | A component that has already opened 2 incidents inside an hour has to stay quiet this long before its incident closes, rather than the normal recovery hold. Gemini opened 179 incidents in two days for one bad afternoon; its median spacing between them was 370 seconds, and this window has to be comfortably longer than the gap it bridges. It never delays the component’s published state and never suppresses a first incident. |
How often every component is re-evaluatedeval_interval_ms | 10000ms | Fast enough for the 60-second detection target. |
Smallest denominator the outage rate may be read offmin_outage_samples | 5 | A rate cannot be resolved by fewer samples than its own reciprocal: four samples cannot express 20%, so one failed request in four would be past the line by arithmetic rather than by evidence. Below this the short window’s outage rule is skipped. It binds only where a cadence is slow enough to make it bind — every exchange endpoint delivers twelve samples a minute. |
Smallest denominator the degraded rate, or an operational grade, may be read offmin_degraded_samples | 20 | The same rule at 5%: twenty samples is the first window in which 5% is a value the sample can take, and so the floor under `operational` as well as under `degraded` — "not degraded" is a claim about the same rate. Below it, a long window no rule fired on is `unknown`, and the p95 is not read either. Together with the one above it is why a component polled every 720 seconds is judged over six hours rather than five minutes — nothing can be detected faster than it is sampled, and the alternative is a window that is permanently `unknown`. The stretched window holds one and a half times this floor, so a component can lose a third of its polls and still be graded: sized to exactly twenty, one late or rate-limited poll would have read `unknown` with every check green. |
Degraded: blocks behind the freshest head up24 sawlag_degraded_blocks | 2 | Two blocks is the smallest gap that is not one block of ordinary skew between two nodes. The reference is the highest block any provider on the same chain reported from the same region in the same polling round, never across regions — a provider can be current in Lauterbourg and behind in Singapore, and that difference is the finding. Fewer than two providers in a round is `unknown`, never zero. Judged in milliseconds, not blocks: see the floor below. |
Degraded: absolute floor under the block-lag rulelag_degraded_floor_ms | 10000ms | Two blocks is 24 seconds on Ethereum, four on Base and half a second on Arbitrum One, and half a second is not a number this instrument can resolve: the comparison is between two readings taken up to five seconds apart, which can manufacture five seconds of apparent lag unaided. Ten seconds is twice that width. Measured against the soak week to 12 September 2026 and unchanged: the largest reading this rule saw in seven days, over six targets and three regions, was 1.152 blocks. |
Stall thresholds, per feed
How long each feed may say nothing before its ten-second window is counted as failed. Registry data rather than an engine constant, because three seconds of silence means something different on a 100 ms order book than on a trade tape that is quiet between trades — each is read off the measured distribution of that feed's own quiet periods over a week of production, at the grain the threshold is enforced at.
Almost every feed has one number wherever it is watched from. Where a feed is measurably quieter from one vantage point than another, it carries that region's own threshold and both are shown — the same reason latency is never averaged across regions.
| Venue | Feed | Channel | Stall after |
|---|---|---|---|
| Binance | ws:trade | btcusdt@trade | 5 min |
| Binance | ws:book | btcusdt@depth@100ms | 7000ms |
| Bybit | ws:trade | publicTrade.BTCUSDT | 5 min |
| Bybit | ws:book | orderbook.50.BTCUSDT | 5000ms |
| Coinbase | ws:trade | matches | 5 min |
| Coinbase | ws:ticker | ticker | 1 min |
| Kraken | ws:trade | trade | 5 min |
| Kraken | ws:book | book | 12s eu-central 40s ap-southeast 12s us-east |
| OKX | ws:trade | trades | 5 min |
| OKX | ws:book | books | 5000ms |
| Bitfinex | ws:trade | trades | 5 min |
| Bitfinex | ws:book | book | 20s |
| Deribit | ws:trade | trades.BTC-PERPETUAL.100ms | 5 min |
| Deribit | ws:book | book.BTC-PERPETUAL.100ms | 5000ms |
| Bitstamp | ws:trade | live_trades_btcusd | 5 min |
| Bitstamp | ws:book | diff_order_book_btcusd | 10s |
| Gemini | ws:trade | trade | 9 min |
| Gemini | ws:book | l2_updates | 20s |
| Crypto.com | ws:trade | trade.BTC_USDT | 5 min |
| Crypto.com | ws:book | book.BTC_USDT.10 | 15s |
| Hyperliquid | ws:trade | trades | 5 min |
| Hyperliquid | ws:book | l2Book | 22s |
| Upbit | ws:trade | trade | 5 min |
| Upbit | ws:book | orderbook | 12s |
| Binance Futures | ws:trade | btcusdt@trade | 5 min |
| Binance Futures | ws:book | btcusdt@depth@100ms | 10s |
| Bybit Futures | ws:trade | publicTrade.BTCUSDT | 5 min |
| Bybit Futures | ws:book | orderbook.50.BTCUSDT | 6000ms |
| OKX Futures | ws:trade | trades | 5 min |
| OKX Futures | ws:book | books | 7000ms |