Documentation

The rules an incident is graded by

Every number the incident engine judges by, read live from /v1/thresholds. A board that grades incidents by numbers nobody can see is asking to be trusted rather than checked. Every change to one of these is dated in the changelog, with its commit and what it moved.

build a13f65a · these are the rules core is running right now, not a copy of them

Engine rules

RuleValueWhy it is that number
Outage: failing checks, in either window
outage_error_rate
20%Not 100%. A venue failing one request in four is unusable for trading, and waiting for total failure means reporting the incident after the desks have already noticed. Read in both the 60-second and the 5-minute window since 12 Sep 2026: before that only the short window could conclude an outage, so a venue failing by timing out — which delivers fewer samples per minute, often too few for the short rule — was graded degraded at 100% failure. See /changelog.
Degraded: failing checks over 5 minutes
degraded_error_rate
5%The quiet window. A 4% error rate is real but is not worth waking anyone at 3am over a 60-second sample of twelve requests.
Outage: how long the failure must hold before the grade is published
sustained_outage_ms
120000msA rate says whether a component is failing, not how long it has been failing, and the top grade is a claim about both. In the month to 12 Sep 2026 up24 published 190 outages, 26 of them shorter than a minute and the shortest 15.2 seconds — each one a page for a fault that was over before the alert could be read. Until the failure has held this long the same evidence is published one grade down: the incident, its permalink and its alert at the degraded floor all exist from the first bad window, and only the word waits. Two minutes is where the data is empty — the longest sub-minute outage ran 50.0 seconds and the next one up ran 230.0 — so no real event sits near the line. On the way up only: an incident that reaches outage keeps the grade afterwards.
Degraded: p95 against the venue’s own 7-day baseline
latency_multiple
4×Relative, never absolute. 200ms is healthy for a venue on another continent and a serious regression for one two racks away, so one absolute number would be simultaneously too tight and too loose. The baseline is the median of the last 7 days’ p95s — a median so that one bad day does not raise the bar against detecting the next one.
Degraded: absolute floor under the latency rule
latency_floor_ms
500msStops the ratio firing on an endpoint that is fast in absolute terms: 4× a 30ms baseline is 120ms, which nobody would call an incident. Raised from 250ms on 28 Aug 2026, before the first 90-day window filled — see /changelog.
Outage: share of a 5-minute window a connected feed was silent
stall_outage_rate
50%A stalled feed is degraded, not down, until the silence is most of the window. One pause straddling a window boundary fails two of them, which is 33% at six windows a minute and an "outage" for a book that went quiet for twelve seconds. Socket-level failures keep the request rules above: a venue refusing the connection is down whatever the feed would say.
Degraded: sequence gaps tolerated over 5 minutes
seq_gap_limit
3Not zero. A single dropped book update during a reconnect is normal; a feed dropping updates repeatedly is a data-quality fault, and that is what the count catches.
Unknown: measurements needed in the short window (60 seconds at a 5 s cadence)
min_short_samples
4A REST endpoint produces 12 samples a minute; this floor is a third of that. Below it and below the long floor as well, the component is `unknown` before any rate is read — absence of evidence is never operational. Clearing it is not enough to be called healthy: that takes the degraded floor below.
Unknown: measurements needed in the long window (five minutes at a 5 s cadence)
min_long_samples
10The same rule over the long window, tolerating a probe that missed a few cycles. It is also the smallest long window that may call an outage without resolving 5%: ten checks can resolve 20%, and a component failing by timeout delivers fewer checks exactly when it is worst — fifteen in the five-minute window of a 5-second endpoint, all failing, used to read operational.
Days of history the latency baseline is drawn from
baseline_days
7 daysThe median of those days’ p95s. Until a venue has one, the latency rule is skipped entirely rather than guessed at.
How long a better state must hold before it is published
recovery_ms
120000msRecovery is the direction where being wrong is expensive: an outage that flickers green for one window and back to red produces two incidents for one event. Getting worse publishes immediately — the windows above already smoothed it. This is the hold on top of recovery, not the whole of it: a component reads better only once its long window’s failure rate has fallen back under the line, so a component polled slower than every five seconds carries a fault for up to its long window after it recovered — 30 minutes for an RPC block read polled every 60 seconds, 6 hours for the heavier methods polled every 720 seconds.
How long evidence must be missing before a component is `unknown`
unknown_after_ms
120000msLonger than a probe deploy takes, so a rolling restart does not blank the board, and short enough that a probe that died stops being reported as healthy within a couple of minutes.
Quiet hold on a component that keeps re-opening
flap_quiet_ms
900000msA component that has already opened 2 incidents inside an hour has to stay quiet this long before its incident closes, rather than the normal recovery hold. Gemini opened 179 incidents in two days for one bad afternoon; its median spacing between them was 370 seconds, and this window has to be comfortably longer than the gap it bridges. It never delays the component’s published state and never suppresses a first incident.
How often every component is re-evaluated
eval_interval_ms
10000msFast enough for the 60-second detection target.
Smallest denominator the outage rate may be read off
min_outage_samples
5A rate cannot be resolved by fewer samples than its own reciprocal: four samples cannot express 20%, so one failed request in four would be past the line by arithmetic rather than by evidence. Below this the short window’s outage rule is skipped. It binds only where a cadence is slow enough to make it bind — every exchange endpoint delivers twelve samples a minute.
Smallest denominator the degraded rate, or an operational grade, may be read off
min_degraded_samples
20The same rule at 5%: twenty samples is the first window in which 5% is a value the sample can take, and so the floor under `operational` as well as under `degraded` — "not degraded" is a claim about the same rate. Below it, a long window no rule fired on is `unknown`, and the p95 is not read either. Together with the one above it is why a component polled every 720 seconds is judged over six hours rather than five minutes — nothing can be detected faster than it is sampled, and the alternative is a window that is permanently `unknown`. The stretched window holds one and a half times this floor, so a component can lose a third of its polls and still be graded: sized to exactly twenty, one late or rate-limited poll would have read `unknown` with every check green.
Degraded: blocks behind the freshest head up24 saw
lag_degraded_blocks
2Two blocks is the smallest gap that is not one block of ordinary skew between two nodes. The reference is the highest block any provider on the same chain reported from the same region in the same polling round, never across regions — a provider can be current in Lauterbourg and behind in Singapore, and that difference is the finding. Fewer than two providers in a round is `unknown`, never zero. Judged in milliseconds, not blocks: see the floor below.
Degraded: absolute floor under the block-lag rule
lag_degraded_floor_ms
10000msTwo blocks is 24 seconds on Ethereum, four on Base and half a second on Arbitrum One, and half a second is not a number this instrument can resolve: the comparison is between two readings taken up to five seconds apart, which can manufacture five seconds of apparent lag unaided. Ten seconds is twice that width. Measured against the soak week to 12 September 2026 and unchanged: the largest reading this rule saw in seven days, over six targets and three regions, was 1.152 blocks.

Stall thresholds, per feed

How long each feed may say nothing before its ten-second window is counted as failed. Registry data rather than an engine constant, because three seconds of silence means something different on a 100 ms order book than on a trade tape that is quiet between trades — each is read off the measured distribution of that feed's own quiet periods over a week of production, at the grain the threshold is enforced at.

Almost every feed has one number wherever it is watched from. Where a feed is measurably quieter from one vantage point than another, it carries that region's own threshold and both are shown — the same reason latency is never averaged across regions.

VenueFeedChannelStall after
Binancews:tradebtcusdt@trade5 min
Binancews:bookbtcusdt@depth@100ms7000ms
Bybitws:tradepublicTrade.BTCUSDT5 min
Bybitws:bookorderbook.50.BTCUSDT5000ms
Coinbasews:tradematches5 min
Coinbasews:tickerticker1 min
Krakenws:tradetrade5 min
Krakenws:bookbook
12s eu-central
40s ap-southeast
12s us-east
OKXws:tradetrades5 min
OKXws:bookbooks5000ms
Bitfinexws:tradetrades5 min
Bitfinexws:bookbook20s
Deribitws:tradetrades.BTC-PERPETUAL.100ms5 min
Deribitws:bookbook.BTC-PERPETUAL.100ms5000ms
Bitstampws:tradelive_trades_btcusd5 min
Bitstampws:bookdiff_order_book_btcusd10s
Geminiws:tradetrade9 min
Geminiws:bookl2_updates20s
Crypto.comws:tradetrade.BTC_USDT5 min
Crypto.comws:bookbook.BTC_USDT.1015s
Hyperliquidws:tradetrades5 min
Hyperliquidws:bookl2Book22s
Upbitws:tradetrade5 min
Upbitws:bookorderbook12s
Binance Futuresws:tradebtcusdt@trade5 min
Binance Futuresws:bookbtcusdt@depth@100ms10s
Bybit Futuresws:tradepublicTrade.BTCUSDT5 min
Bybit Futuresws:bookorderbook.50.BTCUSDT6000ms
OKX Futuresws:tradetrades5 min
OKX Futuresws:bookbooks7000ms