Changelog

Every change to the rules up24 grades by, and every measurement it has withdrawn. Dated, with the commit, and with what it moved.

21 rule changes · 6 corrections · the rules in force are published at /docs#thresholds

A status board that quietly re-grades its own history is worth nothing to anyone quoting it: one silent change and every number it ever published needs an asterisk. So the rules are allowed to move — they have to, because production keeps finding shapes nobody designed for — and every move is written down here instead.

Corrections retract rather than delete. A withdrawn incident leaves every list, count, feed and API response; its permalink keeps serving, marked retracted and carrying the reason. If a measurement here is wrong, say so: reviewed within 48 hours, corrected or explained within a week, and the correction lands on this page.

RULE

Stream counts with no measurement behind them are null, not zero

A stream up24 held no record for published reconnects24h, seqGaps24h and maxGapMs24h (and messages24h per stream) as 0 — no reconnect, no gap, a longest silence of 0 ms, which describes a clean stream — beside a null connectedPct24h. The four fields are now nullable in the API and read null with no stream row behind them.

e15d404 · nothing already published was re-graded

CORRECTION

5 Sep 2026 was published as its last nine hours; it is now the whole day

Migration 0017 wound the day rollup back to 5 Sep 14:45:13, and the re-run started there rather than at midnight, so the day was replaced with 14:45 to 24:00: Binance read 51,840, 19,272 and 51,840 checks for 4, 5 and 6 Sep, and every venue’s 5 Sep read about 19,272. The day is rebuilt from its minute rollups, which are kept for good. Its percentiles are withdrawn rather than rebuilt: a day’s p95 cannot be computed from minute p95s, and the samples it needs were deleted at seven days.

563ac1b · Every venue’s 5 Sep bar, from about 19,272 checks to about 51,780. Two changed grade: Deribit from 100% to 97.28% (degraded) and Hyperliquid from 100% to 99.88% (degraded), failures in the morning the partial day had dropped. Deribit’s 90-day figure fell by about 0.04 points. 5 Sep no longer contributes to any typical p50 or p95. No incident was opened, closed or retracted.

RULE

Percentages are floored to two decimals, so 100.00% means no check failed

Every uptime figure was rounded half-up, so 99.995% and above printed as 100.00%. /is/bybit-down said Bybit answered 100.00% of 2,498,280 checks while its own daily series carried a 99.89% day, and a day with 54 failures in 51,840 (99.8958%) published 99.9 and graded operational beside the “under 99.9 is degraded” rule. A floored figure sits on the same side of 99.9 and 95 as the exact ratio, so a day is now graded on what it measured.

64eb3dd · Read from production before and after the deploy: 243 published day bars across the listed venues moved by 0.01 points, and one changed grade (Hyperliquid, 23 Aug, 99.9 operational to 99.89 degraded). Eleven 90-day figures moved by 0.01; the six that read 100.00% (Binance, Binance-futures, Bybit, Kraken, OKX, OKX-futures) now read 99.99%. The fleet figure moved from 99.97% to 99.96%. No incident was opened, closed or retracted.

RULE

A window of N days sums N UTC days, the same days its strip draws

Every ?days=N window summed N complete days plus today, one day more than the N bars it drew, so the oldest day was in the checks and the uptime and in no bar. /v1/reliability?days=1 for Binance published 97,625 checks beside its one bar of 45,799. The incident count in the same summary now covers the same calendar days too, rather than a rolling N × 24 hours.

8356c9a · No 90-day figure: the longest record was 51 measured days, so the extra day held no data. Every shorter window (?days=1 to 89) lost the day before its first bar. No incident was opened, closed or retracted.

RULE

A region needs 1,000 checks in a day for its reading to vote on that day

Bybit’s 2 Sep 2026 bar read 60.13%. Singapore measured 100% off 51,732 checks and Lauterbourg 60.13% off 33,308. Carlstadt, whose refusals are discounted, still held one timeout, and that single check voted 0%: three regions instead of two, and the median landed on Lauterbourg. The engine opened no incident that day. A region now votes only with at least 1,000 gradeable checks, the fewest that can resolve the 0.1% between a day at 99.9% and a degraded one. On a day no region reaches that (the first minutes after UTC midnight, or a slow-cadence target), everyone who reported still votes.

f2e979c · 43 published days, all on Binance, Binance-futures, Bybit and Bybit-futures, the four venues whose Carlstadt readings are refused. Bybit and Bybit-futures on 2 Sep moved from 60.13% and 60.12% to 100%; the other 41 moved by at most 0.06 points. No incident was opened, closed or retracted.

RULE

A window too thin to judge is graded unknown, never operational

Until this change a component whose long window held 10 to 19 checks and whose short window held fewer than 5 skipped every rate rule and was graded operational, whatever those checks said — the shape of a component failing by timeout, which starves its own windows. Fifteen of fifteen checks failing read green, counted in the regional quorum, and could close an open incident. Now a long window of 10 or more checks past the 20% line is an outage, the degraded rate, the latency rule and operational need 20, and anything else is unknown, except where the block-lag or sequence-gap rule, which read their own series, finds a fault. RPC long windows are sized for 30 polls rather than exactly 20, so one missed poll no longer blanks a row: 30 minutes for the block read, 6 hours for the heavier methods. Exchange windows are unchanged.

90369bc · nothing already published was re-graded

CORRECTION

A caveat said Binance’s WebSocket was unaffected from Carlstadt, and it was not

The block note published on every refused region of binance, binance-futures, bybit and bybit-futures ended “and this venue’s WebSocket from the same box is unaffected”. It was true of three of the four and flatly false of the fourth, on the page that names it. Over the 24 hours to 15:30 UTC on 13 September, us-east stream_health read 0 of 17,230 windows ok on binance, against 16,974 on binance-futures, 17,211 on bybit and 17,211 on bybit-futures. Binance’s socket from Carlstadt has never connected: the venue refuses the upgrade handshake from that country exactly as it refuses the REST calls. The evidence for the sentence had been measured on Bybit, where it holds, and written into a note four targets share — a fact that differs per venue cannot live in a caveat they hold in common, and that is the mistake rather than the wording. The clause is gone; the per-venue numbers stay in the registry’s own documentation, attributed to what each was measured on. It was public for about 24 hours, from 14:59 UTC on 12 September, which bounds the damage and does not change the rule: a published claim is withdrawn where it was made, not deleted.

ca2e116 · One published caveat withdrawn, on four venue pages and in four /v1/exchanges payloads. No percentage, state or incident changes: the sentence described the socket and nothing was ever graded on it.

RULE

A refused vantage point is discounted wherever it was recorded, not only in REST

The 12 September rule said a reading from a region the venue refuses is not evidence about that venue, and it was applied in three queries, all of them over raw_samples. Eight queries read raw_samples or stream_health. So Carlstadt’s REST refusals were discounted from the 12th while the socket half went on voting, and two public reads went on publishing the refusals under both names at once. Over the 24 hours to 15:30 UTC on 13 September, us-east wrote 8,610 windows at 0% ok on binance/book and 8,610 more on binance/trade — 17,220 rows a day of pure dissent — while Lauterbourg and Singapore read 100% and 99.99%. /v1/reliability published binance’s streams at connectedPct24h 66.8, which is two regions delivering and one refused, averaged; /v1/exchanges/binance with region=us-east returned the same 17,239 checks as 0% uptime in endpoints[] and as 51,717 refused-and-not-counted in regions.compare[], in one response, beside sockets reading operational on zero messages. Six of the eight queries now carry the anti-join, matched on (venue, region, class) and the date window as the REST ones are. readStreamErrors abstains deliberately and says so at the query: a breakdown of what the failures were is where http_451 belongs. No new registry entry was needed — the socket records the same class names the REST calls do, so the declaration was never missing and only the join was.

ca2e116 · The socket evidence behind three components on each of four venues, out of the incident engine’s windows and out of the stream day rollups from this deploy forward. On the reads: binance’s published stream figures stop averaging a refused region, and the per-region endpoint and stream detail for a refused region stop reporting 0% over checks the same payload calls uncounted. Already-written stream_rollup_1d rows are not corrected by this deploy — every day from 30 August carries them, and the correction is its own task rather than a silent rewrite.

RULE

The three perpetual venues are back on the board, with their own measured thresholds

They were taken off on 31 August with one measured day behind them and stall thresholds copied from their spot siblings, and the retraction said the week would set both. The week is read: 181,198 ten-second windows per feed, worst-hour p99 of 2,056 ms in Singapore, 3,119 ms in Lauterbourg and 2,454 ms in Carlstadt on binance-futures’ book, 1,101 / 1,249 / 1,733 on bybit-futures and 1,535 / 1,383 / 2,163 on okx-futures. Each threshold is now three times the loudest region’s worst hour — 10,000 ms, 6,000 and 7,000, against the 7,000, 5,000 and 5,000 they borrowed — which is the rule every other stall threshold in the fleet was set by. One number each and no per-region override: the spread between the quietest region and the loudest is 1.5–1.6×, where Kraken’s book was 4× and is the case that per-region numbers exist for. The reading excludes every window within two of a reconnect as well as the disconnected ones, which moved it by 0–1.5% and changed no threshold, so these are the feeds going quiet and not up24 resubscribing. Twelve measured days rather than the seven the rule asks for, because the reading was held to 7 September and taken on the 12th.

de9db3d · Three venues onto the board, the leaderboard from twelve rows to fifteen, and three comparison pages live. The 90-day window’s fill date moves with them: it is the least-measured listed row, which is now one of these three, so the sentence on /reliability reads 28 November 2026 instead of 10 November. Zero stall incidents opened on any of the three in the week measured, against three on binance/book and one on bybit/book, so nothing already published changes grade.

RULE

Head-stream stall thresholds come off the week, not off the chain’s block time

The RPC head streams were given max(3 × block time, 10 s) when they arrived — 36 s on Ethereum and the floor’s 10 s on both L2s — a derivation written before anything had been measured, with the floor there because three block times on a 250 ms chain is 750 ms. A week of production says the formula understates every chain by 2.7–3.2×. Over ~60,200 honest windows per region to 12 September the median hourly p99 is 16.4 s on Ethereum, 2.5 s on Base and 1.2 s on Arbitrum One, and the worst hour is 30,401 / 23,464 / 32,014 ms on Ethereum, 11,230 / 6,970 / 5,580 on Base and 8,546 / 2,941 / 10,499 on Arbitrum One. Three times the loudest region gives 96,000 ms, 34,000 and — because Arbitrum’s 3.6% spread between Lauterbourg and Carlstadt is the Kraken shape exactly, on a chain that only produces a block when there are transactions to sequence — 9,000 with per-region overrides of 26,000 in Singapore and 32,000 in Carlstadt. Ethereum at eight block times looks loose beside a 12-second chain and is what the distribution says: a feed whose median hourly p99 is already 16 s cannot be judged in block times, and the three 120-second silences in the week, none of them with a reconnect behind them, are what the threshold is for. The block-lag rules were measured the same way and did not move: the series the rule reads — the smallest reading in a five-minute window holding at least three minutes — maxed at 1.152 blocks across all six targets and three regions, against two blocks on Ethereum and the ten-second floor’s forty and five on Arbitrum One and Base.

de9db3d · Nothing published: every RPC row is measured and listed nowhere, and M15.4b lists them once the category has a week that measures the providers rather than up24 — see the next entry. The old thresholds had been calling 275, 327 and 373 windows a stall in that week.

RULE

up24 was spending Infura’s per-second limit, and the Base row paid for it

A quota has two windows and only one of them was written down. The day was right — 49.5% of the free plan’s 3,000,000 credits, checked at module load since 4 September — and the second was at roughly 750% of a limit nothing here carried: Infura publishes 500 credits a second on the Core plan beside that daily grant. Every endpoint of every target started its grid at the instant its own loop started, so all three chains’ eth_getBlockByNumber, eth_call and eth_getLogs fired inside 900 ms of each other in registry order, 1,245 credits from one region with three regions spending one key. The last target in that order took the refusals: from 5 to 12 September infura-base answered HTTP 429 to 99.09% of its log queries — 2,504 of 2,527 — and to 16% of its block reads, while infura-ethereum and infura-arbitrum answered 429 to six requests between them. None of it was visible as a fault, which is the part worth recording: a 429 from a keyed provider is excluded from the incident windows and from the rollups because up24’s own spent quota is never published as a provider’s fault, so the row read unknown with nothing behind it rather than wrong. The registry now carries the per-second cap as data beside the daily one, every call that shares a credential gets its own slice of its interval — 27 of them 26.7 s apart, worst second 255 credits, one eth_getLogs alone — and the block-height reads on one chain keep a shared slice so the two providers stay inside the five seconds the block-lag comparison needs. The ceiling is checked against the phases the scheduler actually fires on.

de9db3d · No published number: the RPC rows are not on the board. What it moves is the soak — infura-base’s log query has 23 usable samples for the week it was supposed to be measured over, so M15.4b’s listing waits for a clean week to 20 September rather than publishing a component whose evidence is mostly up24.

RULE

An outage has to last two minutes before it is called one

Between 13 Aug and 12 Sep up24 published 190 outages on listed venues, 26 of them shorter than a minute. The shortest was two Bybit REST components for 15.2 seconds — over and recovered before anybody could have opened the link in the alert about it — and a new account watches every venue at the outage floor, so all 26 were pages. An error rate says whether a component is failing and nothing about how long, which is half of what the word claims. An outage verdict is now published as degraded until the same evidence has held for two minutes. The number came from the distribution rather than from taste: the longest of those 26 ran 50.0 seconds and the next outage up ran 230.0, so nothing ever published falls between one and two minutes. Nothing but the word waits — the component turns unhealthy inside the same 60-second window, the incident and its permalink open on the spot, the alert reaches everyone watching at the degraded floor, and the row carries a clause saying why the worse word is being withheld. The hold is on the way up only.

a8df6e5 · 26 published outages re-graded to degraded by migration 0018 — every closed incident shorter than two minutes, which cannot have held the rules for two minutes whatever the per-tick evidence was. The original grade is in each row’s meta and no permalink changed.

RULE

Both windows can conclude an outage, not just the short one

The 60-second window was the only one allowed to grade an outage and the five-minute window could only ever reach degraded, so the grade depended on which window happened to hold enough samples rather than on how bad the reading was. Timeouts decide that backwards: a check that times out occupies the probe for its whole endpoint timeout, so a venue failing by timeout delivers fewer samples a minute and falls under the short window’s evidence floor exactly when it is at its worst. On 12 Sep production published a Crypto.com REST component at 100% of checks failing for 110 seconds, graded degraded, on the same afternoon as a 15-second Bybit row graded an outage. A rate past the outage line is an outage whichever window resolved it; the long window now reads the same two thresholds the short one does, including the stall carve-out that keeps a quiet feed off the outage grade until most of the window is silent. Nothing already published is re-graded: raw_samples is kept seven days, so the per-tick evidence for older rows is gone, and reconstructing it from a summary string would be a guess printed as a measurement.

a8df6e5 · nothing already published was re-graded

CORRECTION

The three perpetual venues are off the board until their week is measured

binance-futures, bybit-futures and okx-futures went on the board the day they were added, with one measured day behind them and stall thresholds copied from their spot parents and marked provisional in the registry. Ranking a venue with one day of history beside eleven with nineteen is a comparison the leaderboard cannot support, and a threshold inherited from a different feed is a number nobody has measured. All three are now unlisted: they are polled, graded and rolled up exactly as before, and they appear on no list, in no count and in no feed until a seven-day production soak sets each stallMs from gaps.js. Nothing measured is withdrawn — the three opened no incidents and reported 100% of checks ok on the day they were listed, and their own pages, badges and API rows keep answering. What is withdrawn is their place in the fleet numbers.

cec1570 · The leaderboard's fill date, from 28 November 2026 back to 10 November: it reads the least-measured listed venue, and that is a nineteen-day row again rather than a one-day one. No incident and no published measurement changed.

RULE

hyperliquid/book re-read against a clean week, and no longer provisional

The 28 Aug pass set this feed to 20,000 ms and said so with a caveat: it had 21,539 windows behind it rather than the 111,000 every other feed got, because the probe held a stuck socket to Hyperliquid for five days, and the number wanted re-reading against a week that was not half a fault. This is that re-read. 30,262 windows across the two regions old enough to count, and the reference is ap-southeast — the one box with no known contamination, a dedicated machine on a quiet network with seven days behind it — which reads a 7,259 ms worst-hour p99. Three times that, the same rule the pass used, is 21,777, so the threshold is 22,000. eu-central corroborates at 7,390 ms; it is not the basis because it is the co-located box and its p99 carries that noise. us-east was excluded with six hours of history against the seven days the rule needs. The two mature regions agreeing to within 2% is what keeps this a single number: kraken/book needed a per-region threshold because Singapore and Lauterbourg disagreed by 3.5× on the same feed, and splitting a feed the regions agree on would be false precision. What a reader should take from this is less the number than its insensitivity: across 48 mature region-hours the worst gap is under 8,000 ms in 41 of them, then 8,405, 9,576 and 16,054, then four hours at 28,455, 28,638, 28,871 and 32,450 — that last group seen in both regions in the same hour, which makes it Hyperliquid rather than a probe box. Nothing lands between 20,000 and 22,000, so this moves no window on the week behind it. It changed because the published rule is three times the worst-hour p99 and the old figure came from a fault, not because the old figure was firing: hyperliquid ws:book has opened one incident since the pass, on 21 August, before it.

ab75274 · nothing already published was re-graded

CORRECTION

Three and a half hours of Europe latency were our own box, not the venues

Between 17:00 and 20:35 UTC every eu-central latency figure on this site was too high, on all twelve venues at once, and the cause was the probe box rather than anything an exchange did. Binance read a 235 ms p50 before and 377 after; bitstamp went from 20 ms to 164; p95 on binance/rest reached 1,111 ms against the 262 it returned once the move was undone. Twelve unrelated venues on four continents do not degrade in lockstep, and the smallest numbers distorted worst because the delay was additive — the same tell that found the co-location gaps earlier the same day. The new Lauterbourg box, brought up at 16:57, could not hold a steady round trip: measured from the box itself, a TCP handshake to a fixed address 4 ms away returned in 5 ms at best and 190 ms at worst, and the jitter was already present at the first hop inside the datacenter (best 0.4 ms, worst 167 ms). The hop from that box to the core box in the same building — half a millisecond apart — ranged to 122 ms, while the 12,000 km hop from Singapore to the same machine held a 1.5 ms deviation. Two other boxes on the same provider, same image, same product tier and the same 36-endpoint workload were steady throughout, at 0.4% link utilisation, which rules out our own traffic as the cause. Measurement moved back onto the core box at 20:35 and the numbers returned to their prior baseline in the next five-minute bucket. Availability is not affected and no incident was created or withdrawn: every check in the window succeeded, and the failing latency rule is four times a venue’s own baseline, which an inflation of roughly 1.6× never reached. The affected rows are kept rather than deleted, and this entry is what withdraws the claim they make. The provisioning script now refuses to hand a probe host to a box whose handshake spread exceeds 15 ms, measured before it is allowed to publish anything.

0d5b780 · 3 h 35 m of eu-central latency rollups across all twelve venues, in rollup_1m and the 2026-08-29 row of rollup_1d. No incident, no availability figure, and no other region.

RULE

A third vantage point, and what it does to a day’s grade

A third probe box joined the fleet, in Carlstadt, New Jersey — about ten milliseconds from Ashburn, where the US-domiciled venues and most of the software that trades them actually run. Adding it changes how a day is graded, and that is worth stating plainly rather than leaving for a reader to find. A day’s published uptime is the floor(regions / 2) + 1-th worst region: with two regions that is the second worst of two, which is the better of the pair, and with three it is the second worst of three, which is the median. So from today a venue that fails from two regions and answers from one publishes the failure, where two regions would have published the success. Every day already written keeps the two regions it was measured by and is untouched, which does mean the 90-day window filling in November spans two rules. The alternative was pooling the regions, which would let one probe box’s bad afternoon halve a venue’s published uptime forever, and that is the failure that makes extra vantage points a liability rather than an asset.

a6ad0fa · nothing already published was re-graded

RULE

The Europe probe moved off the core box onto a machine of its own

eu-central was measured from the same box that runs Postgres, the API and the web app, and a vantage point that shares a machine with a database is not a vantage point. Seven days of stream_health show it: the worst gap of the week was 44,404 ms on binance/book, 43,247 on binance/trade, 39,941 on bitstamp/book, 38,662 on bitfinex/book, 37,885 on cryptocom/book and 33,700 on deribit/book. Six unrelated exchanges do not go quiet for forty seconds together; one machine does. The probe now runs on its own VPS in Lauterbourg — the same city, so the region id and its baselines carry over unchanged — and the old Dokku app is stopped. Nothing already published is re-graded, but every measurement from today is taken from a machine with nothing else on it, so the gaps that remain are the venues’.

f84ed7e · nothing already published was re-graded

RULE

Stall thresholds re-read off a week of production, and one made per-region

The stall-tuning pass D6.8 asked for, run against 111,000 windows per feed. Nineteen of the twenty-four feeds had comfortable margin and did not move. Of the five that did not: binance/book went 3,000 → 7,000 ms, cryptocom/book 10,000 → 15,000, upbit/book 10,000 → 12,000, and hyperliquid/book 10,000 → 20,000, each set at three times its own measured worst-hour p99 — the figure the previous numbers were sitting inside, which is why they were firing tens of times a week on feeds that were behaving normally. kraken/book is the exception and the interesting one: its worst hour reads a 12,815 ms p99 seen from Singapore and 3,654 ms from Lauterbourg, so one number for both had to be either loose enough to blind Europe or tight enough to manufacture stalls in Asia. It is now 12,000 ms by default and 40,000 in ap-southeast — the first per-region threshold in the registry, on the same principle that stops latency being averaged across regions. hyperliquid/book is provisional: it has 21,539 windows behind it rather than 111,000, because the probe held a stuck socket to it for five days, and it wants re-reading against a clean week.

d6a71fa · about 450 stall windows a week that were normal feed behaviour

RULE

The window after a reconnect is up24’s, and no longer counts against the venue

A disconnect already excluded its own window: the socket was not up throughout, so the row reads `conn` and leaves the statistics. What it did not exclude was the next window, which is fully connected and carries the wait for the first message after resubscribing — a round trip to the venue plus its subscription handling, indistinguishable from the feed going quiet. Binance closes a WebSocket after 24 hours, so the Europe probe reconnected at the same hour every day, and over a week that hour read a 10,549 ms p99 on binance/book against 1,562 ms from Singapore, with seven reconnects behind it. The 3,000 ms threshold called it a stall about forty times a week. Raising the threshold to cover it would have been tuning the stall detector on up24’s own reconnect, which is the mistake the disconnect exclusion exists to prevent. The first fully-connected window after a reconnect is now marked as a local fault: still written, out of every statistic, and visible through `coveragePct24h` and `probeFault`.

d6a71fa · roughly 40 Binance stall windows a week, and the equivalent on any feed that reconnects

RULE

Raise the latency floor from 250 ms to 500 ms

A REST component is degraded on latency when its p95 exceeds both four times its own seven-day baseline and an absolute floor. At 250 ms the floor was doing the deciding rather than the multiple: a venue with a 60 ms baseline tripped at 250 ms, which is four times worse than usual and still an API no desk would call degraded. The multiple is the rule that carries meaning; the floor exists only to stop the arithmetic firing on a 30 ms endpoint. 500 ms is where a REST call starts to cost a fill. Deliberately set before the 90-day window fills rather than after — a re-grade in November would leave every quarterly figure quotable only with an asterisk, which was Priya’s objection exactly. Nothing already published is re-graded: this changes future evaluation only.

6e1bc88 · nothing already published was re-graded

CORRECTION

Delete the incidents that were up24 measuring itself

Migration 0011 removed 21 rows production was still publishing, in two groups. The four Hyperliquid WebSocket rows retracted the day before, whose permalinks were still serving 138.5 and 125.3 hours of outage for a venue whose REST API answered 99.98% of checks throughout — the socket was up24’s own, stuck in CONNECTING. And every incident opened before the Kraken fix landed on 13 Aug, ids 1–18: fifteen of them are 410 seconds long to within 80 ms, which is the 120-second recovery hold rather than an outage, and thirteen sit on ws:book or ws:trade with a connection error class — six real disconnects published as twelve incidents. Deleting rather than retracting, argued in the commit: the board was unlaunched, so no reader had seen the original claim. From the first outside reader on, corrections retract. History starts 13 Aug 13:29 UTC under one rule throughout.

d9fbf3c · 21 incidents removed

RULE

A stalled feed is degraded, not an outage, until half the window is silent

The outage rule was written for requests, where a failed check is a failed check. A stream window fails when its longest silence exceeds the venue’s stall threshold, and one pause straddling a window boundary fails two of them — 33% at six windows a minute, and an "outage" for a book that went quiet for twelve seconds. Deribit’s daily settlement pause at 08:00 UTC opened a four-minute outage every day. A connected socket that stops delivering is now degraded, and an outage only once more than half of the five-minute window has been silent. Socket-level failures keep the request rules: a venue refusing the connection is down whatever the feed would have said. Migration 0010 applied the same grade backwards, keeping the original severity in the row’s meta.

29f9a69 · every stalled outage in the table re-graded to degraded

CORRECTION

Hyperliquid’s four WebSocket incidents withdrawn as up24’s own fault

Migration 0010 added retracted_at and retracted_reason and set them on four Hyperliquid rows that recorded up24’s stuck socket as 138 and 125 hours of venue outage. The mechanism is the one every correction uses from here: the row leaves every list, count, feed and the broadcast queue, and the permalink keeps serving with the reason and a noindex, so the correction is as public as the claim was.

29f9a69 · 4 incidents retracted

RULE

Count one bad afternoon as one incident, not a hundred

179 of the fleet’s last 200 incidents were Gemini — three REST endpoints over two days, while the other eleven venues produced 21 between them. Gemini is not broken; it is spiky: a p95 of 140–160 ms jumping to 600–1100 ms for a few minutes at a time, which is exactly what the latency rule exists to catch, and every verdict was correct. The counting was not. The median gap between consecutive Gemini incidents was 370 seconds and the recovery hold was 120, so every quiet stretch closed an incident and every next spike opened another; the shortest lasted ten seconds. A component that has opened two or more incidents inside an hour now has to stay quiet for fifteen minutes before its incident closes. The board still goes green the moment the venue answers normally, and ended_at is still the last recovery. Migration 0009 applied the rule backwards, keeping the earliest incident of each run so the URL published first still works.

0e3736d · ~180 near-duplicate incidents collapsed into their runs

RULE

One socket drop is one incident, and ended_at is when the venue came back

Kraken opened 11 of the fleet’s first 18 production incidents in a day, and the investigation found two reporting faults of up24’s that mattered more than the venue did, both fleet-wide. The probe denormalises connection state onto every stream row, so one socket drop failed ws:book and ws:trade in lockstep and the engine opened one incident each — four Kraken drops produced eight of the eleven. Socket-scoped faults are now recorded once against ws:conn. And every duration was fiction: ended_at was stamped when the clearing rule finished rather than when the venue came back, so 15 of the first 18 incidents were all exactly 410 seconds long, a number describing the state machine. Kraken’s drops themselves are Cloudflare-initiated disconnects, which Kraken documents and its page now carries with a link.

34bdd86 · 8 of Kraken’s 11 incidents were one drop counted twice; 15 durations were wrong

RULE

Read stall thresholds off the distribution, not off memory

Every stallMs in the registry was picked from a fifteen-minute soak during D2 and never revisited. The one time a number was questioned — Kraken’s book, on a 12.1 s reading — the answer went into PLAN as "the tightest margin in the fleet", which was wrong: Binance’s book is 3 s and three venues sit at 5 s. Nobody had the distribution in front of them. The tool now reports p50 through p99.9 and the worst window per stream at the grain the threshold is enforced at, excluding windows where the socket was not up throughout and windows marked as up24’s own fault — tuning on those would raise every threshold until it was high enough to hide disconnects.

c25639b · nothing already published was re-graded