Documentation

Scrape it

Four gauge families at https://api.up24.app/metrics, in the Prometheus text format. It is the same evidence the board renders — the incident engine decides, and this repeats what it decided in another shape. Nothing here is derived at scrape time.

75 components across 15 venues, from 3 vantage points. No key, no account, and free during the pilot; it is on the pricing page as one of the things that may be paid later.

The job

prometheus.yml
scrape_configs:
  - job_name: up24
    scheme: https
    static_configs:
      - targets: ['api.up24.app']
    # The board is recomputed every ten seconds and cached at the edge for
    # five. Anything faster than 30s buys nothing and is served from cache.
    scrape_interval: 30s

The path is /metrics, which is where a Prometheus job looks by default — so there is no metrics_path line above and this is the one integration on the site that needs no configuration beyond the host. It sits inside the same 600 requests a minute the rest of the API does, counted at the origin; a 30-second job spends about two of them an hour.

What you get

A scrape, abridged
# HELP up24_component_state Published state of one component: 0 operational, 1 degraded, 2 outage, 3 unknown.
# TYPE up24_component_state gauge
up24_component_state{venue="binance",component="rest:ticker"} 0
up24_component_state{venue="kraken",component="ws:book"} 1

# TYPE up24_component_state_region gauge
up24_component_state_region{venue="kraken",component="ws:book",region="eu-central"} 1
up24_component_state_region{venue="kraken",component="ws:book",region="ap-southeast"} 0

# TYPE up24_latency_seconds gauge
up24_latency_seconds{venue="binance",endpoint="ticker",region="eu-central",percentile="50"} 0.041
up24_latency_seconds{venue="binance",endpoint="ticker",region="ap-southeast",percentile="50"} 0.193

# TYPE up24_incident_open gauge
up24_incident_open{venue="kraken",component="ws:book",severity="degraded"} 1
up24_incident_open{venue="kraken",component="ws:book",severity="outage"} 0
FamilyLabelsValue
up24_component_statevenue, component0 operational, 1 degraded, 2 outage, 3 unknown. What the board says.
up24_component_state_regionvenue, component, regionThe same encoding, for what one vantage point saw on its own.
up24_latency_secondsvenue, endpoint, region, percentileThe most recent complete minute’s percentile, in seconds. REST only.
up24_incident_openvenue, component, severity1 while up24 holds an open incident of that severity, 0 otherwise.

Why the state is published twice

A component’s published state is a quorum. Each vantage point classifies it independently and the quorum-th worst reading is what gets published, so that one probe box having a bad afternoon cannot take a venue down. There is no region that owns that number, and putting a region label on it would be the first figure on this site that does not name its own vantage point.

So up24_component_state carries no region and is what the board, the API and the alerts all agree on. When you want the disagreement — the question “is the venue down, or is one probe sulking” — up24_component_state_region is the reading each box made before the quorum combined them. An expression written against the first cannot page you on one region dissenting; the second is there for when that is exactly what you want.

How the numbers are made has the quorum rule in full, and the grading rules have every threshold with the reason it is that number.

Nothing is ever missing, and that is the point

Every venue × component in the registry exports a series on every scrape — all 75 of them — whether or not up24 measured anything. A component with no evidence exports 3, and a region that shipped nothing exports 3, rather than not appearing.

That is deliberate and it is the whole reason to trust this endpoint. up24_component_state > 0 matches nothing when the series have vanished, so a rule written against a board that omits its blind spots resolves itself the moment the evidence disappears. An alert that clears because nobody is looking any more is the single most dangerous thing a monitoring product can ship. The same rule applies to up24 itself: if the incident engine stops writing for five minutes, every state in this document reads 3.

Latency is the one family that does omit, for the opposite reason. There is no value on a latency gauge that means “not measured” — 0 is a claim of zero milliseconds. An absent latency series is honest, and the state families above already say whether the silence is a problem.

Rules worth starting from

up24.rules.yml
groups:
  - name: up24
    rules:
      # The venues you actually trade. Not every venue up24 measures —
      # a page for somebody else's exchange is how a rota learns to ignore one.
      - alert: ExchangeApiDown
        expr: up24_component_state{venue=~"binance|kraken",component=~"rest:.*"} == 2
        for: 2m
        annotations:
          summary: '{{ $labels.venue }} {{ $labels.component }} is out, per up24'

      # Three is "up24 cannot see it", which is not the same as "it is down"
      # and belongs on a different, quieter channel.
      - alert: ExchangeApiUnknown
        expr: up24_component_state == 3
        for: 15m

      # Latency against its own recent normal, in one region — never blended
      # across two. Smoothed, because one minute is about twelve samples.
      - alert: ExchangeApiSlow
        expr: |
          avg_over_time(up24_latency_seconds{percentile="95"}[10m])
            > 3 * avg_over_time(up24_latency_seconds{percentile="95"}[6h] offset 1h)
        for: 10m

      # up24 itself. If the scrape fails, none of the above can fire.
      - alert: Up24ScrapeFailing
        expr: up{job="up24"} == 0
        for: 5m

The last one is not decoration. Every rule above is a statement about a series that has to arrive, and up is the only one that fires when it does not. up24 is one operator with two probe boxes and a small VPS; treat it as a second opinion beside your own instrumentation, not as the thing your rota depends on.

What the latency numbers are, exactly

One minute’s percentile over roughly twelve samples — the probe polls every five seconds — taken from the most recent complete minute, per endpoint and per vantage point. It is a raw figure and it is noisy at percentile="95", where twelve samples make the 95th percentile close to the maximum. Smooth it with avg_over_time before you alert on it.

Seconds, not the milliseconds the rest of this site speaks, and percentile rather than quantile. Both are Prometheus conventions rather than preferences: quantile is reserved for summary metrics, whose _sum and _count counters up24 does not have and will not invent, and base units mean this series sits on a dashboard beside your own instrumentation without a conversion in the middle of the expression.

Never combined across regions, here or anywhere. Latency is a property of a venue and the chair it is measured from: an exchange 12 ms from France is 190 ms from Singapore, and the mean of the two is a number no trader will ever experience. The nearest figure on the JSON API is typicalMs on /v1/status, which is a median of the last 24 hours of per-minute p50s and deliberately not a 24-hour p50 — percentiles do not compose.

If nothing has been rolled up in the last five minutes the series are absent rather than stale. up24_component_state is what says whether that silence means anything.