Measured latency, time to first token and uptime for ~45 AI inference APIs, by region.
Independent, provider-neutral latency & uptime for AI inference APIs — measured, not scraped.
🌐 Live: llmlatency.dev · 📊 JSON API · 🤖 MCP server · 🗓️ Deprecation calendar
Most "AI API latency" numbers come from the providers themselves, or from a benchmark run once and never updated. This project measures it continuously, from multiple regions, and publishes the result as an open dataset.
The site is a self-updating static site (Cloudflare Pages). The value isn't the code — it's the continuously-accumulated, distributed measurement archive. The code is open so the methodology is transparent.
# All regions, provider rankings for the last 24h — measured latency + uptime:
curl https://llmlatency.dev/api/rankings.json
/api/rankings.json · OpenAPI: /openapi.jsonAccept: text/markdown to any page URL, or append .md./llms.txt (index) and /llms-full.txt (full corpus).There's a real MCP server (Streamable HTTP) exposing a get_ai_api_latency tool backed by the live data:
curl -X POST https://llmlatency.dev/mcp \
-H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call",
"params":{"name":"get_ai_api_latency","arguments":{"region":"eu-hetzner"}}}'
The hosted endpoint above needs no setup. If you prefer a local stdio server (or want to build it from source), mcp_server.py is a dependency-free proxy over the same public JSON API:
python3 mcp_server.py # stdio MCP, stdlib only
# or
docker build -t llm-latency-mcp . && docker run -i llm-latency-mcp
Also available: an MCP Server Card (/.well-known/mcp/server-card.json), a browser WebMCP tool, an API catalog (RFC 9727) and an Agent Skills index. Regions: eu-hetzner, us-central, ap-tokyo, sa-east (omit for all).
config.py — registry of providers + this node's REGION (env)
probe.py — network probe (DNS→TCP→TLS→TTFB, stdlib, no key) + inference probe (TTFT, needs key)
run.py — one probe cycle across all providers (run on a schedule)
db.py — SQLite time-series (the accumulated measurement archive)
aggregate.py — measurements → p50 / p95 / uptime rankings per region & provider
sitegen.py — rankings → static site (JSON API, OpenAPI, llms.txt, schema.org, MCP surface)
ingest.py — central endpoint that collects measurements from remote probe nodes
ship.py — probe node → central node shipper (watermark-based, never loses data on outage)
deprecations.py — model deprecation/migration calendar (only verified, sourced entries)
Each probe node runs with its own REGION, measures every provider, and writes to the time-series. For multi-region, remote nodes ship their measurements to a central node that aggregates and builds the site.
git clone https://github.com/mazamaka/llm-latency-tracker
cd llm-latency-tracker
REGION=local python3 run.py # take edge-latency measurements
python3 aggregate.py --region local # see the ranking from this location
Runs on plain Python 3.12+ (standard library). httpx / loguru are optional.
Inference probes (real TTFT):
cp .env.example .env # add keys for the providers you want to measure
pip install -r requirements.txt
REGION=local python3 run.py
python3 aggregate.py --region local --type inference
Build the site locally:
BASE_URL=https://example.com python3 sitegen.py # → ./site/
python3 -m pytest -q # tests
See deploy/ for a container + a generic multi-region deployment guide.
Especially welcome:
Provider(...) entry in config.py (host + public models endpoint is enough for edge probes).pytest + ruff on every push.See CONTRIBUTING.md for dev setup, how to add a provider/region, and PR guidelines. Please keep the project's principle: measured, not scraped, and honest about the dataset's age.
Measured latency across 45 AI inference providers in 4 regions. Method: distributed edge (DNS→TCP→TLS→TTFB) + inference (TTFT) probes, last 24h. License: CC-BY-4.0.
| Region | Fastest provider (p50) | p50 | p95 | Uptime |
|---|---|---|---|---|
| Asia (Tokyo) | fireworks | 20 ms | 67 ms | 100% |
| Europe (Germany) | nscale | 98 ms | 201 ms | 100% |
| South America (São Paulo) | openrouter | 58 ms | 91 ms | 100% |
| US (Central) | fireworks | 28 ms | 75 ms | 100% |
data/rankings/2026-09-12.json (latest)10.5281/zenodo.21954788 — daily aggregates, CC-BY-4.0swh:1:snp:2778cbabd72a70a629ee35fbd5ac536d1ccb7a9aSnapshot generated 2026-09-12T07:48:44Z — this table is regenerated daily.
This listing does not have a supported local package template. Use the maintainer’s documentation for its hosted endpoint, authentication, and client-specific setup. No install command has been inferred.