// opendata

The real numbers
behind the platform.

Real usage data from our inference platform, aggregated and anonymized. No prompts, no content; just request-level counters.

Updated weekly · 22 Jun 2026

333.8B

Tokens processed

cumulative

27.6M

Requests served

cumulative

9

Active models

in production

// models

Tokens by model.

Cumulative tokens processed per model. One open model carries most of the load, and the full stack is always available.

01 Qwen 3.6 76.1% 254 B · 24.6 M
02 DeepSeek V4-Flash 12.4% 41.2 B · 848.6 K
03 MiMo V2.5 8.6% 28.6 B · 414.4 K
04 Gemma 4 2.4% 8 B · 1.2 M
05 Qwen3 Embedding 0.5% 1.8 B · 457.7 K
06 Qwen3 Coder <0.1% 142.8 M · 7.1 K

// tokens

Input vs. output.

Inference here is overwhelmingly read-heavy, long prompts, retrieval and context, with a thin slice of generated tokens.

Input · prompt 329 B 98.6%
Output · generated 4.8 B 1.4%

// usage

Tokens per day.

Daily tokens processed over the last 90 days, peaking at 3.1 B/day.

0 1 B 2 B 3 B 4 B 3.1 B peak
90 days agotoday

// beyond text

Speech & reranking.

The stack is more than LLMs, transcription, synthesis and reranking run on the same API.

Speech-to-text Whisper Large v3 18.5 K requests
Text-to-speech Kokoro 15.8 K requests
Reranking Qwen3 Reranker 2.9 K requests

// clients · last 7 days

How teams connect.

Drop-in OpenAI compatibility in the wild, the official SDK and OpenCode account for the vast majority of traffic.

OpenAI SDK (Python) 49.2% 233 devs
OpenAI SDK (JS) 12.4% 59 devs
Node.js 8.2% 39 devs
curl 6.5% 31 devs
Python httpx 6.1% 29 devs
Vercel AI SDK 5.3% 25 devs
Python aiohttp 4% 19 devs
Python requests 3.6% 17 devs
Others 4.6% 22 devs

// geography · last 7 days

Where requests come from.

84.1% of traffic originates inside the EU, the audience this infrastructure is built for.

Spain 30.5% 949.1 K
Finland 27.5% 857.3 K
Germany 23.9% 745.1 K
Colombia 5.1% 158.1 K
United States 4.9% 152.4 K
United Kingdom 2.1% 66.8 K
Argentina 1.7% 52.3 K
Mexico 1.4% 42 K
France 1.1% 35 K
Netherlands 0.4% 13.1 K
Ireland 0.4% 10.9 K
Poland 0.3% 7.7 K
Others 0.8% 24 K

// performance

Latency & throughput.

Median time to first token and sustained throughput per model, measured on 22 Jun 2026.

Model TTFT p50 Throughput
Qwen 3.6 1.4 s 468 rpm
DeepSeek V4-Flash 6.6 s 33.3 rpm
Gemma 4 174 ms 14.4 rpm
MiMo V2.5 1.4 s 9.5 rpm

TTFT p50 = median time to first token · Throughput = sustained requests per minute.

// who it's for

Two ways to run on this stack.

These numbers come from real workloads across the community and private deployments alike.

Builders & community

Frontier models, fair price, no data sharing.

Access the latest open models at a reasonable cost, without handing over your data, through the NaN community.

nan.builders →
Startups & enterprise

Private, dedicated inference with SLAs.

Dedicated infrastructure, support and contractual SLAs, flat rate, EU data, OpenAI-compatible.

see_pricing →

Methodology. Figures are aggregated, anonymized counters collected at the request level. Helmcode keeps zero logs, no prompt or completion content is ever stored. Cumulative metrics span the platform's lifetime; windowed metrics are labelled per section.

// get started

START BURNING TOKENS

Skip the AI infra work. Deploy your first private inference endpoint today.

Flat rate. EU data. OpenAI API compatible.