// models
Open models.
Frontier results.
The open frontier already covers the work enterprises actually run, retrieval, code, agents, extraction. Benchmarked against the closed labs, proven in production, served only from the EU.
// our models
Models we serve at Helmcode.
The complete catalogue on the Shared cluster: 8 models behind one OpenAI-compatible API, all of them open-weight, all served from the EU. Model IDs are shown as-is, copy them straight into your code.
Language models
glm-5.2 MIT 744B MoE · 40B active · 500K ctx · add-on Long-horizon agentic coding and hard maths. Sold per key, €150 a month. deepseek-v4-flash MIT 284B MoE · 21B active · NVFP4 · 1M ctx Reasoning, agents and long-context analysis. Native tool calling. qwen3.6 Apache 2.0 35B MoE · 3B active · NVFP4 · 256K ctx High-volume RAG, classification and code. Multimodal input. gemma4 Gemma 26B MoE · 4B active · NVFP4 · 256K ctx Efficient assistants, documents and scans. Multimodal input. Embeddings & reranking
qwen3-embedding Apache 2.0 8B · 4096 dims · Float32 · MMTEB 70.58 The retrieval layer of a RAG pipeline, 100+ languages. rerank Apache 2.0 Qwen3-Reranker-8B · BF16 · /v1/rerank Reorders the retrieved top-k by real relevance. Speech
whisper-large-v3 MIT 99+ languages · 3.2% WER (Spanish) Speech to text for calls, meetings and media. kokoro Apache 2.0 82M params · <1s latency · 67 voices Text to speech for live agents and IVR, Spanish included. Every model here runs on Helmcode today. On Dedicated and On-premise plans we also serve custom or fine-tuned open-weight models on hardware reserved for you.
Full model reference in the docs// the 80 / 20
Open models cover 80% of enterprise inference.
The work that runs your business, RAG, classification, code generation, internal assistants, is solved today by open-weight models. The 20% where you truly need a frontier closed model is narrower than it looks.
Bleeding-edge reasoning at the very limit of capability. Real, but specific, and rarely the workload a regulated team needs to keep in-house. We are honest about that boundary instead of pretending it doesn't exist.
// benchmark · artificial analysis
Frontier-class intelligence, open-weight prices.
The Artificial Analysis Intelligence Index is a composite of nine hard evaluations, GPQA Diamond, SciCode, Terminal-Bench, Humanity's Last Exam and more. Open-weight models now sit just under the closed frontier.
Kimi K3 leads the models you can actually download, at 60 against 63 for Claude Opus 5, the closed leader: a three-point gap. GLM-5.3 ties K3 on the index and its weights are not published yet, so it is not a row here. GLM-5.2 follows at 53, level with GPT-5.6 Luna (52), and we serve it: the highest-scoring model on the platform, sold as a per-key add-on. Behind it sit the models your plan covers: DeepSeek V4 Flash 0731 (52), Qwen3.6 35B (32) and Gemma 4 26B (26). V4 Flash scores 52 at $0.22 / $0.66 per million tokens against Claude Opus 5's $5.00 / $25.00, that's ~23× cheaper to read and ~38× cheaper to write, for 83% of the index leader's intelligence.
Source: artificialanalysis.ai · Intelligence Index v4.1.1 · August 2026 · composite of 9 evaluations · index rounded to whole points · price = first-party API, per 1M tokens (input · output). The two DeepSeek figures are the off-peak rate and double during peak hours, 01:00-04:00 and 06:00-10:00 UTC, under the rate card DeepSeek introduced on 16 August 2026. Rows marked on helmcode, GLM-5.2, DeepSeek V4 Flash 0731, Qwen3.6 35B and Gemma 4 26B, are the ones we serve, with GLM-5.2 sold per key rather than covered by the plan. Kimi K3 and DeepSeek V4 Pro 0813 are open-weight but not on the platform: K3 needs roughly 64 GPUs of the H100 or B200 class to serve, and its weights ship under Moonshot's own K3 licence rather than an OSI-approved one.
// proven in production
The production numbers.
On our own platform, near-enough all inference already runs on open models, and most of it on a single 35B one.
333.8B
Tokens in production
cumulative
76%
Run on Qwen 3.6 (35B)
the open workhorse
99.5%
Of tokens on open models
LLM traffic
// the lineup
What actually runs the 80%.
The three language models, ordered by real production token share. One open 35B model carries most of the load, the rest step in for reasoning, scale and multimodal.
qwen3.6 35B MoE · 256K ctx 76.1% High-volume RAG, classification, code deepseek-v4-flash 284B MoE · 1M ctx 12.4% Reasoning, agents, long-context gemma4 26B MoE · 256K ctx 2.4% Efficient assistants, document work Plus embeddings & reranking (qwen3-embedding, rerank) and speech (kokoro, whisper-large-v3), eight models on one API. Full model reference →
// models faq
Open models, answered.
The questions everyone asks before trusting open models in production.
Are open models actually good enough?
For the work enterprises run day to day, yes. On Artificial Analysis’ Intelligence Index, the leading open-weight model, Kimi K3, scores 60 against 63 for the closed leader, Claude Opus 5, and the model we serve, DeepSeek V4 Flash 0731, delivers 83% of leader-level intelligence for cents per million tokens. In production, 99.5% of all tokens on Helmcode already flow through open models. The gap that remains is a narrow set of frontier tasks most teams never hit.
Which model should I use?
Start with Qwen 3.6, it carries three quarters of all production traffic and is the fastest, cheapest path for RAG, classification and code. Move to DeepSeek V4-Flash for hard reasoning, agents or 1M-token context. For image and audio input, Qwen 3.6 and Gemma 4 are both multimodal. Same API, just change the model id.
What about the 20% that genuinely needs a frontier model?
It exists, and it is more specific than most assume, frontier-only reasoning at the very edge of capability. Helmcode is honest about that boundary: we cover the 80% that runs your business, privately and at a flat rate, not the last mile of the leaderboard.
How current are these benchmarks?
Figures are published scores as of July 2026, open models served on Helmcode, closed-model numbers from vendor reports. Benchmarks move every release, so treat them as directional. What does not move is where your data is processed: always the EU, always zero logs.
Can I run a model that is not listed?
On Dedicated and On-premise plans, yes, custom or fine-tuned open-weight models on hardware reserved for you. The Shared cluster serves the curated lineup above.
// get started
START BURNING TOKENS
Skip the AI infra work. Deploy your first private inference endpoint today.
Flat rate. EU data. OpenAI API compatible.