Summarizing long emails
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Low-complexity extraction at high volume: near-zero cost per token, with 1M of context and caching.
Open alternative: Gemma 4 12B
open model guide · july 2026
80 enterprise use cases and the open-weight model that solves each one. No closed APIs: everything here can be downloaded, audited and run under your control.
quick selector
Pick the type of task and your main constraint. The recommendation updates instantly; the 80 cases in detail are just below.
What kind of task?
What is your main constraint?
Qwen3.8-27BAPACHE 2.0
The newest of the Qwen generation and the strongest open writer you can run on one machine: 52 on the Artificial Analysis index, Apache 2.0, and native multilingual with solid Spanish.
the 80 cases
Filter by area or complexity. Each card gives the primary recommendation, the alternative and why. The green dot marks the models served in Helmcode on a flat rate.
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Low-complexity extraction at high volume: near-zero cost per token, with 1M of context and caching.
Open alternative: Gemma 4 12B
Recommended open model
Qwen 3.6 35B
available in Helmcode
Native multilingual with a strong Spanish register, and 3B active parameters out of 35B, which is what makes it affordable at email volume. When a person signs the email, the writing matters.
Open alternative: Qwen3.8-27B for the ones that matter most
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
1M of context takes a whole report, case file or book in one pass, with no chunking. Bigger windows exist on weaker models, and a window only says what fits, not what the model understood.
Open alternative: Kimi K3, the open leader on long-context reasoning
Recommended open model
Voxtral Small
One model transcribes and summarises: Voxtral handles audio understanding, not just words, so a 30-minute meeting comes back as minutes without a second model in the chain. Spanish is one of its eight languages.
Open alternative: Whisper large-v3 + V4 Flash, for meetings past 30 minutes
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
The context is already retrieved: extraction plus writing with precise citations. For confidential docs, self-hosting is the only valid option.
Open alternative: Qwen 3.6 35B
Recommended open model
Gemma 4 12B
Classification against clear criteria in milliseconds. It fits on a modest GPU while it processes every incoming email.
Open alternative: DeepSeek V4 Flash
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Entity extraction (what, who, when) with reliable JSON. At this volume, cost is what decides.
Open alternative: Gemma 4 12B
Recommended open model
Qwen3.8-27B
Follow-ups should read as human, not automated. Same 27B dense footprint as the 3.6 generation it replaces, fourteen points higher on the index.
Open alternative: Qwen 3.6 35B, the MoE we serve on a flat rate
Recommended open model
Qwen 3.6 35B
available in Helmcode
It respects the author’s voice instead of rewriting it, which is the real difference in editing. Multilingual by design and cheap enough to run over everything the team writes.
Open alternative: Qwen3.8-27B, or GLM-5.2
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Frontier-adjacent quality at the lowest cost on the market. For fine cultural adaptation, step up to Qwen 3.6 35B.
Open alternative: Qwen 3.6 35B (ES/EN/FR/DE/PT)
Recommended open model
qwen3-embedding + rerank
available in Helmcode
The forgotten half of every RAG stack: good embeddings plus reranking decide more than the generator model does.
Open alternative: DeepSeek V4 Flash (synthesis)
Recommended open model
qwen3-embedding + rerank
available in Helmcode
Reordering the retrieved top-k multiplies RAG precision at a minimal marginal cost.
Open alternative: Gemma 4 12B as cross-encoder
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Comprehension plus rephrasing into user language. Ideal for manuals and policies that cannot leave for external APIs.
Open alternative: Qwen3.8-27B
Recommended open model
GLM-5.2
available in Helmcode
Data synthesis plus narrative: the best open reasoning following complex structured templates.
Open alternative: Qwen 3.6 35B
Recommended open model
GLM-5.2
available in Helmcode
One conclusion per slide, title as message. Wired to python-pptx or reveal.js: a full data-to-deck pipeline.
Open alternative: Qwen 3.6 35B
Recommended open model
Qwen 3.6 35B
available in Helmcode
Precise language and consistent terminology across long documents. Internal content that should stay at home.
Open alternative: GLM-5.2
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
A 1-3h meeting is 50-150K tokens, so 1M of context holds a whole week of them. Voxtral Small transcribes, V4 Flash writes the agenda and the minutes, both self-hosted.
Open alternative: GLM-5.2 for meetings that end in a decision
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
RSS/scraping → batch summary and classification → dashboard. Cost per article is practically zero.
Open alternative: Gemma 4 26B
Recommended open model
Gemma 4 12B
One of the simplest tasks on the list: millisecond latency, on-premise, on modest hardware.
Open alternative: DeepSeek V4 Flash
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Task status, alerts and summaries. Milestones and owners of client projects never leave the internal network.
Open alternative: Qwen3.8-27B
Recommended open model
GLM-5.2
available in Helmcode
The best open reasoning for synthesizing sources you have already gathered. Pair it with your own scraping for the capture step.
Open alternative: DeepSeek V4 Pro 0813
Recommended open model
DeepSeek V4 Pro 0813
Still the strongest open model for bounded generation, by a margin that has almost closed: 53 on the Artificial Analysis index against Flash 0731’s 52, at 2.3x the price. DeepSeek publishes 80.6% on SWE-bench Verified, its own figure.
Open alternative: Kimi K3, or K2.7 Code on less hardware
Recommended open model
GLM-5.2
available in Helmcode
The open leader in agentic coding: Z.ai publishes 62.1% on SWE-bench Pro and 81.0 on Terminal-Bench 2.1, ahead of DeepSeek on multi-step work over real repos. GLM-5.3 scores higher and its weights are not out, so this is the newest GLM you can run.
Open alternative: Kimi K3, or K2.7 Code on less hardware
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Spots bugs, debt and security issues by reasoning over the full diff. When every pull request goes through it, price and speed decide: $0.11 per task against V4 Pro’s $0.25, and 1.7x the tokens per second. The code never leaves your infrastructure.
Open alternative: DeepSeek V4 Pro 0813 for the hardest diffs
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Edge-case coverage and coherent mocks, at the lowest price per task in the open field. DeepSeek publishes that Flash 0731 beats the V4 Pro preview on all nine of its agent benchmarks, its own figures, run with its own harness.
Open alternative: Kimi K2.7 Code
Recommended open model
GLM-5.2
available in Helmcode
Hours-long multi-file refactors: the area where GLM-5.2 pulls furthest ahead (FrontierSWE, DeepSWE), with the 500K of context we serve. K3 is the open ceiling today if you have the cluster to serve it, around 64 GPUs.
Open alternative: Kimi K3, or K2.7 Code on less hardware
Recommended open model
Qwen3.8-27B
Text-to-SQL over your schema with a few-shot prompt. This is code, so the coding generation matters: Alibaba publishes 90.3 on LiveCodeBench v6 for this model. The database schema is sensitive information: better kept at home.
Open alternative: Qwen 3.6 35B, or V4 Flash
Recommended open model
Kimi K2.7 Code
Reads the code, writes the docs. K2.7 keeps coherence across modules in large codebases.
Open alternative: GLM-5.2
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Resolves 80% of L1 queries by loading the whole KB into context. Employee and systems data: inside the network.
Open alternative: Gemma 4 26B
Recommended open model
GLM-5.2
available in Helmcode
Multi-step reasoning connecting unrelated events, with a method (5 Whys, Ishikawa). Confidential logs: self-host.
Open alternative: DeepSeek V4 Pro 0813
Recommended open model
Qwen3.8-27B
Reads the mockup and writes the component, with native vision and current-generation coding in the same 27B weights. This case needed a 64-GPU cluster until August; now it needs one machine, under Apache 2.0.
Open alternative: Kimi K3 if you already run the cluster
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Training and test datasets at zero marginal cost. The MIT license places no restriction on how you use the outputs.
Open alternative: Qwen 3.6 35B
Recommended open model
Gemma 4 12B
Fine-tuned on your own categories of personal data, it runs at the edge of the pipeline before anything leaves.
Open alternative: DeepSeek V4 Flash
Recommended open model
Voxtral Small
The most accurate open-weights transcription measured by Artificial Analysis: 2.8% word error rate against 4.1% for the best Whisper large-v3 row. Customer calls should never pass through an opaque API.
Open alternative: Whisper large-v3 for wider language coverage on less hardware
Recommended open model
Chatterbox Multilingual
23 languages, Spanish among them, voice cloning from a few seconds of reference, MIT licence and 500M parameters. A full open voicebot: Voxtral → LLM → Chatterbox. Kokoro is lighter and cheaper and only speaks English, which rules it out of a Spanish deployment.
Open alternative: Kokoro, for English only and the smallest footprint
Recommended open model
Qwen3.8-27B
Processes the document image directly, no separate OCR step, and returns the fields as structured JSON. A wrong digit on an ID or an invoice costs more than the inference does, so this is where the stronger vision model earns its keep. Invoices and IDs demand a 100% internal pipeline.
Open alternative: Gemma 4 26B, served on a flat rate
Recommended open model
Qwen3.8-27B
Visual QA, catalog tagging and asset verification, over video as well as stills, and fine-tunable to your own acceptance criteria under Apache 2.0.
Open alternative: Gemma 4 26B, or MiniMax M3 for heavier batches
Recommended open model
Gemma 4 26B
available in Helmcode
Text and image in the same call, fine-tunable to your platform’s criteria for more consistency than zero-shot.
Open alternative: MiniMax M3
Recommended open model
Qwen3.8-27B
Persuasive writing with frameworks (SPIN, Challenger) and a customer-centered narrative, and enough head to hold the argument across a long proposal.
Open alternative: GLM-5.2
Recommended open model
Gemma 4 12B
Fine-tuned on your conversion history, it beats zero-shot by a wide margin: the structural advantage of open source.
Open alternative: DeepSeek V4 Flash
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Briefings from CRM activity. Amounts and clients under negotiation: commercial intelligence that must not leave.
Open alternative: Qwen3.8-27B
Recommended open model
Qwen 3.6 35B
available in Helmcode
Objection handling and personalization by profile. The sales script is confidential strategy: self-host.
Open alternative: GLM-5.2
Recommended open model
GLM-5.2
available in Helmcode
The most solid open strategic reasoning you can actually deploy. K3 scores higher, but 1.6 TB of weights puts it out of reach unless you already run a cluster.
Open alternative: Kimi K3
Recommended open model
Qwen 3.6 35B
available in Helmcode
Creativity, copywriting and adaptation to brand tone. At high volume, the saving over a closed API is substantial.
Open alternative: GLM-5.2
Recommended open model
Qwen3.8-27B
Dozens of variations per campaign at zero marginal cost, and it reads the creative as well as the brief: native vision over image and video comes in the same weights.
Open alternative: Qwen 3.6 35B, or V4 Flash at the highest volumes
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Structured batch generation across the whole catalog or blog. High volume, clear criteria: cost is what rules.
Open alternative: Qwen3.8-27B
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Thousands of opinions in batch with aspect-based analysis. Fine-tuning on your sector’s vocabulary sharpens the result.
Open alternative: Gemma 4 12B (fine-tune)
Recommended open model
Gemma 4 12B
Your platform’s own criteria, fine-tuned, at a practically zero cost per message.
Open alternative: Gemma 4 26B (multimodal)
Recommended open model
Gemma 4 12B
Real-time multi-label, fine-tunable on your ticket history for business-specific precision.
Open alternative: DeepSeek V4 Flash
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
A support chatbot with customer data under GDPR: in regulated sectors, self-hosting is not optional.
Open alternative: Qwen3.8-27B
Recommended open model
Qwen 3.6 35B
available in Helmcode
Empathy plus firmness plus a concrete solution. Complaints contain personal data: process it at home.
Open alternative: GLM-5.2
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Thematic classification and sentiment over thousands of open responses, with the personal data never leaving EU infrastructure.
Open alternative: Gemma 4 12B
Recommended open model
Qwen3.8-27B
Automatic closure of each ticket with a summary for the CRM. High volume, sound writing, one machine.
Open alternative: Qwen 3.6 35B, or V4 Flash at the highest volumes
Recommended open model
Qwen3.8-27B
Telling evidence from generic claims, without bias. Candidate data is GDPR territory: deploy locally.
Open alternative: GLM-5.2
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
All the onboarding documentation fits in 1M of context, with no RAG and without internal policies leaving.
Open alternative: Gemma 4 26B
Recommended open model
Qwen3.8-27B
Constructive feedback with nuance. This data can end up in labor proceedings: self-host plus encryption.
Open alternative: GLM-5.2
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
HR data is the most sensitive in the company: the conversations do not leave the corporate infrastructure.
Open alternative: Qwen3.8-27B
Recommended open model
Qwen3.8-27B
Persuasive writing with a defined structure, in native Spanish, even for confidential roles.
Open alternative: Qwen 3.6 35B
Recommended open model
Qwen 3.6 35B
available in Helmcode
Instructional design adapted to the learner’s level. In high-volume e-learning, cost per module tends to zero.
Open alternative: GLM-5.2
Recommended open model
Qwen3.8-27B
Good tone and structure on one machine. Internal communications stay out of third-party APIs.
Open alternative: Gemma 4 26B, or Qwen 3.6 35B
Recommended open model
Gemma 4 26B
available in Helmcode
Processes the invoice image directly and returns structured JSON. A 100% internal financial pipeline.
Open alternative: MiniMax M3, self-hosted
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
One index point behind V4 Pro at less than half the cost, with the same 1M of context for a full set of accounts and its notes. For listed companies or M&A, self-hosting is practically mandatory.
Open alternative: DeepSeek V4 Pro 0813, or GLM-5.2
Recommended open model
GLM-5.2
available in Helmcode
Complex patterns over the full transaction history (1M ctx). In banking, the data stays in: the only viable option.
Open alternative: DeepSeek V4 Pro 0813
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
High volume of short comparisons: statement lines against ledger entries, with the near-matches flagged for a person. Cost per token is what decides here.
Open alternative: Gemma 4 26B
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Reads the accounts and writes the reasoning behind a limit, ratio by ratio, at $0.11 per file. That is a solvency judgement on a named company: it belongs self-hosted, and a person still signs it.
Open alternative: GLM-5.2 for the files that go to committee
Recommended open model
Qwen 3.6 35B
available in Helmcode
The same message has to escalate from a nudge to a formal notice without losing the client. Register weighs more than reasoning, which is where Qwen leads in Spanish.
Open alternative: Gemma 4 26B
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Not extraction, judgement: this dinner is over the per diem, this taxi has no attendees. The policy travels in the prompt and changes without retraining anything.
Open alternative: Gemma 4 12B
Recommended open model
GLM-5.2
available in Helmcode
The highest open level of nuance comprehension, with MIT license and model card published: auditable end to end.
Open alternative: DeepSeek V4 Pro 0813
Recommended open model
GLM-5.2
available in Helmcode
With the AI Act applying in general since 2 August 2026, an open, traceable stack on EU infrastructure gives your assessment the evidence it needs.
Open alternative: Qwen 3.6 35B
Recommended open model
Qwen 3.6 35B
available in Helmcode
Tagging and triage of case files by matter, jurisdiction and urgency, without a single page leaving the firm.
Open alternative: Gemma 4 26B
Recommended open model
Kimi K3
Thousands of documents that only mean something read together, so what matters is reasoning across them, not window size: K3 leads every open model on AA-LCR, the long-context reasoning eval, at 82.7%. Serving it takes a cluster, around 64 GPUs.
Open alternative: DeepSeek V4 Flash, 1M of context on hardware you already have
Recommended open model
GLM-5.2
available in Helmcode
Reasoning that has to hold a chain of citations without inventing one. Ground it with RAG on your own database of rulings, never on the model memory.
Open alternative: DeepSeek V4 Pro 0813
Recommended open model
Qwen 3.6 35B
available in Helmcode
Generation rather than analysis: assembling a first draft out of clauses legal has already approved, so the team edits instead of starting from nothing.
Open alternative: GLM-5.2
Recommended open model
GLM-5.2
available in Helmcode
Watching what changed in a regulation and which internal policies it touches. A different job from checking compliance: this one runs before anybody is out of it.
Open alternative: DeepSeek V4 Flash
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Comparing terms and flagging expirations. Agreed prices are confidential commercial information.
Open alternative: GLM-5.2 (every contract at once)
Recommended open model
GLM-5.2
available in Helmcode
Structured reasoning over sources you have already gathered. For web capture, combine it with your own scraping.
Open alternative: DeepSeek V4 Flash
Recommended open model
Qwen3.8-27B
Hundreds of requirements answered one by one, each traceable to a document you already hold. The bid stays confidential until the envelope is opened.
Open alternative: GLM-5.2
Recommended open model
GLM-5.2
available in Helmcode
The same two hundred questions arrive from every client with the wording changed. Answered from your own evidence base, with a person signing it off.
Open alternative: DeepSeek V4 Flash
Recommended open model
DeepSeek V4 Flash 0731
available in Helmcode
Thousands of short events a day, each needing a route: delay, damage, wrong address or nothing at all. Volume picks the model.
Open alternative: Gemma 4 26B
Recommended open model
Qwen3.8-27B
Reads the picture off the line, or the clip, and writes the defect report against your own criteria. Production images rarely have permission to leave the plant.
Open alternative: Gemma 4 26B, served on a flat rate
*This licence is not a free and open-source one (use restrictions and specific conditions): read it before you build on it.
In production 99.5% of our tokens go through open models, and not on the easy tasks: classifying, extracting, summarizing, drafting, answering and reasoning over documents nobody else has read. What is left is a narrow set of genuinely frontier problems, and the answer to those is to route them to the big model, not to pay frontier prices for the other 99.5%.
go deeper
The 80 cases above roll up into twelve canonical use cases. Each has its own page with architecture, FAQ and deployment in detail.
get started
Pick a case from the list, point the same code at our endpoint and read both outputs side by side.
The API you already call. One afternoon. Nothing to migrate.
A process with volume and stable criteria can become a small, specialized model with the weights in your name, running on far less hardware. model_specialization →
// cookies
We only use strictly necessary cookies to run the site. No analytics, no advertising, ever — see our Cookie Policy.
// preferences