This is the at-a-glance truth view used to pick a coordination panel. A model needs three things to make the panel: reliable on two axes — we can reach it (not throttled to death) and its replies are usable (they parse, not truncate) — comparable (within a few points of its peers), and diverse-lineage (a different tech tree, so errors decorrelate; rows are tinted by lineage so you can spot it — note that Gemini and Gemma share Google's lineage and are not independent for panel purposes). The columns let you see all of it at once — and where the gaps still are.
The only membership rule is $0 to call — not which gateway. The access path is noted in small print under each model name (OpenRouter, Google AI Studio, …) so the numbers are reproducible on the same endpoints — and where a model has several free endpoints, the one in bold is the path we benched and reliability-checked on (the most reliable); we add any model that is freely accessible by some path, not just via OpenRouter. The public benchmarks (MMLU-Pro, GPQA) are third-party scores; MMLU-Pro-Mini-196 is our own custom stratified benchmark, run identically across every model.
| Current frontier models — comparison only · public access, not free | |||||||
| # | Model (generation) | Lineage | GPQA public | MMLU-Pro public | MMLU-Pro-Mini-196 our custom bench | Reach responds? | Usable reply parses |
|---|---|---|---|---|---|---|---|
| ★ | Claude Fable 5 (Gen 5 · flagship) | Anthropic | 92.6 | 91.5 | — | — | — |
| ★ | GPT-5.5 (Gen 6) | OpenAI | 95.8 | — | — | — | — |
| ★ | Gemini 3.1 Pro (Gen 5) | 94.1 | 91.0 | — | — | — | |
| ★ | MiniMax M3 (Gen 5) | MiniMax | 92.9 | — | — | — | — |
| ★ | Claude Opus 4.8 (Gen 5) | Anthropic | — | 89.5 | 86.2 | — | — |
| ★ | Claude Sonnet 5 (Gen 5) | Anthropic | — | — | 80.6 | — | — |
| ★ | GLM-5.2 (Gen 5) | Zhipu / GLM | — | — | — | — | — |
| ★ | Grok 4.3 (Gen 5) | xAI | ~88 | ~91 | — | — | — |
| ★ | DeepSeek V4 (Gen 5) | DeepSeek | — | — | — | — | — |
| Free & open-source models — used for research · our coordination-panel candidates | |||||||
| # | Model (access) | Lineage | GPQA public | MMLU-Pro public | MMLU-Pro-Mini-196 ours | Reach | Usable |
| 1 | nemotron-ultra-550b (OpenRouter · NVIDIA) | NVIDIA | —6 | —6 | 86.8† | ✓ | ⚠ 96%† |
| 2 | gemini-3-flash (AI Studio) | 90.4 | 89.0 | 20/day cap | ⚠ thin | — | |
| 3 | gemma-4-26b (OpenRouter) | 82.3 | 82.6 | 82.1 | ✓ | ✓ 100%† | |
| 4 | glm-4.5-flash (Z.ai) | Zhipu / GLM | —6 | —6 | 81.8 | ✓ | ⚠ 92%2 |
| 5 | nemotron-super-120b (OpenRouter · NVIDIA) | NVIDIA | —6 | —6 | 81.4† | ✓ | ✓ 96%† |
| 6 | deepseek-v3.2 (SambaNova) | DeepSeek | 82.4 | 85.0 | capped 39/1962 | ✓ | — |
| 7 | gemma-4-31b (OpenRouter · SambaNova) | 84.3 | 85.2 | 80.6 | ✓ | ✓ 100%† | |
| 8 | gemini-2.5-flash (AI Studio) | 82.8 | —6 | — | ⚠ thin | — | |
| 9 | mistral-large-3-675b (NVIDIA) | Mistral | ~44 | 78.0 | 78.7† | ✓ | ⚠ 38%† |
| 10 | nemotron-nano-30b (OpenRouter · NVIDIA) | NVIDIA | 73.0 | 78.3 | 78.5† | ✓ | ⚠ 87%† |
| 11 | qwen3-next-80b (OpenRouter · NVIDIA) | Qwen | 72.9 | 80.6 | 78.1† | ✓ | ✓ 100%† |
| 12 | gpt-oss-120b (OpenRouter · NVIDIA · SambaNova) | OpenAI | 80.8 | 80.8 | 75.0 | ✓ | ✓ 100%† |
| 13 | command-a (Cohere · GitHub) | Cohere | 52.7 | 71.2 | 73.0 | ✓ | ✓ 100% |
| 14 | gpt-4o (GitHub Models) | OpenAI | 53.6 | 74.7 | 58.32 | ⚠ capped | ⚠ 25% |
| 15 | gpt-oss-20b (OpenRouter) | OpenAI | 58.6 | 73.6 | 69.9 | ✓ | ✓ 96%† |
| 16 | phi-4 (GitHub Models) | Microsoft | 56.1 | 70.4 | no data2 | ⚠ capped | — |
| 17 | llama-3.3-70b (OpenRouter · GitHub · NVIDIA · SambaNova) | Meta | 50.5 | 68.9 | — | ✓ | — untested |
| 18 | dolphin-mistral-24b (OpenRouter) | Mistral | 45.3 | 66.3 | — | — | — untested |
| 19 | nemotron-nano-9b (OpenRouter) | NVIDIA | 64.0 | 59.4 | — | ✓ | ✓ 100%† |
| 20 | hermes-3-405b (OpenRouter) | Nous / Llama | 44.8 | 54.1 | — | — | — untested |
★ The top row is the bar, not a panel member. It shows the best public frontier score on each bench as of mid-2026 — MMLU-Pro ~89.5 (Claude Opus 4.5), GPQA ~94.3 (Gemini 3.1 Pro); different models lead each. Paid / public access only — it sits there purely to mark the ceiling the free panel is measured against, never as one of the free models.† Usable re-verified clean on the pinned harness this session — coverage = parseable ÷ answered, over the questions each model actually returned (gemma-31b 36, gemma-26b 78, gpt-oss-120b 77, gpt-oss-20b 75, nemotron-nano-30b 78, nemotron-nano-9b 41). nemotron-30b's MMLU-Pro-Mini-196 78.5 is accuracy on the 88% it answered (≈ its 78.3 public MMLU-Pro — full strength when it concludes); the 12% it repetition-collapses are excluded. Reach is the contact axis: qwen's ✗ 9/196 is exactly why it has no Usable number — unreachable, not unintelligent. On our MMLU-Pro-Mini-196, gemma-4-26b (82.1) edges the bigger gemma-4-31b (80.6) — the small 4B-active MoE is the strongest reliable free single, the best-single reference for a panel to beat. Rows are tinted by lineage. Public MMLU-Pro/GPQA are each model's own-report or most-cited value — eval methodologies vary across sources (0-shot · CoT · reasoning-mode), so the public columns are approximate; MMLU-Pro-Mini-196 is our one consistent measure, run identically across every model. Entries marked "capped / thin / no data / ⚠ low-%" hit a free-tier wall (daily caps, 429s, or 550B/675B timeouts) before a clean 196-run — the number shown is accuracy on the questions answered, with coverage flagged. Row order is smartest→dumbest by our MMLU-Pro-Mini-196 (accuracy-on-answered) where we have it, else by public MMLU-Pro mapped onto that scale (our bench runs ~3–4 pts below public-reported, so public-only and coverage-capped rows are approximate placements, not exact).
Preferred trio (when qwen is up): gemma-4-26b (82.6) · gpt-oss-120b (80.8) · qwen3-next (80.6) — three lineages within 2.0 points.
Backup trio — gemma-4-26b · gpt-oss-120b · nemotron-nano-30b, three lineages. Measured on MMLU-Pro-Mini-196: the majority vote (81.1%) loses to its best member gemma-26b (82.1%, lift −1.0) — the trio is dominated by gemma-26b, and voting a dominated panel hurts (§6.1). Yet the selection ceiling is +4.1 (the right answer sits in the panel 86.2% of the time): the diversity is real, plain voting just can't reach it — the case for synthesis over voting, measured on our own clean data.
Coordinator (for the orchestrator): nemotron-nano-30b — external, comparable, NOT smarter, a different tech tree.
Best-single reference to beat: gemma-4-31b (85.2) — reliable, but out-of-panel (shares Google's lineage with gemma-4-26b).
Every model scored across all 14 MMLU-Pro disciplines (14 questions each). No model wins every discipline — different lineages own different domains, which is the whole precondition for coordination. The two ceiling columns show the best score reachable per discipline: ceiling trio = best of the three trio members (marked ★), ceiling all = best of all seven. Both tower over every single model — that gap is the coordination headroom.
Green intensity = accuracy on that discipline. gold outline = discipline winner. ★ marks the three trio members. ceiling trio = best of the three per discipline; ceiling all = best of all seven — the bar coordination chases.
The other free models on OpenRouter — code, vision, and small edge models. Each is shown against the one benchmark that fits its specialty, with that benchmark's best-known score (any model, paid or free) and the gap — so you can see how far below the field's best each free specialist sits on its own turf. Kept out of the panel table above (a coordination panel wants general reasoning).
| Model (access) | Lineage | Specialises in | Score (relevant benchmark) | Best (any model) | Δ below |
|---|---|---|---|---|---|
| nemotron-3-nano-omni-30b (OpenRouter · NVIDIA) | NVIDIA | omni · reasoning (multimodal) | 77.3 (MMLU-Pro) | 89.8 (Gemini 3) | −12.5 |
| nemotron-nano-12b-vl (OpenRouter · NVIDIA) | NVIDIA | vision-language | ~74 (MMMU) | 86.0 (Qwen3.6) | −12 |
| llama-3.2-90b-vision (GitHub · NVIDIA) | Meta | vision-language | 60.3 (MMMU) | 86.0 (Qwen3.6) | −25.7 |
| lfm-2.5-1.2b-thinking (OpenRouter) | Liquid | on-device reasoning | 49.7 (MMLU-Pro) | 89.8 (Gemini 3) | −40.1 |
| lfm-2.5-1.2b-instruct (OpenRouter) | Liquid | on-device general | 44.4 (MMLU-Pro) | 89.8 (Gemini 3) | −45.4 |
| llama-3.2-3b (OpenRouter) | Meta | small general | 36.57 (MMLU-Pro) | 89.8 (Gemini 3) | −53.3 |
| qwen3-coder (OpenRouter) | Qwen | agentic coding | ~69.6% (SWE-bench) | 88.6 (Claude Opus 4.8)9 | −19.0 |
| north-mini-code (OpenRouter · Cohere) | Cohere | agentic coding | 80.2%8 (SWE-bench) | 88.6 (Claude Opus 4.8)9 | −8.48 |
| laguna-m.1 (OpenRouter) | Poolside | agentic coding | 74.6% (SWE-bench) | 88.6 (Claude Opus 4.8)9 | −14.0 |
| laguna-xs.2 (OpenRouter) | Poolside | agentic coding | 69.9% (SWE-bench) | 88.6 (Claude Opus 4.8)9 | −18.7 |
| nemotron-3.5-content-safety (OpenRouter · NVIDIA) | NVIDIA | safety classifier (guardrail) | — (guardrail F1) | — | n/a |
| whisper-large-v3 (Groq / HF) | OpenAI | speech-to-text (ASR) | 2.7%11 (LibriSpeech WER) | ~2% (best ASR) | ≈ best11 |
| qwen3-embedding-8b (HF) | Qwen | text embeddings (retrieval) | 70.6 (MTEB) | 74.3 (Harrier-27B) | −3.7 |
| rerank-v4.0 (Cohere) | Cohere | retrieval rerank (re-ordering) | — (BEIR nDCG) | — | n/a |
| aya-expanse-32b (Cohere / HF) | Cohere | multilingual (23 langs) | 76.6% WR12 (m-Arena-Hard) | open leader | ≈ best (open)12 |
| lyria-3-pro/clip (OpenRouter) | text → music / audio | — (subjective) | — | n/a10 |
Best · any model = current SOTA on that benchmark by ANY model, paid or free (sourced 2026: MMLU-Pro 89.8 Gemini 3 · SWE-bench Verified 88.6 Claude Opus 4.8 · MMMU 86.0 Qwen3.6 · MTEB 74.3 Harrier · ASR-WER ~2%). Only the 2 router meta-models (openrouter/free, owl-alpha) are left out — they aren't real models. Everything else free is here.
A possible production shape — illustrative, one representative model per type. A request in any language is first translated to English, then screened by two guards — Prompt Guard for injection and a content-safety model — before an omni coordinator routes it to the general panel (three models coordinated into a single answer) or the fitting specialist, then translated back into the user's own language.
Every free LLM API we probed while assembling the panel, and what its free tier actually delivers in practice. The recurring lesson: "free" rarely means "usable at volume" — daily caps, IP blocks and silent free→paid churn make the access layer the real constraint (full account in /problems).
| Provider | Free access | Signup | Our experience | Model availability |
|---|---|---|---|---|
| NVIDIA NIM | ~40 req/min · no daily wall we hit | Free NVIDIA developer account → API key | ✓ Our workhorse for strong models — reliable; 429-throttles only under sustained heavy load | Strong tier: nemotron-super & ultra-550B, qwen3-next, gpt-oss-120B, mistral-large-3, llama-3.3 |
| OpenRouter | ~50 req/day on :free models | Account → key (no card for the free tier) | ✓ Main workhorse — but pin the provider: silent free→paid churn + quant/routing variance | Largest :free catalog: gemma, gpt-oss, nemotron-nano, qwen, dolphin, hermes |
| Google AI Studio | ~20 requests/day, per model | Google account → AI Studio key | ⚠ Frontier Geminis are free, but the daily cap is too thin to finish a 196-run or feed a live panel | gemini-3-flash / 3-pro / 3.1, gemini-2.5-flash |
| Cohere | 1,000 calls/month (trial key) | Account → key | ✓ Clean + reliable; the monthly budget means bench-once, not serve-at-volume | command-a, aya-expanse (multilingual), rerank-v4, command-r |
| SambaNova | Daily-capped (~40 calls, then 429) | Account → key (no card) | ⚠ DeepSeek-V3.2 is strong, but the cap stopped us at 39/196 — can't complete a run | DeepSeek-V3.2, Llama-3.x (MiniMax is paid) |
| Z.ai (GLM) | -flash only · 429-rate-limited | Account → key | ⚠ glm-4.5-flash is strong (~82) but coverage caps ~87%; it does serve live | glm-4.5-flash (glm-4.6 / full are paid) |
| GitHub Models | Small shared daily request budget | GitHub account → personal-access token | ⚠ Capped fast (~48 calls/run, shared across models) — good for a peek, not volume | gpt-4o, phi-4, llama-3.3 + catalog |
| Groq | Broad + very fast | Account → key | ✗ Blocks data-centre IPs (403 from our server) — works from a residential line, not our host | llama-4-scout, qwen3-32B, gpt-oss-120B |
| Cerebras | ~1M tokens/day | Account → key | ✗ Cloudflare WAF (err 1010) blocks our IP, and the 8K context is too small for long prompts | gpt-oss-120B, zai-glm |
✓ usable for real work · ⚠ works but capped/limited · ✗ blocked for us · all free with no card (our hard rule); paid-only models noted but never used.