The Model Table

Every free model × every public benchmark × every internal bench we have ever run them through. Ranked by MMLU-Pro. Caveated results are kept — but flagged, never silently compared.
Green Wick AI Research · Composite Intelligence · rebuilt as runs land

This is the at-a-glance truth view used to pick a coordination panel. A model needs three things to make the panel: reliable on two axes — we can reach it (not throttled to death) and its replies are usable (they parse, not truncate) — comparable (within a few points of its peers), and diverse-lineage (a different tech tree, so errors decorrelate; rows are tinted by lineage so you can spot it — note that Gemini and Gemma share Google's lineage and are not independent for panel purposes). The columns let you see all of it at once — and where the gaps still are.

The only membership rule is $0 to call — not which gateway. The access path is noted in small print under each model name (OpenRouter, Google AI Studio, …) so the numbers are reproducible on the same endpoints — and where a model has several free endpoints, the one in bold is the path we benched and reliability-checked on (the most reliable); we add any model that is freely accessible by some path, not just via OpenRouter. The public benchmarks (MMLU-Pro, GPQA) are third-party scores; MMLU-Pro-Mini-196 is our own custom stratified benchmark, run identically across every model.

reliable ≥90% borderline 80–90% unreliable <80% not measuredpending run in flightorange + ¹ = flagged, see notes Reach = contactable under free limits (✓ responds · ⚠ slow · ✗ blocked)Usable = % of replies that parse cleanly
Current frontier models — comparison only · public access, not free
#Model
(generation)
LineageGPQA
public
MMLU-Pro
public
MMLU-Pro-Mini-196
our custom bench
Reach
responds?
Usable
reply parses
Claude Fable 5
(Gen 5 · flagship)
Anthropic92.691.5
GPT-5.5
(Gen 6)
OpenAI95.8
Gemini 3.1 Pro
(Gen 5)
Google94.191.0
MiniMax M3
(Gen 5)
MiniMax92.9
Claude Opus 4.8
(Gen 5)
Anthropic89.586.2
Claude Sonnet 5
(Gen 5)
Anthropic80.6
GLM-5.2
(Gen 5)
Zhipu / GLM
Grok 4.3
(Gen 5)
xAI~88~91
DeepSeek V4
(Gen 5)
DeepSeek
Free & open-source models — used for research · our coordination-panel candidates
#Model (access)LineageGPQA publicMMLU-Pro publicMMLU-Pro-Mini-196 oursReachUsable
1nemotron-ultra-550b
(OpenRouter · NVIDIA)
NVIDIA6686.8⚠ 96%
2gemini-3-flash
(AI Studio)
Google90.489.020/day cap⚠ thin
3gemma-4-26b
(OpenRouter)
Google82.382.682.1✓ 100%
4glm-4.5-flash
(Z.ai)
Zhipu / GLM6681.8⚠ 92%2
5nemotron-super-120b
(OpenRouter · NVIDIA)
NVIDIA6681.4✓ 96%
6deepseek-v3.2
(SambaNova)
DeepSeek82.485.0capped 39/1962
7gemma-4-31b
(OpenRouter · SambaNova)
Google84.385.280.6✓ 100%
8gemini-2.5-flash
(AI Studio)
Google82.86⚠ thin
9mistral-large-3-675b
(NVIDIA)
Mistral~4478.078.7⚠ 38%
10nemotron-nano-30b
(OpenRouter · NVIDIA)
NVIDIA73.078.378.5⚠ 87%
11qwen3-next-80b
(OpenRouter · NVIDIA)
Qwen72.980.678.1✓ 100%
12gpt-oss-120b
(OpenRouter · NVIDIA · SambaNova)
OpenAI80.880.875.0✓ 100%
13command-a
(Cohere · GitHub)
Cohere52.771.273.0✓ 100%
14gpt-4o
(GitHub Models)
OpenAI53.674.758.32⚠ capped⚠ 25%
15gpt-oss-20b
(OpenRouter)
OpenAI58.673.669.9✓ 96%
16phi-4
(GitHub Models)
Microsoft56.170.4no data2⚠ capped
17llama-3.3-70b
(OpenRouter · GitHub · NVIDIA · SambaNova)
Meta50.568.9— untested
18dolphin-mistral-24b
(OpenRouter)
Mistral45.366.3— untested
19nemotron-nano-9b
(OpenRouter)
NVIDIA64.059.4✓ 100%
20hermes-3-405b
(OpenRouter)
Nous / Llama44.854.1— untested

★ The top row is the bar, not a panel member. It shows the best public frontier score on each bench as of mid-2026 — MMLU-Pro ~89.5 (Claude Opus 4.5), GPQA ~94.3 (Gemini 3.1 Pro); different models lead each. Paid / public access only — it sits there purely to mark the ceiling the free panel is measured against, never as one of the free models.† Usable re-verified clean on the pinned harness this session — coverage = parseable ÷ answered, over the questions each model actually returned (gemma-31b 36, gemma-26b 78, gpt-oss-120b 77, gpt-oss-20b 75, nemotron-nano-30b 78, nemotron-nano-9b 41). nemotron-30b's MMLU-Pro-Mini-196 78.5 is accuracy on the 88% it answered (≈ its 78.3 public MMLU-Pro — full strength when it concludes); the 12% it repetition-collapses are excluded. Reach is the contact axis: qwen's ✗ 9/196 is exactly why it has no Usable number — unreachable, not unintelligent. On our MMLU-Pro-Mini-196, gemma-4-26b (82.1) edges the bigger gemma-4-31b (80.6) — the small 4B-active MoE is the strongest reliable free single, the best-single reference for a panel to beat. Rows are tinted by lineage. Public MMLU-Pro/GPQA are each model's own-report or most-cited value — eval methodologies vary across sources (0-shot · CoT · reasoning-mode), so the public columns are approximate; MMLU-Pro-Mini-196 is our one consistent measure, run identically across every model. Entries marked "capped / thin / no data / ⚠ low-%" hit a free-tier wall (daily caps, 429s, or 550B/675B timeouts) before a clean 196-run — the number shown is accuracy on the questions answered, with coverage flagged. Row order is smartest→dumbest by our MMLU-Pro-Mini-196 (accuracy-on-answered) where we have it, else by public MMLU-Pro mapped onto that scale (our bench runs ~3–4 pts below public-reported, so public-only and coverage-capped rows are approximate placements, not exact).

Table notes — why a flagged result can't be fairly compared

  1. 1 · Truncation artifact. A verbose reasoner ran out of token budget before writing "ANSWER: X"; the strict parser then scored the cut-off answer as wrong. nemotron-super's real intelligence is ~82–87% (its clean MMLU-Pro-Mini-196 = 82.1); the 58.7 / 59.2 / 62.8 values are measurement failure, not capability. Fixed by strict-parse + re-ask + the reliability/intelligence split.
  2. 2 · Rate-limit-invalidated. 429s mid-run left most answers empty → scored as wrong (gemma 33.2, nemotron-super 20.9, …). Not real accuracy — the source file is literally named ..._INVALID_ratelimited.
  3. 3 · Small-N pilot. N=30 (domain-routing, 6 disciplines × 5) and N=32 (the "crossover" hard slice) — 95% CI ≈ ±10–13 pts, easy/hard-skewed subsets → directional only, not comparable to the stratified 196.
  4. 4 · Sub-sample. Some fusion runs scored only the N=168 disagreement subset (nemotron-super 80.2 / gemma 78.6 / gpt-oss-120b 73.2) — a different, smaller sample than the full 196.
  5. 5 · Run-to-run reproducibility variance. Same model + same frozen 196-Q sample, different run → different score (the ±7 swing: gemma 80.6 vs 73.0 vs 66.3; gpt-oss-120b 75.0 vs 74.0). Pre-provider-pinning; the new pinned harness is the fix.
  6. 6 · Unverified public score. NVIDIA's nemotron-super/ultra MMLU-Pro/GPQA could not be confirmed against a primary source (tech-report PDF unparseable, build.nvidia digits unconfirmed) → treated as UNKNOWN, not guessed.

What the gaps say

The gateway models are now benchedThe comparable band spans many lineages now. The backup trio has clean MMLU-Pro-Mini-196 numbers — gemma-26b 82.1 · gpt-oss-120b 75.0 · nemotron-30b 78.5 (the measured panel result is below). The other-gateway rows have now been run on our harness: glm-4.5-flash (81.8 on the 92% it answered) and nemotron-ultra-550b (90.9 on the 84% it answered) landed real numbers, while several hit free-tier walls — DeepSeek-V3.2 capped at 39/196, Mistral-Large-3 at 38% coverage, GPT-4o and phi-4 throttled out, gemini-3-flash at 20 calls/day (all flagged in the table). qwen3-next stays rate-limited (9/196).
nemotron-super — the cautionary tale, partly rehabilitatedRanks high on raw intelligence (81.4), but it earned its reputation for reliability — a score that swung 20.9 → 82.1 across early runs, mostly the old OpenRouter route plus pre-strict-parse truncation. Re-benched clean this session on NVIDIA-direct it answered 96% (✓), so the table now reads as reliable — but the rank reflects intelligence, not shippability: re-verify coverage at scale before trusting it in a live panel, because the historical swing was real. Capability you can't reliably access isn't capability you can ship.

The panel this picks

Preferred trio (when qwen is up): gemma-4-26b (82.6) · gpt-oss-120b (80.8) · qwen3-next (80.6) — three lineages within 2.0 points.
Backup trio — gemma-4-26b · gpt-oss-120b · nemotron-nano-30b, three lineages. Measured on MMLU-Pro-Mini-196: the majority vote (81.1%) loses to its best member gemma-26b (82.1%, lift −1.0) — the trio is dominated by gemma-26b, and voting a dominated panel hurts (§6.1). Yet the selection ceiling is +4.1 (the right answer sits in the panel 86.2% of the time): the diversity is real, plain voting just can't reach it — the case for synthesis over voting, measured on our own clean data.
Coordinator (for the orchestrator): nemotron-nano-30b — external, comparable, NOT smarter, a different tech tree.
Best-single reference to beat: gemma-4-31b (85.2) — reliable, but out-of-panel (shares Google's lineage with gemma-4-26b).

Where each model wins — and the best trio

Every model scored across all 14 MMLU-Pro disciplines (14 questions each). No model wins every discipline — different lineages own different domains, which is the whole precondition for coordination. The two ceiling columns show the best score reachable per discipline: ceiling trio = best of the three trio members (marked ★), ceiling all = best of all seven. Both tower over every single model — that gap is the coordination headroom.

Best complementary trioglm-4.5 · gemma-4-26b · nemotron-super-120b — the 2nd/3rd/4th strongest models, three distinct lineages whose errors decorrelate. Ceiling trio 88.7% (best of the three per discipline); ceiling all 90.4% across all seven — each above the best single model. The win-row markers show why: each member owns disciplines the others miss. Measured this session (MMLU-Pro-Mini-196): the trio's plain majority vote scores 84.2% — +2.0 over the best single member — the program's first selection-tier WIN (the looser backup trio above lost −1.0; clustering the panel in strength is what flips selection positive). The +7.7 selection ceiling stays the open frontier; a reconciling LLM judge did NOT beat the plain vote here — a tight, diverse panel's agreement is a stronger signal than any single free judge's verdict (full write-up: /research §5.1).
No model wins every discipline Accuracy by MMLU-Pro discipline. Winner outlined gold; ★ = best-trio member. nemo-ultra ★glm-4.5 ★gemma-26b ★nemo-super gemma-31b qwen3-next gpt-oss ceiling trio ceiling all biology 93 100 93 93 93 93 86 100 100 business 100 100 93 93 93 93 86 100 100 chemistry 100 100 100 100 100 86 100 100 100 computer science 93 93 93 100 93 93 93 100 100 economics 93 86 79 86 79 86 64 86 93 engineering 100 100 93 71 86 71 64 100 100 health 79 77 50 64 57 79 71 77 79 history 57 50 57 43 50 43 50 57 57 law 79 57 79 71 86 57 50 79 86 math 100 100 100 100 100 100 100 100 100 other 50 64 50 50 36 57 43 64 64 philosophy 93 86 79 86 79 71 64 86 93 physics 100 100 100 93 93 79 86 100 100 psychology 86 71 86 93 86 86 93 93 93 mini-bench · all 196 86.8 81.8 82.1 81.6 80.6 78.1 75.0 88.7 90.4 nemo-ultra tops 9 · glm-4.5 tops 7 · nemo-super tops 4 · gemma-26b tops 4 · gemma-31b tops 3 · gpt-oss tops 3 · qwen3-next tops 2. Different models own different domains — the precondition for coordination. Ceiling trio = best of the three trio models per discipline; ceiling all = best of all seven. Each towers over any single model — that gap is the coordination headroom.

Green intensity = accuracy on that discipline. gold outline = discipline winner. marks the three trio members. ceiling trio = best of the three per discipline; ceiling all = best of all seven — the bar coordination chases.

Specialised models — not panel candidates

The other free models on OpenRouter — code, vision, and small edge models. Each is shown against the one benchmark that fits its specialty, with that benchmark's best-known score (any model, paid or free) and the gap — so you can see how far below the field's best each free specialist sits on its own turf. Kept out of the panel table above (a coordination panel wants general reasoning).

Model
(access)
LineageSpecialises inScore
(relevant benchmark)
Best
(any model)
Δ below
nemotron-3-nano-omni-30b
(OpenRouter · NVIDIA)
NVIDIAomni · reasoning
(multimodal)
77.3
(MMLU-Pro)
89.8
(Gemini 3)
−12.5
nemotron-nano-12b-vl
(OpenRouter · NVIDIA)
NVIDIAvision-language~74
(MMMU)
86.0
(Qwen3.6)
−12
llama-3.2-90b-vision
(GitHub · NVIDIA)
Metavision-language60.3
(MMMU)
86.0
(Qwen3.6)
−25.7
lfm-2.5-1.2b-thinking
(OpenRouter)
Liquidon-device reasoning49.7
(MMLU-Pro)
89.8
(Gemini 3)
−40.1
lfm-2.5-1.2b-instruct
(OpenRouter)
Liquidon-device general44.4
(MMLU-Pro)
89.8
(Gemini 3)
−45.4
llama-3.2-3b
(OpenRouter)
Metasmall general36.57
(MMLU-Pro)
89.8
(Gemini 3)
−53.3
qwen3-coder
(OpenRouter)
Qwenagentic coding~69.6%
(SWE-bench)
88.6
(Claude Opus 4.8)9
−19.0
north-mini-code
(OpenRouter · Cohere)
Cohereagentic coding80.2%8
(SWE-bench)
88.6
(Claude Opus 4.8)9
−8.48
laguna-m.1
(OpenRouter)
Poolsideagentic coding74.6%
(SWE-bench)
88.6
(Claude Opus 4.8)9
−14.0
laguna-xs.2
(OpenRouter)
Poolsideagentic coding69.9%
(SWE-bench)
88.6
(Claude Opus 4.8)9
−18.7
nemotron-3.5-content-safety
(OpenRouter · NVIDIA)
NVIDIAsafety classifier
(guardrail)
— (guardrail F1)n/a
whisper-large-v3
(Groq / HF)
OpenAIspeech-to-text
(ASR)
2.7%11
(LibriSpeech WER)
~2%
(best ASR)
≈ best11
qwen3-embedding-8b
(HF)
Qwentext embeddings
(retrieval)
70.6
(MTEB)
74.3
(Harrier-27B)
−3.7
rerank-v4.0
(Cohere)
Cohereretrieval rerank
(re-ordering)
— (BEIR nDCG)n/a
aya-expanse-32b
(Cohere / HF)
Coheremultilingual
(23 langs)
76.6% WR12
(m-Arena-Hard)
open leader≈ best
(open)12
lyria-3-pro/clip
(OpenRouter)
Googletext → music / audio— (subjective)n/a10

Best · any model = current SOTA on that benchmark by ANY model, paid or free (sourced 2026: MMLU-Pro 89.8 Gemini 3 · SWE-bench Verified 88.6 Claude Opus 4.8 · MMMU 86.0 Qwen3.6 · MTEB 74.3 Harrier · ASR-WER ~2%). Only the 2 router meta-models (openrouter/free, owl-alpha) are left out — they aren't real models. Everything else free is here.

  1. 7 · Method-divergent. llama-3.2-3b's MMLU-Pro reports as 36.5 (HuggingFace community eval) vs ~20 (lower-shot) — shown as 36.5.
  2. 11 · Lower is better. WER (word error rate) is an error metric, so lower = better — Whisper-large-v3 is the leading open ASR model (~2.7% on clean LibriSpeech, 8–12% on real-world audio), a hair off the best specialised systems.
  3. 12 · Relative metric. Aya Expanse 32B is the leading open multilingual model — it beats Llama-3.1-70B (twice its size) on non-English evals; m-Arena-Hard is a head-to-head win-rate across 23 languages, not a vs-SOTA absolute.
  4. 8 · Benchmark-variant. north-mini-code's 80.2% is SWE-bench Verified pass@10; the SOTA + the other coders are best-effort single-attempt — gap is approximate.
  5. 9 · Vendor harness. the 88.6% SWE-bench Verified figure (Claude Opus 4.8) is a vendor agentic-harness number, not a neutral pass@1 — so the code-model gaps are config-approximate. (Fable 5 / Mythos 5 score higher ~95% but aren't openly accessible, so the best accessible model anchors the gap.)
  6. 10 · Not truly free. the Lyria music models carry $0 token pricing but bill per generated clip (~$0.04–0.08), and audio quality has no standard benchmark.

The full stack, assembled

A possible production shape — illustrative, one representative model per type. A request in any language is first translated to English, then screened by two guards — Prompt Guard for injection and a content-safety model — before an omni coordinator routes it to the general panel (three models coordinated into a single answer) or the fitting specialist, then translated back into the user's own language.

INPUT any language OUTPUT user's language TRANSLATOR aya-expanse-32b (via Cohere) non-English English SAFETY GUARDS 1 · Prompt Guard injection · (HF) 2 · nemotron-3.5 content · (NVIDIA) gate in & out OMNI COORDINATOR nemotron-3-nano-omni-30b (via NVIDIA) routes each request ↺ fallback no specialist fits → omni handles it GENERAL PANEL · 3 models acting as 1 gemma-4-26b Google (OpenRouter) glm-4.5 Zhipu (Z.ai) nemotron-super-120b NVIDIA (NVIDIA) coordinate → vote / synthesize → one answer diverse lineages, errors decorrelate SPECIALISTS · one of each, on demand visionnemotron-nano-12b-vlNVIDIA codeqwen3-coderOpenRouter audiolyria-3-proOpenRouter speech→textwhisper-large-v3HF embeddingsqwen3-embedding-8bHF

The free-API landscape — every provider we tried

Every free LLM API we probed while assembling the panel, and what its free tier actually delivers in practice. The recurring lesson: "free" rarely means "usable at volume" — daily caps, IP blocks and silent free→paid churn make the access layer the real constraint (full account in /problems).

ProviderFree accessSignupOur experienceModel availability
NVIDIA NIM~40 req/min · no daily wall we hitFree NVIDIA developer account → API key✓ Our workhorse for strong models — reliable; 429-throttles only under sustained heavy loadStrong tier: nemotron-super & ultra-550B, qwen3-next, gpt-oss-120B, mistral-large-3, llama-3.3
OpenRouter~50 req/day on :free modelsAccount → key (no card for the free tier)✓ Main workhorse — but pin the provider: silent free→paid churn + quant/routing varianceLargest :free catalog: gemma, gpt-oss, nemotron-nano, qwen, dolphin, hermes
Google AI Studio~20 requests/day, per modelGoogle account → AI Studio key⚠ Frontier Geminis are free, but the daily cap is too thin to finish a 196-run or feed a live panelgemini-3-flash / 3-pro / 3.1, gemini-2.5-flash
Cohere1,000 calls/month (trial key)Account → key✓ Clean + reliable; the monthly budget means bench-once, not serve-at-volumecommand-a, aya-expanse (multilingual), rerank-v4, command-r
SambaNovaDaily-capped (~40 calls, then 429)Account → key (no card)⚠ DeepSeek-V3.2 is strong, but the cap stopped us at 39/196 — can't complete a runDeepSeek-V3.2, Llama-3.x (MiniMax is paid)
Z.ai (GLM)-flash only · 429-rate-limitedAccount → key⚠ glm-4.5-flash is strong (~82) but coverage caps ~87%; it does serve liveglm-4.5-flash (glm-4.6 / full are paid)
GitHub ModelsSmall shared daily request budgetGitHub account → personal-access token⚠ Capped fast (~48 calls/run, shared across models) — good for a peek, not volumegpt-4o, phi-4, llama-3.3 + catalog
GroqBroad + very fastAccount → key✗ Blocks data-centre IPs (403 from our server) — works from a residential line, not our hostllama-4-scout, qwen3-32B, gpt-oss-120B
Cerebras~1M tokens/dayAccount → key✗ Cloudflare WAF (err 1010) blocks our IP, and the 8K context is too small for long promptsgpt-oss-120B, zai-glm

usable for real work · works but capped/limited · blocked for us · all free with no card (our hard rule); paid-only models noted but never used.