01 / 06
The idea

Go up a tier.
Without a bigger model.

Coordination, not scale. We take AI from one capability tier to the level above — no giant training run required.
One model wins maths. Another wins law. Another sees images. Another writes code.
So we stopped picking one — we route each job to the model that's best at it, and
slot in a dedicated specialist for any mode the others are weak at:
vision, audio, code. A door no single model can open.
The smart part isn't a bigger brain — it's the orchestration.
aus.wick.pics — the case, in six slides.
02 / 06
The opportunity
No model wins every discipline Accuracy by MMLU-Pro discipline. Winner outlined gold; ★ = best-trio member. nemo-ultra ★glm-4.5 ★gemma-26b ★nemo-super gemma-31b qwen3-next gpt-oss ceiling trio ceiling all biology 93 100 93 93 93 93 86 100 100 business 100 100 93 93 93 93 86 100 100 chemistry 100 100 100 100 100 86 100 100 100 computer science 93 93 93 100 93 93 93 100 100 economics 93 86 79 86 79 86 64 86 93 engineering 100 100 93 71 86 71 64 100 100 health 79 77 50 64 57 79 71 77 79 history 57 50 57 43 50 43 50 57 57 law 79 57 79 71 86 57 50 79 86 math 100 100 100 100 100 100 100 100 100 other 50 64 50 50 36 57 43 64 64 philosophy 93 86 79 86 79 71 64 86 93 physics 100 100 100 93 93 79 86 100 100 psychology 86 71 86 93 86 86 93 93 93 mini-bench · all 196 86.8 81.8 82.1 81.6 80.6 78.1 75.0 88.7 90.4 nemo-ultra tops 9 · glm-4.5 tops 7 · nemo-super tops 4 · gemma-26b tops 4 · gemma-31b tops 3 · gpt-oss tops 3 · qwen3-next tops 2. Different models own different domains — the precondition for coordination. Ceiling trio = best of the three trio models per discipline; ceiling all = best of all seven. Each towers over any single model — that gap is the coordination headroom.
Even a couple of models each leading a couple of categories is enough: the panel’s best-of-breed coverage — the ceiling columns — tower over every single model. That gap is exactly the lift coordination reaches for.
03 / 06
The proof

The diversity is real.
The lever is the open part.

measured loss measured lift ceiling (headroom) projected (?) frontier target selection ceiling 89.8 this lift = a 100% benchmark score (no theoretical upper limit for synthesis) best single model (gemma-4-26b, 82.1%) = 0 +0.6 category routing +2.0 plain vote +2.0 weighted vote +7.7 selection CEILING −1.6 critique fusion +0.5 mixture- of-agents +17.9 synthesis CEILING −8.2 panel-IQ LLM orchestrator −3.1 low IQ LLM orchestrator −0.6 unanimity + judge +4.5 cascade → bigger model +4.6 Opus 4.8 TARGET TIER 1 · SELECTION (mechanical harness) TIER 2 · SYNTHESIS (LLM aggregator) TIER 3 · ORCHESTRATION (LLM runs it) New result — a tighter ★trio (glm-4.5 · gemma-4-26b · nemotron-super, all within ~1pt) flips the selection tier POSITIVE: plain and weighted vote both +2.0, routing +0.6 — reversing the earlier looser trio, where every selection method LOST. The lift is small (95% CI still includes 0) but the direction is the signal. Selection ceiling +7.7 stays mostly unclaimed. Three more aggregators on this trio now measured — every one at or below the trio’s plain vote (84.2), which is the honest headline: mixture-of-agents (one shared-revision round → vote) = 83.2, +0.5 vs its run’s best single (glm 82.7; 95% CI −3.6..+4.6); critique-fusion (an editor rewrites one answer from all three drafts) = 81.1, −1.6 vs that best single (95% CI −6.1..+3.1) and −3.1 below the trio’s plain vote (84.2); unanimity-else-judge (pass through where the panel agrees, adjudicate only where it splits) = 82.1, −0.6 vs best — and it shows why the ladder caps here: its 146 unanimous items score 95.9%, its 50 contested items only 42.0%, and no adjudicator recovers that contested block. DFPE-style reliability-weighted voting (per-category leave-one-out weights) tops out at 84.7, only +0.5 over the plain vote and only at an in-sample-tuned weight — a trio this tight has almost no reliability spread to exploit. Agreement carries the lift; every richer aggregation pays a tax. First Tier-3 rung measured — honestly negative: a bounded LLM orchestrator (nemotron-nano coordinating the ★trio per-question across solo/vote/consult/debate/fuse, N=196) = 79.6, −3.1 vs its run’s best single (glm 82.7; 95% CI −7.7..+2.0, includes 0) — below even its own run’s plain vote (82.7). Its move mix (82 solo, 79 vote, 28 fuse, 5 consult, 2 debate) shows the manager mostly delegating down the ladder and still paying a coordination tax — as the ladder predicts: agreement carries lift, adjudication doesn’t. Second Tier-3 conductor — the smarter-manager hypothesis tested and rejected: swapping the low-IQ conductor for a panel-level, diverse-lineage one (qwen3-next-80b, ~75% solo here) made it WORSE — 74.5, −8.2 vs its run’s best single (glm 82.7; 95% CI −14.3..−2.0, excludes 0 — the ladder’s first statistically significant loss). The competent manager intervened far more (92 fuse, 81 consult, 22 vote, 1 solo vs nano’s delegate-heavy mix) and paid the coordination tax on every intervention; its bar is clipped at the chart floor. THE headline win, and what our members’ AI now runs: an agreement-cascade — keep the trio’s answer where all three agree, escalate only the ~26% of contested items to a bigger free model (nemotron-ultra-550b) — reaches 86.7, a measured +4.6 over the best member (95% CI [+0.5, +8.7], the ladder’s first lift whose interval EXCLUDES 0). That is a dead-heat with the Opus 4.8 reference (also +4.6), so we chart it conservatively at +4.5 — a hair under the frontier rather than claiming an exact tie on a 196-item bench. The first free coordinator to arrive at the frontier’s doorstep, at a fraction of the cost. Synthesis’s ceiling is perfect play (+17.9 to 100%) — the open frontier. Dashed purple bars are ceilings (headroom), not measurements.
The two rightmost bars are the story: an agreement-cascade — escalate only the hard, contested questions to a bigger free model — reaches a measured +4.6 (86.7), charted conservatively at +4.5 to sit a hair under Opus 4.8 rather than claim an exact dead-heat, at a fraction of the cost. It is the coordinator now powering our members’ AI. Each bar is a coordination strategy, scored against the best single model (82.1%). The selection ceiling (+7.7) shows that correct coordination could theoretically reach even Opus 4.8 — but this is far more difficult than it sounds, and likely requires synthesis. Tracing the full ladder this round sharpened how: a judge that re-solves the question reverts to its own ability, and one that merely picks among the drafts scores at chance — so it is the panel's agreement that carries the lift, not the judge's verdict, and the unclaimed headroom sits where the majority is confidently wrong. The open frontier.
04 / 06
Why it climbs

As you climb, the AI
gets promoted.

Same models the whole way up. The only thing that changes is the AI's job in the coordination — and that's where the gains come from.
🛠️
Worker
Every model just answers. A dumb harness tallies the votes — no intelligence does the combining.
routing · voting · weighted vote
✍️
Editor
An AI reads everyone's drafts, finds the mistake, and writes the best final answer. It decides the answer.
critique-fusion · mixture-of-agents
🧠
Manager
An AI runs the whole show — who to ask, how many rounds, when to stop. It decides the process itself.
LLM orchestrator · trained orchestrator
05 / 06
Read it against two ladders

Capability climbs.
The ruler must climb too.

A coordination lift only shows up where the benchmark still has headroom for the models being coordinated — so every result is read against both ladders: which model generation, on which benchmark generation.
Models climbed (capability) — so the ruler climbed too (benchmarks).
model generation≈ BG2 score (MMLU-Pro)
G1 · ’22–23~40% · GPT-3.5
G2 · ’23–24~60% · GPT-4, Claude 2
G3 · ’24~75% · GPT-4o, gemma/oss ◄ here
G4 · ’25~90%+ · o3, Claude 4
accuracy band on MMLU-Pro
benchmark generationheadroom
BG1 · MMLU, ARCsaturated — every model aces it
BG2 · MMLU-Pro, GPQAdiscriminating ◄ we test here
BG3 · FrontierMath, HLEunsaturated even for the best
harder benchmarks = more headroom
It's also why testing frontier coordination needs a BG3 ruler: on a saturated benchmark the best models have no room to lift — change the ruler and the headroom comes back.
06 / 06
The open frontier

The exciting part is
what's not done yet.

The +8.8 selection ceiling proves the prize is real (above the frontier). These are the open questions for capturing it — where the unclaimed wins are.
🧗
Beat the ceiling
Voting can only pick the best answer that already exists. Fusion that produces a correct answer no model gave — manufacturing correctness — is barely solved. The real level-up.
🔓
Unlock new powers
Compose specialists so the system does what no single model can — see an image, run code — zero → working. Categorical, not a few percent.
🏔️
Frontier, done right
On hard-enough benchmarks and across rival labs (different blind spots), does the lift hold? Newly concrete: a single free model (nemotron-ultra-550B) now benches just under our panel's ceiling (87 vs 90.8) — so the real test is coordinating models at that level, not lifting a weaker trio toward it.
🤖
An AI that only conducts
A model trained for nothing but orchestrating a panel of others. The moonshot at the top of the ladder.
The old way up a tier: train a bigger frontier model — $$$$. Our way: coordinate a diverse panel + the orchestration stack — ¢. A sovereign path up the capability ladder, without the frontier price tag.
See the evidence Read the paper →
Composite Intelligence · proof: aus.wick.pics/research · every experiment reproducible, every caveat in the paper.