Coordination, not scale. We take AI from one capability tier to the level above — no giant training run required.
One model wins maths. Another wins law. Another sees images. Another writes code.
So we stopped picking one — we route each job to the model that's best at it, and slot in a dedicated specialist for any mode the others are weak at: vision, audio, code. A door no single model can open. The smart part isn't a bigger brain — it's the orchestration.
aus.wick.pics — the case, in six slides.
02 / 06
The opportunity
Even a couple of models each leading a couple of categories is enough: the panel’s best-of-breed coverage — the ceiling columns — tower over every single model. That gap is exactly the lift coordination reaches for.
03 / 06
The proof
The diversity is real. The lever is the open part.
The two rightmost bars are the story: an agreement-cascade — escalate only the hard, contested questions to a bigger free model — reaches a measured +4.6 (86.7), charted conservatively at +4.5 to sit a hair under Opus 4.8 rather than claim an exact dead-heat, at a fraction of the cost. It is the coordinator now powering our members’ AI. Each bar is a coordination strategy, scored against the best single model (82.1%). The selection ceiling (+7.7) shows that correct coordination could theoretically reach even Opus 4.8 — but this is far more difficult than it sounds, and likely requires synthesis. Tracing the full ladder this round sharpened how: a judge that re-solves the question reverts to its own ability, and one that merely picks among the drafts scores at chance — so it is the panel's agreement that carries the lift, not the judge's verdict, and the unclaimed headroom sits where the majority is confidently wrong. The open frontier.
04 / 06
Why it climbs
As you climb, the AI gets promoted.
Same models the whole way up. The only thing that changes is the AI's job in the coordination — and that's where the gains come from.
🛠️
Worker
Every model just answers. A dumb harness tallies the votes — no intelligence does the combining.
routing · voting · weighted vote
✍️
Editor
An AI reads everyone's drafts, finds the mistake, and writes the best final answer. It decides the answer.
critique-fusion · mixture-of-agents
🧠
Manager
An AI runs the whole show — who to ask, how many rounds, when to stop. It decides the process itself.
LLM orchestrator · trained orchestrator
05 / 06
Read it against two ladders
Capability climbs. The ruler must climb too.
A coordination lift only shows up where the benchmark still has headroom for the models being coordinated — so every result is read against both ladders: which model generation, on which benchmark generation.
Models climbed (capability) — so the ruler climbed too (benchmarks).
model generation≈ BG2 score (MMLU-Pro)
G1 · ’22–23~40% · GPT-3.5
G2 · ’23–24~60% · GPT-4, Claude 2
G3 · ’24~75% · GPT-4o, gemma/oss ◄ here
G4 · ’25~90%+ · o3, Claude 4
accuracy band on MMLU-Pro
benchmark generationheadroom
BG1 · MMLU, ARCsaturated — every model aces it
BG2 · MMLU-Pro, GPQAdiscriminating ◄ we test here
BG3 · FrontierMath, HLEunsaturated even for the best
harder benchmarks = more headroom
It's also why testing frontier coordination needs a BG3 ruler: on a saturated benchmark the best models have no room to lift — change the ruler and the headroom comes back.
06 / 06
The open frontier
The exciting part is what's not done yet.
The +8.8 selection ceiling proves the prize is real (above the frontier). These are the open questions for capturing it — where the unclaimed wins are.
🧗
Beat the ceiling
Voting can only pick the best answer that already exists. Fusion that produces a correct answer no model gave — manufacturing correctness — is barely solved. The real level-up.
🔓
Unlock new powers
Compose specialists so the system does what no single model can — see an image, run code — zero → working. Categorical, not a few percent.
🏔️
Frontier, done right
On hard-enough benchmarks and across rival labs (different blind spots), does the lift hold? Newly concrete: a single free model (nemotron-ultra-550B) now benches just under our panel's ceiling (87 vs 90.8) — so the real test is coordinating models at that level, not lifting a weaker trio toward it.
🤖
An AI that only conducts
A model trained for nothing but orchestrating a panel of others. The moonshot at the top of the ladder.
The old way up a tier: train a bigger frontier model — $$$$. Our way: coordinate a diverse panel + the orchestration stack — ¢. A sovereign path up the capability ladder, without the frontier price tag.