This is the proof made tangible. A coordinator decomposes your task, routes the pieces across a panel of open specialist models, has them debate and verify, and synthesises one answer — open models lifted to near-frontier. Swap the panel for frontier models and the same architecture is built to go beyond frontier. Try it, see how it's measured, and read what funding unlocks.
Behind this single chat box sits a society of specialist models under one coordinator. You ask once; the coordinator plans, routes sub-tasks to the best-suited specialist, runs a round of cross-checking, and returns one synthesised answer. The interface is deliberately simple — the orchestration is the product.
No data is sent yet — this is a UI preview. The live orchestration endpoint is on the build queue.
The whole claim — "orchestration lifts a tier" — only matters if it survives a scoreboard. We measure three groups on the same public benchmarks: the individual open component models that make up our panel, the frontier models we're measured against, and our orchestrated system.
On MMLU-Pro, scored against the objective answer key — no model judging another — a tighter, complementary trio (glm-4.5 · gemma-4-26b · nemotron-super, three lineages clustered within ~1 point) now makes simple majority voting beat the best single model by +1.6 (84%) — the program's first selection-tier win (an earlier, looser trio lost −1.0; clustering the panel in strength is what flips it positive). The selection ceiling (the correct answer present somewhere in the panel) is +7.7 above the best single — most of that headroom is still unclaimed. Notably, a reconciling LLM judge did NOT beat the plain vote here: a strong judge reverts to its own answer (self-preference), a weak external judge is helped by the drafts but still trails — so for this panel the diversity's agreement is a stronger signal than any single free judge's verdict, and reaching the ceiling via synthesis is the open frontier. Full method, every number, the negative results, and the LLM-as-judge literature are in the research paper →. The expandable per-benchmark harness below (more models, more evals, live) is being stood up; until it's wired we show "soon" rather than numbers we haven't actually run.
| Model / system | General reasoning | Hard QA | Math | Code | Aggregate |
|---|---|---|---|---|---|
| Component model A open · single | soon | soon | soon | soon | soon |
| Component model B open · single | soon | soon | soon | soon | soon |
| Component model C open · single | soon | soon | soon | soon | soon |
| Frontier reference 1 closed · reference | soon | soon | soon | soon | soon |
| Frontier reference 2 closed · reference | soon | soon | soon | soon | soon |
| ★ Orchestrated system coordinator + open panel | soon | soon | soon | soon | soon |
| ★ Orchestrated system coordinator + frontier panel — "beyond frontier" | soon | soon | soon | soon | soon |
The read we're after: the orchestrated row over an open panel should sit level with the frontier-reference rows while every individual component sits below them — that's the "lift a tier" result. The orchestrated row over a frontier panel is the "beyond frontier" hypothesis. Column benchmarks shown generically (reasoning / hard-QA / math / code) until the displayed suite is finalised — see the worklist below.
The prototype is built to run at near-zero cost on open, free-tier models — that's the point, it proves the method is cheap to verify. But the same architecture has a steep, fundable improvement curve. Each lever below is a concrete place where capital converts directly into capability. For a government, defence partner or sovereign-AI investor, this is what the spend buys.
Today the coordinator is hard-coded — a fixed decompose → route → verify → aggregate pipeline. Replacing it with a strong LLM acting as coordinator is the single largest jump available: it plans adaptively, recognises when a sub-task needs a different specialist, decides when to debate further versus commit, and writes a better final synthesis. The hard-coded version proves the method on a budget; the LLM coordinator is where it becomes genuinely frontier-grade.
Funding unlocks → per-query inference budget for a capable coordinator model on every request.The current panel is open, non-frontier models — chosen so the lift is clean to demonstrate. Level the panel up so the orchestrated specialists are themselves frontier-class, and the ceiling moves: a coordinated society of frontier models is built to exceed any one of them on its own. This is the "beyond frontier" path — the same architecture, far stronger parts.
Funding unlocks → access to top-tier model APIs / weights as the orchestrated panel.Quality scales with how many rounds of debate, verification and best-of-N sampling the system can afford. Today that's throttled by free-tier rate limits. More compute means more independent attempts, more cross-checking, and longer multi-round reasoning before the system commits — directly lifting reliability on the hardest problems.
Funding unlocks → dedicated inference compute for wider, deeper deliberation.An off-the-shelf model coordinating is good; a model trained on the orchestration task itself — learning which specialist to trust for which problem, how to phrase sub-tasks, when to stop debating — is better and cheaper per query. This is the bridge from the hard-coded coordinator to a learned one, and it compounds every other lever.
Funding unlocks → data collection + training runs for a purpose-built coordinator.The strength of a society of models comes from diversity — different architectures, training data and failure modes that don't all break the same way. Expanding the panel and adding domain specialists (code, maths, retrieval, vision, legal, defence) widens coverage and sharpens routing, so the coordinator always has a strong specialist on hand.
Funding unlocks → a larger, more diverse, sovereign-hosted expert panel.A single "multimodal" model is forced to be a generalist: it spreads one set of weights across text, vision, audio and every domain, and it is precisely in the niche modes — specialist vision, audio, OCR, rare languages, narrow technical fields — where even the best monolithic models go soft. Orchestration removes that compromise. Because the coordinator only routes, each sub-task goes to the single strongest model for that mode — a dedicated vision model, a dedicated audio model, a code specialist — so the system fields the frontier of every modality at once, which no one model can. It also unlocks capability no single model has at all: a text-only panel cannot see an image, but a coordinator that can call a vision specialist can — there the gain is not a few points, it is an entire capability switched on.
Funding unlocks → a roster of best-in-class specialist models, one per modality and niche, under the coordinator.Capability you can't measure isn't fundable. A serious eval harness — held-out suites, contamination controls, frontier-grade judging, red-teaming — turns the benchmark frame above into defensible, repeatable evidence. Paired with the ability to run the whole stack on sovereign infrastructure, it's what lets a government or defence partner actually depend on the system.
Funding unlocks → eval engineering + on-shore, air-gappable deployment.The cheapest capability gains often come from method, not spend. A small in-house research function continuously mines the open literature and our own results for low-cost wins — smarter routing, open-source integration, orchestration improvements, better coordination methods — so the system keeps getting better between funding rounds instead of being built once and left to decay. This is what makes it a sustainable sovereign lab, not a one-off prototype.
Funding unlocks → an ongoing R&D loop that compounds every other lever at minimal extra cost.If you took everything the prototype demonstrates and stood it up at genuine frontier level, here is the projected first-year spend. These are Year-One estimates and the figures are indicative ranges, not quotes — they move with scale, vendor terms and how much of the system you own versus rent. Each line is tagged one-off · capital or ongoing · per year.
Enterprise-grade access to the leading models, at the volume an orchestrated system consumes.
~$120K–$360K / yearTo run the orchestrated system. Two paths — rent the capacity, or own it outright.
Cloud (ongoing): ~$180K–$480K / year — or — Owned hardware (one-off capital): ~$400K–$900K, plus ~$60K / year hostingOptional. Only applies if the coordinator is trained rather than simply configured — the build-it-once cost of a purpose-fit coordinator.
~$80K–$300KA small group of specialists to build, evaluate and operate the system.
~$700K–$1.6M / yearCounsel, IP protection, liability & indemnity structuring, and regulatory compliance — a genuine cost given the litigation and liability landscape around AI, and one that can run high if exposure is large.
~$150K–$500K / year* ≈ in-house / near-$0 if government-funded — public bodies carry their own legal teams and indemnity cover.A small R&D function that keeps the system improving — mining the open literature and our own results for low-cost capability gains (open-source integration, smarter routing, coordination methods) so it stays a sustainable lab, not a build-once prototype.
~$250K–$700K / yearYear one: roughly $1.3–3M for a fully legit, sustainable, world-class orchestrated system — a fraction of what it would cost to train a frontier model from scratch.
Fair question for any funder: if the idea is "orchestrate open models," what stops someone reading this page and rebuilding it? Short answer — the idea alone doesn't reproduce the result. The slide is cheap; the system, the recipe and the operator behind it are not. Here's what stays in when you fund this.
The orchestration approach and the sovereign-AI framing were conceived here — not lifted from anywhere. The strategic vision and the judgement behind it travel with the originator, not the slide. You can photocopy a thesis; you can't photocopy the person who keeps having them.
This is a running system, not a pitch deck. The build, the tuning and the iteration velocity sit with the team that made it. A copier opens a blank repo on day one; we open the one that already works.
The specific orchestration recipe — coordinator design, model mix, and the tuning that makes it actually beat the baselines — is deliberately kept out of the public materials. The headline idea is cheap. The working implementation is the expensive part, and it stays with us.
The kind of person dialled-in enough to spot this early is the kind who spots the next ten. You're not buying one good call — you're backing the instinct that keeps producing them.
Assembling the right specialists and the frontier-lab relationships is half the battle — and that network forms around the architect, not the abstract. Fund the originator and the team and access come attached.
We're already shipping daily. While a fork is still copying the concept, we compound — every day of head start widens into a moat.
Back the architect, not the idea — the idea is the easy part.
Working notes kept inline for now. We'd normally hold these on a separate page — delete this whole section before going fully public.
:free per board policy — no paid models).Funders, partners and first customers to approach, grouped by type. The government / sovereign-AI angles lead (the thesis is built for them); private capital and vertical first-customers run in parallel.
Ordered low → high effort — work top-down. The same warm-referral angle lifts all three.
Email a named person in the right team (not a generic inbox), keep it under ~120 words: one line on who we are, one specific line on why them (sovereign/on-shore AI angle tailored to their mandate), one clear ask ("a 20-minute intro call?"). A subject that signals relevance, not a pitch (e.g. "Sovereign on-shore AI — intro re: NAIC partnerships"). Link to a 1-page brief, don't attach a heavy deck (attachments trip spam filters and lower opens). A warm referral in the first line beats any cold send — get introduced wherever possible.
Call the department's listed line / EA a few days after the email, reference it by subject, and have a 20-second value prop ready ("sovereign frontier-class AI by orchestration — Australian-built, ready to pilot"). Don't pitch on the call; just ask for a short intro meeting with the right person. Calling confirms the email landed and moves you past the inbox.
Meet them where they already are: a relevant conference or industry event where the target speaks or exhibits (defence-innovation, sovereign-AI, deep-tech forums), or via a warm intro at a sector function. A 60-second hallway pitch plus a follow-up email referencing the meeting converts far better than any cold channel — but it's the most time-costly, so reserve it for high-value targets.