★ The Working Prototype

One coordinator.
A society of specialist models.
Frontier-class output.

This is the proof made tangible. A coordinator decomposes your task, routes the pieces across a panel of open specialist models, has them debate and verify, and synthesises one answer — open models lifted to near-frontier. Swap the panel for frontier models and the same architecture is built to go beyond frontier. Try it, see how it's measured, and read what funding unlocks.

Section 1 · Interactive prototype

Chat with the orchestrated model.

Behind this single chat box sits a society of specialist models under one coordinator. You ask once; the coordinator plans, routes sub-tasks to the best-suited specialist, runs a round of cross-checking, and returns one synthesised answer. The interface is deliberately simple — the orchestration is the product.

Orchestrated Prototype — Coordinator + specialist panel demo · read-only preview
Explain, step by step, why a hard-coded coordinator over open models can match a frontier model on a reasoning task.
Orchestrated answer A frontier model carries breadth and verification inside one set of weights. A coordinator reconstructs both externally: it decomposes the problem, routes each part to a specialist that is strong on that part, then has independent models check each other before a final synthesis. The diversity of the panel supplies breadth; the debate-and-verify step supplies the self-correction. No single open model is frontier-class — but the coordinated system, measured end-to-end, can be.
coordinator: plan specialist: reasoning specialist: code-check debate ×2 synthesis
System Live demo coming soon. This is a preview of the interface — the orchestration backend is being wired up. Leave a question below and we'll show you a worked trace once the live endpoint is connected.

No data is sent yet — this is a UI preview. The live orchestration endpoint is on the build queue.

Section 2 · Benchmarks

How the orchestrated model is measured.

The whole claim — "orchestration lifts a tier" — only matters if it survives a scoreboard. We measure three groups on the same public benchmarks: the individual open component models that make up our panel, the frontier models we're measured against, and our orchestrated system.

★ The result we already have

On MMLU-Pro, scored against the objective answer key — no model judging another — a tighter, complementary trio (glm-4.5 · gemma-4-26b · nemotron-super, three lineages clustered within ~1 point) now makes simple majority voting beat the best single model by +1.6 (84%) — the program's first selection-tier win (an earlier, looser trio lost −1.0; clustering the panel in strength is what flips it positive). The selection ceiling (the correct answer present somewhere in the panel) is +7.7 above the best single — most of that headroom is still unclaimed. Notably, a reconciling LLM judge did NOT beat the plain vote here: a strong judge reverts to its own answer (self-preference), a weak external judge is helped by the drafts but still trails — so for this panel the diversity's agreement is a stronger signal than any single free judge's verdict, and reaching the ceiling via synthesis is the open frontier. Full method, every number, the negative results, and the LLM-as-judge literature are in the research paper →. The expandable per-benchmark harness below (more models, more evals, live) is being stood up; until it's wired we show "soon" rather than numbers we haven't actually run.

Component models (open-weight, individual) Frontier models (reference) Our orchestrated system
Model / system General reasoning Hard QA Math Code Aggregate
Component model A open · singlesoonsoonsoonsoonsoon
Component model B open · singlesoonsoonsoonsoonsoon
Component model C open · singlesoonsoonsoonsoonsoon
Frontier reference 1 closed · referencesoonsoonsoonsoonsoon
Frontier reference 2 closed · referencesoonsoonsoonsoonsoon
★ Orchestrated system coordinator + open panelsoonsoonsoonsoonsoon
★ Orchestrated system coordinator + frontier panel — "beyond frontier"soonsoonsoonsoonsoon

The read we're after: the orchestrated row over an open panel should sit level with the frontier-reference rows while every individual component sits below them — that's the "lift a tier" result. The orchestrated row over a frontier panel is the "beyond frontier" hypothesis. Column benchmarks shown generically (reasoning / hard-QA / math / code) until the displayed suite is finalised — see the worklist below.

Section 3 · The case for funding

What more funding unlocks.

The prototype is built to run at near-zero cost on open, free-tier models — that's the point, it proves the method is cheap to verify. But the same architecture has a steep, fundable improvement curve. Each lever below is a concrete place where capital converts directly into capability. For a government, defence partner or sovereign-AI investor, this is what the spend buys.

1 · A high-level LLM coordinator biggest lever

Today the coordinator is hard-coded — a fixed decompose → route → verify → aggregate pipeline. Replacing it with a strong LLM acting as coordinator is the single largest jump available: it plans adaptively, recognises when a sub-task needs a different specialist, decides when to debate further versus commit, and writes a better final synthesis. The hard-coded version proves the method on a budget; the LLM coordinator is where it becomes genuinely frontier-grade.

Funding unlocks → per-query inference budget for a capable coordinator model on every request.

2 · Frontier-class component models beyond frontier

The current panel is open, non-frontier models — chosen so the lift is clean to demonstrate. Level the panel up so the orchestrated specialists are themselves frontier-class, and the ceiling moves: a coordinated society of frontier models is built to exceed any one of them on its own. This is the "beyond frontier" path — the same architecture, far stronger parts.

Funding unlocks → access to top-tier model APIs / weights as the orchestrated panel.

3 · More compute for deeper deliberation depth

Quality scales with how many rounds of debate, verification and best-of-N sampling the system can afford. Today that's throttled by free-tier rate limits. More compute means more independent attempts, more cross-checking, and longer multi-round reasoning before the system commits — directly lifting reliability on the hardest problems.

Funding unlocks → dedicated inference compute for wider, deeper deliberation.

4 · Fine-tuning the coordinator on orchestration specialisation

An off-the-shelf model coordinating is good; a model trained on the orchestration task itself — learning which specialist to trust for which problem, how to phrase sub-tasks, when to stop debating — is better and cheaper per query. This is the bridge from the hard-coded coordinator to a learned one, and it compounds every other lever.

Funding unlocks → data collection + training runs for a purpose-built coordinator.

5 · Scaling the number and diversity of experts breadth

The strength of a society of models comes from diversity — different architectures, training data and failure modes that don't all break the same way. Expanding the panel and adding domain specialists (code, maths, retrieval, vision, legal, defence) widens coverage and sharpens routing, so the coordinator always has a strong specialist on hand.

Funding unlocks → a larger, more diverse, sovereign-hosted expert panel.

6 · Best-of-breed in every modality — even where frontier models are weak opens the door

A single "multimodal" model is forced to be a generalist: it spreads one set of weights across text, vision, audio and every domain, and it is precisely in the niche modes — specialist vision, audio, OCR, rare languages, narrow technical fields — where even the best monolithic models go soft. Orchestration removes that compromise. Because the coordinator only routes, each sub-task goes to the single strongest model for that mode — a dedicated vision model, a dedicated audio model, a code specialist — so the system fields the frontier of every modality at once, which no one model can. It also unlocks capability no single model has at all: a text-only panel cannot see an image, but a coordinator that can call a vision specialist can — there the gain is not a few points, it is an entire capability switched on.

Funding unlocks → a roster of best-in-class specialist models, one per modality and niche, under the coordinator.

7 · Evaluation & sovereign-deployment infrastructure trust

Capability you can't measure isn't fundable. A serious eval harness — held-out suites, contamination controls, frontier-grade judging, red-teaming — turns the benchmark frame above into defensible, repeatable evidence. Paired with the ability to run the whole stack on sovereign infrastructure, it's what lets a government or defence partner actually depend on the system.

Funding unlocks → eval engineering + on-shore, air-gappable deployment.

8 · A standing research arm keeps improving

The cheapest capability gains often come from method, not spend. A small in-house research function continuously mines the open literature and our own results for low-cost wins — smarter routing, open-source integration, orchestration improvements, better coordination methods — so the system keeps getting better between funding rounds instead of being built once and left to decay. This is what makes it a sustainable sovereign lab, not a one-off prototype.

Funding unlocks → an ongoing R&D loop that compounds every other lever at minimal extra cost.
Section 4 · What it costs at frontier scale

What it costs to do this at frontier scale (Year One).

If you took everything the prototype demonstrates and stood it up at genuine frontier level, here is the projected first-year spend. These are Year-One estimates and the figures are indicative ranges, not quotes — they move with scale, vendor terms and how much of the system you own versus rent. Each line is tagged one-off · capital or ongoing · per year.

1 · Frontier model access ongoing · per year

Enterprise-grade access to the leading models, at the volume an orchestrated system consumes.

~$120K–$360K / year

2 · Compute & infrastructure ongoing or one-off

To run the orchestrated system. Two paths — rent the capacity, or own it outright.

Cloud (ongoing): ~$180K–$480K / year — or — Owned hardware (one-off capital): ~$400K–$900K, plus ~$60K / year hosting

3 · Coordinator development one-off · capital

Optional. Only applies if the coordinator is trained rather than simply configured — the build-it-once cost of a purpose-fit coordinator.

~$80K–$300K

4 · Core team ongoing · per year

A small group of specialists to build, evaluate and operate the system.

~$700K–$1.6M / year

5 · Legal & compliance ongoing · per year scales with exposure

Counsel, IP protection, liability & indemnity structuring, and regulatory compliance — a genuine cost given the litigation and liability landscape around AI, and one that can run high if exposure is large.

~$150K–$500K / year* ≈ in-house / near-$0 if government-funded — public bodies carry their own legal teams and indemnity cover.

6 · Standing research arm ongoing · per year

A small R&D function that keeps the system improving — mining the open literature and our own results for low-cost capability gains (open-source integration, smarter routing, coordination methods) so it stays a sustainable lab, not a build-once prototype.

~$250K–$700K / year
Total capital (one-off) ~$80K–$300K on the pure-cloud path (coordinator dev only)
rising to ~$0.5M–$1.2M if compute is owned outright
Total Year-One operating (ongoing / per year) ~$1.25M–$3.1M / year
(legal ≈ in-house if government-funded — see line 5)

Year one: roughly $1.3–3M for a fully legit, sustainable, world-class orchestrated system — a fraction of what it would cost to train a frontier model from scratch.

Section 5 · Why this isn't a copy-paste

Indispensable — the part you can't fork.

Fair question for any funder: if the idea is "orchestrate open models," what stops someone reading this page and rebuilding it? Short answer — the idea alone doesn't reproduce the result. The slide is cheap; the system, the recipe and the operator behind it are not. Here's what stays in when you fund this.

1 · Architect of the thesis the originator

The orchestration approach and the sovereign-AI framing were conceived here — not lifted from anywhere. The strategic vision and the judgement behind it travel with the originator, not the slide. You can photocopy a thesis; you can't photocopy the person who keeps having them.

2 · A working prototype already exists shipping

This is a running system, not a pitch deck. The build, the tuning and the iteration velocity sit with the team that made it. A copier opens a blank repo on day one; we open the one that already works.

3 · The method is the moat — and it isn't in this document withheld

The specific orchestration recipe — coordinator design, model mix, and the tuning that makes it actually beat the baselines — is deliberately kept out of the public materials. The headline idea is cheap. The working implementation is the expensive part, and it stays with us.

4 · Operator instinct next ten

The kind of person dialled-in enough to spot this early is the kind who spots the next ten. You're not buying one good call — you're backing the instinct that keeps producing them.

5 · Team & access network

Assembling the right specialists and the frontier-lab relationships is half the battle — and that network forms around the architect, not the abstract. Fund the originator and the team and access come attached.

6 · Speed compounding

We're already shipping daily. While a fork is still copying the concept, we compound — every day of head start widens into a moat.

Back the architect, not the idea — the idea is the easy part.

Section 6 · Internal working doc

Open decisions, tasks & ideas.

Working notes kept inline for now. We'd normally hold these on a separate page — delete this whole section before going fully public.

⚠ INTERNAL — remove this section before public launch

Component models — which to orchestrate

  • Pick the open panel. Candidates from the free/open-weight tier (general reasoning + a code specialist + a maths specialist). Decide on 3–5 with genuinely different lineages so failure modes don't correlate. Confirm they're reachable on free tiers we already use (OpenRouter :free per board policy — no paid models).
  • Decide the panel-vs-cost tradeoff. More experts = better coverage but more rate-limit pressure on free tiers. Find the smallest panel that still shows the lift.
  • Lock a fixed panel for the headline benchmark so results are reproducible; keep an "experimental panel" lane for testing additions.

Benchmarks — which to display

  • Choose the displayed suite. Candidates: MMLU / MMLU-Pro (broad knowledge), GPQA (hard graduate QA), MATH or AIME (maths), HumanEval / MBPP / LiveCodeBench (code), BBH / ARC (reasoning), GSM8K (grade-school maths sanity). Pick a small honest set — too many columns dilute the story; pick ones where the lift is real, not cherry-picked.
  • Decide contamination handling. Prefer newer / held-out or rotating sets so component models haven't memorised them; note it on the page.
  • Frontier-judge column? Keep Opus-4.8-as-judge (per the index page's proof framing) as a qualitative score alongside the quantitative suite, or drop it from the public table.
  • Decide whether to show absolute scores or deltas-vs-best-component — the delta tells the "lift a tier" story more directly.

Positioning — what level to pitch the prototype at

  • Audience tier: government / sovereign-AI grant bodies, defence contractors, or high-net-worth strategic investors. Each wants a different emphasis (national capability vs deployable system vs upside). Decide the primary audience for v1.
  • Honesty line: keep the "we lift a tier, we did not pretrain a frontier model" framing consistent with the index page. Don't overclaim "beyond frontier" until the frontier-panel row has real numbers.
  • Demo depth: ship the mock chat first, or wait for a live worked-trace endpoint before pitching? Decide minimum-credible demo.

Research — orchestration literature to pull

  • Read deeper into the academic literature on multi-model orchestration / "society of models" / ensembles-beat-single-model / multi-agent debate / mixture-of-agents / LLM-routing. Build a proper annotated bibliography (seeds in Section 7).
  • Map our hard-coded coordinator to a named method in the literature so the approach is defensible and we can cite prior art on capability gains.
  • Find the strongest published evidence that orchestration/ensembling closes a tier gap, and the known failure modes (cost blow-up, latency, error compounding) so the pitch is honest about limits.

Build / to bring the prototype to life

  • Wire the live orchestration endpoint behind the chat box (coordinator → route → debate → synthesise), with a visible per-query trace. (Gated on Chef's "cook the swarm" nod — index-page memory says foothold site first.)
  • Stand up the eval harness to fill the benchmark table — automate run + result ingestion so numbers are reproducible, not hand-entered.
  • Decide hosting for the live demo (which port / tmux session / tunnel route) and rate-limit / abuse protection for a public chat box.
  • Privacy/abuse: sanitise prompt-injection on the public input (consistent with our other public chat surfaces).
Section 8 · Plan to contact

Who to approach (working list)

⚠ INTERNAL — outreach targets & strategy. Not for public / investor-facing display. Hide before launch.

Funders, partners and first customers to approach, grouped by type. The government / sovereign-AI angles lead (the thesis is built for them); private capital and vertical first-customers run in parallel.

Government & defence

Research, grants & deep-tech funds

Private capital

Vertical first-customers (revenue, not just funding)

Three best ways to reach them

Ordered low → high effort — work top-down. The same warm-referral angle lifts all three.

1 · Emaillowest effort

Email a named person in the right team (not a generic inbox), keep it under ~120 words: one line on who we are, one specific line on why them (sovereign/on-shore AI angle tailored to their mandate), one clear ask ("a 20-minute intro call?"). A subject that signals relevance, not a pitch (e.g. "Sovereign on-shore AI — intro re: NAIC partnerships"). Link to a 1-page brief, don't attach a heavy deck (attachments trip spam filters and lower opens). A warm referral in the first line beats any cold send — get introduced wherever possible.

2 · Phonemedium effort

Call the department's listed line / EA a few days after the email, reference it by subject, and have a 20-second value prop ready ("sovereign frontier-class AI by orchestration — Australian-built, ready to pilot"). Don't pitch on the call; just ask for a short intro meeting with the right person. Calling confirms the email landed and moves you past the inbox.

3 · In personhighest effort · last resort

Meet them where they already are: a relevant conference or industry event where the target speaks or exhibits (defence-innovation, sovereign-AI, deep-tech forums), or via a warm intro at a sector function. A 60-second hallway pitch plus a follow-up email referencing the meeting converts far better than any cold channel — but it's the most time-costly, so reserve it for high-value targets.