A tiered model architecture for development with LLMs
Picking the best model is the wrong problem to solve
When developers use LLMs, whether they’re writing back-end code or building strategies, most of them try to choose the best model for the task at hand. It’s a common mistake that misreads the problem, and it also runs the developer straight into the arbitrary limits providers set to ration how much compute each customer can use.
The multi-model architecture we describe here deals with three high-level problems: getting more done, spending less, and keeping a run alive when an account hits its limit. It won’t suit everybody, mainly because keeping several high-end model subscriptions is expensive (you can build the same thing on low-cost commercial models, but you’ll get a lower-quality result).
Tooling and model names date quickly, so what follows describes the structure and leaves the setup details to you.
Three separate roles
A long task that calls several models needs three roles kept apart: an orchestrator, a proxy, and a skill that runs the procedure.
The orchestrator starts and supervises work, and it isn’t a model. It gives each job an isolated workspace (usually its own checkout), injects whatever environment that job needs, and then leaves the session alone until the work finishes or the provider stops it. Conductor is one example: a Mac app that runs several coding agents in parallel, each in its own worktree, though any tool built along the same lines would serve. The orchestrator doesn’t rank implementations or decide a winner.
The proxy sits between the orchestrator and the providers. It holds a pool of subscriptions and presents one endpoint to the workspace, and every task talks to that endpoint. When an account hits a session window or a weekly cap, the proxy cools it and carries on with another. Session affinity matters here: a job should stay on one account until that account is exhausted, so that the prompt cache isn’t rebuilt on every turn. We don’t depend on any particular product for this, and any gateway that can sign in several accounts, fail over on a limit, and keep a conversation pinned will serve.
The skill is the procedure inside the workspace. It freezes a specification, writes tests against that specification before any implementation exists, and fans the same frozen spec out to a pool of models. Anthropic seats, OpenAI seats, and a Grok judge make up the core of our pool, and we supplement with various other models (such as Gemini, Muse, Kimi, Qwen, GLM, and others) that can join the the same fan-out as the builders or reviewers, but none of them acts as the orchestrator.
The three roles fail in different ways: if the orchestrator fails you lose the workspace, if a provider fails you lose one account, and a skill failure is a bad spec, a bad test, or a judge that saw the author’s name. We keep them apart so that a failure in one role stays in that role.
One model, or a portfolio of them
The usual way to use a coding model is the way most people use a chart: pick one, give it the task, and wait for the outcome, with one subscription, one session, and one implementation. If the result is wrong, the remedy is another prompt to the same model, or a manual switch to a different one and a fresh start. There’s no second opinion formed without seeing the first, and no test written before either opinion existed.
We treat that the way a book treats a single strategy. A book holds a portfolio of strategies because any one rule has a regime where it fails, and because the disagreement between rules is itself information; a portfolio of models is the same bet. Several models get the same frozen problem and produce competing solutions in isolation, and then a judge that wrote none of them ranks the field. What we keep is whichever solution survived a comparison its authors didn’t control, even when it came from a model we happen to like less.
That’s why the pool exists, and also why this layout isn’t for everyone. A tournament holds several seats open at once (a spec review, the builders, a gate, two judging passes, and a review panel), and a single subscription gives you one window. If there’s nowhere else to send the next call, the first account to hit its cap takes the run down with it.
A pool turns that cap into a handoff: the workspace keeps one endpoint and the proxy chooses which subscription answers. The OpenAI seats and the Grok judge follow the same pattern against their own subscriptions, because the skill calls those tools and the subscriptions behind them are pooled, so a limit on one login doesn’t end the run.
It’s also expensive, and unevenly so, because you need enough high-end plans that one cooling down still leaves another able to finish the seat. If you have a single subscription, or a metered API budget spent on one model, you don’t need the proxy or the tournament. Anyone that can’t isolate workspaces, or doesn’t keep provider credentials on a machine they control, shouldn’t copy the layout either; we run it as a research setup, with the model credentials properly isolated..
What the procedure asks of the pool
The procedure is a tournament, and it starts by freezing a spec that covers signatures, error behaviour, and the files a builder must return. Tests are written against that spec and then hidden. The builders implement in isolation; in our runs those seats include GPT-6 Astra and Sol and Claude Fable, Opus, and Sonnet, with others on trial. None of them sees the tests, or any of the repo beyond the spec.
Every candidate is compiled and run against the tests written in advance, and the survivors go to a judge that didn’t write any of them. We use Grok for that, twice, with the second pass in reverse order so that a preference for whatever arrives first can’t crown the same seat every time. The winner then goes into the tree with at most one graft, which must come from a ballot that also picked the winner. If two ballots name different donors, we take the donor with the higher mean rank. After that, a panel drawn from outside the winner’s model family attacks the result through separate lenses (correctness, data, concurrency, performance, and test adequacy). The orchestrator owns the tests but has no say in who wins.
The pool’s job is narrow: it has to keep each of those calls alive until the procedure says the run is over, however many account windows close along the way.
A run that shows the split
The example below comes from our own development work, on a task that adjusts equity bars for corporate actions at read time. The task itself doesn’t matter much; it’s here to illustrate the process. The field was six seats, and both ballots named the same winner.
At the end of every run the skill prints a scorecard, and this is the top of the one for this run:
TOURNAMENT SCORECARD: Frozen specification: adjusted equity bars computed at read time (marketfeed GLE-320, round MR6)
run tournament.w4-mr6, 2026-10-05
Spec, tests, integration: Claude Fable 5.1 (orchestrator, high)
Advisor (spec freeze): Claude Fable 5.1 (max), 7 findings, 7 upheld
Builders
model build vet spec tests ranks scores mean rank status
----------------- ----- --- ---------- ------- --------- --------- ------
Claude Fable 5.1 ok ok ok a:1 b:1 a:93 b:91 1/6 winner
Claude Opus 5.5 ok ok ok a:2 b:2 a:90 b:89 2/6 judged
GPT-6 Astra ok ok ok a:4 b:3 a:85 b:85 3.5/6 judged
Claude Sonnet 5.5 ok ok ok a:3 b:4 a:88 b:83 3.5/6 judged
GPT-6 Sol ok ok ok a:5 b:5 a:82 b:79 5/6 judged
Gemini 3.1 Pro ok ok ok a:6 b:6 a:78 b:74 6/6 judged
Judge: Grok 4.7 (xhigh), 2 of 2 ballots valid, winners agree
The advisor, a fresh Fable pass before any builder started, raised seven findings, none of them High, and all seven were applied. The graft was Opus’s same-trade-date factor memo in the stamping loop, which the reverse ballot proposed; Opus also had the higher mean rank of the two donors. The other ballot had proposed Sonnet’s binary search for the previous close, and although we didn’t take it as the graft, the same idea came back later from the review panel as a fix.
Because Fable won, the lenses normally held by Claude models moved to spares: correctness went to Astra, data integrity to Grok, and concurrency to Sol. Performance stayed with Gemini and test adequacy with GPT-6.1 Sol. All five reviewers reported, and the panel didn’t re-review the fixes.
The rest of the scorecard covers the panel and the outcome:
Review panel
lens model result
----------- ------------------- -------------------------------
correctness GPT-6 Astra (spare) 0 findings, 0 upheld
data Grok 4.7 (spare) 0 findings, 0 upheld
concurrency GPT-6 Sol (spare) 0 findings, 0 upheld
performance Gemini 3.1 Pro 2 findings, 1 upheld (High 0/1)
tests GPT-6.1 Sol 7 findings, 6 upheld
Outcome: shipped Claude Fable 5.1; fallback none; final gates pass
The run summary lists each finding and what was done about it:
┌─────────────┬────────────┬──────────────────────────────┬─────────────────────────────────┐
│ Lens │ Severity │ Finding │ Action │
├─────────────┼────────────┼──────────────────────────────┼─────────────────────────────────┤
│ Performance │ High, │ A dates slice kept parallel │ Declined. The bar slice already │
│ Gemini │ downgraded │ to the bars, 24 bytes a bar │ holds every bar, and each bar │
│ │ to Low │ on large requests │ is more than 200 bytes, so the │
│ │ │ │ extra slice is under 15 │
│ │ │ │ percent. Dating before the │
│ │ │ │ anchor and ledger reads is the │
│ │ │ │ refusal order the spec pins. │
│ │ │ │ Correctness confirmed that │
│ │ │ │ order. │
├─────────────┼────────────┼──────────────────────────────┼─────────────────────────────────┤
│ Performance │ Medium │ Previous-close walked the │ Fixed. Binary search over the │
│ Gemini │ │ tape from the end once per │ ascending tape. Closes ascend │
│ │ │ dividend, O(dividends x │ by date, which the spec already │
│ │ │ closes) │ required. │
├─────────────┼────────────┼──────────────────────────────┼─────────────────────────────────┤
│ Tests │ Medium │ No request whose dividends │ Fixed. A review test kills the │
│ Sol 6.1 │ │ all fall after the anchor, │ len(divs) > 0 mutant. │
│ │ │ so an unconditional tape │ │
│ │ │ read still passed │ │
├─────────────┼────────────┼──────────────────────────────┼─────────────────────────────────┤
│ Tests │ Medium │ A NaN previous close was │ Fixed. │
│ Sol 6.1 │ │ untested. !(prev > 0) still │ │
│ │ │ passed │ │
├─────────────┼────────────┼──────────────────────────────┼─────────────────────────────────┤
│ Tests │ Medium │ Every ex-date was UTC, so │ Fixed. New York evening and │
│ Sol 6.1 │ │ local-date extraction still │ Sydney morning, splits and │
│ │ │ passed │ dividends. │
├─────────────┼────────────┼──────────────────────────────┼─────────────────────────────────┤
│ Tests │ Medium │ Early close, │ Fixed. A review test kills a │
│ Sol 6.1 │ │ settlement-close, and an │ rebuilt session block. │
│ │ │ absent settle were not │ │
│ │ │ asserted │ │
├─────────────┼────────────┼──────────────────────────────┼─────────────────────────────────┤
│ Tests │ Medium │ No late dating failure past │ Deferred. A │
│ Sol 6.1 │ │ the first chunk │ datable-then-undatable instant │
│ │ │ │ needs an era-disagreement │
│ │ │ │ calendar the fixture does not │
│ │ │ │ expose cheaply. Atomicity is │
│ │ │ │ structural: one emit, after │
│ │ │ │ every error return. Data and │
│ │ │ │ correctness both confirmed │
│ │ │ │ that. │
├─────────────┼────────────┼──────────────────────────────┼─────────────────────────────────┤
│ Tests │ Medium │ Concurrent requests never │ Fixed. A race test, deep-equal │
│ Sol 6.1 │ │ included paths │ including paths. │
├─────────────┼────────────┼──────────────────────────────┼─────────────────────────────────┤
│ Tests │ Low │ A refused schedule's value │ Fixed. The refusal must return │
│ Sol 6.1 │ │ and the full error text were │ the zero schedule. │
│ │ │ unchecked │ │
└─────────────┴────────────┴──────────────────────────────┴─────────────────────────────────┘
Correctness, data integrity, and concurrency returned no findings. Six of the seven test-lens findings held up, and all six were fixed with review tests, each checked by killing its mutant; the seventh was deferred.
The shipped code is Fable’s, plus that one graft, on the market-data adapter. Every gate on the rebased tree passed, including the race detector on the adjustment package and the adapter, and coverage was 91.1 per cent. Mutation testing against the reference killed 64 of 64 mutants. Against the merged winner it produced 90 mutants: 78 were killed, seven were equivalent, and the review tests killed the five real gaps.
Build the architecture, then swap the models
No single model should be expected to solve problems as complex as infrastructure coding or strategy development. And of course today’s “best” model is tomorrow’s also-ran, so picking the single best model is a problem that never stays solved. The answer is to build the right architecture and keep swapping in the parts that change as they improve (which they do, rapidly!).
In our setup an orchestrator keeps the workspaces isolated, a proxy turns a pool of subscriptions into one endpoint that every workspace calls, and the best models of the day compete to build the solution and then argue over the finer points to arrive at the end goal.
This might all seem slightly intimidating if you’ve just settled into the idea that you can ask ChatGPT to solve a development problem or build a strategy for you. And, while many retail traders might be tempted to think that this could yield them some kind of market edge, the process we’ve described here is what more serious players (aka the competetion!) are likely doing. As is always the case in trading, it’s instructive to think about how the party who might be on the other side of your trade is thinking about the problem.