At Straight Up AI we’ve constructed an inner management aircraft for orchestrating coding brokers. How a lot latitude brokers are given is pushed by three elements:
-
Blast radius of a mistake. Mature manufacturing techniques typically have offline penalties whereas 0-1 MVPs don’t.
-
Mission context. Legacy tasks have much less embedded data that coding brokers make probabilistic judgements on.
-
Mission maturity. Greenfield tasks usually tend to comply with established design patterns that coding brokers can comply with extra simply.
This record is lacking one issue; the price of constructing. We’re a small consultancy. When not utilizing consumer allocations we run the management aircraft via a Claude Max account. As that is priced at £200 month-to-month, the marginal value of unhealthy choice making is time.
That could be a high quality place to show a management aircraft works. Creating it additional with out reviewing its sustainability is harmful. Nothing tells us whether or not we have now constructed one thing that we might genuinely afford to run. So I went and analysed seven weeks of the management aircraft in motion.
Our Utilisation
Between 15 July and 4 September we analysed 44 improvement cycles throughout our portfolio.
To calculate the API equal invoice I transformed each token to input-token equivalents with cached reads at 0.1x, cached writes at 1.25x, and output tokens at 5x.
At API charges we’d be paying 22 instances extra for our utilisation. As our consultancy grows this shortly turns into untenable, and is sort of the price of a mid engineer’s wage.

The second factor bothering me was time slightly than cash. A number of components of a run felt unnecessarily gradual. My finger pointed at our utilization of adversarial evaluation, which spins up a second coding agent to assault each commit/increment. That instinct turned out to be roughly proper, although not for the explanation I assumed.
|
Adversarial evaluation |
Implementation |
Ratio |
|
|
Brokers dispatched |
242 |
204 |
1.19x |
|
Price models |
234.8M |
339.6M |
69% |
|
Agent-hours |
28.0 |
39.2 |
71% |
|
Median per agent |
825k models, 5.6 min |
1.12M models, 7.8 min |
0.74x, 0.72x |
The adversarial evaluation overhead is coming from the amount of reviewers dispatched. It’s not one singular, costly evaluation agent. We submit extra reviewers than implementers with 26% of them demanding adjustments. When adjustments are requested a remediation agent is dispatched to implement the fixes, which spins up one other adversarial evaluation cycle. Much like common code evaluation, if an answer will not be discovered by the second or third spherical extra code isn’t the answer, and as an alternative extra structural adjustments are required. Adversarial evaluation on this context dangers rabbit-holing slightly than taking a step again and contemplating what adjustments could also be required that aren’t essentially within the commit scope.

The Management Airplane

A run has one controller and plenty of employees. The controller is the session you discuss to. It’s job is to:
-
Learn the construct plan and increments
-
Dispatch one employee per increment
-
Validate the increment state (accomplished, reviewed, failed)
-
Dispatch the subsequent job
It’s the solely agent that lives all through the whole job. This structure addresses a limitation with plan and construct setups. When a employee each narrates progress and advances a plan, incorrect progress statements cascade to incorrect increments. The controller structure stops this by appearing as an impartial arbiter of the truthfulness of the implementer. Our execute-plan contract states the rule as soon as and each variant inherits it:
The controller can also be required to remain awake for the length, which issues an ideal deal to the price figures additional down:
As soon as an increment is applied and verified, it goes to adversarial evaluation.
A recent agent receives the repository root, the diff scope (head vs base commit) and is instructed to explicitly discover failure eventualities. To make sure a good check the reviewer can not edit code; if it thinks there’s a logic or kind error it should generate a check to show it. The top is an adversarial report which is then fed right into a remediator agent which implements the fixes.
Adversarial evaluation is a blunt instrument by design. The choice at hand it to an costly coding agent (Fable 5 excessive), can also be intentional. It acts as a gate keeper for all code that might be progressed. A less expensive evaluation with decrease recall will cascade failures all through the system. And in scoping the evaluation to solely the related diffs we restrict the evaluation to solely the code that has modified. This ensures the evaluation agent doesn’t discover transient points or identified/accepted challenge dangers.
The place The Cash Goes
The costliest part within the system is the one which writes no code. The controller prices greater than implementation, planning and evaluation mixed.

That is the place I anticipated the investigation to finish shortly. We suspected context over lengthy classes could be an issue and designed the system to incorporate per-commit compaction. In actuality nevertheless this was not occurring.
|
Compaction occasions throughout all 44 controllers |
34 |
|
Controller API calls |
25,878 |
|
Calls per compaction |
~760 |
|
Controllers that compacted in any respect |
12 of 44 |
|
Largest context noticed |
996,659 tokens |
Throughout 75% of runs the controllers had been by no means compacted. And after we analysed the context home windows of the controllers we discovered that context elevated (kind of) linearly with session length. This factors to our instinct that compaction could be wanted to take care of context, nevertheless it was not being applied reliably.

|
API calls |
Share of calls |
Cache-read tokens |
Share |
Median context per name |
|
|
Controllers |
25,878 |
30.7% |
9.31bn |
58.7% |
312k |
|
Employees |
58,319 |
69.3% |
6.54bn |
41.3% |
101k |
The price of our system was not derived from implementation or evaluation. It was the re-reading and upkeep of the session state which elevated on each flip.
In reviewing the system we discovered aggressive compaction, being applied within the improper place. The person employees had been compacting as they opened on common at 38,000 tokens and closed at 103,000 tokens. The controller was not, and ended up carrying the state of each employee in reminiscence. This was an issue not simply from a price perspective. It additionally posed an issue for the reliability of the system because the degradation of efficiency as context home windows saturate is a identified limitation of LLMs.
An Sincere Look At The Controller
From each a price and reliability perspective, we wanted to know what was happening. We began off with the planning system, which was the obvious goal.
We plan in layers. An structure specification turns into an in depth technical specification, which turns into an implementation scope and a construct plan of numbered duties. The controller reads all of it. Though this chain gives a granular execution and analysis path, we thought sustaining this ledger throughout the whole run conflicted with compaction.
On measuring the paperwork in opposition to a median controller increment of 312,000 tokens we discovered specification stacks to be 2% of the overhead. This was the improper path. It’s also the only most dependable method for retaining LLMs on job and inventing options and implementation paths. A variety of time is spent reviewing these specs previous to dispatch, and having these scopes make it clear what we’re checking the brokers in opposition to.
|
Artifact |
Median size |
Approximate tokens |
|
Structure specification |
3,304 phrases |
4,600 |
|
Implementation scope |
1,322 phrases |
1,850 |
|
Construct plan |
1,041 phrases |
1,460 |
What truly fills the controller is its personal output.

The biggest single part of a controller’s context is its instrument arguments. And 1 / 4 of these had been the dispatch briefs used for sending work to brokers. Throughout the corpus there are 698 of them, with a median size of about 1,390 tokens and a longest of over 6,000.
Scoped planning is paid for right here as soon as, and is marginal. The temporary derived from the plan is written for each increment. The size is necessitated because it comprises the duty overview, acceptance standards, base commit and the constraints.
The difficulty is that these had been endured. On each subsequent increment these briefs are re-read till the run ends. This isn’t wanted as by design, the briefs are scoped solely to the increments they’re being measured in opposition to. Subsequently our system was ingesting a slowly incrementing stack of briefs slightly than treating the briefs as ephemeral and marking completion in opposition to the construct plans.
In reviewing our inner contract this was made worse in two locations. These are deliberate design selections which can be high quality in isolation. The primary is the controller should validate each employee outcome itself. Because of this all state should be evaluated by the controller, necessitating the proof being endured to the controller context.
The second is job monitoring. We request the controller to re-render the duty record on each state agent for simple human evaluation. The duty record subsequently additionally must be maintained. This, alongside job metadata, naturally grows throughout longer classes.
Submit the complete record at run begin… every time any job adjustments state (batched: one refreshed record per chunk of labor, not per instrument name), and within the closing message of each flip that leaves work excellent.
Mixed, these selections lead to a coding agent that accumulates metadata that isn’t shed, that will increase on each increment accomplished by one of many employees.
Having Our Cake and Consuming It?

The controller is stuffed with issues it’s already completed with. Three adjustments would guarantee they aren’t re-read on each increment.
Move briefs by reference, not by worth.
For those who’ve executed some C++/low degree language programming you may be accustomed to this idea. Objects which can be costly to create should not handed from one perform to a different. As an alternative, we move pointers or references to them in order that they are often reutilized downstream.
The run report exists in a persistent state that’s incremented on each flip move. The controller doesn’t have to learn about the whole report. It simply must know the newest state. Utilized recursively this ensures that each one earlier states are within the acceptable state earlier than a choice is made. This ensures the controller solely wants to carry a small quantity of metadata and a pointer to the run state.
Compact at increment boundaries.
As soon as an increment reaches a terminal state, nothing in its implementation transcript adjustments the subsequent choice.
The controller can rehydrate from the construct plan and run report, that are the supply of reality. This reduces context overhead from 65,000 to 487,000 tokens right into a sawtooth that resets each increment.
Throughout the corpus’s roughly sixteen dispatches per session, it will put a typical controller flip close to 100,000 tokens slightly than 360,000. I wish to be plain that it is a projection from what I measured, not a outcome I’ve run.
Scope the reviewer by try, not by increment.
The primary evaluation of an increment ought to see the entire change. The second, after remediation, ought to see the remediation diff and the findings nonetheless open in opposition to it. Like an efficient human code reviewer, this ensures that an more and more slim set of in scope objects undergo the evaluation funnel. In our system it’ll cut back the ratio between reviewers and executors.
Conclusion
We have developed a system that has turn into essential to each consumer supply and inner work. The system nevertheless will not be inexpensive, and leans on the discrepancy between Claude Max and API pricing, which might change shortly.
Up to now we have been centered on whether or not the system delivers. That is the primary time we have requested whether or not we are able to maintain it. If we handle the context bloat we predict we are able to. The structure is usually sound. The issues are derived from holding increment state within the international run context. Fixing this reduces each the price and the propensity for incorrect path following.
Our speculation on the bloat was additionally fully improper, twice. We believed we had been compacting aggressively at each increment. We then assumed the planning chain was the load, for the reason that controller reads an structure spec, a technical spec, an implementation scope and a construct plan on each run. That complete stack is beneath 8,000 tokens, roughly 2% of a median controller flip.
What truly stuffed the controller was the fabric it generated itself. Device-call arguments are 37.7% of all the things it accumulates, and 1 / 4 of these are the 698 dispatch briefs it wrote and by no means put down. The train has demonstrated the significance of finishing evaluation with out priors, and being trustworthy about your work.
















