A current Google paper confirmed spec-driven take a look at technology raised bug detection by 9.8 proportion factors on a pattern from their codebase. However their spec is learn out of the code, which leaves the half I care about unresolved. I’ve been arguing for months {that a} take a look at suite written by the identical mannequin that wrote the code can’t actually disagree with it and to assist my declare I’ve developed a brand new open supply Python library.
First, a bit of background. Software program testing has been transferring in a single course for twenty years, and Specification-Pushed Growth (SDD) is the place that motion has not too long ago arrived. This text argues for yet another step: Impartial SDD, or ISDD.

TDD (Check‑Pushed Growth) mentioned the checks are the specification. Write them first and the design follows. It labored, and it left the specification in a kind solely a programmer might learn, which meant the one who knew what the software program was for couldn’t verify it.
BDD (Behaviour‑Pushed Growth) was the reply to that. Structured, plain language situations written in Gherkin, utilizing “Given, When, Then” to explain system behaviour. Enterprise analysts, builders and testers argue over the identical artifact earlier than any code is written, and what they agree on is the anticipated behaviour.
SDD (Specification‑Pushed Growth) is the model that arrived with the brokers. A proper specification, typically in EARS (Simple Strategy to Necessities Syntax) notation. It’s exact sufficient {that a} machine can plan from it, break it into duties, generate the checks and generate the code. The spec stops being only a formality beside the work and turns into the factor the work is definitely generated from.
What’s fascinating to notice right here is that every step moved the supply of fact additional upstream: from the checks, to a shared description, and eventually to a machine-readable contract.
That is the best course, however my argument is with what occurs in SDD utilizing typical AI workflows. At present, Spec-Pushed Growth divides the work. It doesn’t divide who has the data. The identical specification goes to the planner doing the duty breakdown, the take a look at generator and the coding agent. That is “divide and conquer” utilized to the workflow pipeline. The data base is left complete, and handed to everybody.
Divide and conquer solely works when the road you chop alongside is the road the failure runs throughout. Right here the failure is that the code and the checks come from one studying of the identical ambiguous sentence, therefore the road runs throughout data, not work. My argument is to chop alongside that line: give the coding agent the choices (the necessities) and withhold the results (the acceptance standards), and the take a look at suite turns into one thing that may inform the code it’s mistaken.
One phrase wants pinning down earlier than I get to the paper. A contract is an announcement of what the code should do. Google’s agent writes its contract by studying the implementation. My contract is written earlier than the implementation exists, and by no means derived from it. Similar phrase, reverse provenance, and the provenance is your complete argument.
My contract additionally arrives in two halves. The necessities are the choices anyone made, and the coding agent reads them. The acceptance standards are the results that comply with, and it by no means sees them.
Just a few weeks in the past, a group at Google measured the step earlier than withholding: whether or not writing the contract down helps in any respect. “Grounding AI Brokers in Contracts: An Empirical Analysis of Spec-Pushed Check Technology” (arXiv 2608.17177) does one thing narrower than my argument, and it measures what it does. As a substitute of prompting an agent to write down checks, they first ask it to purpose concerning the code and write down its contract: pre-conditions, post-conditions, and the behaviours which can be merely undefined. That doc turns into what they name a cognitive scaffold and the checks are generated from it.
The outcomes, on manufacturing bugs from Google’s personal codebase

That final quantity implies that greater than half the time, a take a look at suite generated from a written contract beat the checks an individual wrote.
Why I’m happy about outcomes I didn’t produce
The paper checks the half of my declare I’ve not been in a position to take a look at but. Here’s what I’ve been saying, in two components:
First: a take a look at written from a said specification is best than a take a look at written from an impression of the code. That is what Google measured, and a specification written down is what they name a contract.
Second: the contract needs to be hidden from whoever writes the code (in any other case the 2 agree by development and the take a look at can’t fail).
The primary is now measured, by a critical group with an actual bug corpus, actual statistics and no stake in my conclusion. It got here out within the course I claimed, with a p-value. My very own proof for it was restricted to inside runs alone duties, which is why anyone else’s corpus issues.
Why it doesn’t settle the factor I actually care about
Have a look at the place their contract comes from. The agent reads the supply and paperwork what it finds. Pre-conditions, post-conditions, undefined behaviour, all of it derived from the implementation in entrance of it.
So the scaffold is actual and it really works, however it’s downstream of the code implementation. If the implementation resolved an ambiguity the mistaken method, a contract written by studying that implementation information the mistaken decision. The take a look at then confidently passes, and as I defined in Half 1 of this sequence, the inexperienced take a look at suite means nothing for the reason that loop by no means noticed an announcement of right behaviour that was impartial of the factor being judged.
Google’s personal framing is sincere about this. The scaffold is described as enhancing the agent’s reasoning, not as an impartial normal. That may be a declare about consideration, and it’s a good one. Independence is a unique declare, and the paper doesn’t make it.
No one has measured whether or not withholding the standards from the code-writing agent catches extra actual defects than a take a look at suite written with full sight of the code. I’ve measured the neighbouring query, whether or not refining the standards sharpens the take a look at suite, throughout a number of rounds of planted-fault experiments. I’ve not discovered clear proof of an impact but. I nonetheless anticipate one to be there, and greater than as soon as an early run seemed prefer it, however every obvious acquire disappeared after I in contrast the refined-criteria runs and the original-criteria runs on matched artifacts.
I’ve since measured one thing adjoining, and it’s price separating from what I’ve not measured. Throughout twenty runs on two duties, one agent wrote the implementation from the necessities alone, and a second agent wrote the take a look at suite from the acceptance standards, which the primary by no means noticed. The suite disagreed with the code each single time, ten out of ten on every. By disagreed I imply a take a look at failed: the suite asserted one behaviour and the implementation did one other. In every case the coding agent had additionally recorded the judgement name prematurely, which is what makes the disagreement significant: the agent discovered the paradox itself, resolved it, and had no strategy to verify the decision. The suite was the verify.
The 2 duties disagreed in numerous methods. On one it was a plain defect: banker’s rounding the place cash wants half up. On the opposite it was no defect in any respect, however a choice no one had made, sitting unnoticed within the distinction between < and <=.
That may be a smaller declare than the one I would love. It says the suite finds disagreements between the implementation and the specification, and that in each of the instances I checked out, the disagreement was price having: one a defect, one a choice no one had made. It doesn’t say the suite catches extra bugs than a collection written the atypical method, which remains to be the open query and nonetheless no one’s end result.
One shortcoming I seen throughout these experiments is {that a} criterion is just nearly as good as the information that workouts it. A criterion about scientific notation can’t fireplace on information containing none, and a take a look at written from that criterion will go regardless of the code does. That led me so as to add a brand new flag to qikly, --propose-fixtures, which finds the standards your information can’t attain.
My takeaways from the paper
Three issues, and the third one is essential.
-
Writing the specification down beats leaving it implicit. That’s now measured fairly than asserted, by anyone else, and it’s a higher quotation than something I might produce about my very own device.
-
The place the specification comes from is a separate query, and Google’s end result doesn’t contact it. Their contract is learn off the code, which is the most affordable supply and the one supply that may by no means let you know the code is mistaken.
-
And who’s allowed to learn it’s a third. TDD, BDD and SDD every moved the specification additional upstream, and none of them needed to reply this, as a result of an individual who writes each the specification and the code nonetheless meets code assessment, a tester, and a colleague who reads the identical sentence in a different way. An agent meets none of them. Divide the work all you want; the division that decides whether or not inexperienced means something is the one throughout what every agent is aware of.
Getting scooped on half my argument was truly a very good day. I spent a day deciding whether or not this paper undercut what I’ve been constructing. Seems that it truly does the other. It removes a declare I used to be making based mostly on restricted analysis and replaces it with one anyone measured, which leaves me arguing for one factor as an alternative of two. A smaller declare that’s nonetheless standing is price greater than a big one no one has examined.
Impartial Spec-Pushed Growth (ISDD)
In abstract, the evolution I’m proposing is Impartial Spec-Pushed Growth, ISDD: the identical specification, divided in order that whoever writes the code can’t learn the half that judges it.

Whether or not that division catches extra actual defects is the open query, and it’s what qikly is constructed to check. If of labor that measures it, I’d be eager to learn it.
Earlier articles on this sequence
-
Half 1 explains my argument intimately and claims “the inexperienced take a look at suite means nothing”
-
Half 2 walks by a full run of qikly
Concerning the Writer
Gal Arav is the creator of Utilized Statistics for Knowledge Science and maintains qikly, a brand new open-source Python venture for spec-driven take a look at automation.
















