Build the eval harness before the feature
Why we build the eval harness before the feature, and why a green eval is necessary but never sufficient. What TAFI's QA 100 taught us.

When a team sets out to build an AI feature, the eval usually arrives last. The feature gets built, someone demos it, and only then does anyone write a test to confirm it works. We run it the other way around: the eval harness is the first thing we build, before the agent, before the pipeline, before a single artefact is generated.
The reason is practical. An AI system has no compiler error to catch a wrong answer. A model returns fluent, confident output whether it is right or wrong, and the failure modes are silent. If you cannot score the output before you start, you cannot tell whether a change made the system better or worse.
The eval is the specification, made executable
Writing the eval first forces a decision that is easy to defer: what does correct actually mean here? On TAFI, our natural-language app builder for the TechAppForce platform, correct is not “the model produced plausible JSON.” It is “every entity, screen, query and workflow the user described is registered on the live platform, in the right dependency order, with no fabricated identifiers.” You cannot write that eval without first agreeing on the definition.
The eval is the specification, made executable. You cannot grade an output until you have decided, precisely, what correct means.
A score is a measurement, not a guarantee
Here is the part that took a real incident to learn. On TAFI the eval is a deterministic harness: Score = 100 − Σ penalties across every story in the suite. A certification run on an Equipment Tracker app came back QA 100. Every story passed. By the usual bar, that ships.
It did not ship. A second layer, platform verification, checked the build against the real, live state of TechAppForce and found five failure classes the score was blind to.
Each of these is invisible to an eval that grades the generated output. The output looked correct; the platform state was broken.
| Failure class Gate 2 caught | Why the score was blind to it |
|---|---|
| Queries dead-lettered in the message bus | The query was well-formed; the bus still rejected it |
| Screen-container JSON rendering empty | Valid JSON, nothing rendered |
| Workflow bus commands with no TAF consumer | Command shape was correct; nobody was listening |
| Navigation records missing required columns | Passed generation, failed at the platform |
| No audit trail on artefact generation | Invisible to output-level grading entirely |
Necessary, never sufficient
We fixed all five in one epic, then required three consecutive clean runs before the build was certified. The rule we have carried into every project since: a green eval is necessary but never sufficient. The five failure classes did not just get fixed in the product — they became assertions in the harness. The next build cannot regress on them, because the eval now knows to look.
An eval is not a fixed artefact you write once. It is a growing record of every way the system has been caught being wrong.
The uncomfortable version of this lesson: the more impressive your eval score, the more carefully you should ask what it cannot see. QA 100 is exactly the moment to run the second gate, because a perfect score is the most persuasive way for a system to be wrong.
See how the two-gate discipline plays out end to end in the TAFI case file, or tell us what you are building at cal.com/agentix-tech.

