AGENTIX · ENGINEERING LOGEVAL ENGINEERING / 2026
Eval engineering5 min read

Build the eval harness before the feature

Why we build the eval harness before the feature, and why a green eval is necessary but never sufficient. What TAFI's QA 100 taught us.

Two-gate release discipline: an eval harness and platform verification must both agree before a build ships

When a team sets out to build an AI feature, the eval usually arrives last. The feature gets built, someone demos it, and only then does anyone write a test to confirm it works. We run it the other way around: the eval harness is the first thing we build, before the agent, before the pipeline, before a single artefact is generated.

The reason is practical. An AI system has no compiler error to catch a wrong answer. A model returns fluent, confident output whether it is right or wrong, and the failure modes are silent. If you cannot score the output before you start, you cannot tell whether a change made the system better or worse.

100QA score, every story passed
5failure classes the score missed
3clean runs required to certify

The eval is the specification, made executable

Writing the eval first forces a decision that is easy to defer: what does correct actually mean here? On TAFI, our natural-language app builder for the TechAppForce platform, correct is not “the model produced plausible JSON.” It is “every entity, screen, query and workflow the user described is registered on the live platform, in the right dependency order, with no fabricated identifiers.” You cannot write that eval without first agreeing on the definition.

The eval is the specification, made executable. You cannot grade an output until you have decided, precisely, what correct means.

A score is a measurement, not a guarantee

Here is the part that took a real incident to learn. On TAFI the eval is a deterministic harness: Score = 100 − Σ penalties across every story in the suite. A certification run on an Equipment Tracker app came back QA 100. Every story passed. By the usual bar, that ships.

It did not ship. A second layer, platform verification, checked the build against the real, live state of TechAppForce and found five failure classes the score was blind to.

Generated build GATE 1 Eval harness score = 100 GATE 2 Platform verify vs live state Ship Gate 2 caught 5 failure classes Gate 1 could not see Ships only when both gates agree.
A perfect eval score is a floor, not a ceiling. The second gate checks what the score cannot see.

Each of these is invisible to an eval that grades the generated output. The output looked correct; the platform state was broken.

Failure class Gate 2 caught Why the score was blind to it
Queries dead-lettered in the message bus The query was well-formed; the bus still rejected it
Screen-container JSON rendering empty Valid JSON, nothing rendered
Workflow bus commands with no TAF consumer Command shape was correct; nobody was listening
Navigation records missing required columns Passed generation, failed at the platform
No audit trail on artefact generation Invisible to output-level grading entirely

Necessary, never sufficient

We fixed all five in one epic, then required three consecutive clean runs before the build was certified. The rule we have carried into every project since: a green eval is necessary but never sufficient. The five failure classes did not just get fixed in the product — they became assertions in the harness. The next build cannot regress on them, because the eval now knows to look.

An eval is not a fixed artefact you write once. It is a growing record of every way the system has been caught being wrong.

The uncomfortable version of this lesson: the more impressive your eval score, the more carefully you should ask what it cannot see. QA 100 is exactly the moment to run the second gate, because a perfect score is the most persuasive way for a system to be wrong.

See how the two-gate discipline plays out end to end in the TAFI case file, or tell us what you are building at cal.com/agentix-tech.

Work with us

Building a system
that has to hold up?

Tell us what you're building. We'll tell you how we'd architect it, what the eval harness would cover, and what production deployment involves.

AGENTIX TECH · engineering log · Eval engineering