The hidden cost of "it passed the eval"
A green eval suite is necessary but never sufficient. TAFI's platform verification caught five failure classes that a green score alone would have missed.

A team ships an AI feature. The eval suite is green. Score: 100. Confidence: high. Two weeks later, production surfaces a class of failure the eval never modelled. The eval was not wrong — it was incomplete. The gap between "passes the suite" and "safe in production" is where silent failures live, and no scoreboard closes it automatically.
We learned this building TAFI, the AI co-builder for the TechAppForce platform. Its certification pipeline runs 279 eval stories across 13 stages (E0–E12), and three consecutive clean internal runs scored QA 100. But the architecture-level verification layer — contract checkers, drift detectors, session-id propagation — caught five failure classes that the eval score alone would have reported as "green."
The five failure classes
The eval suite validates functional success: did the planner pick the right skill? Did the tool return the expected shape? Platform verification validates structural integrity: did the system uphold its own invariants under real execution?
| Failure class | What the eval saw | What platform verification caught |
|---|---|---|
| Silent-success bug (2026-05-13) | Score 100, all stories green | mcp_call_count=0 — the planner hallucinated a completed run with zero MCP calls |
| Contract violation — pending IDs | Query steps marked success | db_register_query returned pending (id=null); backend materialised IDs 5–10 s later; contract_verifier correctly rejected fabricated IDs |
| LLM-side shape drift | Artifact-only skill produced output | Pydantic invariant rejected non-empty tool_calls on an artifact-only skill (screen_validator) |
| Counter drift | Step counts matched | Drift detector found delta=-6 (TAFI ledger 26 vs MCP 32) — retries inflated HTTP arrivals, not logical steps |
| Async crash in verification | Drift detector "passed" in dry-run | asyncio.run() inside running loop crashed the detector mid-cert; would have emitted false-positive CRITICAL |
A green eval suite proves the happy path works. Platform verification proves the system cannot lie about its own execution.
Why the eval missed them
An eval suite is a specification of expected behaviour. It encodes what you thought to check. Platform verification encodes invariants that must hold regardless of what you thought to check.
The silent-success bug (cert cli-1778695014) is the clearest example: every story passed, the score was 100, but the planner had produced zero MCP calls. The eval checked story outcomes; it did not check execution reality. The drift detector and session-id propagation (E12.F5/F5b) were built precisely because that class of failure is invisible to outcome-based evals.
The query_designer pending-ID failures (cert cli-1778832909, steps 2–3) show the same pattern. The eval marked the steps green because the planner emitted a query object. The contract_verifier — reading the skill manifest’s output.produces — saw query_id: null and rejected it. Rule 1 of id_resolution.md held: NEVER invent UUIDs.
The screen_validator shape drift (cert cli-1778794811, failure 4) was an LLM emitting tool_calls: [...] for a skill declared requires_mcp: false. The eval would have accepted the produced artifact. Pydantic rejected the envelope. The fix was a one-line backstory edit across five artifact-only skills — mechanical, but the eval would never have prompted it.
Two-layer verification in practice
TAFI now runs both layers on every certification:
| Layer | Scope | Tooling | Gate |
|---|---|---|---|
| Eval | 279 stories, functional outcomes | v4_runner.py --v4, plan_evaluator.py |
QA score ≥ 88% |
| Platform | Contracts, drift, invariants, session integrity | contract_verifier, drift detector, X-TAFI-Session-Id propagation |
Zero CRITICAL, zero fabricated IDs, drift WARN only |
The platform layer is not a second eval. It is a set of invariant checkers that run on the same execution traces. They do not ask "did the feature work?" — they ask "did the system obey its own rules while working?"
See the E12 gate report and verification architecture on the TAFI product page. To discuss how two-layer verification fits your AI stack, book a call at cal.com/agentix-tech or reach us via /contact.

