One-line lessons from the field, each linked to the system that proves it.
LOG-001Eval engineeringJune 2026
A passing eval score is not a shipped system.
On TAFI, the second certification run returned QA 100, and platform verification still found five failure classes the eval harness could not see: queries dead-lettered, empty screen containers, workflow commands with no consumer, navigation records missing required columns, and no audit trail. We fixed all five during E13, the fidelity-hardening phase that is still in progress, then ran three consecutive clean internal pipelines. The lesson we carry into every build: an eval number is necessary, never sufficient. The gate has to verify the real platform state, not just the score.
Run deterministic grounding before the model reasons.
VistaGPT, which is in development with internal validation only, is designed to limit hallucination by structure rather than prompt-crafting. Six parallel deterministic ChromaDB stages retrieve evidence with zero LLM involvement, a Reconcile stage drops any element the schema can't confirm, and a sqlglot validator with two auto-retries gates the generated SQL before it runs. This is design intent, not a published accuracy result. Grounding is an architecture decision, not a system-prompt paragraph.
Read-only should be enforced by the framework, not a policy.
Orions AI investigates production incidents across StubHub Pro audit logs, per-tenant MongoDB, and source code, and it cannot write to any system under diagnosis. That guarantee lives at the Step Engine level: no agent instruction can trigger a write, and per-tenant resolution makes cross-tenant access structurally impossible (verified by integration tests). "Safe AI" is not a prompt. It is a boundary the model has no path around.
When Orions moved from a 4-agent linear pipeline to a checkpoint-driven open-world loop, v2 was built to run in shadow mode beside v1 behind a DIAGNOSIS_V2 flag, off by default, with results compared at a single endpoint before any traffic shifts. The production default is still the linear pipeline. A regression gate scores changes against 7 golden incidents; it is run manually today. You earn the cutover with evidence; you don't announce it.