AGENTIX · ENGINEERING LOGPRODUCTION / 2026
Production4 min read

Ship agents in shadow mode first

How an agent earns a production cutover with evidence, not announcements: Orions v2 in shadow mode, gated by a precision and calibration check.

Shadow mode: a candidate agent runs beside production, scored against a CI gate, before any cutover

The most common way an AI system reaches production is an announcement. Someone decides it is ready, a flag flips, and the new behaviour goes live to real users. If it is worse than what it replaced, you find out from the users. This is a poor way to run anything that matters, and an especially poor way to run a system whose failures are quiet.

There is a better pattern, borrowed from how careful teams ship any risky change: run the new thing in shadow, beside the old thing, on real traffic, and let it earn the cutover with evidence.

7golden fixtures gate every merge
~200fixed PRs as ground truth
~80%proven verdict, 0 unsupported

What shadow mode looked like for Orions

Orions diagnoses production incidents for a busy marketplace. When we rebuilt its reasoning from a linear four-agent pipeline into a checkpoint-driven loop that holds competing hypotheses, the rebuild changed the most sensitive part of the system: the part that decides what caused an outage.

So the new pipeline went live behind a DIAGNOSIS_V2 flag and ran in shadow against production traffic. Every real incident was diagnosed by both versions. The v2 output was scored and stored, never posted.

Production incidents v1 — live serves users v2 — shadow scored, stored, silent CI gate precision drop < 5 pts calibration err < 0.15 cut over
Users keep the current behaviour. The candidate earns the cutover on evidence already gathered.

The gate that makes shadow mode mean something

Shadow mode only works if “better” is defined precisely enough to block a regression automatically. Watching two outputs side by side is not a gate; it is a vibe. The gate has to live in CI, and it has to say no on its own.

A system that is confidently wrong is worse than one that is uncertain and says so. The gate checks precision and calibration, and it runs before the code lands, not after.

For Orions, every merge request that touches diagnosis or prompt files is scored against seven golden YAML fixtures — real incidents with known root causes — and the merge is blocked if precision drops more than five points against the v1 baseline, or if the expected calibration error exceeds 0.15. The ground-truth set is drawn from around two hundred already-fixed production PRs, and the judge is a second, independent model with tools disabled and separate billing, so the engine can never grade its own homework.

Why this order is the honest one

The sequence is: build the candidate, run it in shadow on real traffic, grade it against known-correct outcomes with an independent judge, gate the merge on precision and calibration, and only then cut over. Each step produces evidence the next can check. By the time traffic shifts, the decision has already been made by the data.

See the shadow-mode and CI-gate details in the Orions case file, or tell us what you are shipping at cal.com/agentix-tech.

Work with us

Building a system
that has to hold up?

Tell us what you're building. We'll tell you how we'd architect it, what the eval harness would cover, and what production deployment involves.

AGENTIX TECH · engineering log · Production