AGENTIX · ENGINEERING LOGFIELD NOTES / 2026
§ Insights

Insights
& Evidence

No think-pieces. Engineering writing from the production builds behind our work, each argument tied to a system that earned it.

Architecture6 min read

The hidden cost of "it passed the eval"

A green eval suite is necessary but never sufficient. TAFI's platform verification caught five failure classes that a green score alone would have missed.

Read the note →
Product design5 min read

One hub for every Vista desktop tool

How Vista Studio replaces three separate installer paths with one Windows hub for installing, checking and updating the Vista app suite.

Read the note →
Developer tooling5 min read

Measure AI leverage without the spyware

How Stack Signal reads your local AI coding logs into one private dashboard: 45 anti-pattern rules across five harnesses, zero telemetry in the core.

Read the note →
Delivery5 min read

Architecture-first AI, in plain English

A buyer's guide to the six-stage Agentix delivery method: define the work, approve the design, test the system, then deploy.

Read the note →
Grounding5 min read

The assistant factory that cites its sources

How Autoagentix Chat reads a business's system four ways, cites every fact to file and line, and hands over a grounded assistant the business owns.

Read the note →
Architecture5 min read

Provider-agnostic by design

Why we don't bet a codebase on one model vendor: how Autoagentix Dev detects the coding agent you already run and upgrades it, contracts-first.

Read the note →
Eval engineering5 min read

Build the eval harness before the feature

Why we build the eval harness before the feature, and why a green eval is necessary but never sufficient. What TAFI's QA 100 taught us.

Read the note →
Architecture5 min read

Route every agent through typed MCP gateways

Route external reads and writes through a typed MCP gateway instead of direct API calls. Governance, audit and identity live in the boundary.

Read the note →
Grounding5 min read

Context engineering beats prompt engineering

Deterministic grounding before the model reasons: VistaGPT constrains every query to a verified schema so the model cannot invent tables.

Read the note →
Production4 min read

Ship agents in shadow mode first

How an agent earns a production cutover with evidence, not announcements: Orions v2 in shadow mode, gated by a precision and calibration check.

Read the note →
§ Quick notes

One-line lessons from the field, each linked to the system that proves it.

LOG-001Eval engineeringJune 2026

A passing eval score is not a shipped system.

On TAFI, the second certification run returned QA 100, and platform verification still found five failure classes the eval harness could not see: queries dead-lettered, empty screen containers, workflow commands with no consumer, navigation records missing required columns, and no audit trail. We fixed all five during E13, the fidelity-hardening phase that is still in progress, then ran three consecutive clean internal pipelines. The lesson we carry into every build: an eval number is necessary, never sufficient. The gate has to verify the real platform state, not just the score.

Read the TAFI case file →
LOG-002GroundingMay 2026

Run deterministic grounding before the model reasons.

VistaGPT, which is in development with internal validation only, is designed to limit hallucination by structure rather than prompt-crafting. Six parallel deterministic ChromaDB stages retrieve evidence with zero LLM involvement, a Reconcile stage drops any element the schema can't confirm, and a sqlglot validator with two auto-retries gates the generated SQL before it runs. This is design intent, not a published accuracy result. Grounding is an architecture decision, not a system-prompt paragraph.

Read the VistaGPT case file →
LOG-003Safety architectureApril 2026

Read-only should be enforced by the framework, not a policy.

Orions AI investigates production incidents across StubHub Pro audit logs, per-tenant MongoDB, and source code, and it cannot write to any system under diagnosis. That guarantee lives at the Step Engine level: no agent instruction can trigger a write, and per-tenant resolution makes cross-tenant access structurally impossible (verified by integration tests). "Safe AI" is not a prompt. It is a boundary the model has no path around.

Read the Orions case file →
LOG-004MigrationApril 2026

Ship new pipeline versions in shadow mode first.

When Orions moved from a 4-agent linear pipeline to a checkpoint-driven open-world loop, v2 was built to run in shadow mode beside v1 behind a DIAGNOSIS_V2 flag, off by default, with results compared at a single endpoint before any traffic shifts. The production default is still the linear pipeline. A regression gate scores changes against 7 golden incidents; it is run manually today. You earn the cutover with evidence; you don't announce it.

How we work →
Work with us

Have a system that
needs to be provable?

Tell us what you're building. We'll tell you how we'd architect it, what the eval harness would cover, and what production deployment involves.