Building reliable systems when the AI is not reliable
Production engineering should assume model output can be wrong.
Reliability comes from the surrounding system: narrow roles, deterministic verification, independent review, evidence-based gates, and feedback.
Verification · 2026-10-06 · 1 min read · Draft
Model confidence is not system reliability
How sure a model sounds says little about whether its output is correct.
Narrow roles reduce ambiguity
An agent with one clearly defined job is easier to evaluate than one asked to do everything.
Deterministic tools should verify what they can
Compilers, tests, linters, and schema validators check what they can check without relying on a model.
Independent agents can review generated work
A separate agent with a different role can review work it did not produce.
Gates should require evidence
A check should report what it found and where, so that a pass or fail can be inspected.
Failures should improve future checks
Misses and false positives are inputs for refining rules and agents.
Human escalation remains part of the system
Some decisions require judgment, and the system should route them to people.
Frequently asked questions
Production engineering should assume model output can be wrong. Reliability comes from the surrounding system, not from assuming perfect model behavior.
Narrow roles, deterministic checks, independent review, evidence-based gates, feedback from failures, and human escalation.
Yes. Compilers, tests, linters, and schema validators verify what they can without relying on a model.