The Transaction That Succeeds: Why 98% Test Pass Rates Hide Catastrophic Production Failure
If you run engineering or finance at a company building with AI agents, you have almost certainly sat through a demo where a technical lead showed off a 98% evaluation pass rate.
The room claps. The VP of Product smiles. The CEO thinks they are weeks away from cutting support overhead in half.
Then you deploy the agent to real customers on a Friday afternoon.
By Sunday evening, you have an emergency incident channel with fifteen people in it, the CFO is asking why your Stripe refund ledger has thirty unexplained credits totaling $142,000, and senior engineers are manually auditing database rows line by line.
What went wrong?
The answer is simple: your team tested reading, but production runs on writing.
In traditional software engineering, a 98% test pass rate is respectable. If a deterministic function passes 98 tests out of 100, the remaining two fail loudly with an error code, a stack trace, and a clean rollback. Nothing touches the database.
In probabilistic agent architectures, failure does not look like an error code. It looks like a successful transaction that should never have happened.
When an autonomous customer service or sales agent hits a corner case (an expired token, an ambiguous policy clause, or a user phrasing a request in slightly broken English), it does not throw an exception. It does not walk down the hall to ask your Director of Finance for permission.
It guesses.
It picks the closest plausible tool parameter, generates a valid JSON payload, and calls your internal API. Your internal API validates the schema, sees valid parameters, and commits the write to your PostgreSQL database with HTTP 200 OK. The transaction succeeded. The business policy was violated.
At 50,000 automated customer interactions a month, a 98% success rate means exactly 1,000 unverified production mutations happening behind your back every four weeks.
To survive production deployment without blowing up your company's balance sheet, engineering organizations must enforce the 4 Pillars of Agent Governance:
1. Decouple Inference from State Authority: Never allow an LLM or reasoning loop to hold credentials that write directly to production tables. The model produces a candidate mutation; a deterministic execution gateway evaluates the business invariant before committing the transaction.
2. Strict Admissibility Allowlists: Tools exposed to agents should operate on bounded enumerations, not open-ended SQL or arbitrary JSON blobs. If a tool call parameter falls outside verified operational bounds, the execution engine drops it instantly for zero cost.
3. Two-Phase Micro-Reversibility: Every agent-initiated mutation must include an automated compensation transaction. If a multi-step agent workflow fails on step four, steps one through three must unwind deterministically rather than leaving orphan state in your database.
4. The Air Traffic Control Queue: Any transaction exceeding a defined dollar threshold or affecting core accounting ledgers must route to a supervisory human review queue. If your agent is processing enterprise transactions, a Customer Support Manager or Product Ops Lead must retain the ultimate kill switch.
Observability alone is not governance. Watching an autonomous agent issue improper refunds in an OpenTelemetry dashboard is just watching your cash burn in real time. Governance is having a deterministic gate that prevents the transaction from committing in the first place.
Richard Ewing writes on technology, systems economics, and runtime governance as The AI Economist. He is the founder of Exogram.ai (deterministic runtime firewalls for AI) and CareerWin.ai.