Blog→AI Governance
AI Governance7 min read read

The Transaction That Succeeds: Why 98% Test Pass Rates Hide Catastrophic Production Failure

In probabilistic computing, a 98% success rate sounds great to an Engineering Manager. To a CFO, it means 2 out of every 100 transactions are unmonitored capital leakage...

By Richard Ewing·
Share:

The Transaction That Succeeds: Why 98% Test Pass Rates Hide Catastrophic Production Failure

If you run engineering or finance at a company building with AI agents, you have almost certainly sat through a demo where a technical lead showed off a 98% evaluation pass rate.

The room claps. The VP of Product smiles. The CEO thinks they are weeks away from cutting support overhead in half.

Then you deploy the agent to real customers on a Friday afternoon.

By Sunday evening, you have an emergency incident channel with fifteen people in it, the CFO is asking why your Stripe refund ledger has thirty unexplained credits totaling $142,000, and senior engineers are manually auditing database rows line by line.

What went wrong?

The answer is simple: your team tested reading, but production runs on writing.

In traditional software engineering, a 98% test pass rate is respectable. If a deterministic function passes 98 tests out of 100, the remaining two fail loudly with an error code, a stack trace, and a clean rollback. Nothing touches the database.

In probabilistic agent architectures, failure does not look like an error code. It looks like a successful transaction that should never have happened.

When an autonomous customer service or sales agent hits a corner case (an expired token, an ambiguous policy clause, or a user phrasing a request in slightly broken English), it does not throw an exception. It does not walk down the hall to ask your Director of Finance for permission.

It guesses.

It picks the closest plausible tool parameter, generates a valid JSON payload, and calls your internal API. Your internal API validates the schema, sees valid parameters, and commits the write to your PostgreSQL database with HTTP 200 OK. The transaction succeeded. The business policy was violated.

At 50,000 automated customer interactions a month, a 98% success rate means exactly 1,000 unverified production mutations happening behind your back every four weeks.

To survive production deployment without blowing up your company's balance sheet, engineering organizations must enforce the 4 Pillars of Agent Governance:

1. Decouple Inference from State Authority: Never allow an LLM or reasoning loop to hold credentials that write directly to production tables. The model produces a candidate mutation; a deterministic execution gateway evaluates the business invariant before committing the transaction.

2. Strict Admissibility Allowlists: Tools exposed to agents should operate on bounded enumerations, not open-ended SQL or arbitrary JSON blobs. If a tool call parameter falls outside verified operational bounds, the execution engine drops it instantly for zero cost.

3. Two-Phase Micro-Reversibility: Every agent-initiated mutation must include an automated compensation transaction. If a multi-step agent workflow fails on step four, steps one through three must unwind deterministically rather than leaving orphan state in your database.

4. The Air Traffic Control Queue: Any transaction exceeding a defined dollar threshold or affecting core accounting ledgers must route to a supervisory human review queue. If your agent is processing enterprise transactions, a Customer Support Manager or Product Ops Lead must retain the ultimate kill switch.

Observability alone is not governance. Watching an autonomous agent issue improper refunds in an OpenTelemetry dashboard is just watching your cash burn in real time. Governance is having a deterministic gate that prevents the transaction from committing in the first place.

Richard Ewing writes on technology, systems economics, and runtime governance as The AI Economist. He is the founder of Exogram.ai (deterministic runtime firewalls for AI) and CareerWin.ai.

Like this analysis?

Get the weekly engineering economics briefing - one email, every Monday.

Subscribe Free →

More in AI Governance

Related Canonical Concepts

The Systems Governor

A dedicated enterprise role accountable for governing the boundary between what autonomous AI agents propose and what an organization permits them to execute. Reporting directly to the CIO or CEO, the Systems Governor maintains permission allowlists, sets state integrity thresholds, owns the cryptographic audit trail, and translates technical agent error rates into financial liability metrics.

Read Concept →

Deterministic Execution Control

An execution governance architecture formulated by Richard Ewing that enforces hard, cryptographically verified boundary constraints between probabilistic AI models and production enterprise infrastructure. Deterministic Execution Control dictates that probabilistic models are never permitted to execute state-mutating operations (database writes, financial transactions, credential deletions) directly; all operations must pass through deterministic schema allowlists, pre-execution assertions, and rollback ledgers.

Read Concept →

The Transaction That Succeeds

An enterprise AI governance failure mode formulated by Richard Ewing in CIO.com where an automated agent transaction completes with perfect technical execution (glowing green operations dashboards, 240ms latency, zero server errors), but completely violates internal business policy, financial controls, or procurement rules. Examples include automated support agents issuing unapproved corporate credits, procurement agents bypassing $50,000 competitive bid mandates, or sales agents altering contract terms that destroy gross margin. Because technical monitoring verifies mechanics rather than business authorization, enterprises must enforce the 4 Pillars of Agent Governance: Monitoring, Auditability, Authorization, and Accountability.

Read Concept →

Canonical Frameworks

Technical Insolvency Date

The Technical Insolvency Date (TID) is the specific future quarter when an organization's technical debt maintenance will consume 100% of engineering capacity, leaving zero time for new feature development. Every software organization accumulates technical debt over time - shortcuts taken under deadline pressure, aging infrastructure, deprecated dependencies, and code that nobody understands anymore. This debt isn't free. It requires ongoing maintenance hours: bug fixes, security patches, dependency updates, and workarounds for architectural limitations. The critical insight is that maintenance burden grows faster than most leaders realize. If your team currently spends 40% of its time on maintenance and that percentage is growing 3% per quarter, you can calculate the exact quarter when maintenance reaches 100%. That quarter is your Technical Insolvency Date. At the TID, your engineering team is fully consumed by keeping existing systems alive. Feature velocity drops to zero. No new capabilities. No competitive response. No innovation. Your R&D investment becomes pure maintenance spend - you're paying innovation-era salaries for maintenance-era output. The concept draws from financial insolvency: the point where a company's liabilities exceed its assets and it cannot meet its obligations. Technical insolvency is the same idea applied to engineering capacity - the point where your maintenance obligations exceed your available engineering hours. Most organizations don't realize they're approaching the TID because they track technical debt qualitatively rather than quantitatively. Telling a board "we have technical debt" gets deprioritized. Telling a board "we are 8 quarters from technical insolvency - the point where we can no longer ship any new features" gets immediate action and budget allocation.

Read Definition →

Audit Interview

The Audit Interview is a hiring protocol that tests verification skills instead of code generation skills. In the AI age, the scarce human skill is not writing code - it's catching what AI gets wrong. Traditional coding interviews ask candidates to write algorithms on a whiteboard or in a shared editor. This was a reasonable proxy for engineering skill when humans wrote all the code. But in 2026, AI tools like GitHub Copilot, Cursor, and Claude generate code faster and often more correctly than human candidates under interview pressure. When Anthropic discovered that candidates were using Claude to pass their own coding interviews, it proved that traditional interviews are testing the wrong thing. They're testing a skill that AI performs better than humans under artificial conditions. The Audit Interview flips the model. Instead of asking candidates to generate code, it presents them with AI-generated code that contains hidden flaws - security vulnerabilities, logic errors, performance anti-patterns, edge case failures, and architectural problems. The candidate's job is to find the bugs, rank them by severity, and make a ship/no-ship recommendation. The protocol works like this: candidates receive a realistic code review scenario (500-1000 lines of AI-generated code with 3-5 hidden flaws). They have 10 minutes to review the code, identify issues, and present their findings. The evaluation scores 4 dimensions of engineering judgment: 1. Verification: How many bugs did they find? Did they catch the security vulnerability? 2. Prioritization: Did they correctly rank issues by severity? 3. Communication: Can they explain the risk to a non-technical stakeholder? 4. Judgment: Would they ship this code? Under what conditions? With what caveats? The free Audit Interview tool at richardewing.io/tools/audit-interview generates realistic AI-written code with calibrated flaws for interviewers to use immediately.

Read Definition →
📊

Richard Ewing

The AI Economist - Quantifying engineering economics for technology leaders, PE firms, and boards.

⚡

Want to apply this to your organization?

Run a free diagnostic first. If the numbers concern you, book a session to build a remediation plan.

Richard Ewing: AI Economist & Capital Auditor