Blog→AI Governance
AI Governance3 min read read

How to Reduce LLM API Token Costs in Production

Caching Math, Architecture, and Margin Protection When building AI software, every API call to OpenAI, Anthropic, or Google costs real money. If your application relies on multi-agent execution loops,...

By Richard Ewing·
Share:

How to Reduce LLM API Token Costs in Production

Caching Math, Architecture, and Margin Protection

When building AI software, every API call to OpenAI, Anthropic, or Google costs real money. If your application relies on multi-agent execution loops, a single user click can easily trigger 10 to 20 separate model calls.

Without a deliberate caching strategy, your cloud bill will quickly erode your unit economics.

This research note breaks down how we structured semantic caching inside Exogram.ai to cut token spend by over 50% while dramatically improving response speed for end users.

Traditional web caching matches exact URLs or exact database keys, but AI users ask questions using different words for the exact same intent.

Semantic caching generates a vector representation of the incoming query and checks whether a mathematically similar query already exists inside the cache database.

Cosine similarity scores range from 0.0 to 1.0.

0.95+ Threshold: Extremely safe for factual, transactional queries, carrying virtually zero risk of serving an irrelevant answer.

0.88 - 0.92 Threshold: Good for general intent classification, though it requires ongoing monitoring for subtle nuance loss.

Rule: Start at 0.94 and tune downward based on real user feedback.

AI context changes over time, so you should never store semantic cache entries indefinitely.

Static Context (System Rules, FAQs): 30-day TTL.

Dynamic Context (User Career Profiles, Active Workflows): 1-hour to 24-hour TTL.

Real-time State Checks: Do not cache, forcing live deterministic evaluations every time.

During an automated attack or bot scrape, malicious traffic can trigger thousands of LLM API calls per minute. Placing Cloudflare Turnstile and strict edge rate-limiting in front of your caching layer ensures that bot traffic is dropped before touching either your cache database or your LLM endpoints.

By deploying semantic caching and edge pre-filtering across our runtime:

Average Cache Hit Rate: Settled at roughly 38% across routine agent execution loops.

Cost per 1,000 Requests: Dropped from $4.20 down to $1.85.

User-Perceived Speed: Cache hits feel instantaneous, consistently returning under 20ms.

This economic buffer is what allows us to run infrastructure for Exogram.ai and CareerWin.ai efficiently without sacrificing capability or performance.

Built In: Is Anything Standing Between Your AI Agent and Your Database? (https://builtin.com/articles/ai-agent-security-gates)

Systems Infrastructure: Learn how Exogram.ai enforces runtime security and how CareerWin.ai applies context systems to career intelligence.

Join the list to receive our newest posts straight to your inbox.

Enter your email to access the Product Debt Index Calculator, and join the private ledger where I break down the exact strategies bridging technical execution with financial viability for B2B SaaS.

Home

Archive

Login

Reset Password

Update Password

Profile

Search

Like this analysis?

Get the weekly engineering economics briefing - one email, every Monday.

Subscribe Free →

More in AI Governance

Canonical Frameworks

Technical Insolvency Date

The Technical Insolvency Date (TID) is the specific future quarter when an organization's technical debt maintenance will consume 100% of engineering capacity, leaving zero time for new feature development. Every software organization accumulates technical debt over time - shortcuts taken under deadline pressure, aging infrastructure, deprecated dependencies, and code that nobody understands anymore. This debt isn't free. It requires ongoing maintenance hours: bug fixes, security patches, dependency updates, and workarounds for architectural limitations. The critical insight is that maintenance burden grows faster than most leaders realize. If your team currently spends 40% of its time on maintenance and that percentage is growing 3% per quarter, you can calculate the exact quarter when maintenance reaches 100%. That quarter is your Technical Insolvency Date. At the TID, your engineering team is fully consumed by keeping existing systems alive. Feature velocity drops to zero. No new capabilities. No competitive response. No innovation. Your R&D investment becomes pure maintenance spend - you're paying innovation-era salaries for maintenance-era output. The concept draws from financial insolvency: the point where a company's liabilities exceed its assets and it cannot meet its obligations. Technical insolvency is the same idea applied to engineering capacity - the point where your maintenance obligations exceed your available engineering hours. Most organizations don't realize they're approaching the TID because they track technical debt qualitatively rather than quantitatively. Telling a board "we have technical debt" gets deprioritized. Telling a board "we are 8 quarters from technical insolvency - the point where we can no longer ship any new features" gets immediate action and budget allocation.

Read Definition →

Audit Interview

The Audit Interview is a hiring protocol that tests verification skills instead of code generation skills. In the AI age, the scarce human skill is not writing code - it's catching what AI gets wrong. Traditional coding interviews ask candidates to write algorithms on a whiteboard or in a shared editor. This was a reasonable proxy for engineering skill when humans wrote all the code. But in 2026, AI tools like GitHub Copilot, Cursor, and Claude generate code faster and often more correctly than human candidates under interview pressure. When Anthropic discovered that candidates were using Claude to pass their own coding interviews, it proved that traditional interviews are testing the wrong thing. They're testing a skill that AI performs better than humans under artificial conditions. The Audit Interview flips the model. Instead of asking candidates to generate code, it presents them with AI-generated code that contains hidden flaws - security vulnerabilities, logic errors, performance anti-patterns, edge case failures, and architectural problems. The candidate's job is to find the bugs, rank them by severity, and make a ship/no-ship recommendation. The protocol works like this: candidates receive a realistic code review scenario (500-1000 lines of AI-generated code with 3-5 hidden flaws). They have 10 minutes to review the code, identify issues, and present their findings. The evaluation scores 4 dimensions of engineering judgment: 1. Verification: How many bugs did they find? Did they catch the security vulnerability? 2. Prioritization: Did they correctly rank issues by severity? 3. Communication: Can they explain the risk to a non-technical stakeholder? 4. Judgment: Would they ship this code? Under what conditions? With what caveats? The free Audit Interview tool at richardewing.io/tools/audit-interview generates realistic AI-written code with calibrated flaws for interviewers to use immediately.

Read Definition →
📊

Richard Ewing

The AI Economist - Quantifying engineering economics for technology leaders, PE firms, and boards.

⚡

Want to apply this to your organization?

Run a free diagnostic first. If the numbers concern you, book a session to build a remediation plan.

Richard Ewing: AI Economist & Capital Auditor