Compound AI Systems
An architectural approach that builds AI applications using multiple interconnected models, deterministic tools, and external memory.
“The future of AI is not a bigger brain in a jar; it is a highly coordinated assembly line of specialized cognitive tools.”
Relying on a single, massive frontier model for all tasks is economically ruinous and architecturally fragile. It leads to high latency, exorbitant costs, and a single point of failure. Compound AI Systems allow organizations to optimize for cost, speed, and accuracy simultaneously. By breaking down complex tasks into specialized, deterministic workflows guided by smaller, purpose-built models, architects can build highly resilient applications that do not depend entirely on the shifting capabilities of one vendor's API.
Multi-Hop Causal Traversal Engine
Explore how concepts dynamically feed into each other across 1-hop, 2-hop, and 3-hop transitive relationships. Click any node to navigate the causal highway.
Compound AI Systems
An architectural approach that builds AI applications using multiple interconnected models, deterministic tools, and external memory.
Direct Relationships (4)
Transitive Neighbors (Connected via Hop 1)
Extended Causal Ripple Effects
Richard Ewing’s Research Thesis
Do not worship the model. Engineer the system. The orchestration of components is more valuable than the parameter count.
Why This Specification Exists
Single frontier models are too slow, expensive, and fragile for complex enterprise applications.
Building simple wrapper apps around one large LLM.
Monolithic models fail at deterministic routing and specialized sub-tasks.
Orchestrating specialized small models, vector databases, and deterministic state machines.
What Changes If You Believe This?
Architecture shifts to dynamic routing and component orchestration.
Massive reduction in API costs by routing simple queries to small models.
Lower latency improves user experience.
Reduced dependency on a single external vendor API.
Recommended Action by Role
Stop routing every internal request to massive frontier models; orchestrate small, fast models for categorization and reserve expensive models for heavy reasoning.
Improve application responsiveness by breaking single monolithic prompts into discrete steps that return instant feedback to users.
Reduce vendor lock-in and cut downtime by building multi-model pipelines that seamlessly failover to alternative providers.
Map product user flows into deterministic state flows so models execute predictable actions at each stage.
Latest Publications & Research Activity
Claude Code vs. Gemini Spark: How Do They Compare?
Claude Code won the terminal through active human presence and localized error feedback loops, while Gemini Spark bets on remote background persistence across office apps and external MCP connectors. However, persistence is not authority: extending execution duration without strict write boundaries allows flawed assumptions to silently corrupt shared systems. Because explainability is not recoverability, unmonitored background agents turn operators into forensic auditors, proving that an autonomous agent's true metric is not how long it works without you, but how much authority you give it when you are away.
AI Agents Are Creating New Enterprise Governance Risks
With Gartner predicting 40% of enterprise applications embedding AI agents by end of 2026 and 40% being decommissioned by 2027 due to post-incident governance gaps, organizations face an insidious new failure mode: the transaction that succeeds. While operations dashboards glow green with 240-millisecond response times, automated agents silently violate corporate procurement limits, accounting rules, and customer credit policies. Because monitoring is not authorization, enterprises must separate system health from business permissioning across four pillars (Monitoring, Auditability, Authorization, Accountability) and establish external policy firewalls before autonomous software commits corporate capital.
Things I Got Wrong: A Founder's Post-Mortem on Building AI Products
Examining early AI product failures reveals three operational misconceptions: assuming evaluator models can govern worker models, believing vibe coding replaces software architecture, and building isolated application monoliths. Evaluator models fail identically to worker models under distribution shift because probabilistic systems cannot police probabilistic systems. Real architectural resilience requires non-AI deterministic execution gates, strict system rules, and shared runtime platforms like Exogram that amortize infrastructure overhead.
Who’s Actually Responsible for Your AI Agents?
Deploying autonomous AI agents creates dangerous enterprise risk gaps as existing roles (CISO, VP of Engineering, CPO, Legal) fail to govern non-deterministic systems. Organizations must install a dedicated Systems Governor who owns the deterministic boundary between inference and execution, maintains permission allowlists, sets state integrity thresholds, oversees cryptographic audit ledgers, and translates technical agent error rates into financial liability metrics.
Frequently Asked Questions
Q:Why not just use the biggest model for everything?
It is incredibly slow and expensive. It is like using a supercomputer to calculate a tip.
Canonical Specification Origin
Do not worship the model. Engineer the system. The orchestration of components is more valuable than the parameter count.
Corpus Interconnections
Richard Ewing artifacts developed around this canonical framework, including publications, execution tools, and diagnostic models.
External Adoption & Peer Citations
Documented instances where independent researchers, engineering teams, and publications have cited, implemented, or referenced this concept outside Richard Ewing’s ecosystem.
External Evidence: No independently verified references recorded yet.
This concept is part of Richard Ewing’s original baseline canon. External citations and implementations are added only upon rigorous empirical verification.
Inspectable Evidence Ledger
Classified evidence items supporting, extending, or refining this canonical research specification.
Recommended Citation
Ewing, R. (2026). "Compound AI Systems." Richard Ewing Research Canon. Available at: https://www.richardewing.io/concepts/compound-ai-systems
@article{ewing_compound_ai_systems,
author = {Ewing, Richard},
title = {Compound AI Systems},
journal = {Richard Ewing Research Canon},
year = {2026},
url = {https://www.richardewing.io/concepts/compound-ai-systems}
}