Home/Research/Specifications/Eval-Driven Development (EDD)
Canonical Research SpecificationLevel: Intermediate
Verified: August 2026

Eval-Driven Development (EDD)

30-Second Executive Definition

An engineering practice that embeds comprehensive, multi-dimensional AI evaluation suites into the CI/CD pipeline.

“If you are not evaluating your AI application against a golden dataset in your CI/CD pipeline, you are flying blind in a probabilistic storm.”

Why It Matters:

When a foundational model provider updates their weights, your application's behavior can change overnight without a single line of your code being altered. Without an eval-driven approach, these regressions go unnoticed until they reach the end user, causing trust erosion and financial loss. EDD provides the safety net required to deploy non-deterministic systems, allowing engineering teams to confidently ship updates, switch underlying models, and optimize prompts while quantitatively proving that system quality has improved or remained stable.

Who Should Care:
Chief Product Officer (CPO)Director of Quality AssuranceQuality Engineering (QE) ManagerProduct Operations ManagerEngineering Manager (EM)
Infinite Relationship Navigator118-Node Sovereign Knowledge Graph

Multi-Hop Causal Traversal Engine

Explore how concepts dynamically feed into each other across 1-hop, 2-hop, and 3-hop transitive relationships. Click any node to navigate the causal highway.

Current Traversal Path (1 Hops Traveled):
AI GovernanceBridge ConceptConfidence: 95%
Open Full Specification ↗

Eval-Driven Development (EDD)

An engineering practice that embeds comprehensive, multi-dimensional AI evaluation suites into the CI/CD pipeline.

Relationship Filter:
Hop Level 1

Direct Relationships (6)

Hop Level 2

Transitive Neighbors (Connected via Hop 1)

Hop Level 3

Extended Causal Ripple Effects

★ Canonical Research Position

Richard Ewing’s Research Thesis

Evaluation suites must be continuous, multi-dimensional, and treated as first-class citizens in the Agent Development Lifecycle.

Genesis & Intellectual Positioning

Why This Specification Exists

1. The Problem

Silent model updates and semantic drift break AI applications unpredictably.

2. Existing Approaches

Traditional unit testing and manual QA review.

3. The Structural Gap

Deterministic tests cannot evaluate probabilistic text or reasoning.

4. This Specification

Continuous, automated evaluation against golden datasets using LLM-as-a-judge.

Operational Realignment

What Changes If You Believe This?

Engineering

Testing moves to statistical confidence intervals and LLM-as-a-judge frameworks.

Finance & COGS

Evals consume API credits, requiring dedicated testing budgets.

Product Strategy

Ensures tone and brand safety remain consistent across model updates.

Security & Audit

Catches jailbreaks and toxic outputs before production.

Audience-Specific Executive Guidance

Recommended Action by Role

Chief Product Officer (CPO)

Maintain a verified benchmark of real customer prompts to ensure model updates never degrade response accuracy or product tone.

Recommended Next Step →
Director of Quality Assurance

Replace subjective manual spot-checking with automated golden evaluation suites that score model grounding and factual precision on every deploy.

Recommended Next Step →
Quality Engineering (QE) Manager

Build synthetic edge-case test sets that probe model responses under unexpected inputs and high concurrency before shipping to customers.

Recommended Next Step →
Product Operations Manager

Capture production failure cases directly from user feedback tickets and convert them into automated test fixtures within 24 hours.

Recommended Next Step →
Freshness & Research Updates

Latest Publications & Research Activity

Explore Full Corpus (167 Works) →
CIO.com• April 2026

The Hidden Inflation of AI: Why Model Collapse Is a Business Risk

Examines degrading economics and operational risks of recursive AI model training on enterprise margin.

Read Work ↗
Built In• September 21, 2026

Claude Code vs. Gemini Spark: How Do They Compare?

Claude Code won the terminal through active human presence and localized error feedback loops, while Gemini Spark bets on remote background persistence across office apps and external MCP connectors. However, persistence is not authority: extending execution duration without strict write boundaries allows flawed assumptions to silently corrupt shared systems. Because explainability is not recoverability, unmonitored background agents turn operators into forensic auditors, proving that an autonomous agent's true metric is not how long it works without you, but how much authority you give it when you are away.

Read Work ↗
CIO.com• September 2026

AI Agents Are Creating New Enterprise Governance Risks

With Gartner predicting 40% of enterprise applications embedding AI agents by end of 2026 and 40% being decommissioned by 2027 due to post-incident governance gaps, organizations face an insidious new failure mode: the transaction that succeeds. While operations dashboards glow green with 240-millisecond response times, automated agents silently violate corporate procurement limits, accounting rules, and customer credit policies. Because monitoring is not authorization, enterprises must separate system health from business permissioning across four pillars (Monitoring, Auditability, Authorization, Accountability) and establish external policy firewalls before autonomous software commits corporate capital.

Read Work ↗
LinkedIn• September 14, 2026

Things I Got Wrong: A Founder's Post-Mortem on Building AI Products

Examining early AI product failures reveals three operational misconceptions: assuming evaluator models can govern worker models, believing vibe coding replaces software architecture, and building isolated application monoliths. Evaluator models fail identically to worker models under distribution shift because probabilistic systems cannot police probabilistic systems. Real architectural resilience requires non-AI deterministic execution gates, strict system rules, and shared runtime platforms like Exogram that amortize infrastructure overhead.

Read Work ↗
Answer Engine FAQ Matrix

Frequently Asked Questions

Q:What is a golden dataset?

A curated collection of diverse inputs paired with their verified, ideal outputs.

01 • Origin & GenesisProvenance Record

Canonical Specification Origin

Evaluation suites must be continuous, multi-dimensional, and treated as first-class citizens in the Agent Development Lifecycle.

First IntroducedAugust 2026
Primary VenueRichard Ewing
02 • Internal Research Corpusrichardewing.io

Corpus Interconnections

Richard Ewing artifacts developed around this canonical framework, including publications, execution tools, and diagnostic models.

Articles1
Tools0
Specs1
Chapters1
03A • Verified Human External EvidenceAudit Status: Baseline

External Adoption & Peer Citations

Documented instances where independent researchers, engineering teams, and publications have cited, implemented, or referenced this concept outside Richard Ewing’s ecosystem.

External Evidence: No independently verified references recorded yet.

This concept is part of Richard Ewing’s original baseline canon. External citations and implementations are added only upon rigorous empirical verification.

Inspectable Evidence Ledger

Classified evidence items supporting, extending, or refining this canonical research specification.

Evidence ItemPublisherEvidence TypeStrengthRoleAction
The Hidden Inflation of AI: Why Model Collapse Is a Business RiskCIO.comExecutive Essay★★★★★SupportsInspect ↗
The Architecture of Runtime GovernanceBeehiivArchitecture Guide★★★★OriginInspect ↗
Academic & Industry Attribution Standard

Recommended Citation

Canonical Reference String

Ewing, R. (2026). "Eval-Driven Development (EDD)." Richard Ewing Research Canon. Available at: https://www.richardewing.io/concepts/eval-driven-development

BibTeX Citation
@article{ewing_eval_driven_development,
  author = {Ewing, Richard},
  title = {Eval-Driven Development (EDD)},
  journal = {Richard Ewing Research Canon},
  year = {2026},
  url = {https://www.richardewing.io/concepts/eval-driven-development}
}
First Origin & Provenance:Richard Ewing (August 2026)
Current Specification Version:Version 1.0 (Q2 2026 Baseline)