Beyond "Ship and Pray"
Testing Agentic Systems with Geometric Ground Truth
Published on Leanpub in PDF, EPUB and online, with free updates for life.
An agent can score well on average and still fail exactly where it matters.
Averages are the problem. A benchmark number tells you how a system did across a distribution that nobody in production actually inhabits, and it stays quiet about the narrow, expensive cases that decide whether the deployment survives its first audit.
This book replaces the benchmark score with a reproducible methodology built on geometric ground truth — verified knowledge graphs and exact numerical memory, so a test knows the right answer rather than asking a second model for an opinion.
From there it works through what testing an agent actually requires: individual components, final outcomes, the structure of the trajectory that produced them, and behavior when the tools underneath start failing. The capstone applies all of it to a governed banking complaint agent, with companion notebooks so the results reproduce on your machine.
For anyone who has to answer “how do you know it works?” with something more durable than a leaderboard.
Ready to read Beyond "Ship and Pray"?
Published on Leanpub, with free updates as the book evolves.
Buy on LeanpubMore books
Beyond "Prompt and Pray"
Building Governed Agentic Systems with Geometric Memory and Verification
Beyond "Chunk and Pray"
Building Trustworthy RAG with Geometric Knowledge Graphs
Knowledge Graph Embeddings as Geometric Operators
One relation operator built from rotation, stretch and translation