How to Evaluate LLM Agents: 5 Patterns That Catch Failures
Eval shortcuts fail on honesty-trained models. Five assertion patterns that survived production.
July 11, 2026 · 12 min
Cloud Architecture · AI Engineering · Distributed Systems
47 posts · 2025–2026
Eval shortcuts fail on honesty-trained models. Five assertion patterns that survived production.
July 11, 2026 · 12 min
Six retrieval decisions for agent RAG, each settled by a benchmark, not a vibe. Defaults included.
July 10, 2026 · 10 min
pdf_search went from a pivot tool to a terminal tool. The thesis still holds.
July 04, 2026 · 8 min
Your tools work. That doesn't mean an agent can use them. Eight rounds. Seventeen bugs.
June 27, 2026 · 12 min
How agents should navigate documents at production scale, from 32,000+ downloads of one MCP server.
June 24, 2026 · 11 min
Seven Lambdas, two SQS queues, one DynamoDB table. SES for sending. No per-subscriber fee.
June 20, 2026 · 10 min
Why more MCP tools make agents worse, and the pattern that fixes the surface.
June 13, 2026 · 8 min
Take the notes server from localhost STDIO to a secured, containerized HTTP service you can run on any host.
June 09, 2026 · 11 min
Hybrid wins at page grain. BM25 wins at section grain. Granularity decides.
June 06, 2026 · 10 min
Page-mode PDF search costs 2 to 6 extra tool calls per query depending on the document. Section-aware search delivers...
May 30, 2026 · 10 min
Every Claude Desktop session using your MCP server is a free QA pass. You just have to listen to what the LLM is tryi...
May 23, 2026 · 10 min
A prompt change broke my agent silently. Behavioral CI caught 4 regressions before they shipped.
May 16, 2026 · 8 min
No articles found for this filter.