Cross-Document Retrieval for AI Agents Without a Vector Database
BM25 scores don't compare across documents. Ranks do. Fusing 100 PDFs into one hit list.
August 15, 2026 · 12 min
Cloud Architecture · AI Engineering · Distributed Systems
52 posts · 2025–2026
BM25 scores don't compare across documents. Ranks do. Fusing 100 PDFs into one hit list.
August 15, 2026 · 12 min
An open-source MCP server for Redmine. 51 tools, OAuth2, read-only mode, and a live triage board.
August 08, 2026 · 8 min
No pipeline, no vector store. Topic folders, a warm habit, a generated catalog, honest coverage.
August 01, 2026 · 10 min
Two Python libraries fix the 403 and the 80k-token HTML at the same time.
July 25, 2026 · 9 min
Fixing two-column extraction broke academic title pages. The benchmark never noticed.
July 18, 2026 · 8 min
Eval shortcuts fail on honesty-trained models. Five assertion patterns that survived production.
July 11, 2026 · 12 min
Six retrieval decisions for agent RAG, each settled by a benchmark, not a vibe. Defaults included.
July 10, 2026 · 10 min
pdf_search went from a pivot tool to a terminal tool. The thesis still holds.
July 04, 2026 · 8 min
Your tools work. That doesn't mean an agent can use them. Eight rounds. Seventeen bugs.
June 27, 2026 · 12 min
How agents should navigate documents at production scale, from 39,000+ downloads of one MCP server.
June 24, 2026 · 11 min
Seven Lambdas, two SQS queues, one DynamoDB table. SES for sending. No per-subscriber fee.
June 20, 2026 · 10 min
Why more MCP tools make agents worse, and the pattern that fixes the surface.
June 13, 2026 · 8 min
No articles found for this filter.