The model made three tool calls at once. Your backend wasn't ready.
A tool execution planner for runtimes that get several tool calls per turn.
Every deep dive we've published, newest first. Each one stands alone — start anywhere.
A tool execution planner for runtimes that get several tool calls per turn.
The design-for-variance checklist for tests, caches and evals that assume idempotence.
The symptom-to-origin table that tells you which layer to fix the bug in.
The three-layer defence for prompt injection through your RAG pipeline.
What every MCP tutorial skips: auth, tool scoping, timeouts, blast radius.
How to migrate an embedding model with a rollback path and no ranking regression.
A shadow-eval gate that scores the candidate model before your traffic sees it.
The thirty-line agent runtime that keeps a tool-calling loop from becoming a bill.
Who owns the eval set? Who owns the prompt? Who decides when to upgrade the embedding model?
Pure vector search misses exact-match queries.
Your retriever gives you 50 results. Reranking turns the top 5 useful. The difference between a RAG system that answers customer questions and one that hallucinates with confidence is rarely the...
A logging schema, three alert thresholds, and the shadow eval that catches drift before users do.
If your team cannot agree on what "good" means, you do not have a product. You have a demo.
You shipped RAG two weeks ago. The demo was 800ms, production is 4.2 seconds at p95, and the only thing your trace tells you is "LLM call: 3.8s". So you start guessing. Maybe the model is slow today.
We just sent a confirmation link. The cheatsheet lands right after you confirm.