The production RAG eval cheatsheet.
25 failure modes to test before you ship. The symptom, the test, the snippet, the fix — for each one.
One more step — check your inbox.
We just sent a confirmation link. Click it to lock in your subscription and we'll send the cheatsheet right after.
First Token is not another AI newsletter rounding up the week's headlines.
Every issue starts with a real failure mode, debug story, or architectural decision — drawn from production systems serving actual users. If it only works in a demo, it doesn't ship here.
One deep dive lands every Tuesday. Long enough to be useful, short enough to read on the commute.
Tactics from real shipping, not blog posts about blog posts.
Every issue starts with a real failure mode, debug story, or architectural decision — drawn from systems serving actual users. If it works in a demo, it doesn't ship here.
Written for engineers who ship the systems behind the demo.
APIs, queues, retrieval, caching, observability, evals. The boring infrastructure that makes LLM products actually work in production — not prompt engineering hot takes.
A single deep dive every Tuesday. No filler, no daily noise.
Long enough to be useful, short enough to read on the commute. Roughly 1,800 words, one diagram, one runnable snippet per issue. That's the contract.
Recent issues.
Sep 22, 2026
Sep 15, 2026
Your model's weirdest production bugs were trained into it
The symptom-to-origin table that tells you which layer to fix the bug in.
Sep 8, 2026
Retrieved documents are an untrusted input channel
The three-layer defence for prompt injection through your RAG pipeline.
Sep 1, 2026
Shipping an MCP server that survives real traffic
What every MCP tutorial skips: auth, tool scoping, timeouts, blast radius.
Aug 25, 2026
Swapping the embedding model without taking retrieval down
How to migrate an embedding model with a rollback path and no ranking regression.
What do I get when I subscribe?
One production AI deep dive every Tuesday, plus the 25-entry RAG eval cheatsheet — a PDF that lands right after you confirm your email. Both free.
Who is this written for?
Backend engineers shipping LLM features. It assumes you're comfortable with APIs, queues, and databases — and skips the prompt engineering hot takes.
How long is each issue?
Roughly 1,800 words with one diagram and one runnable snippet. Long enough to be useful, short enough for the commute.
Can I read something before subscribing?
Yes — every published issue is free to read in the archive. Start anywhere; each one stands alone.
Production AI engineering. One Tuesday at a time.
Check your inbox.
We just sent a confirmation link. The cheatsheet lands right after you confirm.