The production RAG eval cheatsheet.
25 failure modes to test before you ship. The symptom, the test, the snippet, the fix — for each one.
One more step — check your inbox.
We just sent a confirmation link. Click it to lock in your subscription and we'll send the cheatsheet right after.
First Token is not another AI newsletter rounding up the week's headlines.
Every issue starts with a real failure mode, debug story, or architectural decision — drawn from production systems serving actual users. If it only works in a demo, it doesn't ship here.
One deep dive lands every Tuesday. Long enough to be useful, short enough to read on the commute.
Tactics from real shipping, not blog posts about blog posts.
Every issue starts with a real failure mode, debug story, or architectural decision — drawn from systems serving actual users. If it works in a demo, it doesn't ship here.
Written for engineers who ship the systems behind the demo.
APIs, queues, retrieval, caching, observability, evals. The boring infrastructure that makes LLM products actually work in production — not prompt engineering hot takes.
A single deep dive every Tuesday. No filler, no daily noise.
Long enough to be useful, short enough to read on the commute. Roughly 1,800 words, one diagram, one runnable snippet per issue. That's the contract.
Recent issues.
Aug 4, 2026
Jul 28, 2026
When to fine-tune, when to use RAG, and when to just fix the prompt
Jul 21, 2026
How to structure an AI engineering team
Who owns the eval set? Who owns the prompt? Who decides when to upgrade the embedding model?
Jul 14, 2026
The 90% discount on input tokens that most teams do not claim.
Jul 7, 2026
Streaming LLM responses without breaking your backend
What do I get when I subscribe?
One production AI deep dive every Tuesday, plus the 25-entry RAG eval cheatsheet — a PDF that lands right after you confirm your email. Both free.
Who is this written for?
Backend engineers shipping LLM features. It assumes you're comfortable with APIs, queues, and databases — and skips the prompt engineering hot takes.
How long is each issue?
Roughly 1,800 words with one diagram and one runnable snippet. Long enough to be useful, short enough for the commute.
Can I read something before subscribing?
Yes — every published issue is free to read in the archive. Start anywhere; each one stands alone.
Production AI engineering. One Tuesday at a time.
Check your inbox.
We just sent a confirmation link. The cheatsheet lands right after you confirm.