RAG Eval

Measuring whether your retrieval actually works.

Read the blog →

Your RAG system works. Probably. You can’t currently prove it.

The demo went well. Then someone in the room typed a question you hadn’t tried, got a confidently wrong answer with a citation attached, and now you’re being asked whether the whole thing is broken.

You need a number, and you don’t have one — because “is our RAG good” isn’t a question a vibe check can answer, and the metrics that look like answers mostly aren’t.

This site is about measurement. Not what RAG is. Whether yours works.

What you’ll find here

The discipline this site keeps: retrieval metrics, generation metrics, and end-to-end metrics are three different things. Nearly every confused RAG evaluation is confused because someone averaged across that boundary.

Written for people who already shipped something. Code where it helps, none where it doesn’t.

Read the blog →

No benchmark numbers presented as universal. Whether reranking helps your corpus is an empirical question about your corpus, and anyone quoting you a percentage is quoting someone else’s data.

Latest posts

See all posts →