RAG and Evaluation of Generative Systems

2 min
  1. Why measure retrieval before the quality of the final answer?
  2. What does recall@k mean in a set where the supporting passage is annotated?
  3. At which level should the citations of a generated answer be verified?
  4. Why include unanswerable questions in the evaluation set?
  5. Why apply access rights before document search?

Answers. 1. An end-to-end score does not localize the failure; without a supporting passage in context, the generation cannot be correctly blamed. 2. The proportion of questions for which at least one supporting passage appears among the k results. 3. At the level of each important statement, verifying that an identified passage really supports it. 4. They measure refusal and abstention when evidence is missing. 5. Filtering after generation is too late: a forbidden document may already have entered the model's context and influenced or exposed the answer.

Preview — the rest of the lesson is for enrolled readers.

Already enrolled with a code?

Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.

Sign inNo account yet? Create one
This lesson is part of the “Modern Deep Dives” module

The first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.

Are you a student on this course?

The code is tied to your account: sign in or create an account and it will be applied automatically when you come back.

No account yet? Create one