Week 13 — Lesson 2: Evaluating with a golden set

2 min

You changed the chunk size of the NorthPeak RAG. Is it better or worse? Trying two questions by hand does not tell you. A golden set is a short list of questions with known answers, checked by a human; you run it after every change and compare the scores. This lesson covers the three fields of a golden entry, the two scores that separate a retrieval failure from a reading failure, and the real result on ten questions about the NorthPeak manuals.

Preview — the rest of the lesson is for enrolled readers.

Already enrolled with a code?

Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.

Sign inNo account yet? Create one
This lesson is part of the “Week 13 — Observability and model reliability” module

The first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.

Are you a student on this course?

The code is tied to your account: sign in or create an account and it will be applied automatically when you come back.

No account yet? Create one