Week 4 — Lesson 4: Train, validation, test and the three splits

3 min

NorthPeak will use the fault model next month, on dates it has never seen. A model must therefore be scored on rows it never trained on, and the way you choose those rows decides whether the score means anything. Random, temporal and group splits give different scores on the same data. This lesson explains the three sets, the three ways to cut, and why every AUC of this course uses the temporal split at 2025-10-01.

Preview — the rest of the lesson is for enrolled readers.

Already enrolled with a code?

Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.

Sign inNo account yet? Create one
This lesson is part of the “Week 4 — Data preparation and dataset separation” module

The first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.

Are you a student on this course?

The code is tied to your account: sign in or create an account and it will be applied automatically when you come back.

No account yet? Create one