Evaluation, Calibration and Explainability

2 min

Objective: evaluate the expected system, not just an average score.

Three levels of evaluation

  1. Loss: signal optimized during training.
  2. ML metrics: precision, recall, F1, ROC-AUC, MAE, etc.
  3. System/business metrics: error cost, latency, abstention rate, impact.

The choice depends on the cost of false positives and false negatives. With a rare class, accuracy can be nearly perfect for a useless model.

Calibration

A model is calibrated if, among predictions announced at 80%, about 80% are correct. Softmax confidence is not automatically a reliable probability. Reliability diagram, Brier score and temperature calibration can help.

Robustness and subgroups

Measure by class, source, period and relevant group. Test noise, plausible variations, out-of-distribution data and realistic adversarial examples. Document uncertainty when samples are few.

Preview — the rest of the lesson is for enrolled readers.

Already enrolled with a code?

Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.

Sign inNo account yet? Create one
This lesson is part of the “From Model to System” module

The first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.

Are you a student on this course?

The code is tied to your account: sign in or create an account and it will be applied automatically when you come back.

No account yet? Create one