Outliers and Erroneous Values

56 min
Block 5 — Data cleaning
Objective
separate an erroneous value from a legitimate outlier; master the univariate and multivariate detection methods together with their thresholds and their failure modes; quantify what a single extreme observation does to a mean, a standard deviation, a regression slope and an error metric; and choose a treatment that follows from the origin of the value rather than from its position in the distribution.
Estimated duration
55 minutes
Prerequisites
chapters 014 (descriptive statistics), 015 (visualization), 018 (missing values) and 020 (duplicates and data quality); the notion of a train/test split (chapter 026)
Associated quizzes
021.1-quiz-definitions-and-distinction.md to 021.7-quiz-train-test-scope.md

Preview — the rest of the lesson is for enrolled readers.

Already enrolled with a code?

Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.

Sign inNo account yet? Create one
This lesson is part of the “Cleaning the Data” module

The first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.

Are you a student on this course?

The code is tied to your account: sign in or create an account and it will be applied automatically when you come back.

No account yet? Create one