Duplicates, Inconsistencies and Data Quality

66 min
Block 5 — Data cleaning
Objective
hold an explicit data-quality framework with an executable test behind each dimension; detect and qualify duplicates, value inconsistencies, impossible values and mis-inferred types; run a reproducible quality audit that produces a control report; and understand why failure caused by poor data quality is silent in machine learning.
Estimated duration
50 minutes
Prerequisites
chapters 005 (dataset structure), 013 (exploratory analysis), 014 (descriptive statistics), 018 (missing values)
Associated quizzes
020.1-quiz-quality-dimensions.md to 020.7-quiz-garbage-in-garbage-out.md

Chapter 018 handled a single quality dimension — completeness — and chapter 019 handled what to do about it. This chapter handles the other five, and it handles them the same way: not as advice, but as checks that each return a number.

Preview — the rest of the lesson is for enrolled readers.

Already enrolled with a code?

Your access is tied to your account, not to this link. Sign in with the same email you used in class: your course is waiting, no need to enter the code again.

Sign inNo account yet? Create one
This lesson is part of the “Cleaning the Data” module

The first modules of the course are open to everyone. For the rest you have three options: buy this course once and for all, subscribe, or enter the code handed out in class.

Are you a student on this course?

The code is tied to your account: sign in or create an account and it will be applied automatically when you come back.

No account yet? Create one