Lesson 1

Why Data Gets Messy

Data quality problems don't appear by accident. Learn the common causes of messy data and how to prevent them at the source.

Read Lesson →
Lesson 2

Missing Values

Missing data is one of the most common quality problems. Learn how to identify, assess, and handle missing values appropriately.

Read Lesson →
Lesson 3

Duplicate Records

Duplicates inflate counts and cause confusion. Learn how to find and resolve duplicate records in your datasets.

Read Lesson →
Lesson 4

Formatting Issues

Inconsistent formats — dates, phone numbers, postal codes — are a major source of data problems. Learn how to standardize them.

Read Lesson →
Lesson 5

Address Validation

Addresses are notoriously messy. Learn techniques for validating, parsing, and standardizing address data.

Read Lesson →
Lesson 6

Data Standardization

Standardization makes data consistent and comparable. Learn how to apply standards to names, codes, categories, and values.

Read Lesson →
Lesson 7

Quality Assurance

QA is about preventing problems, not just fixing them. Learn how to build validation rules and quality checks into your data processes.

Read Lesson →
Lesson 8

Building a Data Refinery

A data refinery is a repeatable process for cleaning and improving data. Learn how to design one for your organization.

Read Lesson →