Course progress0%
Course content

Module 1

Module 2

Module 3

Module 4

Module 5

Module 6

Module 7

Module 8

Module 9

Module 10

Module 11

Module 12

Module 13

Module 14

Module 15

Module 16

Module 17

Module 18

35 min

Data Cleaning without Inventing Data

Separate normalization from validation

By the end of this lesson

  • Separate normalization from validation
  • Preserve information when the transformation is uncertain

Cleaning can destroy meaning

A username arrives as " Sara ". A phone identifier arrives as "00123". Trimming surrounding spaces may be agreed behavior; converting the identifier to a number would lose zeros. Ask which transformations preserve the meaning before applying them.

function cleanUsername(rawUsername) {
  if (typeof rawUsername !== "string") {
    return null;
  }
  const username = rawUsername.trim();
  if (username.length === 0) return null;
  return username;
}
console.assert(cleanUsername("  Sara  ") === "Sara");
console.assert(cleanUsername("   ") === null);
console.assert(cleanUsername(null) === null);

Normalization brings equivalent representations into an agreed form. Validation decides whether the result is acceptable. The function rejects unsupported inputs rather than converting everything into text and accepting it.

Guided review: email and search policies

For a search index, lowercasing may be appropriate. For an email address, do not assume arbitrary edits preserve its meaning. Trim surrounding accidental whitespace if agreed; use a defined email-validation policy rather than assuming that includes("@") proves an address is valid. Avoid lowercasing the entire address as a universal rule.

Keep a user's chosen display name when case matters to them. A separate normalized lookup value can serve matching. Never clean a password by silently trimming it: that changes the supplied secret.

Independent lab

Given [" Sara ", "", " Youssef", " "], produce cleaned valid names and count rejected entries. State whether duplicates should be removed. Do not invent that policy.

Correction

const rawNames = [" Sara ", "", "  Youssef", "   "];
const results = rawNames.map(cleanUsername);
const validNames = results.filter((name) => name !== null);
console.assert(validNames.length === 2);
console.assert(results.length - validNames.length === 2);

This fragment uses cleanUsername from above. The valid names are Sara and Youssef. Duplicate handling remains a separate decision; silently deleting names might remove distinct people.

Review

A professional cleaning step should explain what was changed, what was rejected and which uncertainties remain. "The input looks nicer" is not evidence that it is now correct.

Lesson complete?

Your progress is saved on this device.