# AI data cleaning: 20-check review worksheet

OpenMax editorial template · September 4, 2026. Complete with authorized evidence; a filled form is not automatic release approval. Use the companion source packet for the fictional eight-row example.

Dataset / source version: ____
Intended use / recipient: ____
Owner / reviewer / review date: ____
Candidate version / rule-set version: ____
Permitted environment / data-handling review: ____

For every numbered check, record status (pass / fail / unresolved / justified not applicable), evidence location, responsible person and next action. “Not checked” is unresolved, not pass. Where several rows differ, attach a row-level register instead of a single blanket answer.

## 1. Preserve the received extract

- Original payload location and immutable/versioned reference: ____
- Export timestamp, owner and content identity (not only filename): ____
- Working-copy location and recovery test result: ____

## 2. Record extraction provenance

- Export query/filter and upstream transformation version: ____
- Snapshot, incremental update or correction delivery: ____
- Unknown extraction details and person who can resolve them: ____

## 3. Define population and record grain

- One row represents / included period and status: ____
- Expected uniqueness key and legitimate repetitions: ____
- Population completeness evidence and unresolved exclusions: ____

## 4. Validate the schema and dictionary

- Required, new, missing and renamed columns: ____
- Approved field mapping and changed business definitions: ____
- Outputs paused because a definition changed: ____

## 5. Verify types and parsing

- Identifier representation and leading-zero test: ____
- Numeric locale, grouping, scale and failed parses: ____
- Actual import tool/version/settings and observed before/after values: ____

## 6. Classify missing states

- Field-specific empty / unknown / not applicable / redacted codes: ____
- Valid categories resembling missing markers, such as NA: ____
- Raw-token preservation and normalized missing-reason field: ____

## 7. Apply conditional completeness

- State-dependent required-field rule: ____
- Affected source rows and missing measurements: ____
- Owner correction request; distinguish unknown from observed zero: ____

## 8. Resolve key collisions

- All colliding rows, their grain and source versions: ____
- Documented precedence or repeated-delivery evidence: ____
- Retained row, excluded-row references and rejected alternatives: ____

## 9. Review possible entity duplicates

- Candidate matching criteria and supporting fields: ____
- Conflicting facts and legitimate same-name examples: ____
- Authorized identity decision or unresolved hold: ____

## 10. Validate references and keys

- Permitted parent snapshot and timing/status conditions: ____
- Unmatched valid keys versus empty or invalid keys: ____
- Invalid-key quarantine before enrichment and retained-source location: ____

## 11. Resolve dates and periods

- Accepted date/time formats, time zone if needed and cutoff: ____
- Ambiguous raw strings and source clarification request: ____
- Normalized date, original text and period inclusion evidence: ____

## 12. Reconcile units and conversion basis

- Raw unit, normalized unit and conversion definition: ____
- Applicable reference/date/rounding where relevant: ____
- Unavailable conversions and outputs that cannot be summed: ____

## 13. Apply the category crosswalk

- Approved source-to-target mappings, scope and version: ____
- Unmapped labels and potential meaning-erasing collisions: ____
- Raw/derived status fields and owner mapping decisions: ____

## 14. Review text normalization

- Encoding, whitespace and Unicode operations considered separately: ____
- Chosen form and field-specific before/after comparisons: ____
- Original preservation and separate matching representation: ____

## 15. Retest ranges and relationships

- Domain rules and legitimate exception scope: ____
- Post-transformation violations with source-row references: ____
- Defects introduced by parsing/conversion and corrective decision: ____

## 16. Investigate outliers

- Peer population, window and reason for flag: ____
- Source confirmation or unresolved factual uncertainty: ____
- Any analytical exclusion/capping and sensitivity with untreated data: ____

## 17. Separate estimates from observations

- Eligible fields, estimation method and uncertainty: ____
- Observed/estimated flag and preserved original missing state: ____
- For modeling: split boundary, training-only fitting and evaluation discipline: ____

## 18. Check join semantics

- Declared cardinality and uniqueness checks on each input: ____
- Null/empty-key behavior in actual tool/version: ____
- Output counts, unmatched rows, invalid-key holds and suspected wrong matches: ____

## 19. Reconcile every row and known amount

- Received = excluded + eligible + held, with row references: ____
- Known-value subtotals, unknown quantity count and full-total availability: ____
- Metric numerator/denominator and why it is not model accuracy: ____

## 20. Decide versioned release for a stated use

- Complete output or labeled subset; accepted limitations and remaining holds: ____
- Owner decision, time, recipient, candidate/rule versions and recovery reference: ____
- Next-refresh schema check, monitoring owner and reopen condition: ____

## Release cover note

Approved use or reason for continued hold: ____
Candidate, change-log, exception-register and reconciliation locations: ____
Unresolved dependencies and who will resolve them: ____
Prohibited uses, including any unsupported complete-total claim: ____

Never infer approval from a high pass rate. In the teaching packet, 3/7 eligible observations do not justify a complete-volume report, and the unknown quantity is not zero.
