Data quality checks make assumptions visible and testable. The useful checks are the ones tied to how data will be consumed, with a defined response when a check fails.
Start with structure and completeness
- Null checks: verify fields that are required for joins, reporting, or downstream processing.
- Schema validation: check expected columns and compatible data types, and decide how schema changes are reviewed.
- Accepted values and ranges: validate codes, statuses, dates, and numeric boundaries against documented rules.
- Freshness: compare the latest available event or load time with the consumer's expected update window.
Check keys and reconcile movement
- Duplicate checks: test business-key uniqueness at the grain where uniqueness is required.
- Referential integrity: identify dimension or lookup keys that do not resolve.
- Row-count reconciliation: compare source, landed, accepted, rejected, and target counts for the same run boundary.
- Source-to-target validation: compare selected totals or hashes when exact row comparison is not practical.
Make failures useful
A failed check should report its scope, rule, expected behavior, observed result, and run identifier. Quarantine invalid records with reason codes when safe; block publication when a critical contract is broken. Avoid silently dropping records or turning every warning into a pipeline-wide outage.
Put checks into the workflow
Run inexpensive structural checks near ingestion, transformation assertions around business rules, and reconciliation before publishing. Keep the rule definitions with the versioned pipeline and monitor trends so a gradual deterioration is visible.