0%
Data-Intensive Applications
Foundations of Data Systems
Distributed Data
Encoding and Evolution
Batch Processing
Stream Processing
Operational Patterns
Data Validation and Quality
Bad data is worse than no data. When data is missing, you know you cannot make a decision. When data is wrong, you make the wrong decision with confidence. A pricing model trained on duplicate records overestimates demand. A dashboard built on stale inventory counts tells the ops team that warehouses are full when they are actually empty. A customer segmentation pipeline that treats "N/A" strings as valid country codes routes marketing emails to nowhere.
Data quality is not a one-time check you run before launch. It is a continuous discipline, like testing in software engineering. Every time a source system changes, every time a new producer starts writing events, every time a schema migration runs, data quality can degrade silently. The cost of catching problems early (in the pipeline) versus late (in a board meeting when revenue numbers do not match) differs by orders of magnitude.
The data engineering community has converged on five dimensions that together define whether data is fit for its intended use. These are not academic abstractions. Each dimension maps to a specific class of bugs that will eventually hit your pipeline if you do not measure and enforce it.
The cost of bad data
The business impact of poor data quality is concrete and measurable. Gartner estimates that poor data quality costs organizations an average of $12.9 million per year. But the real cost is not just financial. Bad data erodes trust: when a VP sees different revenue numbers in two dashboards, they stop trusting both. Rebuilding trust in data after a quality incident takes months, even if the fix takes hours. Every decision made with bad data has a hidden error bar that grows with each quality dimension you fail to measure.
The five dimensions below give you a systematic framework for identifying, measuring, and preventing quality problems before they reach decision-makers.
Completeness
Completeness measures whether all expected data is present. A user table with 10% of email addresses missing is 90% complete on that column. A daily sales extract that only captured 18 out of 24 hours is 75% complete on time coverage.
Completeness has two levels. Column-level completeness asks "what fraction of rows have a non-null value for this field?" Row-level completeness asks "did we receive all the rows we expected?" The second is harder to measure because you need a reference point: yesterday's count, a source system's reported total, or an upstream reconciliation number.
Missing data is often silent. A LEFT JOIN that drops rows due to a missing foreign key does not throw an error. It just produces fewer results. The dashboard still renders. The report still generates. Nobody notices until someone manually compares the output against the source and finds 15% of records vanished in the join.
Detecting completeness failures requires proactive checks at every stage. At ingestion, compare the number of records received against the number the source system claims to have sent. After transformation, compare row counts before and after joins (an INNER JOIN that drops more than 1% of rows is a warning sign). After loading, compare the total with the previous day's count and flag deviations beyond a configurable threshold (typically 10-20% for daily loads).
Accuracy
Accuracy measures whether data values correctly represent the real-world entity they describe. A customer record with the wrong phone number is inaccurate. A sensor reading of -40 degrees Celsius in Miami in July is inaccurate. An order total of $0.01 for a laptop purchase is inaccurate.
Accuracy is the hardest dimension to validate automatically because it requires knowing the truth. You can check for obvious violations (negative ages, prices above $1 million for a grocery order), but subtle inaccuracies (a shipping address with the wrong apartment number) often require human review or cross-referencing against external sources.
Practical accuracy checks fall into two categories. Deterministic checks flag values that are physically impossible (a person born in the year 2200, a package weight of -3 kg). These are cheap and should always be implemented. Probabilistic checks flag values that are statistically unlikely (an order total that is 10x the historical average for that product category). These require baseline statistics and produce false positives, so they should alert rather than reject.
Consistency
Consistency measures whether the same fact is represented the same way across different systems and datasets. If the CRM says a customer is in "New York" but the billing system says "NY" and the analytics warehouse says "new york," all three are arguably accurate, but they are inconsistent. Joins, aggregations, and deduplication all break when the same entity has different representations.
Consistency problems are especially common after mergers, migrations, and when multiple teams independently model the same domain. Two teams might both track "revenue" but one includes refunds and the other does not. Both numbers are internally consistent, but comparing them produces nonsense.
Fixing consistency requires agreeing on canonical representations and enforcing them at the point of ingestion. A common approach is to define a "golden record" for each entity type: one authoritative source and format that all downstream systems reference. When a new data source is onboarded, the first step is mapping its representations to the canonical format (state abbreviation lookup tables, currency standardization, date format normalization). This mapping logic should be centralized in a shared library, not duplicated across pipelines.
Timeliness
Timeliness measures whether data arrives when it is needed. A fraud detection model that receives transaction data with a 6-hour delay cannot prevent fraud. It can only report it after the fact. A real-time inventory system that updates every 15 minutes is "timely" for a warehouse dashboard but not for a flash sale that sells out in 3 minutes.
Timeliness is relative to the use case. Batch analytics can tolerate hours of delay. Operational dashboards need minutes. Real-time decisioning needs seconds or less. The SLA should be defined per data consumer, not globally.
A subtle form of timeliness failure is "data that arrives on time but reflects a stale state." If a CDC (change data capture) pipeline captures a snapshot of a table every hour, the data is fresh at capture time but could be up to 59 minutes stale at query time. For use cases that require true real-time freshness, event-driven architectures (publishing changes as they happen) are necessary rather than periodic snapshots.
Uniqueness
Uniqueness measures whether each real-world entity appears exactly once in the dataset. Duplicate customer records inflate customer counts, skew segmentation models, and cause duplicate communications (two welcome emails, two invoices). Duplicate transaction records inflate revenue.
Deduplication is straightforward when records have natural keys (email address, order ID). It becomes a fuzzy matching problem when they do not: "John Smith at 123 Main St" and "J. Smith at 123 Main Street" might be the same person or two different people. Probabilistic matching algorithms (Levenshtein distance, Jaro-Winkler similarity) help, but they introduce false positives (merging records that should be separate) and false negatives (missing duplicates).
How the dimensions interact
These five dimensions are not independent. A completeness problem often causes an accuracy problem downstream. If 10% of transaction records are missing (completeness failure), any aggregate computed from those transactions (total revenue, average order value) is inaccurate by definition. Similarly, consistency problems create uniqueness problems: if two systems represent the same customer differently, a merge produces duplicates.
The practical implication is that you cannot fix one dimension in isolation. A data quality program must measure all five dimensions continuously, prioritize based on business impact, and understand the causal relationships between them. A rising null rate in a foreign key column (completeness) will eventually cause row drops in downstream joins (completeness again), which will cause revenue underreporting (accuracy). Catching the null rate trend early prevents the cascade.
Measuring data quality in practice
Measurement starts with profiling: running statistical summaries across every column in every table to establish baselines. What is the normal null rate for each column? What is the expected row count range per daily load? What is the typical value distribution for numeric columns? These baselines become the reference points for anomaly detection.
Tools like Great Expectations, Deequ, and Soda (covered in the data contracts section) automate profiling and ongoing measurement. The key principle is that data quality must be measured continuously, not checked once during development. A column that was 100% complete during testing might develop nulls in production when a new mobile client starts sending partial records. Without ongoing measurement, you discover this when a downstream model starts producing wrong predictions, weeks or months after the problem started.
Mid-level engineers validate individual columns. Senior engineers define quality SLAs per dimension across entire datasets. Staff engineers build frameworks that let every team define, measure, and enforce quality dimensions without writing custom code for each pipeline.