0%
Data-Intensive Applications
Foundations of Data Systems
Distributed Data
Encoding and Evolution
Stream Processing
Data Quality and Governance
Operational Patterns
Lakehouse and Open Table Formats
For most of the 2010s a data platform meant two systems, and everybody pretended that was fine.
On one side was the data lake: Parquet or ORC files sitting in object storage, cheap, open, readable by any engine you cared to point at them. On the other side was the data warehouse: Teradata, then Redshift, then Snowflake or BigQuery, where the data was loaded a second time into a proprietary storage layer in exchange for transactions, fast query planning, and the ability to run UPDATE.
The architecture that resulted was not a choice anybody made deliberately. It was the residue of two things that each solved half the problem.
What the lake actually gave you
Storage decoupled from compute, so you could keep ten years of history without keeping ten years of query capacity running. An open file format, so a Spark job, a Trino cluster, and somebody's pandas notebook could all read the same bytes without an export step. And a cost structure roughly an order of magnitude below warehouse storage.
Those are real advantages and they are the reason the lakehouse movement kept the object store rather than replacing it.
What the lake could not give you
The problem was never the file format. Parquet is an excellent columnar format with page-level statistics, dictionary encoding, and predicate pushdown. The problem was one layer up, in what people called the Hive table format, which is barely a format at all. In Hive, a table is a directory prefix plus a list of partition directories registered in a metastore. The set of files that constitutes the table is whatever happens to be under that prefix at the moment you list it.
Every serious defect follows from that single sentence.
There is no atomic commit. A job writing 4,000 files to a partition makes those files visible one at a time, as each upload completes. A reader that lists the prefix mid-write sees a partition that is 60 percent written and has no way to know it. The classic workaround, writing to a staging directory and renaming at the end, works on HDFS where directory rename is an atomic metadata operation. On S3 there is no rename: a rename is a copy of every object followed by a delete of every object, which is neither atomic nor cheap, and the Spark output committers built on that assumption produced silently incomplete tables for years.
There is no row-level delete. The unit of mutation is a file, and files in object storage are immutable. Honoring one GDPR erasure request against a partitioned table meant finding every partition containing that user and rewriting all of it. Teams batched erasure requests into monthly jobs not because the regulation allowed a month but because the architecture did not allow anything faster.
There are no trustworthy statistics. The metastore knows partition values. It does not know row counts, null counts, or column min/max per file, so the query planner cannot prune below the partition level and cannot estimate join cardinality. The planner is flying on the one dimension somebody remembered to partition by.
Listing is the bottleneck nobody budgets for. Planning a query means recursively listing the prefix. A 10 TB table written as 4 MB files contains 2,621,440 objects. At 1,000 keys per LIST response that is 2,622 sequential round trips before a single byte of data has been read, and object stores rate-limit listing per prefix. The same table written as 128 MB files contains 81,920 objects, which is the first hint that file sizing is a first-class operational concern rather than a tuning detail.
Schema is a suggestion. Schema-on-read means the schema is whatever the reader guesses. A producer that changes a column from int to string breaks consumers at read time, in production, with no failed write to point at.
The cost of running both
Because the lake could not support the workloads the business actually asked for, the data was copied into a warehouse, and the copy is where the real damage was. You paid for storage twice. You maintained a pipeline whose only job was to move data between two systems you already owned. Lineage broke at the boundary, so nobody could answer whether a dashboard number disagreed with a notebook number because of a bug or because the copy was six hours stale. And the two copies diverged in exactly the ways that are hardest to detect: a timezone handled differently on each side, a numeric that became a float in transit, a late-arriving record applied to one copy and not the other.
The lakehouse thesis
The lakehouse proposition is narrower than the marketing suggests, and it is worth stating precisely because the narrowness is the point.
Keep the data files exactly where they are, in open columnar formats in object storage. Replace the one broken idea, that a table is whatever files are under a directory, with an explicit, versioned, immutable manifest of exactly which files constitute the table at each point in time. Store that manifest alongside the data. Make the act of publishing a new version of the table a single atomic pointer swap.
That is the whole mechanism. Atomic commits, snapshot isolation, time travel, row-level deletes, reliable statistics, and planning without listing are all consequences of it, and the rest of this lesson is the derivation.
One distinction to fix now, because it causes more confusion than any other point in this area: an open table format is not a file format. Parquet is a file format; it describes the bytes of one file. Apache Iceberg, Delta Lake, and Apache Hudi are table formats; they describe which files belong to a table and what they mean. A lakehouse table is Parquet data files plus table-format metadata. Nothing about the data files changes, which is why adopting a table format over an existing lake is a metadata operation rather than a migration of petabytes.
The lakehouse did not invent anything the warehouse had not had for decades. What it did was move the transaction boundary out of a proprietary storage engine and into a small open metadata specification that sits on commodity object storage. The interesting engineering is not in the features; it is in getting transactional semantics out of a store whose only strong primitive is single-object atomic replacement.