Oct 07, 2026
Lakehouse – Phase 3 Ingest Pipeline
Phase 1 gave us a metadata catalog that can track tables, columns, files, and statistics. Phase 2 gave us the ability to write and read Parquet files with column statistics. This phase connects them: an INSERT writes data to a Parquet file and then atomically registers that file (with its statistics) in the metadata catalog. The design has one core principle: data files are written before metadata is updated, and the metadata update is atomic.