Repository navigation
parquet@py: write a page index and page checksums by default - #23
Merged
mprammer merged 2 commits intoOct 9, 2026
Merged
Conversation
pyarrow's own defaults write neither a ColumnIndex/OffsetIndex nor page CRCs, so every catalog Parquet written by the default writer had no page index, unlike parquet-java and arrow-rs. With RAINCLOUD_PARQUET_PAGE_INDEX and RAINCLOUD_PARQUET_PAGE_CHECKSUMS unset, parquet@py now passes write_page_index (when write.statistics is on) and write_page_checksum, matching parquet-java's defaults. Setting either to 0 restores the old files. Every parquet@py artifact changes bytes and checksums. Co-authored-by: Isaac <no-reply@databricks.com>
|
|
mprammer
approved these changes
Oct 9, 2026
mprammer
left a comment
Contributor
There was a problem hiding this comment.
Thanks for the PR! Looks good. I'll update the other writer defaults as well so everything matches where it can.
3 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The default Parquet writer,
parquet@py, leaves page settings at pyarrow's defaults, and pyarrow writes no page index (no ColumnIndex and no OffsetIndex) and no page CRCs. So every catalog Parquet file it writes has no page index, unlike the files from parquet-java (parquet@java) and arrow-rs (parquet@rs). That matters for anyone using the catalog to benchmark page pruning or compare readers.parquet-java 1.17.1's defaults (
ParquetProperties):DEFAULT_STATISTICS_ENABLED = true, which writes a column index;DEFAULT_PAGE_WRITE_CHECKSUM_ENABLED = true.pyarrow 25.0.1's defaults (
pq.ParquetWriter):write_page_index=False,write_page_checksum=False.What
exporters._writer_options: whenRAINCLOUD_PARQUET_PAGE_INDEXis unset, passwrite_page_index=<write.statistics>. WhenRAINCLOUD_PARQUET_PAGE_CHECKSUMSis unset, passwrite_page_checksum=True. A set value still wins, so0gives the old files back. A recipe withwrite.statistics: falsestill gets no page index, since a page index is page statistics.sidecars/README.md(capability table), the comment abovespec._PARQUET_SETTINGS, and a CHANGELOG[Unreleased]entry.test_parquet_py_writes_a_page_index_and_page_checksums_when_unset.The other three lanes are unchanged.
Impact
Every Parquet artifact
parquet@pywrites changes bytes and sha256, so the published Parquet would need a rebuild/re-measure. The default is not part ofwriter_toolchain(only set options are), so a failure recorded under the old default is not retried automatically. A page index has not been a failure cause in any lane I know of, but say if you'd rather record it in the toolchain.A ColumnIndex is still absent for chunks without min/max statistics, e.g. the leaves of a VARIANT column, which has no defined sort order. Their OffsetIndex is written. This is the same as
RAINCLOUD_PARQUET_PAGE_INDEX=1today.Testing
pytestwith CI's hermetic extras (dev tui build pandas osm sas excel archives): 1648 passed, 541 skipped (sidecar lanes not installed locally), 0 failed.ruff checkandpython -m raincloud.pipeline.validate_manifestboth pass.python -m raincloud.pipeline.build countries-of-the-world --format parquet: every chunk now has an OffsetIndex, and every chunk with min/max statistics has a ColumnIndex.This pull request and its description were written by Isaac.