Skip to content

parquet@py: write a page index and page checksums by default - #23

Merged
mprammer merged 2 commits into
spiraldb:developfrom
Jiayi-Wang-db:parquet-py-page-index-default
Oct 9, 2026
Merged

mprammer merged 2 commits into
spiraldb:developfrom
Jiayi-Wang-db:parquet-py-page-index-default

Conversation

@Jiayi-Wang-db

Copy link
Copy Markdown
Contributor

Why

The default Parquet writer, parquet@py, leaves page settings at pyarrow's defaults, and pyarrow writes no page index (no ColumnIndex and no OffsetIndex) and no page CRCs. So every catalog Parquet file it writes has no page index, unlike the files from parquet-java (parquet@java) and arrow-rs (parquet@rs). That matters for anyone using the catalog to benchmark page pruning or compare readers.

parquet-java 1.17.1's defaults (ParquetProperties): DEFAULT_STATISTICS_ENABLED = true, which writes a column index; DEFAULT_PAGE_WRITE_CHECKSUM_ENABLED = true.
pyarrow 25.0.1's defaults (pq.ParquetWriter): write_page_index=False, write_page_checksum=False.

What

  • exporters._writer_options: when RAINCLOUD_PARQUET_PAGE_INDEX is unset, pass write_page_index=<write.statistics>. When RAINCLOUD_PARQUET_PAGE_CHECKSUMS is unset, pass write_page_checksum=True. A set value still wins, so 0 gives the old files back. A recipe with write.statistics: false still gets no page index, since a page index is page statistics.
  • Docs updated wherever they said "pyarrow writes no page index": README, AGENTS.md, sidecars/README.md (capability table), the comment above spec._PARQUET_SETTINGS, and a CHANGELOG [Unreleased] entry.
  • New test test_parquet_py_writes_a_page_index_and_page_checksums_when_unset.

The other three lanes are unchanged.

Impact

Every Parquet artifact parquet@py writes changes bytes and sha256, so the published Parquet would need a rebuild/re-measure. The default is not part of writer_toolchain (only set options are), so a failure recorded under the old default is not retried automatically. A page index has not been a failure cause in any lane I know of, but say if you'd rather record it in the toolchain.

A ColumnIndex is still absent for chunks without min/max statistics, e.g. the leaves of a VARIANT column, which has no defined sort order. Their OffsetIndex is written. This is the same as RAINCLOUD_PARQUET_PAGE_INDEX=1 today.

Testing

  • pytest with CI's hermetic extras (dev tui build pandas osm sas excel archives): 1648 passed, 541 skipped (sidecar lanes not installed locally), 0 failed.
  • ruff check and python -m raincloud.pipeline.validate_manifest both pass.
  • python -m raincloud.pipeline.build countries-of-the-world --format parquet: every chunk now has an OffsetIndex, and every chunk with min/max statistics has a ColumnIndex.

This pull request and its description were written by Isaac.

pyarrow's own defaults write neither a ColumnIndex/OffsetIndex nor page
CRCs, so every catalog Parquet written by the default writer had no page
index, unlike parquet-java and arrow-rs. With RAINCLOUD_PARQUET_PAGE_INDEX
and RAINCLOUD_PARQUET_PAGE_CHECKSUMS unset, parquet@py now passes
write_page_index (when write.statistics is on) and write_page_checksum,
matching parquet-java's defaults. Setting either to 0 restores the old
files. Every parquet@py artifact changes bytes and checksums.

Co-authored-by: Isaac <no-reply@databricks.com>
@CLAassistant

CLAassistant commented Oct 9, 2026 •

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you all sign our Contributor License Agreement before we can accept your contribution.
1 out of 2 committers have signed the CLA.

✅ mprammer
❌ jiayi-wang-data
You have signed the CLA already but the status is still pending? Let us recheck it.

@mprammer mprammer left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR! Looks good. I'll update the other writer defaults as well so everything matches where it can.

@mprammer
mprammer merged commit bd7a89b into spiraldb:develop Oct 9, 2026
8 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants