Skip to content

Release 0.3.1: ORC, Avro and Nimble; opt-in formats; encoder settings, including the Parquet page index - #22

Merged
mprammer merged 17 commits into
developfrom
mp/release-0.3.1
Oct 8, 2026
Merged

mprammer merged 17 commits into
developfrom
mp/release-0.3.1

Conversation

@mprammer

@mprammer mprammer commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Cuts 0.3.1. Three file formats join Parquet, Vortex and Arrow — ORC, Avro and Nimble — and every format's writers now take the same encoder settings, so a dataset can be built as Parquet with a page index. Formats become opt-in per install: a plain install builds only Vortex, and a build removes its raw download and canonical Arrow once the files are written.

Encoder settings. Each format's writers read one set of settings from the environment (RAINCLOUD_PARQUET_*, RAINCLOUD_ORC_*, RAINCLOUD_AVRO_*, RAINCLOUD_VORTEX_*), declared once in raincloud/pipeline/spec.py. For Parquet that is a page index on every column or only the first N, statistics for only the first N columns, compression level, page size and rows, dictionaries and page checksums. Unset, each writer library's default applies, and no file changes: with nothing set, every writer was checked to write the bytes it wrote before, except that parquet-java's list of a column's encodings can come out in another order, as it already could from one JVM run to the next. A writer whose library cannot do what a setting asks refuses it, and the format is recorded unavailable with the reason, rather than written another way; sidecars/README.md tabulates which writer honours what. The recipe's write.compression and write.statistics now reach the sidecar writers too, which used to pick their own.

RAINCLOUD_PARQUET_PAGE_INDEX=1 raincloud build uci-iris --format parquet

parquet@java moves to parquet-arrow-java 0.3.0 for a compression level, per-column statistics and LZ4_RAW, and parquet@hardwood ships brotli4j.

New formats. ORC is written by pyarrow and orc-rust, Avro by arrow-avro and Arrow Java's adapter, and Nimble by upstream Nimble's C++, built from source from a pinned commit on a fork carrying build fixes (sidecars/nimble/build.sh). auto still picks Vortex, then Parquet; the new formats are built and loaded only when named.

Also. One comparison rule set shared by the Python, Rust and JVM comparators (a union of null and T is a nullable T; zoned timestamps compare by instant). A generated table builds its whole group. The recipe hash ignores export.formats. v1 catalogs keep their old behaviour.

Docs for agents. A new /raincloud-write-settings skill and matching SKILLS.md playbook; the build, export and load skills, AGENTS.md and the README cover the settings and the new formats.

Compliance. docs/v2/compliance.json is still 0.3.0's measurement; the new lanes and settings are measured in a later release (noted in the changelog).

Tag v0.3.1 and the GitHub release follow on merge.

🤖 Generated with Claude Code

mprammer and others added 17 commits October 1, 2026 19:57
`_registry.FORMATS` now declares every artifact format: its file extension,
whether `auto` may pick it, and the in-process reader that opens it. The
extension map, reader capabilities, `auto` order, snapshot generation, status
columns, publish set, manifest validation and CLI help derive from it instead
of naming parquet and vortex. The loader's per-format open code moves behind
a table in `_readers`, so a format with no in-process reader raises
MissingDependency instead of falling through to the Vortex reader.
sources.schema.json cannot import the registry; a test gates its enums.

Co-Authored-By: Claude <noreply@anthropic.com>
A build now writes only Vortex unless the new `formats` setting
(RAINCLOUD_FORMATS, or "all") or `raincloud build --format` asks for more,
and a load that names another format builds just that one. Every dataset
offers every exported format; a recipe's export.formats only narrows that.
`auto` picks among the formats the install builds, Vortex then Parquet, and
falls back to the canonical Arrow when neither can be had.

Once its files are written, a successful build removes the raw download and
the canonical Arrow unless `keep_raw` / `keep_canonical` keep them. A
canonical from which nothing was written is the dataset's only file and
stays. Generated datasets keep their generator output, which a group shares.

`status` reports against the install's settings; `export` without --format
writes the install's formats. Tests pin the 0.3.0 settings where they
exercise pipeline mechanics, and test_build_formats covers the new defaults.

Co-Authored-By: Claude <noreply@anthropic.com>
ORC is a new opt-in artifact format (`formats = ["orc"]`, `--format orc`,
`load(format="orc")`; never picked by `auto`). Two writer lanes:

- orc@py: pyarrow's ORCWriter (the Apache ORC C++ library), streaming the
  canonical's batches, zstd. Also the loader's reader and the in-process
  conformance reader.
- orc@rs: orc-rust 0.9.0 (arrow 59, sharing the sidecar's =59.2 pin) as the
  `orc-write` / `orc-read` sidecar binaries, zstd.

No column is converted for either library: a type it does not write fails the
write, which the build records as ORC unavailable for that dataset; orc-rust's
panic on such a type is its report. With `formats = "all"`, a format with no
installed writer is left out with a note instead of failing the build.

Co-Authored-By: Claude <noreply@anthropic.com>
Avro object container files are a new opt-in format (`formats = ["avro"]`,
`--format avro`), with two sidecar writer lanes and no Python one, since
pyarrow neither reads nor writes Avro:

- avro@rs: arrow-avro =59.2 as the `avro-write` / `avro-read` binaries.
- avro@java: Arrow Java 19.0.0's Avro adapter over Apache Avro 1.12.1, in
  the new `avro-java` Gradle project.

Both write zstandard and the same fixed sync marker, so a rebuild gives the
same bytes: Avro Java takes the marker as an argument; arrow-avro draws one
at random with no way to choose it, so the Rust lane overwrites it in place
after the header and each block, and a test checks every other byte is
arrow-avro's own. No column is converted for either library.

A format raincloud only serves by path now loads like any other: `path()`
works and its readers raise MissingDependency. `run_exporters` and `plan`
without `formats` write the install's formats.

Co-Authored-By: Claude <noreply@anthropic.com>
Nimble is a new opt-in format (`formats = ["nimble"]`, `--format nimble`),
served by path, with one lane, nimble@cpp: Nimble has one implementation
and no releases, so it is built from source.

- sidecars/nimble: raincloud-nimble, a codec over upstream Nimble's
  VeloxWriter (default options) and VeloxReader, crossing Arrow into Velox
  through the C data interface (nanoarrow 0.9.0 IPC streams, Velox's Arrow
  bridge). build.sh builds it inside a Nimble checkout at a pinned commit
  (upstream acead744 plus host build fixes), with sha256-pinned
  dependencies, and records its provenance.
- sidecars/rust: nimble-write / nimble-read stream the canonical into the
  tool and read the file back, then compare and report like every Rust
  lane. View types travel as their plain equivalents, which nanoarrow's IPC
  reader does not read yet and Velox imports as the same type.

A sidecar cell may name a helper binary (_registry.SIDECAR_HELPERS): the
cell needs it installed, and its sha256 joins the writer's toolchain, so a
rebuilt tool retries a recorded failure. The built-in writer order gains
`cpp` so every format has a writer it names; a test keeps it that way.

Co-Authored-By: Claude <noreply@anthropic.com>
Where equality turns on representation rather than data, the Python, Rust
and JVM comparators now follow one rule set, declared once in
sidecars/compare_cases: pairs of Arrow files and the verdict each must get,
written by generate.py and read by all three lanes' tests.

- A union of exactly null and T is a nullable T, at any depth (how Arrow
  Java's Avro adapter returns Avro's nullable fields).
- Zoned timestamps compare by instant, whatever zone labels them; naive and
  zoned still differ.
- An integer and a scale-0 decimal holding the same values are equal.

Both ORC lanes now widen what ORC cannot hold before writing, always rather
than by the data's range, so a dataset's ORC schema never changes with its
values: uint8 -> int16, uint16 -> int32, uint32 -> int64,
uint64 -> decimal(20, 0), and view types to their plain types. orc-rust
writes no decimals, so a uint64 column is still unavailable there.

Co-Authored-By: Claude <noreply@anthropic.com>
- A generated table is one of a group its generator writes at once:
  building one builds the whole group in the same formats, and the
  generator output is removed once the group has built, unless keep_raw
  keeps it. `--only` builds just the tables named.
- A v1 catalog builds and loads as in 0.3.0: every format its recipe
  lists, nothing cleaned, `auto` in the old vortex, parquet, arrow order.
- `export.formats` is no longer part of a recipe's fingerprint (an install
  chooses what it builds), and the four recipes that restated the old
  default drop it.

Co-Authored-By: Claude <noreply@anthropic.com>
The nimble@cpp lane's binaries are now C++ (raincloud-export-nimble-cpp,
raincloud-read-nimble-cpp): two callbacks over upstream Nimble's
VeloxWriter and VeloxReader, linked to nimble-ffi, a new member of the Rust
sidecar workspace that runs the sidecar contract as every Rust lane does --
it reads the canonical with arrow-rs and hands its batches across through
the Arrow C stream interface, takes the read-back the same way, compares and
reports. Velox's own Arrow bridge imports and exports the batches, so view
types reach Nimble as they are and nothing is cast on the way; build.sh
links both halves against the host's libzstd.

This replaces the IPC-stream transport through a separate codec, its view
cast and the helper-binary mechanism it needed: the lane's toolchain is
again just its binary, which now contains Nimble.

Co-Authored-By: Claude <noreply@anthropic.com>
…ingerprint, v1

Co-Authored-By: Claude <noreply@anthropic.com>
The CHANGELOG heading and CITATION release date are set when the release is cut.

Co-Authored-By: Claude <noreply@anthropic.com>
Nimble's reader can hand back constant vectors (an all-null column) and
dictionary vectors, which Velox exports as run-end-encoded and dictionary
arrays. The read stream declares the plain type once, so such a batch did
not match it; arrow-rs then refused the batch (with an assertion, not an
error). Each batch is flattened, at every depth, before export.

Co-Authored-By: Claude <noreply@anthropic.com>
…tings

Every Parquet writer now gets the same options: the recipe's
write.compression and write.statistics (which the sidecar writers used to
choose for themselves) and three install settings,
RAINCLOUD_PARQUET_PAGE_INDEX, RAINCLOUD_PARQUET_PAGE_BYTES and
RAINCLOUD_PARQUET_PAGE_ROWS. They are resolved once in Python
(spec.ParquetOptions) and passed to the sidecars in one form.

An unset page setting leaves each library's own default, so no file changes
unless a setting is given: pyarrow and Hardwood write no page index, arrow-rs
and parquet-java write one. A set one reaches every writer and is part of its
toolchain, so a recorded failure is retried under different options. A writer
whose library cannot do what is asked fails that export as a measurement
instead of writing something else: Hardwood 1.1.0.Beta1 writes no page index,
has no page row limit and cannot turn statistics off; parquet-java writes a
page index whenever statistics are on; parquet-arrow-java has no LZ4 or Brotli.

Co-Authored-By: Claude <noreply@anthropic.com>
…hecksums as settings

Six more install settings reach every Parquet writer:
RAINCLOUD_PARQUET_COMPRESSION_LEVEL, _STATISTICS_COLUMNS (statistics only
for the first N leaf columns), _PAGE_INDEX_COLUMNS (page statistics only for
the first N leaf columns, chunk statistics for all), _DICTIONARY,
_DICTIONARY_PAGE_BYTES and _PAGE_CHECKSUMS. Each is declared once in
spec.py, which derives the parsing, the sidecar environment and the
toolchain entry from that one list; contradictory settings are refused for
every lane before any writer runs.

Unset, each library's default still applies and no file changes. A writer
whose library cannot honour a set one refuses it as a measured failure:
pyarrow writes a page index for all columns with statistics or none;
arrow-rs writes no page checksums; parquet-arrow-java has no level and no
per-column statistics; Hardwood has neither, nor a dictionary page limit,
and always writes checksums. sidecars/README.md tabulates the lot, and one
test writes every setting with every writer.

Co-Authored-By: Claude <noreply@anthropic.com>
…format

The other exported formats get install settings declared and read the way
Parquet's are (spec.FORMAT_SETTINGS), passed to sidecars in one canonical
form and recorded in the writer's toolchain when set:

- ORC: RAINCLOUD_ORC_COMPRESSION, _COMPRESSION_STRATEGY, _STRIPE_BYTES,
  _COMPRESSION_BLOCK_BYTES.
- Avro: RAINCLOUD_AVRO_COMPRESSION (every Avro codec; avro@rs now builds
  arrow-avro's deflate, snappy, bzip2 and xz, and avro@java ships Avro's
  optional snappy-java 1.1.10.8 and xz 1.10), _COMPRESSION_LEVEL,
  _BLOCK_BYTES.
- Vortex: RAINCLOUD_VORTEX_COMPACT (BtrBlocks' compact encodings, as
  vortex-data's Python writer builds them), _ROW_BLOCK_ROWS,
  _DATA_BLOCK_BYTES.

Unset, every writer writes the bytes it wrote before (the codec stays
zstd). A writer whose library cannot honour a set one refuses it as a
measured failure: orc-rust has no compression strategy, arrow-avro no level
or block size, vortex-data's Python writer no block settings, vortex-jni no
write strategy at all.

Co-Authored-By: Claude <noreply@anthropic.com>
parquet-arrow-java 0.3.0 adds a compression level, per-column statistics
and LZ4_RAW, so parquet@java now honours RAINCLOUD_PARQUET_COMPRESSION_LEVEL,
_STATISTICS_COLUMNS (statistics off past the first N leaf columns of the
file's Parquet schema) and an lz4 recipe, where it used to refuse them. It
still refuses Brotli, and leaving out or limiting the page index, which are
parquet-java's own limits.

parquet@hardwood ships brotli4j 1.23.0 (the version Hardwood's BOM pins)
with its native library for Linux and macOS on x86-64 and aarch64, so a
Brotli recipe no longer fails there. A test now has pyarrow read every
Parquet writer's file in every codec.

Co-Authored-By: Claude <noreply@anthropic.com>
A /raincloud-write-settings skill is the procedure for writing a file with
encoder settings: what unset means, that a setting changes only files this
install writes (so an existing one must be re-exported or rebuilt), which
writer refuses what and how to pick another, contradictions, and how to
check a Parquet page index. SKILLS.md gains the matching playbook; the
build, export and load skills and the skills index point at it and list
every format; the README shows a page-indexed build from the shell and from
Python, both run end to end; AGENTS.md points at the skill. The SKILLS.md
Vortex playbook no longer says the defaults export Parquet and Vortex.

The changelog's 0.3.1 section is dated 2026-10-08 and says the compliance
ledger is still 0.3.0's; CITATION.cff carries the date. The Nimble lane's
docs name the fork its pinned commit lives on.

Co-Authored-By: Claude <noreply@anthropic.com>
…pt-in formats

A load that builds asked the build for just the format it wanted, so a v1
dataset loaded with build=True wrote only Vortex where 0.3.0 wrote Parquet
and Vortex. A v1 catalog predates install formats, so its load-triggered
build now names no format and writes what the recipe lists.

Two opt-in tests still expected 0.3.0's choices: the portable base-install
probe expected auto to pick Parquet, which an install that builds only
Vortex no longer offers (it gets the canonical Arrow), and the real-build
test expected a load to write Parquet (it writes the format it serves).

Co-Authored-By: Claude <noreply@anthropic.com>
@mprammer
mprammer merged commit 1ac425a into develop Oct 8, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant