Repository navigation
Release 0.3.1: ORC, Avro and Nimble; opt-in formats; encoder settings, including the Parquet page index - #22
Merged
Merged
Conversation
`_registry.FORMATS` now declares every artifact format: its file extension, whether `auto` may pick it, and the in-process reader that opens it. The extension map, reader capabilities, `auto` order, snapshot generation, status columns, publish set, manifest validation and CLI help derive from it instead of naming parquet and vortex. The loader's per-format open code moves behind a table in `_readers`, so a format with no in-process reader raises MissingDependency instead of falling through to the Vortex reader. sources.schema.json cannot import the registry; a test gates its enums. Co-Authored-By: Claude <noreply@anthropic.com>
A build now writes only Vortex unless the new `formats` setting (RAINCLOUD_FORMATS, or "all") or `raincloud build --format` asks for more, and a load that names another format builds just that one. Every dataset offers every exported format; a recipe's export.formats only narrows that. `auto` picks among the formats the install builds, Vortex then Parquet, and falls back to the canonical Arrow when neither can be had. Once its files are written, a successful build removes the raw download and the canonical Arrow unless `keep_raw` / `keep_canonical` keep them. A canonical from which nothing was written is the dataset's only file and stays. Generated datasets keep their generator output, which a group shares. `status` reports against the install's settings; `export` without --format writes the install's formats. Tests pin the 0.3.0 settings where they exercise pipeline mechanics, and test_build_formats covers the new defaults. Co-Authored-By: Claude <noreply@anthropic.com>
ORC is a new opt-in artifact format (`formats = ["orc"]`, `--format orc`, `load(format="orc")`; never picked by `auto`). Two writer lanes: - orc@py: pyarrow's ORCWriter (the Apache ORC C++ library), streaming the canonical's batches, zstd. Also the loader's reader and the in-process conformance reader. - orc@rs: orc-rust 0.9.0 (arrow 59, sharing the sidecar's =59.2 pin) as the `orc-write` / `orc-read` sidecar binaries, zstd. No column is converted for either library: a type it does not write fails the write, which the build records as ORC unavailable for that dataset; orc-rust's panic on such a type is its report. With `formats = "all"`, a format with no installed writer is left out with a note instead of failing the build. Co-Authored-By: Claude <noreply@anthropic.com>
Avro object container files are a new opt-in format (`formats = ["avro"]`, `--format avro`), with two sidecar writer lanes and no Python one, since pyarrow neither reads nor writes Avro: - avro@rs: arrow-avro =59.2 as the `avro-write` / `avro-read` binaries. - avro@java: Arrow Java 19.0.0's Avro adapter over Apache Avro 1.12.1, in the new `avro-java` Gradle project. Both write zstandard and the same fixed sync marker, so a rebuild gives the same bytes: Avro Java takes the marker as an argument; arrow-avro draws one at random with no way to choose it, so the Rust lane overwrites it in place after the header and each block, and a test checks every other byte is arrow-avro's own. No column is converted for either library. A format raincloud only serves by path now loads like any other: `path()` works and its readers raise MissingDependency. `run_exporters` and `plan` without `formats` write the install's formats. Co-Authored-By: Claude <noreply@anthropic.com>
Nimble is a new opt-in format (`formats = ["nimble"]`, `--format nimble`), served by path, with one lane, nimble@cpp: Nimble has one implementation and no releases, so it is built from source. - sidecars/nimble: raincloud-nimble, a codec over upstream Nimble's VeloxWriter (default options) and VeloxReader, crossing Arrow into Velox through the C data interface (nanoarrow 0.9.0 IPC streams, Velox's Arrow bridge). build.sh builds it inside a Nimble checkout at a pinned commit (upstream acead744 plus host build fixes), with sha256-pinned dependencies, and records its provenance. - sidecars/rust: nimble-write / nimble-read stream the canonical into the tool and read the file back, then compare and report like every Rust lane. View types travel as their plain equivalents, which nanoarrow's IPC reader does not read yet and Velox imports as the same type. A sidecar cell may name a helper binary (_registry.SIDECAR_HELPERS): the cell needs it installed, and its sha256 joins the writer's toolchain, so a rebuilt tool retries a recorded failure. The built-in writer order gains `cpp` so every format has a writer it names; a test keeps it that way. Co-Authored-By: Claude <noreply@anthropic.com>
Where equality turns on representation rather than data, the Python, Rust and JVM comparators now follow one rule set, declared once in sidecars/compare_cases: pairs of Arrow files and the verdict each must get, written by generate.py and read by all three lanes' tests. - A union of exactly null and T is a nullable T, at any depth (how Arrow Java's Avro adapter returns Avro's nullable fields). - Zoned timestamps compare by instant, whatever zone labels them; naive and zoned still differ. - An integer and a scale-0 decimal holding the same values are equal. Both ORC lanes now widen what ORC cannot hold before writing, always rather than by the data's range, so a dataset's ORC schema never changes with its values: uint8 -> int16, uint16 -> int32, uint32 -> int64, uint64 -> decimal(20, 0), and view types to their plain types. orc-rust writes no decimals, so a uint64 column is still unavailable there. Co-Authored-By: Claude <noreply@anthropic.com>
- A generated table is one of a group its generator writes at once: building one builds the whole group in the same formats, and the generator output is removed once the group has built, unless keep_raw keeps it. `--only` builds just the tables named. - A v1 catalog builds and loads as in 0.3.0: every format its recipe lists, nothing cleaned, `auto` in the old vortex, parquet, arrow order. - `export.formats` is no longer part of a recipe's fingerprint (an install chooses what it builds), and the four recipes that restated the old default drop it. Co-Authored-By: Claude <noreply@anthropic.com>
The nimble@cpp lane's binaries are now C++ (raincloud-export-nimble-cpp, raincloud-read-nimble-cpp): two callbacks over upstream Nimble's VeloxWriter and VeloxReader, linked to nimble-ffi, a new member of the Rust sidecar workspace that runs the sidecar contract as every Rust lane does -- it reads the canonical with arrow-rs and hands its batches across through the Arrow C stream interface, takes the read-back the same way, compares and reports. Velox's own Arrow bridge imports and exports the batches, so view types reach Nimble as they are and nothing is cast on the way; build.sh links both halves against the host's libzstd. This replaces the IPC-stream transport through a separate codec, its view cast and the helper-binary mechanism it needed: the lane's toolchain is again just its binary, which now contains Nimble. Co-Authored-By: Claude <noreply@anthropic.com>
…ingerprint, v1 Co-Authored-By: Claude <noreply@anthropic.com>
The CHANGELOG heading and CITATION release date are set when the release is cut. Co-Authored-By: Claude <noreply@anthropic.com>
Nimble's reader can hand back constant vectors (an all-null column) and dictionary vectors, which Velox exports as run-end-encoded and dictionary arrays. The read stream declares the plain type once, so such a batch did not match it; arrow-rs then refused the batch (with an assertion, not an error). Each batch is flattened, at every depth, before export. Co-Authored-By: Claude <noreply@anthropic.com>
…tings Every Parquet writer now gets the same options: the recipe's write.compression and write.statistics (which the sidecar writers used to choose for themselves) and three install settings, RAINCLOUD_PARQUET_PAGE_INDEX, RAINCLOUD_PARQUET_PAGE_BYTES and RAINCLOUD_PARQUET_PAGE_ROWS. They are resolved once in Python (spec.ParquetOptions) and passed to the sidecars in one form. An unset page setting leaves each library's own default, so no file changes unless a setting is given: pyarrow and Hardwood write no page index, arrow-rs and parquet-java write one. A set one reaches every writer and is part of its toolchain, so a recorded failure is retried under different options. A writer whose library cannot do what is asked fails that export as a measurement instead of writing something else: Hardwood 1.1.0.Beta1 writes no page index, has no page row limit and cannot turn statistics off; parquet-java writes a page index whenever statistics are on; parquet-arrow-java has no LZ4 or Brotli. Co-Authored-By: Claude <noreply@anthropic.com>
…hecksums as settings Six more install settings reach every Parquet writer: RAINCLOUD_PARQUET_COMPRESSION_LEVEL, _STATISTICS_COLUMNS (statistics only for the first N leaf columns), _PAGE_INDEX_COLUMNS (page statistics only for the first N leaf columns, chunk statistics for all), _DICTIONARY, _DICTIONARY_PAGE_BYTES and _PAGE_CHECKSUMS. Each is declared once in spec.py, which derives the parsing, the sidecar environment and the toolchain entry from that one list; contradictory settings are refused for every lane before any writer runs. Unset, each library's default still applies and no file changes. A writer whose library cannot honour a set one refuses it as a measured failure: pyarrow writes a page index for all columns with statistics or none; arrow-rs writes no page checksums; parquet-arrow-java has no level and no per-column statistics; Hardwood has neither, nor a dictionary page limit, and always writes checksums. sidecars/README.md tabulates the lot, and one test writes every setting with every writer. Co-Authored-By: Claude <noreply@anthropic.com>
…format The other exported formats get install settings declared and read the way Parquet's are (spec.FORMAT_SETTINGS), passed to sidecars in one canonical form and recorded in the writer's toolchain when set: - ORC: RAINCLOUD_ORC_COMPRESSION, _COMPRESSION_STRATEGY, _STRIPE_BYTES, _COMPRESSION_BLOCK_BYTES. - Avro: RAINCLOUD_AVRO_COMPRESSION (every Avro codec; avro@rs now builds arrow-avro's deflate, snappy, bzip2 and xz, and avro@java ships Avro's optional snappy-java 1.1.10.8 and xz 1.10), _COMPRESSION_LEVEL, _BLOCK_BYTES. - Vortex: RAINCLOUD_VORTEX_COMPACT (BtrBlocks' compact encodings, as vortex-data's Python writer builds them), _ROW_BLOCK_ROWS, _DATA_BLOCK_BYTES. Unset, every writer writes the bytes it wrote before (the codec stays zstd). A writer whose library cannot honour a set one refuses it as a measured failure: orc-rust has no compression strategy, arrow-avro no level or block size, vortex-data's Python writer no block settings, vortex-jni no write strategy at all. Co-Authored-By: Claude <noreply@anthropic.com>
parquet-arrow-java 0.3.0 adds a compression level, per-column statistics and LZ4_RAW, so parquet@java now honours RAINCLOUD_PARQUET_COMPRESSION_LEVEL, _STATISTICS_COLUMNS (statistics off past the first N leaf columns of the file's Parquet schema) and an lz4 recipe, where it used to refuse them. It still refuses Brotli, and leaving out or limiting the page index, which are parquet-java's own limits. parquet@hardwood ships brotli4j 1.23.0 (the version Hardwood's BOM pins) with its native library for Linux and macOS on x86-64 and aarch64, so a Brotli recipe no longer fails there. A test now has pyarrow read every Parquet writer's file in every codec. Co-Authored-By: Claude <noreply@anthropic.com>
A /raincloud-write-settings skill is the procedure for writing a file with encoder settings: what unset means, that a setting changes only files this install writes (so an existing one must be re-exported or rebuilt), which writer refuses what and how to pick another, contradictions, and how to check a Parquet page index. SKILLS.md gains the matching playbook; the build, export and load skills and the skills index point at it and list every format; the README shows a page-indexed build from the shell and from Python, both run end to end; AGENTS.md points at the skill. The SKILLS.md Vortex playbook no longer says the defaults export Parquet and Vortex. The changelog's 0.3.1 section is dated 2026-10-08 and says the compliance ledger is still 0.3.0's; CITATION.cff carries the date. The Nimble lane's docs name the fork its pinned commit lives on. Co-Authored-By: Claude <noreply@anthropic.com>
…pt-in formats A load that builds asked the build for just the format it wanted, so a v1 dataset loaded with build=True wrote only Vortex where 0.3.0 wrote Parquet and Vortex. A v1 catalog predates install formats, so its load-triggered build now names no format and writes what the recipe lists. Two opt-in tests still expected 0.3.0's choices: the portable base-install probe expected auto to pick Parquet, which an install that builds only Vortex no longer offers (it gets the canonical Arrow), and the real-build test expected a load to write Parquet (it writes the format it serves). Co-Authored-By: Claude <noreply@anthropic.com>
mprammer
force-pushed
the
mp/release-0.3.1
branch
from
October 8, 2026 21:07
6792c7d to
1cb0376
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cuts 0.3.1. Three file formats join Parquet, Vortex and Arrow — ORC, Avro and Nimble — and every format's writers now take the same encoder settings, so a dataset can be built as Parquet with a page index. Formats become opt-in per install: a plain install builds only Vortex, and a build removes its raw download and canonical Arrow once the files are written.
Encoder settings. Each format's writers read one set of settings from the environment (
RAINCLOUD_PARQUET_*,RAINCLOUD_ORC_*,RAINCLOUD_AVRO_*,RAINCLOUD_VORTEX_*), declared once inraincloud/pipeline/spec.py. For Parquet that is a page index on every column or only the first N, statistics for only the first N columns, compression level, page size and rows, dictionaries and page checksums. Unset, each writer library's default applies, and no file changes: with nothing set, every writer was checked to write the bytes it wrote before, except that parquet-java's list of a column's encodings can come out in another order, as it already could from one JVM run to the next. A writer whose library cannot do what a setting asks refuses it, and the format is recorded unavailable with the reason, rather than written another way;sidecars/README.mdtabulates which writer honours what. The recipe'swrite.compressionandwrite.statisticsnow reach the sidecar writers too, which used to pick their own.parquet@javamoves to parquet-arrow-java 0.3.0 for a compression level, per-column statistics and LZ4_RAW, andparquet@hardwoodships brotli4j.New formats. ORC is written by pyarrow and orc-rust, Avro by arrow-avro and Arrow Java's adapter, and Nimble by upstream Nimble's C++, built from source from a pinned commit on a fork carrying build fixes (
sidecars/nimble/build.sh).autostill picks Vortex, then Parquet; the new formats are built and loaded only when named.Also. One comparison rule set shared by the Python, Rust and JVM comparators (a union of
nullandTis a nullableT; zoned timestamps compare by instant). A generated table builds its whole group. The recipe hash ignoresexport.formats. v1 catalogs keep their old behaviour.Docs for agents. A new
/raincloud-write-settingsskill and matchingSKILLS.mdplaybook; the build, export and load skills,AGENTS.mdand the README cover the settings and the new formats.Compliance.
docs/v2/compliance.jsonis still 0.3.0's measurement; the new lanes and settings are measured in a later release (noted in the changelog).Tag
v0.3.1and the GitHub release follow on merge.🤖 Generated with Claude Code