Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
2d5aa3f
formats: declare each artifact format once, in the registry
mprammer Oct 1, 2026
292deb3
formats: opt in per install; a build cleans up after itself
mprammer Oct 2, 2026
0f96b0d
formats: add ORC, written by pyarrow and by orc-rust
mprammer Oct 2, 2026
e04c05e
formats: add Avro, written by arrow-avro and by Arrow Java's adapter
mprammer Oct 2, 2026
7c888c8
formats: add Nimble, written and read by upstream Nimble's C++
mprammer Oct 2, 2026
c8e8429
compare: one rule set for representation, shared by every comparator
mprammer Oct 2, 2026
a94dbac
build: generator groups, v1 as before, offered formats out of the recipe
mprammer Oct 2, 2026
b748d2b
nimble: hand batches to upstream Nimble in memory, not over a pipe
mprammer Oct 2, 2026
a126795
changelog: comparison rules, generator groups, ORC widening, recipe f…
mprammer Oct 2, 2026
59142af
release 0.3.1: version
mprammer Oct 2, 2026
e0b2da9
nimble: read back flat batches, matching the stream's declared schema
mprammer Oct 2, 2026
a85d922
parquet: one set of write options in every writer, page layout as set…
mprammer Oct 7, 2026
68fc7c4
parquet: compression level, per-column statistics, dictionaries and c…
mprammer Oct 7, 2026
836e042
orc, avro, vortex: write settings given alike to every writer of the …
mprammer Oct 7, 2026
785fa1c
parquet@java: parquet-arrow-java 0.3.0; parquet@hardwood writes Brotli
mprammer Oct 7, 2026
589fc9d
docs: encoder settings for users and agents; release 0.3.1 dated
mprammer Oct 8, 2026
1cb0376
load: a v1 dataset builds as 0.3.0 did; wheel and network tests for o…
mprammer Oct 8, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion .agents/skills/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ Wrappers around `python -m raincloud.pipeline.<module>`. Side-effecting ones set
| `/raincloud-build` | `raincloud.pipeline.build` | Full pipeline (fetch → … → write_canonical → validate → run_exporters) for one or more slugs. |
| `/raincloud-fetch` | `raincloud.pipeline.fetch` | Download raw bytes only. |
| `/raincloud-extract` | `raincloud.pipeline.extract` | Unpack archives into the recipe's scratch directory. |
| `/raincloud-export` | `raincloud.pipeline.export` | Re-derive Parquet/Vortex from the canonical Arrow already on disk, without refetching; the refresh path for one format. |
| `/raincloud-export` | `raincloud.pipeline.export` | Re-derive any format from the canonical Arrow already on disk, without refetching; the refresh path for one format, or for new encoder settings. |
| `/raincloud-convert` | `raincloud.pipeline.convert` | Re-encode Vortex with the Python writer (v1 catalogs: from Parquet). For v2, prefer `/raincloud-export --format vortex`. |
| `/raincloud-hydrate` | `raincloud.pipeline.hydrate` | Build a `<parent>-hydrated` dataset (URL columns fetched from the open web) with the safe defaults, or write a scratch sample with non-default options (`--limit`/`--block`/`--urlhaus`/`--max-bytes`/`--timeout`/bypass), never published or served. Side-effecting (outbound HTTP); safety-filter-gated; `disable-model-invocation: true`. |
| `/raincloud-docs` | `raincloud.pipeline.docs` | Regenerate derived docs. *(model-invocable — regen is mostly idempotent.)* |
Expand All @@ -35,6 +35,7 @@ These guide multi-step procedures from [`SKILLS.md`](../context/SKILLS.md). Defa
| `/raincloud-add-handler` | Writing a new transform handler under `raincloud/pipeline/handlers/`. |
| `/raincloud-add-kaggle-tos` | Adding a Kaggle dataset gated behind a one-time ToS click-through. |
| `/raincloud-promote-variant` | JSON → VARIANT via the transform recipe and a rebuild. |
| `/raincloud-write-settings` | Write files with chosen encoder settings: a Parquet page index (all or the first N columns), compression level, statistics, dictionaries, checksums; ORC/Avro codecs; Vortex compact. |
| `/raincloud-debug-build` | Diagnostic checklist for a failing build — isolate which stage broke. |
| `/raincloud-large-build` | Run a memory- or runtime-heavy build safely (caps, nohup, logging). *(side-effecting — `disable-model-invocation: true`.)* |
| `/raincloud-remove-dataset` | Remove a dataset from the manifest and clean up its outputs. *(destructive — `disable-model-invocation: true`.)* |
Expand Down
6 changes: 5 additions & 1 deletion .agents/skills/raincloud-build/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
name: raincloud-build
description: Run the full Raincloud pipeline (fetch → extract → parse → transform → canonical Arrow → validate → exports) for one or more dataset slugs. Use when the user asks to build a dataset, rebuild a slug, or process a batch.
argument-hint: <slug>... | --all [--strict] [--clean-workdir] [--retry-errors]
argument-hint: <slug>... | --all [--format FORMAT] [--only] [--strict] [--clean-workdir] [--retry-errors]
disable-model-invocation: true
allowed-tools: Bash(python -m raincloud.pipeline.build *)
---
Expand All @@ -20,6 +20,10 @@ Modifiers:
- `--strict` — make validation drift an error. By default, row-count mismatches are warnings; use that default for first builds with estimated counts.
- `--clean-workdir` — clear the selected `.recipes/<recipe-hash>/<slug>/` scratch directory after each successful build. Essential for large batch runs (Public BI decompressed CSVs can hit ~100 GB).
- `--retry-errors` — attempt a format even when its writer, with this toolchain, already failed to write it at this recipe (see below). Without it that format is skipped.
- `--format FORMAT` (repeatable or comma-separated) — write these formats instead of the install's `formats` setting (only `vortex` by default; `parquet`, `orc`, `avro`, `nimble`, or `arrow` to keep the canonical). Only the format is taken: a writer suffix (`parquet@rs`) is dropped, and the writer comes from `export.priority`.
- `--only` — for a generated table, build its whole group but keep only the tables named.

Encoder settings — a Parquet page index, statistics for the first N columns, a compression level, dictionaries, page checksums, ORC/Avro codecs, Vortex compact encodings — are environment settings every writer of the format reads (`RAINCLOUD_PARQUET_*`, `RAINCLOUD_ORC_*`, `RAINCLOUD_AVRO_*`, `RAINCLOUD_VORTEX_*`); unset is each library's default. Pass them in the build's environment; see `/raincloud-write-settings` for what each does, which writer refuses which, and how to check the result.

Before running:
- **Confirm with the user** before triggering anything non-trivial. JSONBench 100M ≈ 6 h, Wikipedia Structured Contents → ~70 GB parquet, OSM Germany ~45 min per kind. Small (<100 MB) parquets are fine without asking. (See [AGENTS.md "Confirm before rebuilding"](../../context/AGENTS.md#confirm-before-rebuilding).)
Expand Down
7 changes: 4 additions & 3 deletions .agents/skills/raincloud-export/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
name: raincloud-export
description: Re-derive a dataset's Parquet and/or Vortex files from the canonical Arrow file already on disk, without refetching or re-transforming. Use when a change touches only the export stage (row-group sizing, a codec, a Vortex upgrade) or to refresh one format.
argument-hint: <slug>... | --all [--format parquet|vortex|parquet@rs|...] [--dry-run] [--retry-errors]
description: Re-derive a dataset's Parquet, Vortex, ORC, Avro or Nimble files from the canonical Arrow file already on disk, without refetching or re-transforming. Use when a change touches only the export stage (row-group sizing, a codec, an encoder setting such as a Parquet page index, a writer upgrade) or to refresh one format.
argument-hint: <slug>... | --all [--format parquet|vortex|orc|avro|nimble|parquet@rs|...] [--dry-run] [--retry-errors]
disable-model-invocation: true
allowed-tools: Bash(python -m raincloud.pipeline.export *)
---
Expand All @@ -17,9 +17,10 @@ Selection (one required):
- `--all` — every dataset. Hours of work on a full store; confirm first.

Modifiers:
- `--format FORMAT` (repeatable) — export only this format, replacing the spec's `export.formats` for this run: `parquet` or `vortex`. The writer is chosen by `export.priority` as in a build. It overrides a dataset whose policy leaves the format out, so check `python -m raincloud.pipeline.list_datasets --no-vortex --json` before forcing Vortex.
- `--format FORMAT` (repeatable) — export only this format, replacing the install's formats for this run: `parquet`, `vortex`, `orc`, `avro` or `nimble`. The writer is chosen by `export.priority` as in a build. It overrides a dataset whose policy leaves the format out, so check `python -m raincloud.pipeline.list_datasets --no-vortex --json` before forcing Vortex.
- `--format parquet@rs` (or another `<format>@<writer>`) — use that writer for this run. The file is still `parquet/<slug>.parquet`, but its bytes and sha256 change, so it no longer matches the catalog until a maintainer regenerates it. Confirm before doing this to datasets others read.
- `--dry-run` — list what would be exported and exit.
- Encoder settings come from the environment (`RAINCLOUD_PARQUET_PAGE_INDEX=1`, `RAINCLOUD_PARQUET_STATISTICS_COLUMNS=100`, `RAINCLOUD_PARQUET_COMPRESSION_LEVEL=9`, ...): this is the cheapest way to rewrite a file with them, since the canonical is the input. A writer that cannot honour a set one refuses (`[unavailable]`, `<writer> cannot honour <VARIABLE>=...`); name another with `--format <fmt>@<writer>`. See `/raincloud-write-settings`.
- `--retry-errors` — attempt a format even when its writer, with this toolchain, already failed to write it at this recipe (see below).

Behavior:
Expand Down
10 changes: 8 additions & 2 deletions .agents/skills/raincloud-load/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
---
name: raincloud-load
description: Read prepared Raincloud artifacts, inspect catalog metadata, or load batches from local storage or a configured mirror.
argument-hint: <slug> [--format auto|arrow|parquet|vortex]
argument-hint: <slug> [--format auto|arrow|parquet|vortex|orc|avro|nimble]
disable-model-invocation: true
allowed-tools: Bash(python -m raincloud describe *), Bash(python -m raincloud list *), Bash(python -m raincloud load *), Bash(python -m raincloud config show), Bash(python -m raincloud capabilities), Bash(python examples/use_loader.py *)
---
Expand Down Expand Up @@ -37,7 +37,13 @@ resolve the artifact as needed. `.to_arrow()` materializes the whole table.
selected format for any engine (DuckDB, Polars, pyarrow). `.to_vortex()` requires
`[vortex]` and a Vortex artifact. `.path()`, `.batches()`, `.to_arrow()` and
`.dataset()` retain the selected artifact; `.to_vortex()` resolves the dataset's
Vortex file. Each format is one file; `describe` shows which writer made it.
Vortex file. Each format is one file; `describe` shows which writer made it. ORC reads
through pyarrow; Avro and Nimble are served by `.path()` only.

A load serves the file it finds. Encoder settings (a Parquet page index, compression
level, ...) apply only when this install writes a file, so a file already on disk, in
the store or on a mirror is served as it was written, and `--build` builds only a file
that is missing. To get one written with settings, see `/raincloud-write-settings`.

Use `raincloud capabilities` for installed reader modules and `raincloud config
show` for effective paths. Optional TOML settings and environment overrides share
Expand Down
95 changes: 95 additions & 0 deletions .agents/skills/raincloud-write-settings/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
---
name: raincloud-write-settings
description: Write a dataset's files with chosen encoder settings — a Parquet page index (for every column or the first N), Parquet compression level, statistics, page sizes, dictionaries or page checksums; the ORC or Avro codec and level; Vortex compact encodings or block sizes. Use when the user wants page statistics / a page index / column indexes in Parquet, a different codec or compression level, or to compare encoder settings across writers.
argument-hint: <slug> [setting=value ...]
---

Encoder settings are **install settings in the environment**, read by every writer of the
format alike (`raincloud/pipeline/spec.py`: `_PARQUET_SETTINGS`, `FORMAT_SETTINGS`). They are
not recipe fields and not load options. The full list, with defaults, is the
`RAINCLOUD_PARQUET_*`, `RAINCLOUD_ORC_*`, `RAINCLOUD_AVRO_*` and `RAINCLOUD_VORTEX_*` rows of
the table in [AGENTS.md "Data locations"](../../context/AGENTS.md#data-locations); which
writer honours which is tabulated in [`sidecars/README.md`](../../../sidecars/README.md).

## What to know before changing a setting

- **Unset means each writer library's own default**, and is what every published file was
written with. pyarrow (the default Parquet writer, `parquet@py`) writes **no page
index** unless asked; arrow-rs and parquet-java write one.
- **A setting applies to files this install writes.** It does not change a file already on
disk, in this machine's store, or on a mirror, and `raincloud load` serves whichever of
those it finds first. To get a file with the setting, write it again (below).
- **A file written with a setting differs from the catalog's** (another sha256). The build
record (`<data_dir>/builds.json`) records it, and the loader serves this install's file.
Do not do this to a shared store others read without asking.
- **A writer that cannot honour a set value refuses**, rather than writing something else:
the format is recorded unavailable with the reason (`[unavailable] <slug>/parquet:
parquet@py cannot honour RAINCLOUD_PARQUET_PAGE_INDEX_COLUMNS=100: ...`). When the default
writer refuses, pick another: `export --format parquet@rs` names it for one export; a
build takes only the format, so set the machine's writer preference instead
(`RAINCLOUD_EXPORT_PRIORITY=rs,py`; a recipe's or the catalog's `export.priority` outranks
it). arrow-rs (`rs`) is the writer for a page index on the first N columns only.
- **Contradictions are refused before any writer runs**: a page index while the recipe turns
statistics off, `RAINCLOUD_PARQUET_PAGE_INDEX=0` with `_PAGE_INDEX_COLUMNS`, a level for a
codec without levels, a level out of range. A malformed value names the variable.
- **A set value is part of the writer's toolchain**, so a failure recorded without it is
attempted again, and one recorded with it is skipped until something changes.

## Parquet page index

```bash
# every column; the default writer (pyarrow) does this
export RAINCLOUD_PARQUET_PAGE_INDEX=1
# or: page statistics for the first 100 leaf columns only, chunk statistics for all
# (arrow-rs only: pyarrow writes a page index for all columns or none)
export RAINCLOUD_PARQUET_PAGE_INDEX_COLUMNS=100
# or: all statistics, chunk and page, for the first 100 leaf columns only
# (pyarrow, arrow-rs and parquet-java; Hardwood refuses)
export RAINCLOUD_PARQUET_STATISTICS_COLUMNS=100
```

"Leaf columns" are the Parquet column chunks, in schema order: a struct or list contributes
one per leaf field.

## Writing the file again

```bash
# The canonical Arrow is on disk (keep_canonical, or a maintainer's store): re-derive only
python -m raincloud.pipeline.export <slug> --format parquet
# Otherwise rebuild (refetches unless the raw download was kept)
python -m raincloud.pipeline.build <slug> --format parquet
```

`raincloud load <slug> --format parquet --build` builds only when no file is found, so it
does not rewrite an existing one. Confirm with the user before rebuilding anything large
([AGENTS.md "Confirm before rebuilding"](../../context/AGENTS.md#confirm-before-rebuilding)).

## Checking the result

```python
import pyarrow.parquet as pq, raincloud
meta = pq.ParquetFile(raincloud.load("<slug>", format="parquet").path()).metadata
chunks = [meta.row_group(g).column(c) for g in range(meta.num_row_groups) for c in range(meta.num_columns)]
print(sum(c.has_column_index for c in chunks), "of", len(chunks), "column chunks have a page index")
```

`raincloud describe <slug>` shows the writer that made the file; an `[unavailable]` line or a
`FormatUnavailable` from `load` quotes a writer's refusal.

## Other settings

| want | set |
|---|---|
| Parquet zstd/gzip/brotli level | `RAINCLOUD_PARQUET_COMPRESSION_LEVEL` (zstd 1-22, gzip 0-9, brotli 0-11; Hardwood refuses) |
| Parquet page size / rows per page | `RAINCLOUD_PARQUET_PAGE_BYTES`, `RAINCLOUD_PARQUET_PAGE_ROWS` |
| Parquet without dictionaries | `RAINCLOUD_PARQUET_DICTIONARY=0` |
| Parquet page checksums | `RAINCLOUD_PARQUET_PAGE_CHECKSUMS=1` (arrow-rs refuses) or `=0` (Hardwood refuses) |
| ORC codec / stripes | `RAINCLOUD_ORC_COMPRESSION`, `RAINCLOUD_ORC_STRIPE_BYTES`, `RAINCLOUD_ORC_COMPRESSION_BLOCK_BYTES` |
| Avro codec / level / blocks | `RAINCLOUD_AVRO_COMPRESSION`, `RAINCLOUD_AVRO_COMPRESSION_LEVEL`, `RAINCLOUD_AVRO_BLOCK_BYTES` (the last two Avro Java only) |
| Vortex compact encodings | `RAINCLOUD_VORTEX_COMPACT=1` (vortex@jni refuses) |

The Parquet codec itself is the recipe's `write.compression`, which wins over the
environment; changing it is a recipe edit, which changes the recipe hash.

Context: [AGENTS.md "How a build works"](../../context/AGENTS.md#how-a-build-works),
[SKILLS.md](../../context/SKILLS.md#writing-files-with-chosen-encoder-settings).
6 changes: 4 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -72,8 +72,8 @@ jobs:
17
cache: gradle

- name: Build + test JVM sidecars (conformance-common + parquet@java, parquet@hardwood, vortex@jni lanes)
run: ./sidecars/java/gradlew -p sidecars/java :conformance-common:test :parquet-java:test :parquet-java:installDist :parquet-hardwood:test :parquet-hardwood:installDist :vortex-jni-reader:test :vortex-jni-reader:installDist
- name: Build + test JVM sidecars (conformance-common + parquet@java, parquet@hardwood, vortex@jni, avro@java lanes)
run: ./sidecars/java/gradlew -p sidecars/java :conformance-common:test :parquet-java:test :parquet-java:installDist :parquet-hardwood:test :parquet-hardwood:installDist :vortex-jni-reader:test :vortex-jni-reader:installDist :avro-java:test :avro-java:installDist

- uses: astral-sh/setup-uv@v5
with:
Expand All @@ -89,6 +89,8 @@ jobs:
RAINCLOUD_READER_PARQUET_HARDWOOD: ${{ github.workspace }}/sidecars/java/parquet-hardwood/build/install/raincloud-export-parquet-hardwood/bin/raincloud-read-parquet-hardwood
RAINCLOUD_SIDECAR_VORTEX_JNI: ${{ github.workspace }}/sidecars/java/vortex-jni-reader/build/install/raincloud-read-vortex-jni/bin/raincloud-export-vortex-jni
RAINCLOUD_READER_VORTEX_JNI: ${{ github.workspace }}/sidecars/java/vortex-jni-reader/build/install/raincloud-read-vortex-jni/bin/raincloud-read-vortex-jni
RAINCLOUD_SIDECAR_AVRO_JAVA: ${{ github.workspace }}/sidecars/java/avro-java/build/install/raincloud-export-avro-java/bin/raincloud-export-avro-java
RAINCLOUD_READER_AVRO_JAVA: ${{ github.workspace }}/sidecars/java/avro-java/build/install/raincloud-export-avro-java/bin/raincloud-read-avro-java
run: uv run --no-sync pytest -q tests/test_reader_fidelity.py tests/test_java_reader_schema_compare.py tests/test_jvm_sidecar_lanes.py

wheel:
Expand Down
Loading
Loading