Skip to content

Let fits compress their own output CSVs #1276

Description

@jgabry

#1027/#1217 make read_cmdstan_csv() and as_cmdstan_fit() read gzip and bzip2 compressed CSVs, but there's no good way to get a fit object's own files compressed. If you gzip the files yourself, the fitted model object still points at the plain .csv paths and fit$draws() will break. When #1217 is merged, the advice for users is basically "finish with the fit, compress, then rebuild it with as_cmdstan_fit()", which is awkward.

But we could add a compress = c("none", "gzip", "bzip2") argument in two places to make this easier for users:

  • The fitting methods ($sample(), $optimize(), etc.), alongside output_dir and output_basename. Once CmdStan finishes, cmdstanr compresses the output files in place with gzfile()/bzfile() and records the new paths, so the fit works as usual with no extra call. I guess latent dynamics files should be compressed too when save_latent_dynamics = TRUE. We can probably skip small files (profile, metric, config).
  • $save_output_files() should support compressing after the fact.

This would break anything that gives file paths to a CmdStan binary like $cmdstan_summary(), $cmdstan_diagnose(), generate_quantities(fitted_params = fit), and laplace(mode = fit). Once #1217 is merged this already happens for a fit built with as_cmdstan_fit() from compressed files, since the fit keeps the .csv.gz paths and CmdStan rejects them. We could either error in those cases or we could decompress to a temp file when needed. Or maybe there's another option. Something to think about.

A few other things that aren't ideal, but can just be documented clearly so users are aware:

  • CmdStanMCMC reads the sampler diagnostics right after sampling to print warnings, so that first read requires a decompression. We could avoid by compressing after the fit is build (instead of right after run_cmdstan()) but that's a more complicated implementation and I think we should keep it simpler.
  • Every read of a compressed file is slower than the plain CSV. Just need to document this so users keep in mind the tradeoff.

I think reading in the compressed files needs gzip/bzip2 on the PATH (from RTools on Windows)

Activity

  1. jgabry commented on Sep 17, 2026

    @jgabry
    MemberAuthor

    In addition to normal documentation for this we should add a vignette section

  2. self-assigned this
    on Sep 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

featureNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions