#1027/#1217 make read_cmdstan_csv() and as_cmdstan_fit() read gzip and bzip2 compressed CSVs, but there's no good way to get a fit object's own files compressed. If you gzip the files yourself, the fitted model object still points at the plain .csv paths and fit$draws() will break. When #1217 is merged, the advice for users is basically "finish with the fit, compress, then rebuild it with as_cmdstan_fit()", which is awkward.
But we could add a compress = c("none", "gzip", "bzip2") argument in two places to make this easier for users:
- The fitting methods (
$sample(), $optimize(), etc.), alongside output_dir and output_basename. Once CmdStan finishes, cmdstanr compresses the output files in place with gzfile()/bzfile() and records the new paths, so the fit works as usual with no extra call. I guess latent dynamics files should be compressed too when save_latent_dynamics = TRUE. We can probably skip small files (profile, metric, config).
$save_output_files() should support compressing after the fact.
This would break anything that gives file paths to a CmdStan binary like $cmdstan_summary(), $cmdstan_diagnose(), generate_quantities(fitted_params = fit), and laplace(mode = fit). Once #1217 is merged this already happens for a fit built with as_cmdstan_fit() from compressed files, since the fit keeps the .csv.gz paths and CmdStan rejects them. We could either error in those cases or we could decompress to a temp file when needed. Or maybe there's another option. Something to think about.
A few other things that aren't ideal, but can just be documented clearly so users are aware:
- CmdStanMCMC reads the sampler diagnostics right after sampling to print warnings, so that first read requires a decompression. We could avoid by compressing after the fit is build (instead of right after
run_cmdstan()) but that's a more complicated implementation and I think we should keep it simpler.
- Every read of a compressed file is slower than the plain CSV. Just need to document this so users keep in mind the tradeoff.
I think reading in the compressed files needs gzip/bzip2 on the PATH (from RTools on Windows)
#1027/#1217 make
read_cmdstan_csv()andas_cmdstan_fit()read gzip and bzip2 compressed CSVs, but there's no good way to get a fit object's own files compressed. If you gzip the files yourself, the fitted model object still points at the plain .csv paths andfit$draws()will break. When #1217 is merged, the advice for users is basically "finish with the fit, compress, then rebuild it with as_cmdstan_fit()", which is awkward.But we could add a
compress = c("none", "gzip", "bzip2")argument in two places to make this easier for users:$sample(),$optimize(), etc.), alongsideoutput_dirandoutput_basename. Once CmdStan finishes, cmdstanr compresses the output files in place withgzfile()/bzfile()and records the new paths, so the fit works as usual with no extra call. I guess latent dynamics files should be compressed too whensave_latent_dynamics = TRUE. We can probably skip small files (profile, metric, config).$save_output_files()should support compressing after the fact.This would break anything that gives file paths to a CmdStan binary like
$cmdstan_summary(),$cmdstan_diagnose(),generate_quantities(fitted_params = fit), andlaplace(mode = fit). Once #1217 is merged this already happens for a fit built withas_cmdstan_fit()from compressed files, since the fit keeps the.csv.gzpaths and CmdStan rejects them. We could either error in those cases or we could decompress to a temp file when needed. Or maybe there's another option. Something to think about.A few other things that aren't ideal, but can just be documented clearly so users are aware:
run_cmdstan()) but that's a more complicated implementation and I think we should keep it simpler.I think reading in the compressed files needs gzip/bzip2 on the PATH (from RTools on Windows)