This vignette covers running brms-backed simulation studies: how the
model bank eliminates per-task compilation, how to declare
model-comparison grids with model_grid() /
brms_model(), how to pass sampler arguments, and how to
keep memory flat on long runs. All chunks require brms and a working
CmdStan install, so they are shown but not evaluated here.
The model bank: compile once, fit thousands of times
A naive brms simulation study recompiles the Stan model for every
task — minutes of C++ compilation per seconds of sampling.
BrmsFitter() avoids this with a model bank: at the
start of run_simulation(), each distinct model spec in the
fit_grid is compiled once (via
brms::brm(chains = 0)), and every task reuses the compiled
binary through stats::update(recompile = FALSE).
Two properties control this:
-
precompile(defaultTRUE): build the bank at run start. Set toFALSEto fall back to a freshbrms::brm()per task (only sensible for tiny studies or debugging). - The bank is shipped to mirai daemons once per run, so parallel workers reuse the same binaries.
brms does not warn when
recompile = FALSE is used against structurally incompatible
data — it would silently reuse the binary against the wrong model frame.
bayesim therefore compares the Stan data structure (via
brms::make_standata()) between the compiled template and
each task’s data, and aborts the run loudly on a mismatch. If your
data_grid rows produce data with different shapes
(e.g. varying factor levels), either make them structurally identical or
set precompile = FALSE.
Some brms default priors are data-dependent (for example, the
intercept prior can be centered using the response in the template
dataset). Those constants are embedded in the compiled model and cannot
be refreshed by update(recompile = FALSE). Therefore,
always supply explicit priors in every brms_model() when
precompile = TRUE; bayesim raises a fatal configuration
error when a model-bank row omits them (opt in with
BrmsFitter(allow_default_priors = TRUE) if you deliberately
want template-derived priors). Use precompile = FALSE if
task-specific default priors are part of the study design.
Declaring models: brms_model() and
model_grid()
Hand-building list-columns
(fit_grid$formula <- list(...)) works but is
error-prone. model_grid() assembles a tidy fit grid from
named brms_model() specs, validating each at construction
so mistakes surface before any compilation:
explicit_priors <- c(
brms::prior(normal(0, 2), class = "b"),
brms::prior(normal(0, 5), class = "Intercept"),
brms::prior(exponential(1), class = "sigma")
)
fit_grid <- model_grid(
gaussian = brms_model(y ~ x, brms::brmsfamily("gaussian"), explicit_priors),
student = brms_model(y ~ x, brms::brmsfamily("student"), explicit_priors),
lognormal = brms_model(y ~ x, brms::brmsfamily("lognormal"), explicit_priors)
)
fit_grid
#> # A tibble: 3 x 5
#> model formula family prior stanvars
#> <chr> <list> <list> <list> <list>
#> 1 gaussian <formula> <family> <prior> <NULL>
#> 2 student <formula> <family> <prior> <NULL>
#> 3 lognormal <formula> <family> <prior> <NULL>The model name column lands in the result summary as
fit_model, so per-model aggregation is
summarize_simulation(result) or
performance_measures(result) with no extra arguments.
Case study: comparing likelihoods for skewed data
Which likelihood best describes right-skewed data — Gaussian, Student-t, or lognormal? Generate skewed data, fit all three models, and compare expected log predictive density (ELPD):
skew_generator <- function(data_spec, task_ctx) {
n <- data_spec$n
x <- stats::rnorm(n)
mu <- data_spec$beta * x
y <- exp(mu + stats::rnorm(n, sd = 0.5))
list(
train = data.frame(y = y, x = x),
test = NULL,
response = "y",
true_params = c(beta = data_spec$beta),
vars_of_interest = "beta"
)
}
config <- simulation_config(
data_grid = data.frame(n = 200, beta = 1),
fit_grid = model_grid(
gaussian = brms_model(
y ~ x, brms::brmsfamily("gaussian"), explicit_priors
),
student = brms_model(
y ~ x, brms::brmsfamily("student"), explicit_priors
),
lognormal = brms_model(
y ~ x, brms::brmsfamily("lognormal"), explicit_priors
)
),
data_generator = skew_generator,
fitter = BrmsFitter(chains = 2L, iter = 1000L, warmup = 500L),
metrics = list(
elpd_loo_metric(),
posterior_summary_metric(),
sampler_diagnostics_metric()
),
n_replicates = 50L,
seed = 42L
)
result <- run_simulation(config, workers = 4)
summarize_simulation(result, metrics = "elpd_loo__elpd")Three models compile exactly once each; all 150 fits reuse the binaries.
Sampler arguments: stan_args
BrmsFitter(stan_args = ...) passes sampler controls
through to brms/Stan. adapt_delta and
max_treedepth are mapped into the control
list; init and threads pass through
directly:
fitter <- BrmsFitter(
chains = 4L,
stan_args = list(adapt_delta = 0.95, max_treedepth = 12, init = 0.1)
)Warning-conditional retention
Diagnosing occasional divergent transitions across thousands of tasks needs the fit object — but only for the tasks that misbehaved. Retention specs can be conditional on warnings, keeping heavy artifacts only where something went wrong:
config <- simulation_config(
...,
retain = list(
success = c("metrics", "diagnostics"),
warning = c("metrics", "diagnostics", "fit", "draws")
)
)Clean tasks stay light; tasks that emitted sampler warnings keep their full fit for post-hoc inspection.
SBC with brms generators
For calibration checking of brms models,
prior_predictive_generator() draws the truth from the model
prior (a sample_prior = "only" fit) and
ifs_generator() draws it from a preconditioning posterior.
Both forward-simulate the response — including dependency-ordered
simulation for multivariate models. See
vignette("sbc-and-calibration") for the full workflow.
Next steps
-
vignette("parallel-and-hpc")— daemons, checkpoint/resume, memory. -
vignette("custom-fitters")— raw Stan viaCmdStanFitter()when brms cannot express your model. -
vignette("reproducibility")— what the config fingerprint covers.