Tutorial¶
This tutorial walks you through the full mushin workflow end to end: define a sweep, collect a labeled dataset, compare methods with statistical significance, and interpret the result.
The comparison steps need the eval extra
The sweep → dataset steps run on the core install. The later compare +
significance steps use mushin's optional evaluation layer:
pip install "mushin-py[eval]". See Installation.
Step 1: Define a sweep¶
mushin workflows are subclasses of MultiRunMetricsWorkflow. You implement a
task(...) method that returns a dict of metrics, then call .run(...) with
swept parameters:
The multirun(...) wrapper tells Hydra to create one job per value. Here the
sweep creates 3 × 3 = 9 jobs (three learning rates × three seeds). Each job
runs in its own output directory; the returned dict is collected automatically.
Step 2: Collect the dataset¶
After .run(...) completes, call .to_xarray() to get a labeled xarray.Dataset:
ds = wf.to_xarray()
# <xarray.Dataset> Dimensions: (lr: 3, seed: 3)
# Data variables: accuracy (lr, seed)
ds["accuracy"].mean("seed") # average accuracy per learning rate
ds.sel(lr=0.1) # slice to a single lr
The dimensions come from the swept parameters; the data variables come from the
dict your task returned.
Step 3: Compare methods with statistics¶
Once you have trained models (one list per method, one model per seed), pass
them to compare:
compare evaluates every model on data, assembles an (method × seed)
xarray Dataset of metrics, and runs pairwise Holm-corrected significance tests.
Step 4: Read the statistics¶
result.summary()
# method | metric | mean | ci_low | ci_high | significant_vs_ref
# cnn | accuracy | 0.963 | 0.951 | 0.975 |
# mlp | accuracy | 0.941 | 0.928 | 0.954 | *
result.data # xarray.Dataset, dims (method, seed)
result.comparisons # tidy DataFrame with p-values and effect sizes
"*" in significant_vs_ref means the method differs significantly from the
reference (first method listed) after Holm correction at alpha=0.05.
Next steps¶
- Core concepts — the mental model behind mushin
- Comparing methods — deeper coverage of
compare - Studies — combine training and comparison in one call
- Understanding the statistics — which test to choose