Write an experiment
What goes in an experiment's folder, and how its results show in the app.
An experiment is a folder in experiments/, named for its question. The app finds everything in it
by its file names; nothing needs registering.
louped new does-pushback-flip-answers --domain honesty| File | What it is | Where it shows |
|---|---|---|
README.md | The question, the plan and the result | Experiments, and its own page |
run.py | The script. Its Args dataclass becomes a form. | Launch |
task.py | Inspect evals (@task) | Launch, Runs |
*.yaml | A training config (# louped train sft) or a grid (# louped grid) | Launch |
README
It starts with a short header, then these sections: Question, Observation, Hypotheses, Baseline, Test, Stop if, Run, Result, Next.
---
domain: honesty
status: active
extras: tracking, interp
---statusisactive,parkedoranswered.extrasis optional. A job sent to a cluster installs only these.- The first paragraph of Question and of Result shows on the Experiments page.
- Edit on the experiment's page edits the README in place, with its rendering beside it. So
does Edit on any text file a run logged: its
report.md, examples to judge, or a JSON, JSONL, CSV, YAML or TOML file. JSON, JSONL and TOML must still parse to be saved. The run then lists the files under Edited after in its provenance, so its results never silently change. - The bin button on an experiment, run or job page deletes it after one confirm. Nothing is gone
for good: an experiment's folder, an eval's log and a job move to
.louped/trash/, and MLflow keeps a deleted run untilmlflow gc. A run or job still running must be stopped first.
Domains
A domain groups questions on the Experiments page. louped has eight by default:
| Axis | Domains |
|---|---|
| behavior | mechanisms, honesty, conditioning, agents |
| efficiency | context, inference, specialisation |
| checks | reproduction |
To use your own, list them in louped.toml. This replaces the defaults:
[domains.calibration]
axis = "behavior"
title = "Calibration under feedback"run.py
with start_run(EXPERIMENT, name=args.model, params=asdict(args), seed=args.seed):
...
mlflow.log_metrics({"pressure/accuracy": 0.31})
mlflow.log_artifacts(folder) # raw/<condition>.jsonl, report.mdstart_run records the code version, packages, seed and hardware use with every run. Pin the
model's revision and the seed in Args so they are saved too.
What the app shows depends on what the script writes:
| Write | Shown as |
|---|---|
raw/<condition>.jsonl, an id in every row | Items: each item under every condition |
A claim or question field in every row | Items: an item opens on it as its title |
raw/fields.json, {"field": "meaning"} | Items: a ? on each field |
One JSONL file with a condition column | Artifacts: line up by that column |
| JSONL rows with a time and a kind (traces) | Artifacts: a timeline per request |
report.md | Overview |
Numbers with mlflow.log_metrics | Overview, Runs, Compare |
| Metrics logged over steps | Overview; two runs overlaid on Compare |
A view from louped.analysis.views | Figures (table, line, scatter, heatmap) |
views.vega(title, spec), data inline | Figures: any Vega-Lite chart |
views.plotly(title, data, frames=...) | Figures: 3D and animated charts |
A finding made by hand in Probe or Benchmark becomes a run too: Save as run keeps what the
open tool showed (its figures, the prompt and the settings sent) under one of the experiments,
or under probe or benchmark when it answers none yet.
A run's files can grow after it ends. louped derive <run> <script> runs the script's
derive(files) over a copy of them and logs what it returns into the run: rows keyed like the
records become columns in Items, a figure lands on Figures. The script is logged beside
them. Your agent does this when you ask for a column or chart the run lacks.
Code from a paper
If a paper's code needs its own package versions, run it in its own environment. Clone it at a
fixed commit under .louped/vendor/, call it as a subprocess, and log its per-item records and
report as above.