louped

Write an experiment

What goes in an experiment's folder, and how its results show in the app.

An experiment is a folder in experiments/, named for its question. The app finds everything in it by its file names; nothing needs registering.

louped new does-pushback-flip-answers --domain honesty
FileWhat it isWhere it shows
README.mdThe question, the plan and the resultExperiments, and its own page
run.pyThe script. Its Args dataclass becomes a form.Launch
task.pyInspect evals (@task)Launch, Runs
*.yamlA training config (# louped train sft) or a grid (# louped grid)Launch

README

It starts with a short header, then these sections: Question, Observation, Hypotheses, Baseline, Test, Stop if, Run, Result, Next.

---
domain: honesty
status: active
extras: tracking, interp
---
  • status is active, parked or answered.
  • extras is optional. A job sent to a cluster installs only these.
  • The first paragraph of Question and of Result shows on the Experiments page.
  • Edit on the experiment's page edits the README in place, with its rendering beside it. So does Edit on any text file a run logged: its report.md, examples to judge, or a JSON, JSONL, CSV, YAML or TOML file. JSON, JSONL and TOML must still parse to be saved. The run then lists the files under Edited after in its provenance, so its results never silently change.
  • The bin button on an experiment, run or job page deletes it after one confirm. Nothing is gone for good: an experiment's folder, an eval's log and a job move to .louped/trash/, and MLflow keeps a deleted run until mlflow gc. A run or job still running must be stopped first.

Domains

A domain groups questions on the Experiments page. louped has eight by default:

AxisDomains
behaviormechanisms, honesty, conditioning, agents
efficiencycontext, inference, specialisation
checksreproduction

To use your own, list them in louped.toml. This replaces the defaults:

[domains.calibration]
axis = "behavior"
title = "Calibration under feedback"

run.py

with start_run(EXPERIMENT, name=args.model, params=asdict(args), seed=args.seed):
    ...
    mlflow.log_metrics({"pressure/accuracy": 0.31})
    mlflow.log_artifacts(folder)   # raw/<condition>.jsonl, report.md

start_run records the code version, packages, seed and hardware use with every run. Pin the model's revision and the seed in Args so they are saved too.

What the app shows depends on what the script writes:

WriteShown as
raw/<condition>.jsonl, an id in every rowItems: each item under every condition
A claim or question field in every rowItems: an item opens on it as its title
raw/fields.json, {"field": "meaning"}Items: a ? on each field
One JSONL file with a condition columnArtifacts: line up by that column
JSONL rows with a time and a kind (traces)Artifacts: a timeline per request
report.mdOverview
Numbers with mlflow.log_metricsOverview, Runs, Compare
Metrics logged over stepsOverview; two runs overlaid on Compare
A view from louped.analysis.viewsFigures (table, line, scatter, heatmap)
views.vega(title, spec), data inlineFigures: any Vega-Lite chart
views.plotly(title, data, frames=...)Figures: 3D and animated charts

A finding made by hand in Probe or Benchmark becomes a run too: Save as run keeps what the open tool showed (its figures, the prompt and the settings sent) under one of the experiments, or under probe or benchmark when it answers none yet.

A run's files can grow after it ends. louped derive <run> <script> runs the script's derive(files) over a copy of them and logs what it returns into the run: rows keyed like the records become columns in Items, a figure lands on Figures. The script is logged beside them. Your agent does this when you ask for a column or chart the run lacks.

Code from a paper

If a paper's code needs its own package versions, run it in its own environment. Clone it at a fixed commit under .louped/vendor/, call it as a subprocess, and log its per-item records and report as above.

On this page