A workbench for LLM behavior and efficiency.
Your coding agent writes and runs the experiments, and shapes the app around them. You read the results, item by item. Local, open models, free.


How it works
Make a project
One command sets up a folder for your questions, an example to try, and the files your coding agent needs.
pip install 'louped[server,tracking,interp,agent]' louped init my-research && cd my-research louped serve


Ask your agent
Tell Claude Code, Codex or Cursor what to test. It writes the experiment: the question, the plan and the code.


Run it
Small runs go on your machine. Large ones go to a Slurm cluster as one job, and the results come back to your app.


Read the results
See every item under every condition: what changed, and out of how many. Open one to compare the conditions side by side.


Make it yours
Shift+click any rows, cards or fields and ask your agent about them: it gets their data, and points back at the evidence on screen. Ask it to rename, hide, reorder or add any part, in louped's look; the page updates as it works.


What a run cost
Every run records the GPU energy, power and memory it used, and the code, packages and seed behind it.


Agent traces
Open an agent's trace files as a timeline per request. louped view opens any results you already have.


Use your own agent
louped works with the coding agent you already use. In Claude Code, you can also add it as a plugin.