louped

A workbench for LLM behavior and efficiency.

Your coding agent writes and runs the experiments, and shapes the app around them. You read the results, item by item. Local, open models, free.

Get started
Sixteen questions under three conditions: what pushback and evidence did to each answer

How it works

01

Make a project

One command sets up a folder for your questions, an example to try, and the files your coding agent needs.

pip install 'louped[server,tracking,interp,agent]'
louped init my-research && cd my-research
louped serve
A new project in the app, with its example question
02

Ask your agent

Tell Claude Code, Codex or Cursor what to test. It writes the experiment: the question, the plan and the code.

An experiment's page: its question, hypotheses and test
03

Run it

Small runs go on your machine. Large ones go to a Slurm cluster as one job, and the results come back to your app.

The launch form, with the choice to run here or on a cluster
04

Read the results

See every item under every condition: what changed, and out of how many. Open one to compare the conditions side by side.

One question under baseline, evidence and pushback, side by side
05

Make it yours

Shift+click any rows, cards or fields and ask your agent about them: it gets their data, and points back at the evidence on screen. Ask it to rename, hide, reorder or add any part, in louped's look; the page updates as it works.

A run's page laid out by an agent, with the card being pointed at outlined

What a run cost

Every run records the GPU energy, power and memory it used, and the code, packages and seed behind it.

A run's results with the energy and hardware it used

Agent traces

Open an agent's trace files as a timeline per request. louped view opens any results you already have.

Agent requests as timelines of retrieval, model calls and tools

Use your own agent

louped works with the coding agent you already use. In Claude Code, you can also add it as a plugin.