louped

Recipes

How to run common kinds of experiments.

A model louped can load

Hugging Face models load directly. Interventions, analyses (logit lens, patching, probes, SAE features) and fine-tuning all work on them. Compare versions with a grid.

A model served elsewhere

For GGUF models on llama-server, vLLM, Ollama, or your own engine, use its OpenAI-compatible endpoint:

  • Evals: add it to a grid as endpoint(name, url).
  • Speed and cost: louped endpoint-bench http://127.0.0.1:8080/v1 measures time to first token, latency and throughput at different numbers of requests at once. On the same machine it also reports energy per token.

Code in another language

Call the program from run.py as a subprocess (for example cargo bench), then log its numbers with mlflow.log_metrics and its per-item files as artifacts.

Agents

  • An agent behind an endpoint can be scored in a grid: endpoint(name, url, agent=True).
  • An agent's own trace files (rows with a time like at and a kind like kind) open as timelines. Try louped view path/to/traces.

Retrieval

Log one record per question and setup (no retrieval, retrieved passages) with recall, the answer and its score, then line them up on the Items tab. To test retrieval inside the model, use the inject intervention.

Adapters and small models

Load several LoRA adapters as a bank and switch them per request or partway through a reply. To see what a weight format costs, run louped bench <model> --quants none int8 int4 next to an eval.

On this page