Recipes
How to run common kinds of experiments.
A model louped can load
Hugging Face models load directly. Interventions, analyses (logit lens, patching, probes, SAE features) and fine-tuning all work on them. Compare versions with a grid.
A model served elsewhere
For GGUF models on llama-server, vLLM, Ollama, or your own engine, use its OpenAI-compatible endpoint:
- Evals: add it to a grid as
endpoint(name, url). - Speed and cost:
louped endpoint-bench http://127.0.0.1:8080/v1measures time to first token, latency and throughput at different numbers of requests at once. On the same machine it also reports energy per token.
Code in another language
Call the program from run.py as a subprocess (for example cargo bench), then log its numbers
with mlflow.log_metrics and its per-item files as artifacts.
Agents
- An agent behind an endpoint can be scored in a grid:
endpoint(name, url, agent=True). - An agent's own trace files (rows with a time like
atand a kind likekind) open as timelines. Trylouped view path/to/traces.
Retrieval
Log one record per question and setup (no retrieval, retrieved passages) with recall, the answer
and its score, then line them up on the Items tab. To test retrieval inside the model, use the
inject intervention.
Adapters and small models
Load several LoRA adapters as a bank and switch them per request or partway through a reply. To
see what a weight format costs, run louped bench <model> --quants none int8 int4 next to an eval.