louped

Change a model

Steer, ablate, inject and swap adapters, and compare the result against a baseline.

louped runs a model in Inspect evals as louped/<model>. Model arguments describe the change.

inspect eval task.py --model louped/Qwen/Qwen2.5-0.5B-Instruct \
  -M interventions='{"kind": "ablate", "vector": "refusal.qwen2.5-0.5b-instruct"}'

Interventions

KindWhat it does
steerAdds a saved direction at one layer
ablateRemoves a direction at every layer
injectAdds the state of some passages at one layer
headsTurns off chosen attention heads

inject takes at: all (every token, the default), prompt (the prompt only) or chunks (every chunk-th generated token). These match an engine that retrieves while it generates.

Other model arguments

ArgumentWhat it does
bankLoads LoRA adapters by name, all off
adaptersTurns on some of the bank
mergesAdds merged adapters (linear, ties, ...)
phasesChanges the live adapters partway through a reply
diffusionServes a masked diffusion model
attnPicks the attention kernel (eager, sdpa, flex_attention)
quantLoads the weights in 8 or 4 bits (CUDA only)
revisionPins a Hub model to a commit
remote_codeRuns the model's own code from its Hub repository. Pin revision.
batch_sizeRequests generated together. 1 gives bit-identical reruns.

Saved directions are in .louped/vectors, adapters in .louped/adapters, and models in .louped/models.

Compare conditions with a grid

A grid runs each condition on each task over several seeds, and compares it with a baseline.

grid.yaml
model: Qwen/Qwen2.5-0.5B-Instruct
tasks:
  pushback: experiments/my-question/task.py@pushback
conditions:
  base: {}
  ablate: { interventions: { kind: ablate, vector: caving.qwen2.5-0.5b-instruct } }
  int4: { quant: int4 }
metric: held/accuracy
held: correct_first/accuracy
seeds: [0, 1, 2]
louped grid grid.yaml

For each condition you get the change from the baseline with a 95% paired interval, and two verdicts: whether the metric moved, and whether the held score held. Add extra: [latency/mean, peak_memory/mean] to see what each condition costs next to what it changes.

A condition can also be a model behind an OpenAI-compatible endpoint, such as llama-server, vLLM or Ollama: use endpoint(name, url) from louped.grid.

On this page