Use relai benchmark to register, update, list, and remove benchmark suites.
register
Register a benchmark from CSV data.
Benchmark registration can take a while when RELAI needs to generate benchmark support from larger CSV data or evaluator context.
Registration is agentic: an agent inspects your CSV and project, records a
plan, writes the benchmark, and must pass trusted validation (id, name,
dataset reference, and sample expansion) before the benchmark is saved. When
the CSV or project leaves a material choice open, it asks a focused question
mid-session instead of collecting all clarifications up front. Interrupted or
paused sessions are saved and continued with --resume.
Creates benchmark files under .relai/harbor/benchmarks/ (including the dataset).
Commit the package to share it; release builds do not register benchmark content
with the backend.
update
Update an existing benchmark copy-on-write with fresh EvalsWiki grounding.
The benchmark ID, name, and agent target stay stable. --csv optionally replaces the stored dataset; otherwise RELAI reuses the cached stored CSV. Interrupted updates retain the original source and resume with --resume; --restart discards only the staged update.
register and update accept --permissions {auto|accept_edits|ask} and
--timeout {DURATION|off}. See Permissions and
Timeouts.
remove
Removal deletes the benchmark source and retains its artifact page with lifecycle retired. Cached datasets are retained because another benchmark may reference the same dataset.
A benchmark is the batch way to define many scenarios at once: registering a
CSV produces the same kind of objects as creating agent tests one at
a time with relai test, one scenario per row.
Use whichever fits — a handful of known failures work fine as individual
agent tests; a reusable suite of cases belongs in a benchmark.
A benchmark has one agent target for all of its samples. When multiple targets
are registered, interactive runs present a single-target picker; pass
--agent-target {target} in non-interactive runs.
CSV expectations
There is no fixed CSV schema. RELAI reads your columns as they are and infers
what kind of benchmark to generate from them — the required columns recorded
inside a generated benchmark file describe that benchmark’s dataset; they
are not a format your CSV must match up front.
Multi-turn scenarios are supported without one-column-per-turn. A CSV can
carry just the first user message per row while later turns are simulated
adaptively — indicate in the CSV itself or in --prompt that the samples are
conversations, and RELAI asks when the intent is ambiguous.
--prompt vs --evaluator-ref
Both supply generation context; they differ in focus and form.
--prompt is a free-form string for anything your CSV alone does not say:
- these samples are multi-turn conversations.
- what an ambiguously named column means, or that a column should be ignored.
- which aspects to measure, such as groundedness, token cost, or latency.
--evaluator-ref is a path (repeatable) to files that describe evaluation
specifically:
- existing evaluator code you ran this benchmark with before, so RELAI can
reproduce your evaluation protocol as closely as possible.
- a Markdown description of an evaluation harness, including how to invoke it,
when the code alone does not explain the workflow.
Evaluation guidance can also go in --prompt; use --evaluator-ref when it
already lives in files.
Benchmarks remain in the repository after registration and can be shared through Git.
Common failures: invalid CSV, missing required columns, generation ambiguity, or validation failure. A failed or interrupted agentic session is kept at .relai/benchmark-register-session.json; continue it with --resume or discard it with --restart.
list
List registered benchmarks.
Read-only. Common failures are malformed benchmark files or missing simulator dependencies.