Skip to main content
Use relai benchmark to register, update, list, and remove benchmark suites.

register

Register a benchmark from CSV data.
Benchmark registration can take a while when RELAI needs to generate benchmark support from larger CSV data or evaluator context.
Registration is agentic: an agent inspects your CSV and project, records a plan, writes the benchmark, and must pass trusted validation (id, name, dataset reference, and sample expansion) before the benchmark is saved. When the CSV or project leaves a material choice open, it asks a focused question mid-session instead of collecting all clarifications up front. Interrupted or paused sessions are saved and continued with --resume.
Creates benchmark files under .relai/harbor/benchmarks/ (including the dataset). Commit the package to share it; release builds do not register benchmark content with the backend.

update

Update an existing benchmark copy-on-write with fresh EvalsWiki grounding.
The benchmark ID, name, and agent target stay stable. --csv optionally replaces the stored dataset; otherwise RELAI reuses the cached stored CSV. Interrupted updates retain the original source and resume with --resume; --restart discards only the staged update. register and update accept --permissions {auto|accept_edits|ask} and --timeout {DURATION|off}. See Permissions and Timeouts.

remove

Removal deletes the benchmark source and retains its artifact page with lifecycle retired. Cached datasets are retained because another benchmark may reference the same dataset. A benchmark is the batch way to define many scenarios at once: registering a CSV produces the same kind of objects as creating agent tests one at a time with relai test, one scenario per row. Use whichever fits — a handful of known failures work fine as individual agent tests; a reusable suite of cases belongs in a benchmark. A benchmark has one agent target for all of its samples. When multiple targets are registered, interactive runs present a single-target picker; pass --agent-target {target} in non-interactive runs.

CSV expectations

There is no fixed CSV schema. RELAI reads your columns as they are and infers what kind of benchmark to generate from them — the required columns recorded inside a generated benchmark file describe that benchmark’s dataset; they are not a format your CSV must match up front. Multi-turn scenarios are supported without one-column-per-turn. A CSV can carry just the first user message per row while later turns are simulated adaptively — indicate in the CSV itself or in --prompt that the samples are conversations, and RELAI asks when the intent is ambiguous.

--prompt vs --evaluator-ref

Both supply generation context; they differ in focus and form. --prompt is a free-form string for anything your CSV alone does not say:
  • these samples are multi-turn conversations.
  • what an ambiguously named column means, or that a column should be ignored.
  • which aspects to measure, such as groundedness, token cost, or latency.
--evaluator-ref is a path (repeatable) to files that describe evaluation specifically:
  • existing evaluator code you ran this benchmark with before, so RELAI can reproduce your evaluation protocol as closely as possible.
  • a Markdown description of an evaluation harness, including how to invoke it, when the code alone does not explain the workflow.
Evaluation guidance can also go in --prompt; use --evaluator-ref when it already lives in files. Benchmarks remain in the repository after registration and can be shared through Git. Common failures: invalid CSV, missing required columns, generation ambiguity, or validation failure. A failed or interrupted agentic session is kept at .relai/benchmark-register-session.json; continue it with --resume or discard it with --restart.

list

List registered benchmarks.
Read-only. Common failures are malformed benchmark files or missing simulator dependencies.