Choose the smallest object that fits
Connect an existing benchmark
Register the benchmark so you can select it directly in simulation and Agent Optimizer. CSV is the currently supported input format, but it does not need a fixed RELAI schema: RELAI reads the existing columns, asks when intent is ambiguous, and stores the dataset with the generated benchmark package. Multi-turn cases are supported.- Ask your coding agent
- Run in terminal
Use RELAI to connect my benchmark at <path.csv>. Explain how it interprets the columns, proposed scenarios, and evaluator criteria before registering it.
Use the registered benchmark
- Ask your coding agent
- Run in terminal
Use RELAI to optimize this agent against benchmark . Show me the scope and rollout budget before starting.
Create a global evaluator
A global evaluator runs automatically on every matching simulation and optimizer evaluation. Scope it to the intended target and, when useful, a component. There is no per-run subset selector, so remove it when it should stop applying.- Ask your coding agent
- Run in terminal
Use RELAI to create a global evaluator for this agent: “…”. Show its scope, exact or semantic criteria, and score interpretation before saving it.
Code evaluator
LLM judge
Organize suites and carry guidance forward
RELAI normally attaches tags automatically and maintains optimizer memory. Manage them directly only when you need a named suite or durable generation guidance.- Ask your coding agent
- Run in terminal
Use RELAI to show this agent’s test tags and memory scopes. Explain what is automatic, what is user-authored, and what should remain in EvalsWiki instead.
Storage and lifecycle
Coverage ready?
Command reference
Seerelai benchmark, relai evaluator,
and tags & memory for the detailed command options.