Skip to main content
Use agent tests for individual behaviors, benchmarks for many reusable cases, and evaluators for the criteria that should score them.

Choose the smallest object that fits

Connect an existing benchmark

Register the benchmark so you can select it directly in simulation and Agent Optimizer. CSV is the currently supported input format, but it does not need a fixed RELAI schema: RELAI reads the existing columns, asks when intent is ambiguous, and stores the dataset with the generated benchmark package. Multi-turn cases are supported.

Use RELAI to connect my benchmark at <path.csv>. Explain how it interprets the columns, proposed scenarios, and evaluator criteria before registering it.

Use free-form guidance when the CSV alone does not explain multi-turn behavior or column meaning. Use evaluator-reference files when the evaluation protocol already lives in code or Markdown.

Use the registered benchmark

Use RELAI to optimize this agent against benchmark . Show me the scope and rollout budget before starting.

Create a global evaluator

A global evaluator runs automatically on every matching simulation and optimizer evaluation. Scope it to the intended target and, when useful, a component. There is no per-run subset selector, so remove it when it should stop applying.

Use RELAI to create a global evaluator for this agent: “…”. Show its scope, exact or semantic criteria, and score interpretation before saving it.

Code evaluator

Fits tool calls, arguments, required strings, state, counts, and other deterministic evidence.

LLM judge

Fits meaning, policy, groundedness, tone, and criteria with multiple valid responses. RELAI’s model proxy handles the judge.

Organize suites and carry guidance forward

RELAI normally attaches tags automatically and maintains optimizer memory. Manage them directly only when you need a named suite or durable generation guidance.

Use RELAI to show this agent’s test tags and memory scopes. Explain what is automatic, what is user-authored, and what should remain in EvalsWiki instead.

Memory is not EvalsWiki.EvalsWiki records durable requirements, risks, and open questions that ground coverage. Memory is guidance RELAI considers during later generation or optimization and can be scoped to an agent, tag, component, or evaluator.

Storage and lifecycle

Coverage ready?

Run the smallest selected test, tag, or benchmark and inspect every evaluator outcome.

Command reference

See relai benchmark, relai evaluator, and tags & memory for the detailed command options.