> ## Documentation Index
> Fetch the complete documentation index at: https://cli-docs.relai.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Benchmarks & evaluators

> Use RELAI benchmarks, evaluators, tags, and memory to scale from individual agent tests to reusable evaluation suites.

Use agent tests for individual behaviors, benchmarks for many reusable cases, and evaluators for the criteria that should score them.

## Choose the smallest object that fits

| You need | Use | Why |
| - | - | - |
| One behavior or failed interaction | [Agent test](/agent-tests) | One scenario with its own environment and evaluators |
| An existing set of evaluation cases | Benchmark | Connect the reusable suite to RELAI; CSV is the currently supported registration format |
| A criterion owned by one test or benchmark | Local evaluator | Travels with that specific coverage |
| A criterion that should score every matching run | Global evaluator | Automatically applies to simulations and optimizer evaluations in its scope |
| A reusable group of agent tests | Tag | Select the suite by one name |

## Connect an existing benchmark

Register the benchmark so you can select it directly in simulation and Agent Optimizer. CSV is the currently supported input format, but it does not need a fixed RELAI schema: RELAI reads the existing columns, asks when intent is ambiguous, and stores the dataset with the generated benchmark package. Multi-turn cases are supported.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to connect my benchmark at <path.csv>. Explain how it interprets the columns, proposed scenarios, and evaluator criteria before registering it." actions={["copy"]}>
      Use RELAI to connect my benchmark at \<path.csv>. Explain how it interprets the columns, proposed scenarios, and evaluator criteria before registering it.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai benchmark register --csv <path.csv> --name <name>
    ```
  </Tab>
</Tabs>

Use free-form guidance when the CSV alone does not explain multi-turn behavior or column meaning. Use evaluator-reference files when the evaluation protocol already lives in code or Markdown.

### Use the registered benchmark

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to optimize this agent against benchmark <benchmark-id>. Show me the scope and rollout budget before starting." actions={["copy"]}>
      Use RELAI to optimize this agent against benchmark \<benchmark-id>. Show me the scope and rollout budget before starting.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai simulate --benchmarks <benchmark-id>
    relai optimize --benchmarks <benchmark-id>
    ```
  </Tab>
</Tabs>

## Create a global evaluator

A global evaluator runs automatically on every matching simulation and optimizer evaluation. Scope it to the intended target and, when useful, a component. There is no per-run subset selector, so remove it when it should stop applying.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to create a global evaluator for this agent: “…”. Show its scope, exact or semantic criteria, and score interpretation before saving it." actions={["copy"]}>
      Use RELAI to create a global evaluator for this agent: “…”. Show its scope, exact or semantic criteria, and score interpretation before saving it.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai evaluator create --prompt "<what to score>" --name <name>
    ```
  </Tab>
</Tabs>

<CardGroup cols={2}>
  <Card title="Code evaluator">
    Fits tool calls, arguments, required strings, state, counts, and other deterministic evidence.
  </Card>

  <Card title="LLM judge">
    Fits meaning, policy, groundedness, tone, and criteria with multiple valid responses. RELAI’s model proxy handles the judge.
  </Card>
</CardGroup>

## Organize suites and carry guidance forward

RELAI normally attaches tags automatically and maintains optimizer memory. Manage them directly only when you need a named suite or durable generation guidance.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to show this agent’s test tags and memory scopes. Explain what is automatic, what is user-authored, and what should remain in EvalsWiki instead." actions={["copy"]}>
      Use RELAI to show this agent’s test tags and memory scopes. Explain what is automatic, what is user-authored, and what should remain in EvalsWiki instead.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai tag list --verbose
    relai memory list
    relai memory read agent
    relai simulate --tags <tag>
    ```
  </Tab>
</Tabs>

<Note>
  **Memory is not EvalsWiki.**

  EvalsWiki records durable requirements, risks, and open questions that ground coverage. Memory is guidance RELAI considers during later generation or optimization and can be scoped to an agent, tag, component, or evaluator.
</Note>

## Storage and lifecycle

| Object | Location or behavior |
| - | - |
| Benchmark | `.relai/harbor/benchmarks/`; commit reviewed packages and their stored datasets |
| Global evaluator | `.relai/harbor/evaluators/`; commit reviewed packages |
| Update | Keeps stable identity and target while staging a validated replacement |
| Remove | Destructive for an evaluator; benchmark removal retires its artifact record and retains shared cached data |
| Backend publication | Benchmark and evaluator package contents remain repository files; release builds do not publish that content |

<Card title="Coverage ready?" href="/simulation">
  Run the smallest selected test, tag, or benchmark and inspect every evaluator outcome.
</Card>

## Command reference

See [`relai benchmark`](/cli/benchmark), [`relai evaluator`](/cli/evaluator),
and [tags & memory](/cli/tag-and-memory) for the detailed command options.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.