> ## Documentation Index
> Fetch the complete documentation index at: https://cli-docs.relai.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Simulations

> Run RELAI simulations, select agent tests or benchmarks, and inspect available conversations, tool evidence, evaluator scores, and failures.

Run selected agent tests or benchmarks and inspect evaluator outcomes together with any available conversation, tool, and failure evidence.

<Note>
  **Simulation is useful, but optional before optimization.**

  Use it when you need a baseline or want to understand a failure. Agent Optimizer runs its own separate before-and-after evaluation and does not reuse an earlier simulation result.
</Note>

## Run a focused simulation

Start with one test. Ask for every evaluator outcome and any readable trajectory the run produced—not only an aggregate score.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to simulate agent test <test-id> once. Show every evaluator score or error, plus any available conversation and tool evidence. Identify the earliest supported failure without inventing missing evidence." actions={["copy"]}>
      Use RELAI to simulate agent test \<test-id> once. Show every evaluator score or error, plus any available conversation and tool evidence. Identify the earliest supported failure without inventing missing evidence.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai simulate --tests <test-id>
    ```
  </Tab>
</Tabs>

## Know what enters the run

| Input | Boundary |
| - | - |
| Project source | A fresh sanitized snapshot is mounted read-only at `/workspace`. Host changes made after snapshot creation are not visible in that run. |
| Excluded paths | The snapshot omits `.env` files, `.git`, `.relai`, transient Wiki views, `.venv`, `node_modules`, and other host dependency roots. |
| Image build and verifier | Docker builds do not receive the project mount, and the verifier uses task-owned artifacts rather than the project snapshot. |
| Runtime values | Environment values you provide are injected separately into the Harbor agent phase; they do not arrive through the source mount. |

## Select exactly what to measure

| Selector | Use it for |
| - | - |
| Agent test | One or a few specific behavior scenarios |
| Tag | A reusable group of related agent tests |
| Benchmark | Cases in a registered benchmark; CSV is the currently supported registration format |
| Explicit Harbor task | One unregistered task package during advanced development |

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to show the available tests, tags, and benchmarks for this agent. Recommend the smallest selector that answers: “…”. Do not run it yet." actions={["copy"]}>
      Use RELAI to show the available tests, tags, and benchmarks for this agent. Recommend the smallest selector that answers: “…”. Do not run it yet.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai simulate --tests <test-a>,<test-b>
    relai simulate --tags <tag>
    relai simulate --benchmarks <benchmark-id>
    ```
  </Tab>
</Tabs>

## Read the outcome correctly

<CardGroup cols={3}>
  <Card title="Valid low score">
    The agent completed the case but missed an expectation. Read the evaluator feedback and earliest failing action.
  </Card>

  <Card title="Evaluator error">
    The criterion did not score reliably. A missing score is not a zero and should not be presented as an agent failure.
  </Card>

  <Card title="Runtime failure">
    Docker, dependencies, credentials, mounts, or provider setup prevented a valid behavioral result.
  </Card>
</CardGroup>

Batch simulations continue after individual execution failures and retain every valid result. Always report successes, failures, and evaluator errors separately.

## Runtime values and model access

Use the variable names in `.relai/simulator.env.example`. Put persistent local values in ignored `.relai/simulator.env`; use your shell or CI secret store for temporary values.

Values are resolved in this order: process environment, `.relai/simulator.env`, each `--env-file`, then each `--env`; later values win. RELAI forwards only manifest-declared names from the process environment, but forwards the values supplied by environment files and `--env` to the agent phase. Keep copied commands and chat free of secret values.

<Warning>
  **Two kinds of model access are different.**

  Your agent may need its own provider key. RELAI’s semantic evaluators, simulated personas, and generated component inputs use RELAI’s model proxy through your OAuth sign-in and do not need that provider key.
</Warning>

## Know what simulation may change

Simulation does not intentionally edit your application source. If execution fails, RELAI may use a bounded repair flow for the selected agent test, benchmark, or simulator harness and retry. Those are CLI-owned evaluation artifacts, not agent-source optimization. It cannot repair an unavailable Docker daemon, a full host disk, or missing `uv`, and it stops when two consecutive repairs reproduce the same failure.

A simulation has a three-hour wall-clock limit by default. Use `--timeout <duration|off>` when a selected run needs a different invocation budget.

## Where results live

RELAI keeps completed results and available transcripts under `.relai/.internal/runs/`. Each completed run gets a local result under `results/<kind>/<id>/<run-id>.json`, even without `--result-json`. Projects without `.relai/.internal/layout.json` use the legacy `.relai/runs/` directory. These are local evidence, not EvalsWiki content and not artifacts to commit. Request an explicit machine-readable output file when another tool needs the suite result.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to run <test-id> and save a machine-readable result. Also show me the readable outcome; do not dump raw JSON into chat." actions={["copy"]}>
      Use RELAI to run \<test-id> and save a machine-readable result. Also show me the readable outcome; do not dump raw JSON into chat.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai simulate --tests <test-id> \
      --result-json .relai/.internal/runs/<result-name>.json
    ```
  </Tab>
</Tabs>

<Card title="Found a real behavior gap?" href="/optimization">
  Use Agent Optimizer to improve the selected behavior and run a fresh before-and-after evaluation.
</Card>

## Command reference

See [`relai simulate`](/cli/simulate) for all flags, defaults, and recovery details.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.