Skip to main content
Run selected agent tests or benchmarks and inspect evaluator outcomes together with any available conversation, tool, and failure evidence.
Simulation is useful, but optional before optimization.Use it when you need a baseline or want to understand a failure. Agent Optimizer runs its own separate before-and-after evaluation and does not reuse an earlier simulation result.

Run a focused simulation

Start with one test. Ask for every evaluator outcome and any readable trajectory the run produced—not only an aggregate score.

Use RELAI to simulate agent test once. Show every evaluator score or error, plus any available conversation and tool evidence. Identify the earliest supported failure without inventing missing evidence.

Know what enters the run

Select exactly what to measure

Use RELAI to show the available tests, tags, and benchmarks for this agent. Recommend the smallest selector that answers: “…”. Do not run it yet.

Read the outcome correctly

Valid low score

The agent completed the case but missed an expectation. Read the evaluator feedback and earliest failing action.

Evaluator error

The criterion did not score reliably. A missing score is not a zero and should not be presented as an agent failure.

Runtime failure

Docker, dependencies, credentials, mounts, or provider setup prevented a valid behavioral result.
Batch simulations continue after individual execution failures and retain every valid result. Always report successes, failures, and evaluator errors separately.

Runtime values and model access

Use the variable names in .relai/simulator.env.example. Put persistent local values in ignored .relai/simulator.env; use your shell or CI secret store for temporary values. Values are resolved in this order: process environment, .relai/simulator.env, each --env-file, then each --env; later values win. RELAI forwards only manifest-declared names from the process environment, but forwards the values supplied by environment files and --env to the agent phase. Keep copied commands and chat free of secret values.
Two kinds of model access are different.Your agent may need its own provider key. RELAI’s semantic evaluators, simulated personas, and generated component inputs use RELAI’s model proxy through your OAuth sign-in and do not need that provider key.

Know what simulation may change

Simulation does not intentionally edit your application source. If execution fails, RELAI may use a bounded repair flow for the selected agent test, benchmark, or simulator harness and retry. Those are CLI-owned evaluation artifacts, not agent-source optimization. It cannot repair an unavailable Docker daemon, a full host disk, or missing uv, and it stops when two consecutive repairs reproduce the same failure. A simulation has a three-hour wall-clock limit by default. Use --timeout <duration|off> when a selected run needs a different invocation budget.

Where results live

RELAI keeps completed results and available transcripts under .relai/.internal/runs/. Each completed run gets a local result under results/<kind>/<id>/<run-id>.json, even without --result-json. Projects without .relai/.internal/layout.json use the legacy .relai/runs/ directory. These are local evidence, not EvalsWiki content and not artifacts to commit. Request an explicit machine-readable output file when another tool needs the suite result.

Use RELAI to run and save a machine-readable result. Also show me the readable outcome; do not dump raw JSON into chat.

Found a real behavior gap?

Use Agent Optimizer to improve the selected behavior and run a fresh before-and-after evaluation.

Command reference

See relai simulate for all flags, defaults, and recovery details.