Skip to main content
For a guided workflow, see Simulations. relai simulate runs selected agent tests against the generated simulator.

When to simulate

Simulation measures; it never changes your agent. Run it to see where the agent stands against selected agent tests or benchmarks, or to re-check behavior after changes merge. It is not a prerequisite for optimization. relai optimize runs its own before/after evaluation passes and does not reuse earlier simulation results, so when you already know you want the optimizer, you can skip straight to it.

Select agent tests

At least one selector should identify what to run. For a project initialized with the Harbor runtime, pass one explicit task instead of these selectors:
For each agent phase, Harbor mounts a fresh sanitized snapshot of the current project source read-only at /workspace. The snapshot omits sensitive paths (including .env and .env.example), .git, .relai, transient Wiki views, and host dependency roots such as .venv and node_modules. The verifier uses task-image artifacts and transcripts or other task-owned outputs; it does not receive the project mount. Harbor does not mount the project while building the task image. When the task declares workspace setup, RELAI runs it once against the snapshot before any trial, and trials see its output. Docker never mounts the checkout or nested masks, and host changes after snapshot creation are not reflected in that run. Runtime environment values still use the manifest and the environment options below; the mount does not pass project credentials into the task. Before running the target agent, RELAI checks the agent image for python3. This check also applies to cached and custom images. See Compatibility for image requirements.

Runtime and output

Environment variables

Set the runtime variables your simulator needs before running a simulation or optimization. RELAI’s own model calls need none: LLM-judge evaluators, persona turns, and generated component inputs use its model proxy with your relai setup login. Common examples include a model-provider key such as OPENAI_API_KEY, ANTHROPIC_API_KEY, or GOOGLE_API_KEY, and credentials for tools your agent calls. Use the variable names listed in .relai/simulator.env.example; the examples below illustrate common names. For local development, add required runtime variables to .relai/simulator.env. RELAI loads it automatically, and the values persist across runs. Use the command environment for a temporary run or CI. Use --env-file or --env only when you need a different file or an invocation-specific override.

Persistent local values: .relai/simulator.env

relai init creates .relai/simulator.env with an empty NAME= placeholder, and a comment giving its reason, for each variable your simulator declares. Fill in the values after =. RELAI loads this file automatically for both relai simulate and relai optimize when it exists. If the file already exists, init only appends names it does not mention yet, including names you commented out, and never changes your lines. For a project initialized by an older CLI, create the file yourself, for example by copying .relai/simulator.env.example.
A placeholder left empty counts as unset: it does not hide a value from the command environment, so you can keep a credential in your shell instead. Use this file when the same values are needed across local runs. It accepts one KEY=VALUE assignment per line; blank lines, comments, and an optional export prefix are allowed.

One-off or CI values: the command environment

For temporary values, set a variable inline with the command or export it in the shell before running RELAI. This works for both simulation and optimization and is usually the better choice for CI, where the platform injects secrets as environment variables.

Alternate file or command-specific override

Pass --env-file when the values live outside the default local file, or use --env KEY=VALUE for a value that applies only to that command.
RELAI starts with variables from the process environment, then loads .relai/simulator.env, then each --env-file, and finally each --env value. When the same key appears more than once, the later source takes precedence, except that an empty value in a file does not replace an earlier non-empty one. Keep secrets out of version control. .relai/simulator.env is local-only and should not be committed; .relai/simulator.env.example, which init also writes, lists only variable names and can be committed for teammates. In CI, provide secrets through the platform’s secure environment-variable or secret mechanism rather than a tracked file. The run paths below use the current layout, identified by .relai/.internal/layout.json. In a legacy project without that marker, replace .relai/.internal/ with .relai/.

Examples

What to expect

  • terminal simulation summary, grouped by agent test with every evaluator score and evaluator error.
  • optional result JSON when --result-json is provided.
  • a saved result for each completed run at .relai/.internal/runs/results/{kind}/{id}/{run-id}.json: agent-test/{id} for an agent test, benchmark/{id} (the benchmark ID) for a benchmark sample, and harbor-task/{id} for an explicit --harbor-task. It is the same relai.simulation_result.v1 run result that --result-json wraps, with its completed_at time, and is written whether or not --result-json is given. If it cannot be saved, relai simulate warns and the run’s outcome is unchanged.
  • local run artifacts under .relai/.internal/runs/ when produced by the simulator.
  • a Harbor trajectory warning on stderr when the project recorder did not record a turn’s user input as a user step. The run continues, but evaluators that read user turns may score it incorrectly; repair the recorder with relai init --restart.
  • local result persistence. Optional report and transcript attachments remain local artifacts in release builds.
A simulation is locally successful only after the simulator exits successfully and RELAI parses its result JSON. Release builds keep that local result as the complete successful outcome and retain the raw result and available transcript under .relai/.internal/runs/. Batch runs continue after individual execution failures. Their suite report keeps every valid local result: success_count counts parsed local results and failure_count counts local execution failures. Common failures: missing simulator, unknown tag or benchmark, missing runtime credentials, invalid environment file, missing or invalid simulator result JSON, or simulator dependency failure. relai simulate does not intentionally edit agent source. If simulator execution fails, RELAI may use its bounded CLI-owned repair flow for the selected agent test, benchmark, or simulator harness and retry. A diagnosed CLI verifier-wrapper failure is excluded from task-package repair; fix the CLI runtime instead. RELAI cannot repair an unavailable Docker daemon, a full host disk, or missing uv. When two repairs in a row reproduce the same failure, it stops and reports that repair stalled on that failure instead of using more attempts. Diagnostics include available error details and logs, with artifact locations separate from the cause.