relai simulate runs selected agent tests against the generated simulator.
When to simulate
Simulation measures; it never changes your agent. Run it to see where the agent stands against selected agent tests or benchmarks, or to re-check behavior after changes merge. It is not a prerequisite for optimization.relai optimize
runs its own before/after evaluation passes and does not reuse earlier
simulation results, so when you already know you want the optimizer, you can
skip straight to it.
Select agent tests
At least one selector should identify what to run.
For a project initialized with the Harbor runtime, pass one
explicit task instead of these selectors:
/workspace. The snapshot omits sensitive paths
(including .env and .env.example), .git, .relai, transient Wiki views,
and host dependency roots such as .venv and node_modules. The verifier uses
task-image artifacts and transcripts or other task-owned outputs; it does not
receive the project mount. Harbor does not mount the project while building the
task image. When the task declares workspace setup, RELAI runs it once against
the snapshot before any trial, and trials see its output. Docker never mounts the checkout or nested masks, and host changes
after snapshot creation are not reflected in that run. Runtime environment
values still use the manifest and the environment options below; the mount does
not pass project credentials into the task.
Before running the target agent, RELAI checks the agent image for python3.
This check also applies to cached and custom images. See
Compatibility for image requirements.
Runtime and output
Environment variables
Set the runtime variables your simulator needs before running a simulation or optimization. RELAI’s own model calls need none: LLM-judge evaluators, persona turns, and generated component inputs use its model proxy with yourrelai setup login. Common examples include a
model-provider key such as
OPENAI_API_KEY, ANTHROPIC_API_KEY, or GOOGLE_API_KEY, and credentials for
tools your agent calls. Use the variable names listed in
.relai/simulator.env.example; the examples below illustrate common names.
For local development, add required runtime variables to .relai/simulator.env.
RELAI loads it automatically, and the values persist across runs. Use the command
environment for a temporary run or CI. Use --env-file or --env only when you
need a different file or an invocation-specific override.
Persistent local values: .relai/simulator.env
relai init creates .relai/simulator.env with an empty NAME= placeholder,
and a comment giving its reason, for each variable your simulator declares.
Fill in the values after =. RELAI loads this file automatically for both
relai simulate and relai optimize when it exists. If the file already
exists, init only appends names it does not mention yet, including names you
commented out, and never changes your lines. For a project initialized by an
older CLI, create the file yourself, for example by copying
.relai/simulator.env.example.
KEY=VALUE assignment per line; blank lines, comments, and an optional
export prefix are allowed.
One-off or CI values: the command environment
For temporary values, set a variable inline with the command or export it in the shell before running RELAI. This works for both simulation and optimization and is usually the better choice for CI, where the platform injects secrets as environment variables.Alternate file or command-specific override
Pass--env-file when the values live outside the default local file, or use
--env KEY=VALUE for a value that applies only to that command.
.relai/simulator.env, then each --env-file, and finally each --env value.
When the same key appears more than once, the later source takes precedence,
except that an empty value in a file does not replace an earlier non-empty one.
Keep secrets out of version control. .relai/simulator.env is local-only and
should not be committed; .relai/simulator.env.example, which init also
writes, lists only variable names and can be committed for teammates. In CI, provide secrets through the platform’s secure
environment-variable or secret mechanism rather than a tracked file.
The run paths below use the current layout, identified by
.relai/.internal/layout.json. In a legacy project without that marker, replace
.relai/.internal/ with .relai/.
Examples
What to expect
- terminal simulation summary, grouped by agent test with every evaluator score and evaluator error.
- optional result JSON when
--result-jsonis provided. - a saved result for each completed run at
.relai/.internal/runs/results/{kind}/{id}/{run-id}.json:agent-test/{id}for an agent test,benchmark/{id}(the benchmark ID) for a benchmark sample, andharbor-task/{id}for an explicit--harbor-task. It is the samerelai.simulation_result.v1run result that--result-jsonwraps, with itscompleted_attime, and is written whether or not--result-jsonis given. If it cannot be saved,relai simulatewarns and the run’s outcome is unchanged. - local run artifacts under
.relai/.internal/runs/when produced by the simulator. - a
Harbor trajectorywarning on stderr when the project recorder did not record a turn’s user input as auserstep. The run continues, but evaluators that read user turns may score it incorrectly; repair the recorder withrelai init --restart. - local result persistence. Optional report and transcript attachments remain local artifacts in release builds.
.relai/.internal/runs/.
Batch runs continue after individual execution failures. Their suite report
keeps every valid local result: success_count counts parsed local results and
failure_count counts local execution failures.
Common failures: missing simulator, unknown tag or benchmark, missing runtime
credentials, invalid environment file, missing or invalid simulator result JSON,
or simulator dependency failure.
relai simulate does not intentionally edit agent source. If simulator execution fails, RELAI may use its bounded CLI-owned repair flow for the selected agent test, benchmark, or simulator harness and retry. A diagnosed CLI verifier-wrapper failure is excluded from task-package repair; fix the CLI runtime instead.
RELAI cannot repair an unavailable Docker daemon, a full host disk, or missing
uv. When two repairs in a row reproduce the same failure, it stops and
reports that repair stalled on that failure instead of using more attempts.
Diagnostics include available error details and logs, with artifact locations
separate from the cause.