Run a focused simulation
Start with one test. Ask for every evaluator outcome and any readable trajectory the run produced—not only an aggregate score.- Ask your coding agent
- Run in terminal
Use RELAI to simulate agent test once. Show every evaluator score or error, plus any available conversation and tool evidence. Identify the earliest supported failure without inventing missing evidence.
Know what enters the run
Select exactly what to measure
- Ask your coding agent
- Run in terminal
Use RELAI to show the available tests, tags, and benchmarks for this agent. Recommend the smallest selector that answers: “…”. Do not run it yet.
Read the outcome correctly
Valid low score
Evaluator error
Runtime failure
Runtime values and model access
Use the variable names in.relai/simulator.env.example. Put persistent local values in ignored .relai/simulator.env; use your shell or CI secret store for temporary values.
Values are resolved in this order: process environment, .relai/simulator.env, each --env-file, then each --env; later values win. RELAI forwards only manifest-declared names from the process environment, but forwards the values supplied by environment files and --env to the agent phase. Keep copied commands and chat free of secret values.
Know what simulation may change
Simulation does not intentionally edit your application source. If execution fails, RELAI may use a bounded repair flow for the selected agent test, benchmark, or simulator harness and retry. Those are CLI-owned evaluation artifacts, not agent-source optimization. It cannot repair an unavailable Docker daemon, a full host disk, or missinguv, and it stops when two consecutive repairs reproduce the same failure.
A simulation has a three-hour wall-clock limit by default. Use --timeout <duration|off> when a selected run needs a different invocation budget.
Where results live
RELAI keeps completed results and available transcripts under.relai/.internal/runs/. Each completed run gets a local result under results/<kind>/<id>/<run-id>.json, even without --result-json. Projects without .relai/.internal/layout.json use the legacy .relai/runs/ directory. These are local evidence, not EvalsWiki content and not artifacts to commit. Request an explicit machine-readable output file when another tool needs the suite result.
- Ask your coding agent
- Run in terminal
Use RELAI to run and save a machine-readable result. Also show me the readable outcome; do not dump raw JSON into chat.
Found a real behavior gap?
Command reference
Seerelai simulate for all flags, defaults, and recovery details.