Skip to main content
A low behavioral score, an evaluator error, and a broken runtime need three different responses.

First: classify the outcome

Low score

This is valid behavioral evidence. Read the evaluator feedback and first failing turn; do not repair or blindly rerun it.

Evaluator error

Report it separately. The agent’s behavior may still be usable evidence, but the missing score is not a zero.

Runtime failure

Fix Docker, dependencies, credentials, mounts, or provider setup before interpreting agent quality.
Do not blind-retry agentic workflowsRead the reported session and recovery command first. An unfinished write may need --resume; restarting can discard saved progress or repeat provider cost.

Common first-run problems

Add the named variable to the ignored local .relai/simulator.env file, your shell, or a secure CI secret. Never paste its value into chat, commit it, or put it in a copied command.RELAI OAuth signs the CLI into RELAI; your agent’s provider key is a separate runtime value.
Run the prerequisite audit, then start Docker and confirm its daemon is reachable before retrying the dependent workflow.

Use RELAI to diagnose the missing host prerequisite. Check Docker without changing my project and explain the exact fix.

Host prerequisites cannot be repaired from inside an agent-test image.
Install or refresh the matching local package, then activate it for the host: start a new Codex or Cursor session, run /reload-plugins in Claude Code, or use an interactive Copilot session.See Coding-agent plugins for the exact command and host-specific step.
Initialization creates the simulator and EvalsWiki; it does not need to create an agent test. Inspect the curated risks, then target one.

Use RELAI to show my EvalsWiki risks and propose one high-value agent test. Wait before creating it.

Use the exact resume command printed by RELAI. Use --restart only when you deliberately want to discard the unfinished session and start over.Do not invent a resume command or reuse a stale interaction response.
This is an image dependency problem, not an agent behavior score. Every RELAI agent image needs python3 for the launcher, even when the project itself is TypeScript or Go.Repair the image or re-run initialization so the project-owned runtime setup includes the missing interpreter or package.
Some Docker/Colima configurations do not expose macOS system temporary paths inside Linux containers. Use a host-visible task-specific temporary directory under your user directory, then resume or restart the owning RELAI workflow as instructed.Keep this diagnosis separate from application imports: the same project can work correctly once the host mount is visible.
A timeout is not an agent score. Read stderr to learn whether progress was saved. Resume when RELAI provides a resume path; otherwise rerun only after checking the underlying cause and budget.

Results and artifact confusion

This is expected. EvalsWiki stores durable requirements, risks, decisions, and open questions. Scores and conversations live in run artifacts under .relai/.internal/runs/.See The learning loop for the complete artifact map.
Execution success means RELAI obtained and parsed a valid result. It does not mean the agent passed every evaluator. Read every score and its maximum.
The branch is usually already checked out in the optimizer worktree. Locate it and inspect that directory directly.

Find the RELAI optimizer worktree and show me its branch and diff without changing either checkout.

Balanced mode changes the optimizer’s objective; it does not guarantee that comparable token or dollar telemetry will appear in the report. If the artifacts contain no numeric cost evidence, report Cost change: not measured.

Useful information when asking for help

  • The RELAI version and exact public command used
  • The terminal exit status and sanitized error category
  • Whether the failure occurred before, during, or after agent execution
  • The test ID, evaluator ID, and local artifact path
  • Whether a resumable session or retry constraint was reported
Never include secretsDo not share provider values, OAuth credentials, environment-file contents, or unrelated private project data in a support report.

Return to a bounded first run

Once the prerequisite is fixed, resume the smallest selected workflow.