Skip to main content
RELAI measures the current agent, explores scoped harness changes, and accepts only revisions whose overall evaluation improves before putting them on a reviewable branch.

Optimization includes validation

1

Evaluate before

Measure the selected behavior
2

Explore changes

Improve prompts, tools, workflow, memory, or code within scope
3

Compare evidence

Discard candidates whose weighted overall evaluation does not improve
4

Evaluate after

Commit accepted changes with measured evidence
You do not need to simulate first or add a separate “verify” stage. Agent Optimizer runs fresh before-and-after evaluation; earlier simulation results are not reused. Your job is to review the result and decide whether to adopt it.

Start with a bounded request

Use RELAI Agent Optimizer on in Balanced quality/cost mode. Show the scope and rollout budget first, keep changes local, and report before-and-after evidence.

Select the behavior deliberatelyWithout --tests, --tags, or --benchmarks, Agent Optimizer includes every available agent test and registered benchmark sample. Start with the failures you intend to improve, and select one --agent-target in a multi-target project.

Confirm the optimizer scope

Before provider-backed work begins, review the scope that controls what RELAI may change: The default scope allows structural changes, keeps the current models fixed, blocks new paid or credentialed tools, and disables cost optimization. Your confirmed scope is saved under .relai/.internal/optimizer-scope.json in current projects; legacy layouts use .relai/optimizer-scope.json. A non-interactive run needs that saved decision.

Choose the quality/cost objective

A cost objective is not proof of savingsReport token or dollar reductions only when the before-and-after telemetry is comparable. Otherwise report Cost: not measured.

Know what decides acceptance

Every rollout scores the selected agent test or benchmark evaluators plus all matching global evaluators. RELAI keeps those evaluation packages outside the optimizer’s editable source snapshot, so the target evidence is not part of the code it may change. A candidate is accepted only when the weighted overall evaluation improves. A large gain on one evaluator can outweigh a slight regression on another; an equal gain-for-loss trade does not pass. Review the individual scores as well as the aggregate result.

Size the rollout budget

RELAI uses one third of --total-rollouts for candidate exploration and reserves two thirds for final before-and-after evaluation. The default is the larger of 30 or 6 × batch-size; the minimum is 6 × batch-size, and batch size defaults to 1.
  • Keep the first run focused on one or a few known failures.
  • Every selected sample is evaluated once before scheduling favors the weakest cases, so large benchmarks need more budget or a focused subsample.
  • Use a smaller --batch-size when one simulation produces a long conversation or tool trajectory.
  • Interactive runs ask for confirmation above 50 total rollouts; non-interactive runs warn.
  • Use the optimizer sizing reference for advanced tuning.

Where the optimized agent lives

RELAI creates a separate worktree under ~/.relai/worktrees and a local relai/optimizer/… branch. Your main checkout, index, and HEAD remain unchanged. The worktree starts from your current source state, including staged, unstaged, untracked, deleted, renamed, and Git-ignored files. Publishable pre-existing changes are saved in a baseline commit before accepted optimizer changes, so review git status first: an automatic pull request includes that baseline.

Show me the accepted optimizer branch and diff. Explain what changed and connect each change to the before-and-after evaluation.

The optimizer branch is already checked out in its worktree. Inspect that worktree directly instead of trying to check out the same branch in your main checkout. When authenticated GitHub CLI is available, RELAI pushes accepted changes and opens a pull request. Use --no-pr to keep the result on the local optimizer branch.

Read the optimizer report

Allow time and recover safely

The whole optimizer run defaults to a six-hour timeout. Within it, RELAI waits at most 30 minutes for one backend optimizer task; each Harbor trial’s run phase may use its task’s agent and verifier timeouts plus five minutes, capped at two hours.
Interrupted validation is not a validated resultOn timeout, RELAI stops at a safe step and prints a resume command when progress can continue. If evaluator repair pauses, is cancelled, or fails, accepted changes remain on the branch but final validation is incomplete. Use the exact reported recovery step, then start a fresh optimization after the project is repaired; do not blind-retry or describe the branch as validated.

Continue the learning loop

Review or adopt the accepted branch, then turn a remaining gap or harder risk into the next agent test.
Run incomplete? Separate evaluator and runtime failures from agent behavior before retrying.

Command reference

See relai optimize for all flags, defaults, and recovery details.