Skip to main content
Initialize your project, create one repeatable agent test, simulate current behavior, and let Agent Optimizer improve and validate the agent.
Before the learning loopComplete Install & setup and confirm project compatibility. A coding-agent plugin is recommended for the prompt-driven path, but it is optional—the terminal tabs use the same CLI directly.

1. Initialize the repository

Start at the root of a git-tracked agent project. RELAI creates and validates the simulator harness, registers the agent target, and bootstraps EvalsWiki. Already initialized? Skip to step 2. Rerun initialization only when RELAI reports that the shared runtime needs repair or regeneration.

Use RELAI to initialize this agent project. Keep provider keys out of chat and Git.

If initialization pauses, complete the reported action and use the exact resume command RELAI prints.

2. Create one agent test

Choose the source that matches what you know. Start with one behavior rather than a broad suite.

Use RELAI to find the highest-priority uncovered EvalsWiki risk. Show me one proposed agent-test scenario, then create it after I approve.

See Agent Tests for requirements, failed runs, and how agent tests are stored.

3. Simulate current behavior

Run exactly the test you created. A valid low score is useful behavioral evidence; an evaluator or runtime error is a different outcome.

Use RELAI to simulate that agent test once. Show every evaluator score or error, any available conversation and tool evidence, and the earliest supported failure.

A stopped Docker daemon, a missing provider variable, or a dependency missing from the task image prevents a valid behavioral run. Fix the reported prerequisite and rerun the same test; do not treat infrastructure or evaluator errors as a behavioral score.

4. Use Agent Optimizer

Agent Optimizer runs its own before-and-after evaluation, proposes scoped harness changes in a separate worktree, and validates candidates against the selected tests and evaluators.

Use RELAI Agent Optimizer on the failing test. Show the scope and rollout budget first, keep accepted changes local, and report before-and-after results.

Review the optimizer output

Read the accepted diff, its explanation, the final evaluation, and any remaining gap. Your main checkout stays unchanged.
Your first learning loop is complete whenEvalsWiki is initialized, an agent test has been evaluated, and you have reviewed the optimizer’s before-and-after report. When RELAI accepts a change, it also leaves a local optimizer branch ready for review; a run that accepts no change can still produce a useful report.

Understand what RELAI created

See how EvalsWiki, agent tests, simulations, and optimized code fit together.