1. Initialize the repository
Start at the root of a git-tracked agent project. RELAI creates and validates the simulator harness, registers the agent target, and bootstraps EvalsWiki. Already initialized? Skip to step 2. Rerun initialization only when RELAI reports that the shared runtime needs repair or regeneration.- Ask your coding agent
- Run in terminal
Use RELAI to initialize this agent project. Keep provider keys out of chat and Git.
2. Create one agent test
Choose the source that matches what you know. Start with one behavior rather than a broad suite.- Ask your coding agent
- Run in terminal
Use RELAI to find the highest-priority uncovered EvalsWiki risk. Show me one proposed agent-test scenario, then create it after I approve.
3. Simulate current behavior
Run exactly the test you created. A valid low score is useful behavioral evidence; an evaluator or runtime error is a different outcome.- Ask your coding agent
- Run in terminal
Use RELAI to simulate that agent test once. Show every evaluator score or error, any available conversation and tool evidence, and the earliest supported failure.
4. Use Agent Optimizer
Agent Optimizer runs its own before-and-after evaluation, proposes scoped harness changes in a separate worktree, and validates candidates against the selected tests and evaluators.- Ask your coding agent
- Run in terminal
Use RELAI Agent Optimizer on the failing test. Show the scope and rollout budget first, keep accepted changes local, and report before-and-after results.
Review the optimizer output
Read the accepted diff, its explanation, the final evaluation, and any remaining gap. Your main checkout stays unchanged.Understand what RELAI created
See how EvalsWiki, agent tests, simulations, and optimized code fit together.