From one behavior gap to a reviewed improvement
Initialize the repository before its first loop. Then repeat this evidence-driven cycle whenever behavior needs to improve.1
Bring the evidence
Start from a failure, requirement, trace folder, or existing benchmark.
2
Ground the expectation
Update EvalsWiki, then create or connect reusable coverage.
3
Measure behavior
Simulate the selected cases and inspect the available evidence. Simulation is optional before optimization, which runs its own evaluations.
4
Improve & validate
Review scoped changes and before-and-after results. Adopt accepted changes through your normal review workflow.
Choose where to start
Try the sample agent
Choose Python, TypeScript, or Go, then follow short prompts through setup, testing, simulation, and optimization.
Run your first learning loop
Check compatibility, initialize your repository, create one agent test, and use Agent Optimizer.
What you will have at the end
EvalsWiki
The living map of your agent’s behavior: grounded requirements, risk hypotheses, coverage, evidence, and open questions as Markdown in your repository.
Agent tests
Replayable scenarios, environments, and evaluators you can simulate again as the agent changes.
An optimized agent
When a candidate improves the overall evaluation, Agent Optimizer produces reviewable harness changes with before-and-after evidence.