Turn a behavior gap into reusable coverage and a reviewed improvement.
The loop at a glance
Initialize the repository before its first learning loop. Initialization builds and validates the project-owned runtime; it is setup, not a stage you repeat for every behavior gap.1
Bring the evidence
Start from a failure, requirement, trace folder, or existing benchmark
2
Ground expectations
Update EvalsWiki, then create or connect reusable coverage
3
Measure
Simulate selected behavior and inspect available evidence
4
Improve & validate
Review scoped changes and overall before-and-after results
What each loop creates
EvalsWiki
The living map of your agent’s behavior: requirements, risk hypotheses, coverage, evidence, and open questions as Markdown in your repository.
Agent tests
Repeatable behavioral scenarios and evaluators for behavior you want to learn or preserve, stored in Harbor format.
Optimized agent
Accepted prompt, tool, workflow, memory, configuration, or code changes with before-and-after evaluation.
Where the artifacts live
EvalsWiki and run results serve different purposesEvalsWiki stores durable project knowledge. Scores and conversations remain run evidence; tests connect the knowledge to behavior you can replay.
Three ways to define expected behavior
Agent test
One repeatable scenario. Start here when you have a failure, risk, or behavior worth preserving.
Benchmark
A reusable suite of evaluation cases you can connect to simulation and optimization. CSV is the currently supported registration format.
Evaluator
A scoring rule. It can belong to one test or benchmark, or apply globally across matching runs.
How to read a simulation
Start with the behavioral evidence, not only the aggregate score:1
Scenario
What behavior was the test trying to reveal or preserve?
2
Conversation
When a transcript is available, what did the simulated user and agent do?
3
Tool evidence
Which authoritative tools ran, with what relevant results?
4
Evaluator feedback
Which expected behavior passed or failed, and why?
5
Earliest failure
Which action first caused later behavior to go wrong?
Do not confuse these outcomes
Valid low score
The agent completed the run but missed an expectation. This is evidence Agent Optimizer can use.
Evaluator error
Scoring did not complete reliably. A missing score is not the same as a zero.
Runtime failure
Dependencies, Docker, credentials, mounts, or provider setup prevented a valid behavioral run.
Share knowledge, keep secrets out of chat and Git
Ready to improve selected behavior?
See how Agent Optimizer scopes, evaluates, and stores the change.