Skip to main content
RELAI turns requirements, failures, and feedback into repeatable evals, then uses measured behavior to guide scoped improvements. Every accepted eval stays available for the next loop.
A requirement gap leads to defined expectations, optimization and validation, and a reviewed improvement before the loop repeats.A requirement gap leads to defined expectations, optimization and validation, and a reviewed improvement before the loop repeats.

Turn a behavior gap into reusable coverage and a reviewed improvement.

The loop at a glance

Initialize the repository before its first learning loop. Initialization builds and validates the project-owned runtime; it is setup, not a stage you repeat for every behavior gap.
1

Bring the evidence

Start from a failure, requirement, trace folder, or existing benchmark
2

Ground expectations

Update EvalsWiki, then create or connect reusable coverage
3

Measure

Simulate selected behavior and inspect available evidence
4

Improve & validate

Review scoped changes and overall before-and-after results
Harness means the parts around the model that shape behavior: prompts, context, tools, workflow, memory, configuration, and agent logic.

What each loop creates

EvalsWiki

The living map of your agent’s behavior: requirements, risk hypotheses, coverage, evidence, and open questions as Markdown in your repository.

Agent tests

Repeatable behavioral scenarios and evaluators for behavior you want to learn or preserve, stored in Harbor format.

Optimized agent

Accepted prompt, tool, workflow, memory, configuration, or code changes with before-and-after evaluation.

Where the artifacts live

EvalsWiki and run results serve different purposesEvalsWiki stores durable project knowledge. Scores and conversations remain run evidence; tests connect the knowledge to behavior you can replay.

Three ways to define expected behavior

Agent test

One repeatable scenario. Start here when you have a failure, risk, or behavior worth preserving.

Benchmark

A reusable suite of evaluation cases you can connect to simulation and optimization. CSV is the currently supported registration format.

Evaluator

A scoring rule. It can belong to one test or benchmark, or apply globally across matching runs.

How to read a simulation

Start with the behavioral evidence, not only the aggregate score:
1

Scenario

What behavior was the test trying to reveal or preserve?
2

Conversation

When a transcript is available, what did the simulated user and agent do?
3

Tool evidence

Which authoritative tools ran, with what relevant results?
4

Evaluator feedback

Which expected behavior passed or failed, and why?
5

Earliest failure

Which action first caused later behavior to go wrong?

Do not confuse these outcomes

Valid low score

The agent completed the run but missed an expectation. This is evidence Agent Optimizer can use.

Evaluator error

Scoring did not complete reliably. A missing score is not the same as a zero.

Runtime failure

Dependencies, Docker, credentials, mounts, or provider setup prevented a valid behavioral run.

Share knowledge, keep secrets out of chat and Git

Ready to improve selected behavior?

See how Agent Optimizer scopes, evaluates, and stores the change.

Memory and tags

Memory and tags carry guidance forward and select reusable agent-test groups. EvalsWiki holds requirements, risks, and open questions; memory guides later generation and optimization.