Skip to main content
Each agent test captures one behavior as a scenario, runnable environment, and evaluators. Run it again as your agent changes, or give it to Agent Optimizer as a measurable target.

Choose the right input

Use agent tests for behavioral uncertainty.They fit judgment, policy, tool choice and order, grounding, multi-turn state, and semantic outcomes. Keep deterministic code defects in ordinary unit or integration tests.

Create from a known behavior

Use RELAI to create one agent test for this behavior: “…”. Check EvalsWiki and existing tests first, then summarize the scenario and evaluator criteria.

Create from one failed run

The log must be inside the repository, and feedback must state what should have happened. Logs with tool calls and state transitions produce better grounding than user-visible text alone.

Use RELAI to create an agent test from the local run log at . The expected behavior was: “…”. Summarize the scenario and evaluators.

Create or discover from EvalsWiki

Requirements create direct coverage. Risk discovery proposes and critiques hard scenarios, then promotes failure-revealing or coverage-expanding tests.

Use RELAI to find the highest-priority uncovered EvalsWiki risk and propose one hard agent test. Show me the scenario before building or simulating it.

Risk discovery requires a completed EvalsWiki bootstrap. By default it pauses after proposing and critiquing the scenario, before the expensive build and run; use --no-review only for deliberate automation. Its report separates failures revealed from useful coverage the current agent already passes.

Keep ownership and recovery clear

Creation writes a local package under .relai/harbor/agent-tests/; it does not publish the test to the backend or modify application source. RELAI may repair CLI-owned simulator support while validating the candidate. If the shared runtime itself needs a fix, rerun initialization. Commit a reviewed agent-test package together with .relai/harbor/runtime/ when collaborators should run it. Keep draft sessions, run evidence, and caches local.

Inspect the unfinished RELAI agent-test workflow and resume its saved progress. Explain anything that still needs my action.

Creation and update have a 90-minute limit per invocation; discovery has a four-hour limit. Resume starts a new invocation budget, so repeat any --timeout or permission override you still need. When a correction extends the same behavior, update the existing test. RELAI preserves its ID and target, checks current EvalsWiki grounding, and stops if the request belongs in a new test.

Use RELAI to update agent test for this related correction: “…”. Preserve its identity and stop if this should be a separate test.

Simulate and read the result

A simulation measures behavior; it does not intentionally change application source. Keep valid low scores, evaluator errors, and runtime failures separate.

Use RELAI to simulate agent test once. Show every evaluator score or error and any available conversation or tool evidence. Identify the earliest supported failure without inventing missing evidence.

How agent tests are scored

Code evaluator

Checks facts such as whether a tool ran, which arguments it received, the resulting state, or a required exact value.

Semantic evaluator

Uses an LLM judge for policy, tone, refusal quality, groundedness, or any criterion with many valid phrasings.
RELAI’s LLM judges and simulated personas use its model proxy through your RELAI sign-in. Provider keys are needed only for the agent or tools you are testing.

How an agent test is stored

RELAI saves each test under .relai/harbor/agent-tests/<test-id>/ in Harbor format. Harbor is the execution format; the user-facing object is still an agent test. Every test reuses the shared project runtime created during initialization. If the shared adapter or dependencies need repair, rerun initialization. An individual test should not redefine the project runtime.

Measure the current agent

Run the test and read its available conversation, tool evidence, evaluator scores, and failures.

Command reference

See relai test for all flags, defaults, and recovery details.