Choose the right input
Create from a known behavior
- Ask your coding agent
- Run in terminal
Use RELAI to create one agent test for this behavior: “…”. Check EvalsWiki and existing tests first, then summarize the scenario and evaluator criteria.
Create from one failed run
The log must be inside the repository, and feedback must state what should have happened. Logs with tool calls and state transitions produce better grounding than user-visible text alone.- Ask your coding agent
- Run in terminal
Use RELAI to create an agent test from the local run log at . The expected behavior was: “…”. Summarize the scenario and evaluators.
Create or discover from EvalsWiki
Requirements create direct coverage. Risk discovery proposes and critiques hard scenarios, then promotes failure-revealing or coverage-expanding tests.- Ask your coding agent
- Run in terminal
Use RELAI to find the highest-priority uncovered EvalsWiki risk and propose one hard agent test. Show me the scenario before building or simulating it.
--no-review only for deliberate automation. Its report separates failures revealed from useful coverage the current agent already passes.
Keep ownership and recovery clear
Creation writes a local package under.relai/harbor/agent-tests/; it does not publish the test to the backend or modify application source. RELAI may repair CLI-owned simulator support while validating the candidate. If the shared runtime itself needs a fix, rerun initialization.
Commit a reviewed agent-test package together with .relai/harbor/runtime/ when collaborators should run it. Keep draft sessions, run evidence, and caches local.
- Ask your coding agent
- Run in terminal
Inspect the unfinished RELAI agent-test workflow and resume its saved progress. Explain anything that still needs my action.
--timeout or permission override you still need.
Update related coverage instead of duplicating it
When a correction extends the same behavior, update the existing test. RELAI preserves its ID and target, checks current EvalsWiki grounding, and stops if the request belongs in a new test.- Ask your coding agent
- Run in terminal
Use RELAI to update agent test for this related correction: “…”. Preserve its identity and stop if this should be a separate test.
Simulate and read the result
A simulation measures behavior; it does not intentionally change application source. Keep valid low scores, evaluator errors, and runtime failures separate.- Ask your coding agent
- Run in terminal
Use RELAI to simulate agent test once. Show every evaluator score or error and any available conversation or tool evidence. Identify the earliest supported failure without inventing missing evidence.
How agent tests are scored
Code evaluator
Semantic evaluator
How an agent test is stored
RELAI saves each test under.relai/harbor/agent-tests/<test-id>/ in Harbor format. Harbor is the execution format; the user-facing object is still an agent test. Every test reuses the shared project runtime created during initialization.
Measure the current agent
Command reference
Seerelai test for all flags, defaults, and recovery details.