> ## Documentation Index
> Fetch the complete documentation index at: https://cli-docs.relai.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent tests

> Turn requirements, risks, and feedback into repeatable RELAI agent tests.

Each agent test captures one behavior as a scenario, runnable environment, and evaluators. Run it again as your agent changes, or give it to Agent Optimizer as a measurable target.

## Choose the right input

| You have | Use | Result |
| - | - | - |
| A known behavior | Create from a prompt | One agent test |
| An EvalsWiki requirement | Create grounded coverage | One test linked to that requirement |
| An EvalsWiki risk | Discover a hard scenario | Failure-revealing or useful regression coverage |
| One run plus explicit feedback | Create from a local log | One test for the observed gap |
| A folder of many traces | [Trace Ingestion](/trace-ingestion) | Curated EvalsWiki knowledge—not a test |

<Note>
  **Use agent tests for behavioral uncertainty.**

  They fit judgment, policy, tool choice and order, grounding, multi-turn state, and semantic outcomes. Keep deterministic code defects in ordinary unit or integration tests.
</Note>

## Create from a known behavior

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to create one agent test for this behavior: “…”. Check EvalsWiki and existing tests first, then summarize the scenario and evaluator criteria." actions={["copy"]}>
      Use RELAI to create one agent test for this behavior: “…”. Check EvalsWiki and existing tests first, then summarize the scenario and evaluator criteria.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai test create --prompt "<behavior to test>" \
      --permissions ask
    ```
  </Tab>
</Tabs>

## Create from one failed run

The log must be inside the repository, and feedback must state what should have happened. Logs with tool calls and state transitions produce better grounding than user-visible text alone.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to create an agent test from the local run log at <path>. The expected behavior was: “…”. Summarize the scenario and evaluators." actions={["copy"]}>
      Use RELAI to create an agent test from the local run log at \<path>. The expected behavior was: “…”. Summarize the scenario and evaluators.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai test create --log-file <path/to/run.log> \
      --feedback "<what should have happened>" \
      --permissions ask
    ```
  </Tab>
</Tabs>

## Create or discover from EvalsWiki

Requirements create direct coverage. Risk discovery proposes and critiques hard scenarios, then promotes failure-revealing or coverage-expanding tests.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to find the highest-priority uncovered EvalsWiki risk and propose one hard agent test. Show me the scenario before building or simulating it." actions={["copy"]}>
      Use RELAI to find the highest-priority uncovered EvalsWiki risk and propose one hard agent test. Show me the scenario before building or simulating it.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    # Create from a requirement
    relai test create --requirement <requirement-id> --permissions ask

    # Or discover from a risk
    relai test discover --risk <risk-id>
    ```
  </Tab>
</Tabs>

Risk discovery requires a completed EvalsWiki bootstrap. By default it pauses after proposing and critiquing the scenario, before the expensive build and run; use `--no-review` only for deliberate automation. Its report separates failures revealed from useful coverage the current agent already passes.

## Keep ownership and recovery clear

Creation writes a local package under `.relai/harbor/agent-tests/`; it does not publish the test to the backend or modify application source. RELAI may repair CLI-owned simulator support while validating the candidate. If the shared runtime itself needs a fix, rerun initialization.

Commit a reviewed agent-test package together with `.relai/harbor/runtime/` when collaborators should run it. Keep draft sessions, run evidence, and caches local.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Inspect the unfinished RELAI agent-test workflow and resume its saved progress. Explain anything that still needs my action." actions={["copy"]}>
      Inspect the unfinished RELAI agent-test workflow and resume its saved progress. Explain anything that still needs my action.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai test create --resume --permissions ask

    # Discard the unfinished creation session and start from a new source
    relai test create --restart \
      --prompt "<behavior to test>" \
      --permissions ask
    ```
  </Tab>
</Tabs>

Creation and update have a 90-minute limit per invocation; discovery has a four-hour limit. Resume starts a new invocation budget, so repeat any `--timeout` or permission override you still need.

## Update related coverage instead of duplicating it

When a correction extends the same behavior, update the existing test. RELAI preserves its ID and target, checks current EvalsWiki grounding, and stops if the request belongs in a new test.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to update agent test <test-id> for this related correction: “…”. Preserve its identity and stop if this should be a separate test." actions={["copy"]}>
      Use RELAI to update agent test \<test-id> for this related correction: “…”. Preserve its identity and stop if this should be a separate test.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai test update <test-id> --prompt "<related correction>"
    ```
  </Tab>
</Tabs>

## Simulate and read the result

A simulation measures behavior; it does not intentionally change application source. Keep valid low scores, evaluator errors, and runtime failures separate.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to simulate agent test <test-id> once. Show every evaluator score or error and any available conversation or tool evidence. Identify the earliest supported failure without inventing missing evidence." actions={["copy"]}>
      Use RELAI to simulate agent test \<test-id> once. Show every evaluator score or error and any available conversation or tool evidence. Identify the earliest supported failure without inventing missing evidence.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai simulate --tests <test-id>
    ```
  </Tab>
</Tabs>

## How agent tests are scored

<CardGroup cols={2}>
  <Card title="Code evaluator">
    Checks facts such as whether a tool ran, which arguments it received, the resulting state, or a required exact value.
  </Card>

  <Card title="Semantic evaluator">
    Uses an LLM judge for policy, tone, refusal quality, groundedness, or any criterion with many valid phrasings.
  </Card>
</CardGroup>

RELAI’s LLM judges and simulated personas use its model proxy through your RELAI sign-in. Provider keys are needed only for the agent or tools you are testing.

## How an agent test is stored

RELAI saves each test under `.relai/harbor/agent-tests/<test-id>/` in Harbor format. Harbor is the execution format; the user-facing object is still an agent test. Every test reuses the shared project runtime created during initialization.

| Path | Role |
| - | - |
| `task.toml` | Stable test identity, target, input, timeouts, and evaluator registry |
| `environment/scenario.json` | Turns, fixtures, services, and state for the replayable scenario |
| `tests/evaluate.py` | Exact checks such as tool calls, arguments, and state |
| `tests/evaluator.toml` or `tests/<id>/evaluator.toml` | Semantic evaluator configuration and rubric |
| `instruction.md` | Required Harbor file; normally empty for generated tests |
| `generated/` | CLI-owned packaging metadata; do not hand-edit |
| `.relai/harbor/runtime/` | Shared adapter, dependencies, and launcher owned by initialization |

If the shared adapter or dependencies need repair, rerun initialization. An individual test should not redefine the project runtime.

<Card title="Measure the current agent" href="/simulation">
  Run the test and read its available conversation, tool evidence, evaluator scores, and failures.
</Card>

## Command reference

See [`relai test`](/cli/test) for all flags, defaults, and recovery details.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.