For a guided workflow, see Agent tests.
Use relai test to discover, suggest, create, or update repeatable
scenarios.
discover
relai test discover --count 3 --prompt "focus on unsafe refunds" investigates
up to the requested number of accepted Harbor agent tests. Use
relai test discover --risk {id} to target exactly one curated risk
hypothesis from relai wiki list risks. Each discovery invocation runs for at most
4 hours; --timeout 2h or --timeout off changes that limit (see
Timeouts). When the timeout is reached, discovery saves its session;
continue only with explicit --resume, or replace it with --restart.
--no-review or --yes skips human review. --timeout 45m bounds the full
invocation, including Harbor validation and simulation; each discovery resume
has a new invocation budget, so repeat --timeout on --resume to keep an
override or set RELAI_TIMEOUT.
When no risk is specified, discovery tops up the eligible risk pool before
starting candidate generation. That curation is quality-gated: it may add many
supported hypotheses, but it should stop early instead of inventing filler records.
Discovery requires a completed EvalsWiki bootstrap. relai init performs that bootstrap by
default; projects initialized with relai init --no-wiki must run relai wiki bootstrap first.
Discovery roles can read permitted project source and EvalsWiki material, but cannot read
credentials, RELAI runtime state, staged drafts, dependency caches, or run commands. Each
candidate is built in an isolated Harbor draft, validated and simulated by the host, and
reviewed after simulation as either failure-revealing or coverage-only. Failure-revealing
acceptance requires the exact simulated draft to have at least one task-local evaluator score
below its declared maximum. A valid, useful agent test that the current agent passes may be
accepted as coverage, but it is counted separately and does not confirm the risk as
triggered. Evaluator execution errors pause the candidate rather than being treated as agent
failures. Risk hypotheses are saved as typed EvalsWiki records with Markdown pages and
reciprocal agent test lineage. They are
ranked deterministically by impact, confidence, and triggerability. Rejected means simulation
or review evidence refuted the hypothesis; Retired means it is no longer worth pursuing.
Before accepting a Harbor draft, RELAI checks the task-owned JSON assets named
by its descriptor. Malformed authored input, persona, component-adapter,
generator, or mock assets return through the current validation and repair flow.
Final scenario requirements produced by image-local setup are checked at runtime;
Harbor validates the assets again in the task image because Dockerfile inputs can
differ from the host package.
Interactive discovery shows the proposed agent test before the expensive build/run step.
Choose yes to build, no to skip the candidate without retiring the risk, or revise to return
the spec to proposal with feedback. Without a terminal, review is skipped unless
--agent-interaction is set.
Discovery resolves --agent-target with the same rules as create: a sole configured target is
selected automatically, while multiple configured targets require --agent-target outside an
interactive terminal. Accepted artifacts retain package, result, and transcript digests in their
discovery lineage; no transcript copy is published.
With --agent-interaction, discovery writes one relai.agent_interaction.v1 document to
stdout per invocation and keeps human output on stderr. Proposal review becomes a
clarification with choices build, build_all (approve the rest of this run), and skip;
any free-form answer is revision feedback. Model failures, builder questions and approvals,
and timeouts pause with relai test discover --resume --agent-interaction. Builder questions,
approvals, and builder timeouts come from the shared agent-test builder, so their documents
carry its learning_env_create workflow and session ID; answer them with that session ID.
Declining a builder request retires that candidate and cancels the invocation (exit 130);
--resume continues with the next candidate. A run that ends,
including one short of its --count with nothing left to resume, returns a completed summary
of created tests, evaluator scores, each created test’s simulation trajectory, and next steps.
The terminal summary separates accepted tests into:
- Failures revealed: new tests where the current agent failed.
- Coverage expanded: new tests the current agent already passes, but that are useful for preventing future regressions.
create
Create an agent test from a prompt, from a run log plus feedback, or from a
grounded EvalsWiki requirement.
Agent test generation can take a while when the prompt, run log, or agent context is complex.
The run log does not need a specific structure: any repository-local text format under the
size limit works, including JSONL message logs, plain transcripts, or your
own trace format. RELAI parses it. Logs that include internal detail — tool
calls, retrievals, intermediate reasoning — give RELAI more to work with than
user-facing output alone.
The CLI agent inspects the repository, asks a focused question only when needed,
reports its runtime plan, generates a candidate, and validates both the
candidate and simulator before saving the final file under
.relai/harbor/agent-tests/. The completion summary names the validation run’s
ATIF trajectory (Validation trajectory), so you can see what the target agent did. Clarification continues in the same process. The agent
may repair .relai/.internal/simulator as part of creation when
.relai/.internal/layout.json is present; legacy projects without the marker
use .relai/simulator. It cannot modify application source. External setup
work is explained and saved for --resume.
If a model stops before validation, RELAI checks the staged artifact once and
uses a failed check as repair feedback. Missing external prerequisites remain
paused for --resume.
create, update, suggestion generate, and suggestion accept accept
--permissions {auto|accept_edits|ask}. See
Permissions.
Creation reuses the project-specific component, input, log, and mock guidance
recorded by relai init. For known frameworks it also has bundled RELAI
guidance. For custom frameworks or version-specific behavior, it can inspect
repository source and bounded public declarations or source from the exact
installed dependency. Current project evidence takes precedence over bundled
framework hints.
create saves the agent test under .relai/harbor/agent-tests/. Commit that
package, together with .relai/harbor/runtime/, to share it with collaborators;
RELAI does not publish it to the backend.
An agent test reuses the project runtime that relai init validated: the
same base image, locked dependencies, adapter, and launcher. The test holds only
its scenario, mocks, evaluators, and test-specific setup, and it cannot change
the runtime. If the environment itself needs a fix, rerun relai init.
Each criterion the test scores is recorded as exact or semantic. Exact criteria,
such as a tool call, its arguments, or a required string, can be scored by
code. Semantic criteria, such as refusing an off-topic request, tone, or an
answer’s meaning, are always scored by an LLM judge, because no keyword or
pattern check covers every valid phrasing. RELAI scores judges through
its model proxy with your relai setup login; they need no model-provider
keys. The same holds for a test’s simulated persona turns and generated
component inputs.
Agent tests are scoped to the selected agent target. When multiple
targets are registered, interactive runs present a single-target picker; pass
--agent-target {target} in non-interactive runs.
Common failures: log outside the repository, missing paired --feedback, validation failure,
duplicate explicit name, or required external setup. Use
--resume after completing recoverable setup.
suggestion
Generate suggestions, then list, review, accept, or reject them:
Suggestion runs and decisions are ignored run state scoped to the selected agent
target. accept uses the standard agent test creation workflow, and
clear-memory --yes removes the selected target’s stored suggestion history.
These commands do not register an agent or persist suggestions in the backend.