Skip to main content
For a guided workflow, see Agent tests. Use relai test to discover, suggest, create, or update repeatable scenarios.

discover

relai test discover --count 3 --prompt "focus on unsafe refunds" investigates up to the requested number of accepted Harbor agent tests. Use relai test discover --risk {id} to target exactly one curated risk hypothesis from relai wiki list risks. Each discovery invocation runs for at most 4 hours; --timeout 2h or --timeout off changes that limit (see Timeouts). When the timeout is reached, discovery saves its session; continue only with explicit --resume, or replace it with --restart. --no-review or --yes skips human review. --timeout 45m bounds the full invocation, including Harbor validation and simulation; each discovery resume has a new invocation budget, so repeat --timeout on --resume to keep an override or set RELAI_TIMEOUT. When no risk is specified, discovery tops up the eligible risk pool before starting candidate generation. That curation is quality-gated: it may add many supported hypotheses, but it should stop early instead of inventing filler records. Discovery requires a completed EvalsWiki bootstrap. relai init performs that bootstrap by default; projects initialized with relai init --no-wiki must run relai wiki bootstrap first. Discovery roles can read permitted project source and EvalsWiki material, but cannot read credentials, RELAI runtime state, staged drafts, dependency caches, or run commands. Each candidate is built in an isolated Harbor draft, validated and simulated by the host, and reviewed after simulation as either failure-revealing or coverage-only. Failure-revealing acceptance requires the exact simulated draft to have at least one task-local evaluator score below its declared maximum. A valid, useful agent test that the current agent passes may be accepted as coverage, but it is counted separately and does not confirm the risk as triggered. Evaluator execution errors pause the candidate rather than being treated as agent failures. Risk hypotheses are saved as typed EvalsWiki records with Markdown pages and reciprocal agent test lineage. They are ranked deterministically by impact, confidence, and triggerability. Rejected means simulation or review evidence refuted the hypothesis; Retired means it is no longer worth pursuing. Before accepting a Harbor draft, RELAI checks the task-owned JSON assets named by its descriptor. Malformed authored input, persona, component-adapter, generator, or mock assets return through the current validation and repair flow. Final scenario requirements produced by image-local setup are checked at runtime; Harbor validates the assets again in the task image because Dockerfile inputs can differ from the host package. Interactive discovery shows the proposed agent test before the expensive build/run step. Choose yes to build, no to skip the candidate without retiring the risk, or revise to return the spec to proposal with feedback. Without a terminal, review is skipped unless --agent-interaction is set. Discovery resolves --agent-target with the same rules as create: a sole configured target is selected automatically, while multiple configured targets require --agent-target outside an interactive terminal. Accepted artifacts retain package, result, and transcript digests in their discovery lineage; no transcript copy is published. With --agent-interaction, discovery writes one relai.agent_interaction.v1 document to stdout per invocation and keeps human output on stderr. Proposal review becomes a clarification with choices build, build_all (approve the rest of this run), and skip; any free-form answer is revision feedback. Model failures, builder questions and approvals, and timeouts pause with relai test discover --resume --agent-interaction. Builder questions, approvals, and builder timeouts come from the shared agent-test builder, so their documents carry its learning_env_create workflow and session ID; answer them with that session ID. Declining a builder request retires that candidate and cancels the invocation (exit 130); --resume continues with the next candidate. A run that ends, including one short of its --count with nothing left to resume, returns a completed summary of created tests, evaluator scores, each created test’s simulation trajectory, and next steps. The terminal summary separates accepted tests into:
  • Failures revealed: new tests where the current agent failed.
  • Coverage expanded: new tests the current agent already passes, but that are useful for preventing future regressions.

create

Create an agent test from a prompt, from a run log plus feedback, or from a grounded EvalsWiki requirement.
Agent test generation can take a while when the prompt, run log, or agent context is complex.
The run log does not need a specific structure: any repository-local text format under the size limit works, including JSONL message logs, plain transcripts, or your own trace format. RELAI parses it. Logs that include internal detail — tool calls, retrievals, intermediate reasoning — give RELAI more to work with than user-facing output alone. The CLI agent inspects the repository, asks a focused question only when needed, reports its runtime plan, generates a candidate, and validates both the candidate and simulator before saving the final file under .relai/harbor/agent-tests/. The completion summary names the validation run’s ATIF trajectory (Validation trajectory), so you can see what the target agent did. Clarification continues in the same process. The agent may repair .relai/.internal/simulator as part of creation when .relai/.internal/layout.json is present; legacy projects without the marker use .relai/simulator. It cannot modify application source. External setup work is explained and saved for --resume. If a model stops before validation, RELAI checks the staged artifact once and uses a failed check as repair feedback. Missing external prerequisites remain paused for --resume. create, update, suggestion generate, and suggestion accept accept --permissions {auto|accept_edits|ask}. See Permissions. Creation reuses the project-specific component, input, log, and mock guidance recorded by relai init. For known frameworks it also has bundled RELAI guidance. For custom frameworks or version-specific behavior, it can inspect repository source and bounded public declarations or source from the exact installed dependency. Current project evidence takes precedence over bundled framework hints. create saves the agent test under .relai/harbor/agent-tests/. Commit that package, together with .relai/harbor/runtime/, to share it with collaborators; RELAI does not publish it to the backend. An agent test reuses the project runtime that relai init validated: the same base image, locked dependencies, adapter, and launcher. The test holds only its scenario, mocks, evaluators, and test-specific setup, and it cannot change the runtime. If the environment itself needs a fix, rerun relai init. Each criterion the test scores is recorded as exact or semantic. Exact criteria, such as a tool call, its arguments, or a required string, can be scored by code. Semantic criteria, such as refusing an off-topic request, tone, or an answer’s meaning, are always scored by an LLM judge, because no keyword or pattern check covers every valid phrasing. RELAI scores judges through its model proxy with your relai setup login; they need no model-provider keys. The same holds for a test’s simulated persona turns and generated component inputs. Agent tests are scoped to the selected agent target. When multiple targets are registered, interactive runs present a single-target picker; pass --agent-target {target} in non-interactive runs. Common failures: log outside the repository, missing paired --feedback, validation failure, duplicate explicit name, or required external setup. Use --resume after completing recoverable setup.

suggestion

Generate suggestions, then list, review, accept, or reject them:
Suggestion runs and decisions are ignored run state scoped to the selected agent target. accept uses the standard agent test creation workflow, and clear-memory --yes removes the selected target’s stored suggestion history. These commands do not register an agent or persist suggestions in the backend.