Skip to main content
Use relai evaluator to create or remove local global evaluators.

create

Create a global evaluator from a prompt.
Creates an evaluator under .relai/harbor/evaluators/ as a Harbor package. Commit the package to share it; release builds do not publish evaluator content. Evaluators are scoped to the selected agent target. When multiple targets are registered, interactive runs present a single-target picker; pass --agent-target {target} in non-interactive runs. Creation is local and agentic: RELAI starts from a language-appropriate candidate, records a structured evaluator plan, and iterates through local deterministic validation and repairs before saving. --resume continues an unfinished session and --restart discards it and starts again. The backend is not used to generate or repair evaluator source. The plan records each criterion as exact or semantic. An evaluator that scores any semantic criterion, such as tone, a refusal, or an answer’s meaning, is an LLM judge; a code evaluator covers only exact criteria, such as a tool call or a required string. RELAI scores judges through its model proxy with your relai setup login; they need no model-provider keys. An update keeps the evaluator’s type unless its criteria now need a judge. If a model stops before validation, RELAI runs the evaluator check once and uses a failed result as repair feedback in the same bounded session. Common failures: a required clarification or approval, evaluator validation failure, or duplicate explicit name. create and update accept --permissions {auto|accept_edits|ask} and --timeout {DURATION|off} (default 1 hour). See Permissions and Timeouts.

How global evaluators are applied

Evaluators created here are global: they run in every simulation of the agent whose scope they cover, alongside the local evaluators an agent test or benchmark defines for itself. Because optimization evaluates through the simulator, both kinds also score every optimizer run — see Evaluation during optimization. There is no per-run subset selection: all global evaluators with a matching scope apply whenever you simulate or optimize. When an evaluator should stop applying, remove it. Scope is inferred when the evaluator is created, and evaluators run only for simulations with the same scope. Most evaluators target end-to-end agent behavior and run for every end-to-end simulation. An evaluator whose prompt targets one component — for example, scoring the relevance of retrieved documents — runs only for simulations of that component, such as component-targeted agent tests; it does not run in end-to-end simulations, even when the end-to-end run exercises that component. When the intended scope is ambiguous, RELAI asks during creation; state it in the prompt to avoid the question.

remove

Remove a global evaluator by id.
This is a destructive local operation. Confirm the evaluator id before running it.