Skip to main content
EvalsWiki keeps what the agent should do, where it may fail, and what remains unclear as readable Markdown in your repository.
EvalsWiki is not the test suite.It holds durable knowledge. Agent tests turn that knowledge into executable behavior checks stored in Harbor format, while simulations produce separate scores, feedback, and trajectories.

What belongs in EvalsWiki

Requirements

Clear, independently changeable obligations that can ground evaluation criteria.

Risk hypotheses

Falsifiable ideas about where the agent may fail. A risk is not proof of a failure.

Open questions

Ambiguities that should not become scored requirements until someone resolves them.

Inspect the current knowledge

Start read-only. The list commands provide a curated view; the underlying Markdown remains human-readable and reviewable in your repository.
Paste this in your coding agent at the repository root.

Use RELAI to show me the current EvalsWiki requirements, risks, and open questions. Summarize them without changing anything.

Capture a durable decision from the conversation

When a conversation establishes a lasting requirement, correction, constraint, or preference, ask the coding agent to checkpoint only that evidence. RELAI retrieves related knowledge, deduplicates it, validates the update, and reports either a focused change or no_change. Do not use this for temporary task chatter.

Use RELAI to update EvalsWiki with this durable project decision: “…”. Include only the relevant conversation evidence, preserve unrelated knowledge, and tell me whether RELAI updated anything. Do not include credentials, secrets, or unrelated transcript content.

The coding-agent integration sends a bounded evidence envelope to relai wiki update --agent-interaction --evidence-json - through standard input. Let the integration construct that envelope; do not place private transcript content or secret values in a command argument.

Resolve an open question

Give RELAI the explicit decision. It converts the answer into reviewed durable knowledge and removes the ambiguity.

Use RELAI to resolve this EvalsWiki open question with my answer: “…” Show the proposed knowledge change before applying it.

Keep linked coverage current

Stable knowledge IDs connect requirements to tests, evaluators, and benchmarks. After a requirement changes, reconcile its linked artifacts; RELAI asks before reducing or retiring coverage.

Use RELAI to reconcile the tests, benchmarks, and evaluators linked to my changed requirement. Ask before reducing coverage.

What RELAI stores

Prefer RELAI’s Wiki workflows over direct Markdown edits so stable IDs, conflicts, and grounding links remain correct.

Turn durable intent into coverage

Create a repeatable agent test from a requirement, a risk, or a failed run.

Command reference

See relai wiki for all flags, defaults, and recovery details.