> ## Documentation Index
> Fetch the complete documentation index at: https://cli-docs.relai.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# EvalsWiki

> Use RELAI EvalsWiki to keep durable agent requirements, risk hypotheses, and open questions grounded in your repository.

EvalsWiki keeps what the agent should do, where it may fail, and what remains unclear as readable Markdown in your repository.

| Object | Question it answers |
| - | - |
| EvalsWiki | What should be true? |
| Agent test | Can the agent do it? |
| Run evidence | What happened? |

<Note>
  **EvalsWiki is not the test suite.**

  It holds durable knowledge. Agent tests turn that knowledge into executable behavior checks stored in Harbor format, while simulations produce separate scores, feedback, and trajectories.
</Note>

## What belongs in EvalsWiki

<CardGroup cols={3}>
  <Card title="Requirements">
    Clear, independently changeable obligations that can ground evaluation criteria.
  </Card>

  <Card title="Risk hypotheses">
    Falsifiable ideas about where the agent may fail. A risk is not proof of a failure.
  </Card>

  <Card title="Open questions">
    Ambiguities that should not become scored requirements until someone resolves them.
  </Card>
</CardGroup>

## Inspect the current knowledge

Start read-only. The list commands provide a curated view; the underlying Markdown remains human-readable and reviewable in your repository.

<Tabs>
  <Tab title="Ask your coding agent">
    Paste this in your coding agent at the repository root.

    <Prompt description="Use RELAI to show me the current EvalsWiki requirements, risks, and open questions. Summarize them without changing anything." actions={["copy"]}>
      Use RELAI to show me the current EvalsWiki requirements, risks, and open questions. Summarize them without changing anything.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai wiki list requirements
    relai wiki list risks
    relai wiki list open-questions
    ```
  </Tab>
</Tabs>

## Capture a durable decision from the conversation

When a conversation establishes a lasting requirement, correction, constraint, or preference, ask the coding agent to checkpoint only that evidence. RELAI retrieves related knowledge, deduplicates it, validates the update, and reports either a focused change or `no_change`. Do not use this for temporary task chatter.

<Prompt description="Use RELAI to update EvalsWiki with this durable project decision: “…”. Include only the relevant conversation evidence, preserve unrelated knowledge, and tell me whether RELAI updated anything. Do not include credentials, secrets, or unrelated transcript content." actions={["copy"]}>
  Use RELAI to update EvalsWiki with this durable project decision: “…”. Include only the relevant conversation evidence, preserve unrelated knowledge, and tell me whether RELAI updated anything. Do not include credentials, secrets, or unrelated transcript content.
</Prompt>

The coding-agent integration sends a bounded evidence envelope to `relai wiki update --agent-interaction --evidence-json -` through standard input. Let the integration construct that envelope; do not place private transcript content or secret values in a command argument.

## Resolve an open question

Give RELAI the explicit decision. It converts the answer into reviewed durable knowledge and removes the ambiguity.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to resolve this EvalsWiki open question with my answer: “…” Show the proposed knowledge change before applying it." actions={["copy"]}>
      Use RELAI to resolve this EvalsWiki open question with my answer: “…” Show the proposed knowledge change before applying it.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai wiki resolve open-question <question-id> --answer "<explicit answer>"
    ```
  </Tab>
</Tabs>

## Keep linked coverage current

Stable knowledge IDs connect requirements to tests, evaluators, and benchmarks. After a requirement changes, reconcile its linked artifacts; RELAI asks before reducing or retiring coverage.

<Tabs>
  <Tab title="Ask your coding agent">
    <Prompt description="Use RELAI to reconcile the tests, benchmarks, and evaluators linked to my changed requirement. Ask before reducing coverage." actions={["copy"]}>
      Use RELAI to reconcile the tests, benchmarks, and evaluators linked to my changed requirement. Ask before reducing coverage.
    </Prompt>
  </Tab>

  <Tab title="Run in terminal">
    ```sh theme={"system"}
    relai wiki reconcile
    ```
  </Tab>
</Tabs>

## What RELAI stores

| Path | Purpose | Share? |
| - | - | - |
| `.relai/evals-wiki/` | Reviewed Markdown knowledge | Version with the repository |
| `.relai/evals-wiki/attachments/` | Selected supporting evidence | Review for sensitive data first |
| `.relai/.internal/` | Sessions and managed run artifacts | Keep local |
| `.relai/simulator.env` | Local provider and runtime values | Never commit |

Prefer RELAI’s Wiki workflows over direct Markdown edits so stable IDs, conflicts, and grounding links remain correct.

<Card title="Turn durable intent into coverage" href="/agent-tests">
  Create a repeatable agent test from a requirement, a risk, or a failed run.
</Card>

## Command reference

See [`relai wiki`](/cli/wiki) for all flags, defaults, and recovery details.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.