Skip to main content
Clone a small airline customer-support agent, run it locally, and use short coding-agent prompts to create tests, simulate behavior, and optimize the agent.

Before you start

Every sample needs:
  • Git, Bash, uv, and Docker with Docker Compose
  • A running Docker daemon that can run Linux containers
  • Either OPENAI_API_KEY or ANTHROPIC_API_KEY for the sample agent and simulations
  • RELAI installed, or permission to install it during the machine-setup step below
Plan for a few minutes per stepInitialization, test creation, simulation, and optimization can each take several minutes. Run the prompts one at a time so you can review each result.

Choose a sample

Python sample

Python 3.11+. Open the repository on GitHub.

TypeScript sample

Node.js 20+. Open the repository on GitHub.

Go sample

Go 1.24+. Open the repository on GitHub.

Clone the sample

Choose one runtime, then open the cloned repository in your coding agent.

Follow https://cli-docs.relai.ai/sample-agent to help me choose and clone one RELAI airline-support sample. Check its prerequisites, but stop before setup or initialization. Never ask me to paste a provider key into chat.

Set up RELAI once on this machine

If this machine is already signed in to RELAI and the integration for your coding agent is active, skip this section. Setup is machine-level; it is not repeated for every repository.

Follow https://cli-docs.relai.ai/installation to install and set up RELAI for this coding agent. Check prerequisites, guide me through OAuth sign-in, and stop before initializing the repository. Keep provider keys out of chat.

After setup installs the integration, restart Codex or Cursor, run /reload-plugins in Claude Code, or start a new interactive Copilot session.

Run one learning loop

Open your coding agent in the sample repository. Initialization creates the project runtime; the risk-to-test-to-improvement loop is what you repeat. Paste one prompt, wait for it to finish, and review the result before continuing.
1

Initialize this repository

Use RELAI to initialize this sample. Keep provider keys out of chat.

If a provider key is needed, add the requested KEY=value directly to .relai/simulator.env.
2

Create a high-impact agent test

Use RELAI to identify the highest-impact potential failure in this sample agent and build one agent test for it.

RELAI grounds the behavior in EvalsWiki, checks existing coverage, and reports the created test.
3

Measure the current behavior

Use RELAI to simulate the agent test you just created and show me the readable result.

The result should separate evaluator scores and errors, and include any available conversation, tool evidence, and earliest supported failure.
4

Optimize and review

Use RELAI Agent Optimizer on that test in Balanced quality/cost mode. Keep changes local and show before-and-after results.

Review the scope and rollout budget before provider-backed work begins.
Use the prompts above so the plugin can surface approvals and summarize structured results.
--no-pr keeps accepted changes on a local relai/optimizer/… branch, so this sample does not require GitHub authentication. On your own repository, an authenticated GitHub CLI can publish the branch and open a pull request.

Optional: challenge the agent further

After the first loop, choose the prompt that matches your next goal. These are optional paths, not required steps.

Use RELAI to make that test harder with a multi-turn scenario. Reuse or update existing coverage instead of duplicating it.

Use RELAI to create a five-turn adversarial test for the most important uncovered risk. Increase the pressure realistically on each turn.

Use RELAI Agent Optimizer on the failing tests in Balanced quality/cost mode. Keep changes local and compare before-and-after results.

What this walkthrough creates

EvalsWiki

Grounded requirements, risk hypotheses, coverage, and open questions for the sample agent.

Agent tests

Repeatable single- and multi-turn scenarios with the evaluators RELAI creates for their requirements.

Optimizer result

A before-and-after report and, when a candidate improves the overall evaluation, an optimized agent on a local branch.

Ready for your own agent?

Run the same learning loop in a compatible agent repository.