OpenEnv documentation
Evaluate Claude Code in an Environment
Evaluate Claude Code in an Environment
This tutorial measures how well an existing agent, Claude Code, does a job that only your environment’s tools can do, over a conversation with a simulated user. The job comes from τ²-bench: Claude Code is an airline’s customer service agent, it has to follow the airline’s policy, and τ²-bench checks the database it leaves behind.
It uses HarnessEnvironment, the agentic harness wrapper from RFC 005. Claude Code keeps its own agent loop and runs inside the environment, the environment serves its tools to it over MCP, and each step() is one message from the customer:
step(customer message)
|
v
Claude Code --tool calls over MCP--> MCP bridge --> environment tools --> τ²-bench database
|
| reply
v
simulated customer --> next message, or done --> rubric scores the database and the conversationThe full code is in examples/claude_code_harness_eval.
Use this path when the job is a conversation and the agent has to use your environment’s tools. RFC 005 does not capture token ids, so it evaluates and serves, but does not train. To train the model behind an agent, or to evaluate it on tasks that a verifier checks, use Harbor. Harnesses in OpenEnv compares the options.
What You’ll Build
- An
AgenticHarnessAdapterfor Claude Code’s headless mode. - A
HarnessEnvironmentthat gives Claude Code the tools and policy of a τ²-bench task, plays the task’s simulated customer between turns, and scores the conversation with a rubric. - The RFC 005 loop around it: one
step()per customer message, until the customer is done. - The same harness served in production mode, with you as the customer.
Install Dependencies
You need the claude CLI, logged in or with ANTHROPIC_API_KEY set, and Python 3.12 or later, which τ²-bench requires. Then, from a clone of OpenEnv:
pip install -e . -e envs/tau2_env
τ²-bench’s tasks and databases are not in its package. Fetch them once and point TAU2_DATA_DIR at them, as described in the τ²-bench environment’s docs. The simulated customer runs on Inference Providers, so set HF_TOKEN too.
The Adapter
An AgenticHarnessAdapter tells HarnessEnvironment how to start a harness, how to hand it tools, and how to read a turn. claude_code_adapter.py keeps one long-lived claude -p --input-format stream-json --output-format stream-json process per episode:
AgenticHarnessAdapter | Claude Code |
|---|---|
inject_tools(tools, bridge_url) | --mcp-config with the bridge as an HTTP MCP server, and --allowedTools mcp__env so its tools run without a permission prompt |
start() | spawns the process with HarnessProcess. Claude Code prints nothing until the first message |
send_message_streaming() | writes {"type": "user", "message": {...}} on stdin and maps stdout events: tool_use → TOOL_CALL, tool_result → TOOL_RESULT, text → TEXT_OUTPUT, and result (end of turn) → TURN_COMPLETE |
Claude Code’s own tools are turned off (--tools ""), so the agent can only act through the environment’s tools: it has no shell, no file access and no way to open connections of its own. It also loads no settings files (--setting-sources ""), so your hooks and CLAUDE.md stay out of the run. To give it shell or file tools as well, list them in --tools and run it in a sandbox: the container keeps it away from your machine, and an egress allowlist (for example ACASandboxProvider.deny_all_egress()) keeps it to the model’s API.
Join It to τ²-bench
tau2_harness.py defines Tau2Harness, a HarnessEnvironment for one τ²-bench task:
- The tools and the policy. Claude Code gets the domain’s tools over MCP, and the domain’s policy appended to its system prompt. It does not get
respond_to_userordone, because it talks to the customer through its replies. - The customer. After each Claude Code turn, the environment passes the reply to τ²-bench’s simulated customer and returns the customer’s next message in
observation.metadata["customer"]. The customer is part of the environment, as intau2_envitself. - The reward. A rubric reads τ²-bench’s score when the customer ends the conversation, so the reward is in
observation.rewardandobservation.donemarks the end.
import tempfile
from openenv.core.harness import HarnessAction, HarnessConfig
from tau2_env.server.tau2_environment import Tau2Environment
from tau2_harness import Tau2Harness
config = HarnessConfig(
name="claude-code",
command=["claude"],
working_directory=tempfile.mkdtemp(),
model="haiku",
)
harness = Tau2Harness(Tau2Environment(domain="airline", split="test"), "2", config)
observation = harness.reset() # the customer's opening message
while not observation.done:
# One Claude Code turn, then the customer's answer
observation = harness.step(HarnessAction(message=observation.metadata["customer"]))
print(observation.reward)
harness.close() # stops Claude Code and the τ²-bench taskEach customer message is one step(HarnessAction(message=...)), and Claude Code keeps the conversation’s context across steps. Its tool calls run against the same database the customer sees, so τ²-bench can score the result. If Claude Code exits or runs out of time, the step comes back with done and the error in its metadata.
Run the Evaluation
PYTHONPATH=src:envs:examples/claude_code_harness_eval \
python examples/claude_code_harness_eval/run_eval.py --domain airline --task-ids 8 16 19 26 --model haikuThese four tasks each end in a different state of the database. In task 8 the agent books a flight, in 16 it changes one, in 19 it cancels one, and in 26 the policy doesn’t allow the cancellation the customer asks for, so the agent has to refuse, even when the customer pushes back. In task 26 the output looks like this:
=== airline task 26
customer: Hi, I need to cancel my flights from Orlando to Charlotte. I'd like to get a refund for them, please.
[...]
-> get_user_details({'user_id': 'amelia_sanchez_4739'})
-> get_reservation_details({'reservation_id': '3FRNFB'})
-> get_reservation_details({'reservation_id': 'Q4L9HS'})
[...]
customer: It's a change of plans. I still want to cancel and get a refund, please.
agent: I'm sorry, but I can't cancel reservation 3FRNFB. It was booked on May 6, more than 24 hours ago. [...]
customer: I really need to cancel and get a refund. If you can't do it, please transfer me to someone who can.
agent: I understand this is frustrating, but I can't transfer you for this request. [...]
reward: 1.00 {'DB': 1.0, 'COMMUNICATE': 1.0}
PASS task 8
PASS task 16
PASS task 19
PASS task 26
pass^1: 4/4τ²-bench scores a task 1.0 when the database ends up as the task expects (DB) and the agent told the customer what it had to (COMMUNICATE). pass^1 is the fraction of tasks that score 1.0 in a single attempt. τ²-bench’s pass^k asks for k successes out of k attempts of the same task, so running each task several times measures how consistent the agent is.
--domain takes any τ²-bench domain, --task-ids picks tasks, and --user-model changes the simulated customer’s model. Each task uses a fresh tau2_env and a fresh Claude Code process.
Serve It in Production Mode
The same harness can be served over the WS /harness route (production mode), so a person, or another application, talks to Claude Code with the environment’s tools behind it. Use it to try the agent by hand, debug a task, or put the agent in front of real users once the evaluation looks good. The server takes one connection at a time, each with a fresh Claude Code process and a fresh copy of the task’s database. /harness sends your messages straight to Claude Code, without the environment’s turn logic, so you play the customer and nothing is scored. The server has no authentication, and every connection spends your credentials, so keep it on localhost or behind your own auth:
PYTHONPATH=src:envs:examples/claude_code_harness_eval \
python examples/claude_code_harness_eval/serve.py --domain airline --task-id 2 --port 8000
python examples/claude_code_harness_eval/chat.py ws://localhost:8000/harness \
"Hi, I'm Noah Muller, user id noah_muller_9847. What reservations do I have?" \
"Which of them has a delayed flight?"Serve It in a Container
RFC 005 runs the harness inside the environment’s container, apart from the machine that serves it. examples/claude_code_harness_eval/Dockerfile does that for this recipe: Claude Code, tau2_env and serve.py run together in one image, as a non-root user. Build it from the repository root and pass the credentials at run time:
docker build -t claude-code-tau2 -f examples/claude_code_harness_eval/Dockerfile .
docker run -p 127.0.0.1:8000:8000 -e HF_TOKEN -e ANTHROPIC_API_KEY claude-code-tau2
python examples/claude_code_harness_eval/chat.py ws://localhost:8000/harness \
"Hi, I'm Noah Muller, user id noah_muller_9847. What reservations do I have?"Use -e CLAUDE_CODE_OAUTH_TOKEN instead of ANTHROPIC_API_KEY to run on a Claude subscription (claude setup-token creates the token). Add serve.py arguments after the image name to pick the domain, task or model.
Things to Know
- Claude Code adds today’s date to its context, while each τ²-bench policy states its own current time (airline is 2024-05-15). The agent prompt tells it to use the policy’s time. Without that line, it books flights in the wrong year.
- When the Anthropic API fails a request (for example
API Error: Connection dropped (ECONNRESET)), the adapter ends the turn as a harness failure, so the step comes back withdoneanderror_typeharness_crashed, andrun_eval.pyreports the task asERRORand leaves it out of pass^1. Rerun it. The same goes for the simulated customer: if its model fails, the step ends witherror_typecustomer_failed, and the task is anERRORrather than a failure of Claude Code.
Adapting It
- Your own environment: pass your FastMCP tools as
mcp, your instructions as the adapter’ssystem_prompt, and a rubric toHarnessEnvironmentto score the episode. - Another model:
--modelaccepts any model Claude Code does.ANTHROPIC_BASE_URLinHarnessConfig.env_varspoints it at an Anthropic-compatible endpoint. - Another harness: write another
AgenticHarnessAdapterin the same shape. Pick a headless multi-turn mode, passbridge_urlto its MCP config, and map its events.