Agent Bridge
Overview
While Inspect provides facilities for native agent development, you can also very easily integrate agents created with 3rd party frameworks like OpenAI Agents SDK, Pydantic AI, and LangChain, or use fully custom agents you have developed or ported from a research paper. You can also use CLI based agents that run within sandboxes (e.g. Claude Code, Codex CLI, or Gemini CLI).
Agents are bridged into Inspect such that their native model calling functions are routed through the current Inspect model provider. There are two types of agent bridges supported:
Bridging to Python-based agents that run in the same process as Inspect via the agent_bridge() context manager.
Bridging to agents that run in a sandbox via the sandbox_agent_bridge() context manager (these agents can be written in any language).
We’ll cover each of these configurations in turn below. You can also learn from the following examples:
| OpenAI Agents SDK | Demonstrates using a native Open AI Agents SDK agent to perform Q/A using web search. |
| LangChain | Demonstrates using a native LangChain agent to perform Q/A using the Tavili Search API |
| Pydantic AI | Demonstrates using a native Pydantic AI agent to perform Q/A using web search. |
| Claude Code | Demonstrates using a Claude Code agent to explore a Kali Linux system. |
| Codex CLI | Demonstrates using a Codex CLI agent to explore a Kali Linux system. |
Agent Bridge
The agent_bridge() can bridge agents written against the Python APIs for OpenAI Completions, OpenAI Responses, Anthropic, and Google. To bridge a Python based agent running in the same process as Inspect:
Write your custom Python agent as normal using the OpenAI, Anthropic, or Google connector provided by your agent system, specifying “inspect” as the model name.
Run your custom Python agent within the agent_bridge() context manager which redirects OpenAI calls to the current Inspect model provider.
For example, here we build an agent that uses the OpenAI SDK directly (imaging using your favourite agent framework in its place):
from openai import AsyncOpenAI
from inspect_ai.agent import (
Agent, AgentState, agent, agent_bridge
)
from inspect_ai.model import messages_to_openai
@agent
def my_agent() -> Agent:
async def execute(state: AgentState) -> AgentState:
async with agent_bridge(state) as bridge:
client = AsyncOpenAI()
await client.chat.completions.create(
model="inspect",
messages=messages_to_openai(state.messages)
)
return bridge.state
return execute- 1
-
Use the agent_bridge() context manager to redirect the OpenAI API to the Inspect model provider. Pass the
stateso that the bridge can automatically keep track of changes tomessagesandoutputbased on model calls passing through the bridge. - 2
-
Use the OpenAI API with
model="inspect", which enables Inspect to intercept the request and send it to the Inspect model being evaluated for the task. - 3
-
Convert the
state.messagesinput into native OpenAI messages using the messages_to_openai() function. - 4
-
Return the
statechanges automatically tracked by thebridge.
The OpenAI Agents SDK, PydanticAI LangChain example provide a more in-depth demonstration of using the Python agent bridge with Inspect.
Sandbox Bridge
The sandbox_agent_bridge() can bridge agents written against the OpenAI Completions, OpenAI Responses, Anthropic API, or Google API. To bridge an agent running in a sandbox to Inspect:
Configure your sandbox (e.g. via its Dockerfile) to contain the agent that you want to run. The agent should be configured to talk to the OpenAI, Anthropic, or Gemini API on localhost port 13131 (e.g.
OPENAI_BASE_URL=http://localhost:13131/v1,ANTHROPIC_BASE_URL=http://localhost:13131, orGOOGLE_GEMINI_BASE_URL=http://localhost:13131/v1beta).Write a standard Inspect agent that uses the sandbox_agent_bridge() context manager and the
sandbox().exec()method to invoke the custom agent.
The sandbox bridge works via running a proxy server inside the sandbox container which receives requests for the OpenAI, Anthropic, and Google APIs. This proxy server in turn relays requests to the current Inspect model provider.
The proxy listens on the sandbox’s loopback interface and is reachable only by processes running inside the sandbox, such as the agent binary launched with sandbox().exec(). It is not a browser-facing service: it sends no cross-origin headers, does not answer preflight requests, and accepts only application/json request bodies. A web page rendered inside the sandbox therefore cannot reach the model or bridged tools through the proxy: a request that needs a preflight is refused by the browser, and a request sent without one carries a body type the proxy rejects before any handler runs.
For example, here we build an agent that runs a custom agent binary (passing it input on the command line and reading output from stdout):
from openai import AsyncOpenAI
from inspect_ai.agent import (
Agent, AgentState, agent, sandbox_agent_bridge
)
from inspect_ai.model import user_prompt
from inspect_ai.util import sandbox
@agent
def my_agent() -> Agent:
async def execute(state: AgentState) -> AgentState:
async with sandbox_agent_bridge(state) as bridge:
prompt = user_prompt(state.messages)
result = sandbox().exec(
cmd=[
"/opt/my_agent",
"--prompt",
prompt.text
],
env={"OPENAI_BASE_URL": f"http://localhost:{bridge.port}/v1"}
)
if not result.success:
raise RuntimeError(f"Agent error: {result.stderr}")
return bridge.state
return execute- 1
-
Use the sandbox_agent_bridge() context manager to redirect the OpenAI API to the Inspect model provider. Pass the
stateso that the bridge can automatically keep track of changes tomessagesandoutputbased on model calls passing through the bridge. - 2
- Extract the last user message from the message history with user_prompt().
- 3
- Run the agent, using a CLI argument for input and stdout for output (other agents may use more sophisticated encoding schemes for messages in and out).
- 4
-
Redirect the OpenAI API to talk to a proxy server that communicates back to the current Inspect model provider. Note that we read the
portto listen on from thebridgeyielded by the context manager. - 5
-
Return the
statechanges automatically tracked by thebridge.
The Claude Code and Codex CLI agents in the Inspect SWE package provide more in-depth demonstrations of running custom agents in sandboxes.
Granted Capabilities
Some tools a bridged agent can request are executed outside the sandbox — web_search, web_fetch, code_execution, and remote MCP servers all run at the model provider. An agent that names one of these in a request therefore reaches the network regardless of the sandbox’s own egress policy, so sandbox_agent_bridge() withholds them unless the evaluation grants them:
async with sandbox_agent_bridge(state, web_search=True) as bridge:
...Pass True to use the target model’s internal provider, a WebSearchProviders configuration to select providers explicitly, or leave the default to withhold. code_execution works the same way, and client_mcp_servers=True honors MCP servers named by the agent (prefer Bridged Tools for servers you choose yourself). Whenever a declared tool is withheld a warning is logged, since an absent tool is otherwise indistinguishable from the model declining to call it.
Ordinary function tools are never withheld — those are executed by the bridged agent inside the sandbox and grant it nothing it does not already have.
For the same reason, all media content (images, audio, video, and documents) in a request to a sandbox bridge must be an inline data: URI. A remote URL or host path would otherwise be dereferenced outside the sandbox, on the agent’s behalf.
None of this restricts the in-process agent_bridge(), where the scaffold already runs with the host’s network and filesystem and there is no boundary to defend. Inspect explicitly materializes its non-inline media before model submission, preserving that access without relying on provider-side fetching.
AgentBridge itself defaults to rejecting non-inline media so a new bridge subclass cannot accidentally grant host access. Code that constructs the class directly can opt in with allow_remote_media=True; the agent_bridge() factory makes this grant explicitly.
Bridged Tools
Host-side Inspect tools can be exposed as MCP tools to sandboxed agents using the bridged_tools parameter. This is useful when you have Inspect tools that need to run on the host (e.g. tools that access host resources, databases, or APIs) but want them available to agents running in a sandbox.
To bridge tools, wrap them in a BridgedToolsSpec and pass to sandbox_agent_bridge():
from inspect_ai.tool import tool
from inspect_ai.agent import (
Agent, AgentState, agent, sandbox_agent_bridge, BridgedToolsSpec
)
from inspect_ai.util import sandbox
@tool
def search_database():
async def execute(query: str) -> str:
"""Search the internal database.
Args:
query: The search query.
"""
# Runs on the host, not the sandbox
return f"Results for: {query}"
return execute
@agent
def my_agent() -> Agent:
async def execute(state: AgentState) -> AgentState:
async with sandbox_agent_bridge(
state,
bridged_tools=[
BridgedToolsSpec(
name="host_tools",
tools=[search_database()]
)
]
) as bridge:
# bridge.mcp_server_configs contains resolved MCPServerConfigStdio
# objects that can be passed to CLI agents
return bridge.state
return executeThe bridge handles:
- Starting a host-side service that executes the Inspect tools
- Writing an MCP server script to the sandbox that forwards tool calls to the host
- Returning MCPServerConfigStdio configs that CLI agents can use to connect
Bridged tool results are truncated in the same way as other tool results: at the tool’s own max_output if it declares one, otherwise at max_tool_output from the generate config (16KB by default). Results returned as content (e.g. list[ContentText]) are not truncated, as for other tools.
Execution Contract
Bridged tools execute only for calls the model proposed in a bridged generation, once per proposal. Each tool call in a model response handed to the agent grants one execution of exactly that tool with exactly those arguments (as approved, or as modified by an approval policy); a tools/call request with no matching grant is denied and the error is returned to the agent, which surfaces it to the model. This holds whether or not an approval policy is active. Tool discovery (tools/list) is unaffected.
Two consequences follow:
- An agent that retries a
tools/callafter a transport failure is denied, since the proposal has already executed once. The agent should return the error to the model, which re-proposes the call. - Only the primary choice of a multi-choice response is granted; alternate choices that carry tool calls are dropped (with a warning). Request a single choice (
n=1) from a bridged agent.
An agent that legitimately calls a host tool programmatically, outside a model turn, can opt a server out with require_proposal=False. That server’s tools then execute for any tools/call it receives, and the bridge no longer guarantees that they run only for calls the model made:
BridgedToolsSpec(
name="host_tools",
tools=[search_database()],
require_proposal=False,
)Matching a Proposal
Agents rename MCP tools to their models under their own schemes (Claude Code declares mcp__<server>__<tool>, Codex CLI the bare name inside a Responses API namespace, Gemini CLI mcp_<server>_<tool>, OpenCode <server>_<tool>, some cut or hash long names), so the bridge does not match a proposed call by name. It matches the call’s declaration in the agent’s request to a bridged tool by the description the bridge itself served in tools/list. A declaration whose description matches no bridged tool denotes nothing and the call is denied. Tools an agent discovers through the Responses API’s tool_search (declared to the model inside a tool_search_output item rather than the request’s tool list) count as declarations too. Two or more bridged tools the served content cannot tell apart are each granted by a proposal for any of them: one proposal then authorizes one execution of each, still bound to the proposed arguments. An empty description is matched like any other, so an undocumented bridged tool is indistinguishable from every other undocumented tool. The bridge warns at setup, naming the tools that share a description (an empty one included); give bridged tools distinct docstrings to keep matching one-to-one. The supported agents forward the served description to their models unchanged. An agent that truncates long descriptions is tolerated: a declaration that agrees with a served description for at least 64 characters and adds at most a short trailing marker identifies that tool. An exact description match wins even when that description is a prefix of another bridged tool’s; a truncated declaration that could refer to both grants both. An agent that rewrote the leading text would be denied on every call. An agent-local tool declared with a bridged tool’s exact description is indistinguishable from it and would grant it.
One other shape is recognized: an agent that exposes every MCP tool through a single dispatcher function, Antigravity’s call_mcp_tool, whose arguments carry string ServerName and ToolName values naming a bridged tool and an object Arguments. Such a call denotes that tool with those nested arguments. The dispatcher is recognized by that function name together with the argument shape, so an ordinary tool call cannot denote a bridged tool by naming it in its arguments; a dispatcher with another name or other parameter names is not supported and its calls are denied.
Models
As demonstrated above, an agent reaches the eval’s model by naming the model "inspect".
With the in-process agent_bridge(), you can reach other Inspect models by prefacing the fully qualified model name with inspect/. For example, in a LangChain agent, you would do this to use the Inspect interface to Gemini:
model = ChatOpenAI(model="inspect/google/gemini-1.5-pro")The sandbox_agent_bridge() serves only models the eval chose. A request is routed as follows:
- A name in
model_aliasesgoes to the model it maps to. - A name the
model_resolverreturns a model for goes to that model. "inspect", and the eval’s own model name (e.g.claude-sonnet-4-5for an eval onanthropic/claude-sonnet-4-5), go to the eval’s model.- Any other name, including
inspect/<model>names and model role names, also goes to the eval’s model. The bridge logs a warning once for each such name, and therequested_modelfield of the model event in the transcript records the name the agent sent.
Pass model to send unmatched names (and the eval’s own model name) to another model instead, e.g. model="inspect/openai/gpt-4o".
Agents often call more than one model: Claude Code, for example, makes background calls to a smaller model. These calls are served by the eval’s model unless you map the names. To send a name to another model, add it to model_aliases. Keys are the exact names the agent sends, as shown in the warning. To send a sub-agent tier to a model role you define for it, map the name to get_model(role=...):
from inspect_ai.model import get_model
async with sandbox_agent_bridge(
state,
model_aliases={
"claude-haiku-4-5": get_model("anthropic/claude-haiku-4-5"),
"subagent": get_model(role="subagent"),
},
) as bridge:
...Every alias key is a model the agent can call, so do not alias a role the agent should not reach, such as a grader or monitor.
Names for the eval’s own model need no alias. They already get the eval’s model with its generation config, which an alias to a model name string would not.
Generation Config
Bridged agents typically tune their requests (e.g. max_tokens, temperature, reasoning effort) for the model named in their own configuration, which is not the Inspect model actually serving the request. Those values are therefore targeted for the wrong model, and forwarding them can produce incorrect or even failing requests (for example a small max_tokens combined with reasoning being enabled on the real model).
For this reason, by default the bridge does not forward client generation parameters. Instead, the resolved Inspect model configuration and the provider’s defaults govern generation. Structural parameters that express the agent’s intent — the system prompt, tools, tool choice, response format, stop sequences, and seed — are always forwarded.
The generation parameters dropped by default are: max_tokens, temperature, top_p, top_k, frequency_penalty, presence_penalty, n/num_choices, logprobs, top_logprobs, logit_bias, and reasoning effort/tokens/summary.
If you want the bridge to act as a faithful proxy where the agent’s generation parameters are authoritative, pass forward_generation_config=True:
async with agent_bridge(state, forward_generation_config=True) as bridge:
...This option is available on both agent_bridge() and sandbox_agent_bridge().
Request Settings
Some request fields and tool options change billing, storage at the model provider, how much of the conversation the provider keeps, or what provider tools may do. On both agent_bridge() and sandbox_agent_bridge() the eval’s configuration decides these, not the agent:
service_tier,storeandtruncation(OpenAI Responses) andservice_tier(Anthropic) are not forwarded. Set them with the provider’s model args (for OpenAI,-M service_tier=flexor-M responses_store=true) or withGenerateConfig(extra_body=...). If the agent asks for a different value, a warning is logged once per field.A Responses request with
previous_response_idis refused with a 400 error (raised as aBridgePolicyErrorfrom the client call for agent_bridge()). The bridge keeps no stored responses, so the agent must send the whole conversation ininput.metadata,safety_identifier,prompt_cache_key,prompt_cache_retentionandmax_tool_callsare forwarded.Options the agent sets on a provider tool it declares do not replace the bridge’s
web_searchandcode_executionoptions. This covers the code interpreter’scontainer(OpenAI), and web search filters and allowed or blocked domains (OpenAI and Anthropic). For anything that differs, a warning is logged once per tool.The agent may set a web search
user_location(OpenAI and Anthropic) orsearch_context_size(OpenAI) that the bridge’sweb_searchoptions leave unset. These shape the results but do not widen what can be searched.search_context_sizecan change cost, because OpenAI bills search content tokens; to control that cost, set it in the bridge’sweb_searchoption. When the eval sets one, the eval’s value is used.The agent may also ask for less: a lower Anthropic
max_uses, or OpenAIexternal_web_access: false, is applied.An OpenAI computer tool needs storage at the provider. If the eval turns storage off (
-M responses_store=falseorstore: FalseinGenerateConfig(extra_body=...)), a request that declares the computer tool is refused with a 400 error. When the eval does not set storage, Inspect turns it on for these requests, as it does for any OpenAI computer tool.
OpenAI Chat Completions and Google requests forward none of these fields, and Google tool declarations carry no options into the request.
HTTP headers an agent sends to the sandbox bridge’s proxy are not forwarded; only the request body reaches the model provider. The in-process agent_bridge() forwards the agent’s headers other than credentials and SDK-managed ones.
Tool Approval
Tool calls made by a bridged agent can be governed by approval policies, even though the agent executes its own tools. Approval is applied to the tool calls in each model response before that response reaches the agent, so a rejected call is never run.
Eval-level and task-level policies apply automatically; both agent_bridge() and sandbox_agent_bridge() also accept an approval parameter:
from inspect_ai.approval import ApprovalPolicy, auto_approver, human_approver
async with sandbox_agent_bridge(
state,
approval=[
ApprovalPolicy(human_approver(), "bash"),
ApprovalPolicy(auto_approver(), "*"),
],
) as bridge:
...Rejection works by telling the model rather than by editing the response: the rejected call and an explanation are replayed to the model, which then generates again. The agent sees only the replacement, so its loop continues normally. See Bridged Agents for the full semantics.
For bridged_tools, the approved call is also what the agent may execute on the host: a bridged tool runs once per proposed call, with the arguments as approved or modified (see Execution Contract).
An agent that reaches every bridged tool through one dispatcher function (Antigravity’s call_mcp_tool) is approved as if the model had called the bridged tool directly: policies match the tool’s own name, and approvers see and modify the tool’s own arguments.
Transcript
Custom agents run through a bridge still get most of the benefit of the Inspect transcript and log viewer. All model calls are captured and produce the same transcript output as when using conventional agents.
If you want to use additional features of Inspect transcripts (e.g. spans, markdown output, etc.) you can still import and use the transcript function as normal. For example:
from inspect_ai.log import transcript
transcript().info("custom *markdown* content")