# Standard Tools – Inspect

## Overview

Inspect has built-in tools for computing and agentic planning. Computing tools include:

- [Web Search](./tools-standard.html.md#sec-web-search), which uses a search provider (either built in to the model or external) to execute and summarize web searches.
- [Bash and Python](./tools-standard.html.md#sec-bash-and-python) for executing arbitrary shell and Python code (requires a [sandbox](./sandboxing.html.md)).
- [Bash Session](./tools-standard.html.md#sec-bash-session) for creating a stateful bash shell that retains its state across calls from the model (requires a [sandbox](./sandboxing.html.md)).
- [Text Editor](./tools-standard.html.md#sec-text-editor) which enables viewing, creating and editing text files (requires a [sandbox](./sandboxing.html.md)).
- [Computer](./tools-standard.html.md#sec-computer), which provides the model with a desktop computer (viewed through screenshots) that supports mouse and keyboard interaction (requires a [sandbox](./sandboxing.html.md)).
- [Code Execution](./tools-standard.html.md#sec-code-execution), which gives models a Python code execution environment hosted within the model provider’s infrastructure rather than an Inspect sandbox.
- [Web Browser](./tools-standard.html.md#sec-web-browser), which provides the model with a headless Chromium web browser that supports navigation, history, and mouse/keyboard interactions (requires a [sandbox](./sandboxing.html.md)). Deprecated: use Web Search, Computer, or a [browser MCP server](./tools-mcp.html.md#sandboxes) in the sandbox instead.

Agentic tools include:

- [Skill](./tools-standard.html.md#sec-skill) which provides agent skill specifications to the model with specialized knowledge and expertise for specific tasks (requires a [sandbox](./sandboxing.html.md)).
- [Todo Write](./tools-standard.html.md#sec-todo-write) which helps the model tracks steps and progress across longer horizon tasks.
- [Memory](./tools-standard.html.md#sec-memory) which enables storing and retrieving information through a memory file directory.
- [Think](./tools-standard.html.md#sec-think), which provides models the ability to include an additional thinking step as part of getting to its final answer.
- [Intervention](./tools-standard.html.md#sec-intervention), which enable the model to ask questions or send notifications to the user.

## Web Search

The [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool provides models the ability to enhance their context window by performing a search. Web searches are executed using a provider. Providers are split into two categories:

- Internal providers: `"openai"`, `"anthropic"`, `"gemini"`, `"grok"`, `"mistral"`, and `"perplexity"` - these use the model’s built-in search capability and do not require separate API keys. These work only for their respective model provider (e.g. the “openai” search provider works only for `openai/*` models).

- External providers: `"tavily"`, `"exa"`, and `"google"`. These are external services that work with any model and require separate accounts and API keys. Note that “google” is different from “gemini” - “google” refers to Google’s Programmable Search Engine service, while “gemini” refers to Google’s built-in search capability for Gemini models.

By default, all internal providers are enabled if there are no external providers defined. If an external provider is defined then you need to explicitly enable internal providers that you want to use.

Internal providers will be prioritized if running on the corresponding model (e.g., “openai” provider will be used when running on `openai` models). If an internal provider is specified but the evaluation is run with a different model, a fallback external provider must also be specified.

### Configuration

> **IMPORTANT: Important**
>
> Most providers bill separately for web search, so you should consult their documentation for details before enabling this feature.

You can configure the [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool in various ways:

``` python
from inspect_ai.tool import web_search

# use all internal providers
web_search()

# single external provider
web_search("tavily")

# internal provider and fallback
web_search(["openai", "tavily"])

# multiple internal providers and fallback
web_search(["openai", "anthropic", "gemini", "mistral", "tavily"])

# provider with specific options
web_search({"tavily": {"max_results": 5}})

# multiple providers with options
web_search({
    "openai": True, 
    "google": {"num_results": 5}, 
    "tavily": {"max_results": 5}
})
```

### OpenAI Options

The [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool can use OpenAI’s built-in search capability when running on a limited number of OpenAI models (currently “gpt-4o”, “gpt-4o-mini”, “gpt-4.1”, “o3”, “o4-mini”, and GPT-5 and later, including GPT-6). This provider does not require any API keys beyond what’s needed for the model itself.

For more details on OpenAI’s web search parameters, see [OpenAI Web Search Documentation](https://platform.openai.com/docs/guides/tools-web-search?api-mode=responses).

Note that when using the “openai” provider, you should also specify a fallback external provider (like “tavily”, “exa”, or “google”) if you are also running the evaluation with non-OpenAI model.

### Anthropic Options

The [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool can use Anthropic’s built-in search capability when running on a limited number of Anthropic models (currently “claude-opus-4-20250514”, “claude-sonnet-4-20250514”, “claude-3-7-sonnet-20250219”, “claude-3-5-sonnet-latest”, “claude-3-5-haiku-latest”). This provider does not require any API keys beyond what’s needed for the model itself.

For more details on Anthropic’s web search parameters, see [Anthropic Web Search Documentation](https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/web-search-tool).

Note that when using the “anthropic” provider, you should also specify a fallback external provider (like “tavily”, “exa”, or “google”) if you are also running the evaluation with non-Anthropic model.

### Gemini Options

The [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool can use Google’s built-in search capability (called grounding) when running on Gemini 2.0 models and later. This provider does not require any API keys beyond what’s needed for the model itself.

This is distinct from the “google” provider (described below), which uses Google’s external Programmable Search Engine service and requires separate API keys.

For more details, see [Grounding with Google Search](https://ai.google.dev/gemini-api/docs/grounding).

Note that when using the “gemini” provider, you should also specify a fallback external provider (like “tavily”, “exa”, or “google”) if you are also running the evaluation with non-Gemini models.

> **NOTE: Note**
>
> Gemini 3 and later models can use `web_search("gemini")` alongside other tools. For Gemini 2.x models, Google’s search grounding does not support use with other function tools, so Inspect will raise an error if you attempt to combine them. Use an external search provider such as “tavily”, “exa”, or “google” when you need web search alongside other tools on Gemini 2.x.

### Grok Options

The [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool can use Grok’s built-in live search capability when running on Grok 3.0 models and later. This provider does not require any API keys beyond what’s needed for the model itself.

For more details, see [Live Search](https://docs.x.ai/docs/guides/live-search).

Note that when using the “grok” provider, you should also specify a fallback external provider (like “tavily”, “exa”, or “google”) if you are also running the evaluation with non-Grok models.

### Perplexity Options

The [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool can use Perplexity’s built-in search capability when running on Perplexity models. This provider does not require any API keys beyond what’s needed for the model itself. Search parameters can be passed using the `perplexity` provider options and will be forwarded to the model API.

For more details, see [Perplexity API Documentation](https://docs.perplexity.ai/api-reference/chat-completions-post).

Note that when using the “perplexity” provider, you should also specify a fallback external provider (like “tavily”, “exa”, or “google”) if you are also running the evaluation with non-Perplexity models.

### Tavily Options

The [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool can use [Tavily](https://tavily.com/)’s Research API. To use it you will need to set up your own Tavily account. Then, ensure that the following environment variable is defined:

- `TAVILY_API_KEY` — Tavily Research API key

Tavily supports the following options:

| Option | Description |
|----|----|
| `max_results` | Number of results to return |
| `search_depth` | Can be “basic” or “advanced” |
| `topic` | Can be “general” or “news” |
| `include_domains` / `exclude_domains` | Lists of domains to include or exclude |
| `time_range` | Time range for search results (e.g., “day”, “week”, “month”) |
| `max_connections` | Maximum number of concurrent connections |

For more options, see the [Tavily API Documentation](https://docs.tavily.com/documentation/api-reference/endpoint/search).

### Exa Options

The [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool can use [Exa](https://exa.ai/)’s Answer API. To use it you will need to set up your own Exa account. Then, ensure that the following environment variable is defined:

- `EXA_API_KEY` — Exa API key

Exa supports the following options:

| Option | Description |
|----|----|
| `text` | Whether to include text content in citations (defaults to true) |
| `model` | LLM model to use for generating the answer (“exa” or “exa-pro”) |
| `max_connections` | Maximum number of concurrent connections |

For more details, see the [Exa API Documentation](https://docs.exa.ai/reference/answer).

### Google Options

The [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool can use [Google Programmable Search Engine](https://programmablesearchengine.google.com/about/) as an external provider. This is different from the “gemini” provider (described above), which uses Google’s built-in search capability for Gemini models.

To use the “google” provider you will need to set up your own Google Programmable Search Engine and also enable the [Programmable Search Element Paid API](https://developers.google.com/custom-search/docs/paid_element). Then, ensure that the following environment variables are defined:

- `GOOGLE_CSE_ID` — Google Custom Search Engine ID
- `GOOGLE_CSE_API_KEY` — Google API key used to enable the Search API

Google supports the following options:

| Option | Description |
|----|----|
| `num_results` | The number of relevant webpages whose contents are returned |
| `max_provider_calls` | Number of times to retrieve more links in case previous ones were irrelevant (defaults to 3) |
| `max_connections` | Maximum number of concurrent connections (defaults to 10) |
| `model` | Model to use to determine if search results are relevant (defaults to the model being evaluated) |

## Bash and Python

The [bash()](./reference/inspect_ai.tool.html.md#bash) and [python()](./reference/inspect_ai.tool.html.md#python) tools enable execution of arbitrary shell commands and Python code, respectively. These tools require the use of a [Sandbox Environment](./sandboxing.html.md) for the execution of untrusted code. For example, here is how you might use them in an evaluation where the model is asked to write code in order to solve capture the flag (CTF) challenges:

``` python
from inspect_ai.tool import bash, python

CMD_TIMEOUT = 180

@task
def intercode_ctf():
    return Task(
        dataset=read_dataset(),
        solver=[
            system_message("system.txt"),
            use_tools([
                bash(CMD_TIMEOUT), 
                python(CMD_TIMEOUT)
            ]),
            generate(),
        ],
        scorer=includes(),
        message_limit=30,
        sandbox="docker",
    )
```

We specify a 3-minute timeout for execution of the bash and python tools to ensure that they don’t perform extremely long running operations.

See the [Agents](./agents.html.md) section for more details on how to build evaluations that allow models to take arbitrary actions over a longer time horizon.

### Background Tasks

The [bash()](./reference/inspect_ai.tool.html.md#bash) tool can be configured to encourage the model to run long operations in the background and poll for progress in later calls rather than blocking:

``` python
use_tools([bash(timeout=180, background=True)])
```

The `background` option is prompt-only, it doesn’t change how commands execute, it only augments the tool’s description with guidance to launch long-running commands detached (e.g. `nohup <command> > /tmp/task.log 2>&1 &`), record the process id, and check on progress with subsequent calls (`ps`, `tail`).

This works because a detached process keeps running between [bash()](./reference/inspect_ai.tool.html.md#bash) calls even though each call executes in a fresh shell. For interactive long-running commands (e.g. ones you need to send input to or interrupt), prefer the [Bash Session](#sec-bash-session) tool instead.

## Bash Session

The [bash_session()](./reference/inspect_ai.tool.html.md#bash_session) tool provides a bash shell that retains its state across calls from the model (as distinct from the [bash()](./reference/inspect_ai.tool.html.md#bash) tool which executes each command in a fresh session). The prompt, working directory, and environment variables are all retained across calls. The tool also supports a `restart` action that enables the model to reset its state and work in a fresh session.

Note that a separate bash process is created within the sandbox for each instance of the bash session tool. See the [bash_session()](./reference/inspect_ai.tool.html.md#bash_session) reference docs for details on customizing this behavior.

### Configuration

Bash sessions require the use of a [Sandbox Environment](./sandboxing.html.md) for the execution of untrusted code. Like [bash()](./reference/inspect_ai.tool.html.md#bash), the session runs as the sandbox’s default user unless a `user` is specified.

### Task Setup

A task configured to use the bash session tool might look like this:

``` python
from inspect_ai import Task, task
from inspect_ai.scorer import includes
from inspect_ai.solver import generate, system_message, use_tools
from inspect_ai.tool import bash_session

@task
def intercode_ctf():
    return Task(
        dataset=read_dataset(),
        solver=[
            system_message("system.txt"),
            use_tools([bash_session(timeout=180)]),
            generate(),
        ],
        scorer=includes(),
        sandbox=("docker", "compose.yaml")
    )
```

Note that we provide a `timeout` for bash session commands (this is a best practice to guard against extremely long running commands).

## Text Editor

The [text_editor()](./reference/inspect_ai.tool.html.md#text_editor) tool enables viewing, creating and editing text files. The tool supports editing files within a protected [Sandbox Environment](./sandboxing.html.md) so tasks that use the text editor should have a sandbox defined and configured as described below.

### Configuration

The text editor tools requires the use of a [Sandbox Environment](./sandboxing.html.md). Like [bash()](./reference/inspect_ai.tool.html.md#bash), it runs as the sandbox’s default user unless a `user` is specified. Viewing a directory runs the sandbox’s `find`, which must be installed in `/usr/sbin`, `/usr/bin`, `/sbin` or `/bin`.

### Task Setup

A task configured to use the text editor tool might look like this (note that this task is also configured to use the [bash_session()](./reference/inspect_ai.tool.html.md#bash_session) tool):

``` python
from inspect_ai import Task, task
from inspect_ai.scorer import includes
from inspect_ai.solver import generate, system_message, use_tools
from inspect_ai.tool import bash_session, text_editor

@task
def intercode_ctf():
    return Task(
        dataset=read_dataset(),
        solver=[
            system_message("system.txt"),
            use_tools([
                bash_session(timeout=180),
                text_editor(timeout=180)
            ]),
            generate(),
        ],
        scorer=includes(),
        sandbox=("docker", "compose.yaml")
    )
```

Note that we provide a `timeout` for the bash session and text editor tools (this is a best practice to guard against extremely long running commands).

### Tool Binding

The schema for the [text_editor()](./reference/inspect_ai.tool.html.md#text_editor) tool is based on the standard Anthropic [text editor tool type](https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/text-editor-tool). The [text_editor()](./reference/inspect_ai.tool.html.md#text_editor) works with all models that support tool calling, but when using Claude, the text editor tool will automatically bind to the native Claude tool definition.

## Computer

The [computer()](./reference/inspect_ai.tool.html.md#computer) tool provides models with a computer desktop environment along with the ability to view the screen and perform mouse and keyboard gestures. The computer tool work better with models that have been trained for computer use. As of Q1 2026 the recommended models for computer use include:

| Provider | Models |
|----|----|
| Anthropic | `claude-opus-4-5+`, `claude-sonnet-4-6+`, `claude-fable-5+`, `claude-mythos-5+` |
| Open AI | `gpt-5.4+`, `gpt-5.4-pro+` |
| Google | `gemini-3-flash-preview` |

### Configuration

The [computer()](./reference/inspect_ai.tool.html.md#computer) tool runs within a Docker container. To use it with a task you need to reference the `aisiuk/inspect-computer-tool` image in your Docker compose file. For a task that does not require networking, retain Inspect’s no-network posture explicitly:

    compose.yaml

``` yaml
services:
  default:
    image: aisiuk/inspect-computer-tool
    network_mode: none
```

Omit `network_mode: none` only when the task needs networking. If you’d like to view the model’s interactions with the computer desktop in realtime, you will also need port mapping to enable a VNC connection with the container. See the [VNC Client](#vnc-client) section below for details.

The `aisiuk/inspect-computer-tool` image is based on the [ubuntu:22.04](https://hub.docker.com/layers/library/ubuntu/22.04/images/sha256-965fbcae990b0467ed5657caceaec165018ef44a4d2d46c7cdea80a9dff0d1ea?context=explore) image and includes the following additional applications pre-installed:

- Firefox
- VS Code
- Xpdf
- Xpaint
- galculator

### Task Setup

A task configured to use the computer tool might look like this:

``` python
from inspect_ai import Task, task
from inspect_ai.scorer import match
from inspect_ai.solver import generate, use_tools
from inspect_ai.tool import computer

@task
def computer_task():
    return Task(
        dataset=read_dataset(),
        solver=[
            use_tools([computer()]),
            generate(),
        ],
        scorer=match(),
        sandbox=("docker", "compose.yaml"),
    )
```

To evaluate the task with models tuned for computer use:

``` bash
inspect eval computer.py --model anthropic/claude-sonnet-4-6
inspect eval computer.py --model openai/gpt-5.4
inspect eval computer.py --model google/gemini-3-flash-preview
```

#### Options

The computer tool supports the following options:

| Option | Description |
|----|----|
| `max_screenshots` | The maximum number of screenshots to play back to the model as input. Defaults to 1 (set to `None` to have no limit). |
| `timeout` | Timeout in seconds for computer tool actions. Defaults to 180 (set to `None` for no timeout). |

When a Claude model uses Anthropic’s computer toolset (see *Tool Binding* below), it can issue several actions in one turn as a batch. Inspect runs them in order and stops at the first failure, reporting each later action to the model as `Not executed: an earlier computer action in this turn failed.`, as the toolset’s batch contract specifies. Other models and the legacy Anthropic tool are unaffected.

For example:

``` python
solver=[
    use_tools([computer(max_screenshots=2, timeout=300)]),
    generate()
]
```

#### Examples

Two of the Inspect examples demonstrate basic computer use:

- [computer](https://github.com/UKGovernmentBEIS/inspect_ai/tree/main/examples/computer/computer.py) — Three simple computing tasks as a minimal demonstration of computer use.

  ``` bash
  inspect eval examples/computer
  ```

- [intervention](https://github.com/UKGovernmentBEIS/inspect_ai/tree/main/examples/intervention/intervention.py) — Computer task driven interactively by a human operator.

  ``` bash
  inspect eval examples/intervention -T mode=computer --display conversation
  ```

### VNC Client

You can use a [VNC](https://en.wikipedia.org/wiki/VNC) connection to the container to watch computer use in real-time. This requires some additional port-mapping in the Docker compose file. You can define dynamic port ranges for VNC (5900) and a browser based noVNC client (6080) with the following `ports` entries:

    compose.yaml

``` yaml
services:
  default:
    image: aisiuk/inspect-computer-tool
    # network_mode is omitted because container networking is required for
    # these ports. This also permits outbound Internet access.
    ports:
      - "127.0.0.1::5900"
      - "127.0.0.1::6080"
```

> **WARNING: Warning**
>
> The bundled VNC server does not require a password, and VNC/noVNC traffic is not encrypted. Keep these ports bound to loopback (127.0.0.1).

To connect to the container for a given sample, locate the sample in the **Running Samples** UI and expand the sample info panel at the top:

[![](images/vnc-port-info.png)](images/vnc-port-info.png)

Click on the link for the noVNC browser client, or use a native VNC client to connect to the VNC port. Note that the VNC server will take a few seconds to start up so you should give it some time and attempt to reconnect as required if the first connection fails.

The browser link opens noVNC in view-only mode. This is a client-side setting, not access control: any client that can reach the VNC server can enable keyboard and mouse input. If you use a native VNC client, you should also set it to “view only” so as to not interfere with the model’s use of the computer. For example, for Real VNC Viewer:

[![](images/vnc-view-only.png)](images/vnc-view-only.png)

### Approval

If the container you are using is connected to the Internet, you may want to configure human approval for a subset of computer tool actions. Here are the possible actions (specified using the `action` parameter to the `computer` tool):

- `key`: Press a key or key-combination on the keyboard.
- `type`: Type a string of text on the keyboard.
- `cursor_position`: Get the current (x, y) pixel coordinate of the cursor on the screen.
- `mouse_move`: Move the cursor to a specified (x, y) pixel coordinate on the screen.
- Example: execute(action=“mouse_move”, coordinate=(100, 200))
- `left_click`: Click the left mouse button.
- `left_click_drag`: Click and drag the cursor to a specified (x, y) pixel coordinate on the screen.
- `right_click`: Click the right mouse button.
- `middle_click`: Click the middle mouse button.
- `double_click`: Double-click the left mouse button.
- `screenshot`: Take a screenshot.

Here is an approval policy that requires approval for key combos (e.g. `Enter` or a shortcut) and mouse clicks:

    approval.yaml

``` yaml
approvers:
  - name: human
    tools:
      - computer(action='key'
      - computer(action='left_click'
      - computer(action='middle_click'
      - computer(action='double_click'

  - name: auto
    tools: "*"
```

Note that since this is a prefix match and there could be other arguments, we don’t end the tool match pattern with a parentheses.

You can apply this policy using the `--approval` command line option:

``` bash
inspect eval computer.py --approval approval.yaml
```

### Tool Binding

The computer tool’s schema is a superset of the standard [Anthropic](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool),[OpenAI](https://platform.openai.com/docs/guides/tools-computer-use), and [Google](https://ai.google.dev/gemini-api/docs/computer-use) computer tool schemas. When using models tuned for computer use, the computer tool will automatically bind to the native computer tool definitions.

For Claude models, the binding depends on the model and platform: on the Claude API and Vertex, Claude Opus 5.5, Sonnet 5.5, Fable 5/5.1 and Mythos 5/5.1 use Anthropic’s computer toolset (`computer_toolset_20260801`), where each action is its own tool call and a failed action halts the remaining computer actions in that turn; other Claude models, and every model on Bedrock and Foundry (which offer only the earlier tool), keep the legacy `computer_20251124` tool. See [Computer Use](./providers.html.md#anthropic-computer-use) in the Anthropic provider documentation for details, including the `computer_toolset` model arg that forces either mode.

## Code Execution

### Overview

The [code_execution()](./reference/inspect_ai.tool.html.md#code_execution) tool provides models with the ability to execute Python code within a sandboxed environment. There are two significant differences between code execution and the [python()](./reference/inspect_ai.tool.html.md#python) tool described above:

1.  Code runs in a sandbox on the model provider’s server (as opposed to e.g. a locally managed Docker container).
2.  Code runs in a *stateless* environment (each execution is independent of others and no file-system state is preserved across calls).

Since the code execution tool is stateless, it is more suitable as a means to assist with problem solving that for more stateful agentic tasks.

Here is a simple example using the [code_execution()](./reference/inspect_ai.tool.html.md#code_execution) tool:

``` python
from inspect_ai import Task, task
from inspect_ai.dataset import Sample
from inspect_ai.agent import react
from inspect_ai.tool import code_execution

@task
def code_execution_task():
    return Task(
        dataset=[Sample("Add 435678 + 23457")],
        solver=react(tools=[code_execution()])
    )
```

### Availability

[OpenAI](https://platform.openai.com/docs/guides/tools-code-interpreter), [Anthropic](https://platform.claude.com/docs/en/agents-and-tools/tool-use/code-execution-tool), [Google](https://ai.google.dev/gemini-api/docs/code-execution), and [Grok](https://docs.x.ai/docs/guides/tools/code-execution-tool) models all have support for native server-side Python code execution. Note that Anthropic can additionally execute bash and text editor commands, but the primary execution language used is still Python.

For Gemini models, [code_execution()](./reference/inspect_ai.tool.html.md#code_execution) uses Google’s native code execution tool when the Google provider is enabled. Gemini 3 and later models can use native code execution alongside other tools. For Gemini 2.x models, Google’s native tools do not support use with other function tools, so Inspect will raise an error if you attempt to combine them; disable the Google native provider to use the [python()](./reference/inspect_ai.tool.html.md#python) fallback in that case.

> **IMPORTANT: Important**
>
> Note that some providers bill separately for code execution, so you should consult their documentation for details before enabling this feature.

#### Fallback

If you are using a provider that doesn’t support code execution then a fallback using the [python()](./reference/inspect_ai.tool.html.md#python) tool is provided. Additionally, you can optionally disable code execution for a provider with a native implementation and use the [python()](./reference/inspect_ai.tool.html.md#python) tool instead.

Here are some example configurations:

``` python
# default (native where supported, python as fallback):
code_execution()

# selectively disable native (will fallback to python)
code_execution(providers={ "grok": False, "openai": False })

# disable python fallback
code_execution(providers={ "python": False })

# provide openai container options
code_execution(
    providers={"openai": {"container": {"type": "auto", "memory_limit": "4g" }}}
)
```

When falling back to the [python()](./reference/inspect_ai.tool.html.md#python) provider you should ensure that your [Task](./reference/inspect_ai.html.md#task) has a `sandbox` with access to Python enabled.

## Web Browser

The web browser tools provides models with the ability to browse the web using a headless Chromium browser. Navigation, history, and mouse/keyboard interactions are all supported.

> **WARNING: Warning**
>
> The [web_browser()](./reference/inspect_ai.tool.html.md#web_browser) tool is deprecated and its implementation will be removed around November 2026. Calling [web_browser()](./reference/inspect_ai.tool.html.md#web_browser) now logs a deprecation warning; after removal it will raise an error. Most websites now block headless browsers, so for web information retrieval use the [web_search()](#sec-web-search) tool. For interacting with web applications (local web apps, forms, multi-step flows) use the [computer()](#sec-computer) tool, or run a browser MCP server such as [Playwright MCP](https://github.com/microsoft/playwright-mcp) inside your sandbox with [mcp_server_sandbox()](./tools-mcp.html.md#sandboxes). See [issue \#5497](https://github.com/UKGovernmentBEIS/inspect_ai/issues/5497) for details, and comment there if you rely on [web_browser()](./reference/inspect_ai.tool.html.md#web_browser) in a way these paths do not cover.

### Configuration

Under the hood, the web browser is an instance of [Chromium](https://www.chromium.org/chromium-projects/) orchestrated by [Playwright](https://playwright.dev/), and runs in a [Sandbox Environment](./sandboxing.html.md). In addition, you’ll need some dependencies installed in the sandbox container. Please see **Sandbox Dependencies** below for additional instructions.

Note that Playwright (used for the [web_browser()](./reference/inspect_ai.tool.html.md#web_browser) tool) does not support some versions of Linux (e.g. Kali Linux).

> **NOTE: NoteSandbox Dependencies**
>
> You should add the following to your sandbox `Dockerfile` in order to use the web browser tool:
>
> ``` dockerfile
> RUN apt-get update && apt-get install -y pipx && \
>     apt-get clean && rm -rf /var/lib/apt/lists/*
> ENV PATH="$PATH:/opt/inspect/bin"
> RUN PIPX_HOME=/opt/inspect/pipx PIPX_BIN_DIR=/opt/inspect/bin PIPX_VENV_DIR=/opt/inspect/pipx/venvs \
>     pipx install inspect-tool-support && \
>     chmod -R 755 /opt/inspect && \
>     inspect-tool-support post-install
> ```
>
> If you don’t have a custom Dockerfile, you can alternatively use the pre-built `aisiuk/inspect-tool-support` image. This configuration intentionally permits outbound Internet access so the browser can visit external sites:
>
>     compose.yaml
>
> ``` yaml
> services:
>   default:
>     image: aisiuk/inspect-tool-support
>     init: true
>     # network_mode is omitted because the browser visits external sites.
> ```
>
> The pre-built image runs as root by default and also provides a `nonroot` account (UID/GID 65532) for evaluations that should not run the agent as root; add `user: nonroot` to the service to use it. The web browser tool works under either user.

### Task Setup

A task configured to use the web browser tools might look like this:

``` python
from inspect_ai import Task, task
from inspect_ai.scorer import match
from inspect_ai.solver import generate, use_tools
from inspect_ai.tool import bash, python, web_browser

@task
def browser_task():
    return Task(
        dataset=read_dataset(),
        solver=[
            use_tools([bash(), python()] + web_browser()),
            generate(),
        ],
        scorer=match(),
        sandbox=("docker", "compose.yaml"),
    )
```

Unlike some other tool functions like [bash()](./reference/inspect_ai.tool.html.md#bash), the [web_browser()](./reference/inspect_ai.tool.html.md#web_browser) function returns a list of tools. Therefore, we concatenate it with a list of the other tools we are using in the call to [use_tools()](./reference/inspect_ai.solver.html.md#use_tools).

Note that a separate web browser process is created within the sandbox for each instance of the web browser tool. See the [web_browser()](./reference/inspect_ai.tool.html.md#web_browser) reference docs for details on customizing this behavior.

### Browsing

If you review the transcripts of a sample with access to the web browser tool, you’ll notice that there are several distinct tools made available for control of the web browser. These tools include:

| Tool | Description |
|----|----|
| `web_browser_go(url)` | Navigate the web browser to a URL. |
| `web_browser_click(element_id)` | Click an element on the page currently displayed by the web browser. |
| `web_browser_type(element_id)` | Type text into an input on a web browser page. |
| `web_browser_type_submit(element_id, text)` | Type text into a form input on a web browser page and press ENTER to submit the form. |
| `web_browser_scroll(direction)` | Scroll the web browser up or down by one page. |
| `web_browser_forward()` | Navigate the web browser forward in the browser history. |
| `web_browser_back()` | Navigate the web browser back in the browser history. |
| `web_browser_refresh()` | Refresh the current page of the web browser. |

The return value of each of these tools is a [web accessibility tree](https://web.dev/articles/the-accessibility-tree) for the page, which provides a clean view of the content, links, and form fields available on the page (you can look at the accessibility tree for any web page using [Chrome Developer Tools](https://developer.chrome.com/blog/full-accessibility-tree)).

### Disabling Interactions

You can use the web browser tools with page interactions disabled by specifying `interactive=False`, for example:

``` python
use_tools(web_browser(interactive=False))
```

In this mode, the interactive tools (`web_browser_click()`, `web_browser_type()`, and `web_browser_type_submit()`) are not made available to the model.

## Skill

The [skill()](./reference/inspect_ai.tool.html.md#skill) tool provides models with [agent skills](https://agentskills.io/home) which are folders of instructions, scripts, and resources that agents can discover and use to do things more accurately and efficiently.

Skills were originally created as a feature of Claude Code, but are now widely supported by many agents and agent frameworks. You can learn more about creating skills at:

- [Agent Skills Specification](https://agentskills.io/specification)
- [Claude Code Agent Skills](https://code.claude.com/docs/en/skills)
- [Codex CLI Agent Skills](https://developers.openai.com/codex/skills/)
- [Gemini CLI Agent Skills](https://geminicli.com/docs/cli/skills/)

The [skill()](./reference/inspect_ai.tool.html.md#skill) tool takes a list of paths that contain standard skill specifications, copies them into the sample’s sandbox, and provides a tool description that enumerates the available skills. For example, here we make available “system-info” and “network-info” skills:

``` python
from inspect_ai import Task, task
from inspect_ai.scorer import includes
from inspect_ai.agent import react
from inspect_ai.tool import bash, skill, todo_write

SKILLS_DIR = Path(__file__).parent / "skills"

@task
def intercode_ctf():

    # define skill tool
    skill_tool = skill(
        [
            SKILLS_DIR / "system-info",
            SKILLS_DIR / "network-info",
        ]
    )

    return Task(
        dataset=read_dataset(),
        solver=react(tools=[bash(timeout=180), skill_tool]),
        scorer=includes(),
        sandbox=("docker", "compose.yaml")
    )
```

Note that use of the [skill()](./reference/inspect_ai.tool.html.md#skill) tool requires a that a [sandbox](./sandboxing.html.md) be defined for the task so there is a filesystem to publish the skills within.

## Todo Write

The [todo_write()](./reference/inspect_ai.tool.html.md#todo_write) tool provides models with a way to track steps and progress in longer horizon tasks where it might otherwise lose track of where it is or forget earlier goals as context grows. It can also make agent behavior more interpretable, since you can inspect the plan to understand what the model thinks it’s trying to accomplish.

Note though that for simpler tasks, plan maintenance is just overhead, and some models may fixate on updating the plan rather than actually executing it.

### Task Setup

A task configured to use the todo_write tool might look like this:

``` python
from inspect_ai import Task, task
from inspect_ai.scorer import includes
from inspect_ai.agent import react
from inspect_ai.tool import bash, todo_write

@task
def intercode_ctf():
    return Task(
        dataset=read_dataset(),
        solver=react(tools=[bash(timeout=180), todo_write()]),
        scorer=includes(),
        sandbox=("docker", "compose.yaml")
    )
```

## File Reading

Inspect provides three read-only sandbox tools — [read_file()](./reference/inspect_ai.tool.html.md#read_file), [list_files()](./reference/inspect_ai.tool.html.md#list_files), and [grep()](./reference/inspect_ai.tool.html.md#grep) — for agents that need filesystem access without write capabilities. These are the default tools for [research()](./reference/inspect_ai.agent.html.md#research) and [plan()](./reference/inspect_ai.agent.html.md#plan) subagents in the deep agent system, but are useful in any eval where you want to give a model read-only access.

All three tools require a [Sandbox Environment](./sandboxing.html.md) and accept optional `timeout`, `user`, and `sandbox` parameters matching the [bash()](./reference/inspect_ai.tool.html.md#bash) tool.

### read_file

Read the contents of a file, optionally selecting a range of lines:

``` python
from inspect_ai.tool import read_file

# default configuration
read_file()

# with timeout and user
read_file(timeout=30, user="nobody")
```

The model can specify `offset` (0-indexed line to start from) and `limit` (max lines to read) for pagination. Output includes line numbers for reference.

### list_files

List files and directories, with optional depth control:

``` python
from inspect_ai.tool import list_files

# default configuration (recursive)
list_files()

# with depth limit
list_files(timeout=30)
```

The model can specify a `path` and `depth` parameter. `depth=1` lists only immediate contents; omitting it lists everything recursively.

### grep

Search for patterns in files:

``` python
from inspect_ai.tool import grep

# default configuration
grep()

# with timeout
grep(timeout=60)
```

The model can specify a `pattern`, `path`, optional `glob` filter (e.g. `"*.py"`), `fixed_strings` flag for literal matching, `extended_regexp` flag for extended regex (ERE), and `output_mode` (`"content"`, `"files_with_matches"`, or `"count"`). Patterns use grep’s basic regex (BRE) by default. Results include file paths and line numbers by default.

### Task Setup

A task configured with read-only tools might look like this:

``` python
from inspect_ai import Task, task
from inspect_ai.scorer import includes
from inspect_ai.agent import react
from inspect_ai.tool import read_file, list_files, grep

@task
def code_analysis():
    return Task(
        dataset=read_dataset(),
        solver=react(tools=[read_file(), list_files(), grep()]),
        scorer=includes(),
        sandbox=("docker", "compose.yaml")
    )
```

## Memory

The memory tool enables models to store and retrieve information into a virtual `/memories` file directory. Models can create, read, update, and delete files, enabling them to preserve knowledge over time without keeping everything in the context window.

Note that the [memory()](./reference/inspect_ai.tool.html.md#memory) tool does not require a [Sandbox Environment](./sandboxing.html.md)—despite using file-like paths (e.g. `/memories/notes.md`), it stores all data in-memory using Inspect’s sample store.

### Task Setup

A task configured to use the memory tool might look like this:

``` python
from inspect_ai import Task, task
from inspect_ai.scorer import includes
from inspect_ai.agent import react
from inspect_ai.tool import memory

@task
def intercode_ctf():
    return Task(
        dataset=read_dataset(),
        solver=[
            system_message("system.txt"),
            react(tools=[memory()]),
        ],
        scorer=includes(),
    )
```

### Seeding Memories

You can seed the memories from sample data by passing `initial_data` to the [memory()](./reference/inspect_ai.tool.html.md#memory) tool. For example:

``` python
memory(
    initial_data = {
        "/memories/notes.md": "<text or file path>",
        "/memories/theories.md": "<text or file path>"
    }
)
```

Keys should be valid `/memories` paths (e.g. “/memories/notes.md”). Values are resolved via [resource()](./reference/inspect_ai.util.html.md#resource), supporting inline strings, file paths, or remote resources (s3://, https://). Seeding happens once on first tool execution. The model is prompted to read any pre-seeded memories before beginning work.

### Read-Only Mode

Use `memory(readonly=True)` to provide read-only access to the memory directory. In readonly mode, only the `view` command is available — write operations (`create`, `str_replace`, `insert`, `delete`, `rename`) are not exposed to the model. This is used by [research()](./reference/inspect_ai.agent.html.md#research) and [plan()](./reference/inspect_ai.agent.html.md#plan) subagents in the deep agent system to share context without allowing mutation.

``` python
# read-only memory with pre-seeded data
memory(
    initial_data={"/memories/context.md": "shared context"},
    readonly=True,
)
```

### Separate Stores

By default, all [memory()](./reference/inspect_ai.tool.html.md#memory) tools within a sample share a single `/memories` store — every agent and subagent reads and writes the same files. This is usually what you want: a subagent can build on what the parent recorded, and vice versa.

To run independent memory stores within one sample, pass an `instance` name. Each instance has its own files and is seeded independently:

``` python
# two independent memory stores in the same sample
notes = memory(instance="notes")
scratch = memory(instance="scratch")
```

This mirrors `skill(instance=...)`. Use it when two memory tools must not collide — for example, to give a subagent a private scratchpad, or to compare shared vs. separate memory as conditions in a study.

Note that `instance` namespaces the data within the sample’s in-memory store; it is **not** a security boundary, since everything still lives in the same process. For a hard boundary between agents, isolate them with a [Sandbox Environment](./sandboxing.html.md) instead.

### Tool Binding

The schema for the [memory()](./reference/inspect_ai.tool.html.md#memory) tool is based on the standard Anthropic [memory tool type](https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool). The [memory()](./reference/inspect_ai.tool.html.md#memory) works with all models that support tool calling, but when using Claude, the memory tool will automatically bind to the native Claude tool definition.

## Think

The [think()](./reference/inspect_ai.tool.html.md#think) tool provides models with the ability to include an additional thinking step as part of getting to its final answer.

Note that the [think()](./reference/inspect_ai.tool.html.md#think) tool is not a substitute for reasoning and extended thinking, but rather an alternate way of letting models express thinking that is better suited to some tool use scenarios.

### Usage

You should read the original [think tool article](https://www.anthropic.com/engineering/claude-think-tool) in its entirety to understand where and where not to use the think tool. In summary, good contexts for the think tool include:

1.  Tool output analysis. When models need to carefully process the output of previous tool calls before acting and might need to backtrack in its approach;
2.  Policy-heavy environments. When models need to follow detailed guidelines and verify compliance; and
3.  Sequential decision making. When each action builds on previous ones and mistakes are costly (often found in multi-step domains).

Use the [think()](./reference/inspect_ai.tool.html.md#think) tool alongside other tools like this:

``` python
from inspect_ai import Task, task
from inspect_ai.scorer import includes
from inspect_ai.solver import generate, system_message, use_tools
from inspect_ai.tool import bash_session, text_editor, think

@task
def intercode_ctf():
    return Task(
        dataset=read_dataset(),
        solver=[
            system_message("system.txt"),
            use_tools([
                bash_session(timeout=180),
                text_editor(timeout=180),
                think()
            ]),
            generate(),
        ],
        scorer=includes(),
        sandbox=("docker", "compose.yaml")
    )
```

### Tool Description

In the original [think tool article](https://www.anthropic.com/engineering/claude-think-tool) (which was based on experimenting with Claude) they found that providing clear instructions on when and how to use the [think()](./reference/inspect_ai.tool.html.md#think) tool for the particular problem domain it is being used within could sometimes be helpful. For example, here’s the prompt they used with SWE-Bench:

``` python
from textwrap import dedent

from inspect_ai import Task, task
from inspect_ai.scorer import includes
from inspect_ai.solver import generate, system_message, use_tools
from inspect_ai.tool import bash_session, text_editor, think

@task
def swe_bench():

    tools = [
        bash_session(timeout=180),
        text_editor(timeout=180),  
        think(dedent("""
            Use the think tool to think about something. It will not obtain
            new information or make any changes to the repository, but just 
            log the thought. Use it when complex reasoning or brainstorming
            is needed. For example, if you explore the repo and discover
            the source of a bug, call this tool to brainstorm several unique
            ways of fixing the bug, and assess which change(s) are likely to 
            be simplest and most effective. Alternatively, if you receive
            some test results, call this tool to brainstorm ways to fix the
            failing tests.
        """))
    ])

    return Task(
        dataset=read_dataset(),
        solver=[
            system_message("system.txt"),
            use_tools(tools),
            generate(),
        ),
        scorer=includes(),
        sandbox=("docker", "compose.yaml")
    )
```

### System Prompt

In the article they also found that when tool instructions are long and/or complex, including instructions about the [think()](./reference/inspect_ai.tool.html.md#think) tool in the system prompt can be more effective than placing them in the tool description itself.

Here’s an example of moving the custom [think()](./reference/inspect_ai.tool.html.md#think) prompt into the system prompt (note that this was *not* done in the article’s SWE-Bench experiment, this is merely an example):

``` python
from textwrap import dedent

from inspect_ai import Task, task
from inspect_ai.scorer import includes
from inspect_ai.solver import generate, system_message, use_tools
from inspect_ai.tool import bash_session, text_editor, think

@task
def swe_bench():

    think_system_message = system_message(dedent("""
        Use the think tool to think about something. It will not obtain
        new information or make any changes to the repository, but just 
        log the thought. Use it when complex reasoning or brainstorming
        is needed. For example, if you explore the repo and discover
        the source of a bug, call this tool to brainstorm several unique
        ways of fixing the bug, and assess which change(s) are likely to 
        be simplest and most effective. Alternatively, if you receive
        some test results, call this tool to brainstorm ways to fix the
        failing tests.
    """))

    return Task(
        dataset=read_dataset(),
        solver=[
            system_message("system.txt"),
            think_system_message,
            use_tools([
                bash_session(timeout=180),
                text_editor(timeout=180),  
                think(),
            ]),
            generate(),
        ],
        scorer=includes(),
        sandbox=("docker", "compose.yaml")
    )
```

Note that the effectivess of using the system prompt will vary considerably across tasks, tools, and models, so should definitely be the subject of experimentation.

## Intervention

The `ask_user()` and `notify_user()` tools let models communicate with a human operator during a sample. They pair with Inspect’s [Agent Intervention](./intervention.html.md) features (the `inspect acp` client, the in-process task display, and out-of-band notifications).

Both of these tools take advantage of notifications, which are delivered via [Apprise](https://github.com/caronc/apprise) (Slack, desktop, SMS, email, and many other services) when the eval is configured with a notification target. See the [Notifications](./intervention.html.md#notifications) section of the Agent Intervention article for more details.

### Ask User

The `ask_user()` tool lets the model request structured information from the operator. It uses the [ACP Elicitation](https://agentclientprotocol.com/rfds/elicitation) standard, which supports text, boolean, enum, and other field types.

``` python
from inspect_ai.agent import agent, react
from inspect_ai.tool import ask_user, bash, text_editor

@agent
def ctf_agent():
    return react(
        description="Expert at completing cybersecurity challenges.",
        prompt="You are an expert at CTF challenges.",
        tools=[bash(), text_editor(), ask_user()]
    )
```

The prompt is dispatched to whichever surface is attached: an ACP client (e.g. `inspect acp`) if connected, otherwise the in-process Textual panel or the console. The sample pauses on the call until the operator responds.

### Notify User

The `notify_user()` tool lets the model send fire-and-forget status messages to the operator — useful for long-running agents that want to flag progress or surface a heads-up without waiting for a reply.

``` python
from inspect_ai.agent import agent, react
from inspect_ai.tool import ask_user, notify_user, bash, text_editor

@agent
def ctf_agent():
    return react(
        description="Expert at completing cybersecurity challenges.",
        prompt="You are an expert at CTF challenges.",
        tools=[bash(), text_editor(), ask_user(), notify_user()]
    )
```
