# Model Providers – Inspect

## Overview

Inspect has support for a wide variety of language model APIs and can be extended to support arbitrary additional ones. Support for the following providers is built in to Inspect:

|  |  |
|----|----|
| Lab APIs | [OpenAI](./providers.html.md#openai), [Anthropic](./providers.html.md#anthropic), [Google](./providers.html.md#google), [Grok](./providers.html.md#grok), [Mistral](./providers.html.md#mistral), [DeepSeek](./providers.html.md#deepseek), [Moonshot AI](./providers.html.md#moonshot-ai), [Meta](./providers.html.md#meta), [Perplexity](./providers.html.md#perplexity) |
| Cloud APIs | [AWS Bedrock](./providers.html.md#aws-bedrock), [AWS SageMaker](./providers.html.md#aws-sagemaker), and [Azure AI](./providers.html.md#azure-ai) |
| Open (Hosted) | [Groq](./providers.html.md#groq), [Together AI](./providers.html.md#together-ai), [Fireworks AI](./providers.html.md#fireworks-ai), [Cloudflare](./providers.html.md#cloudflare), [HF Inference Providers](./providers.html.md#hugging-face-inference-providers), [SambaNova](./providers.html.md#sambanova) |
| Open (Local) | [Hugging Face](./providers.html.md#hugging-face), [vLLM](./providers.html.md#vllm), [Ollama](./providers.html.md#ollama), [Lllama-cpp-python](./providers.html.md#llama-cpp-python), [SGLang](./providers.html.md#sglang), [TransformerLens](./providers.html.md#transformer-lens), [nnterp](./providers.html.md#nnterp) |

\

If the provider you are using is not listed above, you may still be able to use it if:

1.  It provides an OpenAI compatible API endpoint. In this scenario, use the Inspect [OpenAI Compatible API](./providers.html.md#openai-api) interface.

2.  It is available via OpenRouter (see the docs on using [OpenRouter](./providers.html.md#openrouter) with Inspect).

3.  It is served by a LiteLLM proxy (see the docs on using a [LiteLLM Proxy](./providers.html.md#litellm-proxy) with Inspect).

You can also create [Model API Extensions](./extensions-model-api.html.md#model-apis) to add model providers using their native interface.

## OpenAI

To use the [OpenAI](https://platform.openai.com/) provider, install the `openai` package, set your credentials, and specify a model using the `--model` option:

``` bash
pip install openai
export OPENAI_API_KEY=your-openai-api-key
inspect eval arc.py --model openai/gpt-4o-mini
```

The following environment variables are supported by the OpenAI provider

| Variable | Description |
|----|----|
| `OPENAI_API_KEY` | API key credentials (required). |
| `OPENAI_BASE_URL` | Base URL for requests (optional, defaults to `https://api.openai.com/v1`) |
| `OPENAI_ORG_ID` | OpenAI organization ID (optional) |
| `OPENAI_PROJECT_ID` | OpenAI project ID (optional) |
| `OPENAI_SAFETY_IDENTIFIER` | Default `safety_identifier` passed with each request (optional; overridden by the `safety_identifier` model arg if set). |

### Model Args

The `openai` provider supports the following custom model args (other model args are forwarded to the constructor of the `AsyncOpenAI` class):

| Model Arg | Description |
|----|----|
| `responses_api` | Use the OpenAI Responses API rather than the Chat Completions API. |
| `responses_store` | Pass `store=True` to the Responses API (defaults to `True`). |
| `responses_phase` | Synthesize missing assistant message `phase` values when replaying Responses API histories. |
| `service_tier` | Processing type used for serving the request (“auto”, “default”, or “flex”). |
| `background` | Execute generate requests asynchronously, polling response objects to check status over time. Defaults to `True` for `gpt-5-pro` and `deep-research` models and for requests with `reasoning_mode="pro"`, and `False` otherwise. |
| `safety_identifier` | A stable identifier used to help detect users of your application. |
| `streaming` | Whether to use streaming responses. The default (“auto”) streams when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate) — except Azure chat-completions requests, which never auto-stream (accumulated streamed responses would lose content-filter stop details); pass `true` or `false` to override. If the server rejects the streaming request itself (a 400 naming the `stream` param — e.g. reasoning models on organizations that haven’t completed verification), auto mode retries the request non-streamed. |
| `prompt_cache_key` | Used by OpenAI to cache responses for similar requests. |
| `prompt_cache_retention` | Retention policy for the prompt cache. |
| `http_client` | Custom instance of `httpx2.AsyncClient` (or legacy `httpx.AsyncClient`) for handling requests. |

For example:

``` bash
inspect eval arc.py --model openai/gpt-4o-mini \ 
   -M responses_api=true
```

Or from Python:

``` python
from inspect_ai import eval

eval(
    "arc.py", model=" openai/gpt-4o-mini", 
    model_args= { "responses_api": True }
)
```

### Responses API

By default, Inspect uses the standard OpenAI Chat Completions API for GPT-4 models and the new [Responses API](https://platform.openai.com/docs/api-reference/responses) for GPT-5 and later (including GPT-6), o-series models, and the `computer_use_preview` model.

If you want to manually enable or disable the Responses API you can use the `responses_api` model argument. For example:

``` bash
inspect eval math.py --model openai/gpt-4o -M responses_api=true
```

Note that certain models including `o1-pro` and `computer_use_preview` *require* the use of the Responses API. Check the Open AI [models documentation](https://platform.openai.com/docs/models) for details on which models are supported by the respective APIs.

### Responses Phase

OpenAI Responses API assistant messages can include a [`phase`](https://developers.openai.com/api/docs/guides/reasoning#phase-parameter) label that distinguishes intermediate commentary from the final answer. Inspect preserves and replays `phase` values returned by OpenAI. To additionally synthesize missing `phase` values for assistant messages constructed outside the Responses API, use the `responses_phase` model argument:

``` bash
inspect eval math.py --model openai/gpt-5.4 -M responses_phase=true
```

When enabled, assistant messages with tool calls are labeled `commentary`; other assistant messages are labeled `final_answer`.

### Responses Store

By default, Inspect’s implementation of the Responses API does not store messages on the server. Reasoning content (which is intended to be opaque to clients) is handled using encrypted payloads (via the “reasoning.encrypted_content” include option). To control this behavior explicitly use the `responses_store` model argument. For example:

``` bash
inspect eval math.py --model openai/o4-mini -M responses_store=True
```

### Responses Metadata

You can attach [`metadata`](https://platform.openai.com/docs/api-reference/responses/create#responses-create-metadata) key-value pairs to Responses API requests via the `extra_body` generation config. For example:

``` bash
inspect eval math.py --model openai/gpt-5.4 \
    -M extra_body='{"metadata": {"experiment": "baseline"}}'
```

The `metadata` returned on responses (the echoed request metadata, which some models augment with additional fields) is surfaced as `ModelOutput.metadata`. Note that request metadata is not sent when `responses_store=True`.

### Flex Processing

[Flex processing](https://platform.openai.com/docs/guides/flex-processing) provides significantly lower costs for requests in exchange for slower response times and occasional resource unavailability (input and output tokens are priced using [batch API rates](https://platform.openai.com/docs/guides/batch) for flex requests).

Note that flex processing is in beta, and currently **only available for o3 and o4-mini models**.

To enable flex processing, use the `service_tier` model argument, setting it to “flex”. For example:

``` bash
inspect eval math.py --model openai/o4-mini -M service_tier=flex
```

OpenAI recommends using a [higher client timeout](https://platform.openai.com/docs/guides/flex-processing#api-request-timeouts) when making flex requests (15 minutes rather than the standard 10). Inspect automatically increases the client timeout to 15 minutes (900 seconds) for flex requests. To specify another value, use the `client_timeout` model argument. For example:

``` bash
inspect eval math.py --model openai/o4-mini \
    -M service_tier=flex -M client_timeout=1200
```

### OpenAI on Azure

The `openai` provider supports OpenAI models deployed on the [Azure AI Foundry](https://ai.azure.com/). To use OpenAI models on Azure AI, specify the following environment variables:

| Variable | Description |
|----|----|
| `AZUREAI_OPENAI_API_KEY` | API key credentials (optional, preferred name). |
| `AZURE_OPENAI_API_KEY` | API key credentials (optional, used as a fallback if `AZUREAI_OPENAI_API_KEY` is unset). |
| `AZUREAI_OPENAI_BASE_URL` | Base URL for requests (required) |
| `AZUREAI_OPENAI_API_VERSION` | OpenAI API version (optional) |
| `AZUREAI_AUDIENCE` | Azure resource URI that the access token is intended for when using managed identity (optional, defaults to `https://cognitiveservices.azure.com/.default`) |

You can then use the normal `openai` provider with the `azure` qualifier and the name of your model deployment (e.g. `gpt-4o-mini`). For example:

``` bash
export AZUREAI_OPENAI_API_KEY=your-api-key
export AZUREAI_OPENAI_BASE_URL=https://your-url-at.azure.com
export AZUREAI_OPENAI_API_VERSION=2025-03-01-preview
inspect eval math.py --model openai/azure/gpt-4o-mini
```

If using managed identity for authentication, install the `azure-identity` package and do not specify `AZUREAI_API_KEY`.

``` bash
pip install azure-identity
export AZUREAI_OPENAI_BASE_URL=https://your-url-at.azure.com
export AZUREAI_AUDIENCE=https://cognitiveservices.azure.com/.default
export AZUREAI_OPENAI_API_VERSION=2025-03-01-preview
inspect eval math.py --model openai/azure/gpt-4o-mini
```

Note that if the `AZUREAI_OPENAI_API_VERSION` is not specified, Inspect will generally default to the latest deployed version, which as of this writing is `2025-03-01-preview`. When using managed identity for authentication, install the `azure-identity` package and leave `AZUREAI_OPENAI_API_KEY` undefined.

### OpenAI on AWS Bedrock

The `openai` provider supports OpenAI models served through [Amazon Bedrock](https://aws.amazon.com/bedrock/). Use the normal `openai` provider with the `bedrock` qualifier, and use a standard OpenAI model identifier (Inspect automatically adds prefixes and suffixes required by Bedrock). For example:

``` bash
export AWS_BEARER_TOKEN_BEDROCK=your-bedrock-api-key
inspect eval arc.py --model openai/bedrock/gpt-5.5
```

You don’t need to set a region for the example above — Inspect defaults to `us-east-2`. The available model ids (e.g. `gpt-5.5`, `gpt-oss-120b`) and the regions they’re offered in vary over time and by account; the [`bedrock-mantle` documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-mantle.html) lists the supported regions.

#### Region

The AWS region is resolved with the following precedence: the `aws_region` model arg (`-M aws_region`), then `AWS_REGION`, then `AWS_DEFAULT_REGION`, and finally a default of `us-east-2`. Region availability varies by model (for example, at the time of writing `gpt-5.5` is only offered in `us-east-2`. Note that Bedrock API keys are region-bound (a key only works in the region it was created in).

If your environment sets a global `AWS_REGION` for other AWS services, you can target a specific region for this model only — without changing that global — using the model arg:

``` bash
inspect eval arc.py --model openai/bedrock/gpt-5.5 -M aws_region=us-east-2
```

#### Authentication

Authentication uses an AWS Bedrock bearer token. There are two ways to provide one:

| Variable | Description |
|----|----|
| `AWS_BEARER_TOKEN_BEDROCK` | Bedrock bearer API key — the AWS-standard name, as used in the AWS and OpenAI documentation. |
| `BEDROCK_OPENAI_API_KEY` | Bedrock bearer API key — Inspect-convention alias (takes precedence if both are set). |
| `BEDROCK_OPENAI_BASE_URL` | Custom endpoint override (optional; the AWS-standard `AWS_BEDROCK_BASE_URL` is also accepted). By default Inspect targets the region’s Mantle endpoint, choosing the model-appropriate path automatically (`/openai/v1` for frontier models like GPT-5 and later and Codex, `/v1` for open-weight models like gpt-oss). |
| `AWS_REGION` / `AWS_DEFAULT_REGION` | AWS region (optional; defaults to `us-east-2`). |

The first option is a static bearer key. Generate an [Amazon Bedrock API key](https://docs.aws.amazon.com/bedrock/latest/userguide/api-keys.html) and set it via `AWS_BEARER_TOKEN_BEDROCK` (the name used in the AWS and OpenAI docs; Inspect also accepts `BEDROCK_OPENAI_API_KEY`).

The second option uses your standard AWS credentials (IAM roles, instance profiles, SSO, AssumeRole, `AWS_ACCESS_KEY_ID`/`AWS_SECRET_ACCESS_KEY`, or a configured `AWS_PROFILE`). If no bearer key is set, Inspect generates short-lived bearer tokens from these credentials. This requires the `aws-bedrock-token-generator` package:

``` bash
pip install aws-bedrock-token-generator
export AWS_PROFILE=your-profile           # or any standard AWS credential source
inspect eval arc.py --model openai/bedrock/gpt-5.5
```

If the package is not installed and no bearer key is available, Inspect raises an error explaining how to proceed.

## Anthropic

To use the [Anthropic](https://www.anthropic.com/api) provider, install the `anthropic` package, set your credentials, and specify a model using the `--model` option:

``` bash
pip install anthropic
export ANTHROPIC_API_KEY=your-anthropic-api-key
inspect eval arc.py --model anthropic/claude-sonnet-4-0
```

For the `anthropic` provider, custom model args (`-M`) are forwarded to the constructor of the `AsyncAnthropic` class.

The following environment variables are supported by the Anthropic provider

| Variable | Description |
|----|----|
| `ANTHROPIC_API_KEY` | API key credentials (required). |
| `ANTHROPIC_BASE_URL` | Base URL for requests (optional, defaults to `https://api.anthropic.com`) |

### Betas

Some Anthropic features require that you include a beta identifier in the `betas` field of model requests. Inspect automatically includes the requisite identifier for beta features it utilizes (e.g. “mcp-client-2025-04-04”, “computer-use-2025-01-24”, etc.).

If there are other beta features you want to enable, use the `betas` model arg (`-M`). For example, to enable [1M token context windows](https://docs.anthropic.com/en/docs/build-with-claude/context-windows#1m-token-context-window) for Sonnet 4.5 and Opus 4.6 models:

``` bash
inspect eval arc.py --model anthropic/claude-sonnet-4-0 -M betas=context-1m-2025-08-07
```

### Refusal Fallback

> **NOTE:**
>
> The model fallback feature described below requires the development version of Inspect. You can install the development version from GitHub with:
>
> ``` bash
> pip install git+https://github.com/UKGovernmentBEIS/inspect_ai
> ```

Claude 5 classifiers can decline a request, returning a refusal (surfaced by Inspect as `stop_reason="content_filter"`). Such a request can usually be served by another Claude model. Set the `fallback_models` generate config to retry refused requests on one or more fallback models (tried in order) within the same request, using Anthropic’s [server-side fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback):

``` bash
inspect eval arc.py --model anthropic/claude-fable-5 --fallback-models claude-opus-4-8
```

Or via the [eval()](./reference/inspect_ai.html.md#eval) / [GenerateConfig](./reference/inspect_ai.model.html.md#generateconfig) API:

``` python
eval("arc.py", model="anthropic/claude-fable-5", fallback_models=["claude-opus-4-8"])
```

This is a feature of the first-party Anthropic API only — it is not supported on Bedrock, Vertex, or Azure, nor with [batch mode](./models-batch.html.md), and is ignored (with a warning) in those cases.

See the [Fallbacks](./fallbacks.html.md) article for complete documentation, including what gets recorded in logs, dataframes, and the viewer when fallbacks occur.

#### Cache Diagnostics

Include the `cache-diagnosis-2026-04-07` beta header to produce diagnostics for prompt caching. Diagnostics are automatically included in `ChatMessageAssistant.metadata["diagnostics"]` (which you can see in the viewer) and a warning message is printed for cache misses. For example:

``` bash
inspect eval arc.py --model anthropic/claude-sonnet-4-6 -M betas=cache-diagnosis-2026-04-07
```

Learn more about cache diagnostics at <https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics>.

### Cache TTL

Inspect enables Anthropic [prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) by default. Use the `cache_ttl` model arg (`-M`) to control the cache TTL. Valid values are:

- “auto” (the default): requests use the standard 5-minute cache TTL, and a sample is automatically switched to the 1-hour TTL for the remainder of its execution when more than 5 minutes pass between requests that use the cache (e.g. due to slow tool calls, long generations, or queuing behind `max_connections`). By that point the 5-minute cache has expired and the full prompt prefix is being rewritten regardless, so escalating protects the rest of the sample from repeated cache misses. Escalations are reported in a log message.

- “5m” or “1h”: pin the TTL for all requests (disables automatic escalation).

``` bash
inspect eval arc.py --model anthropic/claude-sonnet-4-6 -M cache_ttl=1h
```

Note that 1-hour cache writes are billed at 2x the base input token price (vs. 1.25x for 5-minute writes), so the longer TTL pays off only when requests sharing a prefix arrive more than 5 minutes apart. Automatic escalation applies only on the first-party Claude API (not Bedrock, Vertex, or Azure) and does not apply in [batch mode](./models-batch.html.md).

### Streaming

The Anthropic provider supports a `streaming` model arg (`-M`) that controls whether streaming responses are used. The default (“auto”) will automatically use streaming when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate), when thinking is enabled, or for potentially [long requests](https://github.com/anthropics/anthropic-sdk-python?tab=readme-ov-file#long-requests) (requests with \>= 8192 `max_tokens`). Pass `true` or `false` to override the default behavior:

``` bash
inspect eval arc.py --model anthropic/claude-sonnet-4-0 -M streaming=true
```

### Computer Use

When the [computer()](./tools-standard.html.md#sec-computer) tool is used with a Claude model, Inspect binds it to Anthropic’s native computer use support. On the Claude API and Vertex, Claude Opus 5.5 (which rejects the earlier tool there), Fable 5/5.1, and Mythos 5/5.1 use Anthropic’s [computer toolset](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool) (`computer_toolset_20260801`). Other models, including Opus 5, Sonnet 5, and Opus 4.8, keep the legacy `computer_20251124` tool. Anthropic offers the toolset only on the Claude API and Vertex, so on Bedrock and Foundry every model (Opus 5.5 and Fable/Mythos included) uses the legacy tool.

With the toolset, each computer action arrives as its own tool call. Inspect executes the actions in one assistant turn in order and stops at the first failure, reporting the remaining actions to the model as not executed. Logs record each action as a `computer` tool call with an `action` argument, exactly as with the legacy tool.

Use the `computer_toolset` model arg (`-M`) to force either mode where the platform supports it (`false` is rejected only for Opus 5.5 on the Claude API or Vertex, where the legacy tool returns an error). The toolset requires Claude Opus 4.8, Sonnet 5, Opus 5 or later on the Claude API or Vertex; for example, to use it with Opus 5:

``` bash
inspect eval computer.py --model anthropic/claude-opus-5 -M computer_toolset=true
```

### Anthropic on AWS Bedrock

To use Anthropic models on Bedrock, use the normal `anthropic` provider with the `bedrock` qualifier, specifying a model name that corresponds to a model you have access to on Bedrock. For Bedrock, authentication is not handled using an API key but rather your standard AWS credentials (e.g. `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY`). You should also be sure to have specified an AWS region. For example:

``` bash
export AWS_ACCESS_KEY_ID=your-aws-access-key-id
export AWS_SECRET_ACCESS_KEY=your-aws-secret-access-key
export AWS_DEFAULT_REGION=us-east-1
inspect eval arc.py --model anthropic/bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0
```

You can also optionally set the `ANTHROPIC_BEDROCK_BASE_URL` environment variable to set a custom base URL for Bedrock API requests.

### Anthropic on Vertex AI

To use Anthropic models on Vertex, you can use the standard `anthropic` model provider with the `vertex` qualifier (e.g. `anthropic/vertex/claude-3-5-sonnet-v2@20241022`). You should also set two environment variables indicating your project ID and region. Here is a complete example:

``` bash
export ANTHROPIC_VERTEX_PROJECT_ID=project-12345
export ANTHROPIC_VERTEX_REGION=us-east5
inspect eval ctf.py --model anthropic/vertex/claude-3-5-sonnet-v2@20241022
```

Authentication is doing using the standard Google Cloud CLI (i.e. if you have authorised the CLI then no additional auth is needed for the model API).

### Anthropic on Azure

The `anthropic` provider supports Anthropic models deployed on the [Azure AI Foundry](https://ai.azure.com/). To use Anthropic models on Azure AI, specify the following environment variables:

| Variable | Description |
|----|----|
| `AZUREAI_ANTHROPIC_API_KEY` | API key credentials (optional, preferred name). |
| `AZURE_ANTHROPIC_API_KEY` | API key credentials (optional, used as a fallback if `AZUREAI_ANTHROPIC_API_KEY` is unset). |
| `AZUREAI_ANTHROPIC_BASE_URL` | Base URL for requests (required). |

You can then use the normal `anthropic` provider with the `azure` qualifier and the name of your model deployment (e.g. `Claude-4-0-Sonnet-2411`). For example:

``` bash
export AZUREAI_ANTHROPIC_API_KEY=key
export AZUREAI_ANTHROPIC_BASE_URL=https://your-url-at.azure.com/models
inspect eval math.py --model anthropic/azure/Claude-4-0-Sonnet-2411
```

## Google

To use the [Google](https://ai.google.dev/) provider, install the `google-genai` package, set your credentials, and specify a model using the `--model` option:

``` bash
pip install google-genai
export GOOGLE_API_KEY=your-google-api-key
inspect eval arc.py --model google/gemini-2.5-pro
```

For the `google` provider, custom model args (`-M`) are forwarded to the `genai.Client` function. Google GenAI requests use a default SDK transport timeout of 1 hour when `timeout` is not configured; setting `timeout` applies the same value to each Google SDK request attempt and to Inspect’s overall retry budget.

The following environment variables are supported by the Google provider

| Variable | Description |
|----|----|
| `GOOGLE_API_KEY` | API key credentials (required unless using OAuth/ADC). |
| `GOOGLE_BASE_URL` | Base URL for requests (optional) |
| `GOOGLE_USE_ADC` | Set to `true` to authenticate Gemini Developer API models with OAuth/ADC by default (optional). |
| `GOOGLE_CLOUD_QUOTA_PROJECT` | Quota/billing project sent as `x-goog-user-project` when using OAuth/ADC (optional). |

### Gemini Developer API with OAuth / ADC

Some Gemini Developer API deployments (for example, partner-served models) are reachable only via an OAuth bearer token — Application Default Credentials (ADC) — plus a quota-project header, with no API key. Enable this **per model** (so a run can mix OAuth and API-key Google models) with `-M use_adc=true`:

``` bash
gcloud auth application-default login   # or an impersonated service account
inspect eval task.py --model google/your-model \
   -M use_adc=true -M quota_project_id=your-project
```

Alternatively, set `GOOGLE_USE_ADC=true` in the environment (e.g. in `.env`) to make OAuth the default for all Gemini Developer API models, so commands don’t need to differ between API-key and OAuth environments; `-M use_adc=false` overrides it per model, and Vertex models ignore it (Vertex uses ADC natively).

ADC covers all the standard credential sources: user credentials from `gcloud auth application-default login` (optionally with `--impersonate-service-account`), a service-account key or workload identity federation config via `GOOGLE_APPLICATION_CREDENTIALS`, and attached service accounts on GCP compute.

Supported custom model args (`-M`): `use_adc` (bool), `scopes` (list, defaults to `cloud-platform`), and `quota_project_id` (falls back to the `GOOGLE_CLOUD_QUOTA_PROJECT` environment variable). Auth precedence: `-M use_adc` (or, when unset, `GOOGLE_USE_ADC`) selects OAuth; otherwise an explicit API key, then `GOOGLE_API_KEY`, is used. Inspect refreshes the token as needed before each request, so long-running evals are supported. Batch inference is **not** supported in this mode (the batch client is long-lived and cannot refresh the token).

### Gemini on Vertex AI

To use Google Gemini models on Vertex, you can use the standard `google` model provider with the `vertex` qualifier (e.g. `google/vertex/gemini-2.0-flash`). You should also set two environment variables indicating your project ID and region. Here is a complete example:

``` bash
export GOOGLE_CLOUD_PROJECT=project-12345
export GOOGLE_CLOUD_LOCATION=us-east5
inspect eval ctf.py --model google/vertex/gemini-2.0-flash
```

You can alternatively pass the project and location as custom model args (`-M`). For example:

``` bash
inspect eval ctf.py --model google/vertex/gemini-2.0-flash \
   -M project=project-12345 -M location=us-east5
```

Authentication is done using the standard Google Cloud CLI. For example:

``` bash
gcloud auth application-default login
```

If you have authorised the CLI then no additional auth is needed for the model API. Alternatively, if you are running in [Vertex Express Mode](https://cloud.google.com/vertex-ai/generative-ai/docs/start/express-mode/overview), set `VERTEX_API_KEY` to authenticate with an Express Mode API key.

You can optionally specify a custom `GOOGLE_VERTEX_BASE_URL` to override the default base URL for Vertex.

### Safety Settings

Google models make available [safety settings](https://ai.google.dev/gemini-api/docs/safety-settings) that you can adjust to determine what sorts of requests will be handled (or refused) by the model. The five categories of safety settings are as follows:

| Category | Description |
|----|----|
| `civic_integrity` | Election-related queries. |
| `sexually_explicit` | Contains references to sexual acts or other lewd content. |
| `hate_speech` | Content that is rude, disrespectful, or profane. |
| `harassment` | Negative or harmful comments targeting identity and/or protected attributes. |
| `dangerous_content` | Promotes, facilitates, or encourages harmful acts. |

For each category, the following block thresholds are available:

| Block Threshold | Description |
|----|----|
| `none` | Always show regardless of probability of unsafe content |
| `only_high` | Block when high probability of unsafe content |
| `medium_and_above` | Block when medium or high probability of unsafe content |
| `low_and_above` | Block when low, medium or high probability of unsafe content |

By default, Inspect sets all four categories to `none` (enabling all content). You can override these defaults by using the `safety_settings` model argument. For example:

``` python
safety_settings = dict(
  dangerous_content = "medium_and_above",
  hate_speech = "low_and_above"
)
eval(
  "eval.py",
  model_args=dict(safety_settings=safety_settings)
)
```

This also can be done from the command line:

``` bash
inspect eval eval.py -M "safety_settings={'hate_speech': 'low_and_above'}"
```

### Streaming

The Google provider supports a `streaming` model arg (`-M`) that controls whether streaming responses are used. The default (“auto”) uses streaming when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate). Pass `true` to always stream (or `false` to never stream):

``` bash
inspect eval arc.py --model google/gemini-2.5-pro -M streaming=true
```

Streaming is particularly useful for Gemini 3+ models that support thinking/reasoning, as it enables proper capture of reasoning summaries from the streaming API.

## Mistral

To use the [Mistral](https://mistral.ai/) provider, install the `mistral` package, set your credentials, and specify a model using the `--model` option:

``` bash
pip install mistral
export MISTRAL_API_KEY=your-mistral-api-key
inspect eval arc.py --model mistral/mistral-large-latest
```

The following environment variables are supported by the Mistral provider

| Variable | Description |
|----|----|
| `MISTRAL_API_KEY` | API key credentials (required). |
| `MISTRAL_BASE_URL` | Base URL for requests (optional, defaults to `https://api.mistral.ai`) |

By default, the Mistral provider uses the [Conversation API](https://docs.mistral.ai/agents/agents#conversations), which includes features not available in the original completions API including native web search and code execution and support for document input. You can switch back to the completions API with the `conversation_api` custom model arg. For example:

``` bash
inspect eval arc.py --model mistral/mistral-large-latest -M conversation_api=false
```

Additional custom model args (`-M`) are forwarded to the constructor of the `Mistral` class.

### Streaming

The `mistral` provider supports a `streaming` model arg (`-M`) that controls whether streaming responses are used on the completions API path. The default (“auto”) uses streaming when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate) (declining requests that use a `response_schema`). Pass `true` to always stream (or `false` to never stream):

``` bash
inspect eval arc.py --model mistral/mistral-large-latest -M conversation_api=false -M streaming=true
```

Note that streaming applies only to the completions API — the default Conversation API path does not stream (an `on_stream` callback simply receives no events), so pass `-M conversation_api=false` to stream.

### Prompt Caching

When using the completions API (`-M conversation_api=false`, and the default for Voxtral models), Inspect sends Mistral’s [`prompt_cache_key`](https://docs.mistral.ai/studio-api/conversations/advanced/prompt-caching) set to the running sample’s unique id. Mistral uses it as a routing hint so a sample’s later turns reach the replica that cached its earlier ones. The Conversation API does not accept this parameter, so samples using it (the default) do not get the hint.

### Mistral on Azure AI

The `mistral` provider supports Mistral models deployed on the [Azure AI Foundry](https://ai.azure.com/). To use Mistral models on Azure AI, specify the following environment variables:

- `AZURE_MISTRAL_API_KEY`
- `AZUREAI_MISTRAL_BASE_URL`

You can then use the normal `mistral` provider with the `azure` qualifier and the name of your model deployment (e.g. `Mistral-Large-2411`). For example:

``` bash
export AZUREAI_MISTRAL_API_KEY=key
export AZUREAI_MISTRAL_BASE_URL=https://your-url-at.azure.com/models
inspect eval math.py --model mistral/azure/Mistral-Large-2411
```

## DeepSeek

To use the [DeepSeek](https://www.deepseek.com/) provider, install the `openai` package (which the DeepSeek service provides a compatible backend for), set your credentials, and specify a model using the `--model` option:

``` bash
pip install openai
export DEEPSEEK_API_KEY=your-deepseek-api-key
inspect eval arc.py --model deepseek/deepseek-flash
```

The current DeepSeek models are `deepseek-flash` (DeepSeek-V4.1-Flash, which also accepts image input) and `deepseek-v4-pro`. Both think by default. Use `--reasoning-effort` to control thinking (`none` disables it entirely) — see [Reasoning Effort](./reasoning.html.md#reasoning-effort) for details. While thinking is enabled the API rejects forced tool choice, so the `deepseek` provider submits forced tool choices as `"auto"` (disable thinking to force tool use).

The `deepseek-v4-flash` and `deepseek-v4-flash-vision-exp` model names were retired on September 10th, 2026: requests using them are served by V4.1 Flash (`deepseek-flash`) at Flash pricing, so logs dated after that point reflect V4.1 Flash. The legacy `deepseek-chat` and `deepseek-reasoner` model names were retired on July 24th, 2026.

The following environment variables are supported by the DeepSeek provider

| Variable | Description |
|----|----|
| `DEEPSEEK_API_KEY` | API key credentials (required). |
| `DEEPSEEK_BASE_URL` | Base URL for requests (optional, defaults to `https://api.deepseek.com`). |

## Moonshot AI

To use the [Moonshot AI](https://platform.moonshot.ai/) provider (Kimi models), install the `openai` package (which the Moonshot AI service provides a compatible backend for), set your credentials, and specify a model using the `--model` option:

``` bash
pip install openai
export MOONSHOT_API_KEY=your-moonshot-api-key
inspect eval arc.py --model moonshot/kimi-k3
```

Note that Kimi K3 uses fixed sampling (Moonshot recommends omitting sampling parameters), so the `moonshot` provider does not pass `temperature`, `top_p`, or penalty options to K3 models. Similarly, K3’s thinking effort currently only accepts `max`, so other `--reasoning-effort` values are submitted as `max`.

The following environment variables are supported by the Moonshot AI provider

| Variable | Description |
|----|----|
| `MOONSHOT_API_KEY` | API key credentials (required). |
| `MOONSHOT_BASE_URL` | Base URL for requests (optional, defaults to `https://api.moonshot.ai/v1`). |

## Meta

To use the [Meta](https://dev.meta.ai/docs) provider (Muse Spark models on the Meta Model API), install the `openai` package (which the Meta Model API provides a compatible backend for), set your credentials, and specify a model using the `--model` option:

``` bash
pip install openai
export META_API_KEY=your-meta-api-key
inspect eval arc.py --model meta/muse-spark-1.3
```

The provider uses the [Responses API](https://dev.meta.ai/docs/protocols/responses), which is the only Meta endpoint that preserves the model’s reasoning across turns (as encrypted reasoning items that are replayed automatically). Pass `-M responses_api=false` to use Chat Completions instead.

Muse Spark always reasons. Use `--reasoning-effort` (`minimal` through `max`) to control how much. `none` is not supported and is ignored with a warning. `max` is available on standard-tier `muse-spark-1.3` only; for the `-contributor` tier and older versions it is submitted as `xhigh`. The API only accepts `tool_choice="auto"`, so forced tool choices are submitted as `"auto"`. Log probabilities, stop sequences, and logit bias are not supported and are dropped with a warning.

Prompts blocked by Meta’s content policy are reported as `stop_reason="content_filter"` with `stop_details`. The API signals such a block three different ways depending on protocol and streaming (an HTTP 400 with a `content_policy_violation` code, a `content_filter` finish reason, or a `refusal` content part), and the provider normalizes all of them. Note that a refusal the model writes itself, rather than one the policy layer blocks, carries no signal at all and arrives as an ordinary completion. Score those from the completion text. Meta also restricts API keys after repeated policy violations (HTTP 403 with code `user_blocked`); while a key is restricted, streamed requests return a refusal for every prompt, so an eval where every sample stops with `content_filter` is more likely hitting that restriction than the content policy.

The following environment variables are supported by the Meta provider

| Variable | Description |
|----|----|
| `META_API_KEY` | API key credentials (required unless `MODEL_API_KEY` is set). |
| `MODEL_API_KEY` | API key credentials (Meta’s official variable name; used when `META_API_KEY` is not set). |
| `META_BASE_URL` | Base URL for requests (optional, defaults to `https://api.meta.ai/v1`). |

## Grok

To use the [Grok](https://x.ai/) provider, install the `openai` package (which the Grok service provides a compatible backend for), set your credentials, and specify a model using the `--model` option:

``` bash
pip install openai
export XAI_API_KEY=your-grok-api-key
inspect eval arc.py --model grok/grok-3-mini
```

The following environment variables are supported by the Grok provider. The provider reads its API key from `XAI_API_KEY` if set, otherwise from `GROK_API_KEY`; one of them must be defined.

| Variable | Description |
|----|----|
| `XAI_API_KEY` | API key credentials (preferred). |
| `GROK_API_KEY` | API key credentials (fallback if `XAI_API_KEY` is unset). |
| `XAI_BASE_URL` | Base URL for requests (optional, defaults to `api.x.ai`, note no “https://” prefix is used for the base url). |

### Model Args

The `grok` provider supports a `streaming` model argument that controls whether streaming responses are used. The default (“auto”) uses streaming when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate). Pass `true` to always stream (or `false` to never stream):

``` bash
inspect eval arc.py --model grok/grok-3-mini -M streaming=true
```

The `grok` provider also supports a `disable_retry` model argument that disables internal GRPC retries. For example:

``` bash
inspect eval arc.py --model grok/grok-3-mini -M disable_retry=true
```

This might be done if you are attempting to accurately track sample `working_time`—typically HTTP retries are subtracted from working time but the Grok provider uses GRPC which has no hooks available for requests and responses (while other providers do).

The `grok` provider also supports a `service_tier` model argument that selects the xAI processing tier for requests (introduced alongside Grok 4.6; requires `xai_sdk` \>= 1.17). For example, to use [Priority Processing](https://docs.x.ai/developers/grok-4-6) (billed at higher token rates):

``` bash
inspect eval arc.py --model grok/grok-4.6 -M service_tier=priority
```

Note that `service_tier` applies to standard requests only — [batch](./models-batch.html.md) requests are processed on xAI’s own batch tier, so the argument is omitted for them.

Additional custom model args (`-M`) are forwarded to the constructor of the `AsynClient` class.

### Prompt Caching

xAI caches prompt prefixes automatically, but its cache is per-server and requests are otherwise load balanced across servers, so a turn can land on a server that has never seen its prefix. Inspect therefore sends xAI’s [`x-grok-conv-id`](https://docs.x.ai/developers/advanced-api-usage/prompt-caching/maximizing-cache-hits) header, set to the running sample’s unique id, which pins all of that sample’s requests to a single server. Multi-turn agents benefit the most: each turn re-sends the whole conversation so far, and that prefix is already cached on the server the previous turn landed on.

Cached tokens are billed at a lower rate and are reported in the eval log as `input_tokens_cache_read`.

This requires no configuration. Note that the header is not sent for [batch](./models-batch.html.md) requests (which share a single long-lived client and are not multi-turn), or when the model is used outside of a running sample (where there is no conversation to pin).

## AWS Bedrock

To use the [AWS Bedrock](https://aws.amazon.com/bedrock/) provider, set your credentials and specify a model using the `--model` option:

``` bash
export AWS_ACCESS_KEY_ID=access-key-id
export AWS_SECRET_ACCESS_KEY=secret-access-key
export AWS_DEFAULT_REGION=us-east-1
inspect eval bedrock/meta.llama2-70b-chat-v1
```

For the `bedrock` provider, custom model args (`-M`) are forwarded to the `create_client` method of the `aiobotocore.session.AioSession` class, save for the `read_timeout` and `connect_timeout` args which are passed in the `config` parameter.

Note that all models on AWS Bedrock require that you [request model access](https://docs.aws.amazon.com/bedrock/latest/userguide/model-access.html) before using them in a deployment (in some cases access is granted immediately, in other cases it could one or more days).

You should be also sure that you have the appropriate AWS credentials before accessing models on Bedrock. You aren’t likely to need to, but you can also specify a custom base URL for AWS Bedrock using the `BEDROCK_BASE_URL` environment variable.

If you are using Anthropic models on Bedrock, you can alternatively use the [Anthropic provider](#anthropic-on-aws-bedrock) as your means of access. If you are using OpenAI models on Bedrock, access them via the `openai` provider with the `bedrock` qualifier — see [OpenAI on AWS Bedrock](#openai-on-aws-bedrock).

### Streaming

The `bedrock` provider supports a `streaming` model arg (`-M`) that controls whether the [ConverseStream](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_runtime_ConverseStream.html) API is used. The default (“auto”) uses streaming when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate) (declining requests that use a `response_schema`). Pass `true` to always stream (or `false` to never stream):

``` bash
inspect eval arc.py --model bedrock/us.anthropic.claude-sonnet-4-6 -M streaming=true
```

## AWS SageMaker

To use the [AWS SageMaker](https://aws.amazon.com/sagemaker/) provider, set your credentials and specify a SageMaker endpoint name using the `--model` option:

``` bash
export AWS_ACCESS_KEY_ID=access-key-id
export AWS_SECRET_ACCESS_KEY=secret-access-key
inspect eval arc.py --model sagemaker/my-endpoint-name \
  -M region_name=us-west-2
```

Deploy your preferred model via Sagemaker studio jumpstart UI/SDK/CLI ([link](https://docs.aws.amazon.com/sagemaker/latest/dg/deploy-jumpstart-model.html)). The model name after `sagemaker/` is the SageMaker endpoint name.

### Model Args

The following model args are supported:

| Model Arg | Description |
|----|----|
| `region_name` | AWS region where the endpoint is deployed (default: `us-east-1`). |
| `endpoint_url` | Custom SageMaker runtime endpoint URL (required). |
| `read_timeout` | Read timeout in seconds (default: `600`). |
| `connect_timeout` | Connection timeout in seconds (default: `60`). |
| `stream` | Enable streaming responses (default: stream only when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate)). |
| `completion_mode` | Send completions-style payloads for CPT/base models instead of chat-style payloads (default: `false`). |
| `inference_component_name` | Name of the inference component for multi-model endpoints. |
| `prompt_logprobs` | Number of prompt log probabilities to return per token. Used for perplexity scoring with vLLM-backed endpoints. |

For example:

``` bash
inspect eval arc.py --model sagemaker/my-endpoint \
  -M region_name=us-west-2 \
  -M read_timeout=300 \
  -M stream=true
```

### Inference Components

For [multi-model endpoints](https://docs.aws.amazon.com/sagemaker/latest/dg/multi-model-endpoints.html) that use inference components, specify the `inference_component_name` to route requests to a specific component:

``` bash
inspect eval arc.py --model sagemaker/my-endpoint \
  -M region_name=us-west-2 \
  -M inference_component_name=my-inference-component
```

### Completion Mode

For CPT (Continual Pre-Training) or base models that expect completions-style payloads (with a `prompt` field) rather than chat-style payloads (with a `messages` array), enable `completion_mode`:

``` bash
inspect eval arc.py --model sagemaker/my-cpt-endpoint \
  -M region_name=us-west-2 \
  -M completion_mode=true
```

Completion mode supports logprobs via the standard CLI flags:

``` bash
inspect eval arc.py --model sagemaker/my-cpt-endpoint \
  -M region_name=us-west-2 \
  -M completion_mode=true \
  --logprobs \
  --top-logprobs 5
```

> **NOTE: Note**
>
> Completion mode builds a plain text prompt from chat messages. Image content is not supported in this mode and will be ignored with a warning.

### Prompt Logprobs & Perplexity

The SageMaker provider supports prompt log probabilities and the [perplexity()](./reference/inspect_ai.scorer.html.md#perplexity) and [target_perplexity()](./reference/inspect_ai.scorer.html.md#target_perplexity) scorers when backed by a vLLM endpoint.

In **chat mode**, set `prompt_logprobs` via [GenerateConfig](./reference/inspect_ai.model.html.md#generateconfig) or the `-G` CLI flag:

``` python
from inspect_ai import Task, task
from inspect_ai.dataset import Sample
from inspect_ai.model import GenerateConfig
from inspect_ai.scorer import perplexity
from inspect_ai.solver import generate

@task
def perplexity_eval():
    return Task(
        dataset=[Sample(input="The capital of France is Paris")],
        solver=[generate(max_tokens=1)],
        scorer=perplexity(),
        config=GenerateConfig(prompt_logprobs=1),
    )
```

``` bash
inspect eval perplexity_eval.py --model sagemaker/my-endpoint \
  -M region_name=us-west-2
```

In **completion mode**, pass `prompt_logprobs` as a model argument:

``` bash
inspect eval perplexity_eval.py --model sagemaker/my-endpoint \
  -M region_name=us-west-2 \
  -M completion_mode=true \
  -M prompt_logprobs=1
```

> **NOTE: Note**
>
> The [target_perplexity()](./reference/inspect_ai.scorer.html.md#target_perplexity) scorer’s auto-tokenization feature is not available for SageMaker (the vLLM `/tokenize` endpoint is not reachable through `invoke_endpoint`). Provide `num_target_tokens` in sample metadata instead.

Authentication uses your standard AWS credentials (e.g. `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, or an IAM role). The endpoint must be accessible from your environment.

## Azure AI

The `azureai` provider supports models deployed on the [Azure AI Foundry](https://ai.azure.com/).

To use the `azureai` provider, install the `azure-ai-inference` package, set your credentials and base URL, and specify the name of the model you have deployed (e.g. `Llama-3.3-70B-Instruct`). For example:

``` bash
pip install azure-ai-inference
export AZUREAI_API_KEY=api-key
export AZUREAI_BASE_URL=https://your-url-at.azure.com/models
$ inspect eval math.py --model azureai/Llama-3.3-70B-Instruct
```

If using managed identity for authentication, install the `azure-identity` package and do not specify `AZUREAI_API_KEY`.

``` bash
pip install azure-identity
export AZUREAI_AUDIENCE=https://cognitiveservices.azure.com/.default
export AZUREAI_BASE_URL=https://your-url-at.azure.com/models
$ inspect eval math.py --model azureai/Llama-3.3-70B-Instruct
```

For the `azureai` provider, custom model args (`-M`) are forwarded to the constructor of the `ChatCompletionsClient` class.

The following environment variables are supported by the Azure AI provider

| Variable | Description |
|----|----|
| `AZURE_API_KEY` | API key credentials (optional, preferred name). |
| `AZUREAI_API_KEY` | API key credentials (optional, used as a fallback if `AZURE_API_KEY` is unset). |
| `AZUREAI_BASE_URL` | Base URL for requests (required) |
| `AZUREAI_AUDIENCE` | Azure resource URI that the access token is intended for when using managed identity (optional, defaults to `https://cognitiveservices.azure.com/.default`) |

If you are using Open AI or Mistral on Azure AI, you can alternatively use the [OpenAI provider](#openai-on-azure) or [Mistral provider](#mistral-on-azure-ai) as your means of access.

### Tool Emulation

When using the `azureai` model provider, tool calling support can be ‘emulated’ for models that Azure AI has not yet implemented tool calling for. This occurs by default for Llama models. For other models, use the `emulate_tools` model arg to force tool emulation:

``` bash
inspect eval ctf.py -M emulate_tools=true
```

You can also use this option to disable tool emulation for Llama models with `emulate_tools=false`.

### Streaming

The `azureai` provider supports a `streaming` model arg (`-M`) that controls whether streaming responses are used. The default (“auto”) uses streaming when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate). Pass `true` to always stream (or `false` to never stream):

``` bash
inspect eval arc.py --model azureai/Llama-3.3-70B-Instruct -M streaming=true
```

## Together AI

To use the [Together AI](https://www.together.ai/) provider, install the `openai` package (which the Together AI service provides a compatible backend for), set your credentials, and specify a model using the `--model` option:

``` bash
pip install openai
export TOGETHER_API_KEY=your-together-api-key
inspect eval arc.py --model together/MiniMaxAI/MiniMax-M2.7
```

For the `together` provider, you can enable [Tool Emulation](#tool-emulation-openai) using the `emulate_tools` custom model arg (`-M`). Other custom model args are forwarded to the constructor of the `AsyncOpenAI` class.

The `together` provider supports a `stream` model arg (`-M`) that controls whether streaming responses are used. The default (“auto”) uses streaming when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate). Pass `true` to always stream (or `false` to never stream):

``` bash
inspect eval arc.py --model together/MiniMaxAI/MiniMax-M2.7 -M stream=true
```

The following environment variables are supported by the Together AI provider

| Variable | Description |
|----|----|
| `TOGETHER_API_KEY` | API key credentials (required). |
| `TOGETHER_BASE_URL` | Base URL for requests (optional, defaults to `https://api.together.xyz/v1`) |

## Groq

To use the [Groq](https://groq.com/) provider, install the `groq` package, set your credentials, and specify a model using the `--model` option:

``` bash
pip install groq
export GROQ_API_KEY=your-groq-api-key
inspect eval arc.py --model groq/llama-3.1-70b-versatile
```

For the `groq` provider, custom model args (`-M`) are forwarded to the constructor of the `AsyncGroq` class.

The following environment variables are supported by the Groq provider

| Variable | Description |
|----|----|
| `GROQ_API_KEY` | API key credentials (required). |
| `GROQ_BASE_URL` | Base URL for requests (optional, defaults to `https://api.groq.com`) |

### Streaming

The `groq` provider supports a `streaming` model arg (`-M`) that controls whether streaming responses are used. The default (“auto”) uses streaming when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate) (declining requests that use a `response_schema`, as well as compound models, whose server-side executed tools are not carried by the stream). Pass `true` to always stream (or `false` to never stream):

``` bash
inspect eval arc.py --model groq/llama-3.3-70b-versatile -M streaming=true
```

## Fireworks AI

To use the [Fireworks AI](https://fireworks.ai/) provider, install the `openai` package (which the Fireworks AI service provides a compatible backend for), set your credentials, and specify a model using the `--model` option:

``` bash
pip install openai
export FIREWORKS_API_KEY=your-firewrks-api-key
inspect eval arc.py --model fireworks/accounts/fireworks/models/kimi-k3
```

For the `fireworks` provider, you can enable [Tool Emulation](#tool-emulation-openai) using the `emulate_tools` custom model arg (`-M`). Other custom model args are forwarded to the constructor of the `AsyncOpenAI` class.

### Prompt Caching

Fireworks caches prompt prefixes automatically, but a cache lives on a single replica and requests are otherwise load balanced across replicas, so a turn can land on a replica that has never seen its prefix. Inspect therefore sends Fireworks’ [`x-session-affinity`](https://docs.fireworks.ai/guides/prompt-caching) header, set to the running sample’s unique id, which pins all of that sample’s requests to a single replica. Multi-turn agents benefit the most: each turn re-sends the whole conversation so far, and that prefix is already cached on the replica the previous turn landed on.

Cached tokens are billed at a lower rate and are reported in the eval log as `input_tokens_cache_read`.

This requires no configuration, though you can override the header by setting your own value in `GenerateConfig(extra_headers={"x-session-affinity": ...})`. Note that the header is not sent when the model is used outside of a running sample (where there is no conversation to pin).

The following environment variables are supported by the Together AI provider

| Variable | Description |
|----|----|
| `FIREWORKS_API_KEY` | API key credentials (required). |
| `FIREWORKS_BASE_URL` | Base URL for requests (optional, defaults to `https://api.fireworks.ai/inference/v1`) |

## SambaNova

To use the [SambaNova](https://sambanova.ai/) provider, install the `openai` package (which the SambaNova service provides a compatible backend for), set your credentials, and specify a model using the `--model` option:

``` bash
pip install openai
export SAMBANOVA_API_KEY=your-sambanova-api-key
inspect eval arc.py --model sambanova/DeepSeek-V1-0324
```

For the `sambanova` provider, you can enable [Tool Emulation](#tool-emulation-openai) using the `emulate_tools` custom model arg (`-M`). Other custom model args are forwarded to the constructor of the `AsyncOpenAI` class.

The following environment variables are supported by the SambaNova provider

| Variable | Description |
|----|----|
| `SAMBANOVA_API_KEY` | API key credentials (required). |
| `SAMBANOVA_BASE_URL` | Base URL for requests (optional, defaults to `https://api.sambanova.ai/v1`) |

## Cloudflare

To use the [Cloudflare](https://developers.cloudflare.com/workers-ai/) provider, set your account id and access token, and specify a model using the `--model` option:

``` bash
export CLOUDFLARE_ACCOUNT_ID=account-id
export CLOUDFLARE_API_TOKEN=api-token
inspect eval arc.py --model cloudflare/@cf/meta/llama-3.1-70b-instruct
```

Specify the model id exactly as it appears in Cloudflare’s [model catalog](https://developers.cloudflare.com/workers-ai/models/): Workers AI model ids start with `@cf/`, while gateway-hosted models have plain ids (e.g. `cloudflare/moonshotai/kimi-k3`).

For the `cloudflare` provider, custom model args (`-M`) are included as fields in the post body of the chat request.

### Prompt Caching

Cloudflare reuses computed prefixes across requests, but only on the model instance that computed them. Inspect therefore sends Cloudflare’s [`x-session-affinity`](https://developers.cloudflare.com/workers-ai/features/prompt-caching/) header, set to the running sample’s unique id, which routes all of that sample’s requests to the same instance. This requires no configuration; override it by setting your own value in `GenerateConfig(extra_headers={"x-session-affinity": ...})`.

Note that Cloudflare’s OpenAI-compatible endpoint does not report cached token counts, so cache hits do not appear in the eval log.

The following environment variables are supported by the Cloudflare provider:

| Variable | Description |
|----|----|
| `CLOUDFLARE_ACCOUNT_ID` | Account id (required). |
| `CLOUDFLARE_API_TOKEN` | API key credentials (required). |
| `CLOUDFLARE_BASE_URL` | Base URL for requests (optional, defaults to `https://api.cloudflare.com/client/v4/accounts`) |

## Perplexity

To use the [Perplexity](https://www.perplexity.ai/) provider, install the `openai` package (if not already installed), set your credentials, and specify a model using the `--model` option:

``` bash
pip install openai
export PERPLEXITY_API_KEY=your-perplexity-api-key
inspect eval arc.py --model perplexity/sonar
```

The following environment variables are supported by the Perplexity provider

| Variable | Description |
|----|----|
| `PERPLEXITY_API_KEY` | API key credentials (required). |
| `PERPLEXITY_BASE_URL` | Base URL for requests (optional, defaults to `https://api.perplexity.ai`) |

Perplexity responses include citations when available. These are surfaced as [UrlCitation](./reference/inspect_ai.model.html.md#urlcitation)s attached to the assistant message. Additional usage metrics such as `reasoning_tokens` and `citation_tokens` are recorded in `ModelOutput.metadata`.

## Hugging Face

The [Hugging Face](https://huggingface.co/models) provider implements support for local models using the [transformers](https://pypi.org/project/transformers/) package. To use the Hugging Face provider, install the `torch`, `transformers`, and `accelerate` packages and specify a model using the `--model` option:

``` bash
pip install torch transformers accelerate
inspect eval arc.py --model hf/openai-community/gpt2
```

### Batching

Concurrency for REST API based models is managed using the `max_connections` option. The same option is used for `transformers` inference—up to `max_connections` calls to [generate()](./reference/inspect_ai.solver.html.md#generate) will be batched together (note that batches will proceed at a smaller size if no new calls to [generate()](./reference/inspect_ai.solver.html.md#generate) have occurred in the last 2 seconds).

The default batch size for Hugging Face is 32, but you should tune your `max_connections` to maximise performance and ensure that batches don’t exceed available GPU memory. The [Pipeline Batching](https://huggingface.co/docs/transformers/main_classes/pipelines#pipeline-batching) section of the transformers documentation is a helpful guide to the ways batch size and performance interact.

### Device

The PyTorch `cuda` device will be used automatically if CUDA is available (as will the Mac OS `mps` device). If you want to override the device used, use the `device` model argument. For example:

``` bash
$ inspect eval arc.py --model hf/openai-community/gpt2 -M device=cuda:0
```

This also works in calls to [eval()](./reference/inspect_ai.html.md#eval):

``` python
eval("arc.py", model="hf/openai-community/gpt2", model_args=dict(device="cuda:0"))
```

Or in a call to [get_model()](./reference/inspect_ai.model.html.md#get_model)

``` python
model = get_model("hf/openai-community/gpt2", device="cuda:0")
```

### Chat Templates

For Hugging Face models, Inspect will use a tokenizer chat template when available. Use the `chat_template` model arg to override the tokenizer template, and `use_chat_template=false` to bypass chat-template rendering entirely.

For example:

``` bash
inspect eval gsm8k.py --model hf/Qwen/Qwen3-1.7B-Base \
  -M "chat_template={% for message in messages %}{{ message.content }}{% endfor %}" \
  -M use_chat_template=true
```

Or to bypass templates:

``` bash
inspect eval gsm8k.py --model hf/Qwen/Qwen3-1.7B-Base -M use_chat_template=false
```

### Hidden States

If you wish to access hidden states (activations) from generation, use the `hidden_states` model arg. For example:

``` bash
$ inspect eval arc.py --model hf/openai-community/gpt2 -M hidden_states=true
```

Or from Python:

``` python
model = get_model(
    model="hf/meta-llama/Llama-3.1-8B-Instruct",
    hidden_states=True
)
```

Activations are available in the `"hidden_states"` field of `ModelOutput.metadata`. So they can be written to the eval log, the hidden-state tensors from transformers’ [GenerateDecoderOnlyOutput](https://huggingface.co/docs/transformers/main/en/internal/generation_utils#transformers.generation.GenerateDecoderOnlyOutput) are materialized as JSON-serializable nested lists, preserving the `[step][layer]` structure with the per-sample activation values in place of each tensor.

### Sampling

Pass the `do_sample` model arg to override the default sampling behavior (which is `do_sample=True`). For example:

``` bash
$ inspect eval arc.py --model hf/openai-community/gpt2 -M do_sample=false
```

### Trust Remote Code

Some Hugging Face models ship custom Python code in their repositories that the `transformers` library will execute on load when `trust_remote_code=True`. Because executing remote code is a security risk (the model author can run arbitrary code in your evaluation process), Inspect defaults `trust_remote_code` to `False` and will not forward `trust_remote_code` from generic `model_args`. To opt in for a specific model you trust, pass it explicitly:

``` bash
inspect eval arc.py --model hf/some-org/custom-arch-model -M trust_remote_code=true
```

Or from Python:

``` python
eval("arc.py", model="hf/some-org/custom-arch-model", model_args=dict(trust_remote_code=True))
```

The flag is applied to both the model and tokenizer `from_pretrained()` calls.

### Model Class

By default the Hugging Face provider loads models with `AutoModelForCausalLM`. Some architectures are not registered with that auto-class and must be loaded with a different one — for example the Mistral 3 series and other image-text-to-text models require `AutoModelForImageTextToText`. Use the `auto_model_class` model arg to name the `transformers` auto-class to use:

``` bash
inspect eval arc.py --model hf/mistralai/Ministral-3-8B-Instruct-2512 -M auto_model_class=AutoModelForImageTextToText
```

Or from Python:

``` python
eval(
    "arc.py",
    model="hf/mistralai/Ministral-3-8B-Instruct-2512",
    model_args=dict(auto_model_class="AutoModelForImageTextToText"),
)
```

The value must be the name of a class exported by `transformers`.

### Local Models

In addition to using models from the Hugging Face Hub, the Hugging Face provider can also use local model weights and tokenizers (e.g. for a locally fine tuned model). Use `hf/local` along with the `model_path`, and (optionally) `tokenizer_path` arguments to select a local model. For example, from the command line, use the `-M` flag to pass the model arguments:

``` bash
$ inspect eval arc.py --model hf/local -M model_path=./my-model
```

Or using the [eval()](./reference/inspect_ai.html.md#eval) function:

``` python
eval("arc.py", model="hf/local", model_args=dict(model_path="./my-model"))
```

Or in a call to [get_model()](./reference/inspect_ai.model.html.md#get_model)

``` python
model = get_model("hf/local", model_path="./my-model")
```

## vLLM

The [vLLM](https://docs.vllm.ai/) provider also implements support for Hugging Face models using the [vllm](https://github.com/vllm-project/vllm/) package. To use the vLLM provider, install the `vllm` package and specify a model using the `--model` option:

``` bash
pip install vllm
inspect eval arc.py --model vllm/openai-community/gpt2
```

For the `vllm` provider, custom model args (-M) are forwarded to the vllm [CLI](https://docs.vllm.ai/en/stable/serving/openai_compatible_server.html#cli-reference). Top-level model arg names are converted to CLI flag form (for example, `tensor_parallel_size` becomes `--tensor-parallel-size`). Dotted vLLM arguments preserve nested field names after the dot, so `-M speculative-config.num_speculative_tokens=1` is forwarded as `--speculative-config.num_speculative_tokens 1`.

The following environment variables are supported by the vLLM provider:

| Variable | Description |
|----|----|
| `VLLM_BASE_URL` | Base URL for requests (optional, defaults to the server started by Inspect) |
| `VLLM_API_KEY` | API key for the vLLM server (optional, defaults to “local”) |
| `VLLM_DEFAULT_SERVER_ARGS` | JSON string of default server args (e.g., ‘{“tensor_parallel_size”: 4, “max_model_len”: 8192}’) |

You can also access models from ModelScope rather than Hugging Face, see the [vLLM documentation](https://docs.vllm.ai/en/stable/getting_started/quickstart.html) for details on this.

vLLM is generally much faster than the Hugging Face provider as the library is designed entirely for inference speed whereas the Hugging Face library is more general purpose.

### Multiple Servers

`VLLM_BASE_URL` sets a single global endpoint, but a vLLM server only serves one model. If you need different models for different purposes — for example, a small model for the solver and a larger one as a judge for [model-graded scoring](./model-graded.html.md) — start a vLLM server per model and pass a per-model `base_url` rather than relying on the env var.

The most ergonomic path is [model roles](./models.html.md#model-roles), which lets the built-in `model_graded_*` scorers automatically resolve their judge from the `grader` role:

``` bash
inspect eval task.py \
    --model vllm/meta-llama/Llama-3-8B \
    --model-base-url http://gpu1:8000/v1 \
    --model-role 'grader={model: vllm/meta-llama/Llama-3-70B-Instruct, model_args: {base_url: http://gpu2:8000/v1}}'
```

Equivalent from Python:

``` python
from inspect_ai import eval
from inspect_ai.model import get_model

eval(
    "task.py",
    model=get_model("vllm/meta-llama/Llama-3-8B", base_url="http://gpu1:8000/v1"),
    model_roles={
        "grader": get_model(
            "vllm/meta-llama/Llama-3-70B-Instruct",
            base_url="http://gpu2:8000/v1",
        ),
    },
)
```

Any number of roles can be defined this way (e.g. a separate `critic` or `red_team` model); each one can point at its own vLLM server. In CLI role dictionaries, place provider-specific values such as `base_url` and `api_key` under `model_args`; in Python, pass them directly to [get_model()](./reference/inspect_ai.model.html.md#get_model).

Note: Inspect reuses a single server entry per base model name, so two `vllm/<same-model>` instances pointed at different URLs will collapse to the first URL. This caveat does not apply to the typical solver-vs-judge setup since the two models are different.

### Batching

vLLM automatically handles batching, so you generally don’t have to worry about selecting the optimal batch size. However, you can still use the `max_connections` option to control the number of concurrent requests which defaults to 32. If the server has saturated the GPU it may reject requests—these are by default retried after 5 seconds (you can customize this using the `retry_delay` model args, e.g. `-M retry_delay=3`).

### Device

The `device` option is also available for vLLM models, and you can use it to specify the device(s) to run the model on. For example:

``` bash
$ inspect eval arc.py --model vllm/meta-llama/Meta-Llama-3-8B-Instruct -M device='0,1,2,3'
```

### Local Models

Similar to the Hugging Face provider, you can also use local models with the vLLM provider. Use `vllm/local` along with the `model_path`, and (optionally) `tokenizer_path` arguments to select a local model. For example, from the command line, use the `-M` flag to pass the model arguments:

``` bash
$ inspect eval arc.py --model vllm/local -M model_path=./my-model
```

### LoRA Adapters

vLLM supports [LoRA (Low-Rank Adaptation)](https://docs.vllm.ai/en/stable/features/lora.html) adapters, allowing you to use fine-tuned models without duplicating the base model weights. To use a LoRA adapter, append `:adapter-path` to the model name:

``` bash
inspect eval arc.py --model vllm/meta-llama/Llama-3-8B:myorg/my-lora-adapter
```

The adapter path can be a HuggingFace repository (e.g., `myorg/my-lora-adapter`) or a local path (e.g., `./adapters/my-adapter`).

When using LoRA adapters:

- The vLLM server is automatically started with `--enable-lora`
- `max_lora_rank` is auto-detected from the adapter’s `adapter_config.json` (supports both local paths and HuggingFace repos)
- Adapters are dynamically loaded on first request via vLLM’s `/v1/load_lora_adapter` endpoint
- Multiple models sharing the same base model reuse a single vLLM server, even with different adapters

For example, you can evaluate multiple LoRA fine-tunes on the same base model efficiently:

``` python
# These will share the same vLLM server
# max_lora_rank is auto-detected as the max across all adapters
eval(
    "task.py",
    model=["vllm/meta-llama/Llama-3-8B:adapter-a", "vllm/meta-llama/Llama-3-8B:adapter-b"],
)
```

You can also compare a base model against its LoRA fine-tune — LoRA will be auto-enabled for the shared server:

``` python
eval(
    "task.py",
    model=["vllm/meta-llama/Llama-3-8B", "vllm/meta-llama/Llama-3-8B:my-adapter"],
)
```

If you need to override the auto-detected rank (e.g. when using [get_model()](./reference/inspect_ai.model.html.md#get_model) directly with multiple adapters of different ranks), pass `max_lora_rank` explicitly:

``` bash
inspect eval task.py --model vllm/meta-llama/Llama-3-8B:my-adapter -M max_lora_rank=128
```

#### External vLLM Server with LoRA

When using an external vLLM server (`VLLM_BASE_URL`), you have two options:

**Option 1: Pre-load adapters manually**

Load the adapters yourself when starting the server and reference them by name:

``` bash
# Start server with pre-loaded adapter
vllm serve meta-llama/Llama-3-8B --enable-lora \
    --lora-modules my-adapter=path/to/adapter

# Use the adapter name directly (not the path)
inspect eval arc.py --model vllm/meta-llama/Llama-3-8B:my-adapter
```

**Option 2: Enable dynamic loading**

Start the server with `VLLM_ALLOW_RUNTIME_LORA_UPDATING=True` to let Inspect load adapters dynamically:

``` bash
VLLM_ALLOW_RUNTIME_LORA_UPDATING=True vllm serve meta-llama/Llama-3-8B --enable-lora
```

Then use adapter paths as normal:

``` bash
inspect eval arc.py --model vllm/meta-llama/Llama-3-8B:myorg/my-lora-adapter
```

Note: When Inspect starts the vLLM server itself, it automatically sets `VLLM_ALLOW_RUNTIME_LORA_UPDATING=True`.

### Chat Templates

For vLLM models, the `chat_template` model arg is forwarded to the vLLM server’s `--chat-template` flag. Use `use_chat_template=false` to bypass chat-template rendering entirely (useful for base models):

``` bash
inspect eval gsm8k.py --model vllm/Qwen/Qwen3-1.7B-Base -M use_chat_template=false
```

> **NOTE: Note**
>
> `use_chat_template` only takes effect when Inspect starts the vLLM server. When connecting to an existing server via `VLLM_BASE_URL`, set `--chat-template` when starting the server instead.

### Raw Text Completions

Use the `vllm-completions` provider when you want vLLM to receive a raw text prompt rather than chat messages rendered through a chat template:

``` bash
inspect eval task.py --model vllm-completions/EleutherAI/pythia-70m
```

This provider uses vLLM’s `/v1/completions` endpoint. It accepts a single user message, sends that message content as the raw prompt, and is useful for base-model generation and log-probability based evaluations. It additionally manages the vLLM server lifecycle for you — to target an already-running OpenAI-compatible server, the [`openai-api-completions`](#openai-api-completions) provider offers the same behavior for any provider/server.

#### Pre-Tokenized Prompts

If you already have token IDs (custom tokenizer, pre-tokenized dataset, anything where you need exact control over the input sequence), pass them through `ChatMessage.metadata["prompt_token_ids"]` instead of a string. vLLM uses the IDs verbatim and skips re-tokenization, so you avoid an `ids → str → ids` round-trip that can change the sequence for non-bijective tokenizers.

``` python
from inspect_ai.model import ChatMessageUser, get_model

token_ids = my_custom_tokenizer.encode("Hello")
model = get_model("vllm-completions/EleutherAI/pythia-70m")
response = await model.generate(
    input=[ChatMessageUser(content="", metadata={"prompt_token_ids": token_ids})]
)
```

When `prompt_token_ids` is present, only the IDs are sent to vLLM — the message’s `content` is not used as the prompt. The `content` field is still part of the [ChatMessage](./reference/inspect_ai.model.html.md#chatmessage) though, so it shows up in transcripts and is readable by scorers/judges. A common pattern is to put the decoded (or any human-readable) version of the prompt in `content` so downstream tooling has something useful to display:

``` python
ChatMessageUser(
    content="Hello",  # what judges/transcripts see
    metadata={"prompt_token_ids": token_ids},  # what the model actually receives
)
```

vLLM only applies `add_special_tokens` when tokenizing a string prompt, so for `list[int]` prompts the IDs go through as-is regardless of that flag.

### Tool Use and Reasoning

vLLM supports tool use and reasoning; however, the usage is often model dependent and requires additional configuration. See the [Tool Use](https://docs.vllm.ai/en/stable/features/tool_calling.html) and [Reasoning](https://docs.vllm.ai/en/stable/features/reasoning_outputs.html) sections of the vLLM documentation for details.

For vLLM reasoning models, pass the model-specific parser and chat-template kwargs through `-M`. See [Reasoning](./reasoning.html.md#vllm-sglang) for CLI examples.

### Prompt Log Probabilities

vLLM supports returning log probabilities for prompt tokens via the `prompt_logprobs` configuration option. This enables [perplexity-based scoring](./perplexity.html.md) for benchmarks like WikiText, C4, ARC-C, and MMLU:

``` bash
inspect eval perplexity_eval.py --model vllm/meta-llama/Meta-Llama-3-8B \
  --prompt-logprobs 1
```

Or in Python:

``` python
Task(
    dataset=dataset,
    solver=generate(max_tokens=1, prompt_logprobs=1),
    scorer=perplexity(),
)
```

> **NOTE: Note**
>
> Prompt log probabilities are not available when streaming is enabled. Ensure streaming is disabled when using perplexity scorers.

### vLLM Server

Rather than letting Inspect start and stop a vLLM server every time you run an evaluation (which can take several minutes for large models), you can instead start the server manually and then connect to it. To do this, set the model base URL to point to the vLLM server and the API key to the server’s API key. For example:

``` bash
$ export VLLM_BASE_URL=http://localhost:8080/v1
$ export VLLM_API_KEY=<your-server-api-key>
$ inspect eval arc.py --model vllm/meta-llama/Meta-Llama-3-8B-Instruct
```

or

``` bash
$ inspect eval arc.py --model vllm/meta-llama/Meta-Llama-3-8B-Instruct --model-base-url http://localhost:8080/v1 -M api_key=<your-server-api-key>
```

See the vLLM documentation on [Server Mode](https://docs.vllm.ai/en/stable/serving/openai_compatible_server.html) for additional details.

## SGLang

To use the [SGLang](https://docs.sglang.ai/index.html) provider, install the `sglang` package and specify a model using the `--model` option:

``` bash
pip install "sglang[all]>=0.4.4.post2" --find-links https://flashinfer.ai/whl/cu124/torch2.5/flashinfer-python
inspect eval arc.py --model sglang/meta-llama/Meta-Llama-3-8B-Instruct
```

For the `sglang` provider, custom model args (-M) are forwarded to the sglang [CLI](https://docs.sglang.ai/backend/server_arguments.html).

The following environment variables are supported by the SGLang provider:

| Variable | Description |
|----|----|
| `SGLANG_BASE_URL` | Base URL for requests (optional, defaults to the server started by Inspect) |
| `SGLANG_API_KEY` | API key for the SGLang server (optional, defaults to “local”) |
| `SGLANG_DEFAULT_SERVER_ARGS` | JSON string of default server args (e.g., ‘{“tp”: 4, “max_model_len”: 8192}’) |

SGLang is a fast and efficient language model server that supports a variety of model architectures and configurations. Its usage in Inspect is almost identical to the [vLLM provider](#vllm). You can either let Inspect start and stop the server for you, or start the server manually and then connect to it:

``` bash
$ export SGLANG_BASE_URL=http://localhost:8080/v1
$ export SGLANG_API_KEY=<your-server-api-key>
$ inspect eval arc.py --model sglang/meta-llama/Meta-Llama-3-8B-Instruct
```

or

``` bash
$ inspect eval arc.py --model sglang/meta-llama/Meta-Llama-3-8B-Instruct --model-base-url http://localhost:8080/v1 -M api_key=<your-server-api-key>
```

### Tool Use and Reasoning

SGLang supports tool use and reasoning; however, the usage is often model dependent and requires additional configuration. See the [Tool Use](https://docs.sglang.ai/backend/function_calling.html) and [Reasoning](https://docs.sglang.ai/backend/separate_reasoning.html) sections of the SGLang documentation for details.

### Batching

SGLang automatically handles batching, so you generally don’t have to worry about selecting the optimal batch size. However, you can still use the `max_connections` option to control the number of concurrent requests which defaults to 32. If the server has saturated the GPU it may reject requests—these are by default retried after 5 seconds (you can customize this using the `retry_delay` model args, e.g. `-M retry_delay=3`).

## nnterp

The [nnterp](https://ndif-team.github.io/nnterp/index.html) provider enables you to use `StandardizedTransformer` models with Inspect. To use the nnterp provider, install the `nnterp` package:

``` bash
pip install nnterp
```

The `nnterp` provider works with Hugging Face models. For example:

``` bash
inspect eval arc.py --model nnterp/openai-community/gpt2
```

The `nnterp` provider supports the following custom model args (other model args are forwarded to the constructor of the `StandardizedTransformer` class):

| Model Arg | Description | Default |
|----|----|----|
| `dispatch` | Immediately load model into memory at initialization time | True |
| `device_map` | Model device map. | “auto” |
| `dtype` | Torch data type | float16 |
| `hidden_states` | Provide hidden states in `ModelOutput.metadata` | False |

For example:

``` bash
inspect eval arc.py \
    --model nnterp/openai-community/gpt2 \
    -M device_map=0 \
    -M hidden_states=true
```

Or from Python:

``` python
eval(
    task=arc(), 
    model="nnterp/openai-community/gpt2", 
    model_args={"device_map": 0, "hidden_states": True}
)
```

## TransformerLens

The [TransformerLens](https://github.com/neelnanda-io/TransformerLens) provider allows you to use `HookedTransformer` models with Inspect.

To use the TransformerLens provider, install the `transformer_lens` package:

``` bash
pip install transformer_lens
```

### Usage with Pre-loaded Models

Unlike other providers, TransformerLens requires you to first load a `HookedTransformer` model instance and then pass it to Inspect. This is because TransformerLens models expose special hooks for accessing and manipulating internal activations that need to be set up before use in the inspect framework.

You will need to specify the `tl_model` and `tl_generate_args` in the model arguments. The `tl_model` is the `HookedTransformer` instance and the `tl_generate_args` is a dictionary of transformer-lens generation arguments. You can specify the model name as anything, it will not affect the model you are using.

Here’s an example:

``` python
# Create a HookedTransformer model and set up all the hooks
tl_model = HookedTransformer(...)
...

# Create model args with the TransformerLens model and generation parameters
model_args = {
    "tl_model": tl_model,
    "tl_generate_args": {
        "max_new_tokens": 50,
        "temperature": 0.7,
        "do_sample": True,
    }
}

# Use with get_model()
model = get_model("transformer_lens/your-model-name", **model_args)

# Or use directly in eval()
eval("arc.py", model="transformer_lens/your-model-name", model_args=model_args)
```

### Limitations

1.  Please note that tool calling is not yet supported for TransformerLens models.
2.  Since the model is loaded dynamically, it is not possible to use cli arguments to specify the model.

## Ollama

To use the [Ollama](https://ollama.com/) provider, install the `openai` package (which Ollama provides a compatible backend for) and specify a model using the `--model` option:

``` bash
pip install openai
inspect eval arc.py --model ollama/llama3.1
```

Note that you should be sure that Ollama is running on your system before using it with Inspect.

You can enable [Tool Emulation](#tool-emulation-openai) for Ollama models using the `emulate_tools` custom model arg (`-M`).

The following environment variables are supported by the Ollma provider

| Variable | Description |
|----|----|
| `OLLAMA_BASE_URL` | Base URL for requests (optional, defaults to `http://localhost:11434/v1`) |

## Llama-cpp-python

To use the [Llama-cpp-python](https://llama-cpp-python.readthedocs.io/en/latest/) provider, install the `openai` package (which llama-cpp-python provides a compatible backend for) and specify a model using the `--model` option:

``` bash
pip install openai
inspect eval arc.py --model llama-cpp-python/llama3
```

Note that you should be sure that the [llama-cpp-python server](https://llama-cpp-python.readthedocs.io/en/latest/server/) is running on your system before using it with Inspect.

The following environment variables are supported by the llama-cpp-python provider

| Variable | Description |
|----|----|
| `LLAMA_CPP_PYTHON_BASE_URL` | Base URL for requests (optional, defaults to `http://localhost:8000/v1`) |

## OpenAI Compatible

If your model provider makes an OpenAI API compatible endpoint available, you can use it with Inspect via the `openai-api` provider, which uses the following model naming convention:

    openai-api/<provider-name>/<model-name>

Inspect will read environment variables corresponding to the api key and base url of your provider using the following convention (note that the provider name is capitalized):

    <PROVIDER_NAME>_API_KEY
    <PROVIDER_NAME>_BASE_URL

Note that hyphens within provider names will be converted to underscores so they conform to requirements of environment variable names. For example, if the provider is named `awesome-models` then the API key environment variable should be `AWESOME_MODELS_API_KEY`.

### Example

Here is how you would access DeepSeek using the `openai-api` provider:

``` bash
export DEEPSEEK_API_KEY=your-deepseek-api-key
export DEEPSEEK_BASE_URL=https://api.deepseek.com
inspect eval arc.py --model openai-api/deepseek/deepseek-flash
```

### Responses API

You can enable the use of the Responses API with the `openai-api` provider by passing the `responses_api` model arg. For example:

``` bash
$ inspect eval arc.py --model openai-api/<provider>/<model> -M responses_api=true
```

Or using the [eval()](./reference/inspect_ai.html.md#eval) function:

``` python
eval("arc.py", model="openai-api/<provider>/<model>", model_args=dict(responses_api=True))
```

When using the Responses API, `openai-api` also supports the `responses_phase` model arg to synthesize missing assistant message `phase` values when replaying Responses API histories.

### Tool Emulation

When using OpenAI compatible model providers, tool calling support can be ‘emulated’ for models that don’t yet support it. Use the `emulate_tools` model arg to force tool emulation:

``` bash
inspect eval ctf.py --model openai-api/<provider>/<model> -M emulate_tools=true
```

Tool calling emulation works by encoding tool JSON schema in an XML tag and asking the model to make tool calls using another XML tag. This works with varying degrees of efficacy depending on the model and the complexity of the tool schema. Before using tool emulation you should always check if your provider implements native support for tool calling on the model you are using, as that will generally work better.

### Strict Tool Schemas

By default, Inspect sets `"strict": true` on tool function schemas for the `openai-api` provider. This preserves compatibility with providers that require strict tool schemas. You can override this using the `strict_tools` model arg:

``` bash
inspect eval arc.py --model openai-api/<provider>/<model> -M strict_tools=false
```

Or using the [eval()](./reference/inspect_ai.html.md#eval) function:

``` python
eval("arc.py", model="openai-api/<provider>/<model>", model_args=dict(strict_tools=False))
```

### Streaming

The `openai-api` provider streams by default only when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate). You can force streaming on (or off) by passing the `stream` model arg. For example:

``` bash
$ inspect eval arc.py --model openai-api/<provider>/<model> -M stream=true
```

### Completions API

Use the `openai-api-completions` provider when you want the model to receive a raw prompt through the legacy `/v1/completions` endpoint rather than chat messages rendered through a chat template:

``` bash
inspect eval task.py --model openai-api-completions/<provider>/<model>
```

It follows the same naming and environment variable conventions as `openai-api`, accepts a single user message, sends that message content as the raw prompt, and is useful for base-model generation and log-probability based evaluations (echo mode, fill-in-the-middle, perplexity benchmarks).

#### Pre-Tokenized Prompts

If you already have token IDs (custom tokenizer, pre-tokenized dataset, anything where you need exact control over the input sequence), pass them through `ChatMessage.metadata["prompt_token_ids"]` instead of a string:

``` python
from inspect_ai.model import ChatMessageUser, get_model

token_ids = my_custom_tokenizer.encode("Hello")
model = get_model("openai-api-completions/<provider>/<model>")
response = await model.generate(
    input=[ChatMessageUser(content="", metadata={"prompt_token_ids": token_ids})]
)
```

When `prompt_token_ids` is present, only the IDs are sent to the server — the message’s `content` is not used as the prompt. The `content` field is still part of the [ChatMessage](./reference/inspect_ai.model.html.md#chatmessage) though, so it shows up in transcripts and is readable by scorers/judges. A common pattern is to put the decoded (or any human-readable) version of the prompt in `content` so downstream tooling has something useful to display.

Token IDs are passed through verbatim; whether the server adds special tokens to pre-tokenized prompts is server-dependent (vLLM does not — see [vLLM Completions API](#sec-vllm-completions)).

## OpenRouter

To use the [OpenRouter](https://openrouter.ai/) provider, install the `openai` package (which the OpenRouter service provides a compatible backend for), set your credentials, and specify a model using the `--model` option:

``` bash
pip install openai
export OPENROUTER_API_KEY=your-openrouter-api-key
inspect eval arc.py --model openrouter/gryphe/mythomax-l2-13b
```

For the `openrouter` provider, the following custom model args (`-M`) are supported (click the argument name to see its docs on the OpenRouter site):

| Argument | Example |
|----|----|
| [`models`](https://openrouter.ai/docs/features/model-routing#the-models-parameter) | `-M "models=anthropic/claude-3.5-sonnet, gryphe/mythomax-l2-13b"` |
| [`provider`](https://openrouter.ai/docs/features/provider-routing) | `-M "provider={ 'quantizations': ['int8'] }"` |
| [`transforms`](https://openrouter.ai/docs/features/message-transforms) | `-M "transforms=['middle-out']"` |
| [`reasoning_enabled`](https://openrouter.ai/docs/use-cases/reasoning-tokens) | `-M "reasoning_enabled=false"` |

In addition, [Tool Emulation](#tool-emulation-openai) is available for models that don’t yet support tool calling in their API.

For `openrouter/anthropic/*` models, Anthropic [prompt caching](https://docs.claude.com/en/docs/build-with-claude/prompt-caching) is enabled by default: per-block `cache_control` markers are inserted on the last system block, the last tool definition, and a rolling pair of message-level breakpoints (mirroring the placement used by the direct `anthropic` provider). The markers are accepted by OpenRouter across Anthropic-direct, Bedrock, and Vertex routing. Cache writes returned upstream are surfaced as `ModelUsage.input_tokens_cache_write`. Pass `--cache-prompt=false` (or set `cache_prompt=False` in [GenerateConfig](./reference/inspect_ai.model.html.md#generateconfig)) to disable. Single-turn evaluations that never re-issue the same prefix pay a small premium (Anthropic charges ~1.25× for cache writes) with no offsetting cache reads — disable caching for those workloads.

Note that OpenRouter may distribute consecutive requests for the same model across multiple backends (for Anthropic models: Anthropic-direct, Bedrock, Vertex), and each backend maintains its own prompt cache. To keep a sample on one backend, Inspect sends OpenRouter’s [`x-session-id`](https://openrouter.ai/docs/guides/best-practices/prompt-caching) header set to the running sample’s unique id: OpenRouter uses it as the sticky routing key, so every turn of a sample returns to the backend holding its warm cache, from the first request rather than only after a cache hit is observed. Override it by setting your own value in `GenerateConfig(extra_headers={"x-session-id": ...})`. For stricter control you can still pin routing to a single backend via the `provider` model-arg, for example `-M provider='{"order":["anthropic"],"allow_fallbacks":false}'`.

The `cache_control` markers are injected just before the request reaches OpenRouter and so will not appear in the request snapshot recorded in `.eval` log files. Verify caching is active by inspecting the usage line (cache reads/writes) on returned [ModelOutput](./reference/inspect_ai.model.html.md#modeloutput)s.

The following environment variables are supported by the OpenRouter AI provider

| Variable | Description |
|----|----|
| `OPENROUTER_API_KEY` | API key credentials (required). |
| `OPENROUTER_BASE_URL` | Base URL for requests (optional, defaults to `https://openrouter.ai/api/v1`) |

## LiteLLM Proxy

To use models served by a [LiteLLM proxy](https://docs.litellm.ai/docs/simple_proxy), install the `openai` package, set the proxy’s URL and your key, and specify a model using the `model_name` of its deployment in the proxy config. For example, for this deployment:

``` yaml
model_list:
  - model_name: claude-sonnet-5
    litellm_params:
      model: anthropic/claude-sonnet-5
      api_key: os.environ/ANTHROPIC_API_KEY
```

``` bash
pip install openai
export LITELLM_PROXY_BASE_URL=https://litellm.example.com
export LITELLM_PROXY_API_KEY=your-litellm-key
inspect eval arc.py --model litellm-proxy/claude-sonnet-5
```

The `model_name` is chosen by the proxy’s operator and need not be a model id. Inspect reads the proxy’s `/model/info` listing to find the upstream model behind it (`litellm_params.model`, or `model_info.base_model` when set), and uses that model to look up the context window and cost and to shape requests. Virtual keys work; the listing shows only the models a key may use. Aliases defined on your key or its team are resolved to the model they route to.

### Model Info

If Inspect can’t determine the context window of a proxy model, creating the model fails with an error that explains how to fix it. This happens for models the proxy lists without an upstream model Inspect knows (for example, a predeployment codename or an opaque deployment name). The best fix is to add `base_model` to the deployment in the proxy config, naming the known model it is closest to:

``` yaml
model_list:
  - model_name: pillow
    litellm_params:
      model: anthropic/claude-pillow-v1
      api_key: os.environ/ANTHROPIC_API_KEY
    model_info:
      base_model: anthropic/claude-opus-5-5
      supports_reasoning: true
      supports_adaptive_thinking: true
```

`base_model` also lets LiteLLM accept parameters it would otherwise reject for an unknown model, such as reasoning effort and (for Bedrock) tool choice. For Anthropic, OpenAI, Google and xAI models the error message suggests that vendor’s current frontier model.

For a Claude 4.6 or later model under a name LiteLLM doesn’t know (such as a codename), also set `supports_reasoning` and `supports_adaptive_thinking` as above. LiteLLM uses adaptive thinking only for Claude models it recognizes by their upstream name, and `base_model` does not change this. Without the flags, LiteLLM either sends extended thinking with a reasoning effort, which Claude 4.7 and later reject (Inspect then fails with an error naming this fix), or, without `base_model`, refuses the reasoning effort (Inspect then sends none and warns with this fix).

Alternatively, set `max_input_tokens` (and `max_output_tokens`) in the deployment’s `model_info`, register the model’s info with `set_model_info("litellm-proxy/<name>", ModelInfo(...))`, or pass `-M require_model_info=false` to use the model without model info. The last two are the only fixes for an alias from the proxy’s `model_alias_map` setting, which the proxy doesn’t show to clients.

Costs are taken from Inspect’s model database or, where it has none, from the prices in the proxy’s `model_info`.

### Claude Models

When the upstream model is a Claude model (recognized by its name, including models Inspect and LiteLLM don’t know), requests follow the native `anthropic` provider:

- `max_tokens` defaults to 32,000 plus an increment for the reasoning effort, rather than LiteLLM’s 4,096.
- [Prompt caching](https://docs.claude.com/en/docs/build-with-claude/prompt-caching) is enabled with cache breakpoints on the system prompt, the last tool definition and the last two messages. Pass `--cache-prompt=false` to disable. Cache writes are costed at the TTL the proxy reports for them (a proxy can apply a 1 hour TTL).
- Tool schemas keep their full JSON Schema (e.g. `pattern` and `minLength`).
- With a `reasoning_effort`, Claude 4.6 and later models return summarized thinking, which is recorded in the log and streamed during long thinking phases. Earlier LiteLLM versions otherwise return no thinking text for Claude 4.7 and later. Pass `-M thinking_display=omitted` to receive no thinking text.

A model that does not accept a conversation ending with an assistant message (assistant prefill) fails with an error saying so rather than being retried.

### Reasoning Effort

When the proxy or upstream model rejects the requested `reasoning_effort`, Inspect retries with the nearest effort the model accepts (or with no effort, if the model takes none), warns once, and uses that effort for later requests. Reasoning returned by the upstream model (including Claude thinking blocks and Gemini thought signatures) is sent back on the next turn.

### Responses API and Streaming

GPT-5 and later, o-series and Codex models served directly by OpenAI use the proxy’s Responses API, as they do with the `openai` provider, so their reasoning is carried from one turn to the next. This applies when every deployment of the model uses LiteLLM’s `openai` provider with no `api_base` other than OpenAI’s (Azure deployments and OpenAI-compatible servers are not included). A predeployment model qualifies when its deployment sets `model_info.base_model` (see [Model Info](#litellm-proxy-model-info)). Other models use the Chat Completions API. Pass `-M responses_api=true` or `-M responses_api=false` to choose the API. Through the proxy, web search is the only OpenAI built-in tool used on the Responses API; [computer()](./reference/inspect_ai.tool.html.md#computer), [code_execution()](./reference/inspect_ai.tool.html.md#code_execution) and remote MCP servers are sent as ordinary tools, as on the Chat Completions API.

Requests stream by default, so long generations don’t hit client or proxy timeouts, except Responses API requests for models not served by OpenAI. Pass `-M stream=false` to disable streaming.

### Web Search

The [web_search()](./reference/inspect_ai.tool.html.md#web_search) tool uses OpenAI’s built-in search (its `openai` provider) for OpenAI models that use the Responses API. Otherwise it needs an external search provider (`tavily`, `exa` or `google`), for example `web_search(["anthropic", "tavily"])`; other built-in search providers are not supported through the proxy.

### Model Args

| Argument | Description |
|----|----|
| `require_model_info` | Fail if the model has no context window (default: `true`). |
| `model_info` | Read the proxy’s `/model/info` listing (default: `true`). With `false`, Inspect does not look up the upstream model and requests are shaped by the proxy model name alone. |
| `stream` | Stream requests (default: `true`, except Responses API requests for models not served by OpenAI). |
| `responses_api` | Use the Responses API (default: `true` for GPT-5 and later, o-series and Codex models served by OpenAI, `false` otherwise). |
| `thinking_display` | Thinking text Claude 4.6 and later models return when a `reasoning_effort` is set: `summarized` (default) or `omitted`. |

The following environment variables are supported by the LiteLLM Proxy provider:

| Variable | Description |
|----|----|
| `LITELLM_PROXY_API_KEY` | API key credentials (required). `LITELLM_API_KEY` is used when it is not set. |
| `LITELLM_PROXY_BASE_URL` | Base URL of the proxy (required). `LITELLM_PROXY_API_BASE` and `LITELLM_BASE_URL` are used when it is not set. |

## Hugging Face Inference Providers

To use [Hugging Face Inference Providers](https://huggingface.co/docs/inference-providers), install the `openai` package (which provides the compatibility layer), set your Hugging Face token, and specify a model using the `--model` option:

``` bash
pip install openai
export HF_TOKEN=your-huggingface-token
inspect eval arc.py --model hf-inference-providers/openai/gpt-oss-120b
```

The above will automatically select the provider for you. If you want to use a specific provider you can append `:` followed by the provider name. To use cerebras for example, you would do the following:

``` bash
pip install openai
export HF_TOKEN=your-huggingface-token
inspect eval arc.py --model hf-inference-providers/openai/gpt-oss-120b:cerebras
```

HF Inference Providers provides unified access to hundreds of machine learning models through multiple world-class inference providers (Cerebras, Groq, Together AI, etc.) with automatic provider routing and failover.

The following environment variables are supported by the HF Inference Providers:

| Variable   | Description                                                 |
|------------|-------------------------------------------------------------|
| `HF_TOKEN` | Hugging Face token with appropriate permissions (required). |

### Streaming

HF Interference Providers uses streaming by default for requests. You can disable streaming using the `stream` model arg (or pass `auto` to stream only when the caller passes an `on_stream` callback to [generate()](./reference/inspect_ai.solver.html.md#generate)). For example:

``` bash
inspect eval arc.py --model hf-inference-providers/openai/gpt-oss-120b -M stream=false
```

## Custom Models

If you want to support another model hosting service or local model source, you can add a custom model API. See the documentation on [Model API Extensions](./extensions-model-api.html.md#sec-model-api-extensions) for additional details.
