# inspect_ai.model – Inspect

## Generation

### get_model

Get an instance of a model.

Calls to get_model() are memoized (i.e. a call with the same arguments will return an existing instance of the model rather than creating a new one). You can disable this with `memoize=False`.

If you prefer to immediately close models after use (as well as prevent caching) you can employ the async context manager built in to the [Model](../reference/inspect_ai.model.html.md#model) class. For example:

``` python
async with get_model("openai/gpt-4o") as model:
    response = await model.generate("Say hello")
```

In this case, the model client will be closed at the end of the context manager and will not be available in the get_model() cache.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L2230)

``` python
def get_model(
    model: str | Model | None = None,
    *,
    role: str | None = None,
    required: bool = False,
    default: str | Model | None = None,
    config: GenerateConfig | None = None,
    base_url: str | None = None,
    api_key: str | None = None,
    memoize: bool = True,
    **model_args: Any,
) -> Model
```

`model` str \| [Model](../reference/inspect_ai.model.html.md#model) \| None  
Model specification. If [Model](../reference/inspect_ai.model.html.md#model) is passed it is returned unmodified, if `None` is passed then the model currently being evaluated is returned (or if there is no evaluation then the model referred to by `INSPECT_EVAL_MODEL`).

`role` str \| None  
Optional named role for model (e.g. for roles specified at the task or eval level). Provide a `default` as a fallback in the case where the `role` hasn’t been externally specified. Pass `required` to raise an error if the role has not been specified. If the role is bound to a list of models, the first model in the list is returned (use [model_roles()](../reference/inspect_ai.model.html.md#model_roles) to access all of them).

`required` bool  
If a model role is specified, is it required? If required and not present, an error is raised. Otherwise, the current default model is returned.

`default` str \| [Model](../reference/inspect_ai.model.html.md#model) \| None  
Optional. Fallback model in case the specified `model` or `role` is not found.

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig) \| None  
Configuration for model.

`base_url` str \| None  
Optional. Alternate base URL for model.

`api_key` str \| None  
Optional. API key for model.

`memoize` bool  
Use/store a cached version of the model based on the parameters to [get_model()](../reference/inspect_ai.model.html.md#get_model)

`**model_args` Any  
Additional args to pass to model constructor.

### Model

Model interface.

Use [get_model()](../reference/inspect_ai.model.html.md#get_model) to get an instance of a model. Model provides an async context manager for closing the connection to it after use. For example:

``` python
async with get_model("openai/gpt-4o") as model:
    response = await model.generate("Say hello")
```

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L832)

``` python
class Model
```

#### Attributes

`api` [ModelAPI](../reference/inspect_ai.model.html.md#modelapi)  
Model API.

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig)  
Generation config.

`name` str  
Model name.

`explicit_base_url` str \| None  
Base URL explicitly provided by the user (not resolved from env/defaults).

`role` str \| None  
Model role.

#### Methods

\_\_init\_\_  
Create a model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L851)

``` python
def __init__(
    self,
    api: ModelAPI,
    config: GenerateConfig,
    model_args: dict[str, Any] | None = None,
) -> None
```

`api` [ModelAPI](../reference/inspect_ai.model.html.md#modelapi)  
Model API provider.

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig)  
Model configuration.

`model_args` dict\[str, Any\] \| None  
Optional model args

canonical_name  
Canonical model name for model info database lookup.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L910)

``` python
def canonical_name(self) -> str
```

input_tokens_name  
Model name used for looking up model input tokens.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L914)

``` python
def input_tokens_name(self) -> str
```

generate  
Generate output from the model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L934)

``` python
async def generate(
    self,
    input: str | list[ChatMessage],
    tools: Sequence[Tool | ToolDef | ToolInfo | ToolSource] | ToolSource = [],
    tool_choice: ToolChoice | None = None,
    config: GenerateConfig = GenerateConfig(),
    cache: bool | CachePolicy | NotGiven = NOT_GIVEN,
    on_stream: StreamHandler | None = None,
) -> ModelOutput
```

`input` str \| list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Chat message input (if a `str` is passed it is converted to a [ChatMessageUser](../reference/inspect_ai.model.html.md#chatmessageuser)).

`tools` Sequence\[[Tool](../reference/inspect_ai.tool.html.md#tool) \| [ToolDef](../reference/inspect_ai.tool.html.md#tooldef) \| [ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo) \| [ToolSource](../reference/inspect_ai.tool.html.md#toolsource)\] \| [ToolSource](../reference/inspect_ai.tool.html.md#toolsource)  
Tools available for the model to call.

`tool_choice` [ToolChoice](../reference/inspect_ai.tool.html.md#toolchoice) \| None  
Directives to the model as to which tools to prefer.

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig)  
Model configuration.

`cache` bool \| [CachePolicy](../reference/inspect_ai.model.html.md#cachepolicy) \| NotGiven  
Caching behavior for generate responses (defaults to no caching).

`on_stream` [StreamHandler](../reference/inspect_ai.model.html.md#streamhandler) \| None  
Optional async callback receiving incremental [StreamEvent](../reference/inspect_ai.model.html.md#streamevent)s (text / reasoning / tool-call deltas, plus retry boundaries) while the response streams — a side-channel for UI display; the final result is still the returned [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput). Passing a callback is itself a request to stream: providers that support streaming stream the response without any provider-level streaming flag (an explicit provider streaming opt-out still wins). Providers or calls that don’t stream never invoke it (a cache hit, for example, produces no content deltas — though a cache hit on a retry attempt still delivers the retry boundary), and the callback is best treated as display-only: on retry a [StreamRetryEvent](../reference/inspect_ai.model.html.md#streamretryevent) signals that deltas received so far belong to a failed attempt and should be discarded. A callback that raises never fails the model call: the exception is logged and the callback is detached for the remainder of that call (the next call tries it again).

generate_loop  
Generate output from the model, looping as long as the model calls tools.

Similar to [generate()](../reference/inspect_ai.solver.html.md#generate), but runs in a loop resolving model tool calls. The loop terminates when the model stops calling tools. The final [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput) as well the message list for the conversation are returned as a tuple.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L1068)

``` python
async def generate_loop(
    self,
    input: str | list[ChatMessage],
    tools: Sequence[Tool | ToolDef | ToolSource] | ToolSource = [],
    config: GenerateConfig = GenerateConfig(),
    cache: bool | CachePolicy | NotGiven = NOT_GIVEN,
    on_stream: StreamHandler | None = None,
) -> tuple[list[ChatMessage], ModelOutput]
```

`input` str \| list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Chat message input (if a `str` is passed it is converted to a [ChatMessageUser](../reference/inspect_ai.model.html.md#chatmessageuser)).

`tools` Sequence\[[Tool](../reference/inspect_ai.tool.html.md#tool) \| [ToolDef](../reference/inspect_ai.tool.html.md#tooldef) \| [ToolSource](../reference/inspect_ai.tool.html.md#toolsource)\] \| [ToolSource](../reference/inspect_ai.tool.html.md#toolsource)  
Tools available for the model to call.

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig)  
Model configuration.

`cache` bool \| [CachePolicy](../reference/inspect_ai.model.html.md#cachepolicy) \| NotGiven  
Caching behavior for generate responses (defaults to no caching).

`on_stream` [StreamHandler](../reference/inspect_ai.model.html.md#streamhandler) \| None  
Optional async callback receiving incremental [StreamEvent](../reference/inspect_ai.model.html.md#streamevent)s (see [generate()](../reference/inspect_ai.solver.html.md#generate)). Invoked for each generate call in the loop; attempt numbers in [StreamRetryEvent](../reference/inspect_ai.model.html.md#streamretryevent)s are per-call, not cumulative across the loop.

count_tokens  
Estimate token count for input.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L1123)

``` python
async def count_tokens(
    self,
    input: str | list[ChatMessage],
    config: GenerateConfig | None = None,
) -> int
```

`input` str \| list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Input to count tokens for.

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig) \| None  
Optional generation config for provider-specific counting (e.g., reasoning parameters that affect token allocation).

count_tool_tokens  
Count tokens for tool definitions.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L1215)

``` python
async def count_tool_tokens(self, tools: Sequence[ToolInfo]) -> int
```

`tools` Sequence\[[ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo)\]  
List of tool definitions.

compact  
Compact messages using provider-native compaction.

Delegates to the model provider’s native compaction API when available. Automatically tracks token usage and enforces token limits.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L1239)

``` python
async def compact(
    self,
    input: list[ChatMessage],
    tools: list[ToolInfo],
    instructions: str | None = None,
) -> tuple[list[ChatMessage], ModelUsage | None]
```

`input` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Chat message input (if a `str` is passed it is converted to a `ChatUserMessage`).

`tools` list\[[ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo)\]  
Tools available for the model to call.

`instructions` str \| None  
Additional instructions to give the model about compaction (e.g. “Focus on preserving code snippets, variable names, and technical decisions.”)

### ModelRole

Reference to a named model role, including its resolution policy.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_role.py#L6)

``` python
class ModelRole(BaseModel)
```

#### Attributes

`name` str  
Name of the model role.

`required` bool  
Whether a model must be bound to the role.

### ModelRoles

Assignment of models to named roles (e.g. the `model_roles` argument to [eval()](../reference/inspect_ai.html.md#eval) or [Task](../reference/inspect_ai.html.md#task)).

Maps a role name to a model (name or [Model](../reference/inspect_ai.model.html.md#model) instance) or to a list of models. Assigned roles are looked up with `get_model(role=...)` or [model_roles()](../reference/inspect_ai.model.html.md#model_roles) (to *reference* a role, e.g. from a scorer, see [ModelRole](../reference/inspect_ai.model.html.md#modelrole)).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L2844)

``` python
ModelRoles: TypeAlias = Mapping[str, str | Model | Sequence[str | Model]]
```

### GenerateConfig

Model generation options.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_generate_config.py#L204)

``` python
class GenerateConfig(BaseModel)
```

#### Attributes

`max_retries` int \| None  
Maximum number of times to retry request, so e.g. 1 allows two attempts total (defaults to unlimited).

`timeout` int \| None  
Timeout (in seconds) for an entire request (including retries).

`attempt_timeout` int \| None  
Timeout (in seconds) for any given attempt (if exceeded, will abandon attempt and retry according to max_retries).

`stream_idle_timeout` int \| None  
Timeout (in seconds) on silence within a streaming response — if a streaming attempt delivers no chunk for this long, the attempt is abandoned and retried according to max_retries. Setting it requests streaming (like on_stream); it has no effect on calls that do not stream.

`max_connections` int \| None  
Maximum number of concurrent connections to Model API (default is model specific).

`adaptive_connections` bool \| int \| [AdaptiveConcurrency](../reference/inspect_ai.util.html.md#adaptiveconcurrency) \| None  
Adaptive concurrency for model API connections. Defaults to enabled (`None` and `True` both resolve to `AdaptiveConcurrency()` defaults: min=10, start=20, max=100). Pass `False` to opt out (uses static concurrency). Pass an integer `N` as shorthand for `AdaptiveConcurrency(max=N)`. Pass an [AdaptiveConcurrency](../reference/inspect_ai.util.html.md#adaptiveconcurrency) to fully customize bounds and tuning (cooldown_seconds, decrease_factor, scale_up_percent). An explicit `max_connections` or `batch=True` takes precedence and uses static concurrency.

`system_message` str \| None  
Override the default system message.

`max_tokens` int \| None  
The maximum number of tokens that can be generated in the completion (default is model specific).

`top_p` float \| None  
An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass.

`temperature` float \| None  
What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

`stop_seqs` list\[str\] \| None  
Sequences where the API will stop generating further tokens. The returned text will not contain the stop sequence.

`best_of` int \| None  
Generates best_of completions server-side and returns the ‘best’ (the one with the highest log probability per token). vLLM only.

`frequency_penalty` float \| None  
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model’s likelihood to repeat the same line verbatim. OpenAI, Google, Grok, Groq, vLLM, and SGLang only.

`presence_penalty` float \| None  
Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model’s likelihood to talk about new topics. OpenAI, Google, Grok, Groq, vLLM, and SGLang only.

`logit_bias` dict\[int, float\] \| None  
Map token Ids to an associated bias value from -100 to 100 (e.g. “42=10,43=-10”). OpenAI, Grok, Grok, and vLLM only.

`seed` int \| None  
Random seed. OpenAI, Google, Mistral, Groq, HuggingFace, and vLLM only.

`top_k` int \| None  
Randomly sample the next word from the top_k most likely next words. Anthropic, Google, HuggingFace, vLLM, and SGLang only.

`num_choices` int \| None  
How many chat completion choices to generate for each input message. OpenAI, Grok, Google, TogetherAI, vLLM, and SGLang only.

`logprobs` bool \| None  
Return log probabilities of the output tokens. OpenAI, Grok, TogetherAI, Huggingface, llama-cpp-python, vLLM, and SGLang only.

`top_logprobs` int \| None  
Number of most likely tokens (0-20) to return at each token position, each with an associated log probability. OpenAI, Grok, Huggingface, vLLM, and SGLang only.

`prompt_logprobs` int \| None  
Number of log probabilities to return per prompt token (1-20). When greater than 1, top-N alternative tokens are also returned. vLLM only.

`parallel_tool_calls` bool \| None  
Whether to enable parallel function calling during tool use (defaults to True). OpenAI and Groq only.

`internal_tools` bool \| None  
Whether to automatically map tools to model internal implementations (e.g. ‘computer’ for anthropic).

`max_tool_output` int \| None  
Maximum tool output (in bytes). Defaults to 16 \* 1024.

`cache_prompt` Literal\['auto'\] \| bool \| None  
Whether to cache the prompt prefix. Enabled by default. Set to False to disable: on Anthropic and Bedrock Converse (Claude and Nova) this turns off the provider’s own automatic caching; on OpenAI it only turns off explicit `cache_breakpoint` marks — the model’s own implicit caching still applies regardless. Use `ContentText(cache_breakpoint=True)` to mark an explicit cache boundary (e.g. a fixed rubric ahead of a varying item) instead of relying on automatic caching; Anthropic and OpenAI `gpt-5.6`+ only.

`fallback_models` list\[str\] \| None  
Fallback models tried in order when the model’s safety classifiers refuse the request. Anthropic Claude API only (not supported on Bedrock/Vertex/Azure or with batch mode).

`fail_on_refusal` bool \| None  
Raise a `ModelRefusalError` (failing the sample) when the model returns `stop_reason="content_filter"`. Defaults to False.

`verbosity` Literal\['low', 'medium', 'high'\] \| None  
Constrains the verbosity of the model’s response. Lower values will result in more concise responses, while higher values will result in more verbose responses. GPT 5.x models only (defaults to “medium” for OpenAI models).

`effort` Literal\['low', 'medium', 'high', 'xhigh', 'max'\] \| None  
Control how many tokens are used for a response, trading off between response thoroughness and token efficiency. Anthropic Claude Opus 4.5+ only (`max` only supported on 4.6 and 4.7, `xhigh` supported only on 4.7).

`reasoning_effort` Literal\['none', 'minimal', 'low', 'medium', 'high', 'xhigh', 'max'\] \| None  
Constrains effort on reasoning. Defaults vary by provider and model and not all models support all values (please consult provider documentation for details).

`reasoning_mode` Literal\['standard', 'pro'\] \| None  
Reasoning mode. “pro” performs more model work for greater reliability on difficult tasks, at higher latency and token usage. OpenAI GPT-5.6+ models only (“standard” is the default).

`reasoning_tokens` int \| None  
Maximum number of tokens to use for reasoning. Anthropic Claude models only.

`reasoning_summary` Literal\['none', 'concise', 'detailed', 'auto'\] \| None  
Provide summary of reasoning steps (OpenAI reasoning models only). Use ‘auto’ to access the most detailed summarizer available for the current model (defaults to ‘auto’ if your organization is verified by OpenAI).

`reasoning_history` Literal\['none', 'all', 'last', 'auto'\] \| None  
Include reasoning in chat message history sent to generate.

`response_schema` [ResponseSchema](../reference/inspect_ai.model.html.md#responseschema) \| None  
Request a response format as JSONSchema (output should still be validated). OpenAI, Google, Mistral, vLLM, and SGLang only.

`extra_headers` dict\[str, str\] \| None  
Extra headers to be sent with requests. Not supported for AzureAI, Bedrock, and Grok.

`extra_body` dict\[str, Any\] \| None  
Extra body to be sent with requests to OpenAI compatible servers. OpenAI, vLLM, and SGLang only.

`modalities` list\[[OutputModality](../reference/inspect_ai.model.html.md#outputmodality)\] \| None  
Additional output modalities to enable beyond text (e.g. \[“image”\]). OpenAI and Google only.

`cache` bool \| [CachePolicy](../reference/inspect_ai.model.html.md#cachepolicy) \| None  
Policy for caching of model generate output.

`batch` bool \| int \| [BatchConfig](../reference/inspect_ai.model.html.md#batchconfig) \| None  
Use batching API when available. True to enable batching with default configuration, False to disable batching, a number to enable batching of the specified batch size, or a BatchConfig object specifying the batching configuration.

#### Methods

merge  
Merge another model configuration into this one.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_generate_config.py#L379)

``` python
def merge(
    self, other: Union["GenerateConfig", GenerateConfigArgs]
) -> "GenerateConfig"
```

`other` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig) \| [GenerateConfigArgs](../reference/inspect_ai.model.html.md#generateconfigargs)  
Configuration to merge.

### GenerateConfigArgs

Type for kwargs that selectively override GenerateConfig.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_generate_config.py#L78)

``` python
class GenerateConfigArgs(TypedDict, total=False)
```

### GenerateFilter

Filter a model generation.

The first argument is the resolved [Model](../reference/inspect_ai.model.html.md#model) instance. Filters that accept a `str` as the first argument are still supported but deprecated and will receive `model.name` instead.

A filter may substitute for the default model generation by returning a [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput), modify the input parameters by returning a `GenerateInput`, or return `None` to allow default processing to continue.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L2070)

``` python
GenerateFilter: TypeAlias = ModelGenerateFilter | StrGenerateFilter
```

### BatchConfig

Batch processing configuration.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_generate_config.py#L37)

``` python
class BatchConfig(BaseModel)
```

#### Attributes

`size` int \| None  
Target minimum number of requests to include in each batch. If not specified, uses default of 100. Batches may be smaller if the timeout is reached or if requests don’t fit within size limits.

`max_size` int \| None  
Maximum number of requests to include in each batch. If not specified, falls back to the provider-specific maximum batch size.

`send_delay` float \| None  
Maximum time (in seconds) to wait before sending a partially filled batch. If not specified, uses a default of 15 seconds. This prevents indefinite waiting when request volume is low.

`tick` float \| None  
Time interval (in seconds) between checking for new batch requests and batch completion status. If not specified, uses a default of 15 seconds.

When expecting a very large number of concurrent batches, consider increasing this value to reduce overhead from continuous polling since an http request must be made for each batch on each tick.

`max_batches` int \| None  
Maximum number of batches to have in flight at once for a provider (defaults to 100).

`max_consecutive_check_failures` int \| None  
Maximum number of consecutive check failures before failing a batch (defaults to 1000).

### ResponseSchema

Schema for model response when using Structured Output.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_generate_config.py#L20)

``` python
class ResponseSchema(BaseModel)
```

#### Attributes

`name` str  
The name of the response schema. Must be a-z, A-Z, 0-9, or contain underscores and dashes, with a maximum length of 64.

`json_schema` [JSONSchema](../reference/inspect_ai.util.html.md#jsonschema)  
The schema for the response format, described as a JSON Schema object.

`description` str \| None  
A description of what the response format is for, used by the model to determine how to respond in the format.

`strict` bool \| None  
Whether to enable strict schema adherence when generating the output. If set to true, the model will always follow the exact schema defined in the schema field. OpenAI and Mistral only.

### ModelOutput

Output from model generation.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_output.py#L259)

``` python
class ModelOutput(BaseModel)
```

#### Attributes

`model` str  
Model used for generation.

`choices` list\[[ChatCompletionChoice](../reference/inspect_ai.model.html.md#chatcompletionchoice)\]  
Completion choices.

`completion` str  
Model completion.

`usage` [ModelUsage](../reference/inspect_ai.model.html.md#modelusage) \| None  
Model token usage

`fallback` ModelFallback \| None  
Model fallback that served this output (None if served by the requested model).

`time` float \| None  
Time elapsed (in seconds) for call to generate.

`metadata` dict\[str, Any\] \| None  
Additional metadata associated with model output.

`error` str \| None  
Error message in the case of content moderation refusals.

`stop_reason` [StopReason](../reference/inspect_ai.model.html.md#stopreason)  
First message stop reason.

`message` [ChatMessageAssistant](../reference/inspect_ai.model.html.md#chatmessageassistant)  
First message choice.

#### Methods

from_message  
Create ModelOutput from a ChatMessageAssistant.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_output.py#L308)

``` python
@staticmethod
def from_message(
    message: ChatMessage,
    stop_reason: StopReason = "stop",
) -> "ModelOutput"
```

`message` [ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)  
Assistant message.

`stop_reason` [StopReason](../reference/inspect_ai.model.html.md#stopreason)  
Stop reason for generation

from_content  
Create ModelOutput from a `str` or `list[Content]`.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_output.py#L344)

``` python
@staticmethod
def from_content(
    model: str,
    content: str | list[Content],
    stop_reason: StopReason = "stop",
    error: str | None = None,
    stop_details: StopDetails | None = None,
) -> "ModelOutput"
```

`model` str  
Model name.

`content` str \| list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Text content from generation.

`stop_reason` [StopReason](../reference/inspect_ai.model.html.md#stopreason)  
Stop reason for generation.

`error` str \| None  
Error message.

`stop_details` StopDetails \| None  
Additional detail about the stop reason (e.g. refusal).

for_tool_call  
Returns a ModelOutput for requesting a tool call.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_output.py#L375)

``` python
@staticmethod
def for_tool_call(
    model: str,
    tool_name: str,
    tool_arguments: dict[str, Any],
    internal: JsonValue | None = None,
    tool_call_id: str | None = None,
    content: str | None = None,
) -> "ModelOutput"
```

`model` str  
model name

`tool_name` str  
The name of the tool.

`tool_arguments` dict\[str, Any\]  
The arguments passed to the tool.

`internal` JsonValue \| None  
The model’s internal info for the tool (if any).

`tool_call_id` str \| None  
Optional ID for the tool call. Defaults to a random UUID.

`content` str \| None  
Optional content to include in the message. Defaults to “tool call for tool {tool_name}”.

### ModelConfig

Model config.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_config.py#L10)

``` python
class ModelConfig(BaseModel)
```

#### Attributes

`model` str  
Model name.

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig)  
Generate config

`base_url` str \| None  
Model base url.

`args` dict\[str, Any\]  
Model specific arguments.

### ModelCall

Model call (raw request/response data).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_call.py#L41)

``` python
class ModelCall(BaseModel)
```

#### Attributes

`request` dict\[str, JsonValue\]  
Raw data posted to model.

`response` dict\[str, JsonValue\] \| None  
Raw response data from model (None if call is still pending).

`error` bool \| None  
Did this model call result in an error.

`time` float \| None  
Time taken for underlying model call.

`call_refs` list\[tuple\[int, int\]\] \| None  
Call pool references. Each element is a (start, end_exclusive) range.

`call_key` str \| None  
Key under which messages lived in call.request (‘messages’ or ‘contents’).

#### Methods

create  
Create a ModelCall object.

Create a ModelCall from arbitrary request and response objects (they might be dataclasses, Pydandic objects, dicts, etc.). Converts all values to JSON serialiable (exluding those that can’t be)

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_call.py#L64)

``` python
@staticmethod
def create(
    request: Any,
    response: Any | None,
    filter: ModelCallFilter | None = None,
    time: float | None = None,
) -> "ModelCall"
```

`request` Any  
Request object (dict, dataclass, BaseModel, etc.)

`response` Any \| None  
Response object (dict, dataclass, BaseModel, etc.), or None if the call is still pending.

`filter` ModelCallFilter \| None  
Function for filtering model call data.

`time` float \| None  
Time taken for underlying ModelCall

### ModelConversation

Model conversation.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_conversation.py#L7)

``` python
class ModelConversation(Protocol)
```

#### Attributes

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Conversation history.

`output` [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput)  
Model output.

### ModelUsage

Token usage for completion.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_output.py#L16)

``` python
class ModelUsage(BaseModel)
```

#### Attributes

`input_tokens` int  
Input tokens charged at full rate (excludes cached tokens).

This count excludes tokens reported in input_tokens_cache_read and input_tokens_cache_write. The true total input token count is: input_tokens + (input_tokens_cache_read or 0) + (input_tokens_cache_write or 0).

`output_tokens` int  
Total output tokens used.

`total_tokens` int  
Total tokens used.

`input_tokens_cache_write` int \| None  
Number of tokens written to the cache.

`input_tokens_cache_read` int \| None  
Number of tokens retrieved from the cache.

`reasoning_tokens` int \| None  
Number of tokens used for reasoning.

`total_cost` float \| None  
Total cost in dollars for this usage.

### StopReason

Reason that the model stopped or failed to generate.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_output.py#L95)

``` python
StopReason = Literal[
    "stop",
    "max_tokens",
    "model_length",
    "tool_calls",
    "content_filter",
    "unknown",
]
```

### ChatCompletionChoice

Choice generated for completion.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_output.py#L224)

``` python
class ChatCompletionChoice(BaseModel)
```

#### Attributes

`message` [ChatMessageAssistant](../reference/inspect_ai.model.html.md#chatmessageassistant)  
Assistant message.

`stop_reason` [StopReason](../reference/inspect_ai.model.html.md#stopreason)  
Reason that the model stopped generating.

`stop_details` StopDetails \| None  
Additional detail about the stop reason (e.g. refusal category/explanation), when provided.

`logprobs` [Logprobs](../reference/inspect_ai.model.html.md#logprobs) \| None  
Logprobs.

`prompt_logprobs` [Logprobs](../reference/inspect_ai.model.html.md#logprobs) \| None  
Per-prompt-token log probabilities (vLLM only).

Placed on the choice (not [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput)) so scorers access prompt and output logprobs uniformly via `choices[0]`. Perplexity evals use `num_choices=1`, so there is no duplication in practice.

### OutputModality

Output modality type. Either a literal string or an ImageOutput configuration.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_generate_config.py#L74)

``` python
OutputModality = Union[Literal["image"], ImageOutput]
```

### ImageOutput

Image output configuration.

Use the `options` field to pass provider-specific options directly to the underlying API (e.g. OpenAI image_generation tool parameters).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_generate_config.py#L63)

``` python
class ImageOutput(BaseModel)
```

#### Attributes

`options` dict\[Literal\['openai'\], dict\[str, Any\]\] \| None  
Provider-specific image output options, keyed by provider name.

### RetryDecision

Classification of a retryable exception for `ModelAPI.should_retry`.

`should_retry()` may return either a plain `bool` (legacy: any True is treated as a generic transient retry) or a [RetryDecision](../reference/inspect_ai.model.html.md#retrydecision) to additionally classify the retry kind for the adaptive concurrency controller and separately record any server-suggested wait time.

[RetryDecision](../reference/inspect_ai.model.html.md#retrydecision) is truthy iff `retry` is True, so existing callers written against the `bool` return (`if api.should_retry(ex): ...`) keep working unchanged.

Use the `no()`, `transient()`, and `rate_limit()` factory methods rather than constructing directly.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L221)

``` python
@dataclasses.dataclass(frozen=True)
class RetryDecision
```

#### Attributes

`retry` bool  
Whether to retry the request.

`kind` Literal\['rate_limit', 'transient'\]  
How to account for the retry against the adaptive controller.

`rate_limit` triggers a scale-down. `transient` (5xx, timeouts, network errors) only marks the request as retried so the eventual success won’t count toward scale-up — it does not shrink the limit.

`retry_after` float \| None  
Recommended seconds to wait before retrying, if the server provided one (e.g. via `Retry-After`).

Exposed and reserved for future use: nothing currently consumes it — it affects neither Inspect’s retry backoff nor the adaptive concurrency cooldown.

#### Methods

no  
Don’t retry.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L259)

``` python
@classmethod
def no(cls) -> "RetryDecision"
```

transient  
Retry as a transient error (pauses scale-up but doesn’t scale down).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L264)

``` python
@classmethod
def transient(cls, retry_after: float | None = None) -> "RetryDecision"
```

`retry_after` float \| None  

rate_limit  
Retry as a rate-limit error (scales the adaptive controller down).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L269)

``` python
@classmethod
def rate_limit(cls, retry_after: float | None = None) -> "RetryDecision"
```

`retry_after` float \| None  

## Streaming

### StreamHandler

Async callback receiving [StreamEvent](../reference/inspect_ai.model.html.md#streamevent)s during `Model.generate()`.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_stream.py#L130)

``` python
StreamHandler: TypeAlias = Callable[[StreamEvent], Awaitable[None]]
```

### StreamEvent

Incremental event delivered to `on_stream` during `Model.generate()`.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_stream.py#L125)

``` python
StreamEvent = Union[
    StreamTextEvent, StreamReasoningEvent, StreamToolCallEvent, StreamRetryEvent
]
```

### StreamTextEvent

Incremental text delta from a streaming model response.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_stream.py#L69)

``` python
class StreamTextEvent(BaseModel)
```

#### Attributes

`type` Literal\['text'\]  
Event type.

`text` str  
Text fragment (append to previously received text).

### StreamReasoningEvent

Incremental reasoning delta from a streaming model response.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_stream.py#L79)

``` python
class StreamReasoningEvent(BaseModel)
```

#### Attributes

`type` Literal\['reasoning'\]  
Event type.

`reasoning` str  
Reasoning fragment (append to previously received reasoning).

### StreamToolCallEvent

Incremental tool call delta from a streaming model response.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_stream.py#L89)

``` python
class StreamToolCallEvent(BaseModel)
```

#### Attributes

`type` Literal\['tool_call'\]  
Event type.

`id` str \| None  
Identifier of the tool call the fragment belongs to (when reported).

`function` str \| None  
Name of the function being called (when reported).

`arguments` str  
Argument fragment (partial JSON — append to previously received fragments for the same call; complete JSON only once the call finishes).

### StreamRetryEvent

The model call is being retried after a failed attempt.

Emitted before any deltas from the new attempt when a prior attempt already delivered deltas: content received so far belongs to the failed attempt and should be discarded — the final [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput) is produced entirely by the attempt that succeeds.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_stream.py#L106)

``` python
class StreamRetryEvent(BaseModel)
```

#### Attributes

`type` Literal\['retry'\]  
Event type.

`attempt` int  
The attempt about to run (1-based; the first retry is attempt 2). A provider-internal retry that regenerates output within one attempt (e.g. a malformed-function-call retry) re-announces the current attempt number, so consecutive boundaries may carry the same value.

## Messages

### ChatMessage

Message in a chat conversation

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L210)

``` python
ChatMessage = Union[
    ChatMessageSystem, ChatMessageUser, ChatMessageAssistant, ChatMessageTool
]
```

### ChatMessageBase

Base class for chat messages.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L20)

``` python
class ChatMessageBase(BaseModel)
```

#### Attributes

`id` str \| None  
Unique identifer for message.

`content` str \| list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Content (simple string or list of content objects)

`source` Literal\['input', 'generate', 'operator'\] \| None  
Source of message.

`metadata` dict\[str, Any\] \| None  
Additional message metadata.

`text` str  
Get the text content of this message.

ChatMessage content is very general and can contain either a simple text value or a list of content parts (each of which can either be text or an image). Solvers (e.g. for prompt engineering) often need to interact with chat messages with the assumption that they are a simple string. The text property returns either the plain str content, or if the content is a list of text and images, the text items concatenated together (separated by newline)

`content_list` list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Message content as a list of Content objects.

#### Methods

metadata_as  
Metadata as a Pydantic model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L35)

``` python
def metadata_as(self, metadata_cls: Type[MT]) -> MT
```

`metadata_cls` Type\[MT\]  
BaseModel derived class.

### ChatMessageSystem

System chat message.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L140)

``` python
class ChatMessageSystem(ChatMessageBase)
```

#### Attributes

`id` str \| None  
Unique identifer for message.

`content` str \| list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Content (simple string or list of content objects)

`source` Literal\['input', 'generate', 'operator'\] \| None  
Source of message.

`metadata` dict\[str, Any\] \| None  
Additional message metadata.

`text` str  
Get the text content of this message.

ChatMessage content is very general and can contain either a simple text value or a list of content parts (each of which can either be text or an image). Solvers (e.g. for prompt engineering) often need to interact with chat messages with the assumption that they are a simple string. The text property returns either the plain str content, or if the content is a list of text and images, the text items concatenated together (separated by newline)

`content_list` list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Message content as a list of Content objects.

`role` Literal\['system'\]  
Conversation role.

#### Methods

metadata_as  
Metadata as a Pydantic model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L35)

``` python
def metadata_as(self, metadata_cls: Type[MT]) -> MT
```

`metadata_cls` Type\[MT\]  
BaseModel derived class.

### ChatMessageUser

User chat message.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L147)

``` python
class ChatMessageUser(ChatMessageBase)
```

#### Attributes

`id` str \| None  
Unique identifer for message.

`content` str \| list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Content (simple string or list of content objects)

`source` Literal\['input', 'generate', 'operator'\] \| None  
Source of message.

`metadata` dict\[str, Any\] \| None  
Additional message metadata.

`text` str  
Get the text content of this message.

ChatMessage content is very general and can contain either a simple text value or a list of content parts (each of which can either be text or an image). Solvers (e.g. for prompt engineering) often need to interact with chat messages with the assumption that they are a simple string. The text property returns either the plain str content, or if the content is a list of text and images, the text items concatenated together (separated by newline)

`content_list` list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Message content as a list of Content objects.

`role` Literal\['user'\]  
Conversation role.

`tool_call_id` list\[str\] \| None  
ID(s) of tool call(s) this message has the content payload for.

#### Methods

metadata_as  
Metadata as a Pydantic model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L35)

``` python
def metadata_as(self, metadata_cls: Type[MT]) -> MT
```

`metadata_cls` Type\[MT\]  
BaseModel derived class.

### ChatMessageAssistant

Assistant chat message.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L157)

``` python
class ChatMessageAssistant(ChatMessageBase)
```

#### Attributes

`id` str \| None  
Unique identifer for message.

`content` str \| list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Content (simple string or list of content objects)

`source` Literal\['input', 'generate', 'operator'\] \| None  
Source of message.

`metadata` dict\[str, Any\] \| None  
Additional message metadata.

`text` str  
Get the text content of this message.

ChatMessage content is very general and can contain either a simple text value or a list of content parts (each of which can either be text or an image). Solvers (e.g. for prompt engineering) often need to interact with chat messages with the assumption that they are a simple string. The text property returns either the plain str content, or if the content is a list of text and images, the text items concatenated together (separated by newline)

`content_list` list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Message content as a list of Content objects.

`role` Literal\['assistant'\]  
Conversation role.

`tool_calls` list\[ToolCall\] \| None  
Tool calls made by the model.

`model` str \| None  
Model used to generate assistant message.

#### Methods

metadata_as  
Metadata as a Pydantic model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L35)

``` python
def metadata_as(self, metadata_cls: Type[MT]) -> MT
```

`metadata_cls` Type\[MT\]  
BaseModel derived class.

### ChatMessageTool

Tool chat message.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L170)

``` python
class ChatMessageTool(ChatMessageBase)
```

#### Attributes

`id` str \| None  
Unique identifer for message.

`content` str \| list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Content (simple string or list of content objects)

`source` Literal\['input', 'generate', 'operator'\] \| None  
Source of message.

`metadata` dict\[str, Any\] \| None  
Additional message metadata.

`text` str  
Get the text content of this message.

ChatMessage content is very general and can contain either a simple text value or a list of content parts (each of which can either be text or an image). Solvers (e.g. for prompt engineering) often need to interact with chat messages with the assumption that they are a simple string. The text property returns either the plain str content, or if the content is a list of text and images, the text items concatenated together (separated by newline)

`content_list` list\[[Content](../reference/inspect_ai.model.html.md#content)\]  
Message content as a list of Content objects.

`role` Literal\['tool'\]  
Conversation role.

`tool_call_id` str \| None  
ID of tool call.

`function` str \| None  
Name of function called.

`error` [ToolCallError](../reference/inspect_ai.tool.html.md#toolcallerror) \| None  
Error which occurred during tool call.

#### Methods

metadata_as  
Metadata as a Pydantic model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_chat_message.py#L35)

``` python
def metadata_as(self, metadata_cls: Type[MT]) -> MT
```

`metadata_cls` Type\[MT\]  
BaseModel derived class.

### trim_messages

Trim message list to fit within model context.

Trim the list of messages by: - Retaining all system messages. - Retaining the ‘input’ messages from the sample. - Preserving a proportion of the remaining messages (`preserve=0.7` by default). - Ensuring that all assistant tool calls have corresponding tool messages. - Ensuring that the sequence of messages doesn’t end with an assistant message.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_trim.py#L10)

``` python
async def trim_messages(
    messages: list[ChatMessage], preserve: float = 0.7
) -> list[ChatMessage]
```

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
List of messages to trim.

`preserve` float  
Ratio of converation messages to preserve (defaults to 0.7)

### user_prompt

Get the last “user” message within a message history.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_prompt.py#L4)

``` python
def user_prompt(messages: list[ChatMessage]) -> ChatMessageUser
```

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Message history.

### stable_message_ids

Create a function that applies stable message IDs based on content hash.

Messages with identical content receive the same ID within a transcript, enabling cross-event message identity tracking. This is useful when an agent makes multiple LLM calls where subsequent calls include previous messages in the conversation history.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_message_ids.py#L19)

``` python
def stable_message_ids() -> Callable[[Sequence[ChatMessage] | ModelEvent], None]
```

## Content

### Content

Content sent to or received from a model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/content.py#L209)

``` python
Content = Union[
    ContentText,
    ContentReasoning,
    ContentImage,
    ContentAudio,
    ContentVideo,
    ContentData,
    ContentToolUse,
    ContentDocument,
]
```

### ContentText

Text content.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/content.py#L16)

``` python
class ContentText(ContentBase)
```

#### Attributes

`internal` JsonValue \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['text'\]  
Type.

`text` str  
Text content.

`refusal` bool \| None  
Was this a refusal message?

`citations` Sequence\[[Citation](../reference/inspect_ai.model.html.md#citation)\] \| None  
Citations supporting the text block.

`cache_breakpoint` bool \| None  
Place an explicit prompt-cache boundary after this block, replacing the provider’s automatic breakpoints for the whole request. Unsupported placements (e.g. tool results, mid-conversation system messages, empty blocks) or models fall back to normal automatic caching. Anthropic Claude API and OpenAI `gpt-5.6`+ only.

### ContentReasoning

Reasoning content.

See the specification for [thinking blocks](https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking#understanding-thinking-blocks) for Claude models.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/content.py#L35)

``` python
class ContentReasoning(ContentBase)
```

#### Attributes

`internal` JsonValue \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['reasoning'\]  
Type.

`reasoning` str  
Reasoning content.

`summary` str \| None  
Reasoning summary or readable reasoning text, if available.

`signature` str \| None  
Signature for reasoning content (used by some models to ensure that reasoning content is not modified for replay)

`redacted` bool  
Indicates that the explicit content of this reasoning block has been redacted.

`text` str  
Pure text rendering of reasoning (used for replay/interop).

### ContentImage

Image content.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/content.py#L91)

``` python
class ContentImage(ContentBase)
```

#### Attributes

`internal` JsonValue \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['image'\]  
Type.

`image` str  
Either a URL of the image or the base64 encoded image data.

`detail` Literal\['auto', 'low', 'high', 'original'\]  
Specifies the detail level of the image.

Currently only supported for OpenAI. Learn more in the [Vision guide](https://platform.openai.com/docs/guides/vision/low-or-high-fidelity-image-understanding).

### ContentAudio

Audio content.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/content.py#L110)

``` python
class ContentAudio(ContentBase)
```

#### Attributes

`internal` JsonValue \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['audio'\]  
Type.

`audio` str  
Audio file path or base64 encoded data URL.

`format` ContentAudioFormat  
Format of audio data (‘mp3’ or ‘wav’)

### ContentVideo

Video content.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/content.py#L126)

``` python
class ContentVideo(ContentBase)
```

#### Attributes

`internal` JsonValue \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['video'\]  
Type.

`video` str  
Video file path or base64 encoded data URL.

`format` ContentVideoFormat  
Format of video data (‘mp4’, ‘mpeg’, or ‘mov’)

### ContentDocument

Document content (e.g. a PDF).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/content.py#L139)

``` python
class ContentDocument(ContentBase)
```

#### Attributes

`internal` JsonValue \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['document'\]  
Type.

`document` str  
Document file path or base64 encoded data URL.

`filename` str  
Document filename (automatically determined from ‘document’ if not specified).

`mime_type` str  
Document mime type (automatically determined from ‘document’ if not specified).

`citations` bool  
Enable model-generated citations for text or PDF documents.

Anthropic requires citations on all citation-capable documents in a request; the provider enables them on every text or PDF document when any document enables them. Image citations are unsupported. Providers without document- citation support ignore this field.

### ContentData

Model internal.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/content.py#L199)

``` python
class ContentData(ContentBase)
```

#### Attributes

`internal` JsonValue \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['data'\]  
Type.

`data` dict\[str, JsonValue\]  
Model provider specific payload - required for internal content.

### ContentToolUse

Server side tool use.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/content.py#L63)

``` python
class ContentToolUse(ContentBase)
```

#### Attributes

`internal` JsonValue \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['tool_use'\]  
Type.

`tool_type` Literal\['web_search', 'mcp_call', 'code_execution'\]  
The type of the tool call.

`id` str  
The unique ID of the tool call.

`name` str  
Name of the tool.

`context` str \| None  
Tool context (e.g. MCP Server)

`arguments` str  
Arguments passed to the tool.

`result` str  
Result from the tool call.

`error` str \| None  
The error from the tool call (if any).

## Citation

### Citation

A citation sent to or received from a model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/citation.py#L79)

``` python
Citation: TypeAlias = Annotated[
    Union[
        ContentCitation,
        DocumentCitation,
        UrlCitation,
    ],
    Discriminator("type"),
]
```

### CitationBase

Base class for citations.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/citation.py#L6)

``` python
class CitationBase(BaseModel)
```

#### Attributes

`cited_text` str \| tuple\[int, int\] \| None  
The cited text

This can be the text itself or a start/end range of the text content within the container that is the cited text.

`title` str \| None  
Title of the cited resource.

`internal` dict\[str, JsonValue\] \| None  
Model provider specific payload - typically used to aid transformation back to model types.

### UrlCitation

A citation that refers to a URL.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/citation.py#L69)

``` python
class UrlCitation(CitationBase)
```

#### Attributes

`cited_text` str \| tuple\[int, int\] \| None  
The cited text

This can be the text itself or a start/end range of the text content within the container that is the cited text.

`title` str \| None  
Title of the cited resource.

`internal` dict\[str, JsonValue\] \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['url'\]  
Type.

`url` str  
URL of the cited resource.

### DocumentCitation

A citation that refers to a page range in a document.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/citation.py#L59)

``` python
class DocumentCitation(CitationBase)
```

#### Attributes

`cited_text` str \| tuple\[int, int\] \| None  
The cited text

This can be the text itself or a start/end range of the text content within the container that is the cited text.

`title` str \| None  
Title of the cited resource.

`internal` dict\[str, JsonValue\] \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['document'\]  
Type.

`range` DocumentRange \| None  
Range of the document that is cited.

### ContentCitation

A generic content citation.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_util/citation.py#L39)

``` python
class ContentCitation(CitationBase)
```

#### Attributes

`cited_text` str \| tuple\[int, int\] \| None  
The cited text

This can be the text itself or a start/end range of the text content within the container that is the cited text.

`title` str \| None  
Title of the cited resource.

`internal` dict\[str, JsonValue\] \| None  
Model provider specific payload - typically used to aid transformation back to model types.

`type` Literal\['content'\]  
Type.

## Tools

### execute_tools

Perform tool calls in the last assistant message.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_call_tools.py#L226)

``` python
async def execute_tools(
    messages: list[ChatMessage],
    tools: Sequence[Tool | ToolDef | ToolSource] | ToolSource,
    max_output: int | None = None,
    approval: list["ApprovalPolicy"] | None = None,
    review: list["ReviewPolicy"] | None = None,
) -> ExecuteToolsResult
```

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Current message list

`tools` Sequence\[[Tool](../reference/inspect_ai.tool.html.md#tool) \| [ToolDef](../reference/inspect_ai.tool.html.md#tooldef) \| [ToolSource](../reference/inspect_ai.tool.html.md#toolsource)\] \| [ToolSource](../reference/inspect_ai.tool.html.md#toolsource)  
Available tools

`max_output` int \| None  
Maximum output length (in bytes). Defaults to max_tool_output from active GenerateConfig (16 \* 1024 by default).

`approval` list\[[ApprovalPolicy](../reference/inspect_ai.approval.html.md#approvalpolicy)\] \| None  
Approval policies to use for tool calls within this execution. Temporarily replaces any active approval policies for the duration of the call.

`review` list\[[ReviewPolicy](../reference/inspect_ai.review.html.md#reviewpolicy)\] \| None  
Review policies to use for the results of tool calls within this execution. Temporarily replaces any active review policies for the duration of the call.

### ExecuteToolsResult

Result from executing tools in the last assistant message.

In conventional tool calling scenarios there will be only a list of [ChatMessageTool](../reference/inspect_ai.model.html.md#chatmessagetool) appended and no-output. However, if there are [handoff()](../reference/inspect_ai.agent.html.md#handoff) tools (used in multi-agent systems) then other messages may be appended and an `output` may be available as well.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_call_tools.py#L110)

``` python
class ExecuteToolsResult(NamedTuple)
```

#### Attributes

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Messages added to conversation.

`output` [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput) \| None  
Model output if a generation occurred within the conversation.

## Compaction

### compaction

Create a conversation compaction handler.

Call `compact_input()` with the full conversation history before sending input to the model. Send the returned `input` and append the supplemental message returned (if any) to the full history. Call `record_output()` after each generate call to calibrate token estimation.

See the [Compaction](https://inspect.aisi.org.uk/compaction.html) for additional details on using compaction.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/_compaction.py#L59)

``` python
def compaction(
    strategy: CompactionStrategy,
    prefix: list[ChatMessage],
    tools: Sequence[Tool | ToolDef | ToolInfo | ToolSource] | ToolSource | None = None,
    model: str | Model | None = None,
    checkpointer: Checkpointer = _NOOP_CHECKPOINTER,
) -> Compact
```

`strategy` [CompactionStrategy](../reference/inspect_ai.model.html.md#compactionstrategy)  
Compaction strategy (e.g. editing, trimming, summary, etc.)

`prefix` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Chat messages to always preserve in compacted conversations.

`tools` Sequence\[[Tool](../reference/inspect_ai.tool.html.md#tool) \| [ToolDef](../reference/inspect_ai.tool.html.md#tooldef) \| [ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo) \| [ToolSource](../reference/inspect_ai.tool.html.md#toolsource)\] \| [ToolSource](../reference/inspect_ai.tool.html.md#toolsource) \| None  
Tool definitions (included in token count as they consume context).

`model` str \| [Model](../reference/inspect_ai.model.html.md#model) \| None  
Target model for compacted input (defaults to active model).

`checkpointer` [Checkpointer](../reference/inspect_ai.util.html.md#checkpointer)  
Session checkpointer. The handler’s internal state is captured at each checkpoint fire and restored on resume so the resumed session continues from the prior compacted view rather than re-deriving from scratch. Defaults to a no-op session.

### Compact

Interface for compaction strategies.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/types.py#L73)

``` python
class Compact(Protocol)
```

#### Methods

compact_input  
Compact messages for input to the model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/types.py#L76)

``` python
async def compact_input(
    self,
    messages: list[ChatMessage],
    force: bool = False,
) -> tuple[list[ChatMessage], ChatMessageUser | None]
```

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Full message history.

`force` bool  
If True, perform compaction unconditionally (skip the threshold gate). Used by overflow recovery paths after a model_length error.

record_output  
Record the output from a generate call.

Calibrates the compaction’s token estimation against the actual input token count from `output.usage`. This captures API-level overhead (tool definitions, system messages, thinking configuration) that per-message counting cannot.

`input` must be the messages that were passed to `model.generate` — it determines the baseline message ids that produced `output.usage`. This matters when one [Compact](../reference/inspect_ai.model.html.md#compact) instance is shared across concurrent callers (e.g. via AgentBridge).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/types.py#L93)

``` python
async def record_output(
    self, input: list[ChatMessage], output: ModelOutput
) -> None
```

`input` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
The list of messages that was passed to model.generate.

`output` [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput)  
The ModelOutput from the generate call.

### CompactionStrategy

Compaction strategy.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/types.py#L10)

``` python
class CompactionStrategy(abc.ABC)
```

#### Attributes

`memory` bool  
Whether to warn the model to save content to memory before compaction.

`preserve_prefix` bool  
Instruction to orchestrator: preserve prefix messages in compacted output.

When True (default), the orchestration layer will prepend any prefix messages not already in the compacted output.

When False (native compaction), only system messages are prepended since user content is either preserved by the provider (OpenAI) or semantically encoded in the compaction block (Anthropic).

#### Methods

\_\_init\_\_  
Compaction strategy.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/types.py#L13)

``` python
def __init__(
    self,
    *,
    type: Literal["summary", "edit", "trim"],
    threshold: int | float = 0.9,
    memory: bool = True,
)
```

`type` Literal\['summary', 'edit', 'trim'\]  
Type of compaction performed.

`threshold` int \| float  
Token count or percent of context window to trigger compaction.

`memory` bool  
Warn the model to save critical content to memory prior to compaction when the memory tool is available.

compact  
Compact messages.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/types.py#L57)

``` python
@abc.abstractmethod
async def compact(
    self, model: Model, messages: list[ChatMessage], tools: list[ToolInfo]
) -> tuple[list[ChatMessage], ChatMessageUser | None]
```

`model` [Model](../reference/inspect_ai.model.html.md#model)  
Target model for compaction.

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Full message history

`tools` list\[[ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo)\]  
Available tools

### CompactionAuto

Automatic compaction: tries native first, falls back to summary.

This strategy uses efficient provider-native compaction when available, and falls back to summary-based compaction for unsupported providers or models.

This is the recommended default for most use cases, as it automatically adapts to the capabilities of the underlying provider and model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/auto.py#L24)

``` python
class CompactionAuto(CompactionStrategy)
```

#### Attributes

`preserve_prefix` bool  
Instruction to orchestrator: preserve prefix messages in compacted output.

When True (default), the orchestration layer will prepend any prefix messages not already in the compacted output.

When False (native compaction), only system messages are prepended since user content is either preserved by the provider (OpenAI) or semantically encoded in the compaction block (Anthropic).

`memory` bool  
Whether to warn the model to save content to memory before compaction.

#### Methods

\_\_init\_\_  
Initialize automatic compaction strategy.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/auto.py#L33)

``` python
def __init__(
    self,
    threshold: int | float = 0.9,
    instructions: str | None = None,
    memory: bool | Literal["auto"] = "auto",
) -> None
```

`threshold` int \| float  
Token count or percent of context window to trigger compaction.

`instructions` str \| None  
Additional instructions to give the model about compaction (e.g. “Focus on preserving code snippets, variable names, and technical decisions.”)

`memory` bool \| Literal\['auto'\]  
Whether to warn the model to save critical content to memory prior to compaction. “auto” (default) enables warnings for all compaction paths.

compact  
Compact messages using native compaction with summary fallback.

Attempts native compaction first. If the provider doesn’t support native compaction (NotImplementedError), falls back to summary-based compaction.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/auto.py#L85)

``` python
@override
async def compact(
    self, model: Model, messages: list[ChatMessage], tools: list[ToolInfo]
) -> tuple[list[ChatMessage], ChatMessageUser | None]
```

`model` [Model](../reference/inspect_ai.model.html.md#model)  
Target model for compaction.

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Full message history to compact.

`tools` list\[[ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo)\]  
Available tools.

### CompactionNative

Compaction strategy using provider-native compaction APIs.

This strategy delegates compaction to the model provider’s native compaction endpoint when available (e.g., OpenAI Codex models). For providers without native compaction support, this will raise NotImplementedError. Use [CompactionAuto](../reference/inspect_ai.model.html.md#compactionauto) for automatic fallback to summary-based compaction.

The native compaction approach differs from other strategies (edit, summary, trim) in that: - Compaction is performed server-side by the provider - The compacted representation is opaque (encrypted) and provider-specific - Token savings may be more aggressive while preserving semantic meaning

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/native.py#L22)

``` python
class CompactionNative(CompactionStrategy)
```

#### Attributes

`memory` bool  
Whether to warn the model to save content to memory before compaction.

`preserve_prefix` bool  
Instruction to orchestrator: do not preserve prefix messages.

For native compaction, only system messages are prepended since user content is either preserved by the provider (OpenAI) or semantically encoded in the compaction block (Anthropic).

#### Methods

\_\_init\_\_  
Initialize native compaction strategy.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/native.py#L37)

``` python
def __init__(
    self,
    threshold: int | float = 0.9,
    instructions: str | None = None,
    memory: bool = False,
) -> None
```

`threshold` int \| float  
Token count or percent of context window to trigger compaction.

`instructions` str \| None  
Additional instructions to give the model about compaction (e.g. “Focus on preserving code snippets, variable names, and technical decisions.”)

`memory` bool  
Whether to warn the model to save critical content to memory prior to compaction. Default is False.

compact  
Compact messages using the provider’s native compaction API.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/native.py#L71)

``` python
@override
async def compact(
    self, model: Model, messages: list[ChatMessage], tools: list[ToolInfo]
) -> tuple[list[ChatMessage], ChatMessageUser | None]
```

`model` [Model](../reference/inspect_ai.model.html.md#model)  
Target model for compaction.

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Full message history to compact.

`tools` list\[[ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo)\]  
Available tools.

### CompactionEdit

Message editing compaction.

Compact messages by editing the history to remove tool call results and thinking blocks. Tool results receive placeholder to indicate they were removed.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/edit.py#L27)

``` python
class CompactionEdit(CompactionStrategy)
```

#### Attributes

`memory` bool  
Whether to warn the model to save content to memory before compaction.

`preserve_prefix` bool  
Instruction to orchestrator: preserve prefix messages in compacted output.

When True (default), the orchestration layer will prepend any prefix messages not already in the compacted output.

When False (native compaction), only system messages are prepended since user content is either preserved by the provider (OpenAI) or semantically encoded in the compaction block (Anthropic).

#### Methods

\_\_init\_\_  
Message editing compaction.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/edit.py#L35)

``` python
def __init__(
    self,
    threshold: int | float = 0.9,
    memory: bool = True,
    keep_thinking_turns: Literal["all"] | int = 1,
    keep_tool_uses: int = 3,
    keep_tool_inputs: bool = True,
    exclude_tools: list[str] | None = None,
)
```

`threshold` int \| float  
Token count or percent of context window to trigger compaction.

`memory` bool  
Warn the model to save critical content to memory prior to compaction when the memory tool is available.

`keep_thinking_turns` Literal\['all'\] \| int  
Defines how many recent assistant turns to preserve thinking blocks within. Specify N to keep the thinking blocks within the last N turns, or “all” to keep all thinking blocks. Defaults to 1. Note that some providers (e.g. google) do not support thinking compaction.

`keep_tool_uses` int  
Defines how many recent tool use/result pairs to keep after clearing occurs. The oldest tool interactions are removed first, preserving the most recent ones. Tool output is replaced with placeholder text to let the model know that tool result was removed.

`keep_tool_inputs` bool  
Controls whether the tool call parameters are cleared along with the tool results. By default, only the tool results are cleared while keeping the original tool calls visible. When False, both the tool call and result are removed entirely and replaced with a placeholder text.

`exclude_tools` list\[str\] \| None  
List of tool names whose tool uses and results should never be cleared. Useful for preserving important context.

compact  
Compact messages by editing the history.

Removes tool call results and thinking blocks from older turns. Tool results receive placeholder to indicate they were removed.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/edit.py#L89)

``` python
@override
async def compact(
    self, model: Model, messages: list[ChatMessage], tools: list[ToolInfo]
) -> tuple[list[ChatMessage], ChatMessageUser | None]
```

`model` [Model](../reference/inspect_ai.model.html.md#model)  
Target model for compaction.

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Full message history

`tools` list\[[ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo)\]  
Available tools

### CompactionSummary

Conversation summary compaction.

Compact messages by summarizing the conversation.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/summary.py#L27)

``` python
class CompactionSummary(CompactionStrategy)
```

#### Attributes

`memory` bool  
Whether to warn the model to save content to memory before compaction.

`preserve_prefix` bool  
Instruction to orchestrator: preserve prefix messages in compacted output.

When True (default), the orchestration layer will prepend any prefix messages not already in the compacted output.

When False (native compaction), only system messages are prepended since user content is either preserved by the provider (OpenAI) or semantically encoded in the compaction block (Anthropic).

#### Methods

\_\_init\_\_  
Conversation summary compaction.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/summary.py#L33)

``` python
def __init__(
    self,
    *,
    threshold: int | float = 0.9,
    memory: bool = True,
    model: str | Model | None = None,
    instructions: str | None = None,
    prompt: str | None = None,
)
```

`threshold` int \| float  
Token count or percent of context window to trigger compaction.

`memory` bool  
Warn the model to save critical content to memory prior to compaction when the memory tool is available.

`model` str \| [Model](../reference/inspect_ai.model.html.md#model) \| None  
Model to use for summarization (defaults to compaction target model).

`instructions` str \| None  
Additional instructions to give the model about compaction (e.g. “Focus on preserving code snippets, variable names, and technical decisions.”). These instructions will be inserted into the `prompt`.

`prompt` str \| None  
Prompt to use for summarization (fully replaces the summarization prompt). Include an `{addendums}` placeholder in your prompt to include custom `instructions` and a prompt to use the [memory()](../reference/inspect_ai.tool.html.md#memory) tool when its available.

compact  
Compact messages by summarizing the conversation.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/summary.py#L73)

``` python
@override
async def compact(
    self, model: Model, messages: list[ChatMessage], tools: list[ToolInfo]
) -> tuple[list[ChatMessage], ChatMessageUser | None]
```

`model` [Model](../reference/inspect_ai.model.html.md#model)  
Target model for compaction.

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Full message history

`tools` list\[[ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo)\]  
Available tools

### CompactionTrim

Message trimming compaction.

Compact messages by trimming the history to preserve a percentage of messages: - Retain all system messages. - Retain the ‘input’ messages from the sample. - Preserve a proportion of the remaining messages (`preserve=0.8` by default). - Ensure that all assistant tool calls have corresponding tool messages. - Ensure that the sequence of messages doesn’t end with an assistant message.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/trim.py#L14)

``` python
class CompactionTrim(CompactionStrategy)
```

#### Attributes

`memory` bool  
Whether to warn the model to save content to memory before compaction.

`preserve_prefix` bool  
Instruction to orchestrator: preserve prefix messages in compacted output.

When True (default), the orchestration layer will prepend any prefix messages not already in the compacted output.

When False (native compaction), only system messages are prepended since user content is either preserved by the provider (OpenAI) or semantically encoded in the compaction block (Anthropic).

#### Methods

\_\_init\_\_  
Message trimming compaction.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/trim.py#L25)

``` python
def __init__(
    self,
    *,
    threshold: int | float = 0.9,
    memory: bool = True,
    preserve: float = 0.8,
)
```

`threshold` int \| float  
Token count or percent of context window to trigger compaction.

`memory` bool  
Warn the model to save critical content to memory prior to compaction when the memory tool is available.

`preserve` float  
Ratio of conversation messages to preserve (defaults to 0.8).

compact  
Compact messages by trimming the history to preserve a percentage of messages.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_compaction/trim.py#L49)

``` python
@override
async def compact(
    self, model: Model, messages: list[ChatMessage], tools: list[ToolInfo]
) -> tuple[list[ChatMessage], ChatMessageUser | None]
```

`model` [Model](../reference/inspect_ai.model.html.md#model)  
Target model for compaction.

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Full message history

`tools` list\[[ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo)\]  
Available tools

## Model Info

### get_model_info

Get model information including context window, output tokens, etc.

Looks up model information from a local database. Supports standard Inspect model strings and performs case-insensitive matching.

This function first tries direct database lookup, which does not require provider SDKs to be installed. It only falls back to full provider instantiation if direct lookup fails.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_info.py#L288)

``` python
def get_model_info(model: str | Model) -> ModelInfo | None
```

`model` str \| [Model](../reference/inspect_ai.model.html.md#model)  
Model name or Model instance. Standard Inspect model strings are supported (e.g., “together/meta-llama/Llama-3.1-8B-Instruct”). The model is resolved and its canonical name is used for lookup.

#### Examples

``` python
from inspect_ai.model import get_model_info

info = get_model_info("together/meta-llama/Llama-3.1-8B-Instruct")
if info:
    print(f"Context window: {info.context_length}")
```

### set_model_info

Set custom model information for models not in the database.

Use this to register model information for custom or private models that are not included in the built-in database.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_info.py#L430)

``` python
def set_model_info(model: str, info: ModelInfo) -> None
```

`model` str  
Model name to register (e.g., “my-provider/custom-model”)

`info` [ModelInfo](../reference/inspect_ai.model.html.md#modelinfo)  
ModelInfo object with context_length, output_tokens, etc.

#### Examples

``` python
from inspect_ai.model import set_model_info, ModelInfo

set_model_info(
    "my-provider/custom-model",
    ModelInfo(
        context_length=32000,
        output_tokens=4096,
        organization="My Organization"
    )
)
```

### set_model_cost

Set cost data for a model already in the database.

Looks up the model and updates its cost field. Raises if the model is not found in the database or custom registry.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_info.py#L458)

``` python
def set_model_cost(model: str, cost: ModelCost) -> None
```

`model` str  
Model name (e.g. “openai/gpt-4o”)

`cost` [ModelCost](../reference/inspect_ai.model.html.md#modelcost)  
ModelCost with pricing per million tokens.

### compute_model_cost

Compute cost for a model call based on usage and cost data.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L3113)

``` python
def compute_model_cost(
    cost_data: ModelCost, usage: ModelUsage, cache_ttl: str | None = None
) -> float
```

`cost_data` [ModelCost](../reference/inspect_ai.model.html.md#modelcost)  
Per-token pricing for the model.

`usage` [ModelUsage](../reference/inspect_ai.model.html.md#modelusage)  
Token counts for the call.

`cache_ttl` str \| None  
Prompt cache TTL used for the call (e.g. `"1h"`), for providers that bill longer-lived cache writes at a higher rate. `None` (the default) bills cache writes at `input_cache_write`.

### ModelInfo

Model information and metadata

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_data/model_data.py#L104)

``` python
class ModelInfo(BaseModel)
```

#### Attributes

`organization` str \| None  
Model organization (e.g. Anthropic, OpenAI).

`model` str \| None  
Model name (e.g. Gemini 2.5 Flash).

`snapshot` str \| None  
A snapshot (version) string, if available (e.g. “latest” or “20240229”).

`release_date` UtcDate \| None  
The mode’s release date.

`knowledge_cutoff_date` UtcDate \| None  
The model’s knowledge cutoff date.

`context_length` int \| None  
The model’s context length in tokens.

`output_tokens` int \| None  
“The model’s maximum output tokens.

`reasoning` bool \| None  
Is this a reasoning model.

`reasoning_effort_default` str \| None  
Documented provider default for `reasoning_effort` on this model.

Sourced from the provider’s published documentation. May be one of the standard effort values (`minimal`, `low`, `medium`, `high`, `xhigh`, `max`) or a sentinel such as `adaptive` (Anthropic Claude 4.6+, where the model selects effort per-request) or `fixed` (models without an effort scale, e.g. DeepSeek-R1 and Mistral Magistral). `None` means undocumented.

Inspect does not send this value automatically — it is metadata used to generate the per-model defaults table in the docs.

`family` str \| None  
Reference model name used for capability and request-shape detection.

When set (typically via :func:`set_model_info`), provider capability checks match against this string instead of the configured model name. Use this to make a model with a custom alias behave like a known family. This value does not change the model identifier sent to the provider.

`cost` [ModelCost](../reference/inspect_ai.model.html.md#modelcost) \| None  
Cost per million tokens for this model.

`input_tokens` int \| None  
Effective input capacity in tokens.

Returns the explicit input_tokens value if set in model data, otherwise falls back to context_length.

This provides a single property callers can use without needing to know about context_length vs input capacity differences.

### ModelCost

Model cost in \$/million tokens.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_data/model_data.py#L10)

``` python
class ModelCost(BaseModel)
```

#### Attributes

`input` float  
Price per million input tokens.

`output` float  
Price per million output tokens.

`input_cache_write` float  
Price per million input tokens written to cache.

Record the provider’s default-TTL rate here (for Anthropic, the 5-minute rate). Providers that bill longer cache TTLs at a higher rate (e.g. Anthropic’s 1-hour writes at 2x base input) are adjusted at cost computation time based on the configured TTL — do not pre-bake a longer-TTL rate into this field or it will be double-applied.

`input_cache_read` float  
Price per million input tokens read from cache.

### model_roles

Model roles.

Get the model roles defined for the current task. A role maps to a single [Model](../reference/inspect_ai.model.html.md#model), or to a list of models when a list was assigned to the role (e.g. via the `model_roles` argument to [eval()](../reference/inspect_ai.html.md#eval)). Call this method only within a running solver or agent execution (it’s not available during task construction).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L2857)

``` python
def model_roles() -> dict[str, Model | list[Model]]
```

## Logprobs

### Logprob

Log probability for a token.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_output.py#L201)

``` python
class Logprob(BaseModel)
```

#### Attributes

`token` str  
The predicted token represented as a string.

`logprob` float  
The log probability value of the model for the predicted token.

`bytes` list\[int\] \| None  
The predicted token represented as a byte array (a list of integers).

`top_logprobs` list\[[TopLogprob](../reference/inspect_ai.model.html.md#toplogprob)\] \| None  
If the `top_logprobs` argument is greater than 0, this will contain an ordered list of the top K most likely tokens and their log probabilities.

### Logprobs

Log probability information for a completion choice.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_output.py#L217)

``` python
class Logprobs(BaseModel)
```

#### Attributes

`content` list\[[Logprob](../reference/inspect_ai.model.html.md#logprob)\]  
a (num_generated_tokens,) length list containing the individual log probabilities for each generated token.

### TopLogprob

List of the most likely tokens and their log probability, at this token position.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model_output.py#L188)

``` python
class TopLogprob(BaseModel)
```

#### Attributes

`token` str  
The top-kth token represented as a string.

`logprob` float  
The log probability value of the model for the top-kth token.

`bytes` list\[int\] \| None  
The top-kth token represented as a byte array (a list of integers).

## Caching

### CachePolicy

Caching options for model generation.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_cache.py#L58)

``` python
class CachePolicy(BaseModel)
```

#### Attributes

`expiry` str \| None  
The expiry time for cache entries (Default “1W”). This is a string of the format “12h” for 12 hours or “1W” for a week, etc. This is how long we will keep the cache entry, if we access it after this point we’ll clear it. Setting to `None` will cache indefinitely.

`per_epoch` bool  
Default True. By default we cache responses separately for different epochs. The general use case is that if there are multiple epochs, we should cache each response separately because scorers will aggregate across epochs. However, sometimes a response can be cached regardless of epoch if the call being made isn’t under test as part of the evaluation. If False, this option allows you to bypass that and cache independently of the epoch.

`scopes` dict\[str, str\]  
A dictionary of additional metadata that should be included in the cache key. This allows for more fine-grained control over the cache key generation.

### cache_size

Calculate the size of various cached directories and files

If neither `subdirs` nor `files` are provided, the entire cache directory will be calculated.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_cache.py#L384)

``` python
def cache_size(
    subdirs: list[str] = [], files: list[Path] = []
) -> list[tuple[str, int]]
```

`subdirs` list\[str\]  
List of folders to filter by, which are generally model names. Empty directories will be ignored.

`files` list\[Path\]  
List of files to filter by explicitly. Note that return value group these up by their parent directory

### cache_clear

Clear the cache directory.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_cache.py#L299)

``` python
def cache_clear(model: str = "") -> bool
```

`model` str  
Model to clear cache for.

### cache_list_expired

Returns a list of all the cached files that have passed their expiry time.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_cache.py#L413)

``` python
def cache_list_expired(filter_by: list[str] = []) -> list[Path]
```

`filter_by` list\[str\]  
Default \[\]. List of model names to filter by. If an empty list, this will search the entire cache.

### cache_prune

Delete all expired cache entries.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_cache.py#L453)

``` python
def cache_prune(files: list[Path] = []) -> None
```

`files` list\[Path\]  
List of files to prune. If empty, this will search the entire cache.

### cache_path

Path to cache directory.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_cache.py#L319)

``` python
def cache_path(model: str = "") -> Path
```

`model` str  
Path to cache directory for specific model.

## Conversion

### messages_from_openai

Convert OpenAI Completions API messages into Inspect messages.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_openai_convert.py#L32)

``` python
async def messages_from_openai(
    messages: "list[ChatCompletionMessageParam]",
    model: str | None = None,
) -> list[ChatMessage]
```

`messages` 'list\[ChatCompletionMessageParam\]'  
OpenAI Completions API Messages

`model` str \| None  
Optional model name to tag assistant messages with.

### messages_from_openai_responses

Convert OpenAI Responses API messages into Inspect messages.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_openai_convert.py#L49)

``` python
async def messages_from_openai_responses(
    messages: "list[ResponseInputItemParam]",
    model: str | None = None,
) -> list[ChatMessage]
```

`messages` 'list\[ResponseInputItemParam\]'  
OpenAI Responses API Messages

`model` str \| None  
Optional model name to tag assistant messages with.

### messages_from_anthropic

Convert OpenAI Responses API messages into Inspect messages.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_anthropic_convert.py#L12)

``` python
async def messages_from_anthropic(
    messages: "list[MessageParam]", system_message: str | None = None
) -> list[ChatMessage]
```

`messages` list\[MessageParam\]  
OpenAI Responses API Messages

`system_message` str \| None  
System message accompanying messages (optional).

### messages_from_google

Convert Google GenAI Content list into Inspect messages.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_google_convert.py#L35)

``` python
async def messages_from_google(
    contents: "Sequence[Content | ContentDict]",
    system_instruction: str | None = None,
    model: str | None = None,
) -> list[ChatMessage]
```

`contents` Sequence\[[Content](../reference/inspect_ai.model.html.md#content) \| ContentDict\]  
Google GenAI Content objects or dicts that can be converted.

`system_instruction` str \| None  
Optional system instruction string.

`model` str \| None  
Optional model name to tag assistant messages with.

### model_output_from_openai

Convert OpenAI ChatCompletion into Inspect [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput)

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_openai_convert.py#L81)

``` python
async def model_output_from_openai(
    completion: Union["ChatCompletion", dict[str, Any]],
) -> ModelOutput
```

`completion` 'ChatCompletion' \| dict\[str, Any\]  
OpenAI `ChatCompletion` object or dict that can converted into one.

### model_output_from_openai_responses

Convert OpenAI `Response` into Inspect [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput)

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_openai_convert.py#L108)

``` python
async def model_output_from_openai_responses(
    response: Union["Response", dict[str, Any]],
) -> ModelOutput
```

`response` 'Response' \| dict\[str, Any\]  
OpenAI `Response` object or dict that can converted into one.

### model_output_from_anthropic

Convert Anthropic Message response into Inspect [ModelOutput](../reference/inspect_ai.model.html.md#modeloutput)

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_anthropic_convert.py#L33)

``` python
async def model_output_from_anthropic(
    message: Union["Message", dict[str, Any]],
) -> ModelOutput
```

`message` Message \| dict\[str, Any\]  
Anthropic `Message` object or dict that can converted into one.

### model_output_from_google

Convert Google GenerateContentResponse into Inspect ModelOutput.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_google_convert.py#L72)

``` python
async def model_output_from_google(
    response: Union["GenerateContentResponse", dict[str, Any]],
    model: str | None = None,
) -> ModelOutput
```

`response` GenerateContentResponse \| dict\[str, Any\]  
Google GenerateContentResponse object or dict that can be converted.

`model` str \| None  
Optional model name override.

### messages_to_openai

Convert messages to OpenAI Completions API compatible messages.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_openai_convert.py#L15)

``` python
async def messages_to_openai(
    messages: list[ChatMessage],
    system_role: Literal["user", "system", "developer"] = "system",
) -> "list[ChatCompletionMessageParam]"
```

`messages` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
List of messages to convert

`system_role` Literal\['user', 'system', 'developer'\]  
Role to use for system messages (newer OpenAI models use “developer” rather than “system”).

## Provider

### modelapi

Decorator for registering model APIs.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_registry.py#L30)

``` python
def modelapi(name: str) -> Callable[..., type[ModelAPI]]
```

`name` str  
Name of API

### ModelAPI

Model API provider.

If you are implementing a custom ModelAPI provider your `__init__()` method will also receive a `**model_args` parameter that will carry any custom `model_args` (or `-M` arguments from the CLI) specified by the user. You can then pass these on to the approriate place in your model initialisation code (for example, here is what many of the built-in providers do with the `model_args` passed to them: <https://inspect.aisi.org.uk/models.html#model-args>)

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L275)

``` python
class ModelAPI(abc.ABC)
```

#### Attributes

`qualified_model_name` str \| None  
Full `provider/model` name (the string `Model.__str__` renders).

Stamped by [get_model()](../reference/inspect_ai.model.html.md#get_model) right after construction — `model_name` is the provider-stripped name, and process registries keyed by model (e.g. the throughput registry) need the qualified form. None for a ModelAPI constructed outside [get_model()](../reference/inspect_ai.model.html.md#get_model), in which case such registries simply don’t attribute this instance’s traffic.

#### Methods

\_\_init\_\_  
Create a model API provider.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L297)

``` python
def __init__(
    self,
    model_name: str,
    base_url: str | None = None,
    api_key: str | None = None,
    api_key_vars: list[str] = [],
    config: GenerateConfig = GenerateConfig(),
) -> None
```

`model_name` str  
Model name.

`base_url` str \| None  
Alternate base URL for model.

`api_key` str \| None  
API key for model.

`api_key_vars` list\[str\]  
Environment variables that may contain keys for this provider (used for override)

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig)  
Model configuration.

initialize  
Reinitialize the model API client.

This can be used to reinitialize the API keys.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L357)

``` python
def initialize(self) -> None
```

refresh_credentials  
Refresh credentials after an authentication failure.

Providers that can update credentials in place should override this method to avoid interrupting concurrent requests using their client.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L364)

``` python
async def refresh_credentials(self) -> None
```

aclose  
Async close method for closing any client allocated for the model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L373)

``` python
async def aclose(self) -> None
```

close  
Sync close method for closing any client allocated for the model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L377)

``` python
def close(self) -> None
```

canonical_name  
Canonical model name for querying results.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L388)

``` python
def canonical_name(self) -> str
```

service_model_name  
Model name used by the provider service.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L392)

``` python
def service_model_name(self) -> str
```

input_tokens_name  
Model name used for looking up model input tokens.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L396)

``` python
def input_tokens_name(self) -> str
```

model_family  
Model name used only for capability and request-shape detection.

Returns :attr:`ModelInfo.family` if one has been registered for this model via :func:`set_model_info` under the configured or canonical model name, otherwise falls back to the model name sent to the provider service. The returned name must not be used as the wire model identifier.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L400)

``` python
def model_family(self) -> str
```

cache_write_ttl  
Prompt-cache TTL billed for cache writes in the current call context.

Consulted when recording usage after each generate/compact call (“1h” bills cache writes at a higher rate than the default 5m). Providers that bill cache writes at a TTL-dependent rate override this; the TTL may vary per call, so it is a method rather than an attribute.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L417)

``` python
def cache_write_ttl(self) -> str | None
```

generate  
Generate output from the model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L427)

``` python
@abc.abstractmethod
async def generate(
    self,
    input: list[ChatMessage],
    tools: list[ToolInfo],
    tool_choice: ToolChoice,
    config: GenerateConfig,
) -> ModelOutput | tuple[ModelOutput | Exception, ModelCall]
```

`input` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Chat message input (if a `str` is passed it is converted to a `ChatUserMessage`).

`tools` list\[[ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo)\]  
Tools available for the model to call.

`tool_choice` [ToolChoice](../reference/inspect_ai.tool.html.md#toolchoice)  
Directives to the model as to which tools to prefer.

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig)  
Model configuration.

count_tokens  
Estimate token count for input.

This default implementation uses character-based heuristics for text and size-based estimates for media. Model providers can override `count_text_tokens()` and `count_media_tokens()` for more accurate results, or override this method entirely to use their native token counting APIs.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L450)

``` python
async def count_tokens(
    self,
    input: str | list[ChatMessage],
    config: GenerateConfig | None = None,
) -> int
```

`input` str \| list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Input to count tokens for.

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig) \| None  
Optional generation config for provider-specific counting (e.g., reasoning parameters that affect token allocation).

count_text_tokens  
Estimate tokens from text using tiktoken (o200k_base with 10% buffer).

Override this method to use model-specific tokenizers.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L474)

``` python
async def count_text_tokens(self, text: str) -> int
```

`text` str  
Text to count.

count_media_tokens  
Estimate tokens for media content (images, audio, video, documents).

For data URIs, estimates are based on decoded size. For URLs/file paths, uses conservative fixed fallbacks. Override this method for provider-specific media token calculations.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L484)

``` python
async def count_media_tokens(
    self, media: ContentImage | ContentAudio | ContentVideo | ContentDocument
) -> int
```

`media` [ContentImage](../reference/inspect_ai.model.html.md#contentimage) \| [ContentAudio](../reference/inspect_ai.model.html.md#contentaudio) \| [ContentVideo](../reference/inspect_ai.model.html.md#contentvideo) \| [ContentDocument](../reference/inspect_ai.model.html.md#contentdocument)  
Media content to count tokens for.

tokenize  
Tokenize text into token IDs using the model’s tokenizer.

Override in providers that support server-side tokenization (e.g. vLLM’s `/tokenize` endpoint).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L498)

``` python
async def tokenize(self, text: str) -> list[int]
```

`text` str  
Text to tokenize.

max_tokens  
Default max_tokens.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L516)

``` python
def max_tokens(self) -> int | None
```

max_tokens_for_config  
Default max_tokens for a given config.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L520)

``` python
def max_tokens_for_config(self, config: GenerateConfig) -> int | None
```

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig)  
Generation config.

max_connections  
Default max_connections.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L531)

``` python
def max_connections(self) -> int
```

connection_key  
Scope for enforcement of max_connections (and adaptive concurrency).

Two instances of the *same provider* that return the same key share one connection pool. This method only needs to distinguish accounts/models within a provider; the model layer adds the provider namespace on top (see `_connection_pool_key`), so distinct providers never collide even when their `connection_key()` values coincide.

Providers that scope by API key should use `self.initial_api_key` here, NOT the live `self.api_key`: the live key can rotate mid-eval (e.g. a credential hook refreshing it via `initialize()`), and keying on it would discard all learned pool/adaptive state on every rotation. `initial_api_key` is fixed at construction, so the scope stays stable.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L535)

``` python
def connection_key(self) -> str
```

apply_redacted_reasoning_tokens_to_input  
Whether compaction should add `redacted_reasoning_tokens` to its input estimate.

Override and return True for providers whose `usage.input_tokens` omits redacted reasoning content on re-injection (e.g., OpenAI Responses with `store=false` + `include=["reasoning.encrypted_content"]`).

Note: bridge-mediated workloads (`agent_bridge`) reconstruct messages from provider-native input and lose this metadata, so the predictive correction does not apply there. Bridge users rely on the reactive `model_length` recovery.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L552)

``` python
def apply_redacted_reasoning_tokens_to_input(self) -> bool
```

should_retry  
Should this exception be retried?

Returns either a plain `bool` (any True is treated as a transient retry by the adaptive controller) or a [RetryDecision](../reference/inspect_ai.model.html.md#retrydecision) to additionally classify the retry as `rate_limit` vs `transient` and to pass through any server-suggested `retry_after`. Built-in providers return [RetryDecision](../reference/inspect_ai.model.html.md#retrydecision) so the adaptive controller scales only on real rate-limit signals.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L566)

``` python
def should_retry(self, ex: Exception) -> bool | RetryDecision
```

`ex` Exception  
Exception to check for retry

is_auth_failure  
Check if this exception indicates an authentication failure.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L584)

``` python
def is_auth_failure(self, ex: Exception) -> bool
```

`ex` Exception  
Exception to check for authentication failure

collapse_user_messages  
Collapse consecutive user messages into a single message.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L595)

``` python
def collapse_user_messages(self) -> bool
```

collapse_assistant_messages  
Collapse consecutive assistant messages into a single message.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L599)

``` python
def collapse_assistant_messages(self) -> bool
```

collapse_system_messages  
Collapse consecutive system messages into a single message.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L603)

``` python
def collapse_system_messages(self) -> bool
```

tools_required  
Any tool use in a message stream means that tools must be passed.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L607)

``` python
def tools_required(self) -> bool
```

supports_remote_mcp  
Does this provider support remote execution of MCP tools?.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L611)

``` python
def supports_remote_mcp(self) -> bool
```

tool_result_images  
Tool results can contain images

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L615)

``` python
def tool_result_images(self) -> bool
```

tool_result_documents  
Tool results can be replayed to the model with documents.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L619)

``` python
def tool_result_documents(self) -> bool
```

disable_computer_screenshot_truncation  
Some models do not support truncation of computer screenshots.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L623)

``` python
def disable_computer_screenshot_truncation(self) -> bool
```

force_reasoning_history  
Force a specific reasoning history behavior for this provider.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L627)

``` python
def force_reasoning_history(self) -> Literal["none", "all", "last"] | None
```

auto_reasoning_history  
Behavior to use for reasoning_history=‘auto’

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L631)

``` python
def auto_reasoning_history(self) -> Literal["none", "all", "last"]
```

compact_reasoning_history  
Is reasoning history eligible for compation for this provider?

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L635)

``` python
def compact_reasoning_history(self) -> bool
```

compact  
Compact messages using provider-native compaction.

Some model providers (e.g., OpenAI Codex models) support native context compaction, which reduces the token count of a conversation while preserving semantic meaning. This is useful for long conversations that approach the context window limit.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/model/_model.py#L639)

``` python
async def compact(
    self,
    input: list[ChatMessage],
    tools: list[ToolInfo],
    config: GenerateConfig,
    instructions: str | None = None,
) -> tuple[list[ChatMessage], ModelUsage | None]
```

`input` list\[[ChatMessage](../reference/inspect_ai.model.html.md#chatmessage)\]  
Chat message input (if a `str` is passed it is converted to a `ChatUserMessage`).

`tools` list\[[ToolInfo](../reference/inspect_ai.tool.html.md#toolinfo)\]  
Tools available for the model to call.

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig)  
Model configuration.

`instructions` str \| None  
Additional instructions to give the model about compaction (e.g. “Focus on preserving code snippets, variable names, and technical decisions.”)
