# inspect_ai – Inspect

## Evaluation

### eval

Evaluate tasks using a Model.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/eval.py#L124)

``` python
def eval(
    tasks: Tasks,
    model: str | Model | list[str] | list[Model] | None | NotGiven = ...,
    model_base_url: str | None = ...,
    model_args: dict[str, Any] | str = ...,
    model_roles: ModelRoles | None = ...,
    task_args: dict[str, Any] | str = ...,
    sandbox: SandboxEnvironmentType | None = ...,
    sandbox_cleanup: bool | None = ...,
    sandbox_prebuilt: bool | None = ...,
    checkpoint: CheckpointConfig | bool | None = ...,
    acp_server: bool | int | str | None = ...,
    ctl_server: bool | str | None = ...,
    solver: Solver | SolverSpec | Agent | list[Solver] | None = ...,
    scanner: Scanners | None = ...,
    tags: list[str] | None = ...,
    metadata: dict[str, Any] | None = ...,
    trace: bool | None = ...,
    display: DisplayType | None = ...,
    approval: str | list[ApprovalPolicy] | ApprovalPolicyConfig | None = ...,
    review: str | list[ReviewPolicy] | ReviewPolicyConfig | None = ...,
    notification: bool | str | None = ...,
    log_level: str | None = ...,
    log_level_transcript: str | None = ...,
    log_dir: str | None = ...,
    log_format: Literal['eval', 'json'] | None = ...,
    limit: int | tuple[int, int] | None = ...,
    sample_id: str | int | list[str] | list[int] | list[str | int] | None = ...,
    sample_shuffle: bool | int | None = ...,
    epochs: int | Epochs | None = ...,
    fail_on_error: bool | float | None = ...,
    continue_on_fail: bool | None = ...,
    retry_on_error: int | None = ...,
    score_on_error: bool | None = ...,
    debug_errors: bool | None = ...,
    message_limit: int | None = ...,
    token_limit: int | str | TokenLimit | None = ...,
    turn_limit: int | None = ...,
    time_limit: int | None = ...,
    working_limit: int | None = ...,
    cost_limit: float | None = ...,
    model_cost_config: str | dict[str, ModelCost] | None = ...,
    max_samples: int | None = ...,
    max_dataset_memory: int | None = ...,
    max_tasks: int | None = ...,
    max_subprocesses: int | None = ...,
    max_sandboxes: int | None = ...,
    log_samples: bool | None = ...,
    log_realtime: bool | None = ...,
    log_images: bool | None = ...,
    log_model_api: bool | None = ...,
    log_refusals: bool | None = ...,
    log_buffer: int | None = ...,
    log_shared: bool | int | None = ...,
    log_header_only: bool | None = ...,
    run_samples: bool = ...,
    score: bool = ...,
    score_display: bool | None = ...,
    eval_set_id: str | None = ...,
    eval_set_tasks: list[str] | None = ...,
    scan_id: str | None = ...,
    task_retry_attempts: int | None = ...,
    *,
    max_retries: int | None = ...,
    timeout: int | None = ...,
    attempt_timeout: int | None = ...,
    stream_idle_timeout: int | None = ...,
    max_connections: int | None = ...,
    adaptive_connections: bool | int | AdaptiveConcurrency | None = ...,
    system_message: str | None = ...,
    max_tokens: int | None = ...,
    top_p: float | None = ...,
    temperature: float | None = ...,
    stop_seqs: list[str] | None = ...,
    best_of: int | None = ...,
    frequency_penalty: float | None = ...,
    presence_penalty: float | None = ...,
    logit_bias: dict[int, float] | None = ...,
    seed: int | None = ...,
    top_k: int | None = ...,
    num_choices: int | None = ...,
    logprobs: bool | None = ...,
    top_logprobs: int | None = ...,
    prompt_logprobs: int | None = ...,
    parallel_tool_calls: bool | None = ...,
    internal_tools: bool | None = ...,
    max_tool_output: int | None = ...,
    cache_prompt: Literal['auto'] | bool | None = ...,
    fallback_models: list[str] | None = ...,
    fail_on_refusal: bool | None = ...,
    verbosity: Literal['low', 'medium', 'high'] | None = ...,
    effort: Literal['low', 'medium', 'high', 'xhigh', 'max'] | None = ...,
    reasoning_effort: Literal['none', 'minimal', 'low', 'medium', 'high', 'xhigh', 'max'] | None = ...,
    reasoning_mode: Literal['standard', 'pro'] | None = ...,
    reasoning_tokens: int | None = ...,
    reasoning_summary: Literal['none', 'concise', 'detailed', 'auto'] | None = ...,
    reasoning_history: Literal['none', 'all', 'last', 'auto'] | None = ...,
    response_schema: ResponseSchema | None = ...,
    extra_headers: dict[str, str] | None = ...,
    extra_body: dict[str, Any] | None = ...,
    modalities: list[OutputModality] | None = ...,
    cache: bool | CachePolicy | None = ...,
    batch: bool | int | BatchConfig | None = ...,
) -> list[EvalLog]
```

`tasks` [Tasks](../reference/inspect_ai.html.md#tasks)  
Task(s) to evaluate. If None, attempt to evaluate a task in the current working directory

`model` str \| [Model](../reference/inspect_ai.model.html.md#model) \| list\[str\] \| list\[[Model](../reference/inspect_ai.model.html.md#model)\] \| None \| NotGiven  
Model(s) for evaluation. If not specified use the value of the INSPECT_EVAL_MODEL environment variable. Specify `None` to define no default model(s), which will leave model usage entirely up to tasks.

`model_base_url` str \| None  
Base URL for communicating with the model API.

`model_args` dict\[str, Any\] \| str  
Model creation args (as a dictionary or as a path to a JSON or YAML config file)

`model_roles` [ModelRoles](../reference/inspect_ai.model.html.md#modelroles) \| None  
Named roles for use in [get_model()](../reference/inspect_ai.model.html.md#get_model) (a role can also map to a list of models).

`task_args` dict\[str, Any\] \| str  
Task creation arguments (as a dictionary or as a path to a JSON or YAML config file)

`sandbox` SandboxEnvironmentType \| None  
Sandbox environment type (or optionally a str or tuple with a shorthand spec)

`sandbox_cleanup` bool \| None  
Cleanup sandbox environments after task completes (defaults to True)

`sandbox_prebuilt` bool \| None  
Treat sandbox images as prebuilt, skipping builds and failing at task startup when an image is missing (defaults to False)

`checkpoint` [CheckpointConfig](../reference/inspect_ai.util.html.md#checkpointconfig) \| bool \| None  
Checkpoint configuration for this eval, or `True` to enable checkpointing with the default trigger (every 500k tokens) — equivalent to the bare `--checkpoint` CLI flag. Overrides any task- or sample-level `checkpoint` that enables checkpointing when set. A task can opt out with `Task(checkpoint=False)`, which overrides this enable for that task only.

`acp_server` bool \| int \| str \| None  
Expose this eval over an Agent Client Protocol server. `True` enables a default AF_UNIX socket at `<inspect_data_dir>/acp/<pid>.sock`; an integer binds a TCP loopback port; a string is taken as a custom UNIX socket path; `None` (default) does not start an ACP server.

`ctl_server` bool \| str \| None  
Control-channel server for this eval process. `True` or `None` (default) binds the default AF_UNIX socket; `False` disables the control endpoint; `"keep"` additionally keeps the process running after the eval finishes so external clients can still query its state — exit via `inspect ctl process release` (or `POST /release`).

`solver` [Solver](../reference/inspect_ai.solver.html.md#solver) \| [SolverSpec](../reference/inspect_ai.solver.html.md#solverspec) \| [Agent](../reference/inspect_ai.agent.html.md#agent) \| list\[[Solver](../reference/inspect_ai.solver.html.md#solver)\] \| None  
Alternative solver for task(s). Optional (uses task solver by default).

`scanner` [Scanners](../reference/inspect_ai.html.md#scanners) \| None  
Scanner(s) to apply to each sample’s transcript after the sample completes.

`tags` list\[str\] \| None  
Tags to associate with this evaluation run.

`metadata` dict\[str, Any\] \| None  
Metadata to associate with this evaluation run.

`trace` bool \| None  
Trace message interactions with evaluated model to terminal.

`display` [DisplayType](../reference/inspect_ai.util.html.md#displaytype) \| None  
Task display type (defaults to ‘full’).

`approval` str \| list\[[ApprovalPolicy](../reference/inspect_ai.approval.html.md#approvalpolicy)\] \| ApprovalPolicyConfig \| None  
Tool use approval policies. Either a path to an approval policy config file, an ApprovalPolicyConfig, or a list of approval policies. Defaults to no approval policy.

`review` str \| list\[[ReviewPolicy](../reference/inspect_ai.review.html.md#reviewpolicy)\] \| ReviewPolicyConfig \| None  
Tool result review policies. Either a path to a review policy config file, a ReviewPolicyConfig, or a list of review policies. Defaults to no review policy.

`notification` bool \| str \| None  
Enable out-of-band notifications when a human-in-the-loop interaction (`ask_user`, human approval) is posted. Pass `True` to send via the URL(s) in the `INSPECT_EVAL_NOTIFICATION` environment variable (single URL, comma-separated list, or path to an Apprise config file). Alternatively pass a path to an Apprise YAML/text config file. URLs are not accepted directly so secrets never end up in source code, shell history, process listings, or eval logs. Requires the `apprise` package.

`log_level` str \| None  
Level for logging to the console: “debug”, “http”, “sandbox”, “info”, “warning”, “error”, “critical”, or “notset” (defaults to “warning”)

`log_level_transcript` str \| None  
Level for logging to the log file (defaults to “info”)

`log_dir` str \| None  
Output path for logging results (defaults to file log in ./logs directory).

`log_format` Literal\['eval', 'json'\] \| None  
Format for writing log files (defaults to “eval”, the native high-performance format).

`limit` int \| tuple\[int, int\] \| None  
Limit evaluated samples (defaults to all samples).

`sample_id` str \| int \| list\[str\] \| list\[int\] \| list\[str \| int\] \| None  
Evaluate specific sample(s) from the dataset. Use plain ids or preface with task names as required to disambiguate ids across tasks (e.g. `popularity:10`); a prefix that names no task in the run is part of the id, and an empty list selects no samples.

`sample_shuffle` bool \| int \| None  
Shuffle order of samples (pass a seed to make the order deterministic).

`epochs` int \| [Epochs](../reference/inspect_ai.html.md#epochs) \| None  
Epochs to repeat samples for and optional score reducer function(s) used to combine sample scores (defaults to “mean”)

`fail_on_error` bool \| float \| None  
`True` to fail on first sample error (default); `False` to never fail on sample errors; Value between 0 and 1 to fail if a proportion of total samples fails. Value greater than 1 to fail eval if a count of samples fails.

`continue_on_fail` bool \| None  
`True` to continue running and only fail at the end if the `fail_on_error` condition is met. `False` to fail eval immediately when the `fail_on_error` condition is met (default).

`retry_on_error` int \| None  
Number of times to retry samples if they encounter errors (by default, no retries occur).

`score_on_error` bool \| None  
Score samples that error rather than failing the eval mid-run. Errors still count toward the `fail_on_error` threshold for marking the eval log as ‘error’. Only takes effect after retries (if any) are exhausted.

`debug_errors` bool \| None  
Raise task errors (rather than logging them) so they can be debugged (defaults to False).

`message_limit` int \| None  
Limit on total messages used for each sample.

`token_limit` int \| str \| [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) \| None  
Limit on tokens used for each sample. An `int` (or a [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) with type “all”) limits total tokens; a [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) with a `type` limits by output tokens or an arithmetic formula over `input`/`output`. Also accepts strings like “500k”, “1m”, “output:1m”, or “(input\*0.1)+output:1m”.

`turn_limit` int \| None  
Limit on total turns (model generations) used for each sample.

`time_limit` int \| None  
Limit on clock time (in seconds) for samples.

`working_limit` int \| None  
Limit on working time (in seconds) for sample. Working time includes model generation, tool calls, etc. but does not include time spent waiting on retries or shared resources.

`cost_limit` float \| None  
Limit on total cost (in dollars) for each sample. Requires model cost data via set_model_cost() or –model-cost-config.

`model_cost_config` str \| dict\[str, [ModelCost](../reference/inspect_ai.model.html.md#modelcost)\] \| None  
YAML or JSON file with model prices for cost tracking or dict of model -\> [ModelCost](../reference/inspect_ai.model.html.md#modelcost)

`max_samples` int \| None  
Maximum number of samples to run in parallel within each task (default is max_connections)

`max_dataset_memory` int \| None  
Maximum MB of dataset sample data to hold in memory per task. When exceeded, samples are paged to a temporary file on disk (defaults to None, which keeps all samples in memory).

`max_tasks` int \| None  
Maximum number of tasks to run in parallel (defaults to number of models being evaluated)

`max_subprocesses` int \| None  
Maximum number of subprocesses to run in parallel (default is the number of processors available to the eval)

`max_sandboxes` int \| None  
Maximum number of sandboxes (per-provider) to run in parallel.

`log_samples` bool \| None  
Log detailed samples and scores (defaults to True)

`log_realtime` bool \| None  
Log events in realtime (enables live viewing of samples in inspect view). Defaults to True.

`log_images` bool \| None  
Log base64 encoded version of images, even if specified as a filename or URL (defaults to False)

`log_model_api` bool \| None  
Log raw model api requests and responses. True logs all calls, False logs only errors, None (default) logs the first few calls per model plus errors.

`log_refusals` bool \| None  
Log warnings for model refusals.

`log_buffer` int \| None  
Number of samples to buffer before writing log file. If not specified, an appropriate default for the format and filesystem is chosen (10 for most all cases, 100 for JSON logs on remote filesystems).

`log_shared` bool \| int \| None  
Sync sample events to log directory so that users on other systems can see log updates in realtime (defaults to no syncing). Specify `True` to sync every 10 seconds, otherwise an integer to sync every `n` seconds.

`log_header_only` bool \| None  
If `True`, the function should return only log headers rather than full logs with samples (defaults to `False`).

`run_samples` bool  
Run samples. If `False`, a log with `status=="started"` and an empty `samples` list is returned.

`score` bool  
Score output (defaults to True)

`score_display` bool \| None  
Show scoring metrics in realtime (defaults to True)

`eval_set_id` str \| None  
Unique id for eval set (this is passed from [eval_set()](../reference/inspect_ai.html.md#eval_set) and should not be specified directly).

`eval_set_tasks` list\[str\] \| None  
Names of every task in the eval set, so `task:id` sample selectors resolve the same way for a retried subset of tasks (this is passed from [eval_set()](../reference/inspect_ai.html.md#eval_set) and should not be specified directly).

`scan_id` str \| None  
Override the scan-dir identifier (defaults to `eval_set_id` or `run_id`). Set by `eval_retry` to reuse the original eval’s scan dir.

`task_retry_attempts` int \| None  
Number of times to retry tasks (defaults to 0)

`max_retries` int \| None  
Maximum number of times to retry request, so e.g. 1 allows two attempts total (defaults to unlimited).

`timeout` int \| None  
Request timeout (in seconds).

`attempt_timeout` int \| None  
Timeout (in seconds) for any given attempt (if exceeded, will abandon attempt and retry according to max_retries).

`stream_idle_timeout` int \| None  
Timeout (in seconds) on silence within a streaming response — if a streaming attempt delivers no chunk for this long, the attempt is abandoned and retried according to max_retries. Setting it requests streaming (like on_stream); it has no effect on calls that do not stream.

`max_connections` int \| None  
Maximum number of concurrent connections to Model API (default is model specific).

`adaptive_connections` bool \| int \| [AdaptiveConcurrency](../reference/inspect_ai.util.html.md#adaptiveconcurrency) \| None  
Adaptive concurrency for model API connections. Defaults to enabled (`None` and `True` both resolve to `AdaptiveConcurrency()` defaults: min=10, start=20, max=100). Pass `False` to opt out (uses static concurrency). Pass an integer `N` as shorthand for `AdaptiveConcurrency(max=N)`. Pass an [AdaptiveConcurrency](../reference/inspect_ai.util.html.md#adaptiveconcurrency) to fully customize bounds and tuning (cooldown_seconds, decrease_factor, scale_up_percent). An explicit `max_connections` or `batch=True` takes precedence and uses static concurrency.

`system_message` str \| None  
Override the default system message.

`max_tokens` int \| None  
The maximum number of tokens that can be generated in the completion (default is model specific).

`top_p` float \| None  
An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass.

`temperature` float \| None  
What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

`stop_seqs` list\[str\] \| None  
Sequences where the API will stop generating further tokens. The returned text will not contain the stop sequence.

`best_of` int \| None  
Generates best_of completions server-side and returns the ‘best’ (the one with the highest log probability per token). vLLM only.

`frequency_penalty` float \| None  
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model’s likelihood to repeat the same line verbatim. OpenAI, Google, Grok, Groq, and vLLM only.

`presence_penalty` float \| None  
Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model’s likelihood to talk about new topics. OpenAI, Google, Grok, Groq, and vLLM only.

`logit_bias` dict\[int, float\] \| None  
Map token Ids to an associated bias value from -100 to 100 (e.g. “42=10,43=-10”). OpenAI and Grok only.

`seed` int \| None  
Random seed. OpenAI, Google, Mistral, Groq, HuggingFace, and vLLM only.

`top_k` int \| None  
Randomly sample the next word from the top_k most likely next words. Anthropic, Google, and HuggingFace only.

`num_choices` int \| None  
How many chat completion choices to generate for each input message. OpenAI, Grok, Google, and TogetherAI only.

`logprobs` bool \| None  
Return log probabilities of the output tokens. OpenAI, Google, Grok, TogetherAI, Huggingface, llama-cpp-python, and vLLM only.

`top_logprobs` int \| None  
Number of most likely tokens (0-20) to return at each token position, each with an associated log probability. OpenAI, Google, Grok, and Huggingface only.

`prompt_logprobs` int \| None  
Number of log probabilities to return per prompt token (1-20). When greater than 1, top-N alternative tokens are also returned. vLLM only.

`parallel_tool_calls` bool \| None  
Whether to enable parallel function calling during tool use (defaults to True). OpenAI and Groq only.

`internal_tools` bool \| None  
Whether to automatically map tools to model internal implementations (e.g. ‘computer’ for anthropic).

`max_tool_output` int \| None  
Maximum tool output (in bytes). Defaults to 16 \* 1024.

`cache_prompt` Literal\['auto'\] \| bool \| None  
Whether to cache the prompt prefix. Enabled by default. Set to False to disable: on Anthropic and Bedrock Converse (Claude and Nova) this turns off the provider’s own automatic caching; on OpenAI it only turns off explicit `cache_breakpoint` marks — the model’s own implicit caching still applies regardless. Use `ContentText(cache_breakpoint=True)` to mark an explicit cache boundary (e.g. a fixed rubric ahead of a varying item) instead of relying on automatic caching; Anthropic and OpenAI `gpt-5.6`+ only.

`fallback_models` list\[str\] \| None  
Fallback models tried in order when the model’s safety classifiers refuse the request. Anthropic Claude API only (not supported on Bedrock/Vertex/Azure or with batch mode).

`fail_on_refusal` bool \| None  
Raise a `ModelRefusalError` (failing the sample) when the model returns `stop_reason="content_filter"`. Defaults to False.

`verbosity` Literal\['low', 'medium', 'high'\] \| None  
Constrains the verbosity of the model’s response. Lower values will result in more concise responses, while higher values will result in more verbose responses. GPT 5.x models only (defaults to “medium” for OpenAI models).

`effort` Literal\['low', 'medium', 'high', 'xhigh', 'max'\] \| None  
Control how many tokens are used for a response, trading off between response thoroughness and token efficiency. Anthropic Claude Opus 4.5+ only (`max` only supported on 4.6 and 4.7, `xhigh` supported only on 4.7).

`reasoning_effort` Literal\['none', 'minimal', 'low', 'medium', 'high', 'xhigh', 'max'\] \| None  
Constrains effort on reasoning. Defaults vary by provider and model and not all models support all values (please consult provider documentation for details).

`reasoning_mode` Literal\['standard', 'pro'\] \| None  
Reasoning mode. “pro” performs more model work for greater reliability on difficult tasks, at higher latency and token usage. OpenAI GPT-5.6+ models only (“standard” is the default).

`reasoning_tokens` int \| None  
Maximum number of tokens to use for reasoning. Anthropic Claude models only.

`reasoning_summary` Literal\['none', 'concise', 'detailed', 'auto'\] \| None  
Provide summary of reasoning steps (OpenAI reasoning models only). Use ‘auto’ to access the most detailed summarizer available for the current model (defaults to ‘auto’ if your organization is verified by OpenAI).

`reasoning_history` Literal\['none', 'all', 'last', 'auto'\] \| None  
Include reasoning in chat message history sent to generate.

`response_schema` [ResponseSchema](../reference/inspect_ai.model.html.md#responseschema) \| None  
Request a response format as JSONSchema (output should still be validated). OpenAI, Google, and Mistral only.

`extra_headers` dict\[str, str\] \| None  
Extra headers to be sent with requests. Not supported for AzureAI, Bedrock, and Grok.

`extra_body` dict\[str, Any\] \| None  
Extra body to be sent with requests to OpenAI compatible servers. OpenAI, vLLM, and SGLang only.

`modalities` list\[[OutputModality](../reference/inspect_ai.model.html.md#outputmodality)\] \| None  
Additional output modalities to enable beyond text (e.g. \[“image”\]). OpenAI and Google only.

`cache` bool \| [CachePolicy](../reference/inspect_ai.model.html.md#cachepolicy) \| None  
Policy for caching of model generations.

`batch` bool \| int \| [BatchConfig](../reference/inspect_ai.model.html.md#batchconfig) \| None  
Use batching API when available. True to enable batching with default configuration, False to disable batching, a number to enable batching of the specified batch size, or a BatchConfig object specifying the batching configuration.

### eval_retry

Retry a previously failed evaluation task.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/eval.py#L1307)

``` python
def eval_retry(
    tasks: str | EvalLogInfo | EvalLog | list[str] | list[EvalLogInfo] | list[EvalLog],
    log_level: str | None = None,
    log_level_transcript: str | None = None,
    log_dir: str | None = None,
    log_format: Literal["eval", "json"] | None = None,
    max_samples: int | None = None,
    max_tasks: int | None = None,
    max_subprocesses: int | None = None,
    max_sandboxes: int | None = None,
    sandbox_cleanup: bool | None = None,
    sandbox_prebuilt: bool | None = None,
    trace: bool | None = None,
    display: DisplayType | None = None,
    fail_on_error: bool | float | None = None,
    continue_on_fail: bool | None = None,
    retry_on_error: int | None = None,
    score_on_error: bool | None = None,
    debug_errors: bool | None = None,
    log_samples: bool | None = None,
    log_realtime: bool | None = None,
    log_images: bool | None = None,
    log_model_api: bool | None = None,
    log_refusals: bool | None = None,
    log_buffer: int | None = None,
    log_shared: bool | int | None = None,
    score: bool = True,
    score_display: bool | None = None,
    acp_server: bool | int | str | None = None,
    ctl_server: bool | str | None = None,
    scanner: "Scanners | None" = None,
    max_retries: int | None = None,
    timeout: int | None = None,
    attempt_timeout: int | None = None,
    stream_idle_timeout: int | None = None,
    max_connections: int | None = None,
    adaptive_connections: bool | int | AdaptiveConcurrency | None = None,
    checkpoint: CheckpointConfig | bool | None = None,
    incomplete_action: IncompleteAction = "retry",
    incomplete_max: int | float | None = None,
) -> list[EvalLog]
```

`tasks` str \| [EvalLogInfo](../reference/inspect_ai.log.html.md#evalloginfo) \| [EvalLog](../reference/inspect_ai.log.html.md#evallog) \| list\[str\] \| list\[[EvalLogInfo](../reference/inspect_ai.log.html.md#evalloginfo)\] \| list\[[EvalLog](../reference/inspect_ai.log.html.md#evallog)\]  
Log files for task(s) to retry.

`log_level` str \| None  
Level for logging to the console: “debug”, “http”, “sandbox”, “info”, “warning”, “error”, “critical”, or “notset” (defaults to “warning”)

`log_level_transcript` str \| None  
Level for logging to the log file (defaults to “info”)

`log_dir` str \| None  
Output path for logging results (defaults to file log in ./logs directory).

`log_format` Literal\['eval', 'json'\] \| None  
Format for writing log files (defaults to “eval”, the native high-performance format).

`max_samples` int \| None  
Maximum number of samples to run in parallel within each task (default is max_connections)

`max_tasks` int \| None  
Maximum number of tasks to run in parallel (defaults to number of models being evaluated)

`max_subprocesses` int \| None  
Maximum number of subprocesses to run in parallel (default is the number of processors available to the eval)

`max_sandboxes` int \| None  
Maximum number of sandboxes (per-provider) to run in parallel.

`sandbox_cleanup` bool \| None  
Cleanup sandbox environments after task completes (defaults to True)

`sandbox_prebuilt` bool \| None  
Treat sandbox images as prebuilt, skipping builds and failing at task startup when an image is missing (defaults to False)

`trace` bool \| None  
Trace message interactions with evaluated model to terminal.

`display` [DisplayType](../reference/inspect_ai.util.html.md#displaytype) \| None  
Task display type (defaults to ‘full’).

`fail_on_error` bool \| float \| None  
`True` to fail on a sample error (default); `False` to never fail on sample errors; Value between 0 and 1 to fail if a proportion of total samples fails. Value greater than 1 to fail eval if a count of samples fails.

`continue_on_fail` bool \| None  
`True` to continue running and only fail at the end if the `fail_on_error` condition is met. `False` to fail eval immediately when the `fail_on_error` condition is met (default).

`retry_on_error` int \| None  
Number of times to retry samples if they encounter errors (by default, no retries occur).

`score_on_error` bool \| None  
Score samples that error rather than failing the eval mid-run. Errors still count toward the `fail_on_error` threshold for marking the eval log as ‘error’. Only takes effect after retries (if any) are exhausted.

`debug_errors` bool \| None  
Raise task errors (rather than logging them) so they can be debugged (defaults to False).

`log_samples` bool \| None  
Log detailed samples and scores (defaults to True)

`log_realtime` bool \| None  
Log events in realtime (enables live viewing of samples in inspect view). Defaults to True.

`log_images` bool \| None  
Log base64 encoded version of images, even if specified as a filename or URL (defaults to False)

`log_model_api` bool \| None  
Log raw model api requests and responses. True logs all calls, False logs only errors, None (default) logs the first few calls per model plus errors.

`log_refusals` bool \| None  
Log warnings for model refusals.

`log_buffer` int \| None  
Number of samples to buffer before writing log file. If not specified, an appropriate default for the format and filesystem is chosen (10 for most all cases, 100 for JSON logs on remote filesystems).

`log_shared` bool \| int \| None  
Sync sample events to log directory so that users on other systems can see log updates in realtime (defaults to no syncing). Specify `True` to sync every 10 seconds, otherwise an integer to sync every `n` seconds.

`score` bool  
Score output (defaults to True)

`score_display` bool \| None  
Show scoring metrics in realtime (defaults to True)

`acp_server` bool \| int \| str \| None  
Override the original eval’s ACP server transport on retry. `True` enables a default AF_UNIX socket; an integer binds a TCP loopback port; a string is taken as a custom UNIX socket path; `None` (default) replays whatever transport (or no transport) was persisted in the original log’s `EvalConfig.acp_server`.

`ctl_server` bool \| str \| None  
Control-channel server for this eval process. `True` or `None` (default) binds the default AF_UNIX socket; `False` disables the control endpoint; `"keep"` additionally keeps the process running after the eval finishes so external clients can still query its state — exit via `inspect ctl process release` (or `POST /release`).

`scanner` [Scanners](../reference/inspect_ai.html.md#scanners) \| None  
Scanner(s) to apply to each sample’s transcript after the sample completes. When provided, the existing scan dir from the original eval (keyed by its `eval_set_id` or `run_id`) is reused — same resume contract as `eval_set`: matching scanner config attaches, divergent config raises `PrerequisiteError`.

`max_retries` int \| None  
Maximum number of times to retry request.

`timeout` int \| None  
Request timeout (in seconds)

`attempt_timeout` int \| None  
Timeout (in seconds) for any given attempt (if exceeded, will abandon attempt and retry according to max_retries).

`stream_idle_timeout` int \| None  
Timeout (in seconds) on silence within a streaming response (if a streaming attempt delivers no chunk for this long, will abandon attempt and retry according to max_retries).

`max_connections` int \| None  
Maximum number of concurrent connections to Model API (default is per Model API)

`adaptive_connections` bool \| int \| [AdaptiveConcurrency](../reference/inspect_ai.util.html.md#adaptiveconcurrency) \| None  
Adaptive concurrency for Model API connections. Defaults to enabled (resolves to `AdaptiveConcurrency()` defaults: min=10, start=20, max=100). Pass `False` to opt out (uses static concurrency), an integer `N` as shorthand for `AdaptiveConcurrency(max=N)`, or an [AdaptiveConcurrency](../reference/inspect_ai.util.html.md#adaptiveconcurrency) to fully customize bounds and tuning (cooldown_seconds, decrease_factor, scale_up_percent). An explicit `max_connections` or `batch=True` takes precedence and uses static concurrency.

`checkpoint` [CheckpointConfig](../reference/inspect_ai.util.html.md#checkpointconfig) \| bool \| None  
Checkpoint configuration for this retry, or `True` to enable checkpointing with the default trigger (every 500k tokens). Must match the config used on the original eval for resume detection to find the checkpoint files (the original `--checkpoint` is not recorded in the log file).

`incomplete_action` [IncompleteAction](../reference/inspect_ai.log.html.md#incompleteaction)  
Disposition applied when recovering a crashed log before retrying, for samples that were in progress at crash. `"retry"` (default) re-runs them; `"error"` resolves them as operator terminations — if that leaves every expected sample final, the recovered log finalizes as `status="success"` and is returned without retrying. A finalized log lives at `<name>-recovered.eval` alongside the crashed log rather than in `log_dir`, so read its location from `EvalLog.location`.

`incomplete_max` int \| float \| None  
Safety threshold for `incomplete_action="error"` (count if \>= 1, or proportion of expected samples if strictly less than 1, so `1.0` means one sample, not 100%): when more than this many samples are in progress, fall back to the default recover-and-retry behavior. Has no effect (a warning is logged) with `incomplete_action="retry"`.

### eval_set

Evaluate a set of tasks.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/evalset.py#L229)

``` python
def eval_set(
    tasks: Tasks,
    log_dir: str | None = ...,
    retry_attempts: int | None = ...,
    retry_wait: float | None = ...,
    retry_connections: float | None = ...,
    retry_cleanup: bool | None = ...,
    retry_immediate: bool | None = ...,
    incomplete_action: IncompleteAction = ...,
    incomplete_max: int | float | None = ...,
    model: str | Model | list[str] | list[Model] | None | NotGiven = ...,
    model_base_url: str | None = ...,
    model_args: dict[str, Any] | str = ...,
    model_roles: ModelRoles | None = ...,
    task_args: dict[str, Any] | str = ...,
    sandbox: SandboxEnvironmentType | None = ...,
    sandbox_cleanup: bool | None = ...,
    sandbox_prebuilt: bool | None = ...,
    checkpoint: CheckpointConfig | bool | None = ...,
    acp_server: bool | int | str | None = ...,
    ctl_server: bool | str | None = ...,
    solver: Solver | SolverSpec | Agent | list[Solver] | None = ...,
    scanner: Scanners | None = ...,
    tags: list[str] | None = ...,
    metadata: dict[str, Any] | None = ...,
    trace: bool | None = ...,
    display: DisplayType | None = ...,
    approval: str | list[ApprovalPolicy] | ApprovalPolicyConfig | None = ...,
    review: str | list[ReviewPolicy] | ReviewPolicyConfig | None = ...,
    notification: bool | str | None = ...,
    score: bool = ...,
    score_display: bool | None = ...,
    log_level: str | None = ...,
    log_level_transcript: str | None = ...,
    log_format: Literal['eval', 'json'] | None = ...,
    limit: int | tuple[int, int] | None = ...,
    sample_id: str | int | list[str] | list[int] | list[str | int] | None = ...,
    sample_shuffle: bool | int | None = ...,
    epochs: int | Epochs | None = ...,
    fail_on_error: bool | float | None = ...,
    continue_on_fail: bool | None = ...,
    retry_on_error: int | None = ...,
    score_on_error: bool | None = ...,
    debug_errors: bool | None = ...,
    message_limit: int | None = ...,
    token_limit: int | str | TokenLimit | None = ...,
    turn_limit: int | None = ...,
    time_limit: int | None = ...,
    working_limit: int | None = ...,
    cost_limit: float | None = ...,
    model_cost_config: str | dict[str, ModelCost] | None = ...,
    max_samples: int | None = ...,
    max_dataset_memory: int | None = ...,
    max_tasks: int | None = ...,
    max_subprocesses: int | None = ...,
    max_sandboxes: int | None = ...,
    log_samples: bool | None = ...,
    log_realtime: bool | None = ...,
    log_images: bool | None = ...,
    log_model_api: bool | None = ...,
    log_refusals: bool | None = ...,
    log_buffer: int | None = ...,
    log_shared: bool | int | None = ...,
    bundle_dir: str | None = ...,
    bundle_overwrite: bool = ...,
    log_dir_allow_dirty: bool | None = ...,
    eval_set_id: str | None = ...,
    embed_viewer: bool = ...,
    *,
    max_retries: int | None = ...,
    timeout: int | None = ...,
    attempt_timeout: int | None = ...,
    stream_idle_timeout: int | None = ...,
    max_connections: int | None = ...,
    adaptive_connections: bool | int | AdaptiveConcurrency | None = ...,
    system_message: str | None = ...,
    max_tokens: int | None = ...,
    top_p: float | None = ...,
    temperature: float | None = ...,
    stop_seqs: list[str] | None = ...,
    best_of: int | None = ...,
    frequency_penalty: float | None = ...,
    presence_penalty: float | None = ...,
    logit_bias: dict[int, float] | None = ...,
    seed: int | None = ...,
    top_k: int | None = ...,
    num_choices: int | None = ...,
    logprobs: bool | None = ...,
    top_logprobs: int | None = ...,
    prompt_logprobs: int | None = ...,
    parallel_tool_calls: bool | None = ...,
    internal_tools: bool | None = ...,
    max_tool_output: int | None = ...,
    cache_prompt: Literal['auto'] | bool | None = ...,
    fallback_models: list[str] | None = ...,
    fail_on_refusal: bool | None = ...,
    verbosity: Literal['low', 'medium', 'high'] | None = ...,
    effort: Literal['low', 'medium', 'high', 'xhigh', 'max'] | None = ...,
    reasoning_effort: Literal['none', 'minimal', 'low', 'medium', 'high', 'xhigh', 'max'] | None = ...,
    reasoning_mode: Literal['standard', 'pro'] | None = ...,
    reasoning_tokens: int | None = ...,
    reasoning_summary: Literal['none', 'concise', 'detailed', 'auto'] | None = ...,
    reasoning_history: Literal['none', 'all', 'last', 'auto'] | None = ...,
    response_schema: ResponseSchema | None = ...,
    extra_headers: dict[str, str] | None = ...,
    extra_body: dict[str, Any] | None = ...,
    modalities: list[OutputModality] | None = ...,
    cache: bool | CachePolicy | None = ...,
    batch: bool | int | BatchConfig | None = ...,
) -> tuple[bool, list[EvalLog]]
```

`tasks` [Tasks](../reference/inspect_ai.html.md#tasks)  
Task(s) to evaluate. If None, attempt to evaluate a task in the current working directory

`log_dir` str \| None  
Output path for logging results (defaults to INSPECT_LOG_DIR or ./logs). The directory is the eval set’s storage scope, so a set that shares one with another set shares its results.

`retry_attempts` int \| None  
Maximum number of retry attempts before giving up (defaults to 10).

`retry_wait` float \| None  
Time to wait between attempts when `retry_immediate=False`, increased exponentially (defaults to 30, resulting in waits of 30, 60, 120, 240, etc.). Wait time per-retry will in no case be longer than 1 hour. Ignored when `retry_immediate=True`.

`retry_connections` float \| None  
Reduce max_connections at this rate with each retry when `retry_immediate=False` (defaults to 1.0, which results in no reduction). Ignored when `retry_immediate=True`.

`retry_cleanup` bool \| None  
Cleanup failed log files after retries (defaults to True)

`retry_immediate` bool \| None  
If True (the default), immediately retry tasks as they fail without waiting for all tasks to complete; completed samples are reused from logs on retry. If False, wait for all tasks to complete before retrying any tasks (legacy batch-retry behavior). When True, `retry_wait` and `retry_connections` are ignored.

`incomplete_action` [IncompleteAction](../reference/inspect_ai.log.html.md#incompleteaction)  
Disposition applied when recovering a crashed log from a previous execution, for samples that were in progress at crash. `"retry"` (default) re-runs them; `"error"` resolves them as operator terminations — if that leaves every expected sample final, the recovered log finalizes as `status="success"`, the task classifies as complete, and nothing re-runs.

`incomplete_max` int \| float \| None  
Safety threshold for `incomplete_action="error"` (count if \>= 1, or proportion of expected samples if strictly less than 1, so `1.0` means one sample, not 100%): when more than this many samples are in progress, fall back to the default recover-and-retry behavior. Has no effect (a warning is logged) with `incomplete_action="retry"`.

`model` str \| [Model](../reference/inspect_ai.model.html.md#model) \| list\[str\] \| list\[[Model](../reference/inspect_ai.model.html.md#model)\] \| None \| NotGiven  
Model(s) for evaluation. If not specified use the value of the INSPECT_EVAL_MODEL environment variable. Specify `None` to define no default model(s), which will leave model usage entirely up to tasks.

`model_base_url` str \| None  
Base URL for communicating with the model API.

`model_args` dict\[str, Any\] \| str  
Model creation args (as a dictionary or as a path to a JSON or YAML config file)

`model_roles` [ModelRoles](../reference/inspect_ai.model.html.md#modelroles) \| None  
Named roles for use in [get_model()](../reference/inspect_ai.model.html.md#get_model) (a role can also map to a list of models).

`task_args` dict\[str, Any\] \| str  
Task creation arguments (as a dictionary or as a path to a JSON or YAML config file)

`sandbox` SandboxEnvironmentType \| None  
Sandbox environment type (or optionally a str or tuple with a shorthand spec)

`sandbox_cleanup` bool \| None  
Cleanup sandbox environments after task completes (defaults to True)

`sandbox_prebuilt` bool \| None  
Treat sandbox images as prebuilt, skipping builds and failing at task startup when an image is missing (defaults to False)

`checkpoint` [CheckpointConfig](../reference/inspect_ai.util.html.md#checkpointconfig) \| bool \| None  
Checkpoint configuration for this eval set, or `True` to enable checkpointing with the default trigger (every 500k tokens). Overrides any task- or sample-level `checkpoint` when set. A task can opt out with `Task(checkpoint=False)`, which overrides this enable for that task only.

`acp_server` bool \| int \| str \| None  
Override the original eval’s ACP server transport on retry. `True` enables a default AF_UNIX socket; an integer binds a TCP loopback port; a string is taken as a custom UNIX socket path; `None` (default) replays whatever transport (or no transport) was persisted in the original log’s `EvalConfig.acp_server`.

`ctl_server` bool \| str \| None  
Control-channel server for this eval-set process. `True` or `None` (default) binds the default AF_UNIX socket; `False` disables the control endpoint; `"keep"` additionally keeps the process running after the eval-set finishes so external clients (the `inspect ctl` CLI, scripted agents, TUIs) can still query state and read results — exit via `inspect ctl process release` (or `POST /release`). Requires `retry_immediate=True` (the default) for the `"keep"` value.

`solver` [Solver](../reference/inspect_ai.solver.html.md#solver) \| [SolverSpec](../reference/inspect_ai.solver.html.md#solverspec) \| [Agent](../reference/inspect_ai.agent.html.md#agent) \| list\[[Solver](../reference/inspect_ai.solver.html.md#solver)\] \| None  
Alternative solver(s) for evaluating task(s). Optional (uses task solver by default).

`scanner` [Scanners](../reference/inspect_ai.html.md#scanners) \| None  
Scanner(s) to apply to each sample’s transcript after the sample completes.

`tags` list\[str\] \| None  
Tags to associate with this evaluation run.

`metadata` dict\[str, Any\] \| None  
Metadata to associate with this evaluation run.

`trace` bool \| None  
Trace message interactions with evaluated model to terminal.

`display` [DisplayType](../reference/inspect_ai.util.html.md#displaytype) \| None  
Task display type (defaults to ‘full’).

`approval` str \| list\[[ApprovalPolicy](../reference/inspect_ai.approval.html.md#approvalpolicy)\] \| ApprovalPolicyConfig \| None  
Tool use approval policies. Either a path to an approval policy config file, an ApprovalPolicyConfig, or a list of approval policies. Defaults to no approval policy.

`review` str \| list\[[ReviewPolicy](../reference/inspect_ai.review.html.md#reviewpolicy)\] \| ReviewPolicyConfig \| None  
Tool result review policies. Either a path to a review policy config file, a ReviewPolicyConfig, or a list of review policies. Defaults to no review policy.

`notification` bool \| str \| None  
Enable out-of-band notifications when a human-in-the-loop interaction (`ask_user`, human approval) is posted. Pass `True` to send via the URL(s) in the `INSPECT_EVAL_NOTIFICATION` environment variable (single URL, comma-separated list, or path to an Apprise config file). Alternatively pass a path to an Apprise YAML/text config file. URLs are not accepted directly so secrets never end up in source code, shell history, process listings, or eval logs. Requires the `apprise` package.

`score` bool  
Score output (defaults to True)

`score_display` bool \| None  
Show scoring metrics in realtime (defaults to True)

`log_level` str \| None  
Level for logging to the console: “debug”, “http”, “sandbox”, “info”, “warning”, “error”, “critical”, or “notset” (defaults to “warning”)

`log_level_transcript` str \| None  
Level for logging to the log file (defaults to “info”)

`log_format` Literal\['eval', 'json'\] \| None  
Format for writing log files (defaults to “eval”, the native high-performance format).

`limit` int \| tuple\[int, int\] \| None  
Limit evaluated samples (defaults to all samples).

`sample_id` str \| int \| list\[str\] \| list\[int\] \| list\[str \| int\] \| None  
Evaluate specific sample(s) from the dataset. Use plain ids or preface with task names as required to disambiguate ids across tasks (e.g. `popularity:10`); a prefix that names no task in the run is part of the id, and an empty list selects no samples.

`sample_shuffle` bool \| int \| None  
Shuffle order of samples (pass a seed to make the order deterministic).

`epochs` int \| [Epochs](../reference/inspect_ai.html.md#epochs) \| None  
Epochs to repeat samples for and optional score reducer function(s) used to combine sample scores (defaults to “mean”)

`fail_on_error` bool \| float \| None  
`True` to fail on first sample error (default); `False` to never fail on sample errors; Value between 0 and 1 to fail if a proportion of total samples fails. Value greater than 1 to fail eval if a count of samples fails.

`continue_on_fail` bool \| None  
`True` to continue running and only fail at the end if the `fail_on_error` condition is met. `False` to fail eval immediately when the `fail_on_error` condition is met (default).

`retry_on_error` int \| None  
Number of times to retry samples if they encounter errors (by default, no retries occur).

`score_on_error` bool \| None  
Score samples that error rather than failing the eval mid-run. Errors still count toward the `fail_on_error` threshold for marking the eval log as ‘error’. Only takes effect after retries (if any) are exhausted.

`debug_errors` bool \| None  
Raise task errors (rather than logging them) so they can be debugged (defaults to False).

`message_limit` int \| None  
Limit on total messages used for each sample.

`token_limit` int \| str \| [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) \| None  
Limit on tokens used for each sample. An `int` (or a [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) with type “all”) limits total tokens; a [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) with a `type` limits by output tokens or an arithmetic formula over `input`/`output`. Also accepts strings like “500k”, “1m”, “output:1m”, or “(input\*0.1)+output:1m”.

`turn_limit` int \| None  
Limit on total turns (model generations) used for each sample.

`time_limit` int \| None  
Limit on clock time (in seconds) for samples.

`working_limit` int \| None  
Limit on working time (in seconds) for sample. Working time includes model generation, tool calls, etc. but does not include time spent waiting on retries or shared resources.

`cost_limit` float \| None  
Limit on total cost (in dollars) for each sample. Requires model cost data via set_model_cost() or –model-cost-config.

`model_cost_config` str \| dict\[str, [ModelCost](../reference/inspect_ai.model.html.md#modelcost)\] \| None  
YAML or JSON file with model prices for cost tracking.

`max_samples` int \| None  
Maximum number of samples to run in parallel within each task (default is max_connections)

`max_dataset_memory` int \| None  
Maximum MB of dataset sample data to hold in memory per task. When exceeded, samples are paged to a temporary file on disk (defaults to None, which keeps all samples in memory).

`max_tasks` int \| None  
Maximum number of tasks to run in parallel (defaults to the greater of 10 and the number of models being evaluated)

`max_subprocesses` int \| None  
Maximum number of subprocesses to run in parallel (default is the number of processors available to the eval)

`max_sandboxes` int \| None  
Maximum number of sandboxes (per-provider) to run in parallel.

`log_samples` bool \| None  
Log detailed samples and scores (defaults to True)

`log_realtime` bool \| None  
Log events in realtime (enables live viewing of samples in inspect view). Defaults to True.

`log_images` bool \| None  
Log base64 encoded version of images, even if specified as a filename or URL (defaults to False)

`log_model_api` bool \| None  
Log raw model api requests and responses. Note that error requests/responses are always logged.

`log_refusals` bool \| None  
Log warnings for model refusals.

`log_buffer` int \| None  
Number of samples to buffer before writing log file. If not specified, an appropriate default for the format and filesystem is chosen (10 for most all cases, 100 for JSON logs on remote filesystems).

`log_shared` bool \| int \| None  
Sync sample events to log directory so that users on other systems can see log updates in realtime (defaults to no syncing). Specify `True` to sync every 10 seconds, otherwise an integer to sync every `n` seconds.

`bundle_dir` str \| None  
If specified, the log viewer and logs generated by this eval set will be bundled into this directory.

`bundle_overwrite` bool  
Whether to overwrite files in the bundle_dir. (defaults to False).

`log_dir_allow_dirty` bool \| None  
If True, allow the log directory to contain unrelated logs. If False, ensure that the log directory only contains logs for tasks in this eval set (defaults to False).

`eval_set_id` str \| None  
ID for the eval set. If not specified, a unique ID will be generated.

`embed_viewer` bool  
If True, embed a log viewer into the log directory.

`max_retries` int \| None  
Maximum number of times to retry request, so e.g. 1 allows two attempts total (defaults to unlimited).

`timeout` int \| None  
Request timeout (in seconds).

`attempt_timeout` int \| None  
Timeout (in seconds) for any given attempt (if exceeded, will abandon attempt and retry according to max_retries).

`stream_idle_timeout` int \| None  
Timeout (in seconds) on silence within a streaming response — if a streaming attempt delivers no chunk for this long, the attempt is abandoned and retried according to max_retries. Setting it requests streaming (like on_stream); it has no effect on calls that do not stream.

`max_connections` int \| None  
Maximum number of concurrent connections to Model API (default is model specific).

`adaptive_connections` bool \| int \| [AdaptiveConcurrency](../reference/inspect_ai.util.html.md#adaptiveconcurrency) \| None  
Adaptive concurrency for model API connections. Defaults to enabled (`None` and `True` both resolve to `AdaptiveConcurrency()` defaults: min=10, start=20, max=100). Pass `False` to opt out (uses static concurrency). Pass an integer `N` as shorthand for `AdaptiveConcurrency(max=N)`. Pass an [AdaptiveConcurrency](../reference/inspect_ai.util.html.md#adaptiveconcurrency) to fully customize bounds and tuning (cooldown_seconds, decrease_factor, scale_up_percent). An explicit `max_connections` or `batch=True` takes precedence and uses static concurrency.

`system_message` str \| None  
Override the default system message.

`max_tokens` int \| None  
The maximum number of tokens that can be generated in the completion (default is model specific).

`top_p` float \| None  
An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass.

`temperature` float \| None  
What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

`stop_seqs` list\[str\] \| None  
Sequences where the API will stop generating further tokens. The returned text will not contain the stop sequence.

`best_of` int \| None  
Generates best_of completions server-side and returns the ‘best’ (the one with the highest log probability per token). vLLM only.

`frequency_penalty` float \| None  
Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model’s likelihood to repeat the same line verbatim. OpenAI, Google, Grok, Groq, and vLLM only.

`presence_penalty` float \| None  
Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model’s likelihood to talk about new topics. OpenAI, Google, Grok, Groq, and vLLM only.

`logit_bias` dict\[int, float\] \| None  
Map token Ids to an associated bias value from -100 to 100 (e.g. “42=10,43=-10”). OpenAI and Grok only.

`seed` int \| None  
Random seed. OpenAI, Google, Mistral, Groq, HuggingFace, and vLLM only.

`top_k` int \| None  
Randomly sample the next word from the top_k most likely next words. Anthropic, Google, and HuggingFace only.

`num_choices` int \| None  
How many chat completion choices to generate for each input message. OpenAI, Grok, Google, and TogetherAI only.

`logprobs` bool \| None  
Return log probabilities of the output tokens. OpenAI, Google, Grok, TogetherAI, Huggingface, llama-cpp-python, and vLLM only.

`top_logprobs` int \| None  
Number of most likely tokens (0-20) to return at each token position, each with an associated log probability. OpenAI, Google, Grok, and Huggingface only.

`prompt_logprobs` int \| None  
Number of log probabilities to return per prompt token (1-20). When greater than 1, top-N alternative tokens are also returned. vLLM only.

`parallel_tool_calls` bool \| None  
Whether to enable parallel function calling during tool use (defaults to True). OpenAI and Groq only.

`internal_tools` bool \| None  
Whether to automatically map tools to model internal implementations (e.g. ‘computer’ for anthropic).

`max_tool_output` int \| None  
Maximum tool output (in bytes). Defaults to 16 \* 1024.

`cache_prompt` Literal\['auto'\] \| bool \| None  
Whether to cache the prompt prefix. Enabled by default. Set to False to disable: on Anthropic and Bedrock Converse (Claude and Nova) this turns off the provider’s own automatic caching; on OpenAI it only turns off explicit `cache_breakpoint` marks — the model’s own implicit caching still applies regardless. Use `ContentText(cache_breakpoint=True)` to mark an explicit cache boundary (e.g. a fixed rubric ahead of a varying item) instead of relying on automatic caching; Anthropic and OpenAI `gpt-5.6`+ only.

`fallback_models` list\[str\] \| None  
Fallback models tried in order when the model’s safety classifiers refuse the request. Anthropic Claude API only (not supported on Bedrock/Vertex/Azure or with batch mode).

`fail_on_refusal` bool \| None  
Raise a `ModelRefusalError` (failing the sample) when the model returns `stop_reason="content_filter"`. Defaults to False.

`verbosity` Literal\['low', 'medium', 'high'\] \| None  
Constrains the verbosity of the model’s response. Lower values will result in more concise responses, while higher values will result in more verbose responses. GPT 5.x models only (defaults to “medium” for OpenAI models).

`effort` Literal\['low', 'medium', 'high', 'xhigh', 'max'\] \| None  
Control how many tokens are used for a response, trading off between response thoroughness and token efficiency. Anthropic Claude Opus 4.5+ only (`max` only supported on 4.6 and 4.7, `xhigh` supported only on 4.7).

`reasoning_effort` Literal\['none', 'minimal', 'low', 'medium', 'high', 'xhigh', 'max'\] \| None  
Constrains effort on reasoning. Defaults vary by provider and model and not all models support all values (please consult provider documentation for details).

`reasoning_mode` Literal\['standard', 'pro'\] \| None  
Reasoning mode. “pro” performs more model work for greater reliability on difficult tasks, at higher latency and token usage. OpenAI GPT-5.6+ models only (“standard” is the default).

`reasoning_tokens` int \| None  
Maximum number of tokens to use for reasoning. Anthropic Claude models only.

`reasoning_summary` Literal\['none', 'concise', 'detailed', 'auto'\] \| None  
Provide summary of reasoning steps (OpenAI reasoning models only). Use ‘auto’ to access the most detailed summarizer available for the current model (defaults to ‘auto’ if your organization is verified by OpenAI).

`reasoning_history` Literal\['none', 'all', 'last', 'auto'\] \| None  
Include reasoning in chat message history sent to generate.

`response_schema` [ResponseSchema](../reference/inspect_ai.model.html.md#responseschema) \| None  
Request a response format as JSONSchema (output should still be validated). OpenAI, Google, and Mistral only.

`extra_headers` dict\[str, str\] \| None  
Extra headers to be sent with requests. Not supported for AzureAI, Bedrock, and Grok.

`extra_body` dict\[str, Any\] \| None  
Extra body to be sent with requests to OpenAI compatible servers. OpenAI, vLLM, and SGLang only.

`modalities` list\[[OutputModality](../reference/inspect_ai.model.html.md#outputmodality)\] \| None  
Additional output modalities to enable beyond text (e.g. \[“image”\]). OpenAI and Google only.

`cache` bool \| [CachePolicy](../reference/inspect_ai.model.html.md#cachepolicy) \| None  
Policy for caching of model generations.

`batch` bool \| int \| [BatchConfig](../reference/inspect_ai.model.html.md#batchconfig) \| None  
Use batching API when available. True to enable batching with default configuration, False to disable batching, a number to enable batching of the specified batch size, or a BatchConfig object specifying the batching configuration.

### score

Score an evaluation log.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/score.py#L81)

``` python
def score(
    log: EvalLog,
    scorers: "Scorers",
    metrics: list[Metric | dict[str, list[Metric]]]
    | dict[str, list[Metric]]
    | None = None,
    epochs_reducer: ScoreReducers | None = None,
    model: str | Model | None = None,
    model_roles: ModelRoles | None = None,
    action: ScoreAction | None = None,
    display: DisplayType | None = None,
    copy: bool = True,
) -> EvalLog
```

`log` [EvalLog](../reference/inspect_ai.log.html.md#evallog)  
Evaluation log.

`scorers` 'Scorers'  
List of Scorers to apply to log

`metrics` list\[[Metric](../reference/inspect_ai.scorer.html.md#metric) \| dict\[str, list\[[Metric](../reference/inspect_ai.scorer.html.md#metric)\]\]\] \| dict\[str, list\[[Metric](../reference/inspect_ai.scorer.html.md#metric)\]\] \| None  
Alternative metrics (overrides the metrics provided by the specified scorer and log).

`epochs_reducer` ScoreReducers \| None  
Reducer function(s) for aggregating scores in each sample. Defaults to previously used reducer(s).

`model` str \| [Model](../reference/inspect_ai.model.html.md#model) \| None  
Optional. Model used for re-scoring (replaces the primary model reconstructed from the log header).

`model_roles` [ModelRoles](../reference/inspect_ai.model.html.md#modelroles) \| None  
Optional. Named model roles used for re-scoring (merged over the model roles reconstructed from the log header).

`action` ScoreAction \| None  
Whether to append or overwrite this score

`display` [DisplayType](../reference/inspect_ai.util.html.md#displaytype) \| None  
Progress/status display

`copy` bool  
Whether to deepcopy the log before scoring.

## Tasks

### Task

Evaluation task.

Tasks are the basis for defining and running evaluations.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task.py#L81)

``` python
class Task
```

#### Methods

\_\_init\_\_  
Create a task.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task.py#L87)

``` python
def __init__(
    self,
    dataset: Dataset | Sequence[Sample] | SampleSource | None = ...,
    setup: Solver | list[Solver] | None = ...,
    solver: Solver | Agent | list[Solver] = ...,
    cleanup: Callable[[TaskState], Awaitable[None]] | None = ...,
    scorer: 'Scorers' | None = ...,
    metrics: list[Metric | dict[str, list[Metric]]] | dict[str, list[Metric]] | None = ...,
    model: str | Model | None = ...,
    config: GenerateConfig = ...,
    model_roles: ModelRoles | None = ...,
    sandbox: SandboxEnvironmentType | None = ...,
    checkpoint: CheckpointConfig | bool | None = ...,
    on_checkpoint: OnCheckpointCallback | None = ...,
    on_resume: OnResumeCallback | None = ...,
    approval: str | ApprovalPolicyConfig | list[ApprovalPolicy] | None = ...,
    review: str | ReviewPolicyConfig | list[ReviewPolicy] | None = ...,
    epochs: int | Epochs | None = ...,
    fail_on_error: bool | float | None = ...,
    continue_on_fail: bool | None = ...,
    score_on_error: bool | None = ...,
    message_limit: int | None = ...,
    token_limit: int | str | TokenLimit | None = ...,
    turn_limit: int | None = ...,
    time_limit: int | None = ...,
    working_limit: int | None = ...,
    cost_limit: float | None = ...,
    early_stopping: 'EarlyStopping' | None = ...,
    display_name: str | None = ...,
    name: str | None = ...,
    version: int | str = ...,
    metadata: dict[str, Any] | None = ...,
    tags: list[str] | None = ...,
    viewer: ViewerConfig | None = ...,
    headline_metric: HeadlineMetric | str | None = ...,
    *,
    plan: Plan | Solver | list[Solver] = ...,
    tool_environment: str | SandboxEnvironmentSpec | None = ...,
    epochs_reducer: ScoreReducers | None = ...,
    max_messages: int | None = ...,
) -> None
```

`dataset` [Dataset](../reference/inspect_ai.dataset.html.md#dataset) \| Sequence\[[Sample](../reference/inspect_ai.dataset.html.md#sample)\] \| [SampleSource](../reference/inspect_ai.html.md#samplesource) \| None  
Dataset to evaluate, or a [SampleSource](../reference/inspect_ai.html.md#samplesource) that generates samples dynamically while the task runs.

`setup` [Solver](../reference/inspect_ai.solver.html.md#solver) \| list\[[Solver](../reference/inspect_ai.solver.html.md#solver)\] \| None  
Setup step (always run even when the main `solver` is replaced).

`solver` [Solver](../reference/inspect_ai.solver.html.md#solver) \| [Agent](../reference/inspect_ai.agent.html.md#agent) \| list\[[Solver](../reference/inspect_ai.solver.html.md#solver)\]  
Solver or list of solvers. Defaults to generate(), a normal call to the model.

`cleanup` Callable\[\[[TaskState](../reference/inspect_ai.solver.html.md#taskstate)\], Awaitable\[None\]\] \| None  
Optional cleanup function for task. Called after all solvers and scorers have run for each sample (including if an exception occurs during the run)

`scorer` 'Scorers' \| None  
Scorer used to evaluate model output.

`metrics` list\[[Metric](../reference/inspect_ai.scorer.html.md#metric) \| dict\[str, list\[[Metric](../reference/inspect_ai.scorer.html.md#metric)\]\]\] \| dict\[str, list\[[Metric](../reference/inspect_ai.scorer.html.md#metric)\]\] \| None  
Alternative metrics (overrides the metrics provided by the specified scorer).

`model` str \| [Model](../reference/inspect_ai.model.html.md#model) \| None  
Default model for task (Optional, defaults to eval model).

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig)  
Model generation config for default model (does not apply to model roles)

`model_roles` [ModelRoles](../reference/inspect_ai.model.html.md#modelroles) \| None  
Named roles for use in [get_model()](../reference/inspect_ai.model.html.md#get_model) (a role can also map to a list of models).

`sandbox` SandboxEnvironmentType \| None  
Sandbox environment type (or optionally a str or tuple with a shorthand spec)

`checkpoint` [CheckpointConfig](../reference/inspect_ai.util.html.md#checkpointconfig) \| bool \| None  
Checkpoint configuration for this task. `True` (or a [CheckpointConfig](../reference/inspect_ai.util.html.md#checkpointconfig)) enables checkpointing with the default trigger (every 500k tokens) unless overridden; `None` (default) inherits from the eval/CLI level; `False` vetoes checkpointing for this task, overriding an eval-set/CLI `--checkpoint` enable. When enabled, an eval-level `checkpoint` overrides this task’s config, which overrides any sample-level `checkpoint`.

`on_checkpoint` OnCheckpointCallback \| None  
Callback invoked before each checkpoint snapshot is taken, so state it flushes to the sandbox/store is captured by that checkpoint. May fire many times (including the final checkpoint on clean completion); must be idempotent.

`on_resume` OnResumeCallback \| None  
Callback invoked after a sample is restored on resume, before the agent resumes. Receives the TaskState and the resume `attempt` (‘resume’ or ‘resume_for_scoring’). At call time `state.store`, the transcript, and the sandbox are restored, but `state.messages`/`state.output` are NOT yet restored (the agent restores those itself) — use `state.store`/sandbox, not `state.messages`. May return a `ResumeReport` (or a `str` shorthand, or `None`) surfaced to the agent via `checkpointer().restored`.

`approval` str \| ApprovalPolicyConfig \| list\[[ApprovalPolicy](../reference/inspect_ai.approval.html.md#approvalpolicy)\] \| None  
Tool use approval policies. Either a path to an approval policy config file, an ApprovalPolicyConfig, or a list of approval policies. Defaults to no approval policy.

`review` str \| ReviewPolicyConfig \| list\[[ReviewPolicy](../reference/inspect_ai.review.html.md#reviewpolicy)\] \| None  
Tool result review policies. Either a path to a review policy config file, a ReviewPolicyConfig, or a list of review policies. Defaults to no review policy.

`epochs` int \| [Epochs](../reference/inspect_ai.html.md#epochs) \| None  
Epochs to repeat samples for and optional score reducer function(s) used to combine sample scores (defaults to “mean”)

`fail_on_error` bool \| float \| None  
`True` to fail on first sample error (default); `False` to never fail on sample errors; Value between 0 and 1 to fail if a proportion of total samples fails. Value greater than 1 to fail eval if a count of samples fails.

`continue_on_fail` bool \| None  
`True` to continue running and only fail at the end if the `fail_on_error` condition is met. `False` to fail eval immediately when the `fail_on_error` condition is met (default).

`score_on_error` bool \| None  
`True` to score samples that error rather than failing the eval mid-run. Errors still count toward the `fail_on_error` threshold for marking the eval log as ‘error’. Only takes effect after retries (if any) are exhausted.

`message_limit` int \| None  
Limit on total messages used for each sample.

`token_limit` int \| str \| [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) \| None  
Limit on tokens used for each sample. An `int` (or a [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) with type “all”) limits total tokens; a [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) with a `type` limits by output tokens or an arithmetic formula over `input`/`output`. Also accepts strings like “500k”, “1m”, “output:1m”, or “(input\*0.1)+output:1m”.

`turn_limit` int \| None  
Limit on total turns (model generations) used for each sample.

`time_limit` int \| None  
Limit on clock time (in seconds) for samples.

`working_limit` int \| None  
Limit on working time (in seconds) for sample. Working time includes model generation, tool calls, etc. but does not include time spent waiting on retries or shared resources.

`cost_limit` float \| None  
Limit on total cost (in dollars) for each sample. Requires model cost data via set_model_cost() or –model-cost-config.

`early_stopping` 'EarlyStopping' \| None  
Early stopping callbacks.

`display_name` str \| None  
Task display name (e.g. for plotting). If not specified then defaults to the registered task name.

`name` str \| None  
Task name. If not specified is automatically determined based on the registered name of the task.

`version` int \| str  
Version of task (to distinguish evolutions of the task spec or breaking changes to it)

`metadata` dict\[str, Any\] \| None  
Additional metadata to associate with the task.

`tags` list\[str\] \| None  
Tags to associate with the task.

`viewer` [ViewerConfig](../reference/inspect_ai.viewer.html.md#viewerconfig) \| None  
Log viewer configuration for this task (how the log’s samples, scores and scanner results are displayed, and whether its content is trusted to render richly).

`headline_metric` [HeadlineMetric](../reference/inspect_ai.log.html.md#headlinemetric) \| str \| None  
Which score/metric best summarises this task (e.g. for a leaderboard or log listing). A `str` names the scorer, as `"<scorer>"` or `"<scorer>.<score>"` to address one value of a scorer returning a dict of scores. Pass a [HeadlineMetric](../reference/inspect_ai.log.html.md#headlinemetric) to also name the `metric` or `reducer`. Unset fields resolve by convention, so `HeadlineMetric(metric="accuracy")` takes that metric from the first score reporting it; the default is the first metric of the first score.

`plan` Plan \| [Solver](../reference/inspect_ai.solver.html.md#solver) \| list\[[Solver](../reference/inspect_ai.solver.html.md#solver)\]  

`tool_environment` str \| SandboxEnvironmentSpec \| None  

`epochs_reducer` ScoreReducers \| None  

`max_messages` int \| None  

### task_with

Task adapted with alternate values for one or more options.

This function modifies the passed task in place and returns it. If you want to create multiple variations of a single task using [task_with()](../reference/inspect_ai.html.md#task_with) you should create the underlying task multiple times.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task.py#L306)

``` python
def task_with(
    task: Task,
    *,
    dataset: Dataset | Sequence[Sample] | SampleSource | None | NotGiven = NOT_GIVEN,
    setup: Solver | list[Solver] | None | NotGiven = NOT_GIVEN,
    solver: Solver | Agent | list[Solver] | NotGiven = NOT_GIVEN,
    cleanup: Callable[[TaskState], Awaitable[None]] | None | NotGiven = NOT_GIVEN,
    scorer: "Scorers" | None | NotGiven = NOT_GIVEN,
    metrics: list[Metric | dict[str, list[Metric]]]
    | dict[str, list[Metric]]
    | None
    | NotGiven = NOT_GIVEN,
    model: str | Model | NotGiven = NOT_GIVEN,
    config: GenerateConfig | NotGiven = NOT_GIVEN,
    model_roles: ModelRoles | NotGiven = NOT_GIVEN,
    sandbox: SandboxEnvironmentType | None | NotGiven = NOT_GIVEN,
    checkpoint: CheckpointConfig | bool | None | NotGiven = NOT_GIVEN,
    on_checkpoint: OnCheckpointCallback | None | NotGiven = NOT_GIVEN,
    on_resume: OnResumeCallback | None | NotGiven = NOT_GIVEN,
    approval: str
    | ApprovalPolicyConfig
    | list[ApprovalPolicy]
    | None
    | NotGiven = NOT_GIVEN,
    review: str | ReviewPolicyConfig | list[ReviewPolicy] | None | NotGiven = NOT_GIVEN,
    epochs: int | Epochs | None | NotGiven = NOT_GIVEN,
    fail_on_error: bool | float | None | NotGiven = NOT_GIVEN,
    continue_on_fail: bool | None | NotGiven = NOT_GIVEN,
    score_on_error: bool | None | NotGiven = NOT_GIVEN,
    message_limit: int | None | NotGiven = NOT_GIVEN,
    token_limit: int | str | TokenLimit | None | NotGiven = NOT_GIVEN,
    turn_limit: int | None | NotGiven = NOT_GIVEN,
    time_limit: int | None | NotGiven = NOT_GIVEN,
    working_limit: int | None | NotGiven = NOT_GIVEN,
    cost_limit: float | None | NotGiven = NOT_GIVEN,
    early_stopping: EarlyStopping | None | NotGiven = NOT_GIVEN,
    name: str | None | NotGiven = NOT_GIVEN,
    version: int | str | NotGiven = NOT_GIVEN,
    metadata: dict[str, Any] | None | NotGiven = NOT_GIVEN,
    tags: list[str] | None | NotGiven = NOT_GIVEN,
    viewer: ViewerConfig | None | NotGiven = NOT_GIVEN,
    headline_metric: HeadlineMetric | str | None | NotGiven = NOT_GIVEN,
) -> Task
```

`task` [Task](../reference/inspect_ai.html.md#task)  
Task to adapt

`dataset` [Dataset](../reference/inspect_ai.dataset.html.md#dataset) \| Sequence\[[Sample](../reference/inspect_ai.dataset.html.md#sample)\] \| [SampleSource](../reference/inspect_ai.html.md#samplesource) \| None \| NotGiven  
Dataset to evaluate, or a [SampleSource](../reference/inspect_ai.html.md#samplesource) that generates samples dynamically while the task runs.

`setup` [Solver](../reference/inspect_ai.solver.html.md#solver) \| list\[[Solver](../reference/inspect_ai.solver.html.md#solver)\] \| None \| NotGiven  
Setup step (always run even when the main `solver` is replaced).

`solver` [Solver](../reference/inspect_ai.solver.html.md#solver) \| [Agent](../reference/inspect_ai.agent.html.md#agent) \| list\[[Solver](../reference/inspect_ai.solver.html.md#solver)\] \| NotGiven  
Solver or list of solvers. Defaults to generate(), a normal call to the model.

`cleanup` Callable\[\[[TaskState](../reference/inspect_ai.solver.html.md#taskstate)\], Awaitable\[None\]\] \| None \| NotGiven  
Optional cleanup function for task. Called after all solvers and scorers have run for each sample (including if an exception occurs during the run)

`scorer` 'Scorers' \| None \| NotGiven  
Scorer used to evaluate model output.

`metrics` list\[[Metric](../reference/inspect_ai.scorer.html.md#metric) \| dict\[str, list\[[Metric](../reference/inspect_ai.scorer.html.md#metric)\]\]\] \| dict\[str, list\[[Metric](../reference/inspect_ai.scorer.html.md#metric)\]\] \| None \| NotGiven  
Alternative metrics (overrides the metrics provided by the specified scorer).

`model` str \| [Model](../reference/inspect_ai.model.html.md#model) \| NotGiven  
Default model for task (Optional, defaults to eval model).

`config` [GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig) \| NotGiven  
Model generation config for default model (does not apply to model roles)

`model_roles` [ModelRoles](../reference/inspect_ai.model.html.md#modelroles) \| NotGiven  
Named roles for use in [get_model()](../reference/inspect_ai.model.html.md#get_model) (a role can also map to a list of models).

`sandbox` SandboxEnvironmentType \| None \| NotGiven  
Sandbox environment type (or optionally a str or tuple with a shorthand spec)

`checkpoint` [CheckpointConfig](../reference/inspect_ai.util.html.md#checkpointconfig) \| bool \| None \| NotGiven  
Checkpoint configuration for this task. `True` (or a [CheckpointConfig](../reference/inspect_ai.util.html.md#checkpointconfig)) enables checkpointing with the default trigger (every 500k tokens) unless overridden; `None` (default) inherits from the eval/CLI level; `False` vetoes checkpointing for this task, overriding an eval-set/CLI `--checkpoint` enable. When enabled, an eval-level `checkpoint` overrides this task’s config, which overrides any sample-level `checkpoint`.

`on_checkpoint` OnCheckpointCallback \| None \| NotGiven  
Callback invoked before each checkpoint snapshot is taken, so state it flushes to the sandbox/store is captured by that checkpoint. May fire many times (including the final checkpoint on clean completion); must be idempotent.

`on_resume` OnResumeCallback \| None \| NotGiven  
Callback invoked after a sample is restored on resume, before the agent resumes. Receives the TaskState and the resume `attempt` (‘resume’ or ‘resume_for_scoring’). At call time `state.store`, the transcript, and the sandbox are restored, but `state.messages`/`state.output` are NOT yet restored (the agent restores those itself) — use `state.store`/sandbox, not `state.messages`. May return a `ResumeReport` (or a `str` shorthand, or `None`) surfaced to the agent via `checkpointer().restored`.

`approval` str \| ApprovalPolicyConfig \| list\[[ApprovalPolicy](../reference/inspect_ai.approval.html.md#approvalpolicy)\] \| None \| NotGiven  
Tool use approval policies. Either a path to an approval policy config file, an ApprovalPolicyConfig, or a list of approval policies. Defaults to no approval policy.

`review` str \| ReviewPolicyConfig \| list\[[ReviewPolicy](../reference/inspect_ai.review.html.md#reviewpolicy)\] \| None \| NotGiven  
Tool result review policies. Either a path to a review policy config file, a ReviewPolicyConfig, or a list of review policies. Defaults to no review policy.

`epochs` int \| [Epochs](../reference/inspect_ai.html.md#epochs) \| None \| NotGiven  
Epochs to repeat samples for and optional score reducer function(s) used to combine sample scores (defaults to “mean”)

`fail_on_error` bool \| float \| None \| NotGiven  
`True` to fail on first sample error (default); `False` to never fail on sample errors; Value between 0 and 1 to fail if a proportion of total samples fails. Value greater than 1 to fail eval if a count of samples fails.

`continue_on_fail` bool \| None \| NotGiven  
`True` to continue running and only fail at the end if the `fail_on_error` condition is met. `False` to fail eval immediately when the `fail_on_error` condition is met (default).

`score_on_error` bool \| None \| NotGiven  
`True` to score samples that error rather than failing the eval mid-run. Errors still count toward the `fail_on_error` threshold for marking the eval log as ‘error’. Only takes effect after retries (if any) are exhausted.

`message_limit` int \| None \| NotGiven  
Limit on total messages used for each sample.

`token_limit` int \| str \| [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) \| None \| NotGiven  
Limit on tokens used for each sample. An `int` (or a [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) with type “all”) limits total tokens; a [TokenLimit](../reference/inspect_ai.util.html.md#tokenlimit) with a `type` limits by output tokens or an arithmetic formula over `input`/`output`. Also accepts strings like “500k”, “1m”, “output:1m”, or “(input\*0.1)+output:1m”.

`turn_limit` int \| None \| NotGiven  
Limit on total turns (model generations) used for each sample.

`time_limit` int \| None \| NotGiven  
Limit on clock time (in seconds) for samples.

`working_limit` int \| None \| NotGiven  
Limit on working time (in seconds) for sample. Working time includes model generation, tool calls, etc. but does not include time spent waiting on retries or shared resources.

`cost_limit` float \| None \| NotGiven  
Limit on total cost (in dollars) for each sample. Requires model cost data via set_model_cost() or –model-cost-config.

`early_stopping` [EarlyStopping](../reference/inspect_ai.util.html.md#earlystopping) \| None \| NotGiven  
Early stopping callbacks.

`name` str \| None \| NotGiven  
Task name. If not specified is automatically determined based on the name of the task directory (or “task”) if its anonymous task (e.g. created in a notebook and passed to eval() directly)

`version` int \| str \| NotGiven  
Version of task (to distinguish evolutions of the task spec or breaking changes to it)

`metadata` dict\[str, Any\] \| None \| NotGiven  
Additional metadata to associate with the task.

`tags` list\[str\] \| None \| NotGiven  
Tags to associate with the task.

`viewer` [ViewerConfig](../reference/inspect_ai.viewer.html.md#viewerconfig) \| None \| NotGiven  
Log viewer configuration for this task (how the log’s samples, scores and scanner results are displayed, and whether its content is trusted to render richly).

`headline_metric` [HeadlineMetric](../reference/inspect_ai.log.html.md#headlinemetric) \| str \| None \| NotGiven  
Which score/metric best summarises this task (e.g. for a leaderboard or log listing).

### task_identifier

Unique identifier for a task within an eval set.

Identifiers have the form `{task_file}@{task_name}#{args_hash}/{model}/{additional_hash}` (the `{task_file}@` prefix is omitted for tasks without a source file). The additional hash covers the remaining fields that distinguish tasks within an eval set (solver plan, generate config, model args, model roles, task version, and execution limits), excluding runtime/transport options that don’t affect model output (e.g. `max_retries`, `max_connections`).

The same identifier is computed from a `ResolvedTask` (before running) and from the [EvalLog](../reference/inspect_ai.log.html.md#evallog) that running it produces — [eval_set()](../reference/inspect_ai.html.md#eval_set) uses this to pair tasks with their existing log files across retries, and external runners can correlate enumerated tasks with logs the same way. The computation is versioned by `TASK_IDENTIFIER_VERSION`: persisted identifiers must be recomputed when the version changes.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/evalset.py#L2058)

``` python
def task_identifier(
    task: ResolvedTask | EvalLog,
    eval_set_args: EvalSetArgsInTaskIdentifier | None,
) -> str
```

`task` ResolvedTask \| [EvalLog](../reference/inspect_ai.log.html.md#evallog)  
Task to identify (a `ResolvedTask` prior to running or an [EvalLog](../reference/inspect_ai.log.html.md#evallog) from a previous run).

`eval_set_args` EvalSetArgsInTaskIdentifier \| None  
Eval-set level arguments that participate in task identity. Required when `task` is a `ResolvedTask`; pass `None` for an [EvalLog](../reference/inspect_ai.log.html.md#evallog) (the log already carries the resolved values).

### Epochs

Task epochs.

Number of epochs to repeat samples over and optionally one or more reducers used to combine scores from samples across epochs. If not specified the “mean” score reducer is used.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/epochs.py#L4)

``` python
class Epochs
```

#### Attributes

`reducer` list\[[ScoreReducer](../reference/inspect_ai.scorer.html.md#scorereducer)\] \| None  
One or more reducers used to combine scores from samples across epochs (defaults to “mean”)

#### Methods

\_\_init\_\_  
Task epochs.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/epochs.py#L12)

``` python
def __init__(self, epochs: int, reducer: ScoreReducers | None = None) -> None
```

`epochs` int  
Number of epochs

`reducer` ScoreReducers \| None  
One or more reducers used to combine scores from samples across epochs (defaults to “mean”)

### TaskInfo

Task information (file, name, and attributes).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task.py#L509)

``` python
class TaskInfo(BaseModel)
```

#### Attributes

`file` str  
File path where task was loaded from.

`name` str  
Task name (defaults to function name)

`attribs` dict\[str, Any\]  
Task attributes (arguments passed to `@task`)

### Tasks

One or more tasks.

Tasks to be evaluated. Many forms of task specification are supported including directory names, task functions, task classes, and task instances (a single task or list of tasks can be specified). None is a request to read a task out of the current working directory.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/tasks.py#L7)

``` python
Tasks: TypeAlias = (
    str
    | PreviousTask
    | ResolvedTask
    | TaskInfo
    | Task
    | Callable[..., Task]
    | type[Task]
    | TaskSource
    | Callable[..., TaskSource]
    | list[str]
    | list[PreviousTask]
    | list[ResolvedTask]
    | list[PreviousTask | ResolvedTask]
    | list[TaskInfo]
    | list[Task]
    | list[Callable[..., Task]]
    | list[type[Task]]
    | None
)
```

### TaskSource

Drives a running eval from code: a seed plus result-driven follow-ups.

Subclass and override the methods you need. The default implementations are no-ops / empty, so a bare [TaskSource](../reference/inspect_ai.html.md#tasksource) runs nothing — override at least `initial_tasks` and `next_tasks`.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task_source.py#L39)

``` python
class TaskSource
```

#### Methods

initial_tasks  
Tasks to run first (the seed).

Called once, synchronously, before the run starts — so it must return immediately (no awaiting / blocking). The returned tasks drive the run’s up-front setup (concurrency, validation) and are the first batch.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task_source.py#L47)

``` python
def initial_tasks(self) -> list["Task"]
```

next_tasks  
The next batch of tasks to run, or `None` when the run is complete.

Called after each batch finishes (after that batch’s `sample_complete` / `task_complete` notifications). May `await` — for more results or external input — and may block indefinitely; return `None` to end the run.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task_source.py#L56)

``` python
async def next_tasks(self) -> list["Task"] | None
```

sample_complete  
A sample finished — observe it and optionally return follow-up tasks.

`sample` is the completed sample and `task` is the task it ran under (the sample alone doesn’t identify its task). Return a list of tasks to add to the run (equivalent to calling `enqueue_task` with them): they run after the current batch, before the next `next_tasks()`. Return `None` (the default) to add nothing.

Fires for every sample the task logs, including one cancelled individually by an operator (its `error` is then the cancellation, with no scores), but not for samples cancelled by the task itself unwinding (a task-level cancel or ^C) – those reach the source only via the log passed to `task_complete`. A cancelled sample’s `error.message` is the cancellation exception’s repr (it starts with `CancelledError(` or `Cancelled(`), which is how to tell it from a genuine error. A sample cancelled before it produced anything to log is reported via :meth:`sample_abandoned` instead.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task_source.py#L66)

``` python
async def sample_complete(
    self, sample: "EvalSample", task: "Task"
) -> list["Task"] | None
```

`sample` 'EvalSample'  

`task` 'Task'  

sample_abandoned  
A sample was cancelled without ever being logged.

`sample` is a copy of the sample as the task’s dataset held it, `epoch` the epoch that was abandoned and `task` the task it was queued under. Nothing about the run reached the log (no [EvalSample](../reference/inspect_ai.log.html.md#evalsample) exists, and the sample is absent from the log passed to `task_complete`), so unlike :meth:`sample_complete` there is no result to deliver. Return a list of tasks to add to the run (equivalent to calling `enqueue_task` with them) or `None` (the default) to add nothing.

Fires when an operator cancels a sample still waiting in the queue (`inspect ctl sample cancel --action cancel`), when a cancel lands on an errored sample in the window before its `retry_on_error` re-run, or when a graceful task cancel (`--action drain`, `score`, `error`) abandons a queued sample — in each case the task itself keeps running, so a source waiting on the sample must hear that it will never complete. Like :meth:`sample_complete` it does not fire for samples cancelled by the task itself unwinding (a task-level cancel or ^C), nor for a withdrawn requeue (the prior terminal outcome, already reported, stands). With `epochs > 1` it fires once per abandoned epoch, as `sample_complete` fires once per completed one.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task_source.py#L89)

``` python
async def sample_abandoned(
    self, sample: "Sample", epoch: int, task: "Task"
) -> list["Task"] | None
```

`sample` 'Sample'  

`epoch` int  

`task` 'Task'  

task_complete  
A task finished — observe its log and optionally return follow-up tasks.

Return a list of tasks to add to the run (like `enqueue_task`): they run after the current batch. Return `None` (the default) to add nothing.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task_source.py#L118)

``` python
async def task_complete(self, log: "EvalLog") -> list["Task"] | None
```

`log` 'EvalLog'  

from_tasks  
Create a :class:[TaskSource](../reference/inspect_ai.html.md#tasksource) from a seed plus optional callbacks.

A convenience for when subclassing is more than you need: provide the initial tasks directly and, optionally, callbacks that react to results. The `sample_complete` / `task_complete` callbacks may **return** a list of follow-up tasks to add to the run (see those methods); `next_tasks` is the blocking / explicit-pull alternative. Callbacks typically close over shared state (e.g. accumulated scores) to decide what to run next.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/task_source.py#L126)

``` python
@classmethod
def from_tasks(
    cls,
    initial_tasks: list["Task"],
    *,
    next_tasks: Callable[[], Awaitable[list["Task"] | None]] | None = None,
    sample_complete: Callable[
        ["EvalSample", "Task"], Awaitable[list["Task"] | None]
    ]
    | None = None,
    task_complete: Callable[["EvalLog"], Awaitable[list["Task"] | None]]
    | None = None,
    sample_abandoned: Callable[
        ["Sample", int, "Task"], Awaitable[list["Task"] | None]
    ]
    | None = None,
) -> "TaskSource"
```

`initial_tasks` list\['Task'\]  
The seed tasks to run first (see :meth:`initial_tasks`). Required, and resolved up front.

`next_tasks` Callable\[\[\], Awaitable\[list\['Task'\] \| None\]\] \| None  
Optional async callback returning the next batch, or `None` to end the run (see :meth:`next_tasks`). If omitted (and no callback returns tasks), the run stops after the seed — equivalent to passing `initial_tasks` directly to [eval()](../reference/inspect_ai.html.md#eval).

`sample_complete` Callable\[\['EvalSample', 'Task'\], Awaitable\[list\['Task'\] \| None\]\] \| None  
Optional async callback invoked as each sample finishes; may return follow-up tasks to add to the run.

`task_complete` Callable\[\['EvalLog'\], Awaitable\[list\['Task'\] \| None\]\] \| None  
Optional async callback invoked as each task finishes; may return follow-up tasks to add to the run.

`sample_abandoned` Callable\[\['Sample', int, 'Task'\], Awaitable\[list\['Task'\] \| None\]\] \| None  
Optional async callback invoked when a sample is cancelled without being logged (see :meth:`sample_abandoned`); may return follow-up tasks to add to the run.

### SampleSource

Drives a running task from code: a seed plus result-driven follow-ups.

Subclass and override the methods you need. The default implementations are no-ops / empty, so a bare [SampleSource](../reference/inspect_ai.html.md#samplesource) runs nothing — override at least `initial_samples` and `next_samples`.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/sample_source.py#L41)

``` python
class SampleSource
```

#### Methods

initial_samples  
Samples to run first (the seed).

Called once, synchronously, when the [Task](../reference/inspect_ai.html.md#task) is created — so it must return immediately (no awaiting / blocking). The returned samples drive the task’s up-front setup (validation, sandbox startup) and are the first batch. May be empty, in which case the task starts by calling `next_samples()`. The seed isn’t required for sandboxes: a sandbox config first seen in a later-added sample gets the same startup (image build/pull, validation, registered cleanup) before that sample runs.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/sample_source.py#L49)

``` python
def initial_samples(self) -> list["Sample"]
```

next_samples  
More samples to run, or `None` when the task is complete.

Called whenever no samples remain in flight or buffered (after those samples’ `sample_complete` notifications). May `await` — for more results or external input — and may block indefinitely; return `None` to end the task. (If samples were enqueued while a `None` return was in progress they still run, and this method may then be called again.)

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/sample_source.py#L63)

``` python
async def next_samples(self) -> list["Sample"] | None
```

sample_complete  
A sample finished — observe it and optionally return follow-up samples.

Return a list of samples to add to the task (equivalent to calling `enqueue_sample` with them): they start as soon as there is free capacity. Return `None` (the default) to add nothing.

Fires for every sample the task logs, including one cancelled individually by an operator (its `error` is then the cancellation, with no scores), but not for samples cancelled by the task itself unwinding (a task-level cancel or ^C), where any follow-ups could never run. A cancelled sample’s `error.message` is the cancellation exception’s repr (it starts with `CancelledError(` or `Cancelled(`), which is how to tell it from a genuine error. A sample cancelled before it produced anything to log is reported via :meth:`sample_abandoned` instead.

On a task retry this is also called for samples reused from the prior attempt, so a completion-driven source regenerates its follow-ups (returned samples whose ids match the prior attempt are themselves reused rather than re-run).

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/sample_source.py#L74)

``` python
async def sample_complete(self, sample: "EvalSample") -> list["Sample"] | None
```

`sample` 'EvalSample'  

sample_abandoned  
A sample was cancelled without ever being logged.

`sample` is a copy of the sample as it was added to the task and `epoch` the epoch that was abandoned. Nothing about the run reached the log (no [EvalSample](../reference/inspect_ai.log.html.md#evalsample) exists), so unlike :meth:`sample_complete` there is no result to deliver. Return a list of samples to add to the task (equivalent to calling `enqueue_sample` with them) or `None` (the default) to add nothing.

Fires when an operator cancels a sample still waiting in the queue (`inspect ctl sample cancel --action cancel`), when a cancel lands on an errored sample in the window before its `retry_on_error` re-run, or when a graceful task cancel (`--action drain`, `score`, `error`) abandons a queued sample — in each case the task itself keeps running, so a source waiting on the sample must hear that it will never complete. Like :meth:`sample_complete` it does not fire for samples cancelled by the task itself unwinding (a task-level cancel or ^C), nor for a withdrawn requeue (the prior terminal outcome, already reported, stands). With `epochs > 1` it fires once per abandoned epoch, as `sample_complete` fires once per completed one.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/sample_source.py#L98)

``` python
async def sample_abandoned(
    self, sample: "Sample", epoch: int
) -> list["Sample"] | None
```

`sample` 'Sample'  

`epoch` int  

from_samples  
Create a :class:[SampleSource](../reference/inspect_ai.html.md#samplesource) from a seed plus optional callbacks.

A convenience for when subclassing is more than you need: provide the initial samples directly and, optionally, callbacks that react to results. The `sample_complete` callback may **return** a list of follow-up samples to add to the task (see that method); `next_samples` is the blocking / explicit-pull alternative. Callbacks typically close over shared state (e.g. accumulated scores) to decide what to run next.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/sample_source.py#L126)

``` python
@classmethod
def from_samples(
    cls,
    initial_samples: list["Sample"],
    *,
    next_samples: Callable[[], Awaitable[list["Sample"] | None]] | None = None,
    sample_complete: Callable[["EvalSample"], Awaitable[list["Sample"] | None]]
    | None = None,
    sample_abandoned: Callable[["Sample", int], Awaitable[list["Sample"] | None]]
    | None = None,
) -> "SampleSource"
```

`initial_samples` list\['Sample'\]  
The seed samples to run first (see :meth:`initial_samples`).

`next_samples` Callable\[\[\], Awaitable\[list\['Sample'\] \| None\]\] \| None  
Optional async callback returning more samples, or `None` to end the task (see :meth:`next_samples`). If omitted (and no callback returns samples), the task stops after the seed — equivalent to passing `initial_samples` directly as the dataset.

`sample_complete` Callable\[\['EvalSample'\], Awaitable\[list\['Sample'\] \| None\]\] \| None  
Optional async callback invoked as each sample finishes; may return follow-up samples to add to the task.

`sample_abandoned` Callable\[\['Sample', int\], Awaitable\[list\['Sample'\] \| None\]\] \| None  
Optional async callback invoked when a sample is cancelled without being logged (see :meth:`sample_abandoned`); may return follow-up samples to add to the task.

### enqueue_task

Add one or more tasks to the running eval.

The tasks run in this process under the current run’s `run_id` (a fresh `eval_id`/`task_id` each, their own log files), resolved against the run’s models and config.

When the run is driven by a :class:`~inspect_ai.TaskSource`, added tasks are *live*: they start as soon as there is free capacity. Otherwise they run as a follow-up batch, after the in-flight batch of tasks completes.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/enqueue.py#L91)

``` python
def enqueue_task(tasks: "Tasks", *, run_id: str | None = None) -> None
```

`tasks` 'Tasks'  
A [Task](../reference/inspect_ai.html.md#task) (or list of tasks) to add to the running eval.

`run_id` str \| None  
Optionally, the `run_id` the caller believes is running; if given it must match the active run, else the call is rejected.

### enqueue_sample

Add one or more samples to the running task.

The samples run in the current task as soon as there is free capacity (bounded by `max_samples`), each for the task’s configured number of epochs. Samples without an `id` are assigned one automatically.

Only available inside a task driven by a :class:[SampleSource](../reference/inspect_ai.html.md#samplesource) (i.e. a [Task](../reference/inspect_ai.html.md#task) whose `dataset` is a [SampleSource](../reference/inspect_ai.html.md#samplesource)) — a plain task’s sample set is fixed, so there is no loop to run additions. Callable from any code running within such a task — a solver, a scorer, a tool — but it must be called from the task’s event loop (where those all run), not from a worker thread.

When the eval was run with `--limit`, samples beyond the limit are ignored (with a warning); with `--sample-id`, only samples matching the filter run.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/sample_source.py#L259)

``` python
def enqueue_sample(samples: "Sample | list[Sample]") -> None
```

`samples` 'Sample \| list\[[Sample](../reference/inspect_ai.dataset.html.md#sample)\]'  
A [Sample](../reference/inspect_ai.dataset.html.md#sample) (or list of samples) to add to the running task.

## Scanning

### Scanners

Argument shape accepted by `eval_set(scanner=...)`.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/scan.py#L178)

``` python
    Scanners: TypeAlias = (
        Sequence[Scanner[Any] | tuple[str, Scanner[Any]]]
        | dict[str, Scanner[Any]]
        | ScannerConfig
    )
```

### ScannerConfig

Configure scanners attached to an `eval_set` run.

A subset of scout’s `ScanJob` / `ScanJobConfig` schema, narrowed to the fields that make sense when `eval_set` is generating the transcripts.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/scan.py#L53)

``` python
class ScannerConfig(BaseModel)
```

#### Attributes

`scanners` Any  
Scanners to run.

`Sequence[Scanner | tuple[str, Scanner]]` for direct construction, `dict[str, Scanner]` for named scanners, or scout `ScannerSpec` references when loading from YAML/JSON config.

`name` str \| None  
Override the scan name written to `_scan.json` (defaults to “eval_set”).

`scans` str \| None  
Override scan output location. Defaults to `<log_dir>/scans/`.

`tags` list\[str\] \| None  
Tags written to the scan spec.

`metadata` dict\[str, Any\] \| None  
Metadata written to the scan spec.

`filter` str \| list\[str\]  
SQL WHERE clause(s) applied per-sample to skip transcripts that don’t match (e.g. `"error = ''"` to scan only successful samples). Mirrors scout’s `Transcripts.where(...)` semantics.

`model` Any  
Model used by scanners’ [get_model()](../reference/inspect_ai.model.html.md#get_model). Overrides the eval’s active model just for the scanner call. `str | Model | None`.

`model_base_url` str \| None  
Base URL for the scanner-side model API.

`model_args` dict\[str, Any\] \| None  
Model creation args forwarded to scout.

`generate_config` Any  
[GenerateConfig](../reference/inspect_ai.model.html.md#generateconfig) for scanner model calls.

`model_roles` dict\[str, Any\] \| None  
Named roles available to scanners via `get_model(role=...)`.

#### Methods

from_file  
Load a [ScannerConfig](../reference/inspect_ai.html.md#scannerconfig) from a YAML or JSON config file.

Scanner entries in the file are written as `ScannerSpec` references (a registry `name` plus optional `params` and `file`). They are resolved to live `Scanner` objects via scout’s registry, loading any referenced `file` modules. `model_args` may also be a path to a separate YAML/JSON file, which is read and inlined.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/task/scan.py#L115)

``` python
@classmethod
def from_file(cls, path: str) -> "ScannerConfig"
```

`path` str  
Path or URL (e.g. `s3://...`) to a YAML or JSON file.

## View

### view

Run the Inspect View server.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_view/view.py#L24)

``` python
def view(
    log_dir: str | None = None,
    recursive: bool = True,
    host: str = DEFAULT_SERVER_HOST,
    port: int = DEFAULT_VIEW_PORT,
    authorization: str | None = None,
    log_level: str | None = None,
    fs_options: dict[str, Any] = {},
    trusted_origins: tuple[str, ...] = (),
    trusted_hosts: tuple[str, ...] = (),
    unsafe_allow_unauthenticated: bool = False,
    show_shards: bool = False,
    trust_content: bool | None = None,
) -> None
```

`log_dir` str \| None  
Directory to view logs from.

`recursive` bool  
Recursively list files in `log_dir`.

`host` str  
Tcp/ip host (defaults to “127.0.0.1”).

`port` int  
Tcp/ip port (defaults to 7575).

`authorization` str \| None  
Validate requests by checking for this authorization header.

`log_level` str \| None  
Level for logging to the console: “debug”, “http”, “sandbox”, “info”, “warning”, “error”, “critical”, or “notset” (defaults to “warning”)

`fs_options` dict\[str, Any\]  
Additional arguments to pass through to the filesystem provider (e.g. `S3FileSystem`). Use `{"anon": True }` if you are accessing a public S3 bucket with no credentials.

`trusted_origins` tuple\[str, ...\]  
Exact browser origins allowed to use the viewer.

`trusted_hosts` tuple\[str, ...\]  
Additional exact HTTP authorities allowed for non-browser clients.

`unsafe_allow_unauthenticated` bool  
Allow a non-loopback bind without request authorization.

`show_shards` bool  
List shard logs (under `<name>.shards/`) that their merged log already covers. By default these are hidden.

`trust_content` bool \| None  
`False` shows the content of every log as plain text, whatever the log’s own `ViewerConfig(trust_content=...)`. `None` (the default) or `True` defers to each log; it never shows an untrusted log richly.

## Decorators

### task

Decorator for registering tasks.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/registry.py#L124)

``` python
def task(*args: Any, name: str | None = None, **attribs: Any) -> Any
```

`*args` Any  
Function returning [Task](../reference/inspect_ai.html.md#task) targeted by plain task decorator without attributes (e.g. `@task`)

`name` str \| None  
Optional name for task. If the decorator has no name argument then the name of the function will be used to automatically assign a name.

`**attribs` Any  
(dict\[str,Any\]): Additional task attributes.

### task_source

Decorator for registering task sources.

Mirrors `@task`: registers a function that returns a [TaskSource](../reference/inspect_ai.html.md#tasksource) so it can be referenced and loaded by name (e.g. `eval("file.py@my_source")` or `inspect eval file.py@my_source -T arg=value`) and parameterized.

[Source](https://github.com/UKGovernmentBEIS/inspect_ai/blob/93f7182cf2ce9be22724b05e499cd1358d7ed41d/src/inspect_ai/_eval/registry.py#L288)

``` python
def task_source(*args: Any, name: str | None = None, **attribs: Any) -> Any
```

`*args` Any  
Function returning [TaskSource](../reference/inspect_ai.html.md#tasksource) targeted by a plain decorator without attributes (e.g. `@task_source`).

`name` str \| None  
Optional name for the source (defaults to the function name).

`**attribs` Any  
Additional task source attributes.
