Caching

Overview

Caching enables you to cache model output to reduce the number of API calls made, saving both time and expense. Caching is also often useful during development—for example, when you are iterating on a scorer you may want the model outputs served from a cache to both save time as well as for increased determinism.

There are two types of caching available: Inspect local caching and provider level caching. We’ll first describe local caching (which works for all models) then cover provider caching which currently works only for Anthropic models.

Caching Basics

Use the cache option of GenerateConfig to activate the use of the cache. The keys for caching (what determines if a request can be fulfilled from the cache) are as follows:

  • Model name and base URL (e.g. openai/gpt-5)
  • Model prompt (i.e. message history)
  • Epoch number (for ensuring distinct generations per epoch)
  • Generate configuration (e.g. temperature, top_p, etc.)
  • Active tools and tool_choice

If all of these inputs are identical, then the model response will be served from the cache. By default, model responses are cached for 1 week (see Cache Policy below for details on customising this).

Here are some example uses of --cache from the CLI:

inspect eval arc.py --cache     # 7 day cache (default)
inspect eval arc.py --cache 1D  # 1 day cache
inspect eval arc.py --cache 4W  # 4 week cache

Or alternatively from Python when calling eval():

eval("arc.py", cache=True)

You can also use caching with lower-level generate() calls (e.g. a model instance you have obtained with get_model(). For example:

model = get_model("anthropic/claude-sonnet-4-20250514")
output = model.generate(
  input, config=GenerateConfig(cache = True)
)

Model Versions

The model name (e.g. openai/gpt-4-turbo) is used as part of the cache key. Note though that many model names are aliases to specific model versions. For example, gpt-4, gpt-4-turbo, may resolve to different versions over time as updates are released.

If you want to invalidate caches for updated model versions, it’s much better to use an explicitly versioned model name. For example:

$ inspect eval ctf.py --model openai/gpt-4-turbo-2024-04-09

If you do this, then when a new version of gpt-4-turbo is deployed a call to the model will occur rather than resolving from the cache.

Cache Policy

By default, if you specify cache = True then the cache will expire in 1 week. You can customise this by passing a CachePolicy rather than a boolean. For example:

cache = CachePolicy(expiry="3h")
cache = CachePolicy(expiry="4D")
cache = CachePolicy(expiry="2W")
cache = CachePolicy(expiry="3M")

You can use s, m, h, D, W , M, and Y as abbreviations for expiry values.

If you want the cache to never expire, specify None. For example:

cache = CachePolicy(expiry = None)

You can also define scopes for cache expiration (e.g. cache for a specific task or usage pattern). Use the scopes parameter to add named scopes to the cache key:

cache = CachePolicy(
    expiry="1M",
    scopes={"role": "attacker", "team": "red"})
)

As noted above, caching is by default done per epoch (i.e. each epoch has its own cache scope). You can disable the default behaviour by setting per_epoch=False. For example:

cache = CachePolicy(per_epoch=False)

Management

Use the inspect cache command the view the current contents of the cache, prune expired entries, or clear entries entirely. For example:

# list the current contents of the cache
$ inspect cache list

# clear the cache (globally or by model)
$ inspect cache clear
$ inspect cache clear --model openai/gpt-4-turbo-2024-04-09

# prune expired entries from the cache
$ inspect cache list --pruneable
$ inspect cache prune
$ inspect cache prune --model openai/gpt-4-turbo-2024-04-09

See inspect cache --help for further details on management commands.

Cache Directory

By default the model generation cache is stored in the system default location for user cache files (e.g. XDG_CACHE_HOME on Linux). You can override this and specify a different directory for cache files using the INSPECT_CACHE_DIR environment variable. For example:

$ export INSPECT_CACHE_DIR=/tmp/inspect-cache

Provider Caching

Model providers may also provide prompt caching features to optimise cost and performance for multi-turn conversations. The only provider that currently enables you to turn off prompt caching is Anthropic, and you can do this using cache-prompt generation config option. For example:

inspect eval ctf.py --cache-prompt=false # force caching off

Or with the eval() function:

eval("ctf.py", cache_prompt=False)

Cache Scope

Providers will typically provide various means of customising the scope of cache usage. The Inspect cache-prompt option will by default attempt to make maximum use of provider caches (in the Anthropic implementation system messages, tool definitions, and all messages up to the last user message are included in the cache).

The Anthropic implementation places cache breakpoints at the end of the system prompt, on the last tool definition, on the second-to-last message block, and (automatically) at the end of the prompt. The automatic breakpoint cache-writes whatever the final block is, which is wasted when that block changes on every call and is never read back (e.g. an LLM judge with a fixed rubric and a varying item).

ContentText(text=..., cache_breakpoint=True) places an explicit breakpoint on a specific block instead. When any block carries one — a message block or a system-prompt block — Inspect adds no automatic message breakpoints (no lookback marker, no end-of-prompt marker); the system-prompt and tool-definition breakpoints are kept alongside your marks only if the budget allows. Anthropic allows 4 breakpoints per request, so you can mark up to 4 blocks yourself; if your marks and the automatic system/tools breakpoints would together exceed 4, Inspect drops the automatic ones first to make room. Marking more than 4 blocks yourself raises a ValueError naming the count and the limit rather than silently falling back to normal automatic caching — resuming automatic caching would reintroduce exactly the varying-tail write the marks were meant to avoid.

from inspect_ai.model import ChatMessageUser, ContentText

judge_prompt = ChatMessageUser(
    content=[
        ContentText(text=rubric, cache_breakpoint=True),
        ContentText(text=item),
    ]
)

Explicit marks are supported on the leading system message and on ordinary user/assistant message blocks. A mark on a tool result, or on a mid-conversation system message, is not supported — Anthropic has no way to mark a boundary inside either without relocating or widening it, so those placements conservatively fall back rather than being honored. If any mark is in an unsupported placement, Inspect silently discards every mark and falls back to the normal automatic caching behavior for the whole request, rather than relocating, widening, or partially honoring the marked layout. Supplying more than 4 marks is different: rather than falling back, Inspect raises a ValueError, since silently resuming automatic caching would reintroduce the varying-tail write the marks exist to prevent.

OpenAI’s gpt-5.6 and later models support the same ContentText(cache_breakpoint=True) marker on user and system/developer message text blocks. Unmarked requests use OpenAI’s automatic (implicit) caching as they always have; marking a block switches the request to explicit mode for that call, so only the marked boundaries are cache points. OpenAI’s guide distinguishes a 4-write-per-request budget from a much larger lookup history (the first two and latest fifty explicit markers) and documents no cap on the number of markers a request may supply, so Inspect applies none — unlike Anthropic’s genuine 4-breakpoint request limit. Earlier OpenAI models silently keep the model’s normal implicit caching instead of sending the (rejected) explicit fields. When an otherwise-supported explicit request leaves its leading system/developer block unmarked, Inspect also retains a checkpoint at the end of that block (cumulatively covering any preceding tool definitions), so the stable prefix has a boundary to reuse when a later marked block changes; a caller’s own mark on that block takes priority and suppresses this automatic one.

Usage Reporting

When using provider caching, model token usage will be reported with 4 distinct values rather than the normal input and output. For example:

13,684 tokens [I: 22, CW: 1,711, CR: 11,442, O: 509]

Where the prefixes on reported token counts stand for:

I Input tokens
CW Input token cache writes
CR Input token cache reads
O Output tokens

Input token cache writes will typically cost more (in the case of Anthropic roughly 25% more) but cache reads substantially less (for Anthropic 90% less) so for the example above there would have been a substantial savings in cost and execution time. See the Anthropic Documentation for additional details.