| Name | Type | Description |
|---|---|---|
persist_model_state | bool | Default: TrueWhether completed calls should write private resume metadata. Subagent instances disable this because they do not own the parent thread's resume state. |
openai_prompt_cache_key | bool | None | Default: NoneWhether to inject a per-thread OpenAI
|
cli_max_retries | int | None | Default: None |
strict_model_resolution | bool | Default: False |
environ | Mapping[str, str] | None | Default: None |
model_result | ModelResult | None | Default: None |
Swap the model or per-call settings from runtime.context.
Reads two optional keys from the runtime context dict:
'model' — a provider:model spec (e.g. "openai:gpt-5").
When present and different from the current model, the request is
re-routed to the new model.'model_params' — a dict of extra model settings (e.g.
{"temperature": 0}) that are shallow-merged into the
request's model_settings.This middleware is typically the outermost layer so it intercepts every
model call before provider-specific middleware (like
AnthropicPromptCachingMiddleware) runs.
Explicit --max-retries value to retain across
runtime model switches.
Whether invalid runtime model overrides should fail the call instead of falling back to the construction-time model.
Workspace environment retained for lazy model switches.
Construction-time workspace model metadata.