FireworksPromptCachingMiddleware(
self,
*,
unsupported_model_behavior: Literal['ignore', 'warn', 'raise'] = 'warn'
)| Name | Type | Description |
|---|---|---|
unsupported_model_behavior | Literal['ignore', 'warn', 'raise'] | Default: 'warn'Behavior when the request model is not
|
| Name | Type |
|---|---|
| unsupported_model_behavior | Literal['ignore', 'warn', 'raise'] |
Set Fireworks prompt-cache session affinity from the active thread ID.
Fireworks prompt caching is enabled by default. This middleware improves
cache hit rate by pinning session affinity to a SHA-256 hash of
config.configurable.thread_id, so related requests route to the same
replica and reuse its warm cache. The hexadecimal hash keeps affinity values
safe for HTTP headers, including when thread IDs contain Unicode.
The middleware supplies a scoped default that ChatFireworks applies to
prompt_cache_key and extra_headers["x-session-affinity"] when invoking
the API. Explicit user or prompt_cache_key body fields (including
extra_body overrides), or x-session-affinity headers on the selected
model or request take precedence, including on fallback models. No affinity
is added when no thread ID is configured.
Generated affinity stays out of shared request settings, so it is never
forwarded to another provider. This works with either ordering of this
middleware and ModelFallbackMiddleware for a Fireworks primary model.