Constrains effort on reasoning for reasoning models.
For use with the Chat Completions API. Reasoning models only.
Currently supported values are 'minimal', 'low', 'medium', and
'high'. Reducing reasoning effort can result in faster responses and fewer
tokens used on reasoning in a response.
Changing this value part-way through a conversation changes a request-level parameter, which invalidates the cached prompt prefix.
Models that support it (currently GPT-6) can instead carry the new effort
in a configuration_update item attached to the message that should start
using it:
HumanMessage(
[
{"type": "configuration_update", "reasoning": {"effort": "high"}},
{"type": "text", "text": "Analyze the failure modes."},
]
)
The new effort applies from that message onward, until another update overrides it.