Problem (one or two sentences)
Claude Opus 5.5 declares a 128K output limit, but Zoo Code resolves it to 8,192 tokens by default and 16,384 when reasoning is enabled on the direct Anthropic provider. Because Opus 5.5 always uses adaptive thinking, a valid agent turn can consume the entire allowance as thinking and stop before producing text or a tool call.
Context (who is affected and when)
This affects users running claude-opus-5-5 through Zoo Code's Anthropic provider, including Anthropic-compatible base URLs, especially on complex agent tasks.
In the observed failure, the provider stream ended normally with stop_reason: max_tokens. All 16,384 output tokens were reported as thinking tokens, and the next event was message_stop. The transport delivered a valid terminal event; the model had simply spent the configured output allowance before producing visible output.
The model entry already declares maxTokens: 128_000, but it also declares both supportsReasoningBudget: true and supportsReasoningBinary: true. supportsReasoningBinary selects adaptive thinking, while supportsReasoningBudget routes output sizing through Zoo Code's legacy hybrid-reasoning defaults.
Reproduction steps
-
Use Zoo Code 3.84.0 or current main with the Anthropic provider and select claude-opus-5-5.
-
Leave modelMaxTokens unset.
-
Inspect the resolved model with reasoning unset or disabled:
new AnthropicHandler({ apiModelId: "claude-opus-5-5" }).getModel().maxTokens
It resolves to 8192, despite the model entry declaring 128000.
-
Set enableReasoningEffort: true and inspect the resolved model again. It resolves to 16384.
-
Create a message with reasoning enabled. The request uses adaptive thinking but sends max_tokens: 16384.
-
On a sufficiently complex agent turn, observe the response stop at max_tokens before emitting text or a tool call.
Expected result
Claude Opus 5.5 uses its declared 128K output limit on the direct Anthropic provider path. Enabling adaptive thinking does not replace that model limit with the legacy 16K hybrid-reasoning default.
Actual result
getModelMaxOutputTokens() interprets supportsReasoningBudget as a legacy hybrid-budget capability. It returns 16,384 when reasoning is enabled and 8,192 otherwise, ignoring the Opus 5.5 entry's declared 128K output limit.
Variations tried and related work (optional)
App Version
3.84.0 (also reproduced against main at 778ad3e18bd68c95281bc48e61c82122e61df6a4)
API Provider (optional)
Model Used (optional)
Zoo Code Task Links (optional)
N/A; the confirming workload is private.
Relevant logs or errors (optional)
"stop_reason":"max_tokens"
"output_tokens":16384
"output_tokens_details":{"thinking_tokens":16384}
{"type":"message_stop"}
Problem (one or two sentences)
Claude Opus 5.5 declares a 128K output limit, but Zoo Code resolves it to 8,192 tokens by default and 16,384 when reasoning is enabled on the direct Anthropic provider. Because Opus 5.5 always uses adaptive thinking, a valid agent turn can consume the entire allowance as thinking and stop before producing text or a tool call.
Context (who is affected and when)
This affects users running
claude-opus-5-5through Zoo Code's Anthropic provider, including Anthropic-compatible base URLs, especially on complex agent tasks.In the observed failure, the provider stream ended normally with
stop_reason: max_tokens. All 16,384 output tokens were reported as thinking tokens, and the next event wasmessage_stop. The transport delivered a valid terminal event; the model had simply spent the configured output allowance before producing visible output.The model entry already declares
maxTokens: 128_000, but it also declares bothsupportsReasoningBudget: trueandsupportsReasoningBinary: true.supportsReasoningBinaryselects adaptive thinking, whilesupportsReasoningBudgetroutes output sizing through Zoo Code's legacy hybrid-reasoning defaults.Reproduction steps
Use Zoo Code 3.84.0 or current
mainwith the Anthropic provider and selectclaude-opus-5-5.Leave
modelMaxTokensunset.Inspect the resolved model with reasoning unset or disabled:
It resolves to
8192, despite the model entry declaring128000.Set
enableReasoningEffort: trueand inspect the resolved model again. It resolves to16384.Create a message with reasoning enabled. The request uses adaptive thinking but sends
max_tokens: 16384.On a sufficiently complex agent turn, observe the response stop at
max_tokensbefore emitting text or a tool call.Expected result
Claude Opus 5.5 uses its declared 128K output limit on the direct Anthropic provider path. Enabling adaptive thinking does not replace that model limit with the legacy 16K hybrid-reasoning default.
Actual result
getModelMaxOutputTokens()interpretssupportsReasoningBudgetas a legacy hybrid-budget capability. It returns 16,384 when reasoning is enabled and 8,192 otherwise, ignoring the Opus 5.5 entry's declared 128K output limit.Variations tried and related work (optional)
a9ebf1a6a3f2ba9c8be964bf0c198dc8e3def8d1corrected Vertex Opus 5.5's declared output limit from 8K to 128K, but it did not change the direct Anthropic resolution path.App Version
API Provider (optional)
Model Used (optional)
Zoo Code Task Links (optional)
Relevant logs or errors (optional)