Skip to content

[BUG] Claude Opus 5.5 defaults to legacy 8K/16K output budgets on Anthropic #1853

Description

@jaszhix

Problem (one or two sentences)

Claude Opus 5.5 declares a 128K output limit, but Zoo Code resolves it to 8,192 tokens by default and 16,384 when reasoning is enabled on the direct Anthropic provider. Because Opus 5.5 always uses adaptive thinking, a valid agent turn can consume the entire allowance as thinking and stop before producing text or a tool call.

Context (who is affected and when)

This affects users running claude-opus-5-5 through Zoo Code's Anthropic provider, including Anthropic-compatible base URLs, especially on complex agent tasks.

In the observed failure, the provider stream ended normally with stop_reason: max_tokens. All 16,384 output tokens were reported as thinking tokens, and the next event was message_stop. The transport delivered a valid terminal event; the model had simply spent the configured output allowance before producing visible output.

The model entry already declares maxTokens: 128_000, but it also declares both supportsReasoningBudget: true and supportsReasoningBinary: true. supportsReasoningBinary selects adaptive thinking, while supportsReasoningBudget routes output sizing through Zoo Code's legacy hybrid-reasoning defaults.

Reproduction steps

  1. Use Zoo Code 3.84.0 or current main with the Anthropic provider and select claude-opus-5-5.

  2. Leave modelMaxTokens unset.

  3. Inspect the resolved model with reasoning unset or disabled:

    new AnthropicHandler({ apiModelId: "claude-opus-5-5" }).getModel().maxTokens

    It resolves to 8192, despite the model entry declaring 128000.

  4. Set enableReasoningEffort: true and inspect the resolved model again. It resolves to 16384.

  5. Create a message with reasoning enabled. The request uses adaptive thinking but sends max_tokens: 16384.

  6. On a sufficiently complex agent turn, observe the response stop at max_tokens before emitting text or a tool call.

Expected result

Claude Opus 5.5 uses its declared 128K output limit on the direct Anthropic provider path. Enabling adaptive thinking does not replace that model limit with the legacy 16K hybrid-reasoning default.

Actual result

getModelMaxOutputTokens() interprets supportsReasoningBudget as a legacy hybrid-budget capability. It returns 16,384 when reasoning is enabled and 8,192 otherwise, ignoring the Opus 5.5 entry's declared 128K output limit.

Variations tried and related work (optional)

App Version

3.84.0 (also reproduced against main at 778ad3e18bd68c95281bc48e61c82122e61df6a4)

API Provider (optional)

Anthropic

Model Used (optional)

claude-opus-5-5

Zoo Code Task Links (optional)

N/A; the confirming workload is private.

Relevant logs or errors (optional)

"stop_reason":"max_tokens"
"output_tokens":16384
"output_tokens_details":{"thinking_tokens":16384}
{"type":"message_stop"}

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions