Advisor tool max_tokens parameter and no-charge refusals
Advisor tool: Now supports a max_tokens parameter to cap the advisor model's output per call, reducing latency and output token cost for workloads that don't need full-length responses. Set tools[].max_tokens on the advisor tool definition.
No-charge refusals: On the Claude API, you are no longer billed for a request when it returns stop_reason: "refusal" without Claude generating any output.
Fetched July 3, 2026

