PriceIndex

Coding Models / Analysis

Claude 3.7 Sonnet Hybrid Reasoning: API Costs with Extended Thinking Tokens

Claude 3.7 Sonnet introduces hybrid reasoning at standard Sonnet pricing, but extended thinking tokens bill at the full $15/1M output rate. Here is how thinking budgets from 1k to 64k tokens impact API expenditures and how to engineer cost-effective agentic routing.

PriceIndex EditorialPublished Oct 5, 2026Updated Oct 5, 20268 min read
Abstract technical illustration of dual data streams representing standard inference tokens and extended reasoning token expansion.

Abstract technical illustration of dual data streams representing standard inference tokens and extended reasoning token expansion.

Hybrid Architecture and Unit Economics of Claude 3.7 Sonnet

Claude 3.7 Sonnet is a hybrid reasoning model: one model ID serves both fast standard responses and step-by-step reasoning, and the API caller chooses per request. By default it answers directly, with latency and cost similar to a conventional chat model. When the request enables extended thinking, the model first produces reasoning tokens, capped by a budget you set, and then writes the final answer. A dedicated reasoning model reasons on every prompt, so hybrid lets you pay for reasoning only on the requests that need it. Anthropic describes the design in its Claude 3.7 Sonnet announcement.

The base rates are unchanged from earlier Sonnet models. See Claude 3.5 Sonnet for the previous generation. The key billing rule is that thinking tokens have no separate reasoning rate. They are billed as ordinary output tokens at $15.00 per million, so every token spent thinking costs five times an input token. The extended thinking documentation explains how the budget is set and how it interacts with the max output limit.

Prompt caching is supported and works the same way as on other Claude models. A cache hit costs $0.30 per million tokens, a 90% discount on input. A 5-minute cache write costs $3.75 per million, a 25% surcharge. Confirm current figures on the Anthropic pricing page. For background on how input, output and cached rates combine, see our token pricing explainer.

  • Input: $3.00 per 1M tokens
  • Output, including extended thinking tokens: $15.00 per 1M tokens
  • Cache hit: $0.30 per 1M tokens (90% off input)
  • 5-minute cache write: $3.75 per 1M tokens (25% surcharge)

How Thinking Tokens Drive Exponential Output Cost Variance

Extended thinking is switched on per request with the thinking parameter: {"type": "enabled", "budget_tokens": N}. Per the extended thinking documentation, budget_tokens has a minimum of 1,024, and max_tokens must be larger than budget_tokens so the model has headroom for the visible answer after it finishes reasoning. The budget is a ceiling on reasoning, not a guaranteed spend, so the model may use less. Thinking tokens are billed as output tokens at the standard rate on the Anthropic pricing page, which for Claude 3.7 Sonnet is $3 per 1M input and $15 per 1M output.

The cost per request is (input x $3 + (thinking + visible output) x $15) / 1M. In ordinary calls, input dominates the token count. Illustratively, a 4,000-token prompt with a 400-token reply costs $0.012 for input and $0.006 for output, or $0.018 in total. The output-to-input token ratio is 0.1:1, and output is a third of the bill.

Thinking reverses this. Take the same 4,000-token prompt with a 16,000-token budget and a 500-token answer. Output reaches up to 16,500 tokens, or $0.2475, against the same $0.012 for input. The total is about $0.2595, roughly 14x the non-thinking call, and output is about 95% of the spend. The output-to-input ratio moves from 0.1:1 to about 4:1. Because the budget is set per request, output cost varies by an order of magnitude or more with a single parameter, so input-based estimates will understate the bill. For background on the billing units, see AI token pricing explained, and use model comparisons to benchmark non-thinking alternatives.

A clean isometric infographic illustrating the token billing distribution between standard LLM calls and extended thinking calls, highlighting the massive expansion of the output token block in warm amber, modern minimal technical aesthetic, dark background, no text.

How Thinking Tokens Drive Exponential Output Cost Variance

  • Minimum budget_tokens is 1,024, and max_tokens must exceed it.
  • Thinking tokens are billed at the output rate, $15 per 1M for Claude 3.7 Sonnet.
  • Worked figures here are illustrative; actual usage depends on how much of the budget the model uses.

Worked Cost Scenarios: From 1k to 64k Thinking Token Budgets

Key takeaways

  • Thinking tokens are billed at the $15 per 1M output rate, so the thinking budget dominates cost as it grows.
  • Going from instant mode to a 64k budget multiplies per-call cost by 40 in these examples, even with larger inputs.
  • Treat each budget as a worst-case ceiling and measure real thinking-token usage per task type before forecasting spend.

The scenarios below use Claude 3.7 Sonnet's list rates of $3 per 1M input tokens and $15 per 1M output tokens. Thinking tokens are billed as output tokens, so each call costs (input tokens × $3 + (thinking tokens + final answer tokens) × $15) / 1M. Rates can change, so check them against Anthropic's pricing page before you budget. The token counts are illustrative workloads, not measurements. They assume the model uses its entire thinking budget on every call, which makes each figure a ceiling for that budget, not an average.

Moving from instant mode to a 1,024-token budget raises cost per call by about 57%, from $0.027 to $0.0424. A 4,000-token budget takes it to $0.090, and a 16,000-token budget to $0.300. The 64,000-token stress test costs $1.080 per call, 40 times the instant baseline. Input cost stays small throughout. At 64k thinking, output is $1.020 of the $1.080 total, so the thinking budget is the main cost lever, not prompt size.

The budget is a cap, not a target, and the model can stop thinking earlier on easy prompts. The real per-call average therefore usually falls below these ceilings, but it depends on your task mix. Log actual output token counts from API responses and replace the assumptions here with your own distribution. The Claude Developer Platform reports usage per request, and the cost calculators can model different call volumes. If you are comparing against a newer model, Claude Sonnet 4 has its own page with current pricing.

  • Instant mode: 5,000 input + 800 output = $0.015 + $0.012 = $0.027 per call, or $27.00 per 1,000 calls.
  • Light reasoning (1,024 budget): 5,000 input + 1,024 thinking + 800 answer = $0.015 + $0.02736 = $0.04236 per call, or $42.36 per 1,000 calls.
  • Moderate reasoning (4,000 budget): 5,000 input + 4,000 thinking + 1,000 answer = $0.015 + $0.075 = $0.090 per call, or $90.00 per 1,000 calls.
  • Deep reasoning (16,000 budget): 10,000 input + 16,000 thinking + 2,000 answer = $0.030 + $0.270 = $0.300 per call, or $300.00 per 1,000 calls.
  • Maximum reasoning (64,000 budget): 20,000 input + 64,000 thinking + 4,000 answer = $0.060 + $1.020 = $1.080 per call, or $1,080.00 per 1,000 calls.

Cost Compounding in Multi-Turn Agentic Loops

Agents change the cost shape because every step is a new API call. A coding agent cycles through planning, tool execution, output parsing and error correction, and each pass can trigger a fresh thinking phase. Thinking tokens are billed as output at $15/1M for Claude 3.7 Sonnet, per Anthropic's pricing page. An illustrative 10-step loop with an 8k thinking budget per step can generate up to 80k thinking tokens, about $1.20 in output alone, before any visible answer or tool-call text. If a task retries after errors, that number grows with each retry.

Input is the second driver. Whether thinking blocks are carried forward depends on how your agent builds history. During a tool-use turn the API requires the last thinking block to be passed back unmodified, and anything you persist beyond that is re-billed as input at $3.00/1M on every later call. Check the extended thinking documentation for how earlier-turn blocks are handled, and avoid storing reasoning you do not need.

Prompt caching limits the input escalation. Suppose a static 20k-token system prompt and repository map is re-sent across 10 steps (illustrative). Uncached, that is 200k input tokens, or $0.60. With caching, cache reads cost $0.30/1M, plus a one-time write premium, so the same context costs roughly a fifth as much. See AI token pricing explained for the mechanics, and cloud pricing for the surrounding infrastructure costs.

Production Routing and Budget Optimization Best Practices

Treat thinking as an opt-in feature, not a default. Because thinking tokens are billed at the output rate, every request that enables extended thinking can cost far more than the same request in standard mode. Deterministic tasks such as schema validation, JSON conversion, semantic routing and basic code transformations gain little from reasoning, so keep them in instant mode.

For requests that do need reasoning, tier the budget by task instead of using one global value. The API enforces a minimum budget of 1,024 tokens, and the budget must be lower than max_tokens. Anthropic's extended thinking documentation describes it as a target rather than a strict cap, so you still need hard limits elsewhere.

To decide which path a request takes, put a low-cost classifier in front. A small model such as Claude 3.5 Haiku can inspect the incoming prompt and label it as standard or reasoning, and optionally assign a budget tier. Its cost is usually small next to a single unnecessary long thinking trace, but measure the misroute rate on your own traffic before relying on it.

Enforce limits at the gateway as well as in the router. Set organization and workspace spend and rate limits in the Claude developer platform, add request timeouts, and set client-side max_tokens ceilings. Then watch realized costs against list prices on the price tracker.

  • 1,024 to 2,048 thinking tokens: isolated unit test generation
  • 4,096 tokens: multi-file bug triage
  • 16,000+ tokens: reserved for hard architecture, math or long-horizon planning tasks, ideally behind an explicit flag
  • Instant mode: validation, formatting, routing and simple transformations

Comparative Economics: Claude 3.7 Thinking vs. DeepSeek R1 and Claude 4

DeepSeek R1 is the obvious price anchor. At launch its list rate was $0.55 per 1M input tokens and $2.19 per 1M output tokens, which makes raw output roughly 7x cheaper than Claude 3.7 Sonnet's $15 output rate. Because reasoning tokens bill as output on both models, that gap carries straight into thinking-heavy workloads. Rates differ by host and have changed over time, so check the DeepSeek R1 model page and your provider's current price list before committing.

The comparison is not only about unit price. R1 reasons by default and gives you no hard cap on the trace, so even simple prompts can produce long chains of thought that you pay for. Claude 3.7 Sonnet lets you set a thinking budget per request, or turn thinking off, so cost per call is easier to bound. Details are in Anthropic's extended thinking documentation. That control narrows the gap on easy tasks but does not close it on hard ones.

Successors keep the same structure. Claude Sonnet 4 stays at $3/$15, and Opus-tier models are priced higher, so the 3.7 rate card is the baseline for Anthropic's mid-tier reasoning cost. Confirm current figures on Anthropic's site, and run your own token mixes through our model comparison tool.

Sources & further reading

The details

Frequently asked questions

Thinking tokens have no separate reasoning rate. They are billed as ordinary output tokens at $15.00 per million, five times the $3.00 per million input rate. Confirm current figures on Anthropic's pricing page.