Library · Power-user techniques

Extended Thinking: thinking things through

Engineer55 minUpdated: October 2026
46 of 105 in the library

Module: 9. Advanced features | Time: ~25 min theory + 30 min practice


The gist

Before a tough move, a chess player thinks for 5 minutes: runs through options, calculates consequences, throws out bad lines. Extended Thinking (Claude's deep-analysis mode) is the same thing for Claude: it "thinks out loud" before answering instead of going with the first thing that comes to mind. For simple tasks it's overkill. For architecture decisions and complex analysis, it makes a fundamental difference in quality.

Terms in this lesson: extended thinking (Claude's deep-analysis mode), API (application programming interface), token (a unit of text for AI), prompt (a request to the AI), prompt caching (saving a prompt for reuse so you don't pay for it again in full).


Key concepts

  • Extended Thinking: an API mode in which Claude generates an internal monologue before the final answer
  • Thinking tokens: the tokens of that internal reasoning, returned in a separate thinking block
  • budget_tokens: a parameter capping the maximum tokens spent on thinking (manual mode, only for older models: deprecated on 4.6, returns a 400 error on 4.7 and newer)
  • Adaptive Thinking: the automatic mode ("type": "adaptive"), where the model decides whether to think and how deeply. The depth is set by the effort parameter in output_config. The main approach for all current models (as of October 2026: Opus 5.5, Sonnet 5.5, Fable 5.1)
  • display: a display parameter, either "summarized" (a summary of the reasoning) or "omitted" (only a signature, no text). On Opus 5.5, Sonnet 5.5 and Fable 5.1 the default is "omitted"
  • Interleaved Thinking: thinking between each tool call, not only at the start
  • Thinking block: a separate block in the API response with the fields type: "thinking", thinking: "..." and signature: "..."
  • Cost: thinking tokens are billed as output tokens (pricier than input). You're billed for the full thinking tokens, even if display = "summarized"

Theory

How it works under the hood

🎨 Picture this: Without Extended Thinking, Claude answers like a Jeopardy! contestant who buzzes in before the clue is finished and blurts out the first answer. With Extended Thinking, it's like a chess player: looks at the board, calculates the options, throws out bad moves, then moves the piece.

Without Extended Thinking, Claude gets a request and immediately generates a response. It's fast, but the thinking is "flat": the model doesn't get a chance to check alternatives.

With Extended Thinking, the request goes through two stages:

Code
Request → [Thinking phase: Claude works through options] → Final answer

The thinking phase is invisible by default; you only see the final answer. Through the API you can get a summary of the internal monologue (the raw train of thought isn't returned under any settings).

What's new as of October 2026: on current models (Opus 5.5, Sonnet 5.5, Fable 5.1), thinking is already on by default, and the standard way to turn it off (thinking: {"type": "disabled"}) returns a 400 error. So the developer's job has shifted: not "turn thinking on," but "choose the depth" (effort) and decide whether to show the reasoning text. The old manual mode with budget_tokens is only needed for legacy models.

Two modes: Manual vs. Adaptive

The API has two ways to control thinking:

1. Manual (a manual budget): you set the token limit yourself. Works only on legacy models:

python
thinking={"type": "enabled", "budget_tokens": 10000}

2. Adaptive (automatic): the model decides how much to think, and you set the depth with the effort parameter, separately from thinking:

python
thinking={"type": "adaptive"},
output_config={"effort": "medium"}  # low / medium / high and above, depending on the model

🎨 Picture this: Manual is telling a chess player "think for exactly 5 minutes." Adaptive is saying "think a medium amount," and the player decides how much a particular position needs.

Which models support what

Model (as of October 2026) Manual ("enabled") Adaptive ("adaptive") Note
Claude Fable 5.1, Opus 5.5, Sonnet 5.5 ❌ returns a 400 error ✅ adaptive only Thinking is on by default ("disabled" returns 400); display defaults to "omitted"
Claude Opus 4.8 and Opus 4.7 ❌ returns a 400 error ✅ adaptive only Without the thinking field, thinking is off; turn it on explicitly
Claude Opus 4.6, Sonnet 4.6 ⚠️ deprecated ✅ recommended Manual still works
Claude Opus 4.5, Sonnet 4.5, Haiku 4.5 ✅ the only mode ❌ returns a 400 error No adaptive mode

Older models are gradually being retired from the API; current list: What's current.

The trend: Anthropic has moved to Adaptive. For new projects, use "adaptive" and effort.

An API request with Extended Thinking (Manual, for legacy models)

python
import anthropic

client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-haiku-4-5",  # manual mode: legacy models only; on Haiku 4.5 it's the only mode
    max_tokens=16000,
    thinking={
        "type": "enabled",
        "budget_tokens": 10000  # up to 10000 tokens for thinking
    },
    messages=[{
        "role": "user",
        "content": """Design a system architecture for the following case:
        - A SaaS platform for small businesses
        - 1000 active users
        - Multi-tenancy is required
        - Budget: $200/month for infrastructure
        - Team: 1 developer
        
        Evaluate at least 3 approaches with real trade-offs."""
    }]
)

# Go through the response blocks
for block in response.content:
    if block.type == "thinking":
        print("=== INTERNAL REASONING ===")
        print(block.thinking)
        print()
    elif block.type == "text":
        print("=== FINAL ANSWER ===")
        print(block.text)
python
response = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    thinking={
        "type": "adaptive",
        "display": "summarized"  # to see a summary of the reasoning
    },
    output_config={"effort": "high"},  # depth of work: low, medium, high and above (depends on the model)
    messages=[{
        "role": "user",
        "content": "Design the architecture of a multi-tenant SaaS platform..."
    }]
)

With Adaptive, you don't have to guess at budget_tokens: the model decides how deeply to think, and you only adjust effort. Keep in mind that at low effort, the model may skip thinking entirely on a simple request. If the SDK complains about output_config, update the library: pip install -U anthropic.

What you see in a thinking block

An example of a real internal monologue (shortened):

Type this into the chat
Hmm, I need to design an architecture for a SaaS on a tight budget...

Option 1: Shared database schema
- Pros: simple, cheap, one database
- Cons: hard to isolate customer data, risky as it grows
- OK for 1000 users, but what if it grows to 10000?

Option 2: Database per tenant
- Pros: full isolation, easy to roll back one customer
- Cons: $200/month won't cover it if 1000 customers = 1000 databases
- Ruling this out for this budget

Option 3: Schema per tenant (Postgres schemas)
- A compromise: isolation without an explosion in databases
- Row Level Security adds another layer
- Cloudflare Workers + PlanetScale serverless = fits in $200/month

I'll recommend Option 3 as the main one, and explain when to move to Option 2...

This isn't a staged showcase. It's the real process of working through options.

🎨 Picture this: budget_tokens is like a time limit on a meeting. 1,000 tokens is a quick brainstorm; 10,000 is a full strategic review. More time = deeper analysis, but also more expensive. In adaptive mode, effort plays the role of that limit.

The budget_tokens parameter (Manual mode)

budget_tokens When to use Cost (example calculation)
1,024 (minimum) Moderately complex tasks ~$0.005 per request
5,000 Architecture decisions, analysis ~$0.025 per request
10,000 Maximum depth, math ~$0.05 per request
32,000 Extremely complex tasks ~$0.16 per request

The math: thinking tokens × the output token price. The example uses $5 per 1M output tokens (what Haiku 4.5 costs as of October 2026: manual mode still works on it, and it may be retired from the API no earlier than October 15, 2026); thinking tokens are counted as output. Prices differ by model: current prices and versions: What's current. Check your real usage in the usage.output_tokens_details.thinking_tokens field of the response.

Limits: budget_tokens must be at least 1024 and less than max_tokens (except with interleaved thinking, where it can be more). The budget is a guideline, not a hard ceiling: the model may stop earlier; max_tokens sets the hard ceiling. For budgets above 32,000 tokens, the documentation recommends the Batch API: requests like that run long and hit timeouts.

The display parameter: what the user sees

It controls what comes back in the response's thinking block:

display What comes back What it's for
"summarized" A summary of the reasoning (the default on Opus 4.6, Sonnet 4.6 and older) Debugging prompts, understanding the model's logic
"omitted" An empty thinking field, only signature (the default on Opus 5.5, Sonnet 5.5, Fable 5.1) Production (faster time-to-first-token)

The field name is the same in both modes. The text in the block is always a summary, not the raw train of thought.

python
# Production mode: faster, no extra data
response = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    thinking={
        "type": "adaptive",
        "display": "omitted"  # ← only the signature, no reasoning text
    },
    messages=[{"role": "user", "content": "..."}]
)

Important: with "omitted" you still pay for all the thinking tokens. The savings aren't in money but in how fast the answer arrives. When you pass thinking blocks back in a multi-turn conversation (with tools this is required), pass them unchanged: the server decrypts the full reasoning context from the signature field.

Streaming thinking tokens

For long reasoning, streaming makes sense, because you can see the progress:

python
with client.messages.stream(
    model="claude-opus-5-5",
    max_tokens=16000,
    thinking={
        "type": "adaptive",
        "display": "summarized"
    },
    messages=[{"role": "user", "content": "...your complex request..."}]
) as stream:
    for event in stream:
        if event.type == 'content_block_start':
            if event.content_block.type == 'thinking':
                print("[Starting to think...]")
        elif event.type == 'content_block_delta':
            if event.delta.type == 'thinking_delta':
                print(event.delta.thinking, end='', flush=True)
            elif event.delta.type == 'text_delta':
                print(event.delta.text, end='', flush=True)
        elif event.type == 'content_block_stop':
            print("\n[Block finished]")

When streaming, three types of delta events arrive (with display: "omitted", thinking_delta arrives with an empty string and there's no reasoning text):

  • thinking_delta: the reasoning text (in chunks)
  • signature_delta: the cryptographic signature (for multi-turn)
  • text_delta: the final answer

When you need Extended Thinking

Use it for:

  • Architecture decisions (choosing technologies, system structure)
  • Multi-step math problems
  • Analyzing complex trade-offs with several variables
  • Debugging tricky bugs where the cause isn't obvious
  • Writing critical algorithms

Don't use it for:

  • Writing simple text or a short summary
  • Routine API calls and straightforward code
  • Tasks where the first answer is already correct
  • When you need speed and cost matters

🎨 Picture this: Bullet chess versus classical. In bullet you play a whole game in 1 minute: fast, shallow. In classical you think 20 minutes per move. Extended Thinking is classical. You don't need it for a speed game.

Interleaved Thinking: thinking between tool calls

🎨 Picture this: Interleaved Thinking is like a cook who tastes the dish after every step. Added salt, tasted it, decided what's next. Added spices, tasted again. Not cooking blind all the way to the end.

When Claude uses tools, it can think after each tool call, not only at the very start. That's called Interleaved Thinking.

Code
Request → [Thinking] → tool_use: get_weather("Chicago")
       → tool_result: "23°F"
       → [Thinking: "OK, it's cold in Chicago, need to factor that in..."]  ← thinks BETWEEN calls
       → tool_use: get_weather("Miami")
       → tool_result: "78°F"
       → [Thinking: "Miami is warmer, let me compare..."]
       → Final answer

Model support (as of October 2026):

  • Opus 5.5, Sonnet 5.5, Fable 5.1, plus Opus 4.8 and 4.7: interleaved automatically in adaptive mode, no header needed
  • Opus 4.6: only in adaptive mode (not available in manual)
  • Sonnet 4.6: automatic in adaptive mode; in manual it still works through the beta header interleaved-thinking-2025-05-14, but that's deprecated
  • Opus 4.5, Sonnet 4.5 and other Claude 4 models: through the beta header interleaved-thinking-2025-05-14
  • Haiku 4.5: not supported

Important when working with tools: when sending a tool_result back, always pass all the thinking blocks from the previous assistant response. Don't modify or remove them.

Example: with thinking vs. without

🎨 Picture this: Without Extended Thinking, it's advice from a stranger on the street: "go with PostgreSQL, everybody does." With Extended Thinking, it's a consultation with an architect who spent an hour studying your load, budget and team.

Without Extended Thinking, the request: "Pick a database for my SaaS"

The answer comes back in 2-3 seconds. It will most likely recommend PostgreSQL or MongoDB with boilerplate arguments.

With Extended Thinking (adaptive, effort high; in the old manual mode, budget_tokens: 8000):

Claude will spend noticeably more time on analysis (on the order of tens of seconds, depending on the model and load). In the thinking block you'll see it consider your specific parameters, compare the cost of different cloud databases at a load of 1000 users, take into account that you have one developer, and weigh PlanetScale vs. Supabase vs. Neon on real criteria.

The answer is noticeably better, not because the model is "smarter," but because it had time to think.

Limitations of Extended Thinking

Not every API feature is compatible with Extended Thinking:

Feature Compatibility Note
tool_choice: "auto" ✅ Works
tool_choice: "none" ✅ Works
tool_choice: "any" ❌ in manual; works in adaptive, except on Opus 5.5, Sonnet 5.5 and Fable 5.1 On those three models, forcing a tool call always returns a 400 error
tool_choice: {"type": "tool", "name": "..."} ❌ in manual; works in adaptive, except on Opus 5.5, Sonnet 5.5 and Fable 5.1 Same as above
max_tokens: 0 (cache warming) ❌ Incompatible
Prompt Caching (system prompt) ⚠️ When you change the thinking mode, budget_tokens or effort, the system prompt and tools cache can miss too: treat it as starting over
Prompt Caching (messages) ⚠️ Invalidated by any change to the thinking mode, budget_tokens or effort

Multi-turn: pass the thinking blocks along

🎨 Picture this: Thinking blocks in multi-turn are like handing a chess player's notebook to the next move. Without the notes, they'd forget which options they'd already ruled out and why. With the notes, they pick up right where they left off.

In multi-turn conversations, pass along all thinking blocks from previous assistant responses unchanged: inside a tool-use loop this is required, and in a regular conversation it's recommended. Otherwise the model can lose the context of its own reasoning:

python
# Turn 1
response1 = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    thinking={"type": "adaptive"},
    messages=[{"role": "user", "content": "First question?"}],
)

# Turn 2: pass response1.content in full (with the thinking blocks), unchanged
response2 = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    thinking={"type": "adaptive"},
    messages=[
        {"role": "user", "content": "First question?"},
        {"role": "assistant", "content": response1.content},  # ← ALL the blocks
        {"role": "user", "content": "Follow-up question?"},
    ],
)

Practice

Assignment: an architecture decision with Extended Thinking

  1. Pick a real architecture problem from your own project (or use a practice one: "How should I store user data for a SaaS with 500 customers?")

  2. First, request an answer without Extended Thinking and save it:

    python
    response_basic = client.messages.create(
        model="claude-haiku-4-5",  # for comparison: a model without thinking by default
        max_tokens=2000,
        messages=[{"role": "user", "content": YOUR_REQUEST}]
    )
  3. Then the same request with Extended Thinking (Manual: on Haiku 4.5 it's the only thinking mode):

    python
    response_thinking = client.messages.create(
        model="claude-haiku-4-5",
        max_tokens=8000,
        thinking={"type": "enabled", "budget_tokens": 6000},
        messages=[{"role": "user", "content": YOUR_REQUEST}]
    )
  4. And the same request with Adaptive Thinking (a current model, for example Opus 5.5):

    python
    response_adaptive = client.messages.create(
        model="claude-opus-5-5",
        max_tokens=8000,
        thinking={"type": "adaptive", "display": "summarized"},
        output_config={"effort": "high"},
        messages=[{"role": "user", "content": YOUR_REQUEST}]
    )
  5. Print all the final answers and the thinking blocks

  6. Compare: where is the analysis deeper? What did the "fast" answer miss?

  7. Estimate the cost of each option with response.usage

Goal: get a feel for the difference in quality and understand which tasks justify paying extra.


Tools and resources

  • Anthropic API docs: Extended Thinking and Thinking
  • Python SDK: pip install -U anthropic (get a recent version: the output_config and display parameters were added recently)
  • In Claude Code: the effort level is set with the /effort command (or the --effort flag), Ctrl+O shows the thinking (verbose mode), and the word ultrathink in a request asks for deeper thinking for one turn; in Claude Code on Opus 5.5, Sonnet 5.5 and Fable 5.1, thinking can't be switched off. More: model documentation
  • Models: Opus 5.5, Sonnet 5.5, Fable 5.1: adaptive only; Opus 4.6 and Sonnet 4.6: adaptive (manual is deprecated); Opus 4.5, Sonnet 4.5, Haiku 4.5: manual only
  • Prices: What's current, claude.com/pricing

Key takeaways

Extended Thinking isn't magic, it's time to work through options. Quality goes up, and so does cost. Thinking tokens are billed as output: to avoid overpaying, lower effort (in manual mode, budget_tokens), and set a hard ceiling with max_tokens. For all current models, use Adaptive Thinking ("type": "adaptive") and effort in output_config: the model decides how much to think. A manual budget_tokens on new models returns a 400 error. The display: "omitted" parameter speeds up responses in production but doesn't save money: billing is based on the full thinking tokens. Use it for architecture decisions and complex analysis. For simple tasks, it's unnecessary overhead.


What's next

→ Computer Use: controlling the screen

The mark stays in this browser only and is never sent anywhere. My progress