Library · Memory and context: keeping the agent on track

Tokens and context management in Claude Code

Confident user35 minUpdated: October 2026
15 of 105 in the library

Time: about 25 min theory + 10 min practice


The gist

Tokens are the currency of Claude Code. Every word Claude reads or writes costs tokens. If you don't manage what Claude reads, you pay for things you didn't need. Progressive loading is like a smart waiter who doesn't haul the whole menu to your table but first asks, "meat or fish?"


Key concepts

  • A token is a unit of text (a little over half an English word on current Claude models) and the basis of API pricing
  • Everything Claude reads while working costs tokens (claude.md, workflows, tools)
  • Progressive loading (L1/L2/L3): read only what's needed right now
  • A lean claude.md means fewer tokens, so every request is cheaper and faster

Theory

What a token is: the simple explanation

🎨 Picture this: tokens are like seconds on a phone call. Not words, but "ticks" that get used up on every character. Talk for a long time and you pay more. A to-the-point call costs less.

A token is a small piece of text. Not always one word: sometimes part of a word, sometimes a punctuation mark. For practical purposes: 1,000 tokens ≈ 555 words ≈ a little over one page of letter-size text (English text on current Claude models; on older models ≈ 750 words).

Why it matters: Claude models (and all LLMs) count the cost of work in tokens. Every API call has a price: input tokens (what you sent) + output tokens (what the model replied).

Example: if claude.md weighs 2,000 tokens and you run 100 workflows a day, claude.md alone eats 200,000 input tokens a day. At $2 per million input tokens (Sonnet 5.5, as of October 2026), that's $0.40/day just for the system prompt. Cutting claude.md in half saves $0.20/day. Current prices: What's current.


What Claude reads on every request

🎨 Picture this: every request to Claude is like a table set for dinner. Claude.md is the tablecloth (always there), the workflow is the dish on the menu (one at a time), and the conversation history is the dirty plates from earlier courses (they pile up). The more plates, the more you pay the busboy.

When you run a workflow or send a message to an agent, Claude Code assembles the context for the model behind the scenes. A typical mix:

  1. claude.md: the project's system prompt. Read always, on every request.
  2. The workflow file: the workflow being run. Read in full.
  3. Tools: descriptions of the tools the agent can use.
  4. MCP services: descriptions of connected MCPs (if any).
  5. Conversation history: earlier messages in this session.
  6. Supporting files: only when actually needed (scripts, references).

Together this can add up to 10,000–50,000 tokens per request, depending on how complex the project is.


Progressive loading: three levels

This is the key concept for using Claude Code efficiently. Instead of reading everything at once, the system reads the minimum it needs and goes deeper only when necessary.

Level 1 (L1): YAML frontmatter, about 100 tokens

🎨 Picture this: L1 is a book cover. You read the title, the year, the genre. If you need the book, you take it off the shelf and read inside. If not, you put it back. Claude does the same with workflows.

Every workflow or skill file starts with a YAML header:

yaml
---
name: newsletter-generator
description: Generates a weekly newsletter on a given topic
triggers: [newsletter, email-digest, weekly-summary]
---

At L1, Claude reads only this header: the name, the description, the triggers. That's about 50–150 tokens. If the task doesn't match this workflow, the full file isn't read. The agent checks all workflows at L1 and picks the right one.

Analogy: it's like index cards in a filing cabinet. You read the heading on the card, and only if it's the one you need do you pull the full file.

Level 2 (L2): the full workflow file, about 1,000–2,000 tokens

Once L1 matches, Claude reads the whole workflow file. This is where the full instructions live: steps, logic, parameters. It can run 500–2,000 tokens depending on how complex the workflow is.

Level 3 (L3): supporting files, only when actually needed

If a workflow uses an external script (helpers/parse_email.py) or a file of examples, Claude reads it only when it reaches the step that requires it. Not in advance, not every time: only when needed.


The "only what's needed" principle

The bad approach: claude.md holds everything you might ever want to use: every agent role, every workflow described in detail, every example pasted right into the text. All of it gets read on every request, even when you just need to write one email.

The good approach: claude.md holds only what the agent needs to understand its role and the project's structure. Workflow details live in workflow files. Examples live in separate files loaded at L3.


The /context command: an X-ray of the context window

🎨 Picture this: /context is an X-ray. Not just "something hurts," but an exact picture: here's the system prompt at 3%, here's the history at 75%, here's MCP at 20%. You see the problem and you know where to cut.

The /context command shows a map of current token usage. Sample output:

Type this into the chat
Context window usage: 225,000 / 200,000 tokens (112%)

System prompt (claude.md):     8,200 tokens  (3.6%)
MCP tool descriptions:        45,000 tokens (20.0%)
Current workflow:              2,100 tokens  (0.9%)
Conversation history:        169,700 tokens (75.4%)

What this example tells us:

  • MCP descriptions eat 20% of the context: maybe too many MCPs are connected
  • Conversation history takes up 75%: time for /clear or a new conversation
  • The context window is overflowing (112%): the model will "forget" the beginning of the conversation

This is a diagnostic tool. When an agent starts "forgetting" what it did earlier, the first thing you check is /context.


Practical rules for saving tokens

  1. A short claude.md: describe the agent's role and the project structure, not the workflows in detail
  2. Separate files for workflows: L2 loads only when needed
  3. Move examples into /examples: L3, not loaded for nothing
  4. /clear when the conversation gets long: conversation history is often the biggest token eater
  5. Don't connect MCPs you don't need: each MCP adds thousands of tokens of tool descriptions
  6. YAML frontmatter in every file: this is what makes L1 filtering work

The context window: what it means in practice

🎨 Picture this: the context window is like a desk. You can spread out a lot of paper, but the desk has edges. When there's no room left, old sheets fall on the floor: Claude can't see them anymore.

Claude has a limit on how many tokens it can see at once. That's the "context window."

Context window sizes by model (as of October 2026)

Model Context window Practical maximum Cost (input/output per 1M tokens)
Claude Fable 5.1 1,000,000 tokens ~750K (leaving room for the answer) $10 / $50
Claude Opus 5.5 1,000,000 tokens ~750K $4 / $20
Claude Sonnet 5.5 1,000,000 tokens ~750K $2 / $10
Claude Haiku 4.5 200,000 tokens ~150K $1 / $5

For the exact model IDs for the API and current prices, see the What's current page and the Anthropic docs (platform.claude.com/docs/en/about-claude/models/overview).

Important when switching to a new model: different model generations may count tokens differently, and the same text can weigh more on a newer model. Check /context and the API token counter after changing models, before you plan a budget.

What changed by October 2026:

  • 1M tokens of context on Fable 5.1, Opus 5.5 and Sonnet 5.5; Haiku 4.5 has a 200,000-token window
  • Older models (Haiku 3.5, Sonnet 4, Opus 4, Opus 4.1) have been retired from the Claude API
  • Haiku 4.5 may be retired from the API no earlier than 10/15/2026: keep an eye on the model deprecations page

When the window overflows, either the request fails or Claude "forgets" the beginning of the conversation (the model sees only the last N tokens). In long development sessions this is a normal situation, and it's handled with the /clear or /compact command.

Important: /clear wipes the conversation history, not the project files. Your workflows and code stay right where they are.

/compact is the gentler option: it compresses the history while keeping key decisions and context. Use /compact when you want to keep working, and /clear when you're switching to a different task.


Practice

Exercise: a token audit of your project

  1. Open the project from the lesson Claude.md: your project's system prompt in Claude Code
  2. Type the /context command
  3. Study the output: what eats the most tokens?
  4. If claude.md takes more than 3,000 tokens, find what can be moved into separate files
  5. Check your connected MCPs: are all of them actually used?
  6. After optimizing, run /context again and compare

Question to reflect on: if your project grows to 50 workflows, how important will the L1/L2/L3 architecture be?


Tools and resources

  • /context: shows how tokens are distributed in the current session
  • /clear: wipes the conversation history (not files)
  • /compact: smart compression of the history that keeps key decisions
  • /usage: plan limits, cost and session stats (the command used to be called /cost)
  • YAML frontmatter in workflows: makes L1 progressive loading possible
  • Claude Pricing: current plan prices; token prices by model are on the What's current page
  • Claude Console: managing API keys and usage
  • tiktoken: a Python library for counting tokens (it's OpenAI's, but useful for estimates)
  • Anthropic SDK: exact token counts through the Token Counting API (the client.messages.count_tokens() method)

Common mistakes

Mistake 1: Ignoring /context You've been working for 3 hours and the agent starts acting dumb: forgetting instructions, repeating itself. The cause: the context is 90% full. Make it a habit: check /context every 30–40 minutes.

Mistake 2: Everything in CLAUDE.md All the rules, all the examples, all the templates in one 5,000-token file. It gets read on every request. Move examples into separate files and use L2/L3 loading.

Mistake 3: Not using /compact Many people only know /clear. But /clear deletes all the context, so you have to explain the task all over again. /compact keeps the essentials and frees up a good chunk of the context. Use /compact early, before the window fills up, and /clear when you change tasks.


Cross-references


Key takeaways

Tokens are money. Every word Claude reads or writes costs money. Understanding this makes you a thrifty architect instead of a wasteful one.

L1/L2/L3 progressive loading isn't a technical detail; it's a design principle. Build your project so Claude reads only what it needs right now.

/context is an X-ray. When something is slow or expensive, run it first.


Next lesson

→ Context Rot and 28 techniques against degradation: what to do when the context has "gone stale." The /clear, /context and other commands are covered in the lesson Claude Code built-in commands.

The mark stays in this browser only and is never sent anywhere. My progress