The gist
The context window (the text the AI can see) is like a computer's RAM: anything that doesn't fit is unavailable to the agent (an AI that carries out tasks on its own) right now. The art of working on large projects is the art of managing what's in memory at any given moment.
Terms in this lesson: context (the text the AI can see), token (a unit of text for AI), agent (an AI that carries out tasks on its own), prompt (a request to the AI), API (application programming interface, the way programs talk to each other).
Key concepts
- The context window as a limited resource
- 5 context management techniques
- The delegation pattern: subagents as isolated workspaces
- Semantic search across a large codebase
Theory
What the context window is
The size of the context window (measured in tokens, the units of text AI works in) depends on the model. As of October 2026, Claude Opus 5.5, Sonnet 5.5 and Fable 5.1 have a 1-million-token window, and Haiku 4.5 has a smaller one. Current figures: What's current. But even a million isn't infinite: a large real project fills it easily, and every extra token costs money and adds noise:
- 100 files of 500 lines of code each = ~300,000 tokens
- Add the conversation history (50 messages) = another ~20,000 tokens
- Add tool results (bash output, grep results) = another ~30,000 tokens
Total: ~350,000 tokens. For a 200,000-token window that's an overflow, and for a million-token window it's already a third, so the session gets slower and more expensive. The agent starts to "forget" the beginning of the conversation or get mixed up in the details. When the window gets close to its limit, Claude Code compresses the history on its own, but it does so at its own discretion; it's better to manage compression yourself.
Signs of overflow:
- The agent repeats questions it has already asked
- The agent "doesn't know" about files it read 20 messages ago
- Answers get less precise
- Speed drops and cost goes up
Technique 1: A lean CLAUDE.md, the minimal context file
claude.md holds the instructions Claude Code reads at the start of a session. Beginners put everything in there: the architecture, conventions, project history. It turns into a 5,000-word monster.
The right approach: claude.md contains only pointers, not content.
# Project: Invoice Automation ## Architecture See: docs/architecture.md ## API documentation See: docs/api-reference.md ## Current tasks See: docs/current-sprint.md ## Code conventions - Python 3.11+, type hints required - Tests with pytest, coverage >80% - Formatting: black + ruff
When the agent needs the architecture, it reads docs/architecture.md. When it doesn't, it doesn't read it and doesn't clutter the context.
The rule: claude.md should fit on one screen (the Claude Code documentation recommends keeping it under 200 lines). Everything else goes into separate files read on demand, or into rules in .claude/rules/. The file is created with the /init command. Besides that, Claude Code keeps an automatic memory (auto memory): it records lessons from your corrections across sessions on its own. You can view and edit it with the /memory command.
Technique 2: Delegating to subagents
A subagent works in its own isolated context (more in the Subagents lesson). That's the key insight. The exception is a fork: it inherits the whole conversation, so to save context, use a regular subagent.
The delegation pattern:
Instead of reading 50 files yourself and holding everything in memory:
Main agent: "Subagent, explore the src/payments/ folder and bring me back: a list of functions, what they do, and the 3 main problems" Subagent: [reads 20 files, analyzes them, returns a 500-line summary] The main agent gets: a compact summary (500 words), not 20 heavy files
The subagent "burns" its own context on the research. The main agent gets only the distilled result. It's like having a researcher who reads for you and brings back only what matters.
When to delegate:
- Analyzing a large folder or module
- Searching the whole codebase
- Generating a large amount of code (the subagent writes it and returns the finished result)
- Any task where the result is more compact than the process
Technique 3: /clear, a fresh start
When a task is done, use /clear to wipe the conversation history. The next task starts with a clean context.
A typical mistake: keeping one session open all day, moving from task to task. By evening the agent is carrying the whole day's history in memory: slow, expensive, less precise.
The right rhythm:
- Task 1: debug the login →
/clear - Task 2: write tests →
/clear - Task 3: write documentation →
/clear
Each task gets a clean context. The agent doesn't drag the weight of earlier tasks along.
When NOT to use /clear: if the next task depends on decisions made in the current one. In that case, it's better to finish both tasks in one session.
Technique 4: Temporary files for passing data
When several agents need to work together, they exchange data through files, not through context.
Example workflow:
Agent 1 (Researcher):
→ Analyzes competitors
→ Writes to analysis/competitor-research.md
Agent 2 (Writer):
→ Reads analysis/competitor-research.md
→ Writes an article based on the research
Agent 3 (SEO):
→ Reads article-draft.md
→ Optimizes it for SEO
→ Writes to article-final.mdEach agent works independently. No one holds the previous agent's work in memory, only the final result.
Naming temporary files:
_tmp/research-2026-10-04.md _tmp/draft-v1.md _tmp/seo-suggestions.md
The _tmp/ prefix and the date make it clear these are intermediate results. Add _tmp/ to .gitignore if you don't want to commit working files.
Technique 5: Progressive loading
Don't read everything at once. Read only what you need right now.
Wrong:
"Read the whole project and find where the login problem is"
The agent reads everything and clutters the context with files it doesn't need.
Right:
"Which files might be responsible for login? List them, don't read them."
→ [The agent gives a list: auth.py, middleware.py, routes/users.py]
"Now read only auth.py and find the problem"
→ [Reads one file, finds the problem]Progressive loading: map first, then details. Not the other way around.
Semantic search across a large codebase
When a project is large, grep searches for exact text. Semantic search searches by meaning.
Example: you're looking for "where payment errors are handled." Grep will find lines with payment_error or PaymentException, but only if you know the exact words. Semantic search will find all the code related to handling payment errors, even if it's named differently.
By default, Claude Code finds its way around a project by searching file names and text (glob and grep): it tries several word variants and reads what it finds. The documentation doesn't mention a separate semantic index. If you need search by meaning or memory across sessions, you connect external tools, for example claude-mem (a community project) or MCP servers with an index. Heavy searching is easy to hand off to the Explore subagent: it reads lots of files in its own context and returns a short summary.
A practical pattern:
"Find every place in the code where something like retrying failed requests happens"
The agent comes up with different phrasings and searches for them, and it finds the retry pattern even if the functions are called attempt_again, with_backoff or safe_execute.
The metric: the /context command
The /context command in Claude Code shows a breakdown of what's taking up space in the current context by category (system instructions, tools, memory and CLAUDE.md, skills, conversation history) and gives optimization tips. The output format changes from version to version; below is an illustration, not an exact copy:
Context: 45,000 tokens of the model's window System instructions and tools: ... Memory and CLAUDE.md: ... Skills: ... Conversation history: ...
Treat it like your computer's Task Manager. See that the architecture documentation takes up a lot of space and isn't needed anymore? Compress the conversation with /compact and say what to keep (for example, /compact keep the decisions about login), or start a new task with /clear. To compress only part of the conversation, use /rewind: pick a message and choose Summarize from here.
Practice
Task: diagnose the context of a real project
- Open any of your projects, or create a test one with 5-10 files
- Start a Claude Code session and read a few files
- Run
/contextand look at how the tokens are distributed - Find the "heaviest" component (a file or a chunk of history)
- Practice delegation: ask the agent to "call a subagent that will read folder X and return a summary"
- After the summary, run
/clearand start a new task with a clean context - Compare: how much faster does the agent respond at the start of a fresh session vs. the end of a long one?
Context management commands: cheat sheet
| Command | What it does | When to use it |
|---|---|---|
/clear |
Wipes the conversation history and starts a clean session | After finishing a task, before a new topic |
/compact |
Compresses the history: the agent summarizes the context; you can add an instruction about what to keep | When the context is filling up but you need to keep the current task's context |
/context |
Shows a breakdown of context usage by category | Diagnosis: what's taking up space? |
/usage |
Plan limits, cost and statistics (/cost is a synonym) |
Keeping an eye on the budget |
/memory |
Opens the memory files and CLAUDE.md for viewing and editing | Checking what Claude remembers about the project |
CLAUDE.md |
Loaded automatically at startup | Storage for permanent project instructions |
| Subagent | Works in an isolated context | Analyzing a large folder, generating a lot of code |
_tmp/ files |
Data exchange between agents | When several agents work as a pipeline |
Common mistakes
1. Never using /compact If the context overflows, the session degrades. /compact compresses the history while keeping the essentials. Claude Code can compress the context automatically too, but then it decides what's important on its own; a manual /compact with an instruction keeps the key things under your control. Use it in long sessions between stages of work.
2. Loading the whole project into context
❌ "Read all the files in src/ and find the problem" ✅ "Which files in src/ are related to login? List them without reading them."
Progressive loading saves thousands of tokens.
3. One endless session for the whole day By evening the agent is carrying the whole day's history in memory. The rule: one task = one session. /clear between tasks.
4. A huge CLAUDE.md (more than ~200 lines) CLAUDE.md is loaded in EVERY session. If it's 5,000 words, that's 5,000+ tokens right off the bat. Use pointers: "Architecture: see docs/architecture.md".
5. Not using subagents for heavy tasks If you need to analyze 20 files, delegate to a subagent. It will "burn" its own context and bring back the distilled result.
Tools and resources
/context: the built-in command that shows context usage/clear: wipes the conversation history/compact: compresses the history while keeping the essentialsclaude-mem(external repo): a long-term memory system built on files- MCP filesystem: extended access to the file system
/memory: view and edit the automatic memory and CLAUDE.md- Subagents: isolated contexts for heavy tasks (see the Subagents lesson)
Key takeaways
The context window is a limited resource. Spend it on what matters most right now. A subagent "burns" its own context and brings back the distilled result, so the main agent stays fresh. /clear after a task is basic developer hygiene. Don't drag yesterday's baggage into today.
Related lessons
- Tokens and context: the basics of what tokens are and how cost is calculated
- Context rot and 28 hacks: more techniques for working effectively with Claude Code, including context management
What's next
The mark stays in this browser only and is never sent anywhere. My progress