Library · Memory and context: keeping the agent on track

Context Rot and 28 techniques to fight degradation

Confident user55 minUpdated: October 2026
16 of 105 in the library

Module: Tokens and context | Time: about 30 min theory + 25 min practice


The gist

Picture a notebook with a very large but finite number of pages. Every time you send Claude a message, it writes down the question, its answer, every file it read and every tool it used. Once the notebook is more than 60-70% full, Claude starts to get sluggish: it forgets instructions, repeats itself, hallucinates. That's Context Rot: quality degrading because the context is overloaded. The lesson Claude Code's built-in commands showed you that there are /compact and /clear buttons. This lesson explains the system: when, why, and 28 techniques to keep degradation from setting in at all.


Key concepts

  • Context Rot: answer quality degrading as the context window overfills
  • Context window: the maximum amount of information Claude holds in one session (as of October 2026: 1M tokens for Fable 5.1, Opus 5.5 and Sonnet 5.5, 200K for Haiku 4.5; current versions: What's current)
  • Tokens: the unit for measuring text: on current Claude models, ~1 English word ≈ 1.8 tokens (1M tokens ≈ 555K words, per Anthropic's documentation); other languages have a different ratio
  • /compact: smart compression: keeps the essence, drops the details, frees up most of the context
  • /clear: a new conversation with an empty context (doesn't touch your files)
  • Progressive disclosure: loading context in three levels: only what's needed, when it's needed

Theory

What Context Rot is

When you work with Claude Code for a long time in one session, the context window fills up:

Type this into the chat
System prompt (CLAUDE.md):        ~6,000 tokens
MCP tools (2-3 servers):         ~11,000 tokens (if the descriptions are loaded in full)
Skills (5-7 files):               ~8,000 tokens
Conversation history (1-2 hours): ~80,000 tokens
Files read:                      ~50,000 tokens

Total:                           ~155,000 / 200,000 (78%) on a model with a 200K window (Haiku 4.5)
                                 ~155,000 / 1,000,000 (15%) on a model with a 1M window (Opus 5.5, Sonnet 5.5)

The numbers in this block are illustrative, not a measurement.

A note on 1M context (as of October 2026): Fable 5.1, Opus 5.5 and Sonnet 5.5 have a million-token window, with no surcharge for long context in the API. That makes the problem less acute, but it doesn't make it go away. In the author's experience, quality starts degrading at 60-70% of any window. So with a 1M context, Context Rot will set in closer to 600-700K tokens, not 200K. The principles are the same, and the techniques below work at any window size.

One more detail: when the window is almost full, Claude Code compresses the history on its own. But by that point quality has already dropped, so a manual /compact earlier on is more reliable than the automatic one.

Above 60-70% full, the first symptoms of degradation appear. At 85%+, Claude starts doing things that will annoy you.


Symptoms of Context Rot: how to recognize it

🎨 Picture this: Context Rot progresses like fatigue in a marathon runner. At mile 6, fine. At mile 18, the pace drops. At mile 25, form falls apart, the knees hurt, the head is foggy. You can still finish, but the quality isn't there anymore.

Early signals (60-70% full):

  • Answers have gotten a bit longer than they need to be
  • Claude sometimes asks again about things you already discussed
  • Small deviations from the CLAUDE.md instructions

Middle stage (70-80%):

  • Repeats the same points across different answers
  • Forgets details you mentioned 30 minutes ago
  • Starts "guessing" instead of following the rules you set

Late stage (80%+):

  • Hallucinates: names files that don't exist, refers to conversations that never happened
  • Ignores the rules in CLAUDE.md (not because it doesn't want to follow them, but because it literally can't see them in the context anymore)
  • Contradicts itself within a single answer
  • Stops using the right style and tone

Emergency signal: Claude starts writing something that clearly contradicts the instructions you gave at the start of the session → /compact or /clear immediately.

🎨 Picture this: Claude with an overloaded context is like an employee who's been working for 18 hours without sleep. Technically they're still doing something. But the quality isn't there, and you can't trust the result.


Three context management commands

The lesson Claude Code's built-in commands explained what these commands do. Here's exactly when to use them.

/compact: smart compression

🎨 Picture this: /compact is an editor who takes a 50-page draft and hands back a 5-page digest. The essence is kept, the fluff is gone. You keep working from the digest.

What happens inside: Claude reads the whole conversation history, pulls out the key points, writes a condensed summary and replaces the history with it. Details are lost, the essence stays. You can hint at what to keep: /compact keep the database decisions.

When to use it: at 60-65% full (check with /context).

The result: most of the context is freed up. Claude still remembers what you were working on and where you ended up.

An example from practice:

Type this into the chat
Before /compact: window 78% full
After /compact:  roughly 22% full

You keep working in the same session, but with a fresh head.

/clear: a clean start

What happens: a new conversation with an empty context. Your project files are untouched. CLAUDE.md stays and gets loaded again. You start from scratch (you can bring back the previous conversation with /resume).

When to use it:

  • You're switching to a fundamentally different task
  • After finishing a big block of work
  • When you've already used /compact several times and quality still degrades
  • The session has run longer than 3-4 hours

How it differs from /compact: /compact compresses the history, /clear deletes it. After /clear, Claude doesn't remember what you were doing. After /compact, it remembers a short summary.

/context: diagnostics

What it shows (in newer versions it's a color map of the window with tips; below is an illustrative example in list form, the numbers are made up):

Code
Model: (your model's name)
Tokens: 51,300 / 200,000 (25.7%)

Breakdown:
  System prompt:  6,200 (3.1%)
  System tools: 4,800 (2.4%)
  MCP tools:  10,800 (5.4%)
  Skills:  7,200 (3.6%)
  Messages:  22,300 (11.2%)
  Free space:  148,700 (74.4%)

When to run it: every 30-40 minutes of intensive work. Make it a reflex.

What to do with the numbers:

  • 0-50%: all good, keep working
  • 50-65%: warning, you'll need /compact soon
  • 65%+: /compact right away
  • 80%+: /clear and start the session over

28 techniques for fighting Context Rot

Grouped into three categories: Architecture, Prompts, Management.


Group 1: Architecture techniques (prevention)

🎨 Picture this: architecture techniques are a well-designed kitchen layout. The cook works faster not because they run faster, but because the knife, the cutting board and the pot are all within reach in the right order. Nothing to hunt for.

1. CLAUDE.md as the single source of truth Don't repeat rules in every prompt. A well-written CLAUDE.md, written once, weighs less than repeating the same rules ten times in the chat.

2. Progressive Disclosure (L1/L2/L3) Load context in levels, only what the specific task needs:

Type this into the chat
L1 (always):     CLAUDE.md: the core rules, the project structure (~6K tokens)
L2 (per task):   the skill.md of the skill you need + API docs for the tool you need
L3 (on request): specific data files, only when they're truly needed

3. Small files in references/ Split big API documentation into small files by function. Instead of one 800-line ClickUp file, three files of 250 lines each: tasks.md, comments.md, webhooks.md. Claude will read only the one it needs.

4. SKILL.md under 500 lines A recommendation from the Claude Code documentation: keep the SKILL.md file under 500 lines. Details go into separate reference files that the skill pulls in when needed.

5. A context/ folder for long-term memory Write important decisions and agreements down in files instead of leaving them only in the chat history. What's written in a file will survive /clear. What's only in the chat won't. Modern Claude Code can also keep notes between sessions on its own (auto memory, managed with /memory), but anything critical is safer in your own files.

6. Archive old files Files you don't need anymore go into archives/. If they accidentally end up in the context through a search, that's wasted tokens.

7. Separate folders for outputs Everything Claude creates goes into outputs/. Don't mix results with instructions. When Claude looks for something, it looks in the right place and doesn't read extra stuff.

8. .env for secrets only Secrets go in .env, not in CLAUDE.md and not in the chat. Less sensitive material in the main context.


Group 2: Prompt techniques (saving tokens)

9. Break up tasks (Splitting Prompts) One big task = lots of tokens for the whole task context at once. Break it into steps. After each step Claude thinks more clearly, and you can see the progress.

Bad: "Build a CRM with a frontend, a database, an API, a dashboard and deployment" Good: first "create the project structure" → check it → then "add the database" → and so on.

10. Specific files instead of "look at the project" Don't tell Claude "look at the whole project." Point to specific files: "read outputs/newsletter-draft.md and improve the third paragraph." Fewer unnecessary reads = fewer tokens.

11. Ask one question at a time Several questions in one message → Claude tries to keep them all in mind at once → more context goes into internal "reasoning."

12. Use commands, not descriptions Instead of "could you create a file with such-and-such content in such-and-such folder" → "create outputs/report.md: [content]". Shorter = fewer tokens.

13. Point to files, don't paste into the chat Don't paste big blocks of text straight into the prompt. Better: "read context/business-data.md and make a report." Claude will read the file, which also costs tokens, but only once.

14. Clean up outputs after a task is done If outputs/ has piled up lots of files from old tasks, clean them up or archive them. Otherwise Claude might accidentally read them during a search.

15. Don't add MCP servers you don't need Every MCP server adds its tool descriptions to the context. In newer versions of Claude Code, only the tool names are loaded by default, and the full descriptions come in as needed (tool search), so an extra server costs less than it used to. But if you turn on full loading, two servers can easily eat ~11K tokens before you've typed a word. Add an MCP only when you really need it, and remove it when the task is done.

16. Turn MCP functions into Skills Once you've tested an MCP and it works, create a skill that does the same thing through a direct API call. Skills take fewer tokens than keeping an MCP switched on all the time.


Group 3: Session management techniques (response)

17. A /context schedule: every 30 minutes Make it a reflex: every 30-40 minutes of work → /context → check how full it is. Don't wait for symptoms, check preventively.

🎨 Picture this: waiting until 90% to run /compact is like changing your car's oil when the engine is already knocking. The right moment is every few thousand miles, not when something breaks.

18. /compact at 60%, not 90% Most people run /compact too late, when Claude is already degrading. Run it at 60%, before any symptoms appear. The result: not a single minute of degraded work.

19. /clear between big tasks Finished a block of work → /clear → start the next block with a fresh head. Don't try to hold the context of two different tasks in one session.

20. New tab = new agent In VS Code (and other editors with the Claude Code extension) you can open several Claude Code tabs, each with its own clean context. Parallel tasks → parallel agents. You can see background sessions with the claude agents command.

21. Subagents for parallel tasks Instead of one agent doing 5 tasks one after another (and piling up context), use 5 agents that each do one task and don't get in each other's way.

22. A final /compact before key deliverables Before Claude produces a final result (a post, a report, code going to production), run /compact first, then the request. Clean context = cleaner result.

23. Save key decisions in decisions/ Important technical decisions, project rules, agreements: write them into decisions/ right away. That preserves them outside the context window.

24. Regular /clear at the end of the workday At the end of the day, before closing VS Code: /clear. The next session will start clean and load a fresh CLAUDE.md without the baggage of yesterday's conversations.

25. Watch the length of your own messages A long prompt with lots of context inside it means tokens that stay in the conversation history. Decide: is this a rule for CLAUDE.md? Then put it there. Is it data for a specific task? Then put it in a file and have Claude refer to the file.

26. Use /statusline and /usage for monitoring You can set up the status line (/statusline) so that context usage is always visible, and /usage (also available as /cost) shows your spend and plan limits. If the cost of a single request suddenly jumps, that's a sign the context has bloated from extra files or an overly long history.

27. Spec → To-Do → Code (a three-phase approach) For complex tasks: first a spec (what we're building), then a to-do list by milestone, then code one milestone at a time. Each stage gets a clean context and doesn't drag along the tail of the previous one.

28. A subagent for heavy research If you need to go through a large amount of information (300 pages of documentation, 1,000 lines of code), delegate it to a subagent. It works in its own context and returns a condensed summary to you. Your main context stays clean.


Summary table: signal → action

How full Symptoms Action
0-50% All good Keep working
50-65% Minor deviations /compact needed soon, get ready
65-75% Repetition, forgetting details /compact right away
75-85% Ignoring CLAUDE.md rules /compact or /clear
85%+ Hallucinations, contradictions /clear + start the session over

What /compact keeps and what it loses

🎨 Picture this: /compact is like an experienced assistant after a long meeting: "Here's what we agreed on, here's what we dropped, here's the next step." The 50-page transcript goes in the trash, the 1-page summary is kept.

Keeps:

  • Decisions made during the session
  • Key facts about the project
  • The status of completed tasks
  • Major errors and how they were fixed

Loses:

  • Detailed explanations and reasoning
  • Intermediate steps that didn't lead anywhere
  • Long blocks of code that were discussed but not used
  • Small details of the conversation

Bottom line: if something matters, write it to a file. /compact keeps the essence, not the details.


Practice

Assignment: audit a live session

  1. Open Claude Code and work for 20-30 minutes on any task (create a few files, ask for some improvements, explore some code)
  2. Run /context and look at the numbers. Write down how full it is and the breakdown by component
  3. Figure out what takes up the most space: the message history? MCP tools? Files?
  4. Run /compact and look at the new numbers after compression
  5. Compare: what changed? What does Claude still remember? What did it forget?
  6. Write in your notebook which of the 28 techniques you already use and which you want to add

The goal: learn to read context usage like an oil pressure gauge: don't wait for a breakdown, keep an eye on it preventively.


Top 10 most effective techniques (summary table)

# Technique Category Effect Difficulty
18 /compact at 60%, not 90% Management Prevents degradation entirely Easy
17 /context every 30 min Management Catches the problem early Easy
2 Progressive Disclosure L1/L2/L3 Architecture Saves 50-80% of tokens Medium
9 Break up tasks Prompts Cleaner context, better results Easy
19 /clear between big tasks Management Fresh context for a new task Easy
1 CLAUDE.md as the single source Architecture No duplicated rules Medium
5 A context/ folder for memory Architecture Decisions survive /clear Easy
15 Don't add MCP servers you don't need Prompts Fewer unneeded tool descriptions in context Easy
21 Subagents for parallel tasks Management Isolated contexts Advanced
27 Spec → To-Do → Code Management Clean context at every stage Medium

Tools and resources

  • /context: the current token map (run it every 30-40 min)
  • /compact: smart compression of the history (use at 60-65%)
  • /clear: a full reset of the history (between big tasks)
  • /statusline: a status line with the metrics you want, /usage: spend and limits
  • Claude Code command list: full documentation on commands
  • Claude Pricing: official prices; a summary with the date it was checked: What's current
  • context/: a folder for long-term memory (survives /clear)
  • decisions/: a log of key decisions
  • archives/: an archive of outdated files

Common mistakes

Mistake 1: Never clearing the context in long sessions You work 4 hours straight and never once check /context. By the end, Claude is hallucinating, mixing up files, forgetting instructions. The reflex: /context every 30 minutes, /compact at 60%.

Mistake 2: /clear instead of /compact You lost 2 hours of working context with a single /clear. You had to explain all over again what you'd been doing. /compact would have kept the key decisions.

Mistake 3: Connecting 5 MCP servers "just in case" Every MCP server adds its tool descriptions to the context. Five servers can take up tens of thousands of tokens before you've typed a word (especially if the descriptions are loaded in full). Connect only the MCP servers you need for the current task.


Cross-references


Key takeaways

Context Rot isn't a bug, it's physics. The context window is finite. The goal isn't to beat the limit but to work with it smartly: load only what's needed, compress on time, save what matters to files.

/compact at 60% beats /compact at 90%. Prevention beats cure: one /context check every 30 minutes saves an hour of degraded work.

Whatever isn't written to a file won't survive /clear. Design your system so that everything valuable lives in files, and the context is just a working tool that you can and should clear out regularly.


Next lesson

→ RAG: Retrieval Augmented Generation: working with documents

The mark stays in this browser only and is never sent anywhere. My progress