The gist
When you learn something new, a textbook lays out knowledge as a chain: chapter 1 → chapter 2 → chapter 3. That works for a person who reads in order. But an AI agent doesn't read that way. An agent arrives with its own question and wants to pull out exactly the page it needs in 50 milliseconds. Not the first chapter, not the table of contents. One specific piece about one specific thing.
A Knowledge Atlas is a different way to organize knowledge. Not a book but a network. Not a table of contents but a map. Not a chain of chapters but 54 short layers, each about one topic, each with its own coordinates and connections.
In this lesson, we'll look at how to put together an atlas like this yourself: what a layer is made of, how the machine-readable INDEX.json works, why there are three HTML views (cards, tree, mind map), and how an agent finds the knowledge it needs for 150 tokens instead of 13,000.
Key concepts
- Layer book (layer): a short book on one topic (about 7–13 KB of text) with a standard structure: image → what it is → how it works → examples → combos → resources → cross-links
- Frontmatter: machine-readable metadata at the top of each layer (YAML): id, group, keywords, related_layers, summary_50w. The agent reads it first
- INDEX.json: a single map of all the layers in one JSON file. The agent looks here first (~80 tokens), then pulls the layer it needs
- Competency groups: 9 groups of skills (understand / create / work / build / automate / release / earn / mindset / production). Each group = a folder
- Retrieval tiers: 4 levels of reading depth (L0 INDEX → L1 frontmatter → L2 section → L3 full layer). The agent takes the minimum it needs
- Cross-links: each layer links to 3–7 related layers. This turns a set of files into a network
- Curriculum: a sequential route through the atlas for a person (for example, "2 weeks for a beginner": L1 → L2 → L6 → L7 → L8). Atlas for the AI + curriculum for the person = one knowledge base
- Mind map view: a force-directed graph (HTML) where each layer = a node and each connection = an edge. It helps you see the architecture with your own eyes
- Stub layer: a placeholder layer with frontmatter but without full content. It gets filled in as you use it. The atlas grows incrementally
- Median retrieval cost: the target cost of one agent lookup (target: ~150 tokens median). A metric you can measure
Theory
Textbook vs. atlas: two different architectures
A textbook assumes a reader who goes from beginning to end. Chapters are in order, and each builds on the one before. Pull chapter 7 out of the middle and it won't make sense without chapter 6.
A knowledge atlas works differently. Each layer is a self-contained page. To read L8 (Vector databases and RAG), you don't need to read L1–L7 in order. The layer explains itself, with an image, plain words and examples. Connections to other layers are listed at the end, but you don't need them to understand it.
This is critical for an AI agent. An agent doesn't read in order. When a user asks "how do I give my bot long-term memory?", the agent needs all of L8 right now, not a path through L1–L7. The atlas hands over exactly L8.
Anatomy of a single layer
In the course author's atlas, every layer has a strict structure. This isn't for looks; it's for retrieval. When all 54 layers have the same sections in the same order, the agent knows where to look.
Frontmatter (YAML):
---
id: L8-vector-rag
layer: L8
group: understand
title: "Vector databases and RAG · Long-term memory for AI"
keywords: [vector, rag, pinecone, vectorize, qdrant, embedding]
tier: foundation
related_layers: [L7, L51]
related_recipes: [r12, r17]
status: stable
summary_50w: "Long-term memory for AI. A vector database stores meanings as coordinates. RAG = find the piece → put it into the prompt → answer. Solves the context limit and hallucinations."
hot_section: "Image"
---This is the layer's passport. The agent can read just the frontmatter (~150 tokens) and already understand what it is, what it belongs to, what it's connected to and how mature the material is. For 90% of retrieval tasks, the frontmatter is enough.
The body of a layer: 8 sections
- 🎨 Image: a metaphor from everyday life. Without it, the layer doesn't count as done. The image helps you remember and find the layer; this is a rule for every layer
- 📖 What it is: in plain words, no jargon. Every term explained
- 🔬 How it works: the mechanism without the math. With pseudocode or a diagram
- 🛠 Use cases: 3–5 concrete examples
- 🎯 Combos (recipes): which ready-made combinations use this layer
- 📚 Resources: a table comparing options, videos, docs, tools
- 🔗 Connected layers: cross-links with a short explanation
- 📝 My notes: an empty section for the owner's personal notes
One layer is 7–13 KB of text. If it comes out bigger, the layer gets split in two. If it's under 5 KB, the layer is still a stub.
Nine competency groups
The layers are grouped into folders: 01-understand/, 02-create/ and so on. Nine groups isn't a random number. It's the path from understanding to production.
01-understand/ 🧠 Understand AI (L1-L9) — basics: LLMs, models, RAG, MCP 02-create/ 🎨 Create (L10-L17) — content generation, image, video 03-work/ ⚡ Work (L18-L22) — productivity, assistants 04-build/ 🛠 Build (L23-L29) — Claude Code, SaaS, agents 05-automate/ 🤖 Automate (L30-L35) — workflows, cron, integrations 06-release/ 🌐 Release (L36-L41) — deploy, hosting, domains 07-monetize/ 💰 Earn (L42-L48) — pricing, payments, marketing 08-mindset/ 🧭 Mindset (L50) — attitudes, principles 09-production/ 🛡 Production (L51-L54) — reliability, observability, security
Each group answers its own question. 01-understand: "what is this, anyway?" 04-build: "how do I build this?" 07-monetize: "how do I make money with this?" 09-production: "how do I keep it from going down?"
This split is useful for three reasons:
First, navigation for a person. You go into 04-build/, see 7 layers and pick the one you need. No need to scan 54 files by eye.
Second, filtering for the agent. In INDEX.json, every layer has a group field. The agent can say "give me only the production layers" and get 4 candidates instead of 54.
Third, gradual growth. The 02-create group is still thin: 4 finished layers out of 8. When you dig deeper into video generation, you'll fill in the rest. The atlas grows as you use it, not all at once.
INDEX.json: the machine-readable map
INDEX.json is the heart of the atlas. One file that holds metadata about all 54 layers. The agent reads it as the first step of any retrieval.
{
"version": "0.5.3",
"stats": {
"total_layers": 54,
"total_books_with_content": 54,
"active_groups": 9
},
"layers": {
"L8": {
"f": "library/01-understand/L8-vector-rag.md",
"g": "understand",
"k": ["vector", "rag", "embedding", "pinecone"],
"related": ["L7", "L51"],
"status": "stable"
},
"L51": {
"f": "library/09-production/L51-data-pipelines.md",
"g": "production",
"k": ["etl", "pipeline", "chunking"],
"related": ["L8", "L52"],
"status": "stable"
}
},
"curricula": {
"c01": {
"title": "Foundation from zero",
"layers": ["L1", "L2", "L6", "L7", "L8"],
"duration": "2 weeks"
}
}
}Why this format:
- Small: the whole INDEX is 15–30 KB. Reading it costs the agent ~5,000 tokens. That's a one-time cost per session
- Machine-readable: JSON can be parsed in any language. You can filter, search and sort it
- Complete: it contains everything needed to decide "where to go next" without reading the layers themselves
- Versioned:
version: 0.5.3shows how mature it is. It's updated along with the atlas
The 4-level retrieval pattern:
L0: Read INDEX.json once (~5K tokens, once per session) L1: Read frontmatter of layer (~150 tokens, usually enough) L2: Read targeted section (~300 tokens, if frontmatter isn't enough) L3: Read full layer (~3K tokens, only for a deep dive)
The median retrieval cost is 150 tokens. Compare that with ~13,000 tokens to "read the whole textbook chapter." That's an 87x improvement. Not from compression, but from precision.
Three HTML views: cards, tree, mind
The atlas isn't only markdown files for the agent. People need visual views. The course author's atlas has three HTML pages:
library.html: cards. Each layer is a card with its title, its image and a colored group label. Your eyes scan a 9×6 grid and find what you need in 3 seconds. Good when you roughly know what the topic is called but don't remember the exact number.
tree.html: a tree with drill-down. 9 groups → expand → 6 layers in a group → expand → the sections inside a layer. Hierarchical navigation. Good when you need to understand the structure: "how much do I have on production in total?"
mind.html: a force-directed graph. 45 nodes (some layers are merged), with connections showing the cross-links. You can move the graph around and drag nodes. Nodes are colored by group. Connections have different thicknesses: thicker when two layers link to each other. Good for spotting gaps: a lonely node with no connections is a candidate for more work.
INDEX.html is the entry page that ties the three views together and links to the curricula and recipes.
Curriculum: the bridge for people
The atlas is for the AI agent. But a person wants to learn from it too. Learning straight from 54 layers is hard: where do you start?
The answer is a curriculum. It's a sequential list of layers marked "read in this order." There are eight curricula right now:
| ID | Name | Duration | Layers |
|---|---|---|---|
| c01 | Foundation from zero | 2 weeks | L1, L2, L6, L7, L8 |
| c02 | Build a SaaS MVP | 4 weeks | L1, L9, L23, L25, L26, L36, L37, L38 |
| c03 | A local AI business in Ecuador | 6 weeks | ~10 layers |
| c04 | Content in Spanish | 3 weeks | content layers |
| c05 | Anthropic Mastery | 4 weeks | Claude layers |
| c06 | Successor onboarding | 12 weeks | the whole atlas + Persona |
| c07 | AI trader path | 6 weeks | data + automate layers |
| c08 | Academy Launch | 8 weeks | release + monetize |
A curriculum isn't separate new knowledge. It's a route through the atlas. The same layers, read in a particular order for a particular goal.
One atlas serves three audiences at once: the AI agent (through INDEX retrieval), the expert (direct search in the graph) and the beginner (a curriculum).
Cost and measurement
An atlas isn't just a structure; it's a system with metrics. Without metrics, you can't tell whether it's working.
What we measure:
retrieval_countin each layer's frontmatter: how many times the agent looked it up this monthlast_retrieved: when it was last accessed. Layers with a last_retrieved older than 6 months are candidates for deletion or mergingoutcome_score: a rating of answer quality after retrieval (1–5). If a layer keeps producing poor results, rewrite itmy_understanding: green/yellow/red. You mark your own level of knowledge
Target metrics:
- Median retrieval cost: ≤ 200 tokens
- Cross-link coverage: every layer is connected to 3+ others
- Stub ratio: no more than 20% of slots are placeholders
- Update freshness: stable layers are updated once a quarter
This turns the atlas into a living system. Layers nobody reads fade away. Layers people read often get better. The atlas itself shows you where to invest your time.
🧪 Practice
Let's build a mini-atlas of 10 layers on a topic of your own. This exercise takes 30–40 minutes. The topic is up to you: your hobby, your job, your future business. The main thing is that it has 10 different aspects.
Say the topic is "Podcasting from scratch." The mini-atlas will be about launching a podcast: equipment, editing, distribution, monetization.
Step 1: Create the folder and the INDEX
cd ~/Desktop
mkdir -p mini-atlas/library/{01-equipment,02-recording,03-edit,04-publish,05-grow}
cd mini-atlas
# Create an empty INDEX.json
cat > INDEX.json << 'EOF'
{
"version": "0.1.0",
"topic": "Podcasting from scratch",
"updated": "2026-10-04",
"stats": {
"total_layers": 0,
"active_groups": 5
},
"groups": {
"equipment": "Equipment",
"recording": "Recording",
"edit": "Editing",
"publish": "Publishing",
"grow": "Audience"
},
"layers": {}
}
EOFStep 2: Write 10 layers (2 per group)
One layer = one markdown file with frontmatter and 8 sections. Template:
cat > library/01-equipment/L1-microphone.md << 'EOF'
---
id: L1-microphone
layer: L1
group: equipment
title: "Microphone · What a beginner should buy"
keywords: [microphone, mic, podcast, audio, equipment]
related_layers: [L2, L3]
status: stable
summary_50w: "The microphone is a podcaster's main tool. Dynamic mics are better for noisy rooms, condenser mics for a studio. An affordable entry-level mic gives quality that's good enough for most purposes."
---
# L1 · Microphone
## 🎨 Image
The microphone is the listener's eyes. Through it, the listener sees your room, your desk, your breathing. A cheap microphone is a window with dirty glass.
## 📖 What it is
A dynamic microphone picks up sound in a narrow beam, cutting out room noise. A condenser mic picks up widely and cuts out only what isn't there.
## 🔬 How it works
... (diaphragm, coil, USB vs XLR)
## 🛠 Use cases
- Shure MV7: dynamic, USB+XLR
- Samson Q2U: dynamic, USB, an affordable model
- Rode NT1: condenser, XLR
- Check current prices in stores: they change
## 🎯 Combos
- Stack: microphone + audio interface + headphones
## 📚 Resources
- Videos: Podcastage YouTube reviews
- Stores: Sweetwater, B&H
## 🔗 Connected layers
- L2 (Audio interface): XLR microphones need an interface
- L3 (Headphones): needed for monitoring
## 📝 My notes
_[your notes]_
EOFDo the same for 10 topics:
- L1-microphone, L2-interface (equipment)
- L3-room-acoustics, L4-recording-software (recording)
- L5-editing-basics, L6-noise-cleanup (edit)
- L7-hosting, L8-rss-distribution (publish)
- L9-show-notes, L10-promotion (grow)
Step 3: Fill in INDEX.json
cat > INDEX.json << 'EOF'
{
"version": "0.1.0",
"topic": "Podcasting from scratch",
"stats": { "total_layers": 10, "active_groups": 5 },
"layers": {
"L1": { "f": "library/01-equipment/L1-microphone.md", "g": "equipment", "k": ["mic", "audio"], "related": ["L2", "L3"] },
"L2": { "f": "library/01-equipment/L2-interface.md", "g": "equipment", "k": ["xlr", "usb"], "related": ["L1"] },
"L3": { "f": "library/02-recording/L3-room-acoustics.md", "g": "recording", "k": ["echo", "absorption"], "related": ["L1"] },
"L4": { "f": "library/02-recording/L4-recording-software.md", "g": "recording", "k": ["daw", "audacity"], "related": ["L5"] },
"L5": { "f": "library/03-edit/L5-editing-basics.md", "g": "edit", "k": ["cut", "fade"], "related": ["L4", "L6"] },
"L6": { "f": "library/03-edit/L6-noise-cleanup.md", "g": "edit", "k": ["noise", "compressor"], "related": ["L5"] },
"L7": { "f": "library/04-publish/L7-hosting.md", "g": "publish", "k": ["buzzsprout", "anchor"], "related": ["L8"] },
"L8": { "f": "library/04-publish/L8-rss-distribution.md", "g": "publish", "k": ["rss", "apple-podcasts"], "related": ["L7"] },
"L9": { "f": "library/05-grow/L9-show-notes.md", "g": "grow", "k": ["seo", "notes"], "related": ["L10"] },
"L10": { "f": "library/05-grow/L10-promotion.md", "g": "grow", "k": ["social", "marketing"], "related": ["L9"] }
}
}
EOFStep 4: Test retrieval
Open Claude Code in the mini-atlas/ folder. Ask a question:
"What podcast equipment should I buy on a limited budget?"
Claude should:
- Read INDEX.json (it will see the
equipmentgroup with L1 and L2) - Read the L1 frontmatter (it will see the summary about mics for beginners)
- Read the "Use cases" section
- Answer with specific models
If Claude read all 10 files instead of two, retrieval isn't working. That's a sign you need to say it explicitly in CLAUDE.md: "Before searching, always read INDEX.json first."
Step 5: Add cross-links and a mind map
In each layer, list 2–3 connections in the "🔗 Connected layers" section. The goal is a dense graph with no lonely nodes.
For the mind map, you can use a ready-made template based on vis.js or d3.js. A minimal HTML file:
<!DOCTYPE html>
<html><head><title>Mini Atlas</title>
<script src="https://unpkg.com/vis-network/standalone/umd/vis-network.min.js"></script>
</head><body>
<div id="net" style="height:600px;border:1px solid #ccc"></div>
<script>
const nodes = new vis.DataSet([
{id:1, label:'L1 Mic', group:'equipment'},
{id:2, label:'L2 Interface', group:'equipment'},
{id:3, label:'L3 Acoustics', group:'recording'}
// ... all 10
]);
const edges = new vis.DataSet([
{from:1, to:2}, {from:1, to:3}
// ... cross-links from INDEX.json
]);
new vis.Network(document.getElementById('net'), {nodes,edges}, {});
</script></body></html>Open it in a browser and you'll see your first atlas as a graph.
Step 6: Measure the retrieval cost
Ask Claude 5 different questions on this topic. Count the tokens for each answer. If the average retrieval is over 500 tokens, the atlas isn't working efficiently. Trim the layers, tighten the frontmatter, check the INDEX.
The target metric for a 10-layer mini-atlas: 150–250 tokens median.
⚠️ Anti-patterns
❌ One big PDF instead of a network of layers. "I'll put all the knowledge into one 200-page file." This dies at retrieval: the agent reads all 200 pages for every question. A network of 54 short layers beats a single 200-page volume on cost per lookup.
❌ A layer without an image. "I'm an engineer, I don't need images." The image isn't for you. It's for the future reader who comes from a different field. Without an image, the layer doesn't connect with someone new to the topic. The author's rule: a layer without the "🎨 Image" section isn't done.
❌ Layers without cross-links. "Each layer is self-contained, so why connections?" Then what you have isn't an atlas, it's an archive. Connections are what turn 54 files into a network. Without them, the agent can't go from L8 to L51 to get the pipeline.
❌ INDEX.json as a checklist. If the INDEX only has paths and titles, it's useless. The INDEX should contain keywords, related_layers, summary_50w and group. The richer the INDEX, the more often the agent can get by with it alone, without digging into the layers themselves.
❌ An atlas without metrics. Making layers "just in case," without knowing which ones get used. A year later you discover that 60% of the layers were never requested. Measure retrieval_count from the first week.
🔗 Related
- RAG: Retrieval Augmented Generation: embedding search can sit on top of INDEX.json for semantic retrieval
- Obsidian as a second brain: an Obsidian vault can serve as the UI for the atlas if you'd rather have a ready-made tool than markdown + JSON
- Folder philosophy: PARA: PARA organizes projects, the atlas organizes knowledge. They complement each other
- MCPs: extending what Claude Code can do: you can serve the atlas through an MCP server so any Claude client can retrieve from it. The atlas as a service
✅ Checkpoint
Check yourself:
If you're confident on 6 or more, you can move on. If fewer, build another mini-atlas on a different topic. An atlas is built by hand, not by theory.
Sources
- The course author's atlas: 54 layers, 9 competency groups. The structure this lesson walks through
- Its INDEX.json (version 0.5.3): a working example of a machine-readable map
- Anthropic Contextual Retrieval (2024): the frontmatter + chunk pattern for high-quality RAG
- Greg Kamradt, "RAG From Scratch": the basics of retrieval cost
- vis.js / d3.js force graph: mind map visualization
→ Next lesson: Arsenal of Prompts: six reusable prompt modes
The mark stays in this browser only and is never sent anywhere. My progress