The gist
The gap between $20 a month and $200 a month usually isn't 10 times more work. More often, it's 10 times more unoptimized code.
Same chatbot, same content machine, same agent. One developer pays $200 because they send everything to Opus, with no cache, in real time, and pass the entire history in every request. Another pays $20 for the same result, because they route models by difficulty, cache the system prompt, and batch-process anything that isn't urgent.
This lesson is a map of 9 cost engineering techniques. Each one saves 30% to 99% on its own type of workload. Stacked together, they can cut the bill by 5-10 times without losing quality.
🎯 Decision tree: is it worth optimizing right now
Before you spend 4-8 hours on caching and batch processing, check whether it makes sense.
Current costs > $200/month?
→ Yes → OPTIMIZE. The ROI is clear.
→ No → Will it grow to $200+ within 2-3 months?
→ Yes → Set up the infrastructure now (techniques 1 + 5)
→ No → Don't waste the time. Keep building features.The payback rule: the savings should be > $100/month to pay for the setup time (4-8 hours the first time + 2-3 hours debugging cache invalidation + 1 hour of quarterly maintenance).
Key concepts
- Right-sizing: choosing the model to fit the task's difficulty (Haiku → Sonnet → Opus), instead of "send everything to the smartest one just in case"
- Prompt caching: repeating context is cached for 5 minutes or 1 hour, and the price on a cache hit drops by 10 times or more
- Batch API: asynchronous processing of up to 100K requests with a deadline of up to 24 hours gets you a 50% discount
- Response caching: the model's final answer is saved in KV/Redis, and a repeat request returns the result without calling the LLM
- Embeddings classification: classification tasks are done with cosine similarity on embeddings, orders of magnitude cheaper than an LLM call
- Context discipline: managing the context window (RAG, sliding window, summary) instead of "just send the whole history"
- Local models: routine tasks (classification, translation, summarization) go to Ollama, and the costs drop to zero
- Budget alerts: automatic triggers at $50/$100/$200 to catch an anomaly before the end of the month
Theory
Technique 1: Right-size your models (Haiku vs. Sonnet vs. Opus)
Anthropic offers several tiers of models at different prices: Haiku (simple, fast tasks), Sonnet (the main workhorse), Opus (complex tasks) and Fable (the longest and hardest tasks). Using Opus to sort support tickets is like hiring a McKinsey consultant to sort your mail.
Pricing comparison (per 1M tokens, as of October 2026; current prices and versions: What's current):
| Model | Input | Output | Speed | When to use it |
|---|---|---|---|---|
| Haiku 4.5 | $1 | $5 | Very fast | Classification, simple answers, intent detection |
| Sonnet 5.5 | $2 | $10 | Fast | Most tasks: writing, coding, reasoning |
| Opus 5.5 | $4 | $20 | Slower | Complex reasoning, ADRs, complex code, strategic decisions |
| Fable 5.1 | $10 | $50 | The heaviest tier | The longest and hardest tasks; you don't need it for routine work |
Ratios (as of October 2026):
- Opus 5.5 / Sonnet 5.5: 2x (Opus 4.1 cost 5x as much as Sonnet 4), so switching from Opus to Sonnet saves less than it used to
- Sonnet 5.5 / Haiku 4.5: 2x on input and 2x on output, so moving from Sonnet to Haiku still gives noticeable savings
- Opus 5.5 / Haiku 4.5: 4x, Fable 5.1 / Haiku 4.5: 10x: this is the main savings gap: use Haiku properly for high-volume tasks
The 80/15/5 rule (a guideline, not a law):
- 80% of tasks should go to Haiku (classification, simple lookups, brief summaries): this is where the main savings are
- 15% of tasks go to Sonnet (drafts, code, multi-step reasoning)
- 5% of tasks go to Opus (architecture decisions, complex strategy, critical code review), which no longer hurts as much
Implementing routing in code:
// router.ts — picking a model by task type
type TaskType = "classify" | "summarize" | "draft" | "code" | "architect";
const MODEL_MAP: Record<TaskType, string> = {
classify: "claude-haiku-4-5", // cheap and fast
summarize: "claude-haiku-4-5", // cheap and fast
draft: "claude-sonnet-5-5", // needs coherence
code: "claude-sonnet-5-5", // needs correctness
architect: "claude-opus-5-5", // needs deep reasoning
};
function pickModel(task: TaskType): string {
return MODEL_MAP[task];
}
// Usage
const model = pickModel("classify");
const response = await anthropic.messages.create({
model,
max_tokens: 100,
messages: [{ role: "user", content: userMessage }],
});You'll save: a noticeable share of your costs if you route correctly (the key is sending high-volume tasks to Haiku, not Sonnet). See Case 1 below for an example (the numbers are illustrative).
An important warning about the tokenizer: models 4.7 and newer (including Opus 5.5 and Sonnet 5.5) use a new tokenizer: the same text comes out to roughly 30% more tokens than with Sonnet 4.6 and earlier models (according to Anthropic's documentation as of October 2026). Factor this into your math when moving off older models.
About models being retired from the API: as of October 2026, Haiku 4.5 is still in the API, but the earliest possible date for its retirement is 10/15/2026. Keep an eye on the model deprecations page and keep your model names in one place in your code, like in router.ts above.
Technique 2: Prompt caching (90% off)
Anthropic's prompt caching means a repeating block of context (system prompt, documents, brand-voice.md, codebase context) is cached on Anthropic's side. On the next call within 5 minutes, that block costs 10 times less on input (on Opus 5.5 and Fable 5.1, the cache read discount is even bigger).
When it works:
- The block is at least the model's minimum (as of October 2026: 512 tokens for Sonnet 5.5 and Opus 5.5, 4,096 tokens for Haiku 4.5; a shorter block won't be cached)
- The same context is reused within 5 minutes (default TTL) or 1 hour (extended TTL, slightly more expensive to write)
- A cache hit is an exact match on the cached prefix
Multipliers (as of October 2026):
| Operation | Multiplier | Duration |
|---|---|---|
| Cache write, 5-min | 1.25x the base input price | 5 minutes |
| Cache write, 1-hour | 2.0x the base input price | 1 hour |
| Cache read (hit) | 0.1x (90% off; 0.05x on Opus 5.5, 0.025x on Fable 5.1) | until the TTL ends |
Specific amounts for Sonnet 5.5 (base input $2/MTok, as of October 2026):
| Type | Write | Hit | No cache |
|---|---|---|---|
| 5-min TTL | $2.50 / 1M | $0.20 / 1M | $2 / 1M |
| 1h TTL | $4.00 / 1M | $0.20 / 1M | $2 / 1M |
Specific amounts for Haiku 4.5 (base input $1/MTok):
| Type | Write | Hit |
|---|---|---|
| 5-min TTL | $1.25 / 1M | $0.10 / 1M |
| 1h TTL | $2.00 / 1M | $0.10 / 1M |
A write costs 25% more, but a hit is 10 times cheaper than the base price. The cache pays for itself after the first cache hit (5-min) or after two (1-hour).
Implementation (Anthropic SDK):
import anthropic
client = anthropic.Anthropic()
# A large system prompt (longer than the cache minimum): we cache it
SYSTEM_PROMPT = """[long brand voice + style guide + context — 5000 tokens]"""
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1000,
system=[
{
"type": "text",
"text": SYSTEM_PROMPT,
"cache_control": {"type": "ephemeral"} # 5-min cache
}
],
messages=[
{"role": "user", "content": "Write a post about X"}
]
)Where caching pays off most (approximate, depends on how much of the context repeats):
- A chatbot with a large system prompt (FAQ + tone + examples): 70-80% savings
- A code assistant with project context: 50-70%
- A RAG system with document citations: 40-60%
You'll save: usually 40-60% of total costs on a read-heavy workload.
Source: platform.claude.com/docs/en/build-with-claude/prompt-caching
Technique 3: Batch API (50% off)
With Anthropic's Batch API, you send up to 100,000 requests in one batch and get all the answers back within 24 hours. The discount is 50% on all tokens (both input and output).
Batch pricing for current models (as of October 2026):
| Model | Standard input/output | Batch input/output |
|---|---|---|
| Opus 5.5 | $4 / $20 | $2 / $10 |
| Sonnet 5.5 | $2 / $10 | $1 / $5 |
| Haiku 4.5 | $1 / $5 | $0.50 / $2.50 |
Where it works perfectly:
- Generating blog content (10 articles by morning)
- Bulk translation
- Summarizing an archive
- Precomputing embeddings (though embeddings have their own cheap models)
- Test runs of prompts on different models
Where it WON'T work:
- Real-time chat
- Customer support
- Voice agents
- Anything where the user is waiting for an answer in < 1 minute
Implementation:
import anthropic
client = anthropic.Anthropic()
# Create a batch of 1000 summarization requests
requests = []
for article in articles:
requests.append({
"custom_id": f"article-{article.id}",
"params": {
"model": "claude-sonnet-5-5",
"max_tokens": 200,
"messages": [
{"role": "user", "content": f"Condense into 3 bullet points:\n\n{article.text}"}
]
}
})
batch = client.messages.batches.create(requests=requests)
print(f"Batch ID: {batch.id}, status: {batch.processing_status}")
# Check back in an hour or two
result = client.messages.batches.retrieve(batch.id)
if result.processing_status == "ended":
# Download the results
for output in client.messages.batches.results(batch.id):
print(output.custom_id, "".join(b.text for b in output.result.message.content if b.type == "text"))You'll save: 50% on any bulk job. A sample calculation (Sonnet 5.5, as of October 2026): 1,000 articles a month, ~3,000 tokens of input and ~200 of output each. Without batch, that's 3M × $2 + 0.2M × $10 = $8; with batch, $4.
Source: platform.claude.com/docs/en/build-with-claude/batch-processing
Technique 4: Local models (Ollama) for routine tasks
Open-source and open-weights models (the Llama, Qwen, Gemma, Mistral, DeepSeek and gpt-oss families) run locally through Ollama. Free. No bill. For the current list of models and tags, see the Ollama library.
Hardware (rough guidelines; depends on the model and compression):
- About 16 GB of memory (a Mac with unified memory or a graphics card with 12 GB) → 7B-14B models
- About 32 GB → models up to 32B
- 64 GB or more, or a graphics card with 24 GB → 70B in compressed (quantized) form
Where local models work (on simple tasks the quality is close to cloud models):
- Classifying short texts
- Entity extraction (NER)
- Simple translations EN↔︎ES
- Summarizing short documents
- Sentiment analysis
- Intent detection in a chatbot
- Reformatting (JSON → markdown, and back)
Where local models DON'T work (the quality is noticeably lower than the cloud):
- Production-grade code generation
- Long-context reasoning (>32K tokens)
- Complex multi-step reasoning
- High-level creative writing
- Complex tool use / function calling
Installing Ollama:
# Mac (5 minutes)
brew install ollama
ollama serve # starts a local server at http://localhost:11434
# Download a model
ollama pull llama3.3:70b # example: ~40GB, takes 10-30 minutes
ollama pull qwen2.5-coder:32b # example: ~20GB, a coding model
ollama pull qwen3:8b # example: a small general-purpose model
# new models and tags are at ollama.com/library
# Test
ollama run llama3.3 "Classify: 'Bought a ticket' → spam/inbox/promo"Using it through the OpenAI-compatible API:
from openai import OpenAI
# Ollama exposes OpenAI-compatible endpoint
local = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
response = local.chat.completions.create(
model="llama3.3:70b",
messages=[
{"role": "user", "content": "Classify this email: 'Bought a plane ticket' → spam/inbox/promo"}
]
)
print(response.choices[0].message.content)You'll save: 100% of costs on the tasks a local model can handle. We're not counting electricity: a laptop under load draws something like 30W, which is cents a month.
See the Local AI models lesson for a deep dive into local models.
Source: ollama.com
Technique 5: Response caching (KV / Redis)
If 30% of your requests repeat (typical support questions, standard translations, classic FAQs), cache the final answer in a KV store or Redis.
The logic:
// hash request → KV lookup → if found, return cached; if not, call the LLM + save
import { createHash } from "crypto";
async function getCachedOrCall(userMessage: string, env: Env): Promise<string> {
const key = createHash("sha256").update(userMessage.toLowerCase().trim()).digest("hex");
// 1. Lookup in KV
const cached = await env.RESPONSE_CACHE.get(key);
if (cached) {
console.log("Cache HIT");
return cached;
}
// 2. Cache miss: call the LLM
const response = await callClaude(userMessage);
// 3. Save for 30 days
await env.RESPONSE_CACHE.put(key, response, { expirationTtl: 60 * 60 * 24 * 30 });
return response;
}Cloudflare KV pricing: there's a free tier with daily limits on reads and writes, and paid usage beyond that; check the current limits and prices on Cloudflare's pricing page.
Most projects stay within the free tier.
Where it works:
- A customer support chatbot (typical questions are ~30%)
- Multi-language translations of standard phrases
- Search autocomplete
- "Similar products / recommendations," if they're stable
You'll save: up to 30-50% of costs if you have a repetitive workload.
Technique 6: Embeddings instead of an LLM for classification
The question "which category does this text belong to?" doesn't need an LLM. Embeddings + cosine similarity are enough.
Pricing comparison (as of October 2026):
| Approach | Price / 1M tokens | Latency (approximate) |
|---|---|---|
| LLM classify (Sonnet 5.5) | $2 input + $10 output | 1-3 sec |
| OpenAI text-embedding-3-small | a few cents; see OpenAI's pricing page | 100-300 ms |
| Other embedding models (Voyage AI, Cohere and others) | cents; see the provider's website | 100-500 ms |
Embeddings are orders of magnitude cheaper than Sonnet 5.5 on input, not counting output (which embeddings don't have at all).
The logic:
from openai import OpenAI
import numpy as np
client = OpenAI()
# 1. Embed the categories ahead of time
CATEGORIES = {
"spam": "Promotional messages, scams, phishing, unwanted offers",
"billing": "Questions about payments, invoices, refunds, subscriptions",
"tech": "Technical problems, bugs, a feature not working",
"feature": "A new feature request, a suggestion, an idea",
}
def embed(text: str) -> np.ndarray:
r = client.embeddings.create(model="text-embedding-3-small", input=text)
return np.array(r.data[0].embedding)
category_embeddings = {name: embed(desc) for name, desc in CATEGORIES.items()}
# 2. Classify a new message
def classify(message: str) -> str:
msg_emb = embed(message)
similarities = {
name: np.dot(msg_emb, emb) / (np.linalg.norm(msg_emb) * np.linalg.norm(emb))
for name, emb in category_embeddings.items()
}
return max(similarities, key=similarities.get)
print(classify("I can't log in to my account")) # → "tech"You'll save: around 99% on classification compared with an LLM. An example at 100K classifications/month (avg. 30 input + 5 output tokens):
- On Sonnet 5.5: ~$6/month input + $5/month output ≈ $11/month
- On Haiku 4.5: ~$3/month input + $2.50/month output ≈ $5.50/month
- On embeddings: pennies a month (the input is about the same 3M tokens, at an embedding model's price)
Embeddings win by a wide margin even against the cheapest model, Haiku 4.5.
See the RAG lesson for a deep dive into RAG and embeddings.
Technique 7: Streaming + early stopping
If the result can be shorter than max_tokens, use streaming with stop conditions. Output tokens are 5 times more expensive than input, so trimming output pays off.
Example: a classifier that returns one word. Without streaming, you set max_tokens=10 "just in case." With streaming + stop="\n", the model returns one word and stops.
with client.messages.stream(
model="claude-haiku-4-5",
max_tokens=10,
stop_sequences=["\n", ".", ","], # stop at any separator
messages=[
{"role": "user", "content": "Category in one word: 'I can't log in'"}
]
) as stream:
for text in stream.text_stream:
print(text, end="")You'll save: 20-40% on output tokens where the result is short.
Technique 8: Context window discipline
The most common beginner mistake: "let's fix the problem by just sending more context." Every new request pays again for the entire accumulated history, so the bill for a long conversation grows faster than the conversation itself.
Anti-pattern:
# ❌ BAD: sending the whole history with every request
messages = []
for turn in conversation_history: # could be 100+ turns
messages.append({"role": turn.role, "content": turn.text})
messages.append({"role": "user", "content": new_message})
# Result: 50K input tokens for every answer → ~$0.10 per turn (Sonnet 5.5, as of October 2026) → ~$10 for a 100-turn conversationThe right way:
# ✅ GOOD: sliding window + summary
def build_context(history, new_message):
# Last 10 turns in full
recent = history[-10:]
# Older turns get summarized (or use RAG retrieval)
if len(history) > 10:
old_summary = summarize(history[:-10]) # done once, cached
messages = [{"role": "system", "content": f"Context:\n{old_summary}"}]
messages.extend(turn.to_message() for turn in recent)
else:
messages = [turn.to_message() for turn in history]
messages.append({"role": "user", "content": new_message})
return messagesContext discipline techniques:
- Sliding window: the last N messages
- Summarization: the old part → one compact block
- RAG retrieval: pull only the relevant chunks instead of sending the whole knowledge base
/compact: in Claude Code, manually compresses the session history; when the context window fills up, compaction also kicks in automatically (the threshold is set with the/autocompactcommand)
You'll save: 50-70% of costs on long conversations.
Technique 9: Free tier hopping (ethically, not for production)
Free tiers exist, but only for prototyping and personal projects, not for production:
| Provider | What's free (as of October 2026) | Good for |
|---|---|---|
| Gemini API (Google AI Studio) | Flash models have a free tier with limits; 3.1 Pro has no free tier | Experimenting with long context |
| Groq | A free tier with request limits (see the console for exact limits) | Experimenting with speed |
| Cloudflare Workers | 100,000 requests a day on the free plan | A free pet project, a proxy to an LLM |
| Deepgram | Free credits when you sign up (see the website for terms) | Speech-to-text experiments |
| Anthropic credits for startups | By application; terms on Anthropic's website | If you get into the program |
The terms of free tiers for hosting and APIs change often: check the pricing page before building on them. Current prices and versions: What's current.
Ethics and limits:
- Production requires an SLA, and free tiers don't give you one
- Risk of a ban for commercial use where the free tier is personal-only
- Read the terms of service carefully
- Don't use a free tier as permanent infrastructure: it's unfair to the provider and unstable
You'll save: 100% during the exploration phase. After that, find a paid tier with an adequate SLA.
Savings cases
The numbers in the cases below are illustrative: they're practice calculations, not measurements from real projects. Model prices are as of October 2026.
Case 1: Content pipeline (a real estate agency blog)
| Parameter | Before | After |
|---|---|---|
| Volume | 30 posts/month | 30 posts/month |
| Models | Everything on Sonnet 5.5 | Haiku 4.5 for tagging, Sonnet 5.5 for drafts, Opus 5.5 for one final piece a month |
| Caching | None | Prompt cache on brand-voice.md (5K tokens) |
| Batch | Real time | Batch API overnight for 25 of the 30 posts |
| Cost | $180/month | $35/month |
Saved: 80% ($145/month = $1,740/year)
Where the savings come from: the main contribution isn't the cheaper models themselves (Opus is now only twice the price of Sonnet, which isn't critical), but Haiku 4.5 on the high-volume tagging tasks + the Batch API overnight + the prompt cache on brand-voice.md.
Case 2: A customer support bot
| Parameter | Before | After |
|---|---|---|
| Volume | 5,000 conversations/month | 5,000 conversations/month |
| Models | Sonnet 5.5 on every turn | Haiku 4.5 for intent classification, Sonnet 5.5 only for complex cases |
| Response cache | None | KV cache for the 30% of typical questions |
| Embeddings | None | Embeddings for routing + FAQ matching |
| Cost | $250/month | $75/month |
Saved: 70% ($175/month = $2,100/year)
The main driver: embeddings instead of an LLM for routing.
Case 3: Voice support (see the Call Support AI lesson)
| Parameter | Before | After |
|---|---|---|
| Stack | One end-to-end speech model in Vapi | Separate STT + Sonnet 5.5 + TTS (ElevenLabs) |
| Per minute (approximate) | $0.31/min | $0.13/min |
| 1,000 min/month | $310 | $130 |
| 5,000 min/month | $1,550 | $650 |
Saved: 58% per minute ($180/month at 1,000 minutes, $900/month at 5,000 minutes)
The advertised $0.05/min is only Vapi's platform fee: speech recognition, the LLM, the voice and telephony are billed separately at the providers' prices (as of October 2026), so the final price per minute depends on the stack you choose.
Cost monitoring tools
Without monitoring, optimization is flying blind. Set this up from day 1.
| Tool | What it gives you | Price (as of October 2026) |
|---|---|---|
| Anthropic Console | Daily breakdown, model usage, spend limits | Free |
| claude.com/pricing | Current subscription and API prices | Free |
| OpenAI Usage | The same for OpenAI | Free |
| Helicone | Observability + cost per request. In maintenance mode since March 2026 after being acquired by Mintlify: fixes still ship, but no new features (announcement) | See the website for terms |
| LangSmith | Tracing + cost per trace, tied to the LangChain ecosystem | A free Developer plan with a monthly trace limit; see the website |
| Langfuse | An open-source alternative, self-hostable | Open source is free; the cloud has a free Hobby plan with a monthly limit |
| PromptLayer | Prompt versioning + analytics, handy for product teams | A free plan with a monthly request limit; see the website |
| Spend limits in the Anthropic Console | A monthly ceiling on API spending | Free |
Spend limit in the Anthropic Console: Settings → Billing → Spend limits → Adjust limit. There's one limit: a monthly ceiling below your tier's ceiling. When it's reached, the API returns an error until you raise the limit (or a new month starts), so set it with some room above your usual bill. Build the $50 and $100 warnings yourself: a script that reads your spending from the Usage page once a day, or a daily report like the one in the section below.
Budget discipline (working rules)
- Daily budget check: check automatically at the start of the day. If you're already at 80% of the monthly limit, pause or escalate to a human
- Per-feature budget: every feature has an explicit budget. "Customer support bot: $50/month max," written in the README and tracked
- Anomaly detection: a spike > 2x the average daily spend gets investigated within the hour
- Quarterly cost review: what grew over the quarter, what can be cut, which features are overspending
Anti-patterns (DON'T)
- ❌ Using Opus for everything "to be safe"
- ❌ Passing the whole history in every request (use summarization)
- ❌ Not using the prompt cache when you have repeating context > 1,024 tokens
- ❌ Real time when the Batch API would do (24h delay vs. 50% savings)
- ❌ Hidden API calls in a loop with no budget check (you can burn $500 overnight)
- ❌ Free tier hopping for production (no SLA, risk of a ban)
- ❌ Optimizing when the bill is $20/month (negative ROI because of the setup time)
- ❌ Not setting a spend limit and warnings (you'll only find out about an $800 bill at the start of the month)
By audience (where to start)
Beginner (bill of $20-50/month, learning curve):
- Technique 1 (right-size models): covers 40-50% of the overspend
- Technique 8 (context discipline): covers another 20-30%
- Total: 50-70% lower costs for 2 hours of work
Intermediate (bill of $50-200/month, production setup):
- Technique 2 (prompt caching): on 40-60% of input cost
- Technique 3 (Batch API): wherever it doesn't need to be real time
- Technique 5 (response cache): on routine tasks
- Total: another 50% off
Professional (bill of $200+/month, scaling):
- All 9 techniques
- Langfuse/LangSmith monitoring
- Budget alerts set up
- Quarterly cost review on the calendar
- Production discipline = predictable costs
The hidden cost of "saving"
Cost engineering is work. Account for it:
| Cost | Time |
|---|---|
| Setting up caching the first time | 4-8 hours |
| Debugging cache invalidation | 2-3 hours when it doesn't work |
| Monitoring setup (Langfuse or LangSmith + alerts) | 2-4 hours |
| Quarterly review and strategy update | 2-3 hours / quarter |
| Migrating to new models (when they come out) | 4-8 hours / release |
The ROI rule: the savings should be > $100/month to pay for the setup time (8h × $50/h = $400 one-time, which pays for itself in 4 months).
Checklist (✅)
After this lesson, you should have:
Tools and resources
- claude.com/pricing: subscription and API plans
- platform.claude.com/docs/.../pricing: the full API docs with pricing
- What's current: our course's summary of current models and prices
- Prompt caching docs: official cache_control documentation
- Batch API docs: bulk processing at 50% off
- Anthropic Console: the usage tab + spend limits
- OpenAI Usage: the same for an OpenAI stack
- Ollama: local models in 5 minutes (Llama, Qwen, Gemma, gpt-oss and others)
- Helicone: LLM observability; in maintenance mode since March 2026
- LangSmith: tracing and cost per trace
- Langfuse: an open-source observability alternative
- PromptLayer: prompt versioning + analytics
Key takeaways
The gap between $20/month and $200/month isn't 10 times more work: it's the difference between unoptimized and optimized code. The same features, the same reliability, 10 times cheaper. It's a matter of discipline, not talent.
3 techniques cover 70% of the savings potential: right-sizing your models (the 80/15/5 rule), prompt caching on repeating context, and context window discipline. The other 6 techniques are for production-grade teams with a bill of $200+/month.
Optimize when the bill is > $200/month or heading that way. Before that, it's a waste of time: setting up caching takes 8 hours and saves $20/month, which is negative ROI. Features first, optimization second.
Next lesson
→ Graduation: wrapping up the first part. For observability, alerting and incident response for LLM applications: Production Observability
The mark stays in this browser only and is never sent anywhere. My progress