The gist
Two ways to save money at scale: Prompt Caching is like a gym membership (you pay once to get in, then go many times), and the Batch API is like a wholesale order from a factory (cheaper per unit, but you wait up to 24 hours for delivery). Each one on its own gives you a 50-90% discount. Combined, cached input tokens can cost as little as 5% of the base price (and cache reads on Opus 5.5 and Fable 5.1 are even cheaper). Current prices and versions: What's current.
Key concepts
- Prompt Caching: caching the repeating parts of your prompt on Anthropic's servers
- cache_control: the
{"type": "ephemeral"}marker that says what to cache - TTL: how long the cache lives: 5 minutes (standard,
"5m") or 1 hour ("1h", more expensive to write) - Cache pricing: a 5m write costs 1.25x the base price, a 1h write costs 2x, and a read costs 0.1x (90% savings; on Opus 5.5 a read is 0.05x, and on Fable 5.1 it's cheaper still)
- Batch API: send up to 100,000 requests (or 256 MB) at once with a 50% discount
- Batch statuses:
in_progress→ended(most batches finish in under an hour, 24 hours at most)
Theory
Part 1: Prompt Caching
How caching works
Every time you send a request to Claude, you pay for all the tokens: the system prompt, the context, the examples, the user's message. If your system prompt is 3,000 tokens and you make 1,000 requests a day, that's 3 million tokens just for the repeating context.
Prompt Caching saves the prompt on Anthropic's servers. There are two TTLs (how long the cache lives):
| TTL | Write cost | Read cost | When to use it |
|---|---|---|---|
5 minutes ("5m", default) |
1.25x the base price (+25%) | 0.1x the base price (−90%) | Frequent requests, chatbots, real-time APIs |
1 hour ("1h") |
2x the base price (+100%) | 0.1x the base price (−90%) | Batch processing, long tasks with extended thinking (>5 min), infrequent requests |
First request: you pay to write to the cache (1.25x for 5m or 2x for 1h) Requests 2-N (within the TTL): you pay about 10% to read from the cache (5% on Opus 5.5, even less on Fable 5.1) + the full price for uncached tokens
How a cached request is structured
Option 1: Automatic caching (top-level cache_control)
The simplest way: add cache_control at the request level, and the API figures out what to cache on its own:
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=2048,
cache_control={"type": "ephemeral"}, # ← automatic caching
system="You are a specialized assistant for analyzing contracts...",
messages=[{"role": "user", "content": "Analyze this contract: [text]"}]
)Option 2: Explicit breakpoints (precise control)
For precise control, put cache_control on specific content blocks:
# System prompt: 2000+ tokens (long, repeated)
SYSTEM_PROMPT = """
You are a specialized assistant for analyzing commercial lease agreements.
Respond in English only.
ANALYSIS RULES:
1. Always check the lease term, start date and end date
2. Highlight the early termination terms
3. Note any penalties and late fees
4. Check whether there's a rent escalation clause
5. Note each party's responsibility for repairs
...another 1500 tokens of rules and examples...
"""
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=2048,
system=[
{
"type": "text",
"text": SYSTEM_PROMPT,
"cache_control": {"type": "ephemeral"} # ← marker on a specific block
}
],
messages=[
{
"role": "user",
"content": "Analyze this contract: [contract text]"
}
]
)
# Check that the cache is working
usage = response.usage
print(f"Input tokens: {usage.input_tokens}")
print(f"Tokens written to cache: {usage.cache_creation_input_tokens}")
print(f"Tokens read from cache: {usage.cache_read_input_tokens}")Option 3: 1-hour TTL (for batch jobs and long tasks with extended thinking)
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=2048,
cache_control={
"type": "ephemeral",
"ttl": "1h" # ← 1 hour instead of 5 minutes (writing costs more, reading costs the same)
},
system="A long system prompt...",
messages=[{"role": "user", "content": "Request..."}]
)What's worth caching
| Cache it | Don't cache it |
|---|---|
| System prompt | The user's request |
| Few-shot examples (5-10 of them) | Users' personal data |
| Long instructions | Dynamic data (time, IDs) |
| Knowledge base (RAG context) | Short parts that change |
| Legal/technical rules | Variable parts of a template |
Minimum size for caching (depends on the model; as of October 2026):
| Models | Minimum tokens |
|---|---|
| Fable 5.1, Opus 5.5, Sonnet 5.5 | 512 tokens |
| Sonnet 5, Sonnet 4.6, Sonnet 4.5, Opus 4.8 | 1,024 tokens |
| Opus 4.7 | 2,048 tokens |
| Haiku 4.5, Opus 4.6, Opus 4.5 | 4,096 tokens |
Below the minimum, the cache silently isn't created (no error). Check: if cache_creation_input_tokens and cache_read_input_tokens are both 0, caching didn't work.
The caching price model (as of October 2026)
Sonnet 5.5 (a typical choice, $2/MTok input):
| Token type | 5 min TTL | 1 hour TTL |
|---|---|---|
| Regular input tokens | $2 / 1M | $2 / 1M |
| Cache write | $2.50 / 1M (+25%) | $4 / 1M (+100%) |
| Cache read | $0.20 / 1M (−90%) | $0.20 / 1M (−90%) |
| Output tokens | $10 / 1M | $10 / 1M |
All current models (as of October 2026):
| Model | Base input | 5m write | 1h write | Cache read |
|---|---|---|---|---|
| Fable 5.1 | $10/MTok | $12.50/MTok | $20/MTok | see the pricing page |
| Opus 5.5 | $4/MTok | $5/MTok | $8/MTok | $0.20/MTok |
| Sonnet 5.5 | $2/MTok | $2.50/MTok | $4/MTok | $0.20/MTok |
| Haiku 4.5 | $1/MTok | $1.25/MTok | $2/MTok | $0.10/MTok |
The formula: 5m write = 1.25x base, 1h write = 2x base, read = 0.1x base (0.05x on Opus 5.5, and less than that on Fable 5.1). Prices change from version to version, but the principle stays the same: What's current.
On the first request you pay a little more to create the cache. Starting with the second request within the TTL, you save about 90% on cached tokens.
A savings example
Scenario (Sonnet 5.5 prices as of October 2026): 1,000 requests a day, a 3,000-token system prompt, a 500-token response. Requests come often enough that the cache gets refreshed every 5 minutes.
Without caching: Input: 1000 × 3000 = 3,000,000 tokens × $2/1M = $6.00/day Output: 1000 × 500 = 500,000 tokens × $10/1M = $5.00/day Total: $11.00/day With caching (5-minute TTL, one session): Cache write (once): 3000 × $2.50/1M = $0.0075 Cache reads (999 times): 999 × 3000 × $0.20/1M = $0.599 Output (1000 times): $5.00/day Total: $5.61/day Savings: ~49%
The example shows the calculation method. Plug in your own numbers from the What's current page: the savings depend on how much of your total cost comes from repeated input.
Multiple cache points (breakpoints)
You can cache several blocks in one request. The maximum is 4 breakpoints (explicit cache_control):
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1024,
system=[
{
"type": "text",
"text": BASE_RULES, # Base rules (always)
"cache_control": {"type": "ephemeral"} # breakpoint 1
},
{
"type": "text",
"text": DOMAIN_KNOWLEDGE, # Domain knowledge (changes sometimes)
"cache_control": {"type": "ephemeral"} # breakpoint 2
}
],
messages=[{
"role": "user",
"content": user_question
}]
)The order the cache is checked in: the API checks the cache in this order: tools → system → messages. Each breakpoint caches everything up to and including it (cumulatively).
What can and can't be cached
| Can be cached | Can't be cached |
|---|---|
Tool definitions (the tools array) |
Empty text blocks |
| System messages | Thinking blocks with explicit cache_control |
| Text messages (user and assistant) | Sub-content (citations inside documents) |
| Images and documents in user messages | |
| Tool use / tool result blocks |
What invalidates the cache
Changing any of these "breaks" the cache, and the next request creates a new one:
- Tool definitions
- Thinking and effort parameters (on some models)
- Switching tool_choice
- Changing images in the prompt
Part 2: The Batch API
When you need the Batch API
The Batch API is for tasks that don't need an answer right this second. You process a whole stack of requests, get the results a few hours later, and pay half as much.
| Regular API | Batch API |
|---|---|
| Answer in 1-5 seconds | Most < 1 hour, max 24 hours |
| Full price | 50% off everything |
| One request | Up to 100,000 requests (or 256 MB) |
| Synchronous | Asynchronous |
Ideal use cases:
- Classifying 5,000 customer reviews
- Generating descriptions for 2,000 products
- Analyzing 1,000 résumés
- Translating 3,000 articles
- SEO optimization for 500 pages
- Large-scale evaluations (thousands of test cases)
What you can send in a batch: any Messages API request: vision, tool use, system messages, multi-turn, extended thinking, any beta features. The exception: fast mode isn't available in batches. Each request is processed independently, so you can mix different types in one batch.
Tip: for batches with a shared system prompt, use the 1-hour cache ("ttl": "1h"): a batch usually takes longer than 5 minutes, and a 5-minute cache will expire.
Sending a batch
import anthropic
client = anthropic.Anthropic()
# Preparing the requests. Haiku 4.5 may be retired from the API no earlier than 10/15/2026:
# before you run this, check the model deprecations page and use the current cheap model
requests = []
products = load_products_from_db() # your 1000 products
for i, product in enumerate(products):
requests.append({
"custom_id": f"product-{product['id']}", # your ID for matching
"params": {
"model": "claude-haiku-4-5-20251001", # Haiku for batches: cheaper
"max_tokens": 500,
"messages": [{
"role": "user",
"content": f"""Write an SEO description for this product:
Name: {product['name']}
Category: {product['category']}
Specs: {product['specs']}
The description should be 100-150 words and include keywords."""
}]
}
})
# Send the batch
batch = client.messages.batches.create(requests=requests)
print(f"Batch created: {batch.id}")
print(f"Status: {batch.processing_status}") # in_progress
print(f"Requests in batch: {batch.request_counts.processing}")Checking the status and getting the results
import time
batch_id = batch.id
# Wait for it to finish (polling)
while True:
batch_status = client.messages.batches.retrieve(batch_id)
if batch_status.processing_status == "ended":
print("Batch finished!")
print(f"Succeeded: {batch_status.request_counts.succeeded}")
print(f"Errors: {batch_status.request_counts.errored}")
break
print(f"Processing: {batch_status.request_counts.processing} requests...")
time.sleep(60) # check once a minute
# Get the results
results = {}
for result in client.messages.batches.results(batch_id):
if result.result.type == "succeeded":
results[result.custom_id] = "".join(b.text for b in result.result.message.content if b.type == "text")
else:
results[result.custom_id] = None
print(f"Error for {result.custom_id}: {result.result.error}")
# Save to the database
save_descriptions_to_db(results)Batch statuses
in_progress → requests are being processed
ending → wrapping up (some are still running)
ended → all done, results available
For each request:
succeeded → OK, there's a result
errored → error (rate limit, invalid request)
expired → the request wasn't processed within 24 hours
canceled → the batch was canceled manuallyImportant: batch results are available for 29 days after creation. After that you can still see the batch itself, but you can't download the results.
Canceling a batch
# Changed your mind? Cancel while there's still time
client.messages.batches.cancel(batch_id)Canceled and unprocessed requests aren't billed.
Batch API pricing
Everything is 50% of the standard prices, both input and output:
| Model (as of October 2026) | Batch input | Batch output |
|---|---|---|
| Fable 5.1 | $5/MTok | $25/MTok |
| Opus 5.5 | $2/MTok | $10/MTok |
| Sonnet 5.5 | $1/MTok | $5/MTok |
| Haiku 4.5 | $0.50/MTok | $2.50/MTok |
Batch API + the output-300k-2026-03-24 beta header: up to 300,000 output tokens per request for Opus 5.5, Sonnet 5.5 and several earlier models (the usual limit for a synchronous request is 128k on Fable 5.1, Opus 5.5 and Sonnet 5.5, and 64k on Haiku 4.5). A single answer like that can take more than an hour to generate, so plan for the full 24-hour window.
The combo: Batch + Caching = maximum savings
# Shared system prompt, cached with a 1-hour TTL (the batch takes > 5 min!)
ANALYSIS_SYSTEM = """[4500+ tokens of analysis rules: for Haiku 4.5 the cache minimum is 4096 tokens]"""
requests = []
for doc in documents: # 5000 documents
requests.append({
"custom_id": f"doc-{doc['id']}",
"params": {
"model": "claude-haiku-4-5-20251001",
"max_tokens": 300,
"system": [
{
"type": "text",
"text": ANALYSIS_SYSTEM,
"cache_control": {
"type": "ephemeral",
"ttl": "1h" # ← 1 hour! The batch takes longer than 5 minutes
}
}
],
"messages": [{"role": "user", "content": doc['text']}]
}
})
batch = client.messages.batches.create(requests=requests)Why "1h" and not "5m"? A batch is processed asynchronously. If it takes 20 minutes, a 5-minute cache expires after the first few requests, and the remaining 4,500 documents pay full price. The 1-hour cache costs more to write (2x), but less overall.
An illustration using a hypothetical $100 (actual savings depend on how much of your input repeats):
| Method | Discount | Final price |
|---|---|---|
| Regular API | 0% | $100 |
| Batch only | -50% | $50 |
| Caching only | -45% (average) | $55 |
| Batch + Caching | -90% to -95% | $5-10 |
Practice
Assignment: Batch processing with caching
Prepare a list of 10 short texts (reviews, descriptions, anything):
python texts = [ "Great service, I recommend it to everyone!", "Delivery was 3 days late, not great.", # ... 8 more texts ]Create a batch for sentiment classification (positive/negative/neutral):
python SENTIMENT_PROMPT = """ Classify the sentiment of the text. Answer with a single word: positive, negative or neutral. Don't add any explanation. """ # ~50 tokens: below the minimum, so add more rules and examplesAdd
cache_controlto the system prompt (expand it with examples: Sonnet 5.5 needs at least 512 tokens, Haiku 4.5 at least 4,096)Send the batch and start a polling loop to check the status
When it's done: print the
custom_id+ result for each textCheck
usagein the responses: do you seecache_read_input_tokens?
Goal: go through the full Batch API cycle and see the savings in real numbers.
Additional paid API features (as of October 2026)
Besides model tokens, the Claude API has some separately billed services:
| Feature | Price | What it does |
|---|---|---|
| Web Search | $10 / 1,000 searches + tokens | Claude searches the internet while answering |
| Web Fetch | Free (tokens only) | Claude reads a URL you give it |
| Code Execution | billed per container-hour, with a free monthly allowance (see the pricing page) | Runs Python inside the response |
| Code Execution + Web | Free | When used together with Web Search/Fetch |
| Managed Agents | billed per session-hour + tokens (see the pricing page) | Anthropic-hosted agents (you pay only for running time) |
| US-only data residency | a surcharge on all tokens (see the pricing page) | A legal requirement to keep data in the US |
Fast mode (research preview) for Opus 5.5: noticeably faster, but twice as expensive: $8/MTok input and $40/MTok output versus $4 and $20 in regular mode. It doesn't work with the Batch API. Use it when speed matters more than cost.
Tools and resources
- Prompt Caching docs: platform.claude.com/docs: prompt caching
- Batch API docs: platform.claude.com/docs: batch processing
- Pricing: platform.claude.com/docs: pricing
- API Console: platform.claude.com
- Python SDK:
pip install anthropic(the Batch API is included in the SDK) - Batch API limits: up to 100,000 requests or 256 MB; results are kept for 29 days
- The
output-300k-2026-03-24beta header: up to 300k output tokens in Batch for Opus 5.5, Sonnet 5.5 and several earlier models - Current prices and versions: What's current
Key takeaways
Prompt Caching pays for itself on the second request. Two TTLs: 5 minutes (default, write +25%) and 1 hour (write +100%). A read is usually 0.1x (−90%), and even cheaper on Opus 5.5 and Fable 5.1. The cache minimum depends on the model: from 512 to 4,096 tokens. If the prompt is shorter, the cache silently isn't created. The Batch API: up to 100,000 requests at once, 50% off. Most batches finish in under an hour. For batches, use the 1-hour cache (
"ttl": "1h"): a 5-minute cache will expire before the batch finishes. Combine both methods for large-scale jobs: savings of 90-95% off the base price.
Next lesson
The mark stays in this browser only and is never sent anywhere. My progress