The gist
An AI stack in production is like the electrical grid in your house. While it works, you don't notice it. The lights are on, the fridge is cold, the laptop is charging. But once a year there's a blackout.
And that's where houses split into two types:
- A house without a generator: you sit in the dark, the food in the fridge spoils, work stops. You wait for the repair. Hours. Sometimes a whole day.
- A house with a generator: a couple of minutes to switch over, the lights come back, work continues. The neighbors are in the dark, you're working.
An AI stack works the same way. The Anthropic API goes down. OpenAI goes down. Cloudflare goes down. It's not "if," it's "when." The only question is whether you have a plan to switch over, or whether you find out about the problem at the same time as your customers.
This lesson covers 7 typical disaster scenarios, a multi-provider fallback strategy, 3-2-1 backups for AI data, and a 10-minute playbook for when everything breaks.
🎯 Decision tree: what level of preparedness do you need
Personal use (just you):
- ✓ A Git backup of configs and prompts is enough
- ✓ Manual recovery is fine; sitting out a 4-hour outage isn't critical
- ✗ Multi-provider fallback is overkill
SMB (you have paying customers, $1-10K MRR):
- ✓ Multi-provider fallback is a must
- ✓ Automatic nightly backups
- ✓ Basic monitoring (uptime + LLM errors)
- ✓ Communication templates ready to go
Professional ($10-50K MRR, customers depend on you):
- ✓ Full 3-2-1 backup
- ✓ Automated key rotation
- ✓ Quarterly drills
- ✓ A public status page
Enterprise ($50K+ MRR, SLA contracts):
- ✓ Multi-region deployment
- ✓ Automated failover
- ✓ 24/7 monitoring + on-call
- ✓ SOC 2 compliance backups
The default for students of this course is the SMB level. It's enough for most real-world cases.
Key concepts
- RTO (Recovery Time Objective): how long you can afford to be down. For SMB, usually 30 minutes; for enterprise, 5-15 minutes
- RPO (Recovery Point Objective): how much data you can afford to lose. Nightly backup = an RPO of 24 hours; real-time replication = an RPO near zero
- Provider fallback chain: a list of LLM providers in priority order, with automatic switching when one fails
- 3-2-1 backup: 3 copies of your data, 2 different types of storage, 1 copy off-site (a different region/cloud)
- Incident playbook: instructions written in advance for what to do when X breaks. Not "we'll figure it out on the spot"
- Drill: a practice simulation of a disaster to check that the plan really works
- Blast radius: what a specific failure affects. One LLM provider down = a blast radius of "every feature that uses it"
- Status page: a public page with the current state of your service. Customers see it right away and don't write to support
Theory
7 typical disaster scenarios
These aren't theoretical: cases like these happen regularly to providers and developers. For specific dates and details, see the public incident histories (links at the end of the lesson).
Scenario 1: Anthropic API outage
History: Anthropic has both short degraded windows with higher latency and 5xx errors, and longer incidents. See the real history at status.claude.com.
Impact: everything that depends on the Claude API stops working. If your stack has a single provider, your product is down.
Recovery time:
- Without a plan: hours to days (you wait for Anthropic to fix it + switch over manually, if you have somewhere to switch to)
- With a plan: 10 minutes (auto-fallback to OpenAI/Gemini; customers won't notice a thing)
Mitigation: a multi-provider fallback chain + a status page + customer communication.
Scenario 2: OpenAI API outage
History: OpenAI has had major outages when ChatGPT and the API were unavailable for several hours. See the current history at status.openai.com.
Impact: a GPT-based fallback doesn't save you if it's first in the chain. Worse, different providers can go down at the same time because of shared upstream dependencies (for example, the same cloud regions).
Mitigation: don't make OpenAI your only fallback. At least 3 different providers in the chain, in different cloud regions.
Scenario 3: A global Cloudflare/Vercel incident
History:
- Cloudflare, June 21, 2022: a routing configuration error took 19 data centers offline for about an hour and a quarter (Cloudflare's write-up)
- Cloudflare, November 18, 2025: a failure in the bot protection system caused 5xx errors in the CDN, as well as failures in Workers KV, Dashboard and Access; core traffic was restored after about three hours (Cloudflare's write-up)
- Vercel and other platforms have incidents with deploys and edge infrastructure too
Impact: your Worker/Edge Function is down even if the LLM is working fine. Requests never reach your code.
Mitigation: don't put all your eggs in one CDN basket. A backup deployment on another provider (AWS Lambda as a secondary), DNS failover through Cloudflare Health Checks or AWS Route 53.
Scenario 4: Account suspension (terms of service violation)
Typical causes:
- An automated terms-of-service check fires by mistake (a false positive)
- A suspicious billing pattern or a payment dispute
- A violation of the terms of use (sometimes unintentional)
Impact: an instant cut-off. No "you have 30 days." The account becomes unavailable immediately.
Recovery time without a backup account: days to weeks (support tickets, manual review).
Mitigation: a secondary account with every critical provider. Different emails, different cards, different billing addresses if possible. Standby keys already saved and tested. The second account is a spare, not a way to get around a ban for breaking the rules: check the provider's terms.
Scenario 5: Payment failure
Scenarios:
- The card expired → auto-renewal failed → service paused
- The bank's fraud detection blocks a charge → account suspended
- The limit on a company card gets exceeded by a sudden usage spike
Impact: usually 24-48h before customers start seeing downtime. But it's bad if you find out from a customer.
Mitigation:
- Two cards on every account (primary + backup)
- Email alerts on any failed charge
- Spending alerts when usage gets close to the card limit
Scenario 6: Database/KV corruption
Typical causes:
- A storage provider error (KV, vector databases, Postgres)
- A bad migration or a bug in your code
- Data accidentally deleted or overwritten
Impact: customer data is lost, and you need to restore from a backup. If there's no backup, the data is gone forever.
Mitigation: automated nightly backups, point-in-time recovery if the provider supports it, testing the restore procedure quarterly.
Scenario 7: Security breach
Real cases:
- API keys accidentally committed to a public GitHub repo (happens regularly)
- An .env file ended up in a Docker image and got published
- A stolen developer laptop without encryption
Impact: a sharp billing spike within hours (someone else's requests through your API key), reputation damage, potentially a data leak.
Recovery procedure:
- Detect (spending anomaly monitoring or a GitHub secret scanning alert)
- Immediately rotate ALL keys (not just the leaked one, related ones too)
- Audit what was accessed
- Notify customers if data was potentially affected
- Postmortem + prevention
Mitigation: pre-commit hooks (gitleaks, trufflehog), quarterly secret rotation, anomaly detection on billing.
Multi-provider fallback strategy
The basic pattern: a chain of providers with automatic switching.
# Pseudo-code: a provider chain with fallback
providers = [
# model names are an example; for current models and prices, see the course's "What's current" page
{"name": "anthropic", "model": "claude-sonnet-5-5", "priority": 1},
{"name": "openai", "model": "gpt-6.1-sol", "priority": 2},
{"name": "google", "model": "gemini-3.8-flash", "priority": 3},
{"name": "local_ollama", "model": "qwen3:8b", "priority": 4} # last resort
]
def call_llm(prompt, max_retries_per_provider=2):
errors = []
for provider in providers:
try:
response = call(
provider=provider["name"],
model=provider["model"],
prompt=prompt,
timeout=10
)
log_success(provider["name"])
return response
except (Timeout, ServerError, RateLimit) as e:
log_failure(provider["name"], str(e))
errors.append((provider["name"], e))
continue
except AuthenticationError:
# Key revoked: skip without retrying
alert_engineer("API key invalid", provider["name"])
continue
# All providers are down: the worst case
raise AllProvidersDown(errors)Important nuances:
Different model capabilities: models from different providers handle different tasks differently (text, reasoning, multimodal). Fallback may lower the quality, and that's OK, because the customer gets a working product instead of an error.
Prompt portability: your Claude prompts may work poorly on another provider's models (different system prompt format, different reaction to XML tags). Test every prompt on every provider in the chain.
Cost variance: falling back to a more expensive model can cause a billing spike. Set a budget cap on each provider separately.
A local model as the last resort: Ollama with an open model (for example, from the Qwen family) on your Mac/VPS. Lower quality than cloud models, but it works when the whole internet is down. Good for critical paths where "some answer" beats "an error."
3-2-1 backup for AI data
A classic IT principle, adapted for an AI stack:
- 3 copies of the data: production + backup1 + backup2
- 2 different storage types (avoids corruption at a single provider)
- 1 copy off-site (a different region or a different cloud)
What to back up and how:
| Data type | Production | Backup 1 | Backup 2 | Cadence |
|---|---|---|---|---|
| Customer data | Cloudflare D1 | Backblaze B2 (different region) | Local encrypted SSD | Nightly |
| Prompts/configs | Git main branch | GitHub remote | GitLab mirror | On every commit |
| Vector embeddings | Pinecone/Weaviate | Raw text in S3 (can reindex) | — | Daily |
| Audit logs | Append-only KV | S3 Glacier (immutable) | — | Real-time stream |
| API keys | 1Password vault | Cloudflare Secrets | Paper backup in a safe | Quarterly rotation |
A tip on vector embeddings: don't back up embeddings; they're expensive to store and can be recalculated. Back up the raw text + metadata, and in a disaster you'll recalculate the embeddings within an hour.
Managing API keys
Storage rules:
- ❌ Never in code, not even temporarily
- ❌ Never in a .env committed to Git (even if "I'll delete it later")
- ✅ Cloudflare Secrets / Vercel Env / AWS Secrets Manager
- ✅ The 1Password CLI for local development
- ✅ Pre-commit hook scanning (gitleaks)
Rotation schedule:
- Quarterly at minimum for all production keys
- Immediately after: leak detection, an employee leaving, any security incident
- After a major release (if a key could have ended up in a build artifact)
Emergency rotation procedure (documented, tested):
# rotate_all.sh — pseudo-code
# 1. Generate new keys in each provider's dashboard
# 2. Update the secrets in production
wrangler secret put ANTHROPIC_API_KEY # Cloudflare Workers
wrangler secret put OPENAI_API_KEY
# 3. Verify production is using the new ones
curl https://api.yoursite.com/health/llm
# 4. Revoke the old keys in the provider dashboards (manual click)
# 5. Audit log: what was accessed between the leak and the rotation
# Check the provider usage logs for that periodThis script should run in under 10 minutes. It gets tested quarterly.
Multi-account strategy:
- A primary account + a backup account with each critical provider
- Different billing methods (if the primary card gets suspended, the backup works)
- Backup keys already saved in 1Password and tested (not "I'll get to it someday")
Anomaly detection:
- Daily spend > 2× average → alert + investigation
- Spend > monthly budget cap → auto-pause the API key
- Unusual geographic origin (a request from a country where you have no customers) → flag
Incident Response Playbook: 4 steps
When something breaks, you've got adrenaline, panic and a pile of impulses. A playbook gives you structure.
STOP (1-3 minutes)
- Turn off the affected feature (feature flag → off)
- Prevent further damage (if it's data corruption, pause write operations)
- Tell the team "incident in progress" (if you're not alone)
- DON'T start debugging right away: stop the bleeding first
TRIAGE (5-10 minutes)
- What happened? (one specific failure mode, not "something broke")
- When did it start? (a precise timestamp from the logs)
- Scope: which customers are affected? (1 customer / a segment / everyone)
- Blast radius: which features are affected?
- Root cause hypothesis (you may not know for sure yet, but have a direction)
STABILIZE (10-30 minutes)
- Best option: roll back to the previous known-good deploy
- Second: fall back to a secondary provider/feature
- Third: a temporary fix (a quick patch, not the ideal solution)
- Communicate with customers (status page update + email if it's major)
- The team's message: "the product is working, we know what happened, we're fixing it"
POSTMORTEM (1-3 hours, after stabilizing)
- Root cause analysis (5 whys or a fishbone diagram)
- Timeline reconstruction
- What went well / What went badly / What to do differently
- Action items: add monitoring? Prevention? Process improvements?
- A blameless culture: focus on fixing the system, not on "who's to blame"
Communication templates: ready before you go down
In the middle of an incident you don't have time to write a polished email. Prepare the templates in advance.
Customer email during an outage:
Subject: [Service Update] Brief disruption in [Feature] — Status Hi [Customer], We've detected a temporary disruption in [feature/service] that began at [time UTC]. The issue is related to [generic cause: third-party API issue, infrastructure incident]. What we're doing: - [current fix or workaround] - Our engineering team is working on restoring service Estimated recovery: [a conservative time; better to overestimate than miss it] What you can do right now: [a workaround if there is one, or "just wait"] Next update within [30 minutes / 1 hour]. If you need urgent help: [contact]. Sorry for the inconvenience. [Your name]
Status page entry:
[Investigating] LLM API Issues — 2026-02-15 14:23 UTC
We're investigating an elevated error rate in [feature]. Some requests are failing.
Updates to follow.
[Update 14:45] Identified — the root cause is related to an upstream provider outage.
We've activated fallback to a secondary provider. Some users are still seeing errors.
[Monitoring 15:10] Fallback active. The error rate has returned to baseline.
We're monitoring the situation; a postmortem will be published within 24 hours.
[Resolved 16:00] Issue resolved. Postmortem: [link]Internal Slack/Telegram alert:
INCIDENT: [Feature] degraded Severity: [P1 / P2 / P3] Started: [timestamp] Owner: [your name] Status page: [link] Customers affected: ~[number] or [segment] Current action: [stabilizing / investigating / monitoring]
Drill plan: test it quarterly
A plan isn't a plan if it's never been run. Quarterly drill.
Q1: Simulate Anthropic API down
- Block egress traffic to api.anthropic.com in the staging environment
- Verify the OpenAI fallback activates automatically
- Measure: time to switch, customer impact, fallback success rate
- Document gaps → fix → re-test
Q2: Simulate database corruption
- Take a staging database snapshot
- Intentionally corrupt one table
- Practice restoring from the backup
- Verify data integrity after the restore
- Measure: RTO, RPO, data loss if any
Q3: Simulate a key leak
- Pretend ANTHROPIC_API_KEY leaked in Slack
- Run the emergency rotation procedure
- Verify all production systems have switched to the new key
- Verify the old key is revoked in the provider dashboard
- Measure: total rotation time (target <10 min)
Q4: Full disaster simulation
- Production down + LLM provider down + backup account unavailable
- Restore the service on new infrastructure (a new Cloudflare account, new keys)
- Measure: time to restore, % of data preserved, customer communication delivered
A drill costs 4 hours a quarter. It'll save you days when a real incident happens.
Tools for disaster recovery
Not an ad, just practical options. Prices and terms change, so check them on the websites (what you see later may differ from October 2026).
| Category | Tool | Terms | Use case |
|---|---|---|---|
| Backup storage | Backblaze B2 | Paid by volume, price per GB on the site | Object storage, S3-compatible |
| Backup storage | Wasabi | Paid by volume, terms on the site | An alternative to AWS S3 |
| Sync tool | rclone | Free, open source | A CLI for syncing between cloud storage |
| LLM observability | Langfuse | Open source, free; the cloud version has a free Hobby plan (limits on the site) | Logging, monitoring LLM calls |
| Error tracking | Sentry | Has a free plan for small projects | Application errors, stack traces |
| Uptime monitoring | UptimeRobot | Has a free plan | Pings, status checks |
| Uptime monitoring | Pingdom | Paid | More serious monitoring |
| Alerting | Telegram bot or email | Free | Enough for SMB |
| Alerting | PagerDuty | Paid | Enterprise on-call rotation |
| Status pages | Statuspage (Atlassian) | Paid, terms on the site | Hosted, professional |
| Status pages | Cachet | Free (self-hosted) | Open source, your own server |
| Status pages | Instatus | Paid, terms on the site | Modern UI, simple setup |
| Secret scanning | Gitleaks | Free | Pre-commit hook |
| Secret scanning | Trufflehog | Free | Deep repo scans |
The minimum stack for SMB: backup storage + UptimeRobot + Langfuse + Sentry + a status page. Some of these are free; the total depends on the plans you pick.
The cost of a disaster vs the cost of preparedness
A rough calculation for a SaaS with $10K MRR. The numbers are illustrative: they're not statistics or a forecast, plug in your own.
Disaster cost (without preparedness):
| Cost item | Estimate |
|---|---|
| A 4-hour outage = about 0.6% of the month | direct revenue loss ~$60-100 |
| 5-15% customer churn after a bad incident | $500-1500 MRR lost |
| Trust damage, takes 6-12 months to recover | $2000-5000 in lost upsells |
| Engineering time on ad-hoc recovery | 20-40 hours = $1000-2000 |
| Support tickets from confused customers | 30-50 hours = $500-1000 |
| Total for one serious incident | $4000-10000 |
Preparedness cost (annual):
| Cost item | Estimate |
|---|---|
| Multi-provider setup, one-time | 8-16 hours = $400-800 |
| Backup infrastructure | $10-50/month = $120-600/year |
| Monitoring + alerting | $20-100/month = $240-1200/year |
| Drill time 4h × 4 quarters | 16 hours = $800 |
| Total annual | $1560-3400 = $130-280/month |
Bottom line: in this made-up example, $130-280/month of insurance for a $10K MRR business = 1.3-2.8% of revenue, while one real disaster without a plan costs as much as several months of that insurance.
It's not "is it worth it." It's more like "why haven't I done this yet."
Audience breakdown: what you need
Beginner (personal use, hobby projects):
- ✅ Git backup for prompts and configs
- ✅ A pre-commit hook for secrets
- ❌ Multi-provider isn't needed
- ❌ A status page is overkill
- Cost: $0/month
Intermediate (SMB, $1-10K MRR):
- ✅ Multi-provider fallback (at least 2)
- ✅ Nightly backup of customer data
- ✅ Basic monitoring (uptime + LLM errors)
- ✅ Communication templates ready to go
- ✅ 1 simple drill a year
- Cost: $30-50/month
Professional (paying customers, $10-50K MRR):
- ✅ Full 3-2-1 backup
- ✅ A multi-provider chain with 3-4 providers
- ✅ Automated key rotation
- ✅ Quarterly drills
- ✅ A public status page
- ✅ A postmortem culture
- Cost: $100-200/month
Enterprise ($50K+ MRR, SLA contracts):
- ✅ Multi-region deployment
- ✅ Automated failover
- ✅ 24/7 on-call rotation
- ✅ SOC 2 compliance backups
- ✅ Dedicated incident response training
- Cost: $500+/month
Anti-patterns
❌ A single LLM provider in production: Anthropic-only or OpenAI-only. When it goes down, you go down. In 2026, multi-provider is basic hygiene, not a nice-to-have.
❌ Backups you never test restoring: a classic. The backup exists, but nobody's ever tried to restore it. On D-day it turns out the backup is corrupt or the procedure is broken.
❌ No communication with customers during an outage: silence is worse than bad news. The customer writes to support → gets an auto-reply → sees the product is down → doesn't know what's going on → loses trust.
❌ API keys in a .env committed to Git history: even if you later deleted the commit, it's in Git history forever. If it was ever committed, consider it leaked and rotate immediately.
❌ Manual rollback without a documented procedure: "I remember how to do it" works when you're calm. In an incident, with adrenaline at 2 a.m., you'll forget a step. Documentation = a checklist.
❌ "That rarely happens": until it happens once and you lose an enterprise deal. Probability × impact, not probability × wishful thinking.
❌ Backups at the same provider as production: Cloudflare KV primary + Cloudflare R2 backup. Cloudflare goes down → both are unavailable. The backup needs to be with an independent provider.
❌ Hiding an incident from customers ("we'll fix it now, nobody will notice"): they'll notice. And when they find out you hid it, the trust damage is 10× worse than from honest disclosure.
❌ Relying on a provider's SLA as a guarantee: an SLA gives you a refund (often proportional to the downtime); it doesn't prevent downtime. A 99.9% SLA = about 8.8 hours of permitted downtime a year (0.1% × 8,760 hours).
Readiness checklist
✅ Multi-provider fallback works (tested last quarter) ✅ Backups run automatically every night ✅ The last restore test passed < 90 days ago ✅ Communication templates written for the top 3 incident types ✅ Status page set up and tested ✅ Emergency rotation procedure documented ✅ All API keys in a secrets manager, zero in code ✅ A pre-commit hook scans for secrets ✅ Monitoring alerts on an LLM error rate spike ✅ Spending anomaly detection is active ✅ A drill is scheduled for next quarter ✅ A postmortem template is ready
If you have ≤ 6 ✅, you're in the risk zone. ≤ 9 ✅ is standard SMB readiness. 12/12 is the professional level.
Practice
Step 1: A multi-provider fallback in Python
# llm_router.py — a simple fallback router
import os
import time
from typing import Optional
import anthropic
import openai
from google import genai
class LLMRouter:
def __init__(self):
self.anthropic_client = anthropic.Anthropic(
api_key=os.getenv("ANTHROPIC_API_KEY")
)
self.openai_client = openai.OpenAI(
api_key=os.getenv("OPENAI_API_KEY")
)
self.gemini_client = genai.Client(
api_key=os.getenv("GOOGLE_API_KEY")
)
self.providers = [
# model names are an example; see the "What's current" page for current ones
("anthropic", "claude-sonnet-5-5"),
("openai", "gpt-6.1-sol"),
("gemini", "gemini-3.8-flash"),
]
def call(self, prompt: str, max_tokens: int = 1000) -> dict:
errors = []
for provider_name, model in self.providers:
try:
start = time.time()
response = self._call_provider(
provider_name, model, prompt, max_tokens
)
latency = time.time() - start
return {
"provider": provider_name,
"model": model,
"text": response,
"latency_ms": int(latency * 1000),
"fallback_used": provider_name != "anthropic"
}
except Exception as e:
errors.append({
"provider": provider_name,
"error": str(e),
"type": type(e).__name__
})
# Log the failure for monitoring
print(f"[FAIL] {provider_name}: {e}")
continue
raise Exception(f"All providers failed: {errors}")
def _call_provider(self, name, model, prompt, max_tokens):
if name == "anthropic":
r = self.anthropic_client.messages.create(
model=model,
max_tokens=max_tokens,
messages=[{"role": "user", "content": prompt}],
timeout=10
)
return "".join(b.text for b in r.content if b.type == "text")
elif name == "openai":
r = self.openai_client.chat.completions.create(
model=model,
max_completion_tokens=max_tokens, # GPT-5 and newer models use this instead of max_tokens
messages=[{"role": "user", "content": prompt}],
timeout=10
)
return r.choices[0].message.content
elif name == "gemini":
r = self.gemini_client.models.generate_content(
model=model,
contents=prompt
)
return r.text
# Usage
router = LLMRouter()
result = router.call("Explain what disaster recovery is in 3 sentences")
print(f"Provider used: {result['provider']}")
print(f"Fallback used: {result['fallback_used']}")
print(result['text'])Step 2: A backup script for prompts and configs
#!/bin/bash
# backup_ai_stack.sh — nightly backup of the AI infrastructure
set -e
BACKUP_DIR="/backups/$(date +%Y-%m-%d)"
B2_BUCKET="my-ai-backup"
mkdir -p "$BACKUP_DIR"
# 1. Back up prompts/configs (Git already covers this, but an extra copy)
tar czf "$BACKUP_DIR/prompts.tar.gz" ./prompts ./.claude
# 2. Back up customer data from Cloudflare D1
wrangler d1 export my-database --remote --output="$BACKUP_DIR/db.sql"
# 3. Back up KV storage: first the list of keys, then the values
# (see the Wrangler documentation for the file format kv bulk get expects)
wrangler kv key list --namespace-id=$KV_ID --remote > "$BACKUP_DIR/kv-keys.json"
wrangler kv bulk get "$BACKUP_DIR/kv-keys.json" --namespace-id=$KV_ID --remote > "$BACKUP_DIR/kv.json"
# 4. Upload to Backblaze B2 (off-site)
rclone copy "$BACKUP_DIR" "b2:$B2_BUCKET/$(date +%Y-%m-%d)"
# 5. Retention: delete local backups older than 30 days
find /backups -type d -mtime +30 -exec rm -rf {} +
# 6. Verify backup integrity
SIZE=$(du -sh "$BACKUP_DIR" | cut -f1)
echo "Backup completed: $BACKUP_DIR ($SIZE)"
# 7. Notification on success/failure
if [ $? -eq 0 ]; then
echo "Backup OK $(date)" >> /var/log/ai-backup.log
else
echo "Backup FAILED $(date)" | mail -s "ALERT: Backup failed" admin-notifications
fiRun it with cron: 0 3 * * * /scripts/backup_ai_stack.sh
Step 3: A quick incident response checklist (print it and keep it handy)
# INCIDENT RESPONSE — 10 minutes to recovery
## STOP (minutes 0-3)
[ ] Feature flag → off for the affected feature
[ ] Notify the team in Slack: "INCIDENT in progress, owner: [me]"
[ ] Open a status page draft
## TRIAGE (minutes 3-10)
[ ] What broke? (specific failure mode)
[ ] When did it start? (timestamp from the logs)
[ ] Who's affected? (segment / count)
[ ] Severity: P1 / P2 / P3
[ ] Root cause hypothesis (you may not know for sure yet)
## STABILIZE (minutes 10-30)
[ ] Option A: Roll back to the previous good deploy
[ ] Option B: Activate fallback (provider, region)
[ ] Option C: Temporary fix (a quick patch)
[ ] Update the status page: "Investigating" → "Identified" → "Monitoring"
[ ] Send a customer email if P1
## POSTMORTEM (after stabilizing, within 48h)
[ ] Reconstruct the timeline
[ ] 5 whys analysis
[ ] What went well / badly / what to change
[ ] Action items with an owner and a deadline
[ ] Update the playbook if you found a gapStep 4: Test your fallback (a drill)
# Simulate the Anthropic API being down: block egress
# In local development, via /etc/hosts:
echo "127.0.0.1 api.anthropic.com" | sudo tee -a /etc/hosts
# Start your app
python app.py
# Make a request: it should fall back to OpenAI/Gemini
curl -X POST http://localhost:8000/api/generate \
-H "Content-Type: application/json" \
-d '{"prompt": "Test fallback"}'
# Check:
# - Did you get a response? (success criteria)
# - Which provider was used? (should not be anthropic)
# - Is the latency acceptable? (target < 30 sec total)
# - Did the log record the failure + fallback? (audit trail)
# Revert /etc/hosts:
sudo sed -i '' '/api.anthropic.com/d' /etc/hostsIf this test doesn't work in your dev environment, it definitely won't work during a production incident.
Tools and resources
- Claude Status: the official status page for Claude and the Anthropic API
- OpenAI Status: the status of OpenAI services
- Cloudflare Status: Cloudflare's global status
- AWS Well-Architected — DR Objectives: the RTO/RPO framework from AWS
- 3-2-1 Backup Strategy: Backblaze's explanation of the principle
- Backblaze B2: S3-compatible cloud storage
- Statuspage.io: hosted status pages
- Gitleaks: open-source secret scanning
- Langfuse: LLM observability, open source, has a free cloud plan
- UptimeRobot: uptime monitoring, free tier
Key takeaways
It's not "if" your LLM provider goes down, it's "when." Major providers (Anthropic, OpenAI and others) have serious outages: see the history on their status pages. A single-provider stack = just a matter of time until your first downtime.
3-2-1 backup isn't paranoia, it's hygiene. 3 copies, 2 storage types, 1 off-site. In the made-up example above, that's on the order of tens of dollars a month for an SMB, while one real disaster without a backup costs thousands in lost revenue, churn and repair work.
Playbook + quarterly drill > improvised heroics. A plan written in advance and tested 4 times a year turns a 4-hour incident into a 10-minute switch. Pilots train for engine failure not because it happens often, but so they can handle it when it does.
Next lesson
→ Telegram bots with the Claude API. On cutting the costs of your AI stack: Cost Engineering
The mark stays in this browser only and is never sent anywhere. My progress