Library · Reliability: monitoring, failures, backups

Backups and recovery for your AI stack: what to do when the Anthropic API is down

Engineer70 minUpdated: October 2026
91 of 105 in the library

Time: about 30 min reading + 40 min practice


The gist

An AI stack in production is like the electrical grid in your house. While it works, you don't notice it. The lights are on, the fridge is cold, the laptop is charging. But once a year there's a blackout.

And that's where houses split into two types:

  • A house without a generator: you sit in the dark, the food in the fridge spoils, work stops. You wait for the repair. Hours. Sometimes a whole day.
  • A house with a generator: a couple of minutes to switch over, the lights come back, work continues. The neighbors are in the dark, you're working.

An AI stack works the same way. The Anthropic API goes down. OpenAI goes down. Cloudflare goes down. It's not "if," it's "when." The only question is whether you have a plan to switch over, or whether you find out about the problem at the same time as your customers.

This lesson covers 7 typical disaster scenarios, a multi-provider fallback strategy, 3-2-1 backups for AI data, and a 10-minute playbook for when everything breaks.

🎨 Picture this: an airline pilot trains for engine failure in flight. Not because it happens often, but because when it does happen, the pilot needs muscle memory, not panic. A quarterly drill = your engine failure training.


🎯 Decision tree: what level of preparedness do you need

Personal use (just you):

  • ✓ A Git backup of configs and prompts is enough
  • ✓ Manual recovery is fine; sitting out a 4-hour outage isn't critical
  • ✗ Multi-provider fallback is overkill

SMB (you have paying customers, $1-10K MRR):

  • ✓ Multi-provider fallback is a must
  • ✓ Automatic nightly backups
  • ✓ Basic monitoring (uptime + LLM errors)
  • ✓ Communication templates ready to go

Professional ($10-50K MRR, customers depend on you):

  • ✓ Full 3-2-1 backup
  • ✓ Automated key rotation
  • ✓ Quarterly drills
  • ✓ A public status page

Enterprise ($50K+ MRR, SLA contracts):

  • ✓ Multi-region deployment
  • ✓ Automated failover
  • ✓ 24/7 monitoring + on-call
  • ✓ SOC 2 compliance backups

The default for students of this course is the SMB level. It's enough for most real-world cases.


Key concepts

  • RTO (Recovery Time Objective): how long you can afford to be down. For SMB, usually 30 minutes; for enterprise, 5-15 minutes
  • RPO (Recovery Point Objective): how much data you can afford to lose. Nightly backup = an RPO of 24 hours; real-time replication = an RPO near zero
  • Provider fallback chain: a list of LLM providers in priority order, with automatic switching when one fails
  • 3-2-1 backup: 3 copies of your data, 2 different types of storage, 1 copy off-site (a different region/cloud)
  • Incident playbook: instructions written in advance for what to do when X breaks. Not "we'll figure it out on the spot"
  • Drill: a practice simulation of a disaster to check that the plan really works
  • Blast radius: what a specific failure affects. One LLM provider down = a blast radius of "every feature that uses it"
  • Status page: a public page with the current state of your service. Customers see it right away and don't write to support

Theory

7 typical disaster scenarios

These aren't theoretical: cases like these happen regularly to providers and developers. For specific dates and details, see the public incident histories (links at the end of the lesson).


Scenario 1: Anthropic API outage

History: Anthropic has both short degraded windows with higher latency and 5xx errors, and longer incidents. See the real history at status.claude.com.

Impact: everything that depends on the Claude API stops working. If your stack has a single provider, your product is down.

Recovery time:

  • Without a plan: hours to days (you wait for Anthropic to fix it + switch over manually, if you have somewhere to switch to)
  • With a plan: 10 minutes (auto-fallback to OpenAI/Gemini; customers won't notice a thing)

Mitigation: a multi-provider fallback chain + a status page + customer communication.


Scenario 2: OpenAI API outage

History: OpenAI has had major outages when ChatGPT and the API were unavailable for several hours. See the current history at status.openai.com.

Impact: a GPT-based fallback doesn't save you if it's first in the chain. Worse, different providers can go down at the same time because of shared upstream dependencies (for example, the same cloud regions).

Mitigation: don't make OpenAI your only fallback. At least 3 different providers in the chain, in different cloud regions.


Scenario 3: A global Cloudflare/Vercel incident

History:

  • Cloudflare, June 21, 2022: a routing configuration error took 19 data centers offline for about an hour and a quarter (Cloudflare's write-up)
  • Cloudflare, November 18, 2025: a failure in the bot protection system caused 5xx errors in the CDN, as well as failures in Workers KV, Dashboard and Access; core traffic was restored after about three hours (Cloudflare's write-up)
  • Vercel and other platforms have incidents with deploys and edge infrastructure too

Impact: your Worker/Edge Function is down even if the LLM is working fine. Requests never reach your code.

Mitigation: don't put all your eggs in one CDN basket. A backup deployment on another provider (AWS Lambda as a secondary), DNS failover through Cloudflare Health Checks or AWS Route 53.


Scenario 4: Account suspension (terms of service violation)

Typical causes:

  • An automated terms-of-service check fires by mistake (a false positive)
  • A suspicious billing pattern or a payment dispute
  • A violation of the terms of use (sometimes unintentional)

Impact: an instant cut-off. No "you have 30 days." The account becomes unavailable immediately.

Recovery time without a backup account: days to weeks (support tickets, manual review).

Mitigation: a secondary account with every critical provider. Different emails, different cards, different billing addresses if possible. Standby keys already saved and tested. The second account is a spare, not a way to get around a ban for breaking the rules: check the provider's terms.


Scenario 5: Payment failure

Scenarios:

  • The card expired → auto-renewal failed → service paused
  • The bank's fraud detection blocks a charge → account suspended
  • The limit on a company card gets exceeded by a sudden usage spike

Impact: usually 24-48h before customers start seeing downtime. But it's bad if you find out from a customer.

Mitigation:

  • Two cards on every account (primary + backup)
  • Email alerts on any failed charge
  • Spending alerts when usage gets close to the card limit

Scenario 6: Database/KV corruption

Typical causes:

  • A storage provider error (KV, vector databases, Postgres)
  • A bad migration or a bug in your code
  • Data accidentally deleted or overwritten

Impact: customer data is lost, and you need to restore from a backup. If there's no backup, the data is gone forever.

Mitigation: automated nightly backups, point-in-time recovery if the provider supports it, testing the restore procedure quarterly.


Scenario 7: Security breach

Real cases:

  • API keys accidentally committed to a public GitHub repo (happens regularly)
  • An .env file ended up in a Docker image and got published
  • A stolen developer laptop without encryption

Impact: a sharp billing spike within hours (someone else's requests through your API key), reputation damage, potentially a data leak.

Recovery procedure:

  1. Detect (spending anomaly monitoring or a GitHub secret scanning alert)
  2. Immediately rotate ALL keys (not just the leaked one, related ones too)
  3. Audit what was accessed
  4. Notify customers if data was potentially affected
  5. Postmortem + prevention

Mitigation: pre-commit hooks (gitleaks, trufflehog), quarterly secret rotation, anomaly detection on billing.


Multi-provider fallback strategy

The basic pattern: a chain of providers with automatic switching.

python
# Pseudo-code: a provider chain with fallback
providers = [
    # model names are an example; for current models and prices, see the course's "What's current" page
    {"name": "anthropic", "model": "claude-sonnet-5-5", "priority": 1},
    {"name": "openai", "model": "gpt-6.1-sol", "priority": 2},
    {"name": "google", "model": "gemini-3.8-flash", "priority": 3},
    {"name": "local_ollama", "model": "qwen3:8b", "priority": 4}  # last resort
]

def call_llm(prompt, max_retries_per_provider=2):
    errors = []
    for provider in providers:
        try:
            response = call(
                provider=provider["name"],
                model=provider["model"],
                prompt=prompt,
                timeout=10
            )
            log_success(provider["name"])
            return response
        except (Timeout, ServerError, RateLimit) as e:
            log_failure(provider["name"], str(e))
            errors.append((provider["name"], e))
            continue
        except AuthenticationError:
            # Key revoked: skip without retrying
            alert_engineer("API key invalid", provider["name"])
            continue

    # All providers are down: the worst case
    raise AllProvidersDown(errors)

Important nuances:

  1. Different model capabilities: models from different providers handle different tasks differently (text, reasoning, multimodal). Fallback may lower the quality, and that's OK, because the customer gets a working product instead of an error.

  2. Prompt portability: your Claude prompts may work poorly on another provider's models (different system prompt format, different reaction to XML tags). Test every prompt on every provider in the chain.

  3. Cost variance: falling back to a more expensive model can cause a billing spike. Set a budget cap on each provider separately.

  4. A local model as the last resort: Ollama with an open model (for example, from the Qwen family) on your Mac/VPS. Lower quality than cloud models, but it works when the whole internet is down. Good for critical paths where "some answer" beats "an error."


3-2-1 backup for AI data

A classic IT principle, adapted for an AI stack:

  • 3 copies of the data: production + backup1 + backup2
  • 2 different storage types (avoids corruption at a single provider)
  • 1 copy off-site (a different region or a different cloud)

What to back up and how:

Data type Production Backup 1 Backup 2 Cadence
Customer data Cloudflare D1 Backblaze B2 (different region) Local encrypted SSD Nightly
Prompts/configs Git main branch GitHub remote GitLab mirror On every commit
Vector embeddings Pinecone/Weaviate Raw text in S3 (can reindex) — Daily
Audit logs Append-only KV S3 Glacier (immutable) — Real-time stream
API keys 1Password vault Cloudflare Secrets Paper backup in a safe Quarterly rotation

A tip on vector embeddings: don't back up embeddings; they're expensive to store and can be recalculated. Back up the raw text + metadata, and in a disaster you'll recalculate the embeddings within an hour.


Managing API keys

Storage rules:

  • ❌ Never in code, not even temporarily
  • ❌ Never in a .env committed to Git (even if "I'll delete it later")
  • ✅ Cloudflare Secrets / Vercel Env / AWS Secrets Manager
  • ✅ The 1Password CLI for local development
  • ✅ Pre-commit hook scanning (gitleaks)

Rotation schedule:

  • Quarterly at minimum for all production keys
  • Immediately after: leak detection, an employee leaving, any security incident
  • After a major release (if a key could have ended up in a build artifact)

Emergency rotation procedure (documented, tested):

bash
# rotate_all.sh — pseudo-code
# 1. Generate new keys in each provider's dashboard
# 2. Update the secrets in production
wrangler secret put ANTHROPIC_API_KEY  # Cloudflare Workers
wrangler secret put OPENAI_API_KEY

# 3. Verify production is using the new ones
curl https://api.yoursite.com/health/llm

# 4. Revoke the old keys in the provider dashboards (manual click)

# 5. Audit log: what was accessed between the leak and the rotation
# Check the provider usage logs for that period

This script should run in under 10 minutes. It gets tested quarterly.

Multi-account strategy:

  • A primary account + a backup account with each critical provider
  • Different billing methods (if the primary card gets suspended, the backup works)
  • Backup keys already saved in 1Password and tested (not "I'll get to it someday")

Anomaly detection:

  • Daily spend > 2× average → alert + investigation
  • Spend > monthly budget cap → auto-pause the API key
  • Unusual geographic origin (a request from a country where you have no customers) → flag

Incident Response Playbook: 4 steps

When something breaks, you've got adrenaline, panic and a pile of impulses. A playbook gives you structure.

STOP (1-3 minutes)

  • Turn off the affected feature (feature flag → off)
  • Prevent further damage (if it's data corruption, pause write operations)
  • Tell the team "incident in progress" (if you're not alone)
  • DON'T start debugging right away: stop the bleeding first

TRIAGE (5-10 minutes)

  • What happened? (one specific failure mode, not "something broke")
  • When did it start? (a precise timestamp from the logs)
  • Scope: which customers are affected? (1 customer / a segment / everyone)
  • Blast radius: which features are affected?
  • Root cause hypothesis (you may not know for sure yet, but have a direction)

STABILIZE (10-30 minutes)

  • Best option: roll back to the previous known-good deploy
  • Second: fall back to a secondary provider/feature
  • Third: a temporary fix (a quick patch, not the ideal solution)
  • Communicate with customers (status page update + email if it's major)
  • The team's message: "the product is working, we know what happened, we're fixing it"

POSTMORTEM (1-3 hours, after stabilizing)

  • Root cause analysis (5 whys or a fishbone diagram)
  • Timeline reconstruction
  • What went well / What went badly / What to do differently
  • Action items: add monitoring? Prevention? Process improvements?
  • A blameless culture: focus on fixing the system, not on "who's to blame"

Communication templates: ready before you go down

In the middle of an incident you don't have time to write a polished email. Prepare the templates in advance.

Customer email during an outage:

Type this into the chat
Subject: [Service Update] Brief disruption in [Feature] — Status

Hi [Customer],

We've detected a temporary disruption in [feature/service] that began at [time UTC].
The issue is related to [generic cause: third-party API issue, infrastructure incident].

What we're doing:
- [current fix or workaround]
- Our engineering team is working on restoring service

Estimated recovery: [a conservative time; better to overestimate than miss it]
What you can do right now: [a workaround if there is one, or "just wait"]

Next update within [30 minutes / 1 hour].
If you need urgent help: [contact].

Sorry for the inconvenience.
[Your name]

Status page entry:

Code
[Investigating] LLM API Issues — 2026-02-15 14:23 UTC
We're investigating an elevated error rate in [feature]. Some requests are failing.
Updates to follow.

[Update 14:45] Identified — the root cause is related to an upstream provider outage.
We've activated fallback to a secondary provider. Some users are still seeing errors.

[Monitoring 15:10] Fallback active. The error rate has returned to baseline.
We're monitoring the situation; a postmortem will be published within 24 hours.

[Resolved 16:00] Issue resolved. Postmortem: [link]

Internal Slack/Telegram alert:

Type this into the chat
INCIDENT: [Feature] degraded
Severity: [P1 / P2 / P3]
Started: [timestamp]
Owner: [your name]
Status page: [link]
Customers affected: ~[number] or [segment]
Current action: [stabilizing / investigating / monitoring]

Drill plan: test it quarterly

A plan isn't a plan if it's never been run. Quarterly drill.

Q1: Simulate Anthropic API down

  • Block egress traffic to api.anthropic.com in the staging environment
  • Verify the OpenAI fallback activates automatically
  • Measure: time to switch, customer impact, fallback success rate
  • Document gaps → fix → re-test

Q2: Simulate database corruption

  • Take a staging database snapshot
  • Intentionally corrupt one table
  • Practice restoring from the backup
  • Verify data integrity after the restore
  • Measure: RTO, RPO, data loss if any

Q3: Simulate a key leak

  • Pretend ANTHROPIC_API_KEY leaked in Slack
  • Run the emergency rotation procedure
  • Verify all production systems have switched to the new key
  • Verify the old key is revoked in the provider dashboard
  • Measure: total rotation time (target <10 min)

Q4: Full disaster simulation

  • Production down + LLM provider down + backup account unavailable
  • Restore the service on new infrastructure (a new Cloudflare account, new keys)
  • Measure: time to restore, % of data preserved, customer communication delivered

A drill costs 4 hours a quarter. It'll save you days when a real incident happens.


Tools for disaster recovery

Not an ad, just practical options. Prices and terms change, so check them on the websites (what you see later may differ from October 2026).

Category Tool Terms Use case
Backup storage Backblaze B2 Paid by volume, price per GB on the site Object storage, S3-compatible
Backup storage Wasabi Paid by volume, terms on the site An alternative to AWS S3
Sync tool rclone Free, open source A CLI for syncing between cloud storage
LLM observability Langfuse Open source, free; the cloud version has a free Hobby plan (limits on the site) Logging, monitoring LLM calls
Error tracking Sentry Has a free plan for small projects Application errors, stack traces
Uptime monitoring UptimeRobot Has a free plan Pings, status checks
Uptime monitoring Pingdom Paid More serious monitoring
Alerting Telegram bot or email Free Enough for SMB
Alerting PagerDuty Paid Enterprise on-call rotation
Status pages Statuspage (Atlassian) Paid, terms on the site Hosted, professional
Status pages Cachet Free (self-hosted) Open source, your own server
Status pages Instatus Paid, terms on the site Modern UI, simple setup
Secret scanning Gitleaks Free Pre-commit hook
Secret scanning Trufflehog Free Deep repo scans

The minimum stack for SMB: backup storage + UptimeRobot + Langfuse + Sentry + a status page. Some of these are free; the total depends on the plans you pick.


The cost of a disaster vs the cost of preparedness

A rough calculation for a SaaS with $10K MRR. The numbers are illustrative: they're not statistics or a forecast, plug in your own.

Disaster cost (without preparedness):

Cost item Estimate
A 4-hour outage = about 0.6% of the month direct revenue loss ~$60-100
5-15% customer churn after a bad incident $500-1500 MRR lost
Trust damage, takes 6-12 months to recover $2000-5000 in lost upsells
Engineering time on ad-hoc recovery 20-40 hours = $1000-2000
Support tickets from confused customers 30-50 hours = $500-1000
Total for one serious incident $4000-10000

Preparedness cost (annual):

Cost item Estimate
Multi-provider setup, one-time 8-16 hours = $400-800
Backup infrastructure $10-50/month = $120-600/year
Monitoring + alerting $20-100/month = $240-1200/year
Drill time 4h × 4 quarters 16 hours = $800
Total annual $1560-3400 = $130-280/month

Bottom line: in this made-up example, $130-280/month of insurance for a $10K MRR business = 1.3-2.8% of revenue, while one real disaster without a plan costs as much as several months of that insurance.

It's not "is it worth it." It's more like "why haven't I done this yet."


Audience breakdown: what you need

Beginner (personal use, hobby projects):

  • ✅ Git backup for prompts and configs
  • ✅ A pre-commit hook for secrets
  • ❌ Multi-provider isn't needed
  • ❌ A status page is overkill
  • Cost: $0/month

Intermediate (SMB, $1-10K MRR):

  • ✅ Multi-provider fallback (at least 2)
  • ✅ Nightly backup of customer data
  • ✅ Basic monitoring (uptime + LLM errors)
  • ✅ Communication templates ready to go
  • ✅ 1 simple drill a year
  • Cost: $30-50/month

Professional (paying customers, $10-50K MRR):

  • ✅ Full 3-2-1 backup
  • ✅ A multi-provider chain with 3-4 providers
  • ✅ Automated key rotation
  • ✅ Quarterly drills
  • ✅ A public status page
  • ✅ A postmortem culture
  • Cost: $100-200/month

Enterprise ($50K+ MRR, SLA contracts):

  • ✅ Multi-region deployment
  • ✅ Automated failover
  • ✅ 24/7 on-call rotation
  • ✅ SOC 2 compliance backups
  • ✅ Dedicated incident response training
  • Cost: $500+/month

Anti-patterns

❌ A single LLM provider in production: Anthropic-only or OpenAI-only. When it goes down, you go down. In 2026, multi-provider is basic hygiene, not a nice-to-have.

❌ Backups you never test restoring: a classic. The backup exists, but nobody's ever tried to restore it. On D-day it turns out the backup is corrupt or the procedure is broken.

❌ No communication with customers during an outage: silence is worse than bad news. The customer writes to support → gets an auto-reply → sees the product is down → doesn't know what's going on → loses trust.

❌ API keys in a .env committed to Git history: even if you later deleted the commit, it's in Git history forever. If it was ever committed, consider it leaked and rotate immediately.

❌ Manual rollback without a documented procedure: "I remember how to do it" works when you're calm. In an incident, with adrenaline at 2 a.m., you'll forget a step. Documentation = a checklist.

❌ "That rarely happens": until it happens once and you lose an enterprise deal. Probability × impact, not probability × wishful thinking.

❌ Backups at the same provider as production: Cloudflare KV primary + Cloudflare R2 backup. Cloudflare goes down → both are unavailable. The backup needs to be with an independent provider.

❌ Hiding an incident from customers ("we'll fix it now, nobody will notice"): they'll notice. And when they find out you hid it, the trust damage is 10× worse than from honest disclosure.

❌ Relying on a provider's SLA as a guarantee: an SLA gives you a refund (often proportional to the downtime); it doesn't prevent downtime. A 99.9% SLA = about 8.8 hours of permitted downtime a year (0.1% × 8,760 hours).


Readiness checklist

✅ Multi-provider fallback works (tested last quarter) ✅ Backups run automatically every night ✅ The last restore test passed < 90 days ago ✅ Communication templates written for the top 3 incident types ✅ Status page set up and tested ✅ Emergency rotation procedure documented ✅ All API keys in a secrets manager, zero in code ✅ A pre-commit hook scans for secrets ✅ Monitoring alerts on an LLM error rate spike ✅ Spending anomaly detection is active ✅ A drill is scheduled for next quarter ✅ A postmortem template is ready

If you have ≤ 6 ✅, you're in the risk zone. ≤ 9 ✅ is standard SMB readiness. 12/12 is the professional level.


Practice

Step 1: A multi-provider fallback in Python

python
# llm_router.py — a simple fallback router
import os
import time
from typing import Optional
import anthropic
import openai
from google import genai

class LLMRouter:
    def __init__(self):
        self.anthropic_client = anthropic.Anthropic(
            api_key=os.getenv("ANTHROPIC_API_KEY")
        )
        self.openai_client = openai.OpenAI(
            api_key=os.getenv("OPENAI_API_KEY")
        )
        self.gemini_client = genai.Client(
            api_key=os.getenv("GOOGLE_API_KEY")
        )

        self.providers = [
            # model names are an example; see the "What's current" page for current ones
            ("anthropic", "claude-sonnet-5-5"),
            ("openai", "gpt-6.1-sol"),
            ("gemini", "gemini-3.8-flash"),
        ]

    def call(self, prompt: str, max_tokens: int = 1000) -> dict:
        errors = []

        for provider_name, model in self.providers:
            try:
                start = time.time()
                response = self._call_provider(
                    provider_name, model, prompt, max_tokens
                )
                latency = time.time() - start

                return {
                    "provider": provider_name,
                    "model": model,
                    "text": response,
                    "latency_ms": int(latency * 1000),
                    "fallback_used": provider_name != "anthropic"
                }
            except Exception as e:
                errors.append({
                    "provider": provider_name,
                    "error": str(e),
                    "type": type(e).__name__
                })
                # Log the failure for monitoring
                print(f"[FAIL] {provider_name}: {e}")
                continue

        raise Exception(f"All providers failed: {errors}")

    def _call_provider(self, name, model, prompt, max_tokens):
        if name == "anthropic":
            r = self.anthropic_client.messages.create(
                model=model,
                max_tokens=max_tokens,
                messages=[{"role": "user", "content": prompt}],
                timeout=10
            )
            return "".join(b.text for b in r.content if b.type == "text")

        elif name == "openai":
            r = self.openai_client.chat.completions.create(
                model=model,
                max_completion_tokens=max_tokens,  # GPT-5 and newer models use this instead of max_tokens
                messages=[{"role": "user", "content": prompt}],
                timeout=10
            )
            return r.choices[0].message.content

        elif name == "gemini":
            r = self.gemini_client.models.generate_content(
                model=model,
                contents=prompt
            )
            return r.text

# Usage
router = LLMRouter()
result = router.call("Explain what disaster recovery is in 3 sentences")
print(f"Provider used: {result['provider']}")
print(f"Fallback used: {result['fallback_used']}")
print(result['text'])

Step 2: A backup script for prompts and configs

bash
#!/bin/bash
# backup_ai_stack.sh — nightly backup of the AI infrastructure

set -e

BACKUP_DIR="/backups/$(date +%Y-%m-%d)"
B2_BUCKET="my-ai-backup"

mkdir -p "$BACKUP_DIR"

# 1. Back up prompts/configs (Git already covers this, but an extra copy)
tar czf "$BACKUP_DIR/prompts.tar.gz" ./prompts ./.claude

# 2. Back up customer data from Cloudflare D1
wrangler d1 export my-database --remote --output="$BACKUP_DIR/db.sql"

# 3. Back up KV storage: first the list of keys, then the values
# (see the Wrangler documentation for the file format kv bulk get expects)
wrangler kv key list --namespace-id=$KV_ID --remote > "$BACKUP_DIR/kv-keys.json"
wrangler kv bulk get "$BACKUP_DIR/kv-keys.json" --namespace-id=$KV_ID --remote > "$BACKUP_DIR/kv.json"

# 4. Upload to Backblaze B2 (off-site)
rclone copy "$BACKUP_DIR" "b2:$B2_BUCKET/$(date +%Y-%m-%d)"

# 5. Retention: delete local backups older than 30 days
find /backups -type d -mtime +30 -exec rm -rf {} +

# 6. Verify backup integrity
SIZE=$(du -sh "$BACKUP_DIR" | cut -f1)
echo "Backup completed: $BACKUP_DIR ($SIZE)"

# 7. Notification on success/failure
if [ $? -eq 0 ]; then
    echo "Backup OK $(date)" >> /var/log/ai-backup.log
else
    echo "Backup FAILED $(date)" | mail -s "ALERT: Backup failed" admin-notifications
fi

Run it with cron: 0 3 * * * /scripts/backup_ai_stack.sh


Step 3: A quick incident response checklist (print it and keep it handy)

markdown
# INCIDENT RESPONSE — 10 minutes to recovery

## STOP (minutes 0-3)
[ ] Feature flag → off for the affected feature
[ ] Notify the team in Slack: "INCIDENT in progress, owner: [me]"
[ ] Open a status page draft

## TRIAGE (minutes 3-10)
[ ] What broke? (specific failure mode)
[ ] When did it start? (timestamp from the logs)
[ ] Who's affected? (segment / count)
[ ] Severity: P1 / P2 / P3
[ ] Root cause hypothesis (you may not know for sure yet)

## STABILIZE (minutes 10-30)
[ ] Option A: Roll back to the previous good deploy
[ ] Option B: Activate fallback (provider, region)
[ ] Option C: Temporary fix (a quick patch)
[ ] Update the status page: "Investigating" → "Identified" → "Monitoring"
[ ] Send a customer email if P1

## POSTMORTEM (after stabilizing, within 48h)
[ ] Reconstruct the timeline
[ ] 5 whys analysis
[ ] What went well / badly / what to change
[ ] Action items with an owner and a deadline
[ ] Update the playbook if you found a gap

Step 4: Test your fallback (a drill)

bash
# Simulate the Anthropic API being down: block egress
# In local development, via /etc/hosts:
echo "127.0.0.1 api.anthropic.com" | sudo tee -a /etc/hosts

# Start your app
python app.py

# Make a request: it should fall back to OpenAI/Gemini
curl -X POST http://localhost:8000/api/generate \
  -H "Content-Type: application/json" \
  -d '{"prompt": "Test fallback"}'

# Check:
# - Did you get a response? (success criteria)
# - Which provider was used? (should not be anthropic)
# - Is the latency acceptable? (target < 30 sec total)
# - Did the log record the failure + fallback? (audit trail)

# Revert /etc/hosts:
sudo sed -i '' '/api.anthropic.com/d' /etc/hosts

If this test doesn't work in your dev environment, it definitely won't work during a production incident.


Tools and resources


Key takeaways

It's not "if" your LLM provider goes down, it's "when." Major providers (Anthropic, OpenAI and others) have serious outages: see the history on their status pages. A single-provider stack = just a matter of time until your first downtime.

3-2-1 backup isn't paranoia, it's hygiene. 3 copies, 2 storage types, 1 off-site. In the made-up example above, that's on the order of tens of dollars a month for an SMB, while one real disaster without a backup costs thousands in lost revenue, churn and repair work.

Playbook + quarterly drill > improvised heroics. A plan written in advance and tested 4 times a year turns a 4-hour incident into a 10-minute switch. Pilots train for engine failure not because it happens often, but so they can handle it when it does.


Next lesson

→ Telegram bots with the Claude API. On cutting the costs of your AI stack: Cost Engineering

The mark stays in this browser only and is never sent anywhere. My progress