The gist
MLOps sounds like something from a big tech company with a hundred engineers. In practice, it's the answer to one question: "Is my AI working as well as it did a week ago?" For an indie developer, MLOps comes down to three things: keep an eye on cost, keep an eye on quality, and don't break what already works.
Key concepts
- Prompt drift: a prompt gets worse over time without any change to the code
- LangSmith: tracing and monitoring for your Claude calls
- Promptfoo: regression testing for prompts
- Cost monitoring: budget alerts and spending anomalies
- Prompt versioning: Git for prompts, not just for code
- Baseline metrics: quality metrics you compare against over time
Theory
What prompt drift is and why it happens
You wrote a prompt three months ago. The code hasn't changed. But suddenly you notice the answers have gotten longer, or less specific, or a bit different in tone. You didn't touch anything. So what happened?
Causes of prompt drift:
The model changed. Your code uses an alias instead of a specific version, or the old model was retired from the API and you moved to a new one. The behavior shifts a little or a lot. Pin the model version in your code and keep an eye on the list of models and retirements: What's current.
The context changed. Your users started writing differently, the questions changed, new patterns appeared that the prompt didn't account for.
The agent chain shifted. One agent started producing a slightly different format → the next agent stopped parsing it correctly.
The system prompt is out of date. You described a context that no longer applies.
LangSmith: tracing your Claude calls
LangSmith (from LangChain) is an observability platform for AI. It records every call: the input prompt, the response, latency, cost and tags.
Why you need it:
- See what's actually happening inside your agent chain
- Compare answers "now vs. a week ago"
- Find expensive calls you can optimize
- Debug when something breaks in production
Integration with the Anthropic SDK:
import anthropic
from langsmith import traceable
from langsmith.wrappers import wrap_anthropic
# wrap_anthropic records the tokens and cost of every call
client = wrap_anthropic(anthropic.Anthropic())
@traceable(name="content-generator")
def generate_content(topic: str, style: str) -> str:
response = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=1000,
messages=[{
"role": "user",
"content": f"Write content about '{topic}' in a {style} style"
}]
)
return "".join(b.text for b in response.content if b.type == "text")
# Every call automatically shows up in the LangSmith dashboard
result = generate_content("email marketing", "friendly")What you see in the dashboard:
- Every call with the full prompt and response
- Latency percentiles (p50, p90, p99)
- Cost per endpoint
- Anomalies (a call took 10 times longer than usual)
Promptfoo: testing your prompts
Promptfoo is a command-line tool for running regression tests on prompts. It works like unit tests, but for AI. In March 2026, Promptfoo was acquired by OpenAI; the company says the project stays open source, but keep an eye on how it develops.
npm install -g promptfooTest configuration (promptfooconfig.yaml):
prompts:
- "Analyze this review and identify the sentiment: {{review}}"
providers:
- anthropic:messages:claude-haiku-4-5
tests:
- vars:
review: "Great product, very happy with it!"
assert:
- type: contains
value: "positive"
- type: not-contains
value: "negative"
- vars:
review: "Terrible quality, never buying again"
assert:
- type: contains
value: "negative"
- type: llm-rubric
value: "The answer should identify negative sentiment"
- vars:
review: "It's fine, nothing special"
assert:
- type: contains-any
value: ["neutral", "mixed", "unclear"]# Run the tests
promptfoo eval
# Compare two versions of a prompt
promptfoo eval --prompts v1-prompt.txt v2-prompt.txtCI/CD integration:
# .github/workflows/prompt-tests.yml
name: Prompt Regression Tests
on: [push, pull_request]
jobs:
test-prompts:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v5
- run: npm install -g promptfoo
- run: promptfoo eval --no-progress-bar
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}Now the tests run on every push. If the prompt has regressed, promptfoo exits with an error and CI fails.
Cost monitoring: keeping your money under control
For an indie builder, money matters most. One bug in a prompt can multiply your daily spending.
Built-in cost monitoring:
import anthropic
from datetime import datetime, timezone
import json
import os
class CostTracker:
def __init__(self, daily_budget_usd: float = 20.0):
self.client = anthropic.Anthropic()
self.daily_budget = daily_budget_usd
self.today_cost = 0.0
self.log_file = "cost_log.jsonl"
# Prices per 1M tokens as of October 2026. They change: see the current ones on
# the "What's current" page (../actual.html). Add other models the same way.
COSTS = {
"claude-sonnet-5-5": {"input": 2.0, "output": 10.0},
"claude-haiku-4-5": {"input": 1.0, "output": 5.0},
}
def track_call(self, model: str, input_tokens: int, output_tokens: int, endpoint: str):
rates = self.COSTS.get(model, self.COSTS["claude-sonnet-5-5"])
cost = (input_tokens / 1_000_000 * rates["input"] +
output_tokens / 1_000_000 * rates["output"])
self.today_cost += cost
entry = {
"ts": datetime.now(timezone.utc).isoformat(),
"model": model,
"endpoint": endpoint,
"input_tokens": input_tokens,
"output_tokens": output_tokens,
"cost_usd": round(cost, 6)
}
with open(self.log_file, "a") as f:
f.write(json.dumps(entry) + "\n")
# Alert when 80% of the budget is used
if self.today_cost > self.daily_budget * 0.8:
self.send_alert(f"⚠️ Spent ${self.today_cost:.2f} of the ${self.daily_budget} budget")
return cost
def send_alert(self, message: str):
# Notification through an incoming webhook (for example, Slack or Discord)
import requests
requests.post(
os.environ['ALERT_WEBHOOK_URL'],
json={"text": message}
)
tracker = CostTracker(daily_budget_usd=20.0)Spending anomalies:
def detect_cost_anomaly(log_file: str, window_days: int = 7):
"""Compares today's spending with the average for the past week."""
import statistics
with open(log_file) as f:
entries = [json.loads(line) for line in f]
today = datetime.now(timezone.utc).date()
daily_costs = {}
for entry in entries:
date = datetime.fromisoformat(entry['ts']).date()
daily_costs[date] = daily_costs.get(date, 0) + entry['cost_usd']
recent = [cost for date, cost in daily_costs.items()
if (today - date).days <= window_days and date != today]
if len(recent) < 3:
return False
avg = statistics.mean(recent)
std = statistics.stdev(recent)
today_cost = daily_costs.get(today, 0)
# Anomaly: today costs more than 2 standard deviations above the average
if today_cost > avg + 2 * std:
return f"🚨 Anomaly! Today: ${today_cost:.2f}, average: ${avg:.2f}"
return FalseVersioning prompts like code
Prompts are code. They belong in Git with a history of changes.
Repository structure:
prompts/
v1/
system-prompt.md # Version 1.0
user-template.md
v2/
system-prompt.md # Version 2.0: changed the tone
user-template.md
current -> v2/ # Symlink to the current version
CHANGELOG.md # What changed and whyCHANGELOG.md for prompts:
## v2.0 — 2026-04-15
### Changes
- Removed the "be brief" instruction: answers were too short for business emails
- Added formatting examples
- Changed the tone from formal to "friendly-professional"
### Metrics before/after (promptfoo)
- Average answer length: 150 → 280 words ✅
- Test "contains a greeting": 60% → 95% ✅
- Test "no 'To whom it may concern'": 100% (unchanged) ✅
### Rollback
git checkout v1/ if v2 turns out worsePractice
Step 1: Install the tools
# Promptfoo for testing prompts
npm install -g promptfoo
# LangSmith (for tracing)
pip install langsmith
# Environment variables
export LANGSMITH_TRACING="true" # tracing won't turn on without this
export LANGSMITH_API_KEY="ls__xxx"
export LANGSMITH_PROJECT="my-ai-app"Step 2: Write your first prompt tests
Create a promptfooconfig.yaml file for your main prompt. Come up with 5-10 test cases, including edge cases: empty input, very long text, text in different languages, a potentially malicious request.
promptfoo eval
# See what passed and what failedStep 3: Set up cost tracking
# cost_tracker.py
# Copy the code from the lesson and set ALERT_WEBHOOK_URL
# Add a tracker.track_call() call after every messages.create()Step 4: Basic monitoring
Run this once a day:
# Daily total
python -c "
import json
from datetime import datetime, timezone
today = datetime.now(timezone.utc).date().isoformat()
total = 0
calls = 0
with open('cost_log.jsonl') as f:
for line in f:
e = json.loads(line)
if e['ts'].startswith(today):
total += e['cost_usd']
calls += 1
print(f'Today: {calls} calls, \${total:.4f}')
"Step 5: Assignment
- Take any prompt from your project
- Write 10 tests in promptfoo
- Run the tests and record your baseline (how many passed)
- Make a small change to the prompt
- Run them again: did it get better or worse?
- Commit both versions to Git with a description of the changes
Tools and resources
- LangSmith: smith.langchain.com (there's a free plan; check the site for limits)
- Promptfoo: promptfoo.dev,
npm install -g promptfoo - Helicone, an alternative to LangSmith: helicone.ai (in March 2026 the company was acquired by Mintlify; the service is in maintenance mode with no new features, so for a new project it's better to pick another tool)
- Braintrust, another evaluation and monitoring platform: braintrust.dev
- Weights & Biases, enterprise MLOps: wandb.ai
Key takeaways
MLOps for indie builders isn't DevOps. It's discipline. Three rules: test your prompts before you deploy (Promptfoo), track spending with alerts (cost tracker), and log everything in LangSmith so you have somewhere to look when something goes wrong. Prompt drift quietly erodes quality. Without tests, you won't notice it until clients start leaving.
Next lesson
→ Production observability: what to monitor when your agent is live
The mark stays in this browser only and is never sent anywhere. My progress