The gist
Using the Anthropic API is like buying electricity from the city grid. Convenient, reliable, you pay by the meter. You don't have to think about transformers, wiring or the control room.
Self-hosted AI is like building your own power plant next to the building. Expensive up front: a turbine, cooling, an operator. But if you run a factory with 50+ machines going 24/7, your own power plant pays for itself in 1-2 years and makes you independent of the supplier.
When self-hosting pays off:
- Steady LLM spend of $5,000+/month (the hardware break-even point)
- 50+ concurrent users (real workload)
- A regulated industry (compliance requires data residency)
- Geopolitics (sanctions risk, sovereignty)
- Custom models (fine-tuned for a domain)
In every other case, the API is cheaper, simpler and more reliable.
The thresholds and amounts in this lesson are benchmarks for doing the math, not a price list: hardware prices and plans change, so check before you buy. Current API prices: What's current.
🎯 Decision tree: do you need to self-host
Before you spend $50K on hardware, go through this checklist. A "no" at every node means stay on the API.
LLM costs > $5,000/month, steadily, for the last 3 months?
→ Yes → Self-host breaks even in 12-18 months
→ No → The API is cheaper, leave the hardware alone
Regulated industry (healthcare, finance, gov, defense)?
→ Yes → Self-host required for compliance (HIPAA, PCI DSS, strict GDPR)
→ No → next question
Team of 50+ people using AI every day at the same time?
→ Yes → Self-host for cost + privacy + latency
→ No → The API is streamlined, don't overcomplicate it
Need a specific model (fine-tuned on domain data, custom)?
→ Yes → Self-host (cloud providers limit fine-tuning to a list of models and terms)
→ No → next question
Geographic / political restrictions (sanctions exposure, data sovereignty laws)?
→ Yes → Self-host = sovereignty insurance
→ No → The API works
All answers "no"?
→ Stay on the API. Self-hosting = overengineering for you.Key concepts
- Inference engine: the program that takes a request and generates a response from an LLM (vLLM, SGLang, TGI, Ollama). It's the engine in the car
- vLLM: the most popular open-source inference engine, Apache 2.0, with an excellent balance of throughput and ease of use
- SGLang: one of the fastest in throughput, grew out of the LMSYS project (Berkeley), optimized for structured outputs and parallel sampling
- API gateway: a proxy layer that turns local inference into an OpenAI-compatible endpoint (LiteLLM). Existing clients work without changes
- Tensor parallelism: splitting a model across several GPUs. A 70B model needs ~140GB of VRAM in FP16, so one card won't do
- Quantization: compressing a model to 4-bit or 8-bit to cut VRAM by 2-4x with minimal loss of quality
- Throughput: how many tokens per second the system generates (important for multi-user setups)
- Latency: the delay before the first token (TTFT, time to first token). Critical for interactive UX
- Concurrent users: how many users can work at the same time without latency degrading
Theory
Hardware tiers for a company
The size of the hardware is determined by three parameters: model size (B parameters), number of concurrent users, and latency requirements.
Tier 1: Departmental (5-20 users)
The target profile: a development, marketing or support department at a mid-sized company. Non-critical load, an experiment or an internal tool.
| Parameter | Value |
|---|---|
| Hardware | 1 server with 2x RTX 4090 (48GB total VRAM) or 1x A100 40GB |
| Model | an open model in the 30-70B class (for example, Qwen3 32B; 70B in 4-bit quantization) |
| One-time hardware cost | $8-15K |
| Electricity | $50-100/month (~500W under load) |
| Concurrent users | 20-30 |
| Latency (TTFT) | 1-3 sec |
| Throughput | 30-60 tokens/sec per user |
| Setup time | 2-4 days |
Example: a team of 10 developers doing code review and docs generation. The build: two high-end consumer graphics cards with 24GB each, a server case, a Threadripper- or Xeon-class processor, 128GB RAM, a 4TB NVMe drive, a 1500W power supply. Roughly $10K total (a ballpark; check component prices when you buy).
Tier 2: Team (20-100 users)
The target profile: a company where AI has become a production-critical tool. A SaaS startup with built-in AI, an agency with 50 developers, a mid-size fintech.
| Parameter | Value |
|---|---|
| Hardware | 2-4 servers with 4x A100 80GB or 2x H100 80GB per server |
| Model | a 70B-class model at full precision or larger (including MoE models) |
| One-time cost | $50-150K |
| Electricity + cooling | $300-800/month |
| Concurrent users | 100+ |
| Latency (TTFT) | 0.5-2 sec |
| Throughput | 80-150 tokens/sec per user |
| Setup time | 1-2 weeks |
At this level you already need a server room with temperature control (HVAC), a 5-10 kW UPS and dedicated network infrastructure. You need a dedicated ops engineer, at least part-time.
Tier 3: Enterprise (100-1,000+ users)
The target profile: a large company, a bank, a government agency. AI is core infrastructure.
| Parameter | Value |
|---|---|
| Hardware | a cluster of 8-16 servers with H100 80GB / B200 |
| Model | Custom fine-tuned 70B+, possibly several at once |
| One-time cost | $500K to $5M |
| OpEx | $5-50K/month (electricity + cooling + ops team) |
| Concurrent users | 1,000+ |
| Latency (TTFT) | <500ms |
| Setup time | 2-3 months |
This is data center territory. You need a team of 3-5 people: SRE, ML engineer, security, network. An SLA of 99.9%+ requires redundancy at every level.
Inference engines compared (as of October 2026)
The inference engine is the heart of the system. Your choice determines throughput, latency and operational complexity. A realistic shortlist of 5 options:
| Engine | Best for | License | Throughput | Setup time | When to choose it |
|---|---|---|---|---|---|
| vLLM | Most popular, balanced | Apache 2.0 | High | 1-2 days | The default choice. Big community, lots of guides |
| SGLang | Best throughput | Apache 2.0 | Highest | 2-3 days | When you need maximum performance, structured outputs |
| TGI (HuggingFace) | HF ecosystem | Apache 2.0 | High | 1 day | The project is in maintenance mode: for a new deployment, go with vLLM or SGLang |
| Ollama | Easiest, small scale | MIT | Low | 30 minutes | Pilot, prototype, up to 10 users |
| TensorRT-LLM | NVIDIA only, fastest | see the repository | Highest on NVIDIA | 1 week | When it's NVIDIA only and you need the maximum |
vLLM (github.com/vllm-project/vllm) is the gold standard. PagedAttention is its signature technique for using VRAM efficiently. An OpenAI-compatible API out of the box. Most tutorials and production case studies are built on vLLM.
SGLang (github.com/sgl-project/sglang) beats vLLM on throughput in a number of benchmarks, but measure it on your own workload. RadixAttention reuses the KV cache across requests. It's especially good when you have lots of similar system prompts (a typical agent scenario). The downside: the ecosystem is younger, with fewer ready-made guides.
TGI (Text Generation Inference) (github.com/huggingface/text-generation-inference) is a HuggingFace product. It integrates well with the HF Hub and is simple to deploy. But as of October 2026 the project is in maintenance mode: only minor fixes are accepted, and the authors themselves recommend vLLM and SGLang. For a new project, it's better to start with those.
Ollama is for pilots and small teams. It starts with one command and has a nice CLI. It doesn't scale beyond ~10 concurrent users. Details: Local AI models: Ollama, LM Studio and private AI.
A Docker stack for team deployment
A production setup is made of 3-4 containers: inference engine + gateway + UI + auth. The basic configuration for Tier 1:
# docker-compose.yml — Tier 1 team setup (5-20 users)
# Modern Docker Compose doesn't need the version key
services:
# Inference engine
vllm:
image: vllm/vllm-openai:latest
runtime: nvidia
ports:
- "8000:8000"
volumes:
- ./models:/root/.cache/huggingface
command:
# The bf16 model weighs ~65GB: on 2×24GB use a quantized version (AWQ or GPTQ),
# on A100/H100 80GB use it as is
- --model
- Qwen/Qwen3-32B
- --tensor-parallel-size
- "2"
- --gpu-memory-utilization
- "0.90"
- --max-model-len
- "32768"
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 2
capabilities: [gpu]
restart: unless-stopped
# API gateway: an OpenAI-compatible proxy with rate limiting + audit
litellm:
# Pin a specific version instead of main-latest (see below about March 2026)
image: ghcr.io/berriai/litellm:main-latest
ports:
- "4000:4000"
environment:
- DATABASE_URL=postgres://litellm:secret@postgres:5432/litellm
- MASTER_KEY=sk-master-internal-key-replace-me
volumes:
- ./litellm-config.yaml:/app/config.yaml
command: ["--config", "/app/config.yaml", "--port", "4000"]
depends_on:
- postgres
- vllm
restart: unless-stopped
# PostgreSQL for LiteLLM (audit, keys, budgets)
postgres:
image: postgres:16
environment:
- POSTGRES_USER=litellm
- POSTGRES_PASSWORD=secret
- POSTGRES_DB=litellm
volumes:
- postgres-data:/var/lib/postgresql/data
restart: unless-stopped
# User-facing UI
open-webui:
image: ghcr.io/open-webui/open-webui:main
ports:
- "3000:8080"
environment:
- OPENAI_API_BASE_URL=http://litellm:4000/v1
- OPENAI_API_KEY=sk-master-internal-key-replace-me
- WEBUI_AUTH=true
volumes:
- open-webui:/app/backend/data
depends_on:
- litellm
restart: unless-stopped
volumes:
open-webui:
postgres-data:The LiteLLM config (litellm-config.yaml):
model_list:
- model_name: qwen-32b
litellm_params:
model: openai/Qwen/Qwen3-32B
api_base: http://vllm:8000/v1
api_key: dummy # vLLM doesn't require real auth
general_settings:
master_key: sk-master-internal-key-replace-me
database_url: postgres://litellm:secret@postgres:5432/litellm
litellm_settings:
drop_params: true
set_verbose: false
json_logs: true
cache: trueStarting it up:
docker compose up -d
# vLLM will download the Qwen3 32B model (~65GB in bf16) on first start
# Watch the progress: docker compose logs -f vllmAfter startup:
http://localhost:8000/v1: the raw vLLM endpointhttp://localhost:4000/v1: the LiteLLM gateway (use this one)http://localhost:3000: Open WebUI for users
An OpenAI-compatible API gateway: LiteLLM
LiteLLM (github.com/BerriAI/litellm) is a critical component of the stack. Without it, a self-hosted setup stays an internal toy; with it, it becomes an enterprise-grade platform.
What LiteLLM gives you:
- An OpenAI-compatible API: every client that works with OpenAI/Anthropic (Cursor, Continue.dev, Aider, custom scripts) works with your local vLLM without rewriting code
- Per-user API keys: each developer gets their own key. Revoking one doesn't break the rest
- Rate limiting: per-user or per-team limits so one script can't hog all the hardware
- Cost tracking + budgets: who spent how many tokens, alerts when a limit is exceeded
- Audit log: every request in PostgreSQL with user_id, model, tokens, latency
- Multi-model routing: you can route simple requests to a small model and complex ones to a big one
- Fallback to API: if the local vLLM goes down, you can proxy to the Anthropic API
Security of the gateway itself: on March 24, 2026, malicious versions of the litellm package (1.82.7 and 1.82.8) briefly appeared on PyPI and stole credentials. The authors' write-up: Security Update: Suspected Supply Chain Incident. The lesson for any gateway that holds keys: pin the version, install only from trusted sources and signed images, and update deliberately, not automatically.
Creating a user key through the LiteLLM admin:
# Create a team
curl -X POST http://localhost:4000/team/new \
-H "Authorization: Bearer sk-master-internal-key-replace-me" \
-H "Content-Type: application/json" \
-d '{
"team_alias": "backend-team",
"max_budget": 100.0,
"models": ["qwen-32b"]
}'
# Create a key for a developer
curl -X POST http://localhost:4000/key/generate \
-H "Authorization: Bearer sk-master-internal-key-replace-me" \
-H "Content-Type: application/json" \
-d '{
"team_id": "backend-team",
"user_id": "dev_user_42",
"max_budget": 10.0,
"duration": "30d",
"rpm_limit": 60,
"tpm_limit": 100000
}'
# Response: {"key": "sk-1a2b3c4d...", "expires": "..."}The developer uses their key in any OpenAI-compatible tool:
# Aider (the openai/ prefix tells it this is an OpenAI-compatible endpoint)
export OPENAI_API_BASE=http://internal-llm:4000/v1
export OPENAI_API_KEY=sk-1a2b3c4d...
aider --model openai/qwen-32b
# Cursor: set a Custom API in the settings
# Continue.dev in VS Code: config.yaml (config.json is deprecated):
# models:
# - name: Internal Qwen
# provider: openai
# model: qwen-32b
# apiBase: http://internal-llm:4000/v1
# apiKey: sk-1a2b3c4d...
# See the Continue documentation for the exact format.Authentication & access control
Handing users the master_key is like giving every employee the key to the server room. You need SSO + per-user provisioning.
Authentik (goauthentik.io) is an open-source IdP and a direct competitor to Okta. Self-hosted, supports SAML, OAuth2, OIDC. Integrates with Google Workspace, Microsoft 365, LDAP.
Keycloak is the older sibling, enterprise-tested, but harder to set up. If the company already runs a Java stack, it's the natural choice.
A typical flow:
An employee signs in to Open WebUI
↓
Open WebUI redirects to Authentik
↓
Authentik checks via Google Workspace SSO
↓
Authentik returns a JWT with user_id, groups
↓
Open WebUI creates a session
↓
Open WebUI calls LiteLLM with a user-specific API key
↓
LiteLLM checks the key and limits, logs it, passes it to vLLM
↓
vLLM generates the responseKey controls:
- Groups → models: junior developers get only the small model, seniors get access to the big one (70B)
- Per-user budgets: a small default budget per user, increased on request through a manager
- Rate limits: 60 req/min by default, so a stray bash script in a loop can't take the system down
- Audit logging: every request is written to Postgres, with 90-day retention (or longer for compliance)
Monitoring stack
Without monitoring, a self-hosted setup turns into a black box. When users start complaining "it's slow," you need to see the metrics right away.
The minimum stack:
- Prometheus: collects metrics. vLLM exposes a
/metricsendpoint natively - Grafana: dashboards. Ready-made templates for vLLM are at grafana.com/dashboards
- Loki: log aggregation (optional, for big setups)
- OpenTelemetry: distributed tracing (optional, for multi-service setups)
Metrics to track from day 1:
- GPU utilization (% — underused = overpaying, overloaded = a queue)
- VRAM usage (% — approaching 95% = OOM coming soon)
- Tokens/sec (aggregate throughput)
- Time to first token (P50, P95, P99 — the UX metric)
- Queue depth (if it's growing, you need more hardware)
- Cost per user (LiteLLM exports it)
- Error rate (5xx, timeouts)
Extending docker-compose for monitoring:
prometheus:
image: prom/prometheus:latest
ports:
- "9090:9090"
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
restart: unless-stopped
grafana:
image: grafana/grafana:latest
ports:
- "3001:3000"
environment:
- GF_SECURITY_ADMIN_PASSWORD=admin-replace-me
volumes:
- grafana-data:/var/lib/grafana
restart: unless-stoppedprometheus.yml:
global:
scrape_interval: 15s
scrape_configs:
- job_name: vllm
static_configs:
- targets: ['vllm:8000']
metrics_path: /metrics
- job_name: litellm
static_configs:
- targets: ['litellm:4000']
metrics_path: /metricsFine-tuning your own model: an extra option
If the company has domain-specific data (legal documents, medical protocols, code in an internal DSL), fine-tuning can give quality a significant boost.
Baseline numbers (ballparks, check them against your task):
- Hardware: 1x A100 80GB for a small fine-tune (7B-13B), 4x for a large one (70B)
- Data: 1K-10K high-quality examples for a good result
- Tools: Axolotl (github.com/axolotl-ai-cloud/axolotl), LLaMA-Factory (github.com/hiyouga/LLaMA-Factory)
- Time: 2-24 hours depending on model size and data volume
- Cost: $100-1,000 when renting GPUs (RunPod, Lambda Labs)
Fine-tuning details are in the lesson Fine-tuning: when prompts aren't enough. What matters here: after fine-tuning, the model is deployed to the same vLLM/SGLang as the base model. Only the --model parameter changes, to the path of the fine-tuned checkpoint.
Real-world setups: 3 case studies
These are generalized illustrations, not reports from specific companies; the amounts are illustrative. Check data requirements (GDPR, medical and banking regulations) with a lawyer: this lesson is not a substitute for legal advice.
Case 1: A regional fintech (50 developers)
Context: a mid-sized bank building core banking software; data residency requirements in its home country rule out foreign cloud APIs.
Stack:
- 4x A100 80GB (1 server with tensor parallel 4)
- An open 70B-class model (the Qwen family) + LiteLLM + Open WebUI + Authentik
- Uses: automated code review, documentation generation, security analysis on pull requests
Economics:
- Hardware: ~$200K (servers + networking + UPS)
- OpEx: $400/month in electricity, 0.5 FTE ops engineer
- The alternative (Anthropic): ~$50K/year at the current load
- Break-even: ~4.4 years even without the ops engineer's salary ($200K / ($50K − $4.8K electricity) a year)
- Sovereignty: 100%, no data leaves the perimeter
Case 2: European healthcare SaaS
Context: a SaaS for clinics in Germany, GDPR + national medical data law. No request containing patient data can go to the US/UK.
Stack:
- 2 servers with 2x H100 80GB each (redundancy)
- SGLang + an open 70B-class model, fine-tuned on anonymized medical transcripts
- Authentik SSO integrated with the existing Active Directory
- Uses: an assistant for doctors (research, summarization), drafts for patient communication (always under supervision)
Economics:
- Hardware: €300K
- OpEx: €600/month for infrastructure + 1 FTE ops
- Compliance: passed a GDPR audit, medical data certification
- Risk reduction (the potential fine for a data breach): millions of euros
Case 3: Latin American media company
Context: a media group in Mexico City publishing in Spanish, Portuguese, English and Quechua. Large volumes of translation and rewriting.
Stack:
- 2x RTX 4090 (one server)
- Qwen3 32B (quantized) + Open WebUI + LiteLLM
- Uses: a translation pipeline for 5 languages, generating article drafts, A/B headline variants
Economics:
- Hardware: $12K one-time
- OpEx: $80/month in electricity
- The alternative before self-hosting: $1,500/month (DeepL + OpenAI)
- Break-even: ~8.5 months ($12K / ($1,500 − $80) a month)
- An extra win: the ability to fine-tune for Latin American regional Spanish (which commercial translators don't account for)
Operational concerns: what breaks in production
Self-hosting isn't "set it and forget it." Here's the real list of things that need attention every week:
- Uptime: 99.9% requires a redundant setup (2 servers minimum, automatic failover)
- GPU failures: 1-2 cards a year die under 24/7 load. Keep spares
- Cooling: the server room has to stay at 64-72°F (18-22°C). Overheating = accelerated GPU wear
- Power: a UPS with 15+ minutes for a graceful shutdown, a generator for production
- Model updates: new versions of Qwen/Llama/Mistral come out every quarter. Updating a model = 1-2 hours of downtime
- OS / driver updates: NVIDIA drivers need care (incompatibilities with vLLM versions)
- Backup: models weigh 50-200GB, so you need a plan for storing checkpoints
- Maintenance time: budget 1-2 hours of ops time a week, even for a stable setup
Cost comparison: 1 year in detail
A comparison for a team of 50 developers with ~$5K/month in usage if they were on the API:
| Approach | Year 1 cost | Year 2 cost | Sovereignty | Flexibility |
|---|---|---|---|---|
| Anthropic API ($5K/month) | $60K | $66K (usage growth) | ❌ Vendor lock | ✅ Any model |
| Self-host Tier 1 (Qwen 32B) | $20K hardware + $1.2K ops | $1.2K ops | ✅ Full | ⚠ One model |
| Self-host Tier 2 (Llama 70B + redundancy) | $80K + $5K ops | $5K ops | ✅ Full | ✅ Several models |
| Hybrid (API primary + local fallback) | $40K (mix) | $36K (balancing) | ⚠ Partial | ✅ Best of both |
Hybrid often turns out to be the sweet spot: most requests go through the local vLLM (cheap, sovereign), and complex requests go through the Anthropic API (when you need a strong cloud model). LiteLLM can route them automatically.
Audience: how well this applies
Beginner (1 developer, personal project)
You don't need to self-host. Come back when your LLM bills are steadily > $300/month. Until then, use the Anthropic/OpenAI API: saving time matters more than saving money.
If you want to play with local models to learn, use Ollama (see the lesson Local AI models: Ollama, LM Studio and private AI); it's up and running in 10 minutes.
Intermediate (small team, 5-20 people)
A Tier 1 setup makes sense if:
- You steadily spend $1,500+/month on LLMs
- Or you have at least one requirement: privacy, sovereignty, latency
Config: vLLM + LiteLLM + Open WebUI. One server with 2x RTX 4090 or 1x A100. Setup in a week, ops 2-4 hours a week.
Pilot before production: run it for 2 weeks on a cloud GPU (for example, RunPod; check the provider's site for the hourly price), measure the real load, then buy hardware to exact specs.
Professional (medium business, 20-100 people)
Tier 2 is recommended when:
- $5K+/month, steadily
- A production-critical AI use case
- You have a dedicated ops engineer
Config: SGLang + LiteLLM + Authentik SSO + Prometheus/Grafana. 2-4 servers for redundancy. Setup in 2-4 weeks, ops at 1 FTE part-time.
Enterprise (100-1,000+ users)
This is a different weight class. Tier 3 means separate infrastructure, a dedicated team, an SLA, disaster recovery, multi-region. It's beyond the scope of this lesson. See the NVIDIA / Anyscale / RunPod courses on enterprise AI infrastructure.
Anti-patterns: common mistakes
- ❌ Self-hosting with <$2K/month in LLM spend: no ROI. Hardware depreciates over 3+ years, and ops time costs money
- ❌ Ignoring electricity costs: every GPU running 24/7 eats $50-200/month in electricity + cooling
- ❌ No dedicated ops person: perpetual downtime. One person has to own the system, even part-time
- ❌ Skipping authentication: a security breach is inevitable. Handing the master_key to everyone = an invitation to disaster
- ❌ No monitoring: problems stay hidden until a production failure. At minimum, Prometheus + 3 key metrics
- ❌ Only one server: no redundancy = a single point of failure. Tier 1 can still be one box, Tier 2+ should always be redundant
- ❌ Running the latest unstable versions: vLLM/SGLang are evolving fast and breaking changes are common. Pin the version, test upgrades
- ❌ Ignoring model updates: models age. A new generation arrives every few months (Qwen 2.5 has already been replaced by Qwen3 and newer). Budget resources for migration
- ❌ Buying enterprise-tier hardware for a pilot: rent GPUs in the cloud for the first 1-2 months, then buy based on real metrics
Practice
Step 1: A pilot on a cloud GPU before buying hardware
Before buying a $50K server, rent a GPU in the cloud for 1-2 weeks and measure the real load.
# RunPod: the most convenient for a pilot
# Sign up at runpod.io, rent an A100 80GB (check the site for the hourly price)
# SSH into the pod
# Install Docker (if it isn't installed)
curl -fsSL https://get.docker.com | sh
# Run vLLM with the Qwen3 32B model
docker run --gpus all -p 8000:8000 \
-v ~/models:/root/.cache/huggingface \
vllm/vllm-openai:latest \
--model Qwen/Qwen3-32B \
--gpu-memory-utilization 0.90
# Test request
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-32B",
"messages": [{"role": "user", "content": "Explain RAG in three sentences"}]
}'Alternatives to RunPod: Lambda (lambda.ai), Anyscale (anyscale.com), Hyperstack, CoreWeave.
Step 2: A production setup on your own hardware
Once the pilot has confirmed the parameters, build production. The basic checklist:
# A server running Ubuntu Server (the current LTS version)
# Install NVIDIA drivers (pick the version to match CUDA and your vLLM version)
sudo apt update && sudo ubuntu-drivers install
# NVIDIA Container Toolkit for Docker GPU support (per the NVIDIA documentation)
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update && sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
# Check that the GPU is available in Docker (use a current nvidia/cuda image tag)
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
# Clone the stack (create your own repo or use a template)
mkdir -p /opt/internal-ai && cd /opt/internal-ai
# Copy docker-compose.yml from the theory section (Tier 1)
# Start it
docker compose up -d
# Watch the logs
docker compose logs -f vllmStep 3: Provisioning the first users
# Create an admin token through LiteLLM
MASTER_KEY=sk-master-internal-key-replace-me
# Create the development team
curl -X POST http://localhost:4000/team/new \
-H "Authorization: Bearer $MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"team_alias": "dev-team",
"max_budget": 500.0,
"budget_duration": "30d",
"models": ["qwen-32b"]
}'
# Save the team_id from the response
# Create a key for each developer (use pseudonymized IDs)
for user in dev_user_001 dev_user_002 dev_user_003; do
curl -X POST http://localhost:4000/key/generate \
-H "Authorization: Bearer $MASTER_KEY" \
-H "Content-Type: application/json" \
-d "{
\"team_id\": \"<team_id_from_response>\",
\"user_id\": \"$user\",
\"max_budget\": 20.0,
\"duration\": \"90d\",
\"rpm_limit\": 60,
\"tpm_limit\": 100000,
\"metadata\": {\"role\": \"developer\"}
}"
doneHand out the keys through 1Password or another secure channel. Each developer points their tool (Cursor, Continue.dev, Aider) at the internal endpoint.
Step 4: Basic monitoring
# Start Prometheus + Grafana
docker compose up -d prometheus grafana
# Open Grafana
# http://localhost:3001 (admin / admin-replace-me)
# Add data source → Prometheus → http://prometheus:9090
# Import a ready-made dashboard for vLLM
# There's an example dashboard in the vLLM documentation (the section on Prometheus and Grafana)
# Dashboards → Import → upload the JSON from the exampleKey alerts to set up on day one:
# alerts.yml for Prometheus
groups:
- name: vllm-critical
rules:
# Metric names have changed between vLLM versions: check against /metrics for your version
- alert: GPU_OOM_Risk
expr: vllm:gpu_cache_usage_perc > 0.95 # in newer versions kv_cache_usage_perc
for: 2m
annotations:
summary: "KV cache >95% full (risk of queueing and OOM)"
- alert: High_Latency
expr: histogram_quantile(0.95, sum(rate(vllm:time_to_first_token_seconds_bucket[5m])) by (le)) > 5
for: 5m
annotations:
summary: "P95 latency > 5 seconds"
- alert: vLLM_Down
expr: up{job="vllm"} == 0
for: 1m
annotations:
summary: "vLLM endpoint unavailable"Step 5: Fallback to the API (hybrid setup)
So your local infrastructure isn't a single point of failure, set up an automatic fallback to the Anthropic API.
litellm-config.yaml with a fallback:
model_list:
- model_name: smart-assistant
litellm_params:
model: openai/Qwen/Qwen3-32B
api_base: http://vllm:8000/v1
api_key: dummy
- model_name: smart-assistant-fallback
litellm_params:
# current model IDs: see the Anthropic documentation
model: anthropic/claude-sonnet-5-5
api_key: os.environ/ANTHROPIC_API_KEY
router_settings:
fallbacks:
- {"smart-assistant": ["smart-assistant-fallback"]}
context_window_fallbacks:
- {"smart-assistant": ["smart-assistant-fallback"]}
timeout: 30
num_retries: 2Now if vLLM goes down or gets overloaded, requests automatically go to Anthropic. Users won't notice the downtime.
Production readiness checklist (✅)
Tools and resources
Inference engines:
- vLLM: most popular, balanced
- SGLang: best throughput, advanced features
- TGI (HuggingFace): HF ecosystem (maintenance mode)
- Ollama: an easy start for pilots
API gateway & auth:
- LiteLLM: an OpenAI-compatible proxy with rate limiting + audit
- Authentik: an open-source IdP (SSO, SAML, OIDC)
- Keycloak: an enterprise IdP (Java stack)
UI:
- Open WebUI: a ChatGPT-like interface for your team
Fine-tuning:
- Axolotl: a flexible fine-tuning framework
- LLaMA-Factory: UI + CLI for fine-tuning
Cloud GPUs for pilots:
- Anyscale: managed Ray + serving
- RunPod: the most convenient pay-per-hour option
- Lambda: long-term contracts, hardware sales
- CoreWeave: enterprise-tier
- Hyperstack: competitive prices
Monitoring:
- Grafana: free tier for small teams
- Prometheus: the standard for metrics
Key takeaways
Self-hosting pays off at scale: $5K+/month in LLM spend or 50+ concurrent users. Below that threshold, the API is cheaper, simpler and more reliable. Don't buy a tractor for a vegetable garden.
A Tier 1 setup (1 server with two 24GB graphics cards + vLLM + LiteLLM + Open WebUI) takes a week to put together and costs roughly $10-15K in hardware plus electricity (a ballpark). That's enough for a team of 10-20 people.
LiteLLM is critical middleware that turns local inference into an enterprise-grade platform: per-user keys, rate limits, audit logs, cost tracking, automatic fallback to the API. Without it, a self-hosted setup stays an internal toy.
A hybrid setup (the main flow local + complex requests through the Anthropic API) is often better than pure self-hosting. LiteLLM routing makes it transparent to users.
Ops time is the main hidden cost of self-hosting. Budget 1-2 hours a week even for a stable setup. Without a dedicated owner, you get perpetual downtime.
What's next
→ AI Regulation & Compliance: which data and model requirements a team building an internal AI platform needs to take into account
The mark stays in this browser only and is never sent anywhere. My progress