Library · AI on your own computer and server

Self-hosted AI Enterprise Stack: vLLM, SGLang, Docker. An internal AI platform for your team

Engineer100 minUpdated: October 2026
79 of 105 in the library

Module: 13. Professional practice | Time: about 40 min theory + 60 min practice


The gist

Using the Anthropic API is like buying electricity from the city grid. Convenient, reliable, you pay by the meter. You don't have to think about transformers, wiring or the control room.

Self-hosted AI is like building your own power plant next to the building. Expensive up front: a turbine, cooling, an operator. But if you run a factory with 50+ machines going 24/7, your own power plant pays for itself in 1-2 years and makes you independent of the supplier.

When self-hosting pays off:

  • Steady LLM spend of $5,000+/month (the hardware break-even point)
  • 50+ concurrent users (real workload)
  • A regulated industry (compliance requires data residency)
  • Geopolitics (sanctions risk, sovereignty)
  • Custom models (fine-tuned for a domain)

In every other case, the API is cheaper, simpler and more reliable.

The thresholds and amounts in this lesson are benchmarks for doing the math, not a price list: hardware prices and plans change, so check before you buy. Current API prices: What's current.

🎨 Picture this: a fast-food restaurant vs a restaurant with its own farm. Fast food buys ingredients from a supplier: quick and cheap to start. Your own farm is a big investment in land and equipment, but after a few years the cost per plate is noticeably lower and you control quality from seed to plate.


🎯 Decision tree: do you need to self-host

Before you spend $50K on hardware, go through this checklist. A "no" at every node means stay on the API.

Code
LLM costs > $5,000/month, steadily, for the last 3 months?
→ Yes → Self-host breaks even in 12-18 months
→ No → The API is cheaper, leave the hardware alone

Regulated industry (healthcare, finance, gov, defense)?
→ Yes → Self-host required for compliance (HIPAA, PCI DSS, strict GDPR)
→ No → next question

Team of 50+ people using AI every day at the same time?
→ Yes → Self-host for cost + privacy + latency
→ No → The API is streamlined, don't overcomplicate it

Need a specific model (fine-tuned on domain data, custom)?
→ Yes → Self-host (cloud providers limit fine-tuning to a list of models and terms)
→ No → next question

Geographic / political restrictions (sanctions exposure, data sovereignty laws)?
→ Yes → Self-host = sovereignty insurance
→ No → The API works

All answers "no"?
→ Stay on the API. Self-hosting = overengineering for you.

🎨 Picture this: buying a tractor. If your garden is 1,000 sq ft, a shovel is cheaper and faster. A tractor makes sense from 25 acres up. Self-hosting is the tractor. For small operations it's just a way to go broke.


Key concepts

  • Inference engine: the program that takes a request and generates a response from an LLM (vLLM, SGLang, TGI, Ollama). It's the engine in the car
  • vLLM: the most popular open-source inference engine, Apache 2.0, with an excellent balance of throughput and ease of use
  • SGLang: one of the fastest in throughput, grew out of the LMSYS project (Berkeley), optimized for structured outputs and parallel sampling
  • API gateway: a proxy layer that turns local inference into an OpenAI-compatible endpoint (LiteLLM). Existing clients work without changes
  • Tensor parallelism: splitting a model across several GPUs. A 70B model needs ~140GB of VRAM in FP16, so one card won't do
  • Quantization: compressing a model to 4-bit or 8-bit to cut VRAM by 2-4x with minimal loss of quality
  • Throughput: how many tokens per second the system generates (important for multi-user setups)
  • Latency: the delay before the first token (TTFT, time to first token). Critical for interactive UX
  • Concurrent users: how many users can work at the same time without latency degrading

Theory

Hardware tiers for a company

The size of the hardware is determined by three parameters: model size (B parameters), number of concurrent users, and latency requirements.

Tier 1: Departmental (5-20 users)

The target profile: a development, marketing or support department at a mid-sized company. Non-critical load, an experiment or an internal tool.

Parameter Value
Hardware 1 server with 2x RTX 4090 (48GB total VRAM) or 1x A100 40GB
Model an open model in the 30-70B class (for example, Qwen3 32B; 70B in 4-bit quantization)
One-time hardware cost $8-15K
Electricity $50-100/month (~500W under load)
Concurrent users 20-30
Latency (TTFT) 1-3 sec
Throughput 30-60 tokens/sec per user
Setup time 2-4 days

Example: a team of 10 developers doing code review and docs generation. The build: two high-end consumer graphics cards with 24GB each, a server case, a Threadripper- or Xeon-class processor, 128GB RAM, a 4TB NVMe drive, a 1500W power supply. Roughly $10K total (a ballpark; check component prices when you buy).

🎨 Picture this: a compact generator at a vacation cabin. Enough for the house, not enough for the neighbors.

Tier 2: Team (20-100 users)

The target profile: a company where AI has become a production-critical tool. A SaaS startup with built-in AI, an agency with 50 developers, a mid-size fintech.

Parameter Value
Hardware 2-4 servers with 4x A100 80GB or 2x H100 80GB per server
Model a 70B-class model at full precision or larger (including MoE models)
One-time cost $50-150K
Electricity + cooling $300-800/month
Concurrent users 100+
Latency (TTFT) 0.5-2 sec
Throughput 80-150 tokens/sec per user
Setup time 1-2 weeks

At this level you already need a server room with temperature control (HVAC), a 5-10 kW UPS and dedicated network infrastructure. You need a dedicated ops engineer, at least part-time.

Tier 3: Enterprise (100-1,000+ users)

The target profile: a large company, a bank, a government agency. AI is core infrastructure.

Parameter Value
Hardware a cluster of 8-16 servers with H100 80GB / B200
Model Custom fine-tuned 70B+, possibly several at once
One-time cost $500K to $5M
OpEx $5-50K/month (electricity + cooling + ops team)
Concurrent users 1,000+
Latency (TTFT) <500ms
Setup time 2-3 months

This is data center territory. You need a team of 3-5 people: SRE, ML engineer, security, network. An SLA of 99.9%+ requires redundancy at every level.


Inference engines compared (as of October 2026)

The inference engine is the heart of the system. Your choice determines throughput, latency and operational complexity. A realistic shortlist of 5 options:

Engine Best for License Throughput Setup time When to choose it
vLLM Most popular, balanced Apache 2.0 High 1-2 days The default choice. Big community, lots of guides
SGLang Best throughput Apache 2.0 Highest 2-3 days When you need maximum performance, structured outputs
TGI (HuggingFace) HF ecosystem Apache 2.0 High 1 day The project is in maintenance mode: for a new deployment, go with vLLM or SGLang
Ollama Easiest, small scale MIT Low 30 minutes Pilot, prototype, up to 10 users
TensorRT-LLM NVIDIA only, fastest see the repository Highest on NVIDIA 1 week When it's NVIDIA only and you need the maximum

vLLM (github.com/vllm-project/vllm) is the gold standard. PagedAttention is its signature technique for using VRAM efficiently. An OpenAI-compatible API out of the box. Most tutorials and production case studies are built on vLLM.

SGLang (github.com/sgl-project/sglang) beats vLLM on throughput in a number of benchmarks, but measure it on your own workload. RadixAttention reuses the KV cache across requests. It's especially good when you have lots of similar system prompts (a typical agent scenario). The downside: the ecosystem is younger, with fewer ready-made guides.

TGI (Text Generation Inference) (github.com/huggingface/text-generation-inference) is a HuggingFace product. It integrates well with the HF Hub and is simple to deploy. But as of October 2026 the project is in maintenance mode: only minor fixes are accepted, and the authors themselves recommend vLLM and SGLang. For a new project, it's better to start with those.

Ollama is for pilots and small teams. It starts with one command and has a nice CLI. It doesn't scale beyond ~10 concurrent users. Details: Local AI models: Ollama, LM Studio and private AI.

🎨 Picture this: choosing an engine for a car. vLLM is a Toyota (reliable, affordable, everything works). SGLang is a Porsche (faster, but takes experience). Ollama is a moped (starts instantly, won't take you far).


A Docker stack for team deployment

A production setup is made of 3-4 containers: inference engine + gateway + UI + auth. The basic configuration for Tier 1:

yaml
# docker-compose.yml — Tier 1 team setup (5-20 users)
# Modern Docker Compose doesn't need the version key

services:
  # Inference engine
  vllm:
    image: vllm/vllm-openai:latest
    runtime: nvidia
    ports:
      - "8000:8000"
    volumes:
      - ./models:/root/.cache/huggingface
    command:
      # The bf16 model weighs ~65GB: on 2×24GB use a quantized version (AWQ or GPTQ),
      # on A100/H100 80GB use it as is
      - --model
      - Qwen/Qwen3-32B
      - --tensor-parallel-size
      - "2"
      - --gpu-memory-utilization
      - "0.90"
      - --max-model-len
      - "32768"
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 2
              capabilities: [gpu]
    restart: unless-stopped

  # API gateway: an OpenAI-compatible proxy with rate limiting + audit
  litellm:
    # Pin a specific version instead of main-latest (see below about March 2026)
    image: ghcr.io/berriai/litellm:main-latest
    ports:
      - "4000:4000"
    environment:
      - DATABASE_URL=postgres://litellm:secret@postgres:5432/litellm
      - MASTER_KEY=sk-master-internal-key-replace-me
    volumes:
      - ./litellm-config.yaml:/app/config.yaml
    command: ["--config", "/app/config.yaml", "--port", "4000"]
    depends_on:
      - postgres
      - vllm
    restart: unless-stopped

  # PostgreSQL for LiteLLM (audit, keys, budgets)
  postgres:
    image: postgres:16
    environment:
      - POSTGRES_USER=litellm
      - POSTGRES_PASSWORD=secret
      - POSTGRES_DB=litellm
    volumes:
      - postgres-data:/var/lib/postgresql/data
    restart: unless-stopped

  # User-facing UI
  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    ports:
      - "3000:8080"
    environment:
      - OPENAI_API_BASE_URL=http://litellm:4000/v1
      - OPENAI_API_KEY=sk-master-internal-key-replace-me
      - WEBUI_AUTH=true
    volumes:
      - open-webui:/app/backend/data
    depends_on:
      - litellm
    restart: unless-stopped

volumes:
  open-webui:
  postgres-data:

The LiteLLM config (litellm-config.yaml):

yaml
model_list:
  - model_name: qwen-32b
    litellm_params:
      model: openai/Qwen/Qwen3-32B
      api_base: http://vllm:8000/v1
      api_key: dummy  # vLLM doesn't require real auth

general_settings:
  master_key: sk-master-internal-key-replace-me
  database_url: postgres://litellm:secret@postgres:5432/litellm

litellm_settings:
  drop_params: true
  set_verbose: false
  json_logs: true
  cache: true

Starting it up:

bash
docker compose up -d
# vLLM will download the Qwen3 32B model (~65GB in bf16) on first start
# Watch the progress: docker compose logs -f vllm

After startup:

  • http://localhost:8000/v1: the raw vLLM endpoint
  • http://localhost:4000/v1: the LiteLLM gateway (use this one)
  • http://localhost:3000: Open WebUI for users

🎨 Picture this: a Lego set. Each container is a brick with one function. vLLM does the computing, LiteLLM controls access, Postgres remembers who did what, WebUI provides the interface.


An OpenAI-compatible API gateway: LiteLLM

LiteLLM (github.com/BerriAI/litellm) is a critical component of the stack. Without it, a self-hosted setup stays an internal toy; with it, it becomes an enterprise-grade platform.

What LiteLLM gives you:

  • An OpenAI-compatible API: every client that works with OpenAI/Anthropic (Cursor, Continue.dev, Aider, custom scripts) works with your local vLLM without rewriting code
  • Per-user API keys: each developer gets their own key. Revoking one doesn't break the rest
  • Rate limiting: per-user or per-team limits so one script can't hog all the hardware
  • Cost tracking + budgets: who spent how many tokens, alerts when a limit is exceeded
  • Audit log: every request in PostgreSQL with user_id, model, tokens, latency
  • Multi-model routing: you can route simple requests to a small model and complex ones to a big one
  • Fallback to API: if the local vLLM goes down, you can proxy to the Anthropic API

Security of the gateway itself: on March 24, 2026, malicious versions of the litellm package (1.82.7 and 1.82.8) briefly appeared on PyPI and stole credentials. The authors' write-up: Security Update: Suspected Supply Chain Incident. The lesson for any gateway that holds keys: pin the version, install only from trusted sources and signed images, and update deliberately, not automatically.

Creating a user key through the LiteLLM admin:

bash
# Create a team
curl -X POST http://localhost:4000/team/new \
  -H "Authorization: Bearer sk-master-internal-key-replace-me" \
  -H "Content-Type: application/json" \
  -d '{
    "team_alias": "backend-team",
    "max_budget": 100.0,
    "models": ["qwen-32b"]
  }'

# Create a key for a developer
curl -X POST http://localhost:4000/key/generate \
  -H "Authorization: Bearer sk-master-internal-key-replace-me" \
  -H "Content-Type: application/json" \
  -d '{
    "team_id": "backend-team",
    "user_id": "dev_user_42",
    "max_budget": 10.0,
    "duration": "30d",
    "rpm_limit": 60,
    "tpm_limit": 100000
  }'
# Response: {"key": "sk-1a2b3c4d...", "expires": "..."}

The developer uses their key in any OpenAI-compatible tool:

bash
# Aider (the openai/ prefix tells it this is an OpenAI-compatible endpoint)
export OPENAI_API_BASE=http://internal-llm:4000/v1
export OPENAI_API_KEY=sk-1a2b3c4d...
aider --model openai/qwen-32b

# Cursor: set a Custom API in the settings
# Continue.dev in VS Code: config.yaml (config.json is deprecated):
# models:
#   - name: Internal Qwen
#     provider: openai
#     model: qwen-32b
#     apiBase: http://internal-llm:4000/v1
#     apiKey: sk-1a2b3c4d...
# See the Continue documentation for the exact format.

Authentication & access control

Handing users the master_key is like giving every employee the key to the server room. You need SSO + per-user provisioning.

Authentik (goauthentik.io) is an open-source IdP and a direct competitor to Okta. Self-hosted, supports SAML, OAuth2, OIDC. Integrates with Google Workspace, Microsoft 365, LDAP.

Keycloak is the older sibling, enterprise-tested, but harder to set up. If the company already runs a Java stack, it's the natural choice.

A typical flow:

Code
An employee signs in to Open WebUI
  ↓
Open WebUI redirects to Authentik
  ↓
Authentik checks via Google Workspace SSO
  ↓
Authentik returns a JWT with user_id, groups
  ↓
Open WebUI creates a session
  ↓
Open WebUI calls LiteLLM with a user-specific API key
  ↓
LiteLLM checks the key and limits, logs it, passes it to vLLM
  ↓
vLLM generates the response

Key controls:

  • Groups → models: junior developers get only the small model, seniors get access to the big one (70B)
  • Per-user budgets: a small default budget per user, increased on request through a manager
  • Rate limits: 60 req/min by default, so a stray bash script in a loop can't take the system down
  • Audit logging: every request is written to Postgres, with 90-day retention (or longer for compliance)

Monitoring stack

Without monitoring, a self-hosted setup turns into a black box. When users start complaining "it's slow," you need to see the metrics right away.

The minimum stack:

  • Prometheus: collects metrics. vLLM exposes a /metrics endpoint natively
  • Grafana: dashboards. Ready-made templates for vLLM are at grafana.com/dashboards
  • Loki: log aggregation (optional, for big setups)
  • OpenTelemetry: distributed tracing (optional, for multi-service setups)

Metrics to track from day 1:

  • GPU utilization (% — underused = overpaying, overloaded = a queue)
  • VRAM usage (% — approaching 95% = OOM coming soon)
  • Tokens/sec (aggregate throughput)
  • Time to first token (P50, P95, P99 — the UX metric)
  • Queue depth (if it's growing, you need more hardware)
  • Cost per user (LiteLLM exports it)
  • Error rate (5xx, timeouts)

Extending docker-compose for monitoring:

yaml
  prometheus:
    image: prom/prometheus:latest
    ports:
      - "9090:9090"
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml
    restart: unless-stopped

  grafana:
    image: grafana/grafana:latest
    ports:
      - "3001:3000"
    environment:
      - GF_SECURITY_ADMIN_PASSWORD=admin-replace-me
    volumes:
      - grafana-data:/var/lib/grafana
    restart: unless-stopped

prometheus.yml:

yaml
global:
  scrape_interval: 15s

scrape_configs:
  - job_name: vllm
    static_configs:
      - targets: ['vllm:8000']
    metrics_path: /metrics

  - job_name: litellm
    static_configs:
      - targets: ['litellm:4000']
    metrics_path: /metrics

🎨 Picture this: the instrument panel in an airplane. Without it, the pilot finds out about a problem when the engine is already smoking. With it, 10 minutes before that.


Fine-tuning your own model: an extra option

If the company has domain-specific data (legal documents, medical protocols, code in an internal DSL), fine-tuning can give quality a significant boost.

Baseline numbers (ballparks, check them against your task):

  • Hardware: 1x A100 80GB for a small fine-tune (7B-13B), 4x for a large one (70B)
  • Data: 1K-10K high-quality examples for a good result
  • Tools: Axolotl (github.com/axolotl-ai-cloud/axolotl), LLaMA-Factory (github.com/hiyouga/LLaMA-Factory)
  • Time: 2-24 hours depending on model size and data volume
  • Cost: $100-1,000 when renting GPUs (RunPod, Lambda Labs)

Fine-tuning details are in the lesson Fine-tuning: when prompts aren't enough. What matters here: after fine-tuning, the model is deployed to the same vLLM/SGLang as the base model. Only the --model parameter changes, to the path of the fine-tuned checkpoint.


Real-world setups: 3 case studies

These are generalized illustrations, not reports from specific companies; the amounts are illustrative. Check data requirements (GDPR, medical and banking regulations) with a lawyer: this lesson is not a substitute for legal advice.

Case 1: A regional fintech (50 developers)

Context: a mid-sized bank building core banking software; data residency requirements in its home country rule out foreign cloud APIs.

Stack:

  • 4x A100 80GB (1 server with tensor parallel 4)
  • An open 70B-class model (the Qwen family) + LiteLLM + Open WebUI + Authentik
  • Uses: automated code review, documentation generation, security analysis on pull requests

Economics:

  • Hardware: ~$200K (servers + networking + UPS)
  • OpEx: $400/month in electricity, 0.5 FTE ops engineer
  • The alternative (Anthropic): ~$50K/year at the current load
  • Break-even: ~4.4 years even without the ops engineer's salary ($200K / ($50K − $4.8K electricity) a year)
  • Sovereignty: 100%, no data leaves the perimeter

Case 2: European healthcare SaaS

Context: a SaaS for clinics in Germany, GDPR + national medical data law. No request containing patient data can go to the US/UK.

Stack:

  • 2 servers with 2x H100 80GB each (redundancy)
  • SGLang + an open 70B-class model, fine-tuned on anonymized medical transcripts
  • Authentik SSO integrated with the existing Active Directory
  • Uses: an assistant for doctors (research, summarization), drafts for patient communication (always under supervision)

Economics:

  • Hardware: €300K
  • OpEx: €600/month for infrastructure + 1 FTE ops
  • Compliance: passed a GDPR audit, medical data certification
  • Risk reduction (the potential fine for a data breach): millions of euros

Case 3: Latin American media company

Context: a media group in Mexico City publishing in Spanish, Portuguese, English and Quechua. Large volumes of translation and rewriting.

Stack:

  • 2x RTX 4090 (one server)
  • Qwen3 32B (quantized) + Open WebUI + LiteLLM
  • Uses: a translation pipeline for 5 languages, generating article drafts, A/B headline variants

Economics:

  • Hardware: $12K one-time
  • OpEx: $80/month in electricity
  • The alternative before self-hosting: $1,500/month (DeepL + OpenAI)
  • Break-even: ~8.5 months ($12K / ($1,500 − $80) a month)
  • An extra win: the ability to fine-tune for Latin American regional Spanish (which commercial translators don't account for)

Operational concerns: what breaks in production

Self-hosting isn't "set it and forget it." Here's the real list of things that need attention every week:

  • Uptime: 99.9% requires a redundant setup (2 servers minimum, automatic failover)
  • GPU failures: 1-2 cards a year die under 24/7 load. Keep spares
  • Cooling: the server room has to stay at 64-72°F (18-22°C). Overheating = accelerated GPU wear
  • Power: a UPS with 15+ minutes for a graceful shutdown, a generator for production
  • Model updates: new versions of Qwen/Llama/Mistral come out every quarter. Updating a model = 1-2 hours of downtime
  • OS / driver updates: NVIDIA drivers need care (incompatibilities with vLLM versions)
  • Backup: models weigh 50-200GB, so you need a plan for storing checkpoints
  • Maintenance time: budget 1-2 hours of ops time a week, even for a stable setup

🎨 Picture this: a fish tank. You bought the fish, and now you have to feed them, clean the tank, keep the temperature right, give them medicine if they get sick. It's not "set it up and admire it."


Cost comparison: 1 year in detail

A comparison for a team of 50 developers with ~$5K/month in usage if they were on the API:

Approach Year 1 cost Year 2 cost Sovereignty Flexibility
Anthropic API ($5K/month) $60K $66K (usage growth) ❌ Vendor lock ✅ Any model
Self-host Tier 1 (Qwen 32B) $20K hardware + $1.2K ops $1.2K ops ✅ Full ⚠ One model
Self-host Tier 2 (Llama 70B + redundancy) $80K + $5K ops $5K ops ✅ Full ✅ Several models
Hybrid (API primary + local fallback) $40K (mix) $36K (balancing) ⚠ Partial ✅ Best of both

Hybrid often turns out to be the sweet spot: most requests go through the local vLLM (cheap, sovereign), and complex requests go through the Anthropic API (when you need a strong cloud model). LiteLLM can route them automatically.


Audience: how well this applies

Beginner (1 developer, personal project)

You don't need to self-host. Come back when your LLM bills are steadily > $300/month. Until then, use the Anthropic/OpenAI API: saving time matters more than saving money.

If you want to play with local models to learn, use Ollama (see the lesson Local AI models: Ollama, LM Studio and private AI); it's up and running in 10 minutes.

Intermediate (small team, 5-20 people)

A Tier 1 setup makes sense if:

  • You steadily spend $1,500+/month on LLMs
  • Or you have at least one requirement: privacy, sovereignty, latency

Config: vLLM + LiteLLM + Open WebUI. One server with 2x RTX 4090 or 1x A100. Setup in a week, ops 2-4 hours a week.

Pilot before production: run it for 2 weeks on a cloud GPU (for example, RunPod; check the provider's site for the hourly price), measure the real load, then buy hardware to exact specs.

Professional (medium business, 20-100 people)

Tier 2 is recommended when:

  • $5K+/month, steadily
  • A production-critical AI use case
  • You have a dedicated ops engineer

Config: SGLang + LiteLLM + Authentik SSO + Prometheus/Grafana. 2-4 servers for redundancy. Setup in 2-4 weeks, ops at 1 FTE part-time.

Enterprise (100-1,000+ users)

This is a different weight class. Tier 3 means separate infrastructure, a dedicated team, an SLA, disaster recovery, multi-region. It's beyond the scope of this lesson. See the NVIDIA / Anyscale / RunPod courses on enterprise AI infrastructure.


Anti-patterns: common mistakes

  • ❌ Self-hosting with <$2K/month in LLM spend: no ROI. Hardware depreciates over 3+ years, and ops time costs money
  • ❌ Ignoring electricity costs: every GPU running 24/7 eats $50-200/month in electricity + cooling
  • ❌ No dedicated ops person: perpetual downtime. One person has to own the system, even part-time
  • ❌ Skipping authentication: a security breach is inevitable. Handing the master_key to everyone = an invitation to disaster
  • ❌ No monitoring: problems stay hidden until a production failure. At minimum, Prometheus + 3 key metrics
  • ❌ Only one server: no redundancy = a single point of failure. Tier 1 can still be one box, Tier 2+ should always be redundant
  • ❌ Running the latest unstable versions: vLLM/SGLang are evolving fast and breaking changes are common. Pin the version, test upgrades
  • ❌ Ignoring model updates: models age. A new generation arrives every few months (Qwen 2.5 has already been replaced by Qwen3 and newer). Budget resources for migration
  • ❌ Buying enterprise-tier hardware for a pilot: rent GPUs in the cloud for the first 1-2 months, then buy based on real metrics

Practice

Step 1: A pilot on a cloud GPU before buying hardware

Before buying a $50K server, rent a GPU in the cloud for 1-2 weeks and measure the real load.

bash
# RunPod: the most convenient for a pilot
# Sign up at runpod.io, rent an A100 80GB (check the site for the hourly price)
# SSH into the pod

# Install Docker (if it isn't installed)
curl -fsSL https://get.docker.com | sh

# Run vLLM with the Qwen3 32B model
docker run --gpus all -p 8000:8000 \
  -v ~/models:/root/.cache/huggingface \
  vllm/vllm-openai:latest \
  --model Qwen/Qwen3-32B \
  --gpu-memory-utilization 0.90

# Test request
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-32B",
    "messages": [{"role": "user", "content": "Explain RAG in three sentences"}]
  }'

Alternatives to RunPod: Lambda (lambda.ai), Anyscale (anyscale.com), Hyperstack, CoreWeave.


Step 2: A production setup on your own hardware

Once the pilot has confirmed the parameters, build production. The basic checklist:

bash
# A server running Ubuntu Server (the current LTS version)
# Install NVIDIA drivers (pick the version to match CUDA and your vLLM version)
sudo apt update && sudo ubuntu-drivers install

# NVIDIA Container Toolkit for Docker GPU support (per the NVIDIA documentation)
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | \
  sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
  sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
  sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update && sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

# Check that the GPU is available in Docker (use a current nvidia/cuda image tag)
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

# Clone the stack (create your own repo or use a template)
mkdir -p /opt/internal-ai && cd /opt/internal-ai
# Copy docker-compose.yml from the theory section (Tier 1)

# Start it
docker compose up -d

# Watch the logs
docker compose logs -f vllm

Step 3: Provisioning the first users

bash
# Create an admin token through LiteLLM
MASTER_KEY=sk-master-internal-key-replace-me

# Create the development team
curl -X POST http://localhost:4000/team/new \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "team_alias": "dev-team",
    "max_budget": 500.0,
    "budget_duration": "30d",
    "models": ["qwen-32b"]
  }'
# Save the team_id from the response

# Create a key for each developer (use pseudonymized IDs)
for user in dev_user_001 dev_user_002 dev_user_003; do
  curl -X POST http://localhost:4000/key/generate \
    -H "Authorization: Bearer $MASTER_KEY" \
    -H "Content-Type: application/json" \
    -d "{
      \"team_id\": \"<team_id_from_response>\",
      \"user_id\": \"$user\",
      \"max_budget\": 20.0,
      \"duration\": \"90d\",
      \"rpm_limit\": 60,
      \"tpm_limit\": 100000,
      \"metadata\": {\"role\": \"developer\"}
    }"
done

Hand out the keys through 1Password or another secure channel. Each developer points their tool (Cursor, Continue.dev, Aider) at the internal endpoint.


Step 4: Basic monitoring

bash
# Start Prometheus + Grafana
docker compose up -d prometheus grafana

# Open Grafana
# http://localhost:3001 (admin / admin-replace-me)
# Add data source → Prometheus → http://prometheus:9090

# Import a ready-made dashboard for vLLM
# There's an example dashboard in the vLLM documentation (the section on Prometheus and Grafana)
# Dashboards → Import → upload the JSON from the example

Key alerts to set up on day one:

yaml
# alerts.yml for Prometheus
groups:
  - name: vllm-critical
    rules:
      # Metric names have changed between vLLM versions: check against /metrics for your version
      - alert: GPU_OOM_Risk
        expr: vllm:gpu_cache_usage_perc > 0.95  # in newer versions kv_cache_usage_perc
        for: 2m
        annotations:
          summary: "KV cache >95% full (risk of queueing and OOM)"

      - alert: High_Latency
        expr: histogram_quantile(0.95, sum(rate(vllm:time_to_first_token_seconds_bucket[5m])) by (le)) > 5
        for: 5m
        annotations:
          summary: "P95 latency > 5 seconds"

      - alert: vLLM_Down
        expr: up{job="vllm"} == 0
        for: 1m
        annotations:
          summary: "vLLM endpoint unavailable"

Step 5: Fallback to the API (hybrid setup)

So your local infrastructure isn't a single point of failure, set up an automatic fallback to the Anthropic API.

litellm-config.yaml with a fallback:

yaml
model_list:
  - model_name: smart-assistant
    litellm_params:
      model: openai/Qwen/Qwen3-32B
      api_base: http://vllm:8000/v1
      api_key: dummy

  - model_name: smart-assistant-fallback
    litellm_params:
      # current model IDs: see the Anthropic documentation
      model: anthropic/claude-sonnet-5-5
      api_key: os.environ/ANTHROPIC_API_KEY

router_settings:
  fallbacks:
    - {"smart-assistant": ["smart-assistant-fallback"]}
  context_window_fallbacks:
    - {"smart-assistant": ["smart-assistant-fallback"]}
  timeout: 30
  num_retries: 2

Now if vLLM goes down or gets overloaded, requests automatically go to Anthropic. Users won't notice the downtime.


Production readiness checklist (✅)


Tools and resources

Inference engines:

API gateway & auth:

  • LiteLLM: an OpenAI-compatible proxy with rate limiting + audit
  • Authentik: an open-source IdP (SSO, SAML, OIDC)
  • Keycloak: an enterprise IdP (Java stack)

UI:

  • Open WebUI: a ChatGPT-like interface for your team

Fine-tuning:

Cloud GPUs for pilots:

Monitoring:


Key takeaways

Self-hosting pays off at scale: $5K+/month in LLM spend or 50+ concurrent users. Below that threshold, the API is cheaper, simpler and more reliable. Don't buy a tractor for a vegetable garden.

A Tier 1 setup (1 server with two 24GB graphics cards + vLLM + LiteLLM + Open WebUI) takes a week to put together and costs roughly $10-15K in hardware plus electricity (a ballpark). That's enough for a team of 10-20 people.

LiteLLM is critical middleware that turns local inference into an enterprise-grade platform: per-user keys, rate limits, audit logs, cost tracking, automatic fallback to the API. Without it, a self-hosted setup stays an internal toy.

A hybrid setup (the main flow local + complex requests through the Anthropic API) is often better than pure self-hosting. LiteLLM routing makes it transparent to users.

Ops time is the main hidden cost of self-hosting. Budget 1-2 hours a week even for a stable setup. Without a dedicated owner, you get perpetual downtime.


What's next

→ AI Regulation & Compliance: which data and model requirements a team building an internal AI platform needs to take into account

The mark stays in this browser only and is never sent anywhere. My progress