Library · AI on your own computer and server

Local AI agents: minimal agents, OpenClaw, Hermes, Qwen-Agent

Builder65 minUpdated: October 2026
78 of 105 in the library

Module: 13. Professional practice | Time: ~35 min theory + 30 min practice


The gist

Claude Code is like electricity from the grid. You pay a subscription and the lights come on (current prices and versions: What's current). It's convenient and reliable, but the power only flows while the bill is paid and the internet is up.

By October 2026 there are plenty of mature frameworks that let you run agent systems entirely on your own computer. No API. No internet. With tool use, memory, MCP servers and multi-agent collaboration.

The lesson Local AI models: Ollama, LM Studio and private AI covered basic local models (Ollama, LM Studio): those were the "light bulbs." This lesson is about a full "power plant": when a model + an agent framework + tools work together and solve real tasks without sending a single byte outside.

We'll go through 8 key frameworks, examples of models for agent tasks, the hardware requirements, and an honest bottom line: where local wins and where it loses to an API.

🎨 Picture this: the Claude API is electricity from the city grid. You pay for every kWh, everything is stable, but you depend on the utility. Local agents are solar panels on your roof. You install them once and generate your own power from then on. They cover household needs (90% of tasks). But when you run industrial machinery (the level of the strongest cloud models), you switch to the grid. It's not "either-or," it's "both," for different tasks.


🎯 Decision tree: when local beats an API

The main question isn't "which is cooler," it's "which fits your task":

Code
→ Production B2B with an SLA, support, uptime guarantees?
   → API (Anthropic/OpenAI). Local isn't built for this.

→ Working with confidential data (healthcare, legal, corporate code)?
   → LOCAL, required. Lawyers and doctors don't send client data outside.

→ Personal automation without the internet (a home bot, notes, translations)?
   → LOCAL. No point paying for private tasks.

→ Hobby projects and experiments?
   → LOCAL. Free and no limits.

→ High volume + simple tasks (classification, translating thousands of records)?
   → LOCAL is cheaper. An API will eat your budget fast.

→ Complex reasoning, ADRs, architecture decisions, complex code?
   → API (strong cloud models). Local isn't there yet.

→ Access to cloud AI APIs is limited where you or your client are, or you can't depend on a vendor's policies?
   → LOCAL = independence. No one can switch it off. (Not every country is on the supported-countries lists of Anthropic and OpenAI.)

→ Not sure whether you need maximum performance or savings?
   → Hybrid. Local for routine work, an API for the hard stuff.

🎨 Picture this: don't choose between a bike and a car: keep both. Ride the bike to the store (local), drive the car to another city (API). Most builders run a mixed stack.


Key concepts

  • Local agent framework: software that turns a local model (Llama, Qwen, Mistral) into a full agent with tools, memory and planning
  • Function calling: a model's ability not just to answer in text but to call tools with the right arguments. A basic ability of agent models
  • MCP (Model Context Protocol): an open protocol for connecting tools to models. It started at Anthropic and is now supported by local frameworks
  • Inference backend: the engine that runs the model: Ollama (simple), vLLM (fast, production), SGLang (max throughput), llama.cpp (minimal dependencies)
  • Quantization: compressing a model to run on smaller hardware. Q4_K_M lets a 70B model fit in 40GB instead of 140GB
  • Open weights: the model is published with its weights (you can download and run it). Not the same as open source (open training code)
  • Tool-native model: a model trained to call functions out of the box. Modern families (Qwen3, gpt-oss, Hermes 4) are tool-native. Older base models without a separate fine-tune often aren't
  • Hybrid stack: a local + API combination: routine work locally, hard tasks through an API
  • MoE (Mixture of Experts): an architecture where only some of the parameters are active for each request. For example, DeepSeek-V3 has 671B parameters with 37B active; Qwen3-Coder 30B activates about 3.3B. Big-model quality at mid-size speed

Theory

When local agents beat an API

Local isn't a replacement for an API. It's a different tool with different economics. Three situations where local really wins:

1. Privacy and sovereignty. Medical records, corporate code, personal notes: this is data that shouldn't leave your computer. APIs and subscriptions have different terms on training with your data, and a client's lawyers may say "no" regardless. Local solves this at the architecture level: the data physically never leaves.

2. Cost at high volume. If you have tens of thousands of text classification tasks a month, your API bill grows with the volume, while a local model in the 7–14B class costs only electricity. Do the math with your own volumes: token prices are on the What's current page.

3. Independence. Sanctions, outages, policy changes. For builders in some countries this isn't a hypothetical risk: not every country is on Anthropic's and OpenAI's supported-countries lists, and access to their APIs is limited there. Local always works.

When local loses: complex reasoning (the strongest cloud models), high-end multimodal, long context (hundreds of thousands of tokens and up), the newest features (computer use, voice). Cloud models are usually ahead here.

🎨 Picture this: local is your vegetable garden. The tomatoes are fresh, free and yours. But you can't grow mangoes in your backyard: for mangoes, you go to the supermarket (the API).


Hardware requirements: what you actually need

The main limit of local AI is hardware. The size of the model determines the minimum RAM/VRAM:

Scenario Mac Windows/Linux PC Budget
Starter (7–8B models) A Mac with an M-series chip, 16 GB of memory a graphics card with 12 GB of memory or more minimal
Standard (13–14B + tools) A Pro/Max-class Mac, 32 GB a 16 GB graphics card medium
Advanced (30–34B + agent stack) A Max-class Mac, 64 GB a 24 GB graphics card or two with 12–16 GB high
Professional (70B+) Mac Studio, 128–192 GB a 48 GB graphics card or two with 24 GB very high

Hardware prices fluctuate a lot, so the table has no dollar amounts: check with sellers before you buy. The specific graphics card and chip models change every year; what matters is the amount of memory.

Important rules:

  • A 70B-parameter model with Q4 quantization needs ~40GB of RAM/VRAM. On 32GB it'll run slowly (swap), and on 16GB it won't run at all
  • The Mac advantage: unified memory. An M2 Max with 64GB behaves like 64GB of VRAM. On a PC, RAM and GPU VRAM are separate
  • Older GPUs (GTX 1080, RTX 2060) work, but slowly. For a real-time agent you need a graphics card with at least 12 GB of memory
  • Electricity: a powerful graphics card at full load draws 300–450 W. At 8 hours a day, that's a noticeable addition to your bill; calculate it with your own rate

🎨 Picture this: if you want to cook at home instead of eating out, you need a stove. A camping burner will boil pasta. An induction cooktop will handle a steak. A professional kitchen can do a banquet for 50. Figure out the menu first, then buy the kitchen.


8 key local agent frameworks (as of October 2026)

Let's go through them in order. Each has its own niche.


Framework 1: nanoClaude: a minimal agent for learning

What: Minimal Claude Code-style agent implementations that fit in a small amount of code. It's not one project but a whole genre: GitHub has several independent community projects (for example, nanoclaude and nano-claude-code). The idea is close to Karpathy's nanoGPT. Don't confuse it with NanoClaw: that's a different project, a lightweight containerized alternative to OpenClaw.

Best for: Education. Understanding how an agent works from the inside, without framework wrappers.

Runs on: Usually Python. They connect to a cloud model or to a local one through Ollama (see the specific project's README).

Cost: $0.

Status: Community-built; quality and maintenance vary. Not for production: it's a learning tool.

Source: https://github.com/karpathy/nanoGPT (the idea), https://github.com/CohleM/nanoclaude (an example of a learning agent).

🎨 Picture this: a Lego set. You build it yourself and understand every brick. Not for a house you'd live in, but for understanding "how houses are built."


Framework 2: OpenClaw: an open-source personal agent

What: An open-source, self-hosted personal AI agent (MIT license). It lives on your computer, talks to you through messaging apps (WhatsApp, Telegram, Discord, Slack and others), remembers past conversations, and can work with your email, calendar, browser and files. It works with both cloud models (Claude, GPT) and local ones. The project is run by the nonprofit OpenClaw Foundation, and there's no paid version.

Best for: People who want a personal agent under their own control and don't want to hand their data to someone else's service. For coding work, Claude Code, Codex and IDE agents are a better fit than OpenClaw.

Runs on: Mac, Windows, Linux. Installed with a script, through npm, or as a desktop app. The backend is any model, including a local one through Ollama.

Cost: $0 (open source); you only pay for the cloud model you choose, if you use one.

Status: Actively developed; check the official website for the current version.

Source: https://openclaw.ai and https://github.com/openclaw/openclaw.

🎨 Picture this: a personal assistant who lives at your place. You text them like a real person, and they open the email, calendar and files on your computer. Full control, but also full responsibility: grant access carefully, as little as possible.


Framework 3: Hermes: Nous Research's models and agent

What: Under the Hermes name, Nous Research releases two related things. First, the Hermes family of fine-tuned models (the current line is Hermes 4), trained on function calling and agent tasks. Second, Hermes Agent, an open-source agent framework (MIT license) with persistent memory, messaging app support and delegation to subagents. You can connect the agent to your own model or to Nous Portal.

Models: Hermes 4 on Hugging Face: 14B, 36B (version 4.3), 70B and 405B (405B needs a server GPU).

Best for: Function calling without an API, multi-step reasoning offline. When you need a model that calls tools correctly out of the box.

Runs on: Ollama / vLLM / TGI. Any framework that supports models of the corresponding family.

Cost: $0 (models and agent) + electricity for compute.

Source: https://huggingface.co/NousResearch and https://hermes-agent.nousresearch.com.

🎨 Picture this: an ordinary base model is a generalist. Hermes is a model that went through a "how to work with tools" course. The difference is like a rookie vs. an experienced server: both will take your order, but one will forget half of it.


Framework 4: Qwen-Agent: Alibaba's native framework

What: An agent framework built by the Qwen team for its own models. Tight integration with Qwen models (Qwen3 and newer). Supports function calling, MCP, a code interpreter and RAG.

Models: Qwen3 (from 0.6B to 235B parameters) and Qwen3-Coder (30B and 480B). Check the project page for newer Qwen generations.

Best for: Complex coding tasks locally, multi-step workflows. If the task is "write and test a function," Qwen3-Coder is the first one to try.

Special: The framework itself is under the Apache 2.0 license, and many Qwen models also have open weights and Apache 2.0; check the specific model's license.

Runs on: Ollama / SGLang / vLLM. Native support.

Cost: $0.

Source: https://github.com/QwenLM/Qwen-Agent.

🎨 Picture this: a workhorse: not the prettiest vehicle, not the loudest marketing videos. But it keeps going where its flashier competitors get stuck. Qwen is the workhorse of open-source AI.


Framework 5: DeepSeek: a reasoning model offline

What: Open models from the Chinese lab DeepSeek. DeepSeek-R1 (early 2025) was one of the first reasoning models with open weights. As of October 2026, the current line is DeepSeek V4: V4.1-Flash came out on September 10, 2026, and its weights are on Hugging Face.

Models: V4.1-Flash is a large MoE model; as a rule you can't run it on ordinary home hardware and need a server. For home hardware, look at the smaller distilled versions of R1 (7B, 14B, 32B, 70B) and small models from other families.

Best for: Reasoning tasks offline. Math problems, complex code logic, multi-step planning.

Special: Open weights. There's also a paid API with low prices (current prices: What's current).

Cost: $0 (local, if your hardware is enough) or at the API rate.

Source: https://github.com/deepseek-ai.

🎨 Picture this: a reasoning model is a calculator that shows its work. These models used to be available only behind a closed subscription. DeepSeek said, "here are the weights, do what you want."


Framework 6: Open WebUI + tools: a full local ChatGPT stack

What: A ChatGPT-like web UI for local models. The backend is Ollama. Supports tools, RAG, web search through SearxNG (local), a code interpreter, voice (TTS/STT through Whisper), and connecting MCP servers.

Best for: A chat interface on your own computer. When you need a "local ChatGPT" with advanced features.

Features: Multi-user (your family or team can use it), RAG (upload documents and the model can see them), pipelines (custom logic). One of the most popular projects in this niche. Check the license in the repository: it has terms about keeping the branding.

Runs on: Docker (starts with one command).

Cost: $0.

Source: https://github.com/open-webui/open-webui.

🎨 Picture this: the ChatGPT interface, but everything runs on your Mac. Your family uses it, documents never leave, and the API bill is zero.


Framework 7: smolagents: Hugging Face's minimal framework

What: A minimalist agent framework from Hugging Face. The main idea: the agent writes code instead of calling tools through JSON schemas. Less overhead, more capability.

Best for: Experimentation, fast prototyping, education. When you want to see "what if the agent writes its own Python instead of using a limited set of tools."

Runs on: Any model (local through Ollama / API through providers).

Cost: $0 (framework).

Source: https://github.com/huggingface/smolagents.

🎨 Picture this: a thin client. The framework stays out of the way and the model works directly. Compare it with CrewAI, which has lots of wrappers and configuration: smolagents puts the minimum amount of code between you and the model.


Framework 8: AutoGen Studio + local models

What: AutoGen is a Microsoft framework for multi-agent collaboration. AutoGen Studio is a GUI on top of it. It works with local models through an OpenAI-compatible API. Important: as of October 2026, AutoGen is in maintenance mode, there won't be new features, and its successor is the Microsoft Agent Framework. AutoGen Studio is fine for prototypes, but not for production.

Best for: Multi-agent collaboration offline. When you need 3+ agents that "talk" to each other to solve a task.

Runs on: Ollama / LM Studio / any OpenAI-compatible endpoint. AutoGen Studio provides a visual builder.

Cost: $0 (framework).

Source: https://github.com/microsoft/autogen.

🎨 Picture this: a team of specialists working in your basement with no internet. An architect, a coder, a tester: each one a specialist, talking to each other and solving the task.


Also frequently mentioned:

  • Continue.dev: an IDE agent for VS Code and JetBrains. Works with local models. Open source. An alternative for people who don't want a SaaS editor
  • LM Studio: a GUI for managing local models + basic agent features. Good for beginners
  • Jan.ai: an open-source ChatGPT alternative. Fully local. A simple UI
  • CrewAI local: regular CrewAI with an Ollama backend instead of OpenAI

Examples of local models for agent tasks (as of October 2026)

Models are the heart of a local agent. Without the right model, even the best framework is useless:

Model Size For Tool calling Download size in Ollama
Qwen3-Coder 30B MoE (~3.3B active) Coding tasks ✅ see the model page
Qwen3 8B / 14B / 30B / 32B General-purpose, fast agents ✅ 5.2 / 9.3 / 19 / 20 GB
gpt-oss (OpenAI, Apache 2.0) 20B / 120B General + reasoning with adjustable effort ✅ 14 GB (needs 16 GB of RAM or more) / 65 GB (80 GB GPU)
Hermes 4 14B / 36B / 70B / 405B Function calling, agent tasks ✅ depends on size and quantization
DeepSeek V4.1-Flash large MoE Reasoning, coding on a server see the documentation needs a server

The Ollama download sizes were checked against the model pages as of October 2026. How to choose: don't trust percentages like "almost as good as Sonnet" from other people's comparisons. Run your own task on two or three models (see Step 5) and compare the results yourself. On complex architecture tasks, the gap with strong cloud models is bigger.

Practical picks:

  • Mac with 32GB → Qwen3 14B or gpt-oss 20B (fits comfortably)
  • Mac with 64GB+ or a 24GB graphics card → Qwen3 30B / Qwen3-Coder 30B / Qwen3 32B (the sweet spot for quality/speed)
  • Mac Studio with 128GB+ or a server → gpt-oss 120B, Hermes 4 70B and larger models

Setup walkthrough: a quick start in 30 minutes

A full local agent stack takes one evening to set up:

bash
# Step 1: Install Ollama (Mac/Linux)
curl -fsSL https://ollama.com/install.sh | sh

# Windows: download the installer from ollama.com

# Step 2: Download a model (once, ~19GB)
ollama pull qwen3:30b

# Alternatives by size:
# ollama pull qwen3:14b   # ~9GB, for 16GB RAM
# ollama pull qwen3:8b    # ~5GB, the minimum
# Current names and sizes: ollama.com/library

# Step 3: Install Open WebUI with Docker
docker run -d -p 3000:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui --restart always \
  ghcr.io/open-webui/open-webui:main

# Step 4: Open http://localhost:3000 — your local ChatGPT is ready

After setup, you have:

  • A local ChatGPT-like interface
  • A model that covers a good share of everyday tasks (code, writing, translation). Check the quality on your own tasks
  • Tool calling (the model can call functions)
  • RAG (upload a PDF and the model answers based on its contents)
  • Full privacy (not a byte leaves your machine)

Real-world cases: what actually works locally

Not theory: patterns that already work:

Case 1: A privacy-first translator (a lawyer)

  • Stack: Qwen3 14B + Open WebUI
  • Use: translating clients' confidential documents EN↔︎ES
  • Speed: depends on the hardware; measure it on your own computer
  • Cost: no token charges after the initial setup
  • Why local: client confidentiality rules out third-party services like DeepL or the Claude API

Case 2: Code review for a proprietary codebase

  • Stack: Qwen3-Coder + Continue.dev in VS Code
  • Use: code review without sending proprietary code to the Anthropic API
  • Quality: good enough for standard reviews; complex architecture reviews are better left to a strong cloud model
  • Cost: no token charges (you need a Mac with 64GB or a PC with a 24GB graphics card)
  • Why local: the company's IP never leaves, and the NDA isn't violated

Case 3: Personal automation (a home bot)

  • Stack: a small model (Qwen3 8B or Hermes 4 14B) + Open WebUI + custom tools
  • Use: smart home (Home Assistant integration), personal notes, reminders
  • Privacy: 100% offline
  • Hardware: a compact Mac or a mini PC with 16GB of memory (one-time)
  • Cost: no monthly fee

Case 4: A local agent for content (a newsletter)

  • Stack: Qwen3 14B + smolagents + RAG over a local Obsidian vault
  • Use: writing posts in the style of your notes and previous posts
  • Cost: no token charges after setup
  • Trade-off: the quality is lower than strong cloud models. For a personal newsletter, it's fine
  • Why local: your drafts and ideas stay with you

🎨 Picture this: a restaurant vs. your home kitchen. The chef at the restaurant cooks better, but at home you can cook as often as you like, with your own recipe, and no check at the end.


Limits and trade-offs: honestly

No rose-colored glasses. Where local is genuinely weaker:

What an API does better What local does better
Reasoning quality (the strongest cloud models) Privacy + sovereignty
Speed on simple queries Cost at high volume
The newest features (computer use, voice) Offline capability
No setup hassle Customization (fine-tuning, prompts)
High-end multimodal (vision, audio) Independence from a vendor
Long context (hundreds of thousands to a million tokens) Predictable (fixed) cost
Tool ecosystem (MCP, plugins, marketplace) Hardware = a one-time cost
Speed on small tasks No rate limits

The honest bottom line as of October 2026: for most professional tasks, an API still wins. Local is for specific use cases (privacy / cost / offline / sovereignty). But the gap is narrowing: the latest generations of open models (Qwen3, gpt-oss, DeepSeek V4) cover many everyday tasks. The strongest cloud models are still ahead on complex reasoning and long tasks.

🎨 Picture this: local is an early-generation electric car. Cheaper to run, but with less range and few charging stations. Every year it gets closer to gas.


Setup cost: one-time vs. ongoing

The most common question is "when does it pay for itself":

Setup One-time Monthly ongoing
Minimum (Mac with 16GB + Ollama + a 7–8B model) $0 (if you already have the Mac) $0 + a small amount of electricity
Standard (Mac with 32GB + a 14–30B model + Open WebUI) $0 for software + the price of a new Mac $0 + electricity
Professional (PC with a 24GB+ graphics card + a 70B model + full stack) the cost of building the PC $0 + electricity (noticeably more because of the GPU)

Compared with an API:

  • A subscription to a cloud assistant costs twelve months' worth per year (current prices: What's current)
  • The payback formula: hardware cost ÷ (monthly API bill − electricity). Hardware only pays for itself at high task volume, or when you need privacy
  • Local doesn't replace an API 100%. In practice, builders use a hybrid: local for routine work, an API for the hard stuff
  • Hybrid economics: local takes the high-volume simple tasks, and the cloud model's bill covers only the hard ones

Anti-patterns: what NOT to do

The rakes beginners step on:

❌ Using local 7B models for production B2B. The quality isn't there. The client will notice. For production, use strong cloud models, period.

❌ Running a 70B model on 8GB of RAM. It'll grind to a halt in swap. The minimum for 70B is 40GB of unified memory or 48GB of VRAM.

❌ Ignoring electricity costs. A GPU at full load for many hours a day raises the bill noticeably. In some scenarios, this eats up the savings vs. an API.

❌ Trying to replicate the strongest cloud models locally. The gap is real. Don't torture your hardware.

❌ Skipping security. Local ≠ automatically safer. A model from an unverified Hugging Face repository may carry malicious code in its scripts. Download from trusted sources, and don't enable trust_remote_code without reading the code.

❌ One local instance for the whole team without queue management. A bottleneck. If 5 people send requests at the same time, the model will stall. You need a load balancer (vLLM) or separate instances.

❌ Comparing local with an API on one task and drawing conclusions. Local wins on volume and privacy. An API wins on complex reasoning. Compare on a specific use case.


By audience: who should use what

A progression by level:

Beginner (just getting started):

  • Stack: Ollama + Qwen3 14B + Open WebUI
  • Covers: most personal tasks (translation, writing/editing, research)
  • Hardware: a computer with 16GB of memory

Intermediate (has a coding background):

  • Stack: + smolagents + Continue.dev in the IDE
  • Covers: an agent-style local workflow for development
  • Hardware: a Mac with 32GB or a PC with a 16GB graphics card

Professional (builder, agency):

  • Stack: + AutoGen or the Microsoft Agent Framework / Hermes 4 70B + custom offline MCP servers
  • Covers: a full multi-agent local stack for proprietary projects
  • Hardware: a Mac Studio with 64GB+ or a PC with a 24GB graphics card

Enterprise (team):

  • Stack: vLLM server + load balancer + fine-tuned models + auth
  • Covers: an internal AI platform for the whole team
  • Hardware: a dedicated server with several professional GPUs

What becomes possible every 6 months:

Smaller, faster. What was frontier a couple of years ago now runs on a laptop. Models of the same quality keep getting smaller.

Tool-native by default. New models (Qwen3, gpt-oss) are trained with function calling out of the box, so separate fine-tunes like Hermes are needed less often.

Multimodal local. Local models can work with images (for example, Qwen3-VL). Audio (local Whisper) is standard. Video is catching up.

Reasoning offline. DeepSeek-R1 showed that reasoning capability is possible with open weights. Today gpt-oss, Qwen3 and DeepSeek all have reasoning modes, and they increasingly fit on consumer hardware.

Mobile AI. Small models (1–4B) already run on smartphones. The next step is agents on mobile without the cloud.

MCP standardization. MCP has become the standard for connecting tools. Local frameworks (Open WebUI, smolagents, Qwen-Agent) support MCP servers, so the same tools work with the Claude API and with local models.

🎨 Picture this: local AI is the personal computer in its early days. Expensive, complicated, not for everyone. Over time it'll get simpler and more accessible, the way computers and smartphones did.


Practice

Step 1: Install Ollama and your first model

bash
# Mac/Linux, with one command
curl -fsSL https://ollama.com/install.sh | sh

# Check the installation
ollama --version
# It should show a version number

# Download Qwen3 14B (a good balance for 32GB of RAM)
ollama pull qwen3:14b

# If you have less memory, start with 8B
ollama pull qwen3:8b

# Test the model
ollama run qwen3:14b "Hi, how are you?"
# It should answer in English

Step 2: Test function calling with Python

python
# test_function_calling.py — checking that the model can call tools
import ollama

# Describe the tool as a function
def get_weather(city: str) -> str:
    """Returns the weather for a city (a stub)"""
    weather_db = {
        "Chicago": "23°F, snow",
        "Cuenca": "64°F, sunny",
        "city_x": "59°F, fog"
    }
    return weather_db.get(city, "No data")

# The tool description for the model
tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for the given city",
        "parameters": {
            "type": "object",
            "properties": {
                "city": {
                    "type": "string",
                    "description": "City name"
                }
            },
            "required": ["city"]
        }
    }
}]

# Send the request
response = ollama.chat(
    model='qwen3:14b',
    messages=[{
        'role': 'user',
        'content': "What's the weather in Cuenca?"
    }],
    tools=tools
)

# Look at the result
print(response['message'])
# It should contain tool_calls with a call to get_weather(city="Cuenca")

# Simulate running the tool
if response['message'].get('tool_calls'):
    for tool_call in response['message']['tool_calls']:
        if tool_call['function']['name'] == 'get_weather':
            city = tool_call['function']['arguments']['city']
            result = get_weather(city)
            print(f"Tool result: {result}")

Step 3: Install Open WebUI with Docker

bash
# Make sure Docker is installed
docker --version

# Start Open WebUI
docker run -d -p 3000:8080 \
  --add-host=host.docker.internal:host-gateway \
  -v open-webui:/app/backend/data \
  --name open-webui --restart always \
  ghcr.io/open-webui/open-webui:main

# Wait ~30 seconds, then open it
open http://localhost:3000

# On first launch:
# 1. Create an admin account (local, it doesn't go anywhere)
# 2. In Settings → Connections → Ollama should be picked up automatically
# 3. In Chat, select the qwen3:14b model
# 4. Start a conversation

Step 4: A simple local agent with smolagents

python
# local_agent.py — a minimal agent with smolagents and Ollama
from smolagents import CodeAgent, OpenAIModel, tool

# Connect to Ollama through the OpenAI-compatible endpoint
# (in older versions of smolagents this class was called OpenAIServerModel)
model = OpenAIModel(
    model_id="qwen3:14b",
    api_base="http://localhost:11434/v1",
    api_key="ollama"  # Ollama ignores the key, but requires a string
)

# Create a custom tool
@tool
def calculate_compound_interest(principal: float, rate: float, years: int) -> float:
    """
    Calculate compound interest.
    
    Args:
        principal: the starting amount
        rate: the annual rate (for example, 0.05 for 5%)
        years: the number of years
    """
    return principal * ((1 + rate) ** years)


@tool  
def read_file(path: str) -> str:
    """
    Read a file from disk.
    
    Args:
        path: the path to the file
    """
    with open(path, 'r', encoding='utf-8') as f:
        return f.read()


# Create the agent
agent = CodeAgent(
    tools=[calculate_compound_interest, read_file],
    model=model,
    add_base_tools=True  # Adds python_interpreter and others
)

# Run it
result = agent.run(
    "If I put $1000 in at 7% annual interest for 10 years, "
    "how much will I have at the end? Explain the calculation."
)

print("\n" + "="*50)
print("RESULT:")
print("="*50)
print(result)
bash
# Installation
pip install smolagents

# Run
python local_agent.py

# The agent should:
# 1. Understand the task
# 2. Call calculate_compound_interest(1000, 0.07, 10)
# 3. Get ~1967.15
# 4. Explain the calculation in English

Step 5: Compare local vs. API on your own task

python
# compare_local_vs_api.py — an honest comparison on one task
import time
import ollama
from anthropic import Anthropic  # pip install anthropic

# Your task: pick something real
prompt = """Write a Python function that:
1. Takes a list of numbers
2. Returns the top 3 largest values
3. Includes type hints
4. Has a docstring with an example
5. Handles edge cases (empty list, < 3 elements)"""

# Option 1: Local through Ollama
print("=" * 50)
print("LOCAL: Qwen3 14B")
print("=" * 50)
start = time.time()
local_response = ollama.chat(
    model='qwen3:14b',
    messages=[{'role': 'user', 'content': prompt}]
)
local_time = time.time() - start
print(local_response['message']['content'])
print(f"\nTime: {local_time:.1f}s | Cost: $0")

# Option 2: API through Anthropic
print("\n" + "=" * 50)
print("API: Claude Sonnet 5.5")
print("=" * 50)
client = Anthropic()  # Requires ANTHROPIC_API_KEY in env
start = time.time()
api_response = client.messages.create(
    model="claude-sonnet-5-5",  # current model IDs: Anthropic's documentation
    max_tokens=1000,
    messages=[{"role": "user", "content": prompt}]
)
api_time = time.time() - start
input_tokens = api_response.usage.input_tokens
output_tokens = api_response.usage.output_tokens
# Prices per 1M tokens: Sonnet 5.5 is $2 input, $10 output (as of October 2026).
# Current prices: the What's current page on the Academy website
PRICE_IN, PRICE_OUT = 2, 10
cost = (input_tokens * PRICE_IN + output_tokens * PRICE_OUT) / 1_000_000
print("".join(b.text for b in api_response.content if b.type == "text"))
print(f"\nTime: {api_time:.1f}s | Cost: ${cost:.4f}")

# Summary
print("\n" + "=" * 50)
print("COMPARISON:")
print("=" * 50)
print(f"Local: {local_time:.1f}s, $0")
print(f"API:   {api_time:.1f}s, ${cost:.4f}")
print(f"\nIf you run 1000 tasks like this a month:")
print(f"Local: $0 + electricity")
print(f"API:   ${cost * 1000:.2f}")

What this test will show:

  • On simple coding tasks, local is often comparable in quality
  • Local latency is usually higher than an API's
  • At high task volume, local starts to win on cost (plug your own volume into the calculation)
  • The quality of the final code: judge it yourself by reading both answers

Tools and resources

  • Ollama: the simplest backend for local models. Mac/Linux/Windows
  • Open WebUI: a ChatGPT-like UI for local models
  • Hermes (Nous Research): the tool-native Hermes 4 models
  • Hermes Agent: an open-source agent from Nous Research
  • OpenClaw: an open-source personal agent in your messaging apps
  • Qwen Team: the main page for Qwen models and frameworks
  • Qwen-Agent: the agent framework for Qwen
  • smolagents (Hugging Face): a minimal Python agent framework
  • DeepSeek: DeepSeek models with open weights
  • AutoGen: Microsoft's framework for multi-agent systems (maintenance mode; the successor is the Microsoft Agent Framework)
  • Continue.dev: an open-source IDE agent
  • LM Studio: a GUI for managing local models
  • Jan.ai: an open-source local ChatGPT alternative
  • vLLM: a production-grade inference engine
  • SGLang: a max-throughput inference framework
  • Karpathy nanoGPT: the educational foundation for nano agents
  • Prices and versions: What's current

Key takeaways

Local agents aren't a replacement for an API: they're a different tool with different economics. Privacy, cost at volume, offline, independence: four reasons local really wins. For everything else, an API is still ahead.

Hardware is the bottleneck. A computer with 16–32GB of memory can run an 8–14B-class model and cover many personal tasks. A Mac Studio or a PC with a 24GB+ graphics card is for serious work with 30B models and up. Electricity counts too.

The realistic path is a hybrid stack. Local for routine work (classification, translation, simple coding, RAG over your personal documents), an API for the hard stuff (architecture, complex reasoning, multimodal). Don't pick one: combine them.


What's next

→ Self-hosted AI enterprise stack: vLLM, SGLang, Docker: how to build an internal AI platform for a team when one computer is no longer enough

The mark stays in this browser only and is never sent anywhere. My progress