Library · Assistants and chatbots

Chatbot managers: NLP, intent routing and smart handoff to a human

Builder60 minUpdated: October 2026
58 of 105 in the library

Module: 19. Voice & Real-Time AI | Time: about 25 min reading + 35 min practice


The gist

🎨 Picture this: a chatbot manager is an airport air traffic controller. Every message from a customer is a plane asking to land. The controller instantly decides: this flight has a pricing question (runway A), this is an unhappy passenger (runway B), this is a VIP client with a big deal (runway C, hand over to manual control right away). Without a controller, it's chaos and collisions. With one, every plane gets the right instructions and lands where it should.

Most business owners build chatbots like this: the bot answers FAQs and that's it. That's an airport with one runway. In this lesson you'll learn to build a system with smart routing: the bot understands the intent, pulls out the key details, and knows when it can answer on its own and when to hand off to a real person immediately.


Key concepts

  • Intent classification: automatically figuring out what the user wants (a question, a complaint, a purchase, a request for a human) using Claude Haiku or rules
  • Entity extraction: pulling specific details out of a message: names, amounts, dates, product SKUs
  • Intent routing: the logic that steers the conversation: rules (keywords) vs. AI classification vs. a hybrid approach
  • Escalation triggers: the set of conditions under which the bot hands the conversation to a human: customer anger, a direct request, a complex question, a large amount
  • State management: keeping the conversation's context between messages with Redis or Supabase
  • Multi-channel unification: one routing logic for Telegram, WhatsApp and a website widget
  • Fallback handling: what to do when the bot doesn't understand the intent: ask a clarifying question, offer options, or escalate right away
  • Quality metrics: escalation rate (% handed to agents), resolution rate (% resolved by the bot), CSAT (the customer's rating)

Theory

What intent classification is and why you need it

When a customer writes "I want to return an item," it's not just text. It's an intent (a return) that calls for a specific response flow. When they write "are you people serious?!", that's an alarm: the customer is irritated, it's risky for the bot to answer, and a real person is needed.

Intent classification means automatically assigning an intent label to every incoming message. Without it, the bot works like a cashier on their first day: it doesn't understand what's being asked and tries to answer everything with the same template.

The main intent categories for a business bot:

  • FAQ: questions about prices, shipping terms, business hours
  • Complaint: a complaint, a problem, negativity
  • Purchase: interest in buying, a request for advice
  • Support: a technical question, a problem with an order
  • Human request: an explicit request to talk to an agent ("let me talk to a manager")
  • Out of scope: an irrelevant request, spam

How classification works: three approaches

Approach 1: Rules (keywords)

The simplest. A dictionary of keywords for each intent. "return", "refund" → return. "price", "how much" → a pricing FAQ. Fast, cheap, predictable. The downside: it's fragile. "I'd like to return to your services" will wrongly trigger a return.

Approach 2: AI classification (Claude Haiku)

You pass the message to the model and ask it to classify it. It's more accurate and understands context and sarcasm. It costs more than rules (you pay for tokens) and is a bit slower. In return, it makes noticeably fewer mistakes with real-world language. The cost per message is small and is calculated in tokens: current prices are on the What's current page.

Approach 3: Hybrid (recommended)

First check the rules: if one matches with high confidence, use the rule (cheap). If not, send it to Claude Haiku (accurate). This is the best balance of speed and cost for high-traffic bots.

🎨 Picture this: hybrid routing is like the sorting line at a post office. Standard letters go down the automatic belt (rules). Odd-shaped packages go to a human sorter (AI). High throughput, few mistakes.

Entity extraction: what's hidden in the text

Besides the intent, you need to pull out the specific details. A customer writes: "I want to order 3 boxes by Friday, delivered to 15 Main St." Here:

  • Quantity: 3
  • Deadline: Friday
  • Address: 15 Main St

Without entity extraction, the bot answers "okay, we'll set that up" and loses all that information. With it, the bot pulls the details into the CRM, checks stock, and says whether you can make Friday.

For entity extraction we also use Claude or regex. Claude handles fuzzy phrasing better ("the day after tomorrow," "around five grand").

Escalation triggers: when to hand off to a human

This is the most important part of the system. Escalate too rarely and the customer leaves angry. Too often and your agent drowns in simple questions.

Smart escalation works off several signals:

1. An explicit request for an agent "I want to talk to a manager," "get me a human," "call me": the bot escalates immediately, no questions asked.

2. A frustration detector Analyze the tone: all caps, exclamation points, marker words ("nightmare," "terrible," "never again," "give me my money back"). Claude Haiku returns a sentiment score. If it's above the threshold, escalate.

3. A complex request The intent wasn't recognized twice in a row. The question needs information from several systems. A legal or financial question (per company policy).

4. A high-value lead The amount in the request is > $N. The customer asks about a business plan. They mentioned a competitor, which calls for a personal touch.

5. A system fallback After 3 failed attempts to answer, escalate automatically. Better to hand off to a human than to annoy the customer with an endless "I didn't understand."

State management: the conversation's memory

🎨 Picture this: a good server remembers what you ordered, even if they got pulled away by other tables. A bad bot with no state asks "how can I help you?" every time, even in the middle of a conversation.

To store state we use Redis (fast, for active sessions) or Supabase (persistent, for history). Customer conversations contain personal data: store only what you need, limit how long you keep it, and don't send the model anything extra. In the state we store:

python
{
    "session_id": "uuid",
    "user_id": "telegram_chat_id",
    "history": [...],       # the last N messages
    "current_intent": "complaint",
    "entities": {"order_id": "12345"},
    "frustration_score": 0.3,
    "escalated": False,
    "turns_without_resolution": 1
}

With every new message we update the state and make the routing decision based on the whole history, not just the last line.

Multi-channel: one engine, many channels

The intent routing logic shouldn't be tied to Telegram or WhatsApp. The right architecture:

Code
[Telegram] → adapter → [Intent Engine] → action → [Telegram response]
[WhatsApp] → adapter → [Intent Engine] → action → [WhatsApp response]
[Web Widget] → adapter → [Intent Engine] → action → [Web response]

The adapter normalizes the incoming message into one format. The Intent Engine works the same way for everyone. The response is formatted for the right channel. This lets you support 3 channels with one codebase.

Alternatives: Botpress, Dialogflow, Rasa

Botpress: an open-source platform with a visual conversation builder. Good for complex conversation trees, with built-in NLU. Choose it if your team isn't technical and you need a visual editor.

Dialogflow (Google): an enterprise solution with powerful NLU. Integrates well with Google Workspace. More expensive, with vendor lock-in. Choose it for large corporate projects.

Rasa: open source with maximum flexibility. Requires ML expertise. Choose it if you need on-premise hosting and full control over your data.

Claude directly (our approach): the best choice for solo founders and small teams. Quick to start, low cost, flexible. Haiku for classification, Sonnet for complex answers.

Monitoring: three key metrics

Escalation rate: the percentage of conversations handed to an agent. There's no single standard: it depends on your niche and how complex the questions are. If almost every conversation goes to an agent, the bot is useless. If almost none do, the bot may not be escalating when it should.

Resolution rate: the percentage of questions the bot closes without an agent. Set your target based on your own pilot, not on other people's numbers. It grows as the knowledge base improves.

CSAT (Customer Satisfaction Score): the customer's rating after the conversation ends (1-5). Track it separately for bot sessions and agent sessions. The gap shows you where the weak spot is.

Build the monitoring dashboard in Grafana, or use a simple /stats command in a chat for the owner.

Code: an intent routing pipeline in Python + FastAPI

python
import os
import json
import redis
from fastapi import FastAPI, Request
from anthropic import Anthropic

app = FastAPI()
client = Anthropic()
r = redis.Redis(host='localhost', port=6379, decode_responses=True)

ESCALATION_PHRASES = [
    "talk to a manager", "agent", "real person",
    "human", "call me", "speak to someone"
]

FRUSTRATION_WORDS = [
    "terrible", "nightmare", "never again", "scam",
    "money back", "ripoff", "unacceptable"
]


def get_session(session_id: str) -> dict:
    data = r.get(f"session:{session_id}")
    if data:
        return json.loads(data)
    return {
        "history": [],
        "frustration_score": 0.0,
        "turns_without_resolution": 0,
        "escalated": False,
        "current_intent": None
    }


def save_session(session_id: str, state: dict):
    r.setex(f"session:{session_id}", 3600, json.dumps(state))


def check_escalation_rules(message: str, state: dict) -> tuple[bool, str]:
    """Escalation rules: a quick check without AI."""
    msg_lower = message.lower()
    
    # An explicit request for an agent
    if any(phrase in msg_lower for phrase in ESCALATION_PHRASES):
        return True, "explicit_request"
    
    # Frustration detector based on words
    frustration_hit = sum(1 for w in FRUSTRATION_WORDS if w in msg_lower)
    if frustration_hit >= 2:
        return True, "frustration_detected"
    
    # Frustration accumulated over the session
    if state["frustration_score"] > 0.7:
        return True, "cumulative_frustration"
    
    # The bot failed to help 3 times in a row
    if state["turns_without_resolution"] >= 3:
        return True, "repeated_fallback"
    
    return False, ""


def classify_intent(message: str, history: list) -> dict:
    """Claude Haiku classifies the intent and extracts entities."""
    history_text = "\n".join([
        f"{m['role']}: {m['content']}" for m in history[-4:]
    ])
    
    response = client.messages.create(
        model="claude-haiku-4-5",   # check that the model is still available in the API; current models: the "What's current" page
        max_tokens=300,
        system="""You are an intent classifier for an online store's chatbot.
Return JSON with these fields:
- intent: one of [faq_price, faq_delivery, faq_return, complaint, purchase_intent, order_status, out_of_scope]
- confidence: 0.0-1.0
- entities: an object with the extracted data (order_id, amount, date, product)
- sentiment: positive/neutral/negative
- frustration_score: 0.0-1.0

JSON only, no explanations.""",
        messages=[{
            "role": "user",
            "content": f"Conversation history:\n{history_text}\n\nNew message: {message}"
        }]
    )
    
    try:
        return json.loads("".join(b.text for b in response.content if b.type == "text"))
    except Exception:
        return {
            "intent": "out_of_scope",
            "confidence": 0.0,
            "entities": {},
            "sentiment": "neutral",
            "frustration_score": 0.3
        }


def generate_bot_response(intent: str, message: str, entities: dict, history: list) -> str:
    """Generate a response based on the intent."""
    
    intent_prompts = {
        "faq_price": "Answer the question about the product's price. If no specific product is mentioned, ask which one.",
        "faq_delivery": "Answer about shipping: 2-5 days, free on orders over $50.",
        "faq_return": "Answer about returns: 14 days, no questions asked, receipt required.",
        "complaint": "Acknowledge the complaint with empathy. Ask for the details needed to resolve it.",
        "purchase_intent": "Help with the purchase. Ask for details and offer to add it to the cart.",
        "order_status": "Ask for the order number if it wasn't given.",
    }
    
    prompt = intent_prompts.get(intent, "Politely say you didn't understand the question and offer some options.")
    
    response = client.messages.create(
        model="claude-haiku-4-5",
        max_tokens=200,
        system=f"You are a polite assistant for an online store. {prompt}. Keep it short: 1-3 sentences.",
        messages=[{"role": "user", "content": message}]
    )
    return "".join(b.text for b in response.content if b.type == "text")


async def notify_operator(session_id: str, message: str, reason: str, state: dict):
    """Notify the agent in Telegram."""
    import httpx
    
    bot_token = os.environ.get("TELEGRAM_BOT_TOKEN")
    operator_chat_id = os.environ.get("OPERATOR_CHAT_ID")
    
    history_preview = "\n".join([
        f"{'👤' if m['role']=='user' else '🤖'} {m['content']}"
        for m in state["history"][-3:]
    ])
    
    text = (
        f"🚨 Escalation | {reason}\n"
        f"Session: {session_id}\n\n"
        f"Recent messages:\n{history_preview}\n\n"
        f"Latest: {message}"
    )
    
    async with httpx.AsyncClient() as http:
        await http.post(
            f"https://api.telegram.org/bot{bot_token}/sendMessage",
            json={"chat_id": operator_chat_id, "text": text}
        )


@app.post("/chat")
async def chat_endpoint(request: Request):
    body = await request.json()
    session_id = body["session_id"]
    message = body["message"]
    
    state = get_session(session_id)
    
    # Quick check of the escalation rules
    should_escalate, escalation_reason = check_escalation_rules(message, state)
    
    if should_escalate and not state["escalated"]:
        state["escalated"] = True
        state["history"].append({"role": "user", "content": message})
        save_session(session_id, state)
        await notify_operator(session_id, message, escalation_reason, state)
        return {
            "response": "Got it. I'm connecting you with an agent, who will reply within 5 minutes.",
            "escalated": True,
            "reason": escalation_reason
        }
    
    # AI intent classification
    classification = classify_intent(message, state["history"])
    intent = classification.get("intent", "out_of_scope")
    confidence = classification.get("confidence", 0.0)
    
    # Update the state
    state["current_intent"] = intent
    state["frustration_score"] = max(
        state["frustration_score"],
        classification.get("frustration_score", 0.0)
    )
    
    # Low confidence: count it as unresolved
    if confidence < 0.5 or intent == "out_of_scope":
        state["turns_without_resolution"] += 1
    else:
        state["turns_without_resolution"] = 0
    
    # Generate the response
    bot_response = generate_bot_response(
        intent, message,
        classification.get("entities", {}),
        state["history"]
    )
    
    # Save the history
    state["history"].append({"role": "user", "content": message})
    state["history"].append({"role": "assistant", "content": bot_response})
    state["history"] = state["history"][-20:]  # Keep the last 20 messages
    save_session(session_id, state)
    
    return {
        "response": bot_response,
        "intent": intent,
        "confidence": confidence,
        "escalated": False
    }

This code runs as a FastAPI service. A Telegram bot, a WhatsApp adapter or a website widget sends POST requests to /chat and gets back a response with the intent, confidence and an escalation flag.


Practice

Step 1: Set up the environment (5 min)

bash
mkdir chatbot-manager && cd chatbot-manager
python -m venv venv && source venv/bin/activate
pip install fastapi uvicorn anthropic redis httpx python-dotenv

# .env file
echo "ANTHROPIC_API_KEY=your_key" >> .env
echo "TELEGRAM_BOT_TOKEN=your_bot_token" >> .env
echo "OPERATOR_CHAT_ID=your_chat_id" >> .env

# Start Redis locally (or use Redis Cloud)
docker run -d -p 6379:6379 redis:alpine

Step 2: Set up a knowledge base for FAQs (10 min)

Create a knowledge_base.json file with answers to your business's common questions. The structure:

json
{
  "faq_price": "Our products start at $10. Current catalog: shop.example.com/catalog",
  "faq_delivery": "Shipping takes 2-5 business days. Free on orders over $50.",
  "faq_return": "Returns within 14 days. Receipt and original packaging required. Refunds are issued within 3-5 days."
}

Replace intent_prompts in the code with real answers from your knowledge base.

Step 3: Run it and test the routing (10 min)

bash
uvicorn main:app --reload --port 8000

Test with curl or Postman:

bash
# A regular question
curl -X POST http://localhost:8000/chat \
  -H "Content-Type: application/json" \
  -d '{"session_id": "test1", "message": "how much is shipping?"}'

# An escalation test
curl -X POST http://localhost:8000/chat \
  -H "Content-Type: application/json" \
  -d '{"session_id": "test2", "message": "this is TERRIBLE, a total ripoff, I want my money back NOW!!!"}'

# An explicit request for an agent
curl -X POST http://localhost:8000/chat \
  -H "Content-Type: application/json" \
  -d '{"session_id": "test3", "message": "I want to talk to a real person"}'

Check that the Telegram notification arrives when there's an escalation.

Step 4: Connect it to a Telegram bot (8 min)

python
# telegram_adapter.py
from telegram.ext import Application, MessageHandler, filters
import httpx, os

async def handle_message(update, context):
    session_id = str(update.effective_chat.id)
    message = update.message.text
    
    async with httpx.AsyncClient() as client:
        resp = await client.post(
            "http://localhost:8000/chat",
            json={"session_id": session_id, "message": message}
        )
        data = resp.json()
    
    await update.message.reply_text(data["response"])

app = Application.builder().token(os.environ["TELEGRAM_BOT_TOKEN"]).build()
app.add_handler(MessageHandler(filters.TEXT, handle_message))
app.run_polling()

The same /chat endpoint works for other channels: a WhatsApp or website widget adapter just sends the message in the same format.

Step 5: Set up monitoring (5 min)

Add a /stats endpoint for a quick look at the metrics in Redis:

python
@app.get("/stats")
async def get_stats():
    keys = r.keys("session:*")
    total = len(keys)
    escalated = sum(1 for k in keys if json.loads(r.get(k)).get("escalated"))
    return {
        "total_sessions": total,
        "escalated": escalated,
        "escalation_rate": f"{escalated/total*100:.1f}%" if total else "0%",
        "active_last_hour": total  # simplified
    }

Tools and resources

Tool What it's for Price
Claude Haiku Intent classification, entity extraction Pay per token; as of October 2026: $1 / $5 per million tokens (input / output)
Redis Storing session state (fast) Free self-hosted; cloud at the provider's rates
Supabase Persistent conversation history Has a free tier with limits
FastAPI Backend for webhook endpoints Free, open source
Botpress Visual conversation builder (an alternative) Has a free plan; pricing on its website
Dialogflow Enterprise NLU (when you need a platform) Pay as you go; pricing on the Google Cloud website
Rasa On-premise, maximum control Free, open source
Telegram Bot API A channel + agent alerts Free
Grafana A metrics monitoring dashboard Free self-hosted

The recommended starter stack: FastAPI + Claude Haiku + Redis + Telegram. It comes together quickly and cheaply. Estimate the model cost like this: two Haiku calls per message (classification and response) × tokens × price per million tokens. Run a pilot on a hundred messages and look at usage in the API responses.


Key takeaways

"A bot without intent routing is a cashier who answers everything with 'combo number three.' A bot with routing is a dispatcher who knows where to send every request."

"The escalation rule is simple: when in doubt, hand it to a human. Better to spend an agent's time than to lose a customer over a bad bot answer."

"Escalation rate is your bot's honest KPI. If it's through the roof, the bot isn't working. If it's close to zero, the bot may be escalating too rarely and customers are leaving without a word."


Next lesson

→ Product analytics: PostHog, Mixpanel and smart insights

The mark stays in this browser only and is never sent anywhere. My progress