The gist
A batch request to an LLM is the letter: Claude thinks, assembles the entire answer, and only then sends it. Real-time streaming is the call: the first words show up on screen very quickly, while Claude is still "thinking" about the end of the sentence.
The difference in how it feels to the user is noticeable. In this lesson we'll go through how to build a streaming architecture, from a simple SSE endpoint all the way to a full voice agent with WebRTC.
Key concepts
- Streaming: sending the LLM's answer word by word, as the tokens are generated, instead of waiting for the full response
- Server-Sent Events (SSE): a one-way protocol. The server pushes data to the client over a regular HTTP connection, and the browser reads it with
EventSource - WebSocket: a two-way protocol with a persistent connection. Both client and server can send data at any time; ideal for chats and interactive agents
- WebRTC: the standard for peer-to-peer media streams (audio/video) in the browser with minimal delay; the foundation of modern voice agents
- TTFT (Time To First Token): the time from sending a request to receiving the first token of the answer; a critical UX metric, and the lower the better
- TPOT (Time Per Output Token): the average time to generate one token; affects how quickly text fills the screen
- LiveKit: open-source WebRTC infrastructure for voice agents; people build voice products on it, and it has a Python framework called LiveKit Agents
- OpenAI Realtime API: an API for two-way streaming audio; lets you build voice assistants without the intermediate STT → LLM → TTS steps. For the current realtime model and its prices, check OpenAI's documentation and pricing page (see also What's current)
Theory
Why streaming is critical for UX
A user staring at a blank screen for several seconds already thinks something broke. A user who sees the first words almost immediately and watches the text "type itself out" is engaged and perceives the system as alive.
Perceived speed matters as much as actual speed. Two services with the same total response time will feel different if one starts streaming immediately and the other holds a pause. For support chatbots and assistants, this usually shows up in user feedback; measure the effect on retention with your own data.
The Claude Streaming API: basic Python
The Anthropic SDK supports streaming out of the box. The key is using stream() instead of the regular messages.create():
import anthropic
client = anthropic.Anthropic()
# Streaming with a context manager
with client.messages.stream(
model="claude-sonnet-5-5", # current models: the What's current page
max_tokens=1024,
messages=[{"role": "user", "content": "Explain quantum entanglement in simple terms"}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
# Get the final message after the stream
final_message = stream.get_final_message()
print(f"\n\nStop reason: {final_message.stop_reason}")
print(f"Input tokens: {final_message.usage.input_tokens}")
print(f"Output tokens: {final_message.usage.output_tokens}")The text_stream method is the simplest option: it returns only the text deltas and ignores housekeeping events. If you need full control over events (content_block_start, ping, message_delta), use stream directly:
with client.messages.stream(...) as stream:
for event in stream:
if event.type == "content_block_delta":
if event.delta.type == "text_delta":
yield event.delta.text
elif event.type == "message_stop":
breakServer-Sent Events: streaming to the browser
SSE is the simplest way to get a streaming response to the browser. It's a regular HTTP GET that the server keeps open and periodically writes data into, in the format data: ...\n\n.
FastAPI with the Anthropic SDK:
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
import anthropic
import json
app = FastAPI()
client = anthropic.Anthropic()
async def generate_stream(prompt: str):
"""Token generator for SSE"""
with client.messages.stream(
model="claude-sonnet-5-5",
max_tokens=1024,
messages=[{"role": "user", "content": prompt}],
) as stream:
for text in stream.text_stream:
# SSE format: data: <payload>\n\n
yield f"data: {json.dumps({'text': text})}\n\n"
# End-of-stream signal
yield "data: [DONE]\n\n"
@app.get("/stream")
async def stream_response(prompt: str):
return StreamingResponse(
generate_stream(prompt),
media_type="text/event-stream",
headers={
"Cache-Control": "no-cache",
"X-Accel-Buffering": "no", # Turn off nginx buffering
}
)The browser side uses the native EventSource:
const eventSource = new EventSource(`/stream?prompt=${encodeURIComponent(userInput)}`);
const output = document.getElementById('output');
eventSource.onmessage = (event) => {
if (event.data === '[DONE]') {
eventSource.close();
return;
}
const { text } = JSON.parse(event.data);
output.textContent += text;
};
eventSource.onerror = () => {
eventSource.close();
console.error('Stream ended or error occurred');
};SSE limitations: GET requests only (long prompts won't fit), and no native support for two-way communication. For interactive chats you need WebSocket.
WebSocket: a two-way stream
WebSocket solves the SSE problem: the client can send messages at any time, and the server can reply with a stream. It's the foundation for chat interfaces.
from fastapi import FastAPI, WebSocket, WebSocketDisconnect
import anthropic
import json
app = FastAPI()
client = anthropic.Anthropic()
@app.websocket("/ws/chat")
async def websocket_chat(websocket: WebSocket):
await websocket.accept()
conversation_history = []
try:
while True:
# Get the user's message
data = await websocket.receive_text()
user_message = json.loads(data)
conversation_history.append({
"role": "user",
"content": user_message["text"]
})
# Stream Claude's answer back over the WebSocket
# (the sync client blocks the event loop; in production use AsyncAnthropic and async with)
full_response = ""
with client.messages.stream(
model="claude-sonnet-5-5",
max_tokens=2048,
system="You are a helpful assistant. Answer in English.",
messages=conversation_history,
) as stream:
for text in stream.text_stream:
full_response += text
await websocket.send_json({
"type": "delta",
"text": text
})
# End-of-answer signal
await websocket.send_json({"type": "done"})
# Add the answer to the history for the next turn
conversation_history.append({
"role": "assistant",
"content": full_response
})
except WebSocketDisconnect:
print("Client disconnected")TTFT and TPOT: the metrics that matter
TTFT (Time To First Token) is the time from sending the request to the first token appearing. This is what the user feels as "the delay before the answer." Rough benchmarks (illustrative; tune them to your product):
- Good: under half a second
- Acceptable: about a second
- Bad: noticeably more than a second
TPOT (Time Per Output Token) is the average time between tokens. It determines how "smoothly" text appears on screen. For comfortable reading, text should appear faster than a person reads it.
What affects TTFT:
- Model size: as a rule, small models (Haiku) respond faster than large ones
- Prompt length: long system prompts and history = slower
- API load: peak hours = higher latency
- Region: closer to the data center = less network latency
- Prompt caching: cached tokens aren't recomputed, so for long contexts TTFT drops noticeably
Measure these metrics in production:
import time
start_time = time.time()
first_token_time = None
token_count = 0
with client.messages.stream(...) as stream:
for text in stream.text_stream:
if first_token_time is None:
first_token_time = time.time()
ttft = first_token_time - start_time
print(f"TTFT: {ttft * 1000:.0f}ms")
token_count += 1
total_time = time.time() - start_time
tpot = (total_time - first_token_time) / token_count * 1000
print(f"TPOT: {tpot:.1f}ms/token")WebRTC and voice agents
WebRTC (Web Real-Time Communication) is a browser standard for peer-to-peer audio and video. It's what Google Meet, Zoom on the web and Discord run on. For voice AI agents, WebRTC lets you:
- Capture the user's microphone and stream the audio
- Receive the agent's synthesized voice
- Keep latency to a minimum (much lower than with regular HTTP requests)
The basic architecture of a voice agent:
User speaks → [Microphone] → [WebRTC] → [STT: Whisper / Deepgram]
↓
[Claude (text)]
↓
[TTS: ElevenLabs / OpenAI TTS]
↓
[User hears] ← [WebRTC] ← [Audio]The problem with this architecture is that latency piles up at every step: speech recognition + Claude's first token + generation + voice synthesis. Together that can easily add up to several seconds, which is critical in a conversation.
LiveKit: infrastructure for voice
LiveKit is an open-source WebRTC server and SDK that handles the hard infrastructure parts: TURN/STUN servers, media routing, session recording. People build their own voice products on LiveKit; ready-made platforms like Vapi (see the lesson Voice AI agents) solve the same problem as a service.
A Python SDK voice agent with LiveKit:
from livekit import agents
from livekit.agents import AgentServer, AgentSession, Agent
from livekit.plugins import anthropic, openai, silero
class VoiceAssistant(Agent):
def __init__(self):
super().__init__(
instructions="""You are a voice assistant.
Answer in English, briefly and to the point.
Use images and analogies in your explanations."""
)
server = AgentServer()
@server.rtc_session()
async def entrypoint(ctx: agents.JobContext):
session = AgentSession(
stt=openai.STT(language="en"), # Speech-to-Text
llm=anthropic.LLM(model="claude-sonnet-5-5"), # LLM; current models: the What's current page
tts=openai.TTS(voice="alloy"), # Text-to-Speech
vad=silero.VAD.load(), # Voice Activity Detection
)
# Noise cancellation and other room parameters are set through room_options
# (see the LiveKit Agents documentation for your version)
await session.start(room=ctx.room, agent=VoiceAssistant())
# Greeting
await session.generate_reply(
instructions="Greet the user in English and offer to help"
)
if __name__ == "__main__":
agents.cli.run_app(server)The plugins are installed separately (the livekit-agents, livekit-plugins-openai, livekit-plugins-silero and livekit-plugins-anthropic packages). The LiveKit Agents API has changed from version to version, so check the documentation.
LiveKit handles all the hard parts: echo cancellation, the jitter buffer, packet loss concealment. You focus on the agent's logic.
OpenAI Realtime API
OpenAI offers a fundamentally different approach: their Realtime API works over WebSocket (there's also a WebRTC option) and lets you send audio directly and get an audio answer back. There's no intermediate STT/TTS; the model handles voice natively. Event names and session fields changed when it came out of beta, so older examples from blogs may not work:
import asyncio
import json
import websockets
import base64
async def realtime_voice_session():
# The model name in the URL is an example: take the current one from the model list in OpenAI's documentation
url = "wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1"
headers = {
"Authorization": f"Bearer {OPENAI_API_KEY}",
# The OpenAI-Beta: realtime=v1 header isn't needed for the current version of the interface
}
async with websockets.connect(url, additional_headers=headers) as ws:
# Configure the session. Audio format, voice and end-of-speech detection
# are set in session.audio: see the Realtime API events reference for the exact fields
await ws.send(json.dumps({
"type": "session.update",
"session": {
"type": "realtime",
"instructions": "You are a voice assistant. Keep your answers short.",
}
}))
# Listen for events
async for message in ws:
event = json.loads(message)
if event["type"] == "response.output_audio.delta":
# Receive a chunk of the audio answer
audio_data = base64.b64decode(event["delta"])
# Play the audio...
elif event["type"] == "response.done":
print("Answer finished")The advantage of the Realtime API is native voice processing without the STT → LLM → TTS chain, so the latency is lower than with a cascade and the conversation sounds more natural. The drawback: OpenAI models only. Claude isn't available in this format (use LiveKit with Claude for the cascade setup).
Reducing latency
A few techniques for minimizing latency in production:
1. Prompt caching. If your system prompt is long, cache it (the minimum length for caching depends on the model; see the documentation). Repeat requests won't recompute the prompt, and TTFT drops:
messages = client.messages.create(
model="claude-sonnet-5-5",
system=[{
"type": "text",
"text": long_system_prompt,
"cache_control": {"type": "ephemeral"} # Cache it
}],
messages=conversation,
max_tokens=1024,
)2. A smaller model for the first answer. The "cascade" strategy: start with Haiku (fast, cheap) and switch to Sonnet for complex questions.
3. Edge deployment. Cloudflare Workers AI or AWS Lambda@Edge: run inference closer to the user.
4. Streaming with early TTS. For voice agents, don't wait for the full text; pass each completed sentence to TTS:
buffer = ""
async for text in stream.text_stream:
buffer += text
# If we've collected a complete sentence, send it straight to TTS
if any(buffer.endswith(p) for p in ['.', '!', '?', '...']):
await tts_engine.synthesize_and_play(buffer.strip())
buffer = ""A real-world case: a streaming support chat
Picture this: an online store, hundreds of inquiries a day, customers waiting several minutes for an agent. You roll out a Claude agent with WebSocket streaming.
What changes: the user sees the first words of the answer almost immediately, and the wait stops feeling like a wait. The agent handles the routine questions, and human staff only see the complicated cases. What share of inquiries gets resolved without a human, and how much each answer costs, depends on your knowledge base and model: measure it in a pilot, and calculate the tokens using the prices on the What's current page.
The key architectural choice here is WebSocket instead of REST. One persistent connection per session, two-way exchange, no overhead of setting up a connection for every message.
Practice
Task: a streaming chat API with FastAPI and Claude
Goal: build a WebSocket server that streams Claude's answers in real time, with a simple HTML client.
Step 1: Install the dependencies
pip install fastapi uvicorn anthropic websockets python-dotenvStep 2: Backend, a FastAPI WebSocket server
Create a server.py file:
import os
import json
import asyncio
from fastapi import FastAPI, WebSocket, WebSocketDisconnect
from fastapi.responses import HTMLResponse
import anthropic
from dotenv import load_dotenv
load_dotenv()
app = FastAPI(title="Streaming Chat API")
client = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
SYSTEM_PROMPT = """You are a helpful AI assistant. Answer in English,
use structure and images in your explanations. Be brief."""
# A simple HTML client for testing
HTML = """
<!DOCTYPE html>
<html>
<head><title>Streaming Chat</title></head>
<body>
<h1>Streaming Chat with Claude</h1>
<div id="chat" style="height:400px;overflow-y:auto;border:1px solid #ccc;padding:10px;"></div>
<div style="margin-top:10px">
<input id="input" type="text" style="width:80%" placeholder="Type a message..." />
<button onclick="sendMessage()">Send</button>
</div>
<script>
const ws = new WebSocket(`ws://${location.host}/ws/chat`);
const chat = document.getElementById('chat');
let currentBubble = null;
ws.onmessage = function(event) {
const data = JSON.parse(event.data);
if (data.type === 'start') {
currentBubble = document.createElement('div');
currentBubble.style.cssText = 'margin:8px 0;padding:8px;background:#e8f4fd;border-radius:8px';
currentBubble.innerHTML = '<strong>Claude:</strong> ';
chat.appendChild(currentBubble);
} else if (data.type === 'delta') {
currentBubble.innerHTML += data.text;
chat.scrollTop = chat.scrollHeight;
} else if (data.type === 'done') {
currentBubble = null;
}
};
function sendMessage() {
const input = document.getElementById('input');
const text = input.value.trim();
if (!text) return;
const userBubble = document.createElement('div');
userBubble.style.cssText = 'margin:8px 0;padding:8px;background:#f0f0f0;border-radius:8px;text-align:right';
userBubble.innerHTML = `<strong>You:</strong> ${text}`;
chat.appendChild(userBubble);
ws.send(JSON.stringify({text}));
input.value = '';
}
document.getElementById('input').addEventListener('keypress', (e) => {
if (e.key === 'Enter') sendMessage();
});
</script>
</body>
</html>
"""
@app.get("/")
async def get():
return HTMLResponse(HTML)
@app.websocket("/ws/chat")
async def websocket_endpoint(websocket: WebSocket):
await websocket.accept()
history = []
try:
while True:
data = await websocket.receive_text()
user_msg = json.loads(data)
history.append({
"role": "user",
"content": user_msg["text"]
})
# Start-of-answer signal
await websocket.send_json({"type": "start"})
full_response = ""
# Stream Claude's answer
with client.messages.stream(
model="claude-haiku-4-5", # Haiku for speed (check that the model is still available in the API)
max_tokens=1024,
system=SYSTEM_PROMPT,
messages=history,
) as stream:
for text in stream.text_stream:
full_response += text
await websocket.send_json({
"type": "delta",
"text": text
})
# End of the answer
await websocket.send_json({"type": "done"})
# Save it to the history
history.append({
"role": "assistant",
"content": full_response
})
# Limit the history (last 10 pairs)
if len(history) > 20:
history = history[-20:]
except WebSocketDisconnect:
print(f"Client disconnected, history length: {len(history)}")
except Exception as e:
print(f"Error: {e}")
await websocket.close()Step 3: Start the server
# Make sure ANTHROPIC_API_KEY is in your .env file
uvicorn server:app --reload --port 8000Open http://localhost:8000 in your browser. You'll see a chat interface with instant streaming.
Step 4: Measure TTFT in code
Add latency measurement to the WebSocket handler:
import time
start = time.time()
first_token = True
token_count = 0
with client.messages.stream(...) as stream:
for text in stream.text_stream:
if first_token:
ttft = (time.time() - start) * 1000
print(f"TTFT: {ttft:.0f}ms")
first_token = False
token_count += 1
# ... send to the client
total = (time.time() - start) * 1000
tpot = total / token_count if token_count > 0 else 0
print(f"TPOT: {tpot:.1f}ms/token | Total tokens: {token_count}")Step 5: Add an SSE endpoint for comparison
from fastapi.responses import StreamingResponse
@app.get("/stream")
async def stream_endpoint(prompt: str):
async def generate():
with client.messages.stream(
model="claude-haiku-4-5", # a fast model for real time
max_tokens=512,
messages=[{"role": "user", "content": prompt}],
) as stream:
for text in stream.text_stream:
yield f"data: {json.dumps({'text': text})}\n\n"
yield "data: [DONE]\n\n"
return StreamingResponse(
generate(),
media_type="text/event-stream",
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"}
)Test it: curl "http://localhost:8000/stream?prompt=What+is+AI%3F". You'll see the tokens arrive in real time right in your terminal.
Tools and resources
| Tool | Purpose | Documentation |
|---|---|---|
| Anthropic SDK (Python) | Streaming API with stream() and text_stream |
platform.claude.com/docs/en/build-with-claude/streaming |
| FastAPI | WebSocket and SSE endpoints | fastapi.tiangolo.com |
| LiveKit | Open-source WebRTC for voice agents | docs.livekit.io/agents |
| LiveKit Agents (Python) | A framework on top of LiveKit with Claude integration | github.com/livekit/agents |
| OpenAI Realtime API | Native voice streaming | developers.openai.com/api/docs/guides/realtime |
| websockets (Python) | WebSocket client/server | websockets.readthedocs.io |
| Deepgram | Fast streaming speech recognition, an alternative to Whisper | developers.deepgram.com |
| ElevenLabs | High-quality TTS with streaming output | elevenlabs.io/docs |
| uvicorn | ASGI server for FastAPI | www.uvicorn.org |
Key takeaways
"Streaming isn't a technical detail, it's a UX decision. The user sees the first token almost immediately and feels the system is alive."
"TTFT is your main metric in real-time AI. If the first token arrives too late, the user loses patience. Cache your prompts and pick a fast model for the first touch."
"WebSocket for chat, SSE for one-way updates, WebRTC for voice: don't mix these patterns up. Each tool solves its own problem and has its own setup cost."
Next lesson
→ Chatbot managers: NLP, intent routing and smart escalation
The mark stays in this browser only and is never sent anywhere. My progress