●2.1.283 — No Claude Code release over the weekend; 2.1.283 from September 25 is still the latest. Its new auto-mode default deserves a check against your allow rules●PLANS — Pro and Team Standard now default to Opus instead of Sonnet, a change shipped in Claude Code 2.1.280 that shifts how you compare plans●10/07 — Nine days left until the old spellings of the Claude Desktop / Cowork managed config keys stop being accepted; after that they fail closed●AUTO — Users are asking why commands covered by allow rules get rejected in auto mode. The fix is shaping the invoked command to match the rule●NEW — The week Opus 5.5 became the default: three places an alias was still pointing at the old model●CURSOR — Cursor credits burn in proportion to model cost, so routing Claude through Cursor gets expensive fast. Above ~50 requests a day, flat-rate Claude Code Pro is the steadier choice●2.1.283 — No Claude Code release over the weekend; 2.1.283 from September 25 is still the latest. Its new auto-mode default deserves a check against your allow rules●PLANS — Pro and Team Standard now default to Opus instead of Sonnet, a change shipped in Claude Code 2.1.280 that shifts how you compare plans●10/07 — Nine days left until the old spellings of the Claude Desktop / Cowork managed config keys stop being accepted; after that they fail closed●AUTO — Users are asking why commands covered by allow rules get rejected in auto mode. The fix is shaping the invoked command to match the rule●NEW — The week Opus 5.5 became the default: three places an alias was still pointing at the old model●CURSOR — Cursor credits burn in proportion to model cost, so routing Claude through Cursor gets expensive fast. Above ~50 requests a day, flat-rate Claude Code Pro is the steadier choice
Production Voice Agents with Claude API: Latency Budgets, Cost, and Fallbacks
Orchestrating Whisper/Deepgram, Claude API, and TTS into a voice agent that survives production — latency budgets measured on Cloudflare Workers and Cloud Run, per-session cost math, three-tier fallbacks, and barge-in handling.
Claude API123voice agentsspeech-to-texttext-to-speechproduction architecture
✦ Premium Article
The 300ms that broke the conversation
The first time I put a voice agent in front of real users, the replies that had felt snappy in testing sounded noticeably slow on a real phone. The measurement said p95 was 1.8 seconds — only 300ms over budget.
Listeners still felt the pause. They repeated themselves, the repeat produced a duplicate transcript, and the duplicate made the next turn slower. A small overshoot on paper, a broken conversation in practice.
The hard part of voice agents is not model quality. It is how you spend time. Claude API handles the language reasoning; speech recognition and speech synthesis live in separate services, and milliseconds pile up at every seam between them.
What follows is a full stack — Whisper or Deepgram for input, Claude API for reasoning, several TTS engines for output — organized around four decisions: latency budget, cost math, failover, and monitoring. Everything is TypeScript/Node.js, with code you can run as written.
Voice Agent Architecture Overview
A production voice agent system spans multiple integrated layers:
What follows is each piece in the order I hit trouble with it, starting with speech recognition.
✦
Thank you for reading this far.
Continue Reading
What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.
WHAT YOU'LL LEARN
✦How I split a 1500ms voice-to-voice latency budget into STT 600ms / Claude Haiku 400ms / TTS 400ms, and compressed real-world p50 to 910ms with Deepgram Streaming
✦Reduced per-session cost from $0.024 to $0.011 across 4 specific decisions, with the Sonnet routing rule that finally worked after Sonnet-judges-itself failed
✦Why Cloudflare Workers cannot host the inference layer, and the Durable Objects + Cloud Run split I run in production with signed JWT session tokens
Secure payment via Stripe · Cancel anytime
✦
Unlock This Article
Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.
Whisper API: Best for accuracy, supports 99+ languages, good for offline-first scenarios
Deepgram Nova-2: Best latency (50-300ms), real-time streaming, conversational focus
Local Whisper: Most cost-effective at scale, but requires GPU infrastructure
Intelligent Response Generation with Claude API
Conversation Context Management
Claude's strength lies in understanding nuance. To leverage this in voice agents, maintain rich conversation context:
// src/services/claude-agent.tsimport Anthropic from '@anthropic-ai/sdk';interface ConversationMessage { role: 'user' | 'assistant'; content: string;}interface VoiceAgentConfig { apiKey: string; systemPrompt: string; maxContextMessages?: number; temperature?: number;}export class VoiceAgent { private client: Anthropic; private systemPrompt: string; private maxContextMessages: number; private temperature: number; constructor(config: VoiceAgentConfig) { this.client = new Anthropic({ apiKey: config.apiKey }); this.systemPrompt = config.systemPrompt; this.maxContextMessages = config.maxContextMessages || 10; this.temperature = config.temperature || 0.7; } async generateResponse( userInput: string, conversationHistory: ConversationMessage[], options: { signal?: AbortSignal } = {} ): Promise<string> { // Keep only recent messages to optimize token usage const recentHistory = conversationHistory.slice(-this.maxContextMessages); const messages = [ ...recentHistory, { role: 'user' as const, content: userInput } ]; try { const response = await this.client.messages.create( { model: 'claude-sonnet-5', max_tokens: 1024, temperature: this.temperature, system: this.systemPrompt, messages }, // Request options take a signal, so a turn folded by barge-in // stops generating instead of billing to the end. { signal: options.signal } ); // Don't hardcode content[0]: add thinking or tool use and the first // block stops being text. const textBlock = response.content.find( (block): block is Extract<typeof block, { type: 'text' }> => block.type === 'text' ); if (!textBlock) { throw new Error('No text block in Claude response'); } return textBlock.text; } catch (error: unknown) { console.error('Claude API call failed:', error); throw new Error(`Response generation failed: ${toMessage(error)}`); } } async generateResponseStream( userInput: string, conversationHistory: ConversationMessage[], onChunk: (chunk: string) => void ): Promise<string> { const recentHistory = conversationHistory.slice(-this.maxContextMessages); const messages = [ ...recentHistory, { role: 'user' as const, content: userInput } ]; let fullResponse = ''; try { const stream = await this.client.messages.stream({ model: 'claude-sonnet-5', max_tokens: 1024, temperature: this.temperature, system: this.systemPrompt, messages }); for await (const chunk of stream) { if (chunk.type === 'content_block_delta' && chunk.delta?.type === 'text_delta') { const text = chunk.delta.text; fullResponse += text; onChunk(text); // Stream to client } } return fullResponse; } catch (error) { console.error('Claude streaming failed:', error); throw new Error(`Streaming response failed: ${error.message}`); } }}
Optimized System Prompts for Voice
Voice agents need brief, action-oriented system prompts. Users can't easily re-read long responses:
const VOICE_AGENT_SYSTEM_PROMPT = `You are a helpful, conversational voice assistant.Guidelines:- Respond naturally, as if speaking to someone. Keep sentences short (under 20 words when possible).- Use conversational language. Avoid jargon unless the user introduced it first.- If unsure, admit it. Don't speculate.- Break complex information into bullet points (max 3 items per response).- No emojis. No markdown formatting. Speak like a real person.- Be warm and encouraging while remaining professional.`;
Text-to-Speech Implementation and Optimization
Multi-Provider TTS Adapter Pattern
Production systems need failover. Implement a provider-agnostic interface:
One thing cost me several days here, so I'll put it first. If you hand ws.on('message') an async listener, the library never awaits the promise it returns. Every frame starts its own run, and whichever turn finishes generating first is the one that lands in the conversation history.
Speech arrives in small pieces. A short acknowledgement and a long question can be tens of milliseconds apart, and when the order flips, Claude receives a transcript where the answer precedes the question. What I saw on my side was replies that felt slightly off — I only found the cause after dumping the history to a log and reading the order with my own eyes.
The fix is a single serialized chain per session, plus an AbortController so a new utterance folds the previous turn. That cancellation is the server-side half of the barge-in handling discussed later on.
// src/websocket/voice-agent-handler.tsimport WebSocket from 'ws';import { VoiceAgent } from '../services/claude-agent';import { DeepgramStreamingSTT } from '../services/stt/deepgram-streaming';import { TTSCache } from '../services/tts/tts-cache';interface StreamingSessionConfig { sessionId: string; userId: string; voiceAgent: VoiceAgent; sttService: DeepgramStreamingSTT; ttsCache: TTSCache;}type Turn = { role: 'user' | 'assistant'; content: string };export class VoiceAgentStreamHandler { private sessionConfig: StreamingSessionConfig; private conversationHistory: Turn[] = []; // One chain per session. History is only ever mutated inside it. private chain: Promise<void> = Promise.resolve(); // The turn currently in flight. A new frame aborts it (barge-in). private current: AbortController | null = null; constructor(config: StreamingSessionConfig) { this.sessionConfig = config; } async handleWebSocketConnection(ws: WebSocket): Promise<void> { console.log(`Session ${this.sessionConfig.sessionId} connected`); // The listener itself is synchronous, so arrival order survives. ws.on('message', (data: Buffer) => { this.current?.abort(); const ac = new AbortController(); this.current = ac; this.chain = this.chain.then(() => this.runTurn(ws, data, ac.signal)); }); ws.on('close', () => { console.log(`Session ${this.sessionConfig.sessionId} closed`); this.current?.abort(); this.cleanupSession(); }); ws.on('error', (error: unknown) => { console.error(`WebSocket error: ${toMessage(error)}`); }); } private async runTurn(ws: WebSocket, audioChunk: Buffer, signal: AbortSignal): Promise<void> { if (signal.aborted) return; // superseded while it waited its turn try { const transcript = await this.transcribeAudioChunk(audioChunk); if (!transcript || signal.aborted) return; const response = await this.sessionConfig.voiceAgent.generateResponse( transcript, this.conversationHistory, { signal } ); if (signal.aborted) return; // Only turns that survived to here get written to history. this.conversationHistory.push( { role: 'user', content: transcript }, { role: 'assistant', content: response } ); const audioBuffer = await this.generateSpeech(response); if (signal.aborted) return; ws.send(JSON.stringify({ type: 'response', transcript, response, audioBuffer: audioBuffer.toString('base64') })); } catch (error: unknown) { if (signal.aborted) return; // abort-driven rejections aren't user-facing console.error('Stream processing error:', error); ws.send(JSON.stringify({ type: 'error', message: toMessage(error) })); } } private async transcribeAudioChunk(audioData: Buffer): Promise<string> { // Integration with Deepgram streaming service return ''; } private async generateSpeech(text: string): Promise<Buffer> { // Integration with TTS cache service return Buffer.alloc(0); } private cleanupSession(): void { this.conversationHistory = []; }}
toMessage is a small helper that turns a caught value into a string safely. Since TypeScript 4.4, strict types the catch binding as unknown, so writing error.message directly doesn't type-check at all. Running TypeScript 7.0.2 with --strict --noEmit over that line stops with error is of type 'unknown'.
// src/utils/to-message.tsexport function toMessage(error: unknown): string { if (error instanceof Error) return error.message; if (typeof error === 'string') return error; try { return JSON.stringify(error); } catch { return String(error); }}
Other snippets in this article still interpolate error.message for brevity; in the running code they all go through toMessage.
What the same three frames look like with and without serialization
You don't need a WebSocket server to see this. ws emits message through EventEmitter, so an EventEmitter with an async listener and per-turn generation times reproduces it exactly.
// Run on Node.js v22.23.2import { EventEmitter } from 'node:events';const em = new EventEmitter();const history = [];const generate = (t) => new Promise(r => setTimeout(() => r(`assistant:${t}`), 60 - t * 10));em.on('message', async (t) => { // the original shape const resp = await generate(t); history.push(`user:${t}`, resp);});for (const t of [1, 2, 3]) setTimeout(() => em.emit('message', t), t * 5);setTimeout(() => console.log(JSON.stringify(history)), 300);
The first row came back exactly reversed, because the fastest turn reached push first. In the third row the superseded turns fold without generating anything and only the last utterance gets an answer — which is what you want from a voice agent. It closes the failure where a user rephrases and then hears a reply to the sentence they abandoned.
The moment you pass async to an event listener, arrival order stops being a guarantee. If a sequence has to stay in order, build one place that waits. I'd rather put that in before any of the other optimizations here.
Conversation Context and Memory Management
PostgreSQL Session Persistence
Store conversations for later retrieval and user continuity:
Claude 3.5 Sonnet costs $3 per million input tokens and $15 per million output tokens. Track every call:
// src/monitoring/cost-tracker.tsimport { CloudWatch } from 'aws-sdk';export interface APICost { service: string; // 'claude', 'whisper', 'tts-google', etc inputUnits: number; // tokens, minutes, characters outputUnits: number; costUSD: number; timestamp: Date;}export class CostTracker { private cloudwatch: CloudWatch; private costs: APICost[] = []; constructor() { this.cloudwatch = new CloudWatch(); } // Rates as published on 2026-08-30 (USD / MTok). // Scattering literals through the metering code means hunting them down // every time a model generation turns over. private static readonly RATES = { 'claude-opus-5': { input: 5, output: 25 }, 'claude-sonnet-5': { input: 2, output: 10 }, 'claude-haiku-4-5-20251001': { input: 1, output: 5 }, } as const; // Pass `usage` straight through. `input_tokens` excludes cached tokens, so // skipping the cache fields produces a number smaller than your invoice. recordClaudeCall( model: keyof typeof CostTracker.RATES, usage: { input_tokens?: number; output_tokens?: number; cache_read_input_tokens?: number; cache_creation_input_tokens?: number; } ): void { const rate = CostTracker.RATES[model]; const fresh = usage.input_tokens ?? 0; const cacheRead = usage.cache_read_input_tokens ?? 0; const cacheWrite = usage.cache_creation_input_tokens ?? 0; const out = usage.output_tokens ?? 0; // Cache reads bill at 0.1x the input rate; 5-minute writes at 1.25x. const costUSD = (fresh * rate.input + cacheRead * rate.input * 0.1 + cacheWrite * rate.input * 1.25 + out * rate.output) / 1_000_000; this.costs.push({ service: `claude:${model}`, inputUnits: fresh + cacheRead + cacheWrite, outputUnits: out, costUSD, timestamp: new Date() }); } recordWhisperCall(duration: number): void { // Whisper: $0.02/minute const cost = (duration / 60) * 0.02; this.costs.push({ service: 'whisper', inputUnits: Math.ceil(duration), outputUnits: 0, costUSD: cost, timestamp: new Date() }); } recordTTSCall(characters: number, provider: string): void { let cost = 0; if (provider === 'google') { // Google Cloud TTS: $16 per million characters cost = (characters / 1_000_000) * 16; } else if (provider === 'elevenlabs') { // ElevenLabs: $0.30 per 10,000 characters cost = (characters / 10_000) * 0.30; } this.costs.push({ service: `tts-${provider}`, inputUnits: characters, outputUnits: 0, costUSD: cost, timestamp: new Date() }); } async publishMetrics(sessionId: string): Promise<void> { const totalCost = this.costs.reduce((sum, c) => sum + c.costUSD, 0); await this.cloudwatch.putMetricData({ Namespace: 'VoiceAgent', MetricData: [ { MetricName: 'SessionCost', Value: totalCost, Unit: 'None', Dimensions: [{ Name: 'SessionId', Value: sessionId }] }, { MetricName: 'APICallCount', Value: this.costs.length, Unit: 'Count' } ] }).promise(); console.log(`Session ${sessionId} cost: $${totalCost.toFixed(4)}`); }}
Metrics Collection with Prometheus
// src/monitoring/metrics.tsimport prom from 'prom-client';export const voiceAgentMetrics = { totalSessions: new prom.Counter({ name: 'voice_agent_total_sessions', help: 'Total number of sessions', labelNames: ['status'] // success, failed, timeout }), apiCallsTotal: new prom.Counter({ name: 'voice_agent_api_calls_total', help: 'Total API calls by service', labelNames: ['service'] // claude, whisper, tts }), sessionDurationSeconds: new prom.Histogram({ name: 'voice_agent_session_duration_seconds', help: 'Session duration in seconds', buckets: [10, 30, 60, 300, 600] }), apiLatencyMs: new prom.Histogram({ name: 'voice_agent_api_latency_ms', help: 'API latency in milliseconds', labelNames: ['service'], buckets: [50, 100, 200, 500, 1000, 2000] }), activeSessions: new prom.Gauge({ name: 'voice_agent_active_sessions', help: 'Number of currently active sessions' })};export function recordSessionMetric(durationSeconds: number, success: boolean): void { voiceAgentMetrics.totalSessions.inc({ status: success ? 'success' : 'failed' }); voiceAgentMetrics.sessionDurationSeconds.observe(durationSeconds);}
Seven lessons that aren't in the official docs
The sections above are the design story. Below are the things I only learned by running this stack in production as a solo developer — the failure modes that never appear in the official docs.
1. Measure the latency budget in three layers (1500ms total)
End-to-end voice-to-voice latency above 1500ms breaks the conversational rhythm. Users start asking "did it cut out?" and re-speak before the agent can answer. That number is the threshold I keep walking back to across every voice product I have shipped.
My production budget split, measured on Cloudflare from Tokyo to us-east-1:
Layer
Budget
p50
p95
STT (Deepgram Streaming)
600ms
280ms
480ms
Claude Haiku response
400ms
320ms
620ms
TTS (ElevenLabs Flash v2)
400ms
240ms
410ms
Network round-trip
100ms
70ms
130ms
Total
1500ms
910ms
1640ms
If you call Whisper REST naively, inference only fires after the full audio clip lands, which adds 600 to 900ms after the last syllable. Deepgram Streaming uses VAD to predict the endpoint, which cuts perceived latency roughly in half. I started with Whisper REST for simplicity, watched p95 exceed 1800ms, and migrated to Deepgram Streaming three weeks later.
// Production budget checker — emit Sentry warnings on any over-budget sessioninterface LatencyBudget { stt: { budget: 600; actual?: number }; llm: { budget: 400; actual?: number }; tts: { budget: 400; actual?: number }; network: { budget: 100; actual?: number };}export function assertBudget(b: LatencyBudget, sessionId: string) { const total = (b.stt.actual ?? 0) + (b.llm.actual ?? 0) + (b.tts.actual ?? 0) + (b.network.actual ?? 0); if (total > 1500) { console.warn(`[budget-exceeded] session=${sessionId} total=${total}ms`, b); } return total <= 1500;}
2. How I cut per-session cost from $0.024 to $0.011
I price every feature in dollars-per-session before writing the first line of code — a habit left over from years of running ad-supported apps. The first build (Whisper + Sonnet + ElevenLabs Multilingual) ran roughly $0.024 per 3-minute session. Over 9 weeks I brought it to $0.011 with four decisions.
Decision
Before
After
Reduction
STT: Whisper → Deepgram Nova-2
$0.006/min
$0.0043/min
-28%
LLM first-pass: Sonnet → Haiku
$3/1M tok
$0.25/1M tok
-91%
Sonnet only for "complex" queries
100% Sonnet
18% Sonnet
-82%
TTS: ElevenLabs Multilingual → Flash v2
$0.30/1k chars
$0.10/1k chars
-67%
Those four numbers come from the rates in effect in March 2026. Model generations have turned over since, and the rates moved with them. Priced against the official page as of 2026-08-30, the comparison looks different.
Dropping from Sonnet to Haiku now saves 50% on input and 50% on output, not the 91% it saved back then, because the rate on the larger model came down. That shrinking margin is not bad news. It means running Sonnet on the primary turn is now a viable budget, which it was not before. These days I only fall back to Haiku 4.5 for acknowledgements and small talk, and keep Sonnet 5 for everything else.
The Sonnet routing rule deserves its own warning. My first attempt was to let Claude itself judge complexity, but the judge ran on Sonnet too, so the savings vanished. The rule that actually works in production is dumb on purpose: input over 80 characters, OR a technical term in the last 3 turns, OR the system prompt explicitly escalated. Simple rules beat clever judges.
3. Cloudflare Workers cannot be the inference layer
My Next.js sites already run on Cloudflare Workers + OpenNext, so my first instinct was "let me put the voice agent there too." It does not work, for three concrete reasons.
30-second CPU limit: every conversation turn burns 1–2s of CPU, and long sessions exceed the limit. Durable Objects share the same quota.
WebSocket constraints: Workers WebSockets cap individual frames at 32KB and disconnect at 16 hours. Bidirectional audio streaming requires Durable Objects + Hibernation API, which is much more design overhead than people assume.
Audio library bundling: ffmpeg WASM builds usually push you past the 10MB Worker bundle limit.
What I run today: Cloudflare Workers for signaling and session management on Durable Objects, and Google Cloud Run for the audio processing + Anthropic API path. The Cloudflare side still owns membership gating (premium_token cookie), and only paid users get a signed JWT for the Cloud Run WebSocket.
// Workers issues a short-lived JWT; Cloud Run verifies it on WebSocket upgradeimport { SignJWT } from 'jose';export async function issueVoiceSessionToken(env: Env, userId: string) { const secret = new TextEncoder().encode(env.VOICE_SESSION_SECRET); return await new SignJWT({ sub: userId, scope: 'voice-session' }) .setProtectedHeader({ alg: 'HS256' }) .setIssuedAt() .setExpirationTime('15m') .sign(secret);}
4. A three-tier fallback so I never get paged at 2am
A failure in a voice pipeline reaches the user as silence. The three seconds a spinner buys you in a text chat becomes a dead pause in a conversation. So this is the one layer where I refuse to start shipping until every dependency has three tiers of fallback behind it.
Claude failure: Sonnet → Haiku → pre-recorded "Sorry, could you say that again" TTS (cost near zero)
TTS failure: ElevenLabs → OpenAI TTS → Browser Web Speech API (audio quality hit, still usable)
Whether to tell the user about the degradation is a product decision. For free assistants I stay silent; for membership users I display "running in simplified mode" so the trust signal stays intact. Honesty is the foundation of paid membership.
5. Hold conversation history as a graph, not a flat array
Voice users cannot consciously segment context the way chat users do. "Wait, not that one, the earlier one" happens constantly. My first version stored a flat array and dumped it into the Claude context window. Even when I stayed under 200k tokens, answer quality drifted because recent turns started leaking influence into older turns.
I now hold the conversation as three node types:
Question node: the user's intent plus the answer the agent gave
Topic node: an intermediate node that groups question nodes about the same subject
State node: an explicit state transition like "booking → awaiting confirmation → canceled"
Each Claude call receives the last 6 turns in full + the current topic-node summary + the state node only. Effective context fits in 4k–8k tokens and my eval set shows 30–40% accuracy improvement on multi-topic conversations.
6. Split UX and cost dashboards in Grafana
Prometheus + Grafana is the obvious choice, but the production lesson is: never put UX metrics and cost metrics on the same board. When p95 latency spikes, you need to ask "should I roll back the Haiku-first routing for cost?" without the cost number pulling your eye.
Cost board: avg cost per session, STT/LLM/TTS share, Sonnet ratio, free vs. premium unit-cost gap
A Grafana variable for tier=free|premium makes it fast to ask "does cost-cutting hurt the paying users?" — the question I check first every morning on anything with a paid tier behind it.
7. Kill barge-in at the playback buffer, not at the TTS call
The loudest complaints in production were not about latency or recognition accuracy. They were about what happens when a user starts talking while the agent is still speaking.
My first fix was the obvious one: abort the server-side TTS request as soon as VAD detects speech. It barely helped. The client already holds 400 to 800ms of audio in its playback buffer, so stopping the source does nothing for audio that has already been delivered — the agent talks over the user anyway.
The right place to cut is the playback side.
// Client: when VAD fires, kill local playback firstexport function createBargeInController(ws: WebSocket) { let current: AudioBufferSourceNode | null = null; return { play(node: AudioBufferSourceNode) { current = node; node.start(); }, onUserSpeechStart() { current?.stop(); // 1. silence locally, within ~20ms current = null; ws.send(JSON.stringify({ type: 'barge_in' })); // 2. then tell the server }, };}
Order matters. Notify the server first and the 70–130ms round trip stays audible, which users hear as "my interruption did nothing."
On the server, receiving barge_in means aborting the TTS stream and rewriting conversation history to contain only the text the user actually heard. Store the full response and Claude will build the next turn on content nobody listened to, which is how voice conversations go subtly off the rails.
I attach a character offset to every TTS streaming chunk, persist the assistant turn only up to the interruption offset, and append (interrupted here — user began speaking). That single line is enough for Claude to pick the thought back up instead of restarting it.
Can you buy latency? Deciding whether Fast mode belongs in the budget
Once the budget table is drawn and every cheap saving is taken, one option is left: buy the speed outright. The Claude API offers Fast mode as a research preview. Add speed: "fast" to a request and output generation on Opus 5 and Opus 4.8 gets visibly quicker. When four hundred milliseconds are being fought over, that is the first lever anyone reaches for.
The rate doubles in exchange.
Model / mode
Input
Output
Claude Opus 5 (standard)
$5 / MTok
$25 / MTok
Claude Opus 5 (Fast mode)
$10 / MTok
$50 / MTok
The multiplier covers the entire context window — no discount kicks in past 200k input tokens. Prompt caching multipliers stack on top, so a cache read under Fast mode costs 0.1x the input rate, which works out to $1 / MTok. Fast mode does not combine with the Batch API and is not available through partner-operated clouds. If your inference sits on Bedrock or Google Cloud, the option is off the table before the pricing question even comes up — worth checking first.
Decide what a single turn is worth before you decide the mode
Here is the arithmetic against my own setup. Escalations — the turns my router judges complex enough for a larger model — run at 18% of all turns, and a session is about 12 turns, so roughly 2.2 escalations. Assume 1,500 input tokens and 160 output tokens on those turns.
Escalation target
Per turn
Per session (2.2 turns)
Claude Sonnet 5 (standard)
$0.0046
$0.0101
Claude Opus 5 (standard)
$0.0115
$0.0253
Claude Opus 5 (Fast mode)
$0.0230
$0.0506
Against the $0.015 per-session line I hold myself to, Fast mode escalations alone come to more than three times the whole budget — before STT and TTS are added back. As an indie developer paying the invoice out of my own pocket, I could not write a story in which that difference earns itself back.
Turn it around and the answer flips. If one response decides whether a booking completes or a deal moves forward, $0.023 a turn is cheap. The deciding factor is neither speed nor price but whether you have already priced a single turn. Trying Fast mode without that line drawn leaves you with a pleasant impression of speed and an unpleasant invoice.
It quietly does nothing on one generation
There is a second trap: support splits three ways across generations.
Model
Behavior with speed: "fast"
Claude Opus 5 / Opus 4.8
Runs in Fast mode, billed at Fast rates
Claude Opus 4.7
Returns an error
Claude Opus 4.6
Runs at standard speed, billed at standard rates, no error
That last row is the awkward one. The response comes back 200, the invoice looks normal, and an investigation into "we turned on Fast mode but p95 never moved" finds nothing in the application logs or the billing breakdown. A setting that is accepted and then ignored burns time precisely because there is nowhere obvious to look.
So I closed the door at the type level: passing a model that does not support Fast mode should fail to compile, not fail at three in the morning.
// src/services/model-policy.tsconst RATES = { 'claude-opus-5': { input: 5, output: 25, fastInput: 10, fastOutput: 50 }, 'claude-sonnet-5': { input: 2, output: 10 }, 'claude-haiku-4-5-20251001': { input: 1, output: 5 },} as const;type ModelId = keyof typeof RATES;// Narrows to the models that carry a fast rate — today, only 'claude-opus-5'type FastCapable = { [K in ModelId]: 'fastInput' extends keyof (typeof RATES)[K] ? K : never;}[ModelId];export async function escalateFast( client: Anthropic, model: FastCapable, // passing Sonnet or Haiku is a type error messages: Anthropic.MessageParam[], system: string) { return client.messages.create({ model, speed: 'fast', max_tokens: 512, system, messages, });}
FastCapable is derived from the shape of RATES, so updating the rate table immediately updates which models the call sites will accept. I wanted the table and the code to be unable to drift apart without the compiler noticing.
I also verify the effect on the metering side. Median generation latency on escalation turns is recorded separately by the value of speed, and if the two distributions sit on top of each other, Fast mode is not doing anything. Carrying a setting that you merely believe is active is more expensive than checking.
What I rely on when I translate the design into code
Numbers and code matter, but the question I lead with is "what am I trading my own time for?" My budget — p95 1500ms, $0.015 per session — exists so I can decide in 30 seconds whether a 2am alert needs me out of bed. Without a line drawn in advance, every alert feels equally urgent, and the deciding itself is what wears you down.
Voice agents demand more "humanness" than text chat. Being fast, cheap, and reliable at the same time is fundamentally hard, but layering budgets, holding three tiers of fallback, and splitting the dashboards keeps the on-call rotation survivable.
Thank you for reading this far. If you are building voice agents on a similar stack, I hope these numbers and decisions save you a few weekends.
Share
Thank You for Reading
Claude Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.