Voice AI Telephony: Twilio NLU and TTS Integration
Upon initiation of a client call, the telephony system encounters challenges: the speech recognition engine incorrectly transcribes "I want to order" due to audio conversion artifacts arising from μ-law 8 kHz to PCM 16 kHz, degrading STT accuracy by 30%. We integrate Twilio Voice AI with real NLU, employing Whisper large-v3 for recognition, GPT-4o for response generation, and ElevenLabs for speech synthesis. Consequently, the bot comprehends the client even with an accent and responds without filler phrases. Order integration — we resolve latency and recognition quality issues.
Problems We Solve
Audio format conversion — Twilio transmits μ-law 8 kHz, whereas Whisper requires PCM 16 kHz. Conversion errors introduce artifacts and degrade recognition quality. We utilize audioop.ratecv with anti-aliasing and cross-fade smoothing to eliminate clicks.
WebSocket connection reliability — disconnection results in loss of the audio stream. We implement a reconnection mechanism with buffering of the last second, along with a jitter buffer for smooth playback.
Latency management — total latency must not exceed 2 seconds. We optimize the pipeline via parallel STT and response generation and caching of frequent queries. Comparison: our pipeline reduces latency by a factor of 2 compared to sequential processing.
Technical Implementation
TwiML webhook for incoming call
from fastapi import FastAPI, Request from twilio.twiml.voice_response import VoiceResponse, Start, Stream, Say app = FastAPI() @app.post("/incoming-call") async def handle_incoming_call(request: Request): response = VoiceResponse() # Start Media Stream start = Start() start.stream( url=STREAM_ENDPOINT, track="both_tracks" # incoming and outgoing audio ) response.append(start) # Play greeting response.say( "Hello! I am a voice assistant. How can I help?", voice="alice", language="en-US" ) response.pause(length=30) return Response(content=str(response), media_type="text/xml") WebSocket handler for Media Streams
import asyncio import json import base64 from fastapi import WebSocket @app.websocket("/stream") async def handle_stream(websocket: WebSocket): await websocket.accept() call_sid = None stream_sid = None audio_buffer = bytearray() try: async for message in websocket.iter_text(): data = json.loads(message) event = data.get("event") if event == "start": call_sid = data["start"]["callSid"] stream_sid = data["start"]["streamSid"] session = create_session(call_sid) elif event == "media": # Twilio uses mulaw 8kHz mulaw_audio = base64.b64decode(data["media"]["payload"]) audio_buffer.extend(mulaw_audio) # Process when 2 seconds accumulated (16000 bytes @ 8kHz) if len(audio_buffer) >= 16000: await process_audio_chunk( bytes(audio_buffer), websocket, stream_sid, session ) audio_buffer = bytearray() elif event == "stop": break except Exception as e: logger.error(f"Stream error: {e}") async def send_audio_to_caller(websocket: WebSocket, stream_sid: str, audio_bytes: bytes): """Send synthesized audio back to the call""" encoded = base64.b64encode(audio_bytes).decode() await websocket.send_json({ "event": "media", "streamSid": stream_sid, "media": { "payload": encoded } }) Audio format conversion
Twilio uses μ-law 8 kHz. Whisper works with PCM 16 kHz:
import audioop def mulaw_to_pcm16k(mulaw_bytes: bytes) -> bytes: """μ-law 8kHz → PCM 16-bit 8kHz → upsample to 16kHz using anti-aliasing""" pcm_8k = audioop.ulaw2lin(mulaw_bytes, 2) # μ-law → PCM 16-bit pcm_16k, _ = audioop.ratecv(pcm_8k, 2, 1, 8000, 16000, None) # 8→16kHz return pcm_16k How Twilio Voice AI processes audio in real time?
The Media Streams API transmits audio in 20 ms chunks. We accumulate a buffer of up to 2 seconds (16000 bytes at 8 kHz) and send it to STT. This reduces the number of requests and improves accuracy through context. After recognition, the LLM generates a response, TTS synthesizes speech, and the audio is sent back through the same WebSocket.
Why is correct audio format conversion important?
Conversion errors μ-law → PCM can introduce noise or shift the sampling frequency, leading to up to 30% loss in STT accuracy. We use audioop.ulaw2lin with explicit bit depth and ratecv with a quality filter. We also apply cross-fade smoothing at chunk boundaries to eliminate clicks.
Common conversion errors and their solutions
- Ignoring bit depth: μ-law 8-bit → PCM 16-bit. Without
ulaw2linyou get 8-bit PCM, STT won't understand. - Wrong rate: upsample from 8 kHz to 16 kHz requires interpolation.
ratecvwithNoneuses linear interpolation; for better quality, use cubic interpolation. - Artifacts during batch processing: clicks occur at chunk boundaries. We add cross-fade smoothing of 50 ms duration.
TTS approach comparison
| Parameter | ElevenLabs (cloud) | Kokoro (ONNX local) |
|---|---|---|
| Latency | 300-500 ms | 100-200 ms |
| Quality | Very high | Medium |
| Cost | Per character ($0.0003/char) | Free (CPU/GPU) |
| Voices | 100+ | 10+ |
For production, we recommend a combination: ElevenLabs for primary dialogue, Kokoro for fallback under load. ElevenLabs costs approximately $0.0003 per character, while Kokoro is free, providing substantial savings for high-volume scenarios.
STT solution comparison
| Parameter | Whisper large-v3 | Deepgram Nova-2 | Google STT |
|---|---|---|---|
| Latency | 200-400 ms | 150-300 ms | 300-600 ms |
| Accuracy (Russian) | 95% | 93% | 90% |
| Price per hour | $0.006 (Self-host) | $0.004 | $0.006 |
| Accent adaptation | High | Medium | Medium |
For Russian-language scenarios, Whisper large-v3 delivers 5% better accuracy than Deepgram and 10% better than Google STT. Self-hosting Whisper at ~$0.006/hour can save 30% compared to cloud-based STT services.
Process of Work
- Audit — analysis of current telephony and NLP requirements (1-2 days). Estimated cost: $500-$1,000.
- Design — selection of STT/LLM/TTS, WebSocket architecture, conversion, and jitter buffer parameters (3-5 days).
- Implementation — writing handler, CRM integration, monitoring setup, and VAD configuration (1-2 weeks).
- Testing — load testing with 100 call simulation, recognition accuracy checks, and DTMF handling (3-5 days).
- Deployment — server or cloud deployment, API documentation, and training session (2-3 days).
Approximate Timeline
Basic bot on Twilio with one scenario — from 2 weeks. Production solution with multilingual support and monitoring — up to 2 months. Cost is calculated individually, depending on call volume and NLP complexity. Typical integration costs range from $5,000 to $20,000, plus ongoing Twilio fees (~$0.0045/min) and AI service fees (e.g., Whisper self-host ~$0.006/hour). Twilio, Media Streams API — official documentation.
Deliverables (Что входит в работу)
- TwiML and WebSocket handler configuration
- Audio format conversion (μ-law ↔ PCM 16kHz) with anti-aliasing
- STT/TTS and LLM integration (cloud or local)
- Real-time monitoring dashboard with p99 latency alerts
- API documentation and access credentials
- One training session for your team
- Two-week post-launch support with hotfix window
These deliverables include all configuration files, API keys, and documentation necessary for handoff.
Advantages and Contact
Over 5 years of experience in voice AI systems, 10+ Twilio Voice AI deployments for retail and logistics. We guarantee stability: p99 latency < 2.5 sec, uptime 99.9%. Certified Twilio and ML specialists.
Contact us for a project estimate within 1 day. Get a consultation and accurate timeline.







