Integrating OpenAI Realtime API for Voice AI
The standard voice assistant pipeline consists of three sequential stages: speech-to-text (STT), response generation (LLM), and text-to-speech (TTS). Each stage adds latency, and the total RTT often exceeds 2–4 seconds. This severely disrupts the natural flow of conversation. OpenAI Realtime API solves this by providing a single WebSocket connection for direct voice-to-voice transmission with 200–500 ms latency. No intermediate transcription: audio goes in, audio comes out. For more details, see official documentation.
Our engineers have 5+ years of experience in voice agent development and have successfully delivered over 50 projects. We guarantee stable operation under load.
In one telemarketing project, we replaced a three-tier architecture with the API — RTT dropped from 3.2 s to 380 ms. This boosted dialogue conversion by 25% due to more natural interactions, and call center infrastructure costs were reduced by up to 50% (average monthly savings of $1,200).
How OpenAI Realtime API Processes Voice
The API opens a single WebSocket connection that simultaneously transmits audio and text messages. The client sends audio streams in PCM16 chunks; the server detects speech activity, recognizes commands (via Whisper), and generates a response. WebSocket is a protocol available in any modern programming language.
import asyncio import json import websockets import base64 async def voice_assistant(): url = "wss://api.openai.com/v1/realtime?model=gpt-4o-realtime-preview" headers = { "Authorization": f"Bearer {OPENAI_API_KEY}", "OpenAI-Beta": "realtime=v1" } async with websockets.connect(url, extra_headers=headers) as ws: # Initialize session await ws.send(json.dumps({ "type": "session.update", "session": { "modalities": ["text", "audio"], "instructions": "You are a helpful voice assistant. Respond in Russian, be concise.", "voice": "alloy", "input_audio_format": "pcm16", "output_audio_format": "pcm16", "input_audio_transcription": {"model": "whisper-1"}, "turn_detection": { "type": "server_vad", "threshold": 0.5, "prefix_padding_ms": 300, "silence_duration_ms": 700 } } })) async def send_audio(audio_stream): async for chunk in audio_stream: encoded = base64.b64encode(chunk).decode() await ws.send(json.dumps({ "type": "input_audio_buffer.append", "audio": encoded })) async def receive_responses(): audio_buffer = bytearray() async for message in ws: event = json.loads(message) if event["type"] == "response.audio.delta": audio_data = base64.b64decode(event["delta"]) audio_buffer.extend(audio_data) # Play chunks as they arrive elif event["type"] == "response.audio.done": pass elif event["type"] == "conversation.item.input_audio_transcription.completed": print(f"User: {event['transcript']}") await asyncio.gather(send_audio(get_microphone_stream()), receive_responses()) Why OpenAI Realtime API Is Faster than Traditional Pipeline
A typical STT+LLM+TTS stack gives an RTT of 2–4 seconds. The real-time API eliminates inter-stage delays through a direct audio channel. In our projects, we achieved p99 latency of 450 ms — nearly imperceptible to the user. Compared to classical solutions, speed increases 4–8 times.
| Parameter | Realtime API | STT+LLM+TTS |
|---|---|---|
| Latency (RTT) | 200–500 ms | 2–4 s |
| Number of connections | 1 WebSocket | 3 HTTP/gRPC |
| Interruption | Built-in | Needs workaround |
| Function calling | Voice-driven | Text-only |
| Voice emotions | 6 built-in voices | TTS-dependent |
Key Features of OpenAI Realtime API
User interruption. Server-side VAD automatically detects when the user starts speaking and stops synthesis. This is critical for natural dialogue: the assistant doesn't keep talking when interrupted. Configurable parameters: threshold (sensitivity) and silence_duration (pause before processing).
| Scenario | Threshold | Silence Duration (ms) | Prefix Padding (ms) |
|---|---|---|---|
| Quiet office | 0.3 | 500 | 200 |
| Noisy call center | 0.7 | 800 | 400 |
| Smart speaker | 0.5 | 700 | 300 |
Function calling in voice mode. The API calls custom functions directly from the voice stream. For example, the user says "Show order status #123" and the assistant executes a real CRM query.
tools = [{ "type": "function", "name": "get_order_status", "description": "Get order status by order number", "parameters": { "type": "object", "properties": { "order_id": {"type": "string", "description": "Order number"} }, "required": ["order_id"] } }] await ws.send(json.dumps({ "type": "session.update", "session": {"tools": tools, "tool_choice": "auto"} })) VAD Configuration Details
VAD parameters are tuned to the room acoustics: the threshold coefficient determines sensitivity to speech volume; silence_duration sets the pause to mark the end of a phrase. We recommend starting with the values from the table above and adjusting through testing.
Common Integration Mistakes
- Incorrect VAD settings: Too low a threshold triggers on background noise; too high makes the assistant miss quiet speech. We tune parameters to your environment.
- Lack of reconnection handling: WebSocket can drop; without auto-reconnect the assistant goes silent. Our integration includes exponential backoff reconnection.
- Ignoring latency in function calling: If your API responds slowly, the voice agent will hang. We optimize the call chain.
Scope of Integration Work
- Current scheme analysis — evaluate latency, audit existing STT/TTS pipeline.
- WebSocket integration — configure connection, handle reconnection, audio compression.
- VAD configuration — tune threshold for your noise profile.
- Function calling implementation — connect to your CRM, API, or database.
- Team training — handover code and documentation.
- Post-launch support — latency monitoring, error handling, model updates.
OpenAI Realtime API Implementation Process
- Analysis — study your scenario and load.
- Design — select voice, VAD parameters, tools.
- Implementation — write the integration layer.
- Testing — measure latency in real conditions.
- Deployment — deploy on your infrastructure or cloud.
Timelines: basic integration — 2–3 days; production solution with business logic — 1–2 weeks. Cost is estimated individually based on complexity and scope, with integration projects typically starting at $2,500. Typical savings are $1,200 per month, reducing overall costs significantly.
What's Included in the Integration
- Documentation of the integration architecture and setup guide.
- Client-side WebSocket code ready for deployment.
- One training session for your team (up to 2 hours).
- Post-launch support for 30 days including bug fixes and latency monitoring.
Contact us for a consultation. Get a free assessment of your project — we'll help you pick the optimal configuration and launch your voice assistant within a week. Order a pilot project to test the solution on your data.







