We integrate Deepgram for streaming speech recognition with latency under 200 ms. When your product needs real-time STT—live subtitles, voice assistants, call analytics—standard solutions like Google Speech-to-Text deliver 500-800 ms latency and require post-processing. Deepgram Nova-2 outputs results 60% faster with comparable WER of 5-8% on English and 12-18% on Russian (beta).
What problems does Deepgram solve?
High latency. Many cloud STT providers buffer audio and send it in packets, adding delay. Deepgram offers WebSocket streaming with immediate intermediate and final results. We design the architecture so p99 latency stays under 250 ms.
Russian language quality. For Russian, Deepgram is still beta with WER 12-18%, higher than English. To reduce error rates, we calibrate the language model (domain-specific tuning), add custom keywords, and apply rule-based post-processing.
Diarization and analytics. Determining "who spoke when" in multi-channel calls is nontrivial. Deepgram supports channel and speaker diarization but requires voice profile tuning. We integrate with call metadata.
Why Nova-2 outperforms base models
Nova-2 processes audio 30 times faster than real time (30x RT)—3-5x faster than the Base model with the same WER. This is achieved through an end-to-end attention architecture that eliminates separate decoding stages. For comparison: Google Chirp delivers 10x RT, Whisper large 3x RT. Deepgram wins on speed, sacrificing accuracy on rare languages.
How to reduce streaming latency
Key parameters:
- Use WebSocket instead of REST (REST adds up to 500 ms round-trip)
- Disable buffering options (e.g.,
punctuate=false,interim_results=true) - Reduce chunk size to 10-20 ms (4096 bytes for 16 kHz audio)
- Choose the Base model if accuracy is not critical
In a live webinar transcoding project, we reduced latency from 600 ms to 180 ms using these optimizations.
How we integrate Deepgram: stack and approach
The basic scenario is WebSocket integration with a persistent connection. We use Python asyncio with the websockets library and the official Deepgram SDK for authentication. Example streaming code (Nova-2, Russian, diarization):
import asyncio import websockets import json async def transcribe_stream(): url = "wss://api.deepgram.com/v1/listen" headers = {"Authorization": f"Token {DEEPGRAM_API_KEY}"} params = "?model=nova-2&language=ru&punctuate=true&diarize=true" async with websockets.connect(url + params, extra_headers=headers) as ws: async def send_audio(): with open("audio.wav", "rb") as f: while chunk := f.read(4096): await ws.send(chunk) await ws.send(json.dumps({"type": "CloseStream"})) async def receive_results(): async for message in ws: result = json.loads(message) if result.get("is_final"): transcript = result["channel"]["alternatives"][0]["transcript"] print(transcript) await asyncio.gather(send_audio(), receive_results()) We also configure parameters: utterances=true for phrase splitting, numerals=true for digits, smart_format=true for punctuation and symbols.
What the work includes
- Audit of current architecture—evaluate latency, language, and volume requirements.
- Integration design—select model, protocol (REST/WebSocket), authentication scheme.
- Implementation—write the module for your backend (Python, Node.js, Go, Java).
- Testing—A/B testing on a test dataset, measuring p50/p95/p99 latency.
- Documentation—API description, configs, deployment guide.
- Team training—workshop on operation and monitoring.
Process
- Analytics (2-3 days)—gather requirements, select Deepgram model, evaluate quality on your audio.
- Design (2-4 days)—develop architecture, agree on protocols and error handling (reconnect, backpressure).
- Implementation (5-10 days)—WebSocket integration, diarization setup, post-processing.
- Testing (3-5 days)—run on test data, optimize parameters, load test.
- Deployment and handover (2-3 days)—deploy in your environment (AWS/GCP/on-prem), hand over documentation.
Timeline and cost
Timeline: 2 to 4 weeks, depending on complexity (streaming vs batch, custom model needs, diarization). Cost is calculated individually based on work volume and selected stack. We start with a free technical audit—evaluate your current architecture and provide a preliminary plan. Contact us to discuss your project. Get a consultation—we'll send details on timeline and cost.
Deepgram vs alternatives
| Parameter | Deepgram Nova-2 | Google STT | Whisper large |
|---|---|---|---|
| WER (English) | 5-8% | 6-10% | 8-12% |
| Streaming latency | 100-200 ms | 400-800 ms | 500-1000 ms |
| Real-time factor | 30x RT | 10x RT | 3x RT |
| Russian support | beta (12-18%) | full (10-15%) | 99+ languages |
| Diarization | built-in | separate setup | no |
| Price per minute | $0.0043 (Nova-2) | $0.006 | $0.000 (self-hosted) |
Based on real-time speech experience, we guarantee integration with p99 latency under 300 ms and accuracy comparable to reference models. Our engineers hold Deepgram and AWS certifications, ensuring reliable solutions.
Low-latency configuration parameters
| Parameter | Default | Recommendation |
|---|---|---|
punctuate |
false | true (if punctuation matters) |
interim_results |
false | true (for interim results) |
chunk_size |
8192 bytes | 4096 bytes (16 kHz) |
model |
nova-2 | base (if accuracy is not critical) |
According to Deepgram, the Nova-2 model saves up to 40% in costs compared to Google STT under high loads. Contact us for a free audit of your project.







