Deepgram integration for low-latency streaming STT

We integrate Deepgram for streaming speech recognition with latency under 200 ms. When your product needs real-time STT—live subtitles, voice assistants, call analytics—standard solutions like Google Speech-to-Text deliver 500-800 ms latency and require post-processing. Deepgram Nova-2 outputs resul

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1301
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1264
  • image_logo-advance_0.webp
    B2B Advance company logo design
    712
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1002
  • image_logo-aider_0.webp
    AIDER company logo development
    942
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1056

We integrate Deepgram for streaming speech recognition with latency under 200 ms. When your product needs real-time STT—live subtitles, voice assistants, call analytics—standard solutions like Google Speech-to-Text deliver 500-800 ms latency and require post-processing. Deepgram Nova-2 outputs results 60% faster with comparable WER of 5-8% on English and 12-18% on Russian (beta).

What problems does Deepgram solve?

High latency. Many cloud STT providers buffer audio and send it in packets, adding delay. Deepgram offers WebSocket streaming with immediate intermediate and final results. We design the architecture so p99 latency stays under 250 ms.

Russian language quality. For Russian, Deepgram is still beta with WER 12-18%, higher than English. To reduce error rates, we calibrate the language model (domain-specific tuning), add custom keywords, and apply rule-based post-processing.

Diarization and analytics. Determining "who spoke when" in multi-channel calls is nontrivial. Deepgram supports channel and speaker diarization but requires voice profile tuning. We integrate with call metadata.

Why Nova-2 outperforms base models

Nova-2 processes audio 30 times faster than real time (30x RT)—3-5x faster than the Base model with the same WER. This is achieved through an end-to-end attention architecture that eliminates separate decoding stages. For comparison: Google Chirp delivers 10x RT, Whisper large 3x RT. Deepgram wins on speed, sacrificing accuracy on rare languages.

How to reduce streaming latency

Key parameters:

  • Use WebSocket instead of REST (REST adds up to 500 ms round-trip)
  • Disable buffering options (e.g., punctuate=false, interim_results=true)
  • Reduce chunk size to 10-20 ms (4096 bytes for 16 kHz audio)
  • Choose the Base model if accuracy is not critical

In a live webinar transcoding project, we reduced latency from 600 ms to 180 ms using these optimizations.

How we integrate Deepgram: stack and approach

The basic scenario is WebSocket integration with a persistent connection. We use Python asyncio with the websockets library and the official Deepgram SDK for authentication. Example streaming code (Nova-2, Russian, diarization):

import asyncio import websockets import json async def transcribe_stream(): url = "wss://api.deepgram.com/v1/listen" headers = {"Authorization": f"Token {DEEPGRAM_API_KEY}"} params = "?model=nova-2&language=ru&punctuate=true&diarize=true" async with websockets.connect(url + params, extra_headers=headers) as ws: async def send_audio(): with open("audio.wav", "rb") as f: while chunk := f.read(4096): await ws.send(chunk) await ws.send(json.dumps({"type": "CloseStream"})) async def receive_results(): async for message in ws: result = json.loads(message) if result.get("is_final"): transcript = result["channel"]["alternatives"][0]["transcript"] print(transcript) await asyncio.gather(send_audio(), receive_results()) 

We also configure parameters: utterances=true for phrase splitting, numerals=true for digits, smart_format=true for punctuation and symbols.

What the work includes

  • Audit of current architecture—evaluate latency, language, and volume requirements.
  • Integration design—select model, protocol (REST/WebSocket), authentication scheme.
  • Implementation—write the module for your backend (Python, Node.js, Go, Java).
  • Testing—A/B testing on a test dataset, measuring p50/p95/p99 latency.
  • Documentation—API description, configs, deployment guide.
  • Team training—workshop on operation and monitoring.

Process

  1. Analytics (2-3 days)—gather requirements, select Deepgram model, evaluate quality on your audio.
  2. Design (2-4 days)—develop architecture, agree on protocols and error handling (reconnect, backpressure).
  3. Implementation (5-10 days)—WebSocket integration, diarization setup, post-processing.
  4. Testing (3-5 days)—run on test data, optimize parameters, load test.
  5. Deployment and handover (2-3 days)—deploy in your environment (AWS/GCP/on-prem), hand over documentation.

Timeline and cost

Timeline: 2 to 4 weeks, depending on complexity (streaming vs batch, custom model needs, diarization). Cost is calculated individually based on work volume and selected stack. We start with a free technical audit—evaluate your current architecture and provide a preliminary plan. Contact us to discuss your project. Get a consultation—we'll send details on timeline and cost.

Deepgram vs alternatives

Parameter Deepgram Nova-2 Google STT Whisper large
WER (English) 5-8% 6-10% 8-12%
Streaming latency 100-200 ms 400-800 ms 500-1000 ms
Real-time factor 30x RT 10x RT 3x RT
Russian support beta (12-18%) full (10-15%) 99+ languages
Diarization built-in separate setup no
Price per minute $0.0043 (Nova-2) $0.006 $0.000 (self-hosted)

Based on real-time speech experience, we guarantee integration with p99 latency under 300 ms and accuracy comparable to reference models. Our engineers hold Deepgram and AWS certifications, ensuring reliable solutions.

Low-latency configuration parameters
Parameter Default Recommendation
punctuate false true (if punctuation matters)
interim_results false true (for interim results)
chunk_size 8192 bytes 4096 bytes (16 kHz)
model nova-2 base (if accuracy is not critical)

According to Deepgram, the Nova-2 model saves up to 40% in costs compared to Google STT under high loads. Contact us for a free audit of your project.