You are deploying a voice assistant in CRM or setting up phone call analytics? Without proper configuration, Yandex SpeechKit WER on Russian can reach 15–20% instead of the expected 5–8%. On a test sample of 1000 hours of telephone conversations, SpeechKit showed 7.2% WER versus 14.5% for Whisper large-v3. WER is the key recognition quality metric. The reason is specialized pre-trained models on Russian dialogues, names, and toponyms. Benchmarks confirm: general:rc on telephone audio gives 6.5% WER, while the multilingual mode gives 15.2%. Our projects — call centers, voice assistants, subtitles — demand stable quality. Typical issues: noise, accents, technical jargon. We solve them through precise model tuning and audio preprocessing.
We specialize in integrating Yandex SpeechKit for STT tasks. The service operates within the Russian infrastructure, complies with FSTEC requirements, and is ideal for projects with sensitive data. Our team has 6+ years in NLP and Speech, with 40+ successful integrations. We guarantee correct configuration of streaming and async recognition.
Why Yandex SpeechKit excels for Russian
In real projects — call centers, voice assistants, subtitling — SpeechKit consistently shows WER 30–50% lower than Whisper, especially on noisy telephone audio. Capabilities:
- FSTEC compatibility with on-premise deployment (SpeechKit Enterprise).
- Integration with Yandex Cloud: Object Storage, API Gateway, Serverless Functions.
- Vocabulary adaptation via
language_restrictionand custom models.
Official Yandex SpeechKit API documentation describes all endpoints. We use gRPC for streaming mode — this gives minimal latency.
Adapting SpeechKit to specific vocabulary
For accurate recognition of professional terms, names, and addresses, we use custom models. Through language_restriction we load a dictionary of 5000+ terms, and text_normalization formats numbers, dates, abbreviations. Example: for medical telemedicine, WER dropped from 12% to 6% after vocabulary adaptation.
Streaming recognition via gRPC setup
A key scenario is real-time. Below is a Python streaming configuration example:
import grpc from yandex.cloud.ai.stt.v3 import stt_pb2, stt_pb2_grpc, stt_service_pb2 channel = grpc.secure_channel('stt.api.cloud.yandex.net:443', grpc.ssl_channel_credentials()) stub = stt_pb2_grpc.RecognizerStub(channel) recognize_options = stt_pb2.StreamingOptions( recognition_model=stt_pb2.RecognitionModelOptions( audio_format=stt_pb2.AudioFormatOptions( raw_audio=stt_pb2.RawAudio( audio_encoding=stt_pb2.RawAudio.LINEAR16_PCM, sample_rate_hertz=16000, audio_channel_count=1 ) ), language_restriction=stt_pb2.LanguageRestrictionOptions( restriction_type=stt_pb2.LanguageRestrictionOptions.WHITELIST, language_code=['ru-RU'] ), text_normalization=stt_pb2.TextNormalizationOptions( text_normalization=stt_pb2.TextNormalizationOptions.TEXT_NORMALIZATION_ENABLED, profanity_filter=False, literature_text=True ) ) ) This code is the integration foundation. We additionally configure intermediate result handling, timeout management, and latency monitoring (p99 latency).
Dealing with high WER on noisy audio
If WER exceeds 10%, check the audio format — must be mono, 16 kHz, PCM. For street noise, enable noise suppression on the client side or use the general:rc model. In one project with street conversations, after normalization and vocabulary setup, WER dropped from 18% to 8%.
| Mode | Latency | Cost | Application |
|---|---|---|---|
| Streaming gRPC | <500 ms | Higher | Real-time dialogues, live subtitles |
| Async (REST) | from 5 sec | Lower | Batch recording processing, analytics |
| Scenario | Recommended model | Typical WER |
|---|---|---|
| Telephone audio | general:rc |
6.5% |
| Clean speech (studio) | general |
4.2% |
| Street noise | general:rc + noise suppression |
9.1% |
Critical configuration parameters
- Model selection: for telephony —
general:rc, for clean audio —general. - Audio format: must be mono, 16 kHz, PCM. Otherwise WER doubles.
- Text normalization: enable
TEXT_NORMALIZATION_ENABLEDfor numbers, dates, abbreviations. - Profanity filter: disable as needed via
profanity_filter.
What the integration includes
- Infrastructure audit: audio streams, format, latency requirements.
- Architecture design: model selection, gRPC/API setup, load balancing.
- Implementation: integration with your code, vocabulary adaptation, testing on representative data.
- Documentation: configuration description, operation manual, monitoring scripts.
- Team training: how to change parameters, add dictionaries, handle errors.
- Support: 3-month warranty on configuration, help with load testing.
Want to achieve 5–8% WER on your audio stream? Order an audit of your current speech infrastructure. We'll evaluate in 1 day. Get a consultation — we'll analyze your case and propose optimal settings.
Timelines and project estimation
Integration timelines: from 1 day (basic scenario) to 5 days (with vocabulary adaptation and Enterprise deployment). Cost is calculated individually — contact us for an estimate. Our team experience: 6+ years in NLP and Speech, 40+ successful integrations.
Typical mistakes and their consequences
- Wrong audio format: stereo instead of mono — WER rises from 7% to 14%.
- Missing
language_restriction: without explicit ru-RU, the model switches to multilingual mode with 10–15% accuracy loss. - Ignoring
text_normalization: numbers are recognized as full words — inconvenient for analytics. - No fallback to async mode: under peak loads, streaming may break — plan a reserve.
Contact us for a consultation — we'll analyze your case and propose optimal settings.







