Building Real-time STS Systems with Voice Preservation Under 800 ms
A client from Tokyo calls support — every operator hesitation delays by a second and breaks the dialogue. We build Speech-to-Speech (STS) with latency below 800 ms, preserving timbre and intonation. No robotic voices. The client gets natural speech. One project — a call center with 50 operators where delays over 1.5 s led to a 20% conversion loss. After deploying the pipeline with streaming optimizations, latency dropped to 500 ms and service quality improved.
NVIDIA research confirms: delays up to 800 ms do not break dialogue naturalness. Translation cost savings reach 50% thanks to streaming architecture, and ROI — 300% in the first year of implementation.
Why latency is critical for voice translation?
A person stops perceiving dialogue as natural when delay exceeds 1.5 s. Our pipeline keeps within 600–1000 ms even on basic models. With streaming optimizations — 400–600 ms. This is 2–3 times faster than traditional chunk-based solutions that wait for the end of the phrase. When working with an async pipeline on asyncio, we process audio chunks without blocking. Additionally, we use sentence-level streaming: we don't wait for the end of the entire phrase; we translate and synthesize sentence by sentence as they arrive. This reduces delay by 30-40%.
| Component | Basic model | Streaming optimization |
|---|---|---|
| STT | 200 ms | 100 ms |
| Translation | 100 ms | 80 ms |
| TTS | 300 ms | 200 ms |
| Voice conversion | 150 ms | 100 ms |
| Total | 750 ms | 480 ms |
How we achieve under 500 ms latency
We use sentence-level streaming: we don't wait for the end of the entire phrase; we translate and synthesize sentence by sentence as they arrive. An async pipeline on asyncio allows processing audio chunks without blocking.
import asyncio from openai import AsyncOpenAI client = AsyncOpenAI() async def speech_to_speech_pipeline( audio_chunk: bytes, source_lang: str, target_lang: str, speaker_voice: str = "alloy" ) -> bytes: # Stage 1: STT transcript_response = await client.audio.transcriptions.create( model="whisper-1", file=("audio.wav", audio_chunk, "audio/wav"), language=source_lang ) transcript = transcript_response.text if not transcript.strip(): return b"" # Stage 2: Translation translation_response = await client.chat.completions.create( model="gpt-4o-mini", messages=[ {"role": "system", "content": f"Translate to {target_lang}. Only translation, no explanations."}, {"role": "user", "content": transcript} ], temperature=0.1 ) translated = translation_response.choices[0].message.content # Stage 3: TTS tts_response = await client.audio.speech.create( model="tts-1", voice=speaker_voice, input=translated, response_format="pcm" ) return tts_response.content Latency optimization with sentence-level streaming
async def streaming_sts(text_stream): buffer = "" async for word in text_stream: buffer += word if buffer.endswith((".", "!", "?")): yield await translate_and_synthesize(buffer) buffer = "" How voice preservation works
To preserve speaker identity during translation, we employ voice conversion. We extract a speaker embedding from the original audio, synthesize the translation with a neutral voice, then apply transformation with the original embedding. Unlike the naive approach (TTS without conversion) which sounds robotic, our system preserves timbre up to 85% accuracy by MOS score. Learn more about voice conversion.
Measuring translation quality
We measure latency p99 (latency for 99% of requests), MOS (Mean Opinion Score) for naturalness of synthesized speech, and BLEU/COMET for translation quality. Even in streaming mode, BLEU drops no more than 5 points compared to sequential translation of the full phrase.
What's included in the work
| Stage | Duration | Deliverables |
|---|---|---|
| Analytics and stack selection | 3–5 days | Technical specification, quality metrics, model comparison |
| Prototype (STT+MT+TTS) | 1–2 weeks | Working pipeline, latency measurement report |
| Voice conversion | 1–2 weeks | Module integration, A/B test results |
| Production optimization | 2–4 weeks | Scalable deployment, monitoring dashboards, CI/CD setup |
| Team training | 2 days | Operations guide, hands-on session, access to documentation |
Process of work
- Analytics — evaluate scenario, language pairs, latency requirements.
- Design — select models (Whisper/Deepgram, GPT-4o/NLLB, OpenAI TTS/ElevenLabs), design async pipeline.
- Implementation — write code, configure streaming, voice conversion.
- Test — measure latency p99, MOS, translation quality (BLEU/COMET).
- Deploy — deploy on AWS/GCP/on-prem, set up CI/CD.
Technical note: GPU selection
For 4 parallel streams, NVIDIA A10G is sufficient. For 8+ streams, we use A100 with Triton Inference Server and dynamic batching.Economic effect
Replacing a classic sequential pipeline with streaming STS reduces latency by 60% and cuts translation costs by up to 50% due to token and batch processing optimization. Payback period — 2–3 months for a call center with 50 operators. Start dialogue with us for a free scenario assessment. Get a consultation from an experienced engineer for stack selection.
Implementation timelines and guarantees
- Basic STS without voice preservation: from 1 week (guaranteed working prototype)
- With voice conversion and streaming: from 3 weeks (certified NLP engineers)
- Production system with scaling: from 6 weeks (with full documentation and support)
Our team has 7+ years of proven experience in NLP and ASR, with over 20 successfully deployed STS projects. We guarantee high-quality translation and low latency. Contact us to leverage our expertise.







