Voice Assistant Implementation in Mobile Apps with ≤1.5s Latency

A user dictates a command, waits for a response — and after 3 seconds gets the wrong thing. Sound familiar? In a voice assistant, speed and recognition accuracy are paramount. We solve this by building a pipeline from VAD, STT, NLU, and TTS with a total latency from the end of the phrase to the resp

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Voice Assistant Implementation in Mobile Apps with ≤1.5s Latency
Complex
~1-2 weeks

Our competencies:

Frequently Asked Questions

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    894
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    782
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1216
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1079
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    1002
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    597

A user dictates a command, waits for a response — and after 3 seconds gets the wrong thing. Sound familiar? In a voice assistant, speed and recognition accuracy are paramount. We solve this by building a pipeline from VAD, STT, NLU, and TTS with a total latency from the end of the phrase to the response ≤1.5 seconds. This is a technical constraint we overcome by optimizing each stage: from voice detection to speech synthesis. We handle the entire cycle: design, development, integration with your CRM or ERP, and app store publication. Over 7+ years, we have deployed voice assistants in 15+ projects for iOS and Android. Evaluate our approach — contact us for a consultation.

How the Voice Assistant Pipeline Works

Microphone → VAD → STT → NLU → Logic → TTS → Speaker. Each component contributes to latency.

VAD — Voice Activity Detection, cuts silence. We use WebRTCVAD or SileroVAD (ONNX/TFLite, ~1 MB). This reduces empty STT requests and saves traffic.

Example VAD configuration on Android```kotlin val vad = SileroVAD.create(context) vad.start { frame -> if (vad.isVoice(frame)) { // send audio to STT } } ```

STT — Speech-to-Text. Options: native SFSpeechRecognizer (iOS) or Android Speech for simple scenarios; for high accuracy Russian — Yandex SpeechKit or OpenAI Whisper API. Apple Developer Documentation recommends using SFSpeechRecognizer for basic commands. Cloud recognition costs ~$0.006 per audio minute for Whisper API; on-device is free. A typical request lasts 3–5 seconds, so costs are minimal.

Why Latency ≤1.5s Is Critical

Users expect instant response. Delay over 2 seconds feels like a hang. We achieve this through parallel requests, local processing, and caching frequent intents. Each component has its typical delay:

Component Typical Latency Cost (per audio minute)
VAD (Silero) 30–50 ms $0 (on-device)
STT (Whisper API) 200–400 ms ~$0.006
NLU (Rasa) 200–400 ms $0 (self-hosted)
TTS (Yandex SpeechKit) 200–500 ms ~$0.002

Total — up to 1.5 seconds. Replacing NLU with an LLM (GPT-4) can increase latency to 4 seconds, unacceptable for real-time.

Intent Recognition: What Actually Works

For a limited domain (smart home, internet banking) — Rasa NLU or Dialogflow with 50–200 training examples per intent. For an open domain — LLM with function calling. Rasa NLU is better than Dialogflow for confidential data since it runs on your server and does not send speech to the cloud.

Feature Rasa NLU Dialogflow LLM (GPT-4)
Privacy Full Google Cloud Cloud (prompt)
Domain Accuracy 90%+ 85%+ 95%+ (but slower)
Setup Difficulty Medium Low High (prompt engineering)
Latency 200–400 ms 500–800 ms 1–4 sec

Case Study from Our Practice

Corporate assistant for field employees: voice task creation in CRM without unlocking the phone. Stack: SileroVAD on-device -> Yandex SpeechKit streaming -> Rasa NLU (self-hosted, 23 intents) -> CRM REST API -> Yandex SpeechKit TTS. Latency: median 1.1 s, p95 2.3 s. Rasa NLU provided full data control. The client estimated time savings of ~25% for employees.

How to Implement a Voice Assistant: Step-by-Step Plan

  1. Scenario Analysis — define command list and contexts (up to 3 days).
  2. Component Selection — STT, NLU, TTS considering language, privacy, and budget.
  3. Pipeline Integration — connect modules, tune VAD parameters and timeouts.
  4. Testing on Real Data — record dialogues, A/B tests, optimization.
  5. App Store Release — prepare metadata, test with TestFlight/Internal Track.

When On-Device STT Is Needed?

If the app must work offline or requires minimal latency — choose on-device. Accuracy is lower (80–90%), but latency is 300–500 ms and no API call costs. For Russian, on-device still lags behind cloud solutions but works for a limited set of phrases.

What’s Included in the Work

  • Architecture documentation and API specifications
  • Configured CI/CD for build and deployment
  • Source code and repository access
  • Training your team on the voice pipeline
  • Post-release support for 2 weeks

Estimated Timelines

  • Basic pipeline STT + NLU + TTS: 2–3 weeks
  • With wake word and context: 4–6 weeks
  • With integration into existing infrastructure: individually defined

Cost is determined after analyzing your requirements. Get a consultation for your scenario — contact us. We will evaluate your project in 1–2 days.