Smart Reply for iOS and Android: From ML Kit to Custom LLMs

How Smart Reply Solves the Speed Problem in Chat Communication? We integrate Smart Reply into messengers, CRMs, and e-commerce platforms, and we see how this feature dramatically accelerates conversations. The user gets three ready-made reply options below each message — one tap is enough. Withou

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Smart Reply for iOS and Android: From ML Kit to Custom LLMs
Medium
~1-2 weeks

Our competencies:

Frequently Asked Questions

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    897
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    784
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1216
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1081
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    1004
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    598

How Smart Reply Solves the Speed Problem in Chat Communication?

We integrate Smart Reply into messengers, CRMs, and e-commerce platforms, and we see how this feature dramatically accelerates conversations. The user gets three ready-made reply options below each message — one tap is enough. Without Smart Reply, they must manually type "Okay", "Got it", "Thanks". This slows down the dialogue and increases the number of unfinished messages. Research by Google ML Kit shows that the feature boosts engagement and sent message volume by 15–30%. For businesses, this means faster query processing and higher customer loyalty. In our practice, clients save 2–3 seconds per reply, which accumulates to hours saved per day for the entire support team. In one project for an online store, implementing Smart Reply reduced the average operator response time from 45 to 15 seconds, and the number of messages per dialogue increased by 20%. If you are evaluating Smart Reply implementation, contact us for a preliminary assessment.

Why ML Kit Smart Reply Doesn't Fit Russian Language?

Google's ready-made SmartReply model works only with English. For Android, integration takes an hour:

val smartReply = SmartReply.getClient() val conversation = messages.takeLast(10).map { msg -> if (msg.isFromUser) { TextMessage.createForLocalUser(msg.text, msg.timestamp) } else { TextMessage.createForRemoteUser(msg.text, msg.timestamp, msg.senderId) } } smartReply.suggestReplies(conversation) .addOnSuccessListener { result -> if (result.status == SmartReplySuggestionResult.STATUS_SUCCESS) { val suggestions = result.suggestions.map { it.text } showSuggestions(suggestions) } } .addOnFailureListener { /* hide UI */ } 

Plus side: speed <20 ms. Minus: limited template set and only English. For iOS, the alternative via Natural Language or Apple Intelligence API exists but offers poorer functionality. If your app targets Russian-speaking audience, ML Kit is not suitable — a custom model is needed. Also consider App Store Review Guidelines: section 5.1 permits on-device ML without special permission, but custom LLMs processing user data must comply with privacy rules.

How to Implement Contextual Smart Reply on a Custom LLM?

For Russian and specific domains (support, healthcare, B2B), we use an LLM with a prompt. Example in Swift:

func generateReplySuggestions( lastMessages: [ChatMessage], count: Int = 3 ) async -> [String] { let context = lastMessages.suffix(5) .map { "\($0.role): \($0.text)" } .joined(separator: "\n") let prompt = """ You help the user quickly reply to a chat message. Dialogue history: \(context) Suggest \(count) short reply options for the user. Each reply is one sentence, maximum 10 words. Format: JSON array of strings. """ let response = try await llmClient.complete(prompt: prompt, maxTokens: 100) return parseJSONArray(response) ?? [] } 

A similar implementation on Android with Kotlin and ML Kit or a custom model:

suspend fun generateReplySuggestions( lastMessages: List<ChatMessage>, count: Int = 3 ): List<String> { val context = lastMessages.takeLast(5) .joinToString("\n") { "${it.role}: ${it.text}" } val prompt = """ You help the user quickly reply to a chat message. Dialogue history: $context Suggest $count short reply options for the user. Each reply is one sentence, maximum 10 words. Format: JSON array of strings. """ val response = llmClient.complete(prompt, maxTokens = 100) return parseJSONArray(response) ?: emptyList() } 

Latency of 1–2 seconds is acceptable. We preload options while the user reads — by the time they're ready to reply, suggestions are ready. Custom LLM is 50x slower than ML Kit but provides Russian language and context flexibility. On-device models (TensorFlow Lite, Core ML) are faster but require more memory and are less configurable.

What to Choose: ML Kit or Custom LLM?

Criterion ML Kit Smart Reply Custom LLM
Russian language support No Yes
Latency <20 ms 1–2 sec
Customization Low High (prompt, context)
Network dependency No Yes
Integration complexity Low (1 day) Medium (5–8 days)
Cost Free API or inference costs

For English and simple scenarios, ML Kit is better (50x faster). For Russian and specific needs, custom model.

How Does Smart Reply Differ on iOS and Android?

Platform Stack Native Smart Reply Custom Smart Reply
iOS Swift/SwiftUI NaturalLanguage (limited) LLM + CoreML
Android Kotlin/Compose ML Kit (English only) LLM + TensorFlow Lite

When to Show Smart Reply?

Smart Reply appears after an incoming message and disappears when the user starts typing. Three suggestions is optimal (Google Research). More overloads, less gives no choice. We use chips (horizontal scroll): MaterialChip on Android, custom Chip in SwiftUI.

// Android: hide when typing editText.addTextChangedListener(object : TextWatcher { override fun onTextChanged(s: CharSequence?, start: Int, before: Int, count: Int) { smartReplyChips.isVisible = s.isNullOrEmpty() } override fun afterTextChanged(s: Editable?) {} override fun beforeTextChanged(s: CharSequence?, start: Int, count: Int, after: Int) {} }) 

Also important to set up analytics: track how many users use suggested replies, and A/B test the number of chips.

How Does the Implementation Process Work?

We work in stages:

  1. Analyze Smart Reply usage scenarios in your app.
  2. Choose approach: ML Kit or custom LLM.
  3. Design architecture (preloading, caching).
  4. Implement on Android (Kotlin/Compose) and iOS (Swift/SwiftUI).
  5. Integrate with chat and analytics system.
  6. Code documentation and instructions for your team.
  7. Test on real dialogues.
  8. Post-implementation support: bug fixes and refinements.

Deliverables include:

  • Integration and setup documentation.
  • Access to test environment for validation.
  • Team training (1–2 hours) on using the solution.
  • One month of technical support after launch.

What Are the Timelines and Results?

Smart Reply via ML Kit (Android, English) — 1–2 days. Custom on LLM with context classification — 5–8 days. Full integration on both platforms — up to 2 weeks. Based on our projects, implementation increases engagement by 20–40% and reduces average response time by 2–3 times. For a support team of 10 people, time savings amount to up to 500 person-hours per month, converting to financial savings of $3,000 to $10,000 monthly. Contact us to discuss your app's needs.

Our Experience

We have implemented Smart Reply for messengers, CRMs, and e-commerce. We work with ML Kit and custom models. Five years in the market, over 50 projects. We guarantee correct operation on Android and iOS. Request a consultation — we'll evaluate your scenario and propose the optimal solution. Get demo access to a working example today.