Imagine a mobile app that must process images, text, or sound without network access. Server latency is unacceptable, data privacy is critical. Every extra megabyte of traffic costs the user. On-device ML solves these issues, and ONNX Runtime is the key tool for cross-platform deployment. We integrate it so that the model runs equally fast on iOS and Android. Switching to on-device can reduce server infrastructure costs by up to 90%, saving hundreds of thousands of rubles monthly under high loads.
ONNX Runtime Mobile appeals with one argument: one model, both platforms. Convert PyTorch or TensorFlow to ONNX, add onnxruntime-android and onnxruntime-objc, run the same .onnx file. In practice, the difference in execution providers between iOS and Android still requires platform-specific code, but the model itself is unified. Our experience: over 5 years in mobile ML, dozens of on-device inference projects. Contact us for an assessment of your model and a preliminary quote.
How to Prepare a Model for Mobile
Standard ONNX export from PyTorch:
import torch import onnx from onnxsim import simplify # onnx-simplifier for graph optimization model = MyModel(); model.eval() dummy = torch.zeros(1, 3, 224, 224) torch.onnx.export( model, dummy, "model.onnx", opset_version=17, input_names=["input"], output_names=["output"], dynamic_axes={"input": {0: "batch_size"}, "output": {0: "batch_size"}} ) # Simplify graph — removes redundant reshapes, transposes, makes graph cleaner model_onnx = onnx.load("model.onnx") model_simplified, check = simplify(model_onnx) onnx.save(model_simplified, "model_simplified.onnx") Additionally for mobile — quantization via onnxruntime.quantization:
from onnxruntime.quantization import quantize_dynamic, QuantType quantize_dynamic( "model_simplified.onnx", "model_int8.onnx", weight_type=QuantType.QInt8 ) # Model size reduces ~4× compared to FP32 | Quantization Type | Model Size (FP32 → Int8) | Accuracy Loss | Speed on CPU |
|---|---|---|---|
| Dynamic | ~75% smaller | <1% | ~30% faster |
| Static (calibration) | ~75% smaller | 0.5-2% | ~40% faster |
Which Execution Provider Delivers Maximum Performance?
Android: NNAPI vs XNNPACK
On Android, the choice of Execution Provider depends on hardware. NNAPI delegates operations to NPU/DSP, providing up to 2x acceleration on supported ops. XNNPACK is an optimized CPU backend using SIMD instructions, speeding up to 2x on CPU but without NPU access. On a project with object detection on MediaTek Dimensity, we got 45 ms on NNAPI vs 80 ms on XNNPACK. We recommend using NNAPI for devices with NPU, XNNPACK as fallback.
iOS: CoreML Execution Provider
appendCoreMLExecutionProvider on iOS 13+ delegates supported operations to Core ML, gaining access to ANE. Operations not supported by Core ML automatically run on CPU. In tests on iPhone 12, we got 35% speedup over CPU on ResNet-50. CoreML EP is convenient for fast cross-platform deployment, but for maximum performance, consider native Core ML.
When is ONNX Runtime Better Than Native Formats?
Use ONNX Runtime for prototyping, cross-platform projects, models with custom ops that coremltools cannot convert, and frequent model updates without rebuilding the conversion pipeline. If you need maximum performance on a single platform, choose the native format: on iOS — Core ML with full ANE acceleration (usually 20–40% faster than ORT+CoreML EP), on Android — TFLite + GPU Delegate (sometimes faster than ORT+NNAPI). For single-platform deployment and critical performance, native is preferable.
How to Integrate ONNX Runtime on Android and iOS
Android: Setup and Inference
// build.gradle implementation("com.microsoft.onnxruntime:onnxruntime-android:1.18.0") // Create session val sessionOptions = OrtSession.SessionOptions().apply { // NNAPI Execution Provider for Android NPU/DSP addNnapi(NNAPIFlags.USE_FP16) // FP16 mode in NNAPI // Or: addXnnpack(mapOf()) for XNNPACK (CPU SIMD) setOptimizationLevel(OrtSession.SessionOptions.OptLevel.ALL_OPT) setIntraOpNumThreads(4) } val env = OrtEnvironment.getEnvironment() val session = env.createSession( context.assets.open("model_simplified.onnx").readBytes(), sessionOptions ) // Inference val inputTensor = OnnxTensor.createTensor( env, FloatBuffer.wrap(preprocessedArray), longArrayOf(1, 3, 224, 224) ) val results = session.run(mapOf("input" to inputTensor)) val outputArray = (results["output"]?.value as Array<FloatArray>)[0] // Resource cleanup — mandatory inputTensor.close() results.close() Leaks from unclosed OnnxTensor and OrtSession.Result are common. In Kotlin, use use {} block: results.use { ... }.
iOS: ObjC/Swift Integration
// Package.swift or Podfile: pod 'onnxruntime-objc' import onnxruntime_objc // Setup let env = try ORTEnv(loggingLevel: ORTLoggingLevel.warning) let options = try ORTSessionOptions() try options.setIntraOpNumThreads(4) // On iOS — CoreML Execution Provider try options.appendCoreMLExecutionProvider(withFlags: [.enableOnSubgraphs]) let session = try ORTSession( env: env, modelPath: Bundle.main.path(forResource: "model_simplified", ofType: "onnx")!, sessionOptions: options ) // Prepare input let inputShape: [NSNumber] = [1, 3, 224, 224] let inputData = Data(bytes: preprocessedFloats, count: preprocessedFloats.count * MemoryLayout<Float>.size) let inputTensor = try ORTValue( tensorData: NSMutableData(data: inputData), elementType: .float, shape: inputShape ) let outputs = try session.run( withInputs: ["input": inputTensor], outputNames: ["output"], runOptions: nil ) let outputTensor = outputs["output"]! let outputData = try outputTensor.tensorData() as Data let floats = outputData.withUnsafeBytes { Array($0.bindMemory(to: Float.self)) } Why Quantization is Critical for Mobile Inference
Quantization reduces model size by 4× (50 MB → 12 MB), lowers memory consumption, and speeds up CPU inference by 30-40%. Dynamic quantization does not require calibration data but yields slightly less speed gain than static. In practice, we use static quantization with a representative dataset — it gives stable improvement without significant accuracy loss (0.5-2%).
What If an Operation is Not Supported?
# Check which ops NNAPI Execution Provider supports python -m onnxruntime.tools.check_nnapi_supported_ops --model model.onnx # If an op is not supported — it runs on CPU (fallback) # This is not a crash but can nullify all NNAPI acceleration To identify bottlenecks, use the ORT Profiling API. It records per-operator timing. Enable via options.enableProfiling("ort_profile") — generates JSON viewable in Chrome chrome://tracing. Profiling on target devices helps choose the optimal execution provider. For example, on one project we switched from NNAPI to XNNPACK for a model with 80% unsupported ops, reducing inference from 300 ms to 120 ms.
What Our Work Includes
- Export and simplify ONNX graph, quantize to Int8.
- Integrate ONNX Runtime on iOS and Android with optimal Execution Providers.
- Profile performance on a fleet of 10+ real devices, including older ones.
- Compare with native formats (Core ML, TFLite) and recommend the best solution.
- Documentation for building and updating the model, integration source code.
- Guarantee stable operation and lock in inference time.
Our Experience in Mobile ML
Over 5 years deploying on-device ML in commercial applications — from retail to healthcare. Completed 20+ projects with ONNX Runtime, Core ML, and TFLite. Our engineers hold Apple and Google certifications. We guarantee the model will work on all stated devices. Get a consultation on ONNX Runtime integration — we'll assess your project and propose the optimal turnkey solution. Order ONNX Runtime integration for your app — let's discuss the details.
Timeline Estimates
Basic cross-platform ONNX Runtime integration: 2–3 weeks. With EP optimization, profiling, testing on a device fleet: 4–6 weeks. Cost calculated individually after analyzing the model and performance requirements.







