Implementing AI Video Quality Enhancement in Mobile Apps
Clients often come with video files shot on old cameras or compressed to save space: 480p, 1 Mbps bitrate, noise in shadows. The task is to boost to 1080p, remove compression artifacts, stabilize shakes. Photo processing won't work: video is 30 frames per second, each frame must be processed quickly while maintaining consistency between them.
The solution splits into two modes: post-processing (a 1-minute clip in 2–3 minutes) and real-time (33 ms per frame at 30 fps). We have implemented dozens of such pipelines for iOS and Android — from simple upscaling to complex temporal enhancement. With 10+ years of experience, over 50 successful projects, and 5 years on the market, we guarantee quality.
Order AI video enhancement development: we'll select the architecture for your scenario.
How to Choose a Model for a Mobile Device?
The model choice is determined by the target resolution and available hardware. For post-processing we use per-frame models (e.g., Real-ESRGAN) — they are easy to deploy but suffer from temporal flickering.
Temporal models (BasicVSR++ with a 3–5 frame window) deliver smooth results but require batch processing and more memory. Temporal models eliminate flickering, achieving 2x better visual quality than per-frame models. On mobile devices we use lightweight versions with 256×256 tiling. For real-time — lightweight custom models (2–4 MB) designed for 360–480p input and accelerated via Core ML or TFLite GPU Delegate.
Implementation Steps
- Choose the model type based on target resolution and hardware: per-frame for simple upscaling, temporal for flicker-free results.
- Set up video decoding with AVAssetReader (iOS) or MediaCodec (Android), ensuring efficient color space conversion.
- Implement ML inference using Core ML on iOS or TFLite GPU Delegate on Android, optimizing for tiling if needed.
- Apply temporal post-processing: average activations across 3–5 neighboring frames to eliminate flickering.
- Encode the processed frames with AVAssetWriter or MediaCodec, and copy the original audio track with PTS synchronization.
- Test on multiple devices and edge cases, including non-standard resolutions and rotation.
Implementation on iOS and Android
iOS: We use AVAssetReader to decode into CVPixelBuffer, convert YUV→RGB via Metal shader, run through a Core ML model, convert back, and write with AVAssetWriter. Conversion via Metal is critical — on CPU it takes 15–20 ms per frame just for color space.
For denoising we use Real-ESRGAN or temporal models. Temporal flickering is removed by post-processing: averaging activations from neighboring frames with a weight of 0.1–0.2.
Android: Decoding via MediaCodec into a Surface (OpenGL texture), processing with TFLite GPU Delegate (works directly with textures via setExternalContext()). For Full HD, 256×256 tiling yields ~48 tiles; at 15 ms/tile inference — 720 ms per frame. For quick-enhance, a lightweight model without tiling, downscale to 540p.
Why Temporal Consistency Matters
Independent processing of each frame leads to flickering: details appear and disappear at boundaries. We solve this in two ways: using temporal models with a 3–5 frame window or adding post-processing with activation averaging.
The result is natural-looking video without flickering. Temporal consistency is the key differentiator of a professional solution from simple upscaling.
Real-time and Audio
For real-time we work at reduced resolution: on iPhone 14 Pro — ESRGAN x2 at 480p (~28 ms via ANE), on Snapdragon 8 Gen 2 — via GPU Delegate. We use CameraX with ImageAnalysis and KEEP_ONLY_LATEST strategy.
Audio is copied unchanged: via AVAssetReaderTrackOutput or MediaExtractor, PTS synchronization one-to-one. A common mistake is forgetting to synchronize PTS, which leads to audio drift.
Comparison of Approaches
Comparison of Approaches
| Parameter | Per-frame Model | Temporal Model (5-frame window) |
|---|---|---|
| Quality | Good, but flickering | Excellent, no artifacts |
| Speed on Full HD | ~720 ms/frame | ~1.5 s/frame (batch of 5) |
| Memory | ~50 MB | ~200 MB |
| Deployment Complexity | High | Medium, requires fine-tuning |
What's Included in the Work
| Stage | Duration | Description |
|---|---|---|
| Scenario Analysis | 1–2 days | Define modes (post-processing/real-time), target devices, measure performance on reference devices. |
| Pipeline Design | 2–3 days | Select model, tiling format, decode/encode architecture, audio synchronization method. |
| Implementation | 3–8 weeks | Code decoder, ML inference, encoder, temporal post-processing, integration into the project. |
| Testing | 1–2 weeks | Verify on 10+ devices (including HDR, non-standard resolutions, rotation), stress test edge cases. |
| Deployment & Documentation | 2–3 days | Code Signing, TestFlight / Play Console, API description and maintenance recommendations. |
Timeline Estimates
Post-processing on a single platform with a per-frame model — 3–5 weeks. Both platforms with temporal consistency and real-time mode — 8–14 weeks. Typical post-processing pipeline costs between $8,000 and $15,000.
With 10+ years of experience, over 50 successful projects, and 5 years on the market, we guarantee quality. Get a consultation for your project — we'll analyze your scenarios, select the optimal architecture, and provide accurate timelines. We work turnkey with a result guarantee.







