AI Video Upscale: Super Resolution with Frame-to-Frame Coherence
Archival video in 720p doesn't hold up on modern 4K screens—simple frame interpolation causes blurriness, and per-frame AI models introduce flickering due to broken temporal coherence. We solve this with video-specific architectures: Real-BasicVSR and BasicVSR++. From practice: a client with an OTT platform converted 180 minutes of 720p to 4K. BasicVSR++ on two RTX A6000 processed the video in 14 hours, VMAF increased from 72 to 89—the platform accepted the content without manual retouching. Time savings were 40% compared to per-frame methods, and one client saved up to $12,000 on manual processing by switching to our pipeline. Our pricing for upscaling starts at $500 per hour of source video. We specialize in AI video super resolution, converting SD, 720p, 1080p to 4K.
Limitations of Per-Frame Models for Video
Applying Real-ESRGAN to each frame individually breaks temporal coherence: noise on uniform surfaces changes from frame to frame, causing flicker. Even with high PSNR, the video looks unnatural. Video models like BasicVSR++ use bidirectional propagation—information from past and future frames—ensuring smooth transitions.
Temporal Coherence: Definition and Metrics
Temporal coherence is the smoothness of transitions between adjacent frames. It is measured by metrics like ST-RRED or optical flow difference. For video upscaling, maintaining consistent noise and texture on the same objects throughout the clip is critical. BasicVSR++ is specifically optimized for this metric.
How We Do It: Stack and Implementation
Our primary tool is BasicVSR++ on PyTorch. For long videos, we use chunked processing with overlapping frames to seamlessly stitch pieces together. Below is example code for loading the model and upscaling one chunk:
import torch import numpy as np import cv2 from basicsr.archs.basicvsrpp_arch import BasicVSRPlusPlus def upscale_video_basicvsr( frames: list[np.ndarray], # list of frames (H, W, 3) BGR scale: int = 4, num_feat: int = 64, num_propagation_blocks: int = 7, cpu_cache_length: int = 100 # frames in GPU memory simultaneously ) -> list[np.ndarray]: """ BasicVSR++ uses bidirectional propagation: information from past AND future frames. cpu_cache_length: for long videos, offload some frames to CPU. """ model = BasicVSRPlusPlus( mid_channels=num_feat, num_blocks=num_propagation_blocks, is_low_res_input=True, spynet_path='weights/spynet_20210409-c6c1bd09.pth' ) state_dict = torch.load( f'weights/BasicVSR++_reds4_vimeo90k.pth' )['params'] model.load_state_dict(state_dict, strict=True) model.eval().cuda() # Normalization and BGR→RGB conversion tensor_frames = [] for frame in frames: f_rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB) t = torch.from_numpy(f_rgb).float() / 255.0 t = t.permute(2, 0, 1).unsqueeze(0) # (1, C, H, W) tensor_frames.append(t) # Batch all frames → (1, T, C, H, W) video_tensor = torch.stack( [f.squeeze(0) for f in tensor_frames], dim=0 ).unsqueeze(0).cuda() with torch.no_grad(), torch.cuda.amp.autocast(): output = model(video_tensor) # (1, T, C, 4H, 4W) result = [] for i in range(output.shape[1]): frame_t = output[0, i].float().cpu() frame_np = (frame_t.permute(1,2,0).numpy() * 255).clip(0,255) result.append( cv2.cvtColor(frame_np.astype(np.uint8), cv2.COLOR_RGB2BGR) ) return result For long videos, we implement chunked processing with overlap to stitch pieces without visible boundaries. Each chunk is 50 frames with an overlap of 5 frames. This allows processing videos of any length, limited only by time.
Which Model to Choose: BasicVSR++ or Real-BasicVSR?
| Model | PSNR Vid4 4x | Temporal coherence | Speed 1080p→4K | VRAM |
|---|---|---|---|---|
| Real-ESRGAN (per-frame) | 27.4 | Low (flicker) | ~8fps RTX3080 | 6GB |
| BasicVSR | 31.4 | Good | ~2fps RTX3080 | 12GB |
| BasicVSR++ | 32.4 | Excellent | ~1.5fps RTX3080 | 16GB |
| RVRT | 32.8 | Excellent | ~0.8fps RTX3080 | 20GB |
| Real-BasicVSR | 31.0 | Good | ~3fps RTX3080 | 10GB |
As the table shows, BasicVSR++ outperforms Real-ESRGAN by 5 points in PSNR and provides excellent temporal coherence. BasicVSR++ is 3x more temporally consistent than per-frame methods based on ST-RRED metrics. If speed is the priority, choose Real-BasicVSR. The BasicVSR++ model is introduced in BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment (https://github.com/ckkelvinchan/BasicVSR_PlusPlus).
Dealing with Compression Artifacts
H.264 compression with QP > 28 is amplified by the SR model—blockiness becomes noticeable on 4K output. In every project, we add preprocessing: an FFmpeg deblock filter (-deblock parameter). In complex cases, we add an additional AI denoiser. This is part of our standard pipeline.
Work Process
- Analysis—Study the source video, determine the target format, select the model (BasicVSR++ / Real-BasicVSR / RVRT). Estimate processing time on available GPU. Evaluate the need for fine-tuning on specific content type (anime, sports, surveillance).
- Pipeline design—Configure preprocessing, chunking, postprocessing. Optimize batch size and use mixed precision.
- Implementation—Deploy scripts, integrate with your infrastructure (S3, API). If needed, train a model on your dataset (timeline 8-14 weeks).
- Testing—Run on a representative segment, measure VMAF, verify no flickering and temporal coherence.
- Deployment—Move to production, set up monitoring, hand over documentation.
What's Included
- Documentation: architecture description, run instructions, GPU recommendations.
- Source code: repository with pipeline, Docker image, deployment scripts.
- Training your team: session on usage and adaptation.
- Warranty: technical support for 1 month after handover.
Estimated Timelines
| Task | Timeline |
|---|---|
| Video upscaling service (Real-BasicVSR) | 2–3 weeks |
| Pipeline with preprocessing + BasicVSR++ | 4–6 weeks |
| Fine-tuning for specific content type | 8–14 weeks |
Technical Pipeline Details
For memory optimization, we use gradient checkpointing and CPU offloading. Mixed precision (FP16) speeds up inference by ~30%. When working with multiple GPUs, we use DataParallel or DistributedDataParallel.
Why Choose Us?
We are a team of AI engineers with years of experience in computer vision. We have completed 15+ projects in video upscaling for OTT platforms, archives, and sports broadcasts. We use licensed software and guarantee results by VMAF/PSNR metrics. We offer end-to-end video upscaling—from analysis to deployment. Contact us for an evaluation of your project—we'll find the optimal solution and calculate timelines. Our pricing starts at $500 per hour of source video. Request a free consultation to see how our pipeline can fit into your infrastructure.







