AI Video Upscale: Super Resolution with Temporal Consistency

AI Video Upscale: Super Resolution with Frame-to-Frame Coherence

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1285
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1241
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    983
  • image_logo-aider_0.webp
    AIDER company logo development
    919
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1033

AI Video Upscale: Super Resolution with Frame-to-Frame Coherence

Archival video in 720p doesn't hold up on modern 4K screens—simple frame interpolation causes blurriness, and per-frame AI models introduce flickering due to broken temporal coherence. We solve this with video-specific architectures: Real-BasicVSR and BasicVSR++. From practice: a client with an OTT platform converted 180 minutes of 720p to 4K. BasicVSR++ on two RTX A6000 processed the video in 14 hours, VMAF increased from 72 to 89—the platform accepted the content without manual retouching. Time savings were 40% compared to per-frame methods, and one client saved up to $12,000 on manual processing by switching to our pipeline. Our pricing for upscaling starts at $500 per hour of source video. We specialize in AI video super resolution, converting SD, 720p, 1080p to 4K.

Limitations of Per-Frame Models for Video

Applying Real-ESRGAN to each frame individually breaks temporal coherence: noise on uniform surfaces changes from frame to frame, causing flicker. Even with high PSNR, the video looks unnatural. Video models like BasicVSR++ use bidirectional propagation—information from past and future frames—ensuring smooth transitions.

Temporal Coherence: Definition and Metrics

Temporal coherence is the smoothness of transitions between adjacent frames. It is measured by metrics like ST-RRED or optical flow difference. For video upscaling, maintaining consistent noise and texture on the same objects throughout the clip is critical. BasicVSR++ is specifically optimized for this metric.

How We Do It: Stack and Implementation

Our primary tool is BasicVSR++ on PyTorch. For long videos, we use chunked processing with overlapping frames to seamlessly stitch pieces together. Below is example code for loading the model and upscaling one chunk:

import torch import numpy as np import cv2 from basicsr.archs.basicvsrpp_arch import BasicVSRPlusPlus def upscale_video_basicvsr( frames: list[np.ndarray], # list of frames (H, W, 3) BGR scale: int = 4, num_feat: int = 64, num_propagation_blocks: int = 7, cpu_cache_length: int = 100 # frames in GPU memory simultaneously ) -> list[np.ndarray]: """ BasicVSR++ uses bidirectional propagation: information from past AND future frames. cpu_cache_length: for long videos, offload some frames to CPU. """ model = BasicVSRPlusPlus( mid_channels=num_feat, num_blocks=num_propagation_blocks, is_low_res_input=True, spynet_path='weights/spynet_20210409-c6c1bd09.pth' ) state_dict = torch.load( f'weights/BasicVSR++_reds4_vimeo90k.pth' )['params'] model.load_state_dict(state_dict, strict=True) model.eval().cuda() # Normalization and BGR→RGB conversion tensor_frames = [] for frame in frames: f_rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB) t = torch.from_numpy(f_rgb).float() / 255.0 t = t.permute(2, 0, 1).unsqueeze(0) # (1, C, H, W) tensor_frames.append(t) # Batch all frames → (1, T, C, H, W) video_tensor = torch.stack( [f.squeeze(0) for f in tensor_frames], dim=0 ).unsqueeze(0).cuda() with torch.no_grad(), torch.cuda.amp.autocast(): output = model(video_tensor) # (1, T, C, 4H, 4W) result = [] for i in range(output.shape[1]): frame_t = output[0, i].float().cpu() frame_np = (frame_t.permute(1,2,0).numpy() * 255).clip(0,255) result.append( cv2.cvtColor(frame_np.astype(np.uint8), cv2.COLOR_RGB2BGR) ) return result 

For long videos, we implement chunked processing with overlap to stitch pieces without visible boundaries. Each chunk is 50 frames with an overlap of 5 frames. This allows processing videos of any length, limited only by time.

Which Model to Choose: BasicVSR++ or Real-BasicVSR?

Model PSNR Vid4 4x Temporal coherence Speed 1080p→4K VRAM
Real-ESRGAN (per-frame) 27.4 Low (flicker) ~8fps RTX3080 6GB
BasicVSR 31.4 Good ~2fps RTX3080 12GB
BasicVSR++ 32.4 Excellent ~1.5fps RTX3080 16GB
RVRT 32.8 Excellent ~0.8fps RTX3080 20GB
Real-BasicVSR 31.0 Good ~3fps RTX3080 10GB

As the table shows, BasicVSR++ outperforms Real-ESRGAN by 5 points in PSNR and provides excellent temporal coherence. BasicVSR++ is 3x more temporally consistent than per-frame methods based on ST-RRED metrics. If speed is the priority, choose Real-BasicVSR. The BasicVSR++ model is introduced in BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment (https://github.com/ckkelvinchan/BasicVSR_PlusPlus).

Dealing with Compression Artifacts

H.264 compression with QP > 28 is amplified by the SR model—blockiness becomes noticeable on 4K output. In every project, we add preprocessing: an FFmpeg deblock filter (-deblock parameter). In complex cases, we add an additional AI denoiser. This is part of our standard pipeline.

Work Process

  1. Analysis—Study the source video, determine the target format, select the model (BasicVSR++ / Real-BasicVSR / RVRT). Estimate processing time on available GPU. Evaluate the need for fine-tuning on specific content type (anime, sports, surveillance).
  2. Pipeline design—Configure preprocessing, chunking, postprocessing. Optimize batch size and use mixed precision.
  3. Implementation—Deploy scripts, integrate with your infrastructure (S3, API). If needed, train a model on your dataset (timeline 8-14 weeks).
  4. Testing—Run on a representative segment, measure VMAF, verify no flickering and temporal coherence.
  5. Deployment—Move to production, set up monitoring, hand over documentation.

What's Included

  • Documentation: architecture description, run instructions, GPU recommendations.
  • Source code: repository with pipeline, Docker image, deployment scripts.
  • Training your team: session on usage and adaptation.
  • Warranty: technical support for 1 month after handover.

Estimated Timelines

Task Timeline
Video upscaling service (Real-BasicVSR) 2–3 weeks
Pipeline with preprocessing + BasicVSR++ 4–6 weeks
Fine-tuning for specific content type 8–14 weeks
Technical Pipeline Details

For memory optimization, we use gradient checkpointing and CPU offloading. Mixed precision (FP16) speeds up inference by ~30%. When working with multiple GPUs, we use DataParallel or DistributedDataParallel.

Why Choose Us?

We are a team of AI engineers with years of experience in computer vision. We have completed 15+ projects in video upscaling for OTT platforms, archives, and sports broadcasts. We use licensed software and guarantee results by VMAF/PSNR metrics. We offer end-to-end video upscaling—from analysis to deployment. Contact us for an evaluation of your project—we'll find the optimal solution and calculate timelines. Our pricing starts at $500 per hour of source video. Request a free consultation to see how our pipeline can fit into your infrastructure.