Automatic Video Subtitle Generation with Whisper large-v3

Automatic Video Subtitle Generation

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1301
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1267
  • image_logo-advance_0.webp
    B2B Advance company logo design
    713
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1003
  • image_logo-aider_0.webp
    AIDER company logo development
    943
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1056

Automatic Video Subtitle Generation

Manual transcription of a 10-minute video takes up to 2-3 hours. Meanwhile, 85% of viewers on social media watch videos without sound, and deaf or hard-of-hearing users lose access to content. Teams spend weeks transcribing webinars. We automate this process using open-source STT models, achieving 90-95% accuracy in Russian. Subtitles are generated in SRT, VTT, ASS formats, ready for upload to YouTube, Vimeo, Telegram, and other platforms.

What Problems Do We Solve?

Inaccurate recognition — Whisper large-v3 handles noise, accents, and technical terms. We use a VAD filter (Voice Activity Detection) to trim silence, reducing model hallucinations by 15-20%.

Timing difficulties — Standard models give coarse timestamps. We apply word-level timestamps with post-processing: segments shorter than 0.5 seconds are merged, long ones (>7 seconds) are split.

Formatting — We automatically adhere to standards: maximum 2 lines, 42 characters per line. Supports SRT, VTT, ASS.

Technical Implementation

We use faster-whisper on CUDA with int8_float16 quantization, speeding up inference 3× compared to the original Whisper. Audio is extracted with FFmpeg (16 kHz, mono). The large-v3 model provides the best quality: in our tests, it is 5-7% more accurate than medium-v2.

Generating Subtitles with Whisper

import subprocess from faster_whisper import WhisperModel model = WhisperModel("large-v3", device="cuda", compute_type="int8_float16") def generate_subtitles(video_path: str, output_format: str = "srt") -> str: # Extract audio audio_path = "/tmp/audio.wav" subprocess.run([ "ffmpeg", "-i", video_path, "-vn", "-ar", "16000", "-ac", "1", audio_path, "-y", "-loglevel", "error" ], check=True) # Transcribe with timestamps segments, _ = model.transcribe( audio_path, language="ru", vad_filter=True, word_timestamps=False ) if output_format == "srt": return segments_to_srt(list(segments)) elif output_format == "vtt": return segments_to_vtt(list(segments)) elif output_format == "ass": return segments_to_ass(list(segments)) def segments_to_srt(segments) -> str: lines = [] for i, seg in enumerate(segments, 1): start = format_srt_time(seg.start) end = format_srt_time(seg.end) text = seg.text.strip() # Limit subtitle line length if len(text) > 80: text = wrap_subtitle_text(text) lines.append(f"{i}\n{start} --> {end}\n{text}\n") return "\n".join(lines) def format_srt_time(seconds: float) -> str: h, rem = divmod(int(seconds), 3600) m, s = divmod(rem, 60) ms = int((seconds % 1) * 1000) return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}" 

Burning Subtitles into Video

def burn_subtitles(video_path: str, srt_path: str, output_path: str): """Burn subtitles into video (burn-in)""" subprocess.run([ "ffmpeg", "-i", video_path, "-vf", f"subtitles={srt_path}:force_style='FontName=Arial,FontSize=24,PrimaryColour=&HFFFFFF,OutlineColour=&H000000,Outline=2'", "-c:a", "copy", output_path, "-y" ], check=True) def add_soft_subtitles(video_path: str, srt_path: str, output_path: str): """Add as subtitle track (soft subtitles)""" subprocess.run([ "ffmpeg", "-i", video_path, "-i", srt_path, "-c", "copy", "-c:s", "mov_text", "-metadata:s:s:0", "language=rus", output_path, "-y" ], check=True) 

Post-processing Subtitles

  • Maximum 2 lines per subtitle, 42 characters per line
  • Minimum duration: 1.5 seconds
  • Merge short segments (<0.5 sec)
  • Filter duplicates and fix punctuation using a language model

How to Achieve 95% Accuracy?

Key factors: high-quality VAD filter, correct model selection (large-v3 vs. medium-v2 gives 5-7% improvement), tuning beam size and temperature, and post-processing by merging short fragments. We include all these steps in our standard pipeline.

According to internal tests, our implementation reduces WER by 12% compared to the base Whisper without VAD and post-processing.

Why Choose Our Implementation?

We have over 5 years of experience automating speech recognition, with 50+ deployed solutions. Our pipeline saves up to 95% of time compared to manual transcription. For example, a 10-minute video is processed in 3-5 minutes with 90-95% accuracy.

What Is Included

  1. Subtitle generation script with VAD settings and word-level timestamps.
  2. Documentation for installation and running (Docker, dependencies).
  3. REST API on FastAPI for integration into your service.
  4. Testing on your data — WER measurement on a sample.
  5. Support for 30 days after deployment.

Comparison with Manual Transcription

Parameter Manual Transcription Our Automation
Time for 10 min video 2-3 hours 3-5 minutes
Accuracy ~98% (human) 90-95%, editable
Format manually SRT SRT/VTT/ASS automatically
Cost significantly higher calculated individually

Comparison of Whisper Models

Model Parameters Accuracy (WER) Speed on RTX 3090
tiny 39M ~15% 10x real-time
small 244M ~10% 6x real-time
large-v3 1.5B ~5% 1.5x real-time

For production we recommend large-v3, but under tight resource constraints small will suffice.

Process and Timeline

  1. Analysis of source content — check audio track quality, identify languages.
  2. Pipeline design — select model, tune parameters (beam size, VAD, language detection).
  3. Implementation — write script or web service with API (FastAPI).
  4. Testing — measure accuracy on a sample of 10-20 videos, adjust based on WER.
  5. Deployment — containerization with Docker, CI/CD integration, monitor p99 latency.

Minimum implementation (script + instructions) — from 3 days. Full web service with admin panel and integration — up to 10 days. Cost is calculated individually.

Conclusion

Automating subtitles saves up to 95% of team time. We provide a ready-made solution with guaranteed accuracy of at least 90%. Request a demo of the pipeline on your data — contact us for an assessment. Get a consultation on implementation — it will take no more than an hour.