Automatic Court Hearing Transcription: >98% Accuracy

Implementation of Automatic Court Hearing Transcription

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1301
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1267
  • image_logo-advance_0.webp
    B2B Advance company logo design
    713
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1003
  • image_logo-aider_0.webp
    AIDER company logo development
    943
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1056

Implementation of Automatic Court Hearing Transcription

A secretary spends up to 40% of their time on transcription: each minute of audio requires 5–10 minutes of manual work. A mistake in the protocol risks overturning a decision. Our automatic court hearing transcription system combines on-premise ASR with speaker diarization and legal vocabulary to achieve WER less than 2% on clean recordings. The model is fine-tuned on a corpus of 100+ hours of court hearings, including regional specifics. This is not just recognition — it is a complete ML pipeline for courts, ensuring confidentiality and accuracy that cloud services cannot match.

Why On-Premise Is Safer Than Cloud

Cloud services transmit audio to a third party — a violation of the secrecy of the deliberation room (Article 241 of the Criminal Procedure Code). On-premise architecture guarantees that data never leaves the court's perimeter. Moreover, commercial STT services yield 10–15% WER on legal vocabulary, while our solution achieves <2% (5–10 times more accurate). Our on-premise solution outperforms cloud services by 5–10x in accuracy on legal vocabulary.

How We Achieve >98% Accuracy

We take Whisper large-v3 and fine-tune it on a corpus of court hearings (100+ hours, 20+ courts). We use LoRA adapters for rapid adaptation to a specific speaker's voice. The dictionary includes 5,000+ legal terms and patterns: "article one hundred fifty-two" → "Art. 152", "part one of article" → "Part 1 of Art.". The normalizer also processes dates, names, and abbreviations (CPC, CCP, APC).

class LegalTextNormalizer: def normalize(self, text: str) -> str: text = re.sub(r'article (\d+)', r'Art. \1', text) text = re.sub(r'part (\w+)', lambda m: f'Part {ROMAN_TO_INT[m.group(1)]}', text) return text 

Additionally, we use Voice Activity Detection (VAD) to filter noise and pauses, reducing WER by another 1–2%. Without VAD, the model "hears" background conversations and generates phantom phrases.

How Much Can You Save on Transcription?

Metric Manual transcription Our system
Time per 1 hour of audio 5–10 hours 15–20 minutes (post-editing)
Accuracy 100% (but slow) >98% (WER<2%)
Secretary workload 40% of time 5–10% (only control)

Payroll savings at typical load amount to over $11k–16k per year for a court with 10 judges. The exact value depends on the volume of hearings and region.

How Does Diarization Handle Overlapping Speech?

We use pyannote/speaker-diarization-3.1 with a threshold of 0.6. When overlapping occurs, the label SPEAKER_00+SPEAKER_01 is assigned, and the utterance is flagged for manual verification. Attribution accuracy is 92% for 2–4 participants, 85% for 5+.

Что входит в работу

  • Fine-tuning Whisper on your recordings (5–10 hours, 2–3 weeks).
  • On-premise deployment on a server with GPU (NVIDIA A10G, L40S, etc.).
  • Integration with GAS Justice via REST API or XML exchange.
  • Training secretaries in post-editing (1 day).
  • Handover of model source code, configs, and documentation.
  • Access to model repository and support for 6 months.

On-Premise vs Cloud STT: Key Differences

Parameter On-premise (our solution) Cloud STT
Confidentiality Data stays within perimeter Audio sent to third party
WER on legal vocabulary <2% 10–15%
Customer-specific fine-tuning Yes No
Diarization PyAnnote 3.1 (92% accuracy) Basic (70–80%)
Integration with GAS Certified module Requires adapter

Typical Mistakes and How to Avoid Them

  • Skipping VAD filtering: The model "hears" noise and generates phantom phrases — WER jumps to 20%. Our VAD removes silence and noise, leaving only speech.
  • Using the model without fine-tuning: Standard Whisper is not adapted for legal vocabulary, yielding 10–15% WER. Fine-tuning is mandatory to achieve the stated accuracy.
  • Ignoring post-editing: Even at 98% accuracy, manual control is needed for complex sections. We include a post-editing interface that highlights uncertain fragments — this speeds up verification by 2–3 times.
Implementation Process
  1. Infrastructure audit and collection of 5+ hours of audio.
  2. Data labeling and fine-tuning (3–4 weeks).
  3. Development of integration modules (2–3 weeks).
  4. Testing on a control sample (1 week).
  5. Deployment and staff training (1 week).
  6. Pilot operation with our support (2 weeks).

Timelines and How to Get Started

Basic system: from 4 weeks. Full cycle with fine-tuning and GAS integration: up to 12 weeks. Contact us for an individual proposal — we will assess your project and prepare a turnkey quote. Order a demo and verify accuracy on your own recordings — get a free engineer consultation.

Architecture described in Whisper and pyannote/speaker-diarization-3.1. Our company has 5+ years of NLP experience and 20+ transcription projects for courts and law firms.