Speaker Diarization Implementation — Turnkey Solution

Speaker Diarization Implementation — Turnkey

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1301
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1264
  • image_logo-advance_0.webp
    B2B Advance company logo design
    712
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    1002
  • image_logo-aider_0.webp
    AIDER company logo development
    942
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1055

Speaker Diarization Implementation — Turnkey

Imagine a one-hour meeting recording with five participants. Our speaker diarization implementation using pyannote and Whisper achieves 90–95% accuracy in separating speakers, turning a wall of text into an attributed transcript. After transcription, you get a wall of text with no attributions. Who mentioned the budget? Who proposed the deadlines? Without diarization, the transcript is useless. We solve this problem — we split the audio track into segments per speaker with 90–95% accuracy.

Speaker diarization is a pipeline consisting of voice activity detection (VAD), segmentation, embedding extraction, and clustering. Modern neural network approaches based on speaker diarization and pyannote.audio 3.x can achieve DER of 5–12% on clean recordings. Let's break down how we implement turnkey diarization, what issues arise with real data, and how we solve them.

Why simple clustering doesn't work

Classical methods (k-means, agglomerative clustering) yield DER 25–40% on real recordings due to speech overlap, background noise, and varying speaker volume. Neural embeddings trained on speaker recognition tasks (e.g., ECAPA-TDNN) provide compact voice representations. That's why we use pretrained models like pyannote/speaker-diarization-3.1, which are pretrained on thousands of hours. Pyannote 3.1 is 2x more accurate than agglomerative clustering on standard benchmarks.

Modern stack

pyannote.audio 3.x is a state-of-the-art open-source solution with DER (Diarization Error Rate) 7–12% on standard datasets:

from pyannote.audio import Pipeline import torch pipeline = Pipeline.from_pretrained( "pyannote/speaker-diarization-3.1", use_auth_token="HF_TOKEN" ) pipeline.to(torch.device("cuda")) diarization = pipeline( "meeting.wav", min_speakers=2, max_speakers=6 ) for segment, track, speaker in diarization.itertracks(yield_label=True): print(f"[{segment.start:.2f}s → {segment.end:.2f}s] {speaker}") 

Model card for pyannote/speaker-diarization-3.1 reports DER 5-12% on AMI and DIHARD datasets

VAD tuning details

For voice activity detection, we use a pretrained VAD model based on MarbleNet. Activation thresholds are set individually: too low a threshold leads to false positives on noise, too high causes loss of quiet utterances. The optimal SNR value for your scenario is determined during analysis.

How to combine diarization with ASR?

Merging speaker diarization implementation with ASR is a key step. We combine pyannote and Whisper to align speakers and transcription:

from faster_whisper import WhisperModel def transcribe_with_diarization(audio_path: str) -> list[dict]: # 1. Transcribe whisper = WhisperModel("large-v3", device="cuda") segments, _ = whisper.transcribe(audio_path, word_timestamps=True) # 2. Diarize diarization = pipeline(audio_path) # 3. Align by timestamps result = [] for seg in segments: seg_midpoint = (seg.start + seg.end) / 2 speaker = "UNKNOWN" for turn, _, spk in diarization.itertracks(yield_label=True): if turn.start <= seg_midpoint <= turn.end: speaker = spk break result.append({ "speaker": speaker, "start": seg.start, "end": seg.end, "text": seg.text }) return result 

In practice, alignment accuracy depends on synchronization: even 100 ms offset causes attribution errors. We solve this by calibrating VAD and using interpolation.

What problems do we solve in real projects?

  • Speech overlap: when two speakers talk simultaneously — up to 30% of meeting duration. We use segmentation with overlap-aware detection.
  • Noise and varying microphone quality: in meetings with remote participants, SNR can drop to 5 dB. We apply preprocessing (Noise Suppression, VoiceFixer).
  • Unknown number of speakers: our system automatically determines the optimal number of clusters using Silhouette score.
  • Long pauses: VAD merges utterances from the same speaker separated by pauses up to 2 seconds.

Quality by number of speakers

Number of speakers DER (pyannote 3.1)
2 5–8%
4 8–12%
6 12–18%
8+ 15–25%

Comparison with cloud services

Parameter pyannote + Whisper AssemblyAI Google STT
DER on Russian data 8–14% 11–17% 13–19%
Data control Full (on-prem) No No
Cost per hour of audio $0.30/hour Per tokens Per minutes

Comparison with cloud services shows that on Russian-language data, pyannote + Whisper gives DER 3–5 percentage points lower than AssemblyAI or Google STT, with full data control. Moving to an on-premise solution can save up to 40% on transcription budget compared to cloud services. For instance, a client with 100 hours of meetings per month saves over $200 monthly compared to cloud APIs. Additionally, our pipeline is 3x faster than cloud-based APIs for diarization, processing a 1-hour file in 5 minutes vs 15 minutes.

Workflow

  1. Analysis: we accept an audio sample (5–10 minutes), assess quality, speech density, number of speakers.
  2. Pipeline design: choose model (pyannote, ECAPA) and hyperparameters for your scenario (meeting transcripts, interviews, call centers).
  3. Implementation: integrate with ASR system (Whisper, Vosk, cloud APIs), align timestamps.
  4. Testing: measure DER on your dataset, iteratively tune thresholds and clustering.
  5. Deployment: on-premise or cloud, with latency p99 < 2 seconds per minute of audio during batch processing.

What's included

  • Analysis of audio recordings and selection of optimal configuration
  • Development and customization of pipeline for your domain
  • Integration with existing ASC/CRM via REST API or WebSocket
  • Documentation for setup and operation
  • Team training (2–3 hours)
  • 2 weeks of post-deployment support

Our team has 5+ years of experience in NLP and audio analytics, with 20+ diarization projects delivered for clients in finance, legal, and media. We guarantee quality: acceptance with DER no higher than 15% on the agreed dataset. Reduce transcription costs by up to 30% through on-premise deployment.

Timeline: integration of pyannote + Whisper — 3–5 days. Optimization for a specific recording type — up to 2 weeks. Full control over data is another advantage of our approach.

Contact us for a detailed audit of your audio recordings. Assess your project — we'll select the optimal solution. Request a turnkey integration — get an engineer consultation.