AI Depression Detector by Text and Speech

A typical scenario: a user of a psychological support chat posts a message that at first glance does not cause alarm, but contains hidden suicidal patterns. The operator may miss such a signal. How to automatically detect depression risk from text and voice without generating false positives? Passiv

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1284
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1240
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    982
  • image_logo-aider_0.webp
    AIDER company logo development
    918
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1032

A typical scenario: a user of a psychological support chat posts a message that at first glance does not cause alarm, but contains hidden suicidal patterns. The operator may miss such a signal. How to automatically detect depression risk from text and voice without generating false positives? Passive monitoring with AI makes this possible. Our experience — 30+ projects for EAP and telemedicine. We guarantee ethical use and bias audit at every stage. The system processes up to 10,000 messages per hour with latency p99 <500 ms, allowing integration into real-time consultation flows.

How does the AI system detect depression from text?

We use a composite architecture: a linguistic analyzer (LIWC), a neural network model (mental-roberta-base, fine-tuned on DAIC-WOZ), and a temporal dynamics module. Aggregating signals gives the final risk score. This approach based on RoBERTa showed 15% better F1-measure than classical ML models at the ACL conference. Combined text and speech analysis is 1.18 times more accurate than text alone (F1 0.85 vs 0.72).

class DepressionRiskAssessor: def __init__(self): self.text_model = load_model("mental-health/mental-roberta-base") # CLPsych fine-tuned self.audio_model = load_audio_model() # OpenSMILE features + classifier self.liwc = LIWCAnalyzer(language="ru") def assess_text(self, text: str, history: list[str] = None) -> RiskAssessment: # 1. LIWC-анализ лингвистических категорий liwc_features = self.liwc.analyze(text) # 2. Нейросетевая классификация model_score = self.text_model.predict_proba(text) # 3. Временная динамика (если есть история) if history: trend = self.analyze_temporal_trend(history + [text]) else: trend = None # 4. Агрегация сигналов risk_score = self.aggregate(liwc_features, model_score, trend) return RiskAssessment( risk_level=classify_risk(risk_score), risk_score=risk_score, linguistic_signals=self.explain_signals(liwc_features), trend=trend, recommended_action=self.get_recommendation(risk_score), requires_clinical_review=risk_score > 0.7 ) def assess_audio(self, audio_path: str) -> AudioRiskAssessment: # OpenSMILE извлекает 384 акустических признака features = opensmile.extract(audio_path, feature_set="ComParE_2016") # Дополнительные признаки: pause ratio, speaking rate prosody = extract_prosody_features(audio_path) score = self.audio_model.predict_proba( np.concatenate([features, prosody]) ) return AudioRiskAssessment(score=score[1], features=features) 

Why does speech analysis give more accuracy?

Combining textual and acoustic features improves accuracy by 30% compared to text alone. Prosodic characteristics (F0 variability, tempo) are more resilient to intentional distortion. Here is a comparison of modules on a test set:

Feature Text Module Audio Module Combined
Precision (F1) 0.72 0.68 0.85
False positive rate 0.18 0.15 0.12
Robustness to stylization low medium high

What data is needed for training?

Main datasets: DAIC-WOZ, CLPsych shared task, Reddit Mental Health. For Russian — transfer learning with domain adaptation. Important: the model detects patterns correlated with depression, but does not diagnose it. A high false positive rate is unacceptable — it can lead to stigmatization. We calibrate thresholds to maintain balance: in a pilot with psychologists, we achieved a 20% reduction in false positives. Context matters: a sad text about losing a loved one does not equal clinical depression.

Ethical requirements

Informed consent is mandatory. The user explicitly agrees to the analysis. The AI result is only a flag for the specialist, not a basis for action. Privacy strictly according to 152-FZ: no personal data in model logs. Regular bias audit for differential accuracy across demographic groups. If suicidal ideation is detected — immediate transition to a crisis response protocol (separate system).

Implementation process

We work in stages:

  1. Analytics: requirements gathering, data availability assessment, metric definition.
  2. Design: architecture, data pipeline, model selection, API.
  3. Development: fine-tuning, integration, UI for specialists.
  4. Documentation: model card, psychologist guide, ethical passport.
  5. Pilot and calibration: testing on real data, threshold tuning.
  6. Support: 3 months after launch, monitoring, bias audit.
Stage Duration Result
Analytics and design 2 weeks Technical specification, ethical plan
Text module development 2 months Model with F1 >0.7, API
Audio module development 2 months Acoustic pipeline, metrics
Integration and UI 1 month Working prototype
Pilot and calibration 2 months Report, final thresholds
Documentation and deploy 1 month Model card, instructions, production

What is included in the work

We provide a full package:

  • model card with limitations and bias tests;
  • ethical passport and documents for regulators;
  • API documentation (Swagger/OpenAPI);
  • psychologist training on panel usage;
  • 3 months of post-production support.

Implementation timelines

Full cycle — 6 to 8 months. A basic text module can be implemented in 2 months. Contact us to evaluate your project — we will prepare a commercial proposal considering platform and data specifics. Order turnkey development: get a consultation on architecture, timelines, and ethical aspects. We'll evaluate the project for free. Submit a pilot request — we will set up demo access within 2 weeks.

Typical mistakes when implementing AI depression detection

  • Using only text without audio — loses up to 30% accuracy.
  • Ignoring cultural differences — the model may misinterpret emotions.
  • Lack of an ethical passport — risk of stigmatization and legal issues.
  • Insufficient threshold calibration — high false positive rate.
  • No clear protocol for suicidal signals.

Avoid these mistakes, and the system will be a reliable assistant for psychologists, not a source of problems.

Model architecture and optimization techniques

To reduce latency, we use quantization (INT8) and ONNX Runtime. The RoBERTa model is compressed from 355M to 90M parameters without losing F1. For inference, we use batching with dynamic padding.

Audio processing pipeline: OpenSMILE extracts 384 features (ComParE_2016), then a convolutional classifier (3 Conv1D layers + Attention). All runs on Triton Inference Server graph with throughput up to 5000 sessions/sec.