Text Analytics for Chat Dialogs: Analysis with SLA Guarantee

Text Analytics for Chat Dialogs: Analysis with SLA Guarantee

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1285
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1241
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    982
  • image_logo-aider_0.webp
    AIDER company logo development
    919
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1033

Text Analytics for Chat Dialogs: Analysis with SLA Guarantee

Operator review of dialogs consumes up to 40% of QA team time — while hidden trends remain unnoticed. We build a Text Analytics system that automatically turns hundreds of thousands of chats into structured business intelligence: not just statistics of contacts, but deep understanding — what customers are talking about, how topics change, where systemic problems arise.

Our experience: 5+ years of NLP solution development, 30+ projects for retail, fintech, and telecom. We guarantee SLA of 99.9% on processing time and classification accuracy of at least 90%. We process up to 5000 dialogs per hour on a single GPU — sufficient for Enterprise streaming.

Why Standard Reports Don't Give the Full Picture?

Traditional dashboards show the number of contacts but do not reveal context. A sharp increase in dialog count may be caused by an outage that is not visible in first-level metrics. Our system detects anomalies and links them to changes in topic, tone, and resolution — allowing response before the problem becomes widespread. On one project, we identified a 12% increase in product negativity three days before operators noticed it.

Compare two approaches: simple keyword-based frequency classification gives 50-60% accuracy and misses up to 40% of hidden patterns. Fine-tuned ML models (ruBERT, SDG) yield 80-94% F1 but require labeled data. LLM (GPT-4) achieves 95%+ without labels, but token cost is 5-10 times higher for mass processing. Our solution is hybrid: ML for routine classification (70% of volume) and LLM for complex cases (30%) — this gives the optimal price/quality ratio.

Architecture of the Analytics System

[Dialogs from helpdesk/CRM/bots] → [Enrichment: topic, sentiment, NER, summary] → [Storage: ClickHouse / BigQuery] → [Aggregation and data marts] → [Dashboards: Superset / Metabase / Grafana] + [Ad-hoc analysis: Jupyter / Python API] 

Stack: ClickHouse for storage, Celery for batches, GPU batching for embeddings. Models: ruBERT for topics, fine-tuned SDG for sentiment, NER based on SpaCy. Batch processing time for 10,000 dialogs — 2 hours on a single V100.

Comparison of Approaches to Dialog Analysis

Approach Accuracy (F1) Speed of Implementation Scalability Cost per 1K Dialogs
Keywords 50-60% Days High $0.01
ML classification (ruBERT) 90% Weeks Medium $0.05
LLM (GPT-4) 95%+ Hours Token-limited $0.50

The hybrid scheme yields 93% accuracy at $0.08 — 6 times cheaper than pure LLM with comparable quality.

Text Analytics Deployment Stages

Stage Duration Result
Analysis and design 3-5 days Architecture, requirements, KPIs
Pipeline setup 1-2 weeks Basic enrichments and data marts
Model customization 2-4 weeks Fine-tuning to business specifics
Integration and double-run 1-2 weeks Parallel operation, validation
Switchover and training 3-5 days Production launch
Technical Pipeline Details

For batch processing, Celery is used with task distribution on GPU nodes. Embeddings are computed in batches of 512 dialogs. In streaming mode, Kafka partitions distribute the load, and PySpark ensures exactly-once semantics. All metrics are monitored via Prometheus + Grafana.

How to Implement Text Analytics Without Stopping Current Processes?

We use a double-run approach: first, the system runs in parallel with existing processes, we validate results for two weeks, then switch main dashboards to new data. This avoids downtime and provides accuracy confirmation on real data. Pilot project on 10,000 dialogs — 5 days.

What's Included in the Work?

  • Architectural documentation and pipeline description.
  • Configured dashboards (3-5 data marts) for key metrics.
  • Fine-tuned models with model card and accuracy report.
  • API access for ad-hoc queries.
  • Team training (2-3 workshops) and technical support for the first 2 months.

Dialog Enrichment

Each dialog goes through an enrichment pipeline:

@dataclass class EnrichedDialog: dialog_id: str timestamp: datetime topic: str # first-level topic subtopic: str # detail sentiment_score: float # -1 to 1 sentiment_trend: str # "improving" | "stable" | "worsening" resolution: bool # was the problem solved escalated: bool key_entities: dict # products, numbers, amounts summary: str # brief description in 1-2 sentences emotion: str # frustrated | satisfied | neutral | confused agent_id: str duration_seconds: int message_count: int 

Batch processing: 10,000 dialogs overnight via Celery + GPU batching for embedding tasks. In real-time mode, Kafka stream processes a dialog in 2-5 seconds.

Analytics Data Marts

Topic Mart: aggregation by topic, subtopic, date, segment. Answers: what causes contacts, how it changes over time. Allows noticing seasonal peaks two weeks ahead.

Sentiment Mart: average score by segment, product, operator. Where customers are most/least satisfied. For example, the mart showed that NPS for the Premium product dropped by 8 points due to response delays over 10 minutes.

Problem Mart: dialogs with resolution=False + negative sentiment — source of systemic issues. Typical insight: 30% of such dialogs contain the word "document" — meaning document processing needs automation.

Operator Mart: performance metrics per agent for QA. Comparing sentiment before and after training shows an average improvement of 15%.

Anomalies and Alerts

Automatic anomaly detection in dialog stream:

  • Sudden spike in topic contacts → likely incident
  • Sharp deterioration in product sentiment → quality problem
  • Increase in resolution=False share by type → procedure changed or knowledge base info missing

Alerts in Slack/Telegram: automatically when thresholds are exceeded. Incident confirmation time — less than 1 minute.

What Does Implementation Give?

  • Reduce dialog analysis time by 80% (from 40% manual work to automation).
  • Save QA budget up to 40% through automation.
  • Cut incident response time from hours to minutes.
  • Increase classification accuracy from 50-60% (keywords) to 93% (hybrid).

SLA 99.9% on processing time — guarantee of pipeline stability.

Order a pilot analysis: process 10,000 dialogs in 5 days and see first insights. Get a consultation — we'll explain how the system fits into your current stack.