Extract Data from Medical Records: On-Premise Medical NLP

Extract Data from Medical Records: Medical NLP

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1285
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1241
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    983
  • image_logo-aider_0.webp
    AIDER company logo development
    919
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1033

Extract Data from Medical Records: Medical NLP

Physicians spend up to 30% of their working time on paperwork, yet key clinical data remains in narrative text. We turn these texts into structured tables: diagnoses, medications, lab results, procedures. In 6–8 months we build a fully custom NLP pipeline for your MIS. The model runs on-premise, data security complies with 152-FZ. Contact us for a free consultation and project evaluation.

How we extract data from medical records using NLP

We use a combination of methods: Clinical NER based on BERT, assertion detection for context, temporal reasoning for chronology, and term normalization to standard classifiers. Each stage is optimized for Russian-language medical texts.

What problems we solve

Clinical NER: not just word search, but context

Standard regex cannot distinguish "angina" in a list of past diseases from the current diagnosis. We use Clinical BERT, fine-tuned on a labeled corpus of 10,000+ sentences from physicians. Entity-level F1 is 94% for the most common entities (diagnosis, medication, dosage).

Assertion detection: where is the truth?

The note "no chest pain" is the opposite of "angina". The model identifies four contexts: present, absent, uncertain, history. An error here is critical — we achieve 97% accuracy on the test set.

Temporal reasoning: treatment chronology

Physicians care about chronology: when was diabetes detected, when was Metformin prescribed, did the dosage change? Without a timeline, analytics is useless. We build an event timeline using a CRF layer on top of BERT.

Term normalization: from text to code

"SD2", "type 2 diabetes", "Diabetes mellitus type 2" — all the same. Mapping to ICD-10 (E11) and UMLS CUI (C0011860) gives a unified picture across the clinic.

How we do it

We use DeepPavlov/rubert-base-cased as the base model. We fine-tune on a corpus of discharge notes (5,000 documents labeled by physicians). Inference is on an on-premise cluster of 4× A100 — p99 latency 120 ms per medical record.

Why on-premise is mandatory

Medical data is a special category (152-FZ, Article 10). Transfer to a foreign vendor's cloud is impossible without de-identification. We deploy the model on your server or in a Russian cloud (Yandex Cloud, SberCloud). No cross-border data transfer.

How we fight LLM hallucinations

For safety-critical tasks we follow the rule: LLM only for summary generation, not for extraction. NER is done with a separate BERT classifier. Manual audit is mandatory until F1 > 95%.

What works better: BERT or LLM for extraction?

Criterion BERT (token classification) LLM (prompting)
F1 score 94–97% 70–85% (model-dependent)
Latency per document 120 ms 2–5 s
Inference cost Low High (tokens)
Hallucinations None Yes (up to 15% on rare terms)
152-FZ compliance Yes (on-premise) Requires fine-tuning

BERT outperforms LLM by an average of 20% in accuracy and 30x in speed. Therefore, we use BERT for data extraction and LLM only for human-supervised summary generation.

What is included in the work

The implementation process consists of five key stages:

  1. Corpus labeling by physicians (1–2 months) — you get a labeled corpus of 5,000 documents and a baseline NER model with F1 ~90%.
  2. Full pipeline training (3–4 months) — NER + assertion + normalization + de-identification. A ready Docker image with REST API.
  3. Integration with MIS via HL7 FHIR or REST (5–6 months) — pilot on 100 medical records.
  4. Deployment of CDSS module and dashboards (7–8 months) — drug interaction checks, analytics.
  5. Staff training and documentation — handover of model source code.
Stage What you get
1–2 months Labeled corpus of 5,000 documents, baseline NER model F1 ~90%
3–4 months Full pipeline: NER + assertion + normalization + de-identification. Docker image with REST API
5–6 months Integration with MIS (HL7 FHIR or REST setup), pilot on 100 records
7–8 months CDSS module (drug interaction checks), analytics dashboard, documentation, staff training
Technical requirements
  • OS: Ubuntu 20.04+
  • RAM: 64 GB+
  • GPU: 4× NVIDIA A100 80GB
  • CUDA 11.8
  • Docker 20.10+

Typical mistakes and how to avoid them

  • Using ready-made models for English: on Russian texts F1 drops by 30–40%. Fine-tuning is required.
  • Ignoring de-identification: without it, data cannot be used for training and analytics.
  • Relying solely on LLMs: they hallucinate on rare terms — human-in-the-loop is mandatory.

Quality assessment

Metrics: F1 for each entity type, assertion detection accuracy, drug-dosage linkage accuracy. For safety-critical applications — human validation with F1 > 95%.

Our expertise: 5 years of experience in medical NLP, 20+ implementations in private and public clinics. Certified specialists (NVIDIA DLI, Bioinformatics).

Timeline and cost

Timeline: from 6 to 8 months depending on MIS complexity and number of entity types. Cost is calculated individually. Typical savings from automation: 3 to 8 million rubles per year per clinic. Get a free assessment of your project — request a consultation. You keep the model, code, documentation, and trained physicians.