Extract Data from Medical Records: Medical NLP
Physicians spend up to 30% of their working time on paperwork, yet key clinical data remains in narrative text. We turn these texts into structured tables: diagnoses, medications, lab results, procedures. In 6–8 months we build a fully custom NLP pipeline for your MIS. The model runs on-premise, data security complies with 152-FZ. Contact us for a free consultation and project evaluation.
How we extract data from medical records using NLP
We use a combination of methods: Clinical NER based on BERT, assertion detection for context, temporal reasoning for chronology, and term normalization to standard classifiers. Each stage is optimized for Russian-language medical texts.
What problems we solve
Clinical NER: not just word search, but context
Standard regex cannot distinguish "angina" in a list of past diseases from the current diagnosis. We use Clinical BERT, fine-tuned on a labeled corpus of 10,000+ sentences from physicians. Entity-level F1 is 94% for the most common entities (diagnosis, medication, dosage).
Assertion detection: where is the truth?
The note "no chest pain" is the opposite of "angina". The model identifies four contexts: present, absent, uncertain, history. An error here is critical — we achieve 97% accuracy on the test set.
Temporal reasoning: treatment chronology
Physicians care about chronology: when was diabetes detected, when was Metformin prescribed, did the dosage change? Without a timeline, analytics is useless. We build an event timeline using a CRF layer on top of BERT.
Term normalization: from text to code
"SD2", "type 2 diabetes", "Diabetes mellitus type 2" — all the same. Mapping to ICD-10 (E11) and UMLS CUI (C0011860) gives a unified picture across the clinic.
How we do it
We use DeepPavlov/rubert-base-cased as the base model. We fine-tune on a corpus of discharge notes (5,000 documents labeled by physicians). Inference is on an on-premise cluster of 4× A100 — p99 latency 120 ms per medical record.
Why on-premise is mandatory
Medical data is a special category (152-FZ, Article 10). Transfer to a foreign vendor's cloud is impossible without de-identification. We deploy the model on your server or in a Russian cloud (Yandex Cloud, SberCloud). No cross-border data transfer.
How we fight LLM hallucinations
For safety-critical tasks we follow the rule: LLM only for summary generation, not for extraction. NER is done with a separate BERT classifier. Manual audit is mandatory until F1 > 95%.
What works better: BERT or LLM for extraction?
| Criterion | BERT (token classification) | LLM (prompting) |
|---|---|---|
| F1 score | 94–97% | 70–85% (model-dependent) |
| Latency per document | 120 ms | 2–5 s |
| Inference cost | Low | High (tokens) |
| Hallucinations | None | Yes (up to 15% on rare terms) |
| 152-FZ compliance | Yes (on-premise) | Requires fine-tuning |
BERT outperforms LLM by an average of 20% in accuracy and 30x in speed. Therefore, we use BERT for data extraction and LLM only for human-supervised summary generation.
What is included in the work
The implementation process consists of five key stages:
- Corpus labeling by physicians (1–2 months) — you get a labeled corpus of 5,000 documents and a baseline NER model with F1 ~90%.
- Full pipeline training (3–4 months) — NER + assertion + normalization + de-identification. A ready Docker image with REST API.
- Integration with MIS via HL7 FHIR or REST (5–6 months) — pilot on 100 medical records.
- Deployment of CDSS module and dashboards (7–8 months) — drug interaction checks, analytics.
- Staff training and documentation — handover of model source code.
| Stage | What you get |
|---|---|
| 1–2 months | Labeled corpus of 5,000 documents, baseline NER model F1 ~90% |
| 3–4 months | Full pipeline: NER + assertion + normalization + de-identification. Docker image with REST API |
| 5–6 months | Integration with MIS (HL7 FHIR or REST setup), pilot on 100 records |
| 7–8 months | CDSS module (drug interaction checks), analytics dashboard, documentation, staff training |
Technical requirements
- OS: Ubuntu 20.04+
- RAM: 64 GB+
- GPU: 4× NVIDIA A100 80GB
- CUDA 11.8
- Docker 20.10+
Typical mistakes and how to avoid them
- Using ready-made models for English: on Russian texts F1 drops by 30–40%. Fine-tuning is required.
- Ignoring de-identification: without it, data cannot be used for training and analytics.
- Relying solely on LLMs: they hallucinate on rare terms — human-in-the-loop is mandatory.
Quality assessment
Metrics: F1 for each entity type, assertion detection accuracy, drug-dosage linkage accuracy. For safety-critical applications — human validation with F1 > 95%.
Our expertise: 5 years of experience in medical NLP, 20+ implementations in private and public clinics. Certified specialists (NVIDIA DLI, Bioinformatics).
Timeline and cost
Timeline: from 6 to 8 months depending on MIS complexity and number of entity types. Cost is calculated individually. Typical savings from automation: 3 to 8 million rubles per year per clinic. Get a free assessment of your project — request a consultation. You keep the model, code, documentation, and trained physicians.







