Artificial intelligence in orthodontics: where planning breaks down
Classic planning starts with manual markup of images: on a cephalogram, the doctor places dozens of landmarks, calculates angles, proportions, and tooth position relative to bone structures.
At scale, this takes dozens of minutes per patient, and the bottleneck becomes the specialist's time, not a lack of data.
Then discrepancy appears. Landmarks are placed manually, and the result depends on experience, fatigue, and habitual methodology. The same image, re-annotated, produces displacement of some points, and along with them the angles and conclusions derived from them “drift.”
Hence the discrepancy in the treatment plan: from one image, two specialists get different values and different tactics — from the duration of wearing the appliance to the decision on tooth extraction.
Measurement reproducibility drops, and discussion of the plan turns into an argument about markup rather than the clinical situation.
On follow-up images this is most evident: when measurements diverge, it is unclear what to revisit — the tactics themselves or the landmarks. The patient sees a change in decision instead of predictable progress, and the doctor spends time on repeated review.
This is exactly where artificial intelligence in orthodontics works: automatic cephalometry fixes a unified set of landmarks and rules, and measurement reproducibility becomes a measurable parameter rather than the luck of a particular appointment. We break down how this is addressed in practice.
Measurable effect: planning time, accuracy, rejections
Manual preparation of an orthodontic plan is 2–3 hours of work per case: placing points on the cephalogram, calculations using the methods, and assembling a forecast of tooth movement.
Automatic placement of anatomical landmarks reduces treatment plan preparation time to 20–30 minutes, and the doctor edits not the markup from scratch but deviations from the proposed variant — there are significantly fewer edits during treatment.
The second effect is the share of images that go for retake. Errors in patient positioning, blurred contrast, and lost landmarks used to be detected already at the planning stage, when the visit had to be scheduled again.
Input image quality checking cuts off such frames immediately: the share of repeat images drops from 15–20% to a few percent, and the “image — plan” cycle stops stretching into an additional appointment.
We measure the accuracy of tooth movement forecasting not by eye: we calculate the deviation of actual position from predicted at control points and the share of cases where the forecast was confirmed within the accepted tolerance.
Discrepancies are grouped into understandable classes of planning errors — landmark displacement on unclear anatomy, dental arch segmentation error, incorrect occlusion interpretation, unaccounted growth.
The PAR Index scale is calculated automatically on the same data, so the “before and after” comparison does not require manual recalculation.
We take metrics on a holdout set and version the models: a drop in accuracy on a new class of images is visible before the error reaches the clinic. We analyze which of these numbers are achievable in your case.
Comparison of approaches: manual markup, templates, custom model
At the start, a clinic usually has images and doctors' experience, not an annotated dataset. Therefore, the choice of approach is not a question of technology but of how much data and quality control you are willing to invest at the start. Below are three modes that produce different results on the same flow of studies.
Manual cephalometry remains the gold standard of accuracy, but the spread between doctors is noticeable: the same points are placed with deviation, and time per case grows.
Ready-made markup templates remove routine, but the limitations of ready-made solutions are that they describe average anatomy rather than your planning school.
| Option | When it fits | Limitations |
|---|---|---|
| Manual markup by a doctor | Complex and nonstandard cases, single patients | Variation between specialists, time for each image, knowledge does not accumulate |
| Ready-made markup templates | Quick start, typical cases, even flow | Do not account for the anatomy of a specific clinic; every point has to be checked and edited |
| Training a model on clinic data | Your own volume of images and your own planning logic | Requires a labeled set and a regulation for markup quality control |
| Hybrid planning mode | When you need a quick start and gradual accuracy growth | Requires agreement on what exactly the doctor confirms and what they accept as is |
Training a model on clinic data removes the limitations of templates but requires a labeled set and a verification regulation.
In practice, the working point becomes hybrid planning mode: the model proposes markup and forecast, the doctor confirms or edits, and the edits replenish the set — accuracy grows with the flow, without stopping appointments.
Implementation stages: from image audit to model handover
Implementation begins not with model training but with analyzing the already accumulated archive: what images exist, how they are stored, how consistently doctors place cephalometric points. Until this is done, any metric remains a number without connection to real appointments.
- Audit of the image archive and current process. We review the completeness of studies, correctness of calibration, and variation in markup between specialists. The output is how many cases are actually suitable for training and what will need to be captured additionally.
- Data processing scheme. We describe the path of an image: de-identification, normalization, markup storage, access rights. Separately, we fix the points of contact between the contour and the clinic's information system.
- Corpus collection and markup. Points are annotated by several doctors independently; discrepancies are reviewed jointly — the reference reflects agreed practice, not the opinion of one specialist. For 3D, we separately annotate occlusion and root position.
- Training and internal validation. We calculate error for each point separately, not as a single average: averaging hides failure on incisors or chin. The forecast learns to rank plan options by result stability.
- Testing on a holdout set. Studies that did not participate in training are run blindly and compared with doctors' conclusions, separately for complex groups: metal artifacts, pronounced crowding. We record the share of cases where the discrepancy falls within clinically acceptable.
- Pilot operation and handover to doctors. The solution works as a second look; the final conclusion remains with the doctor. We hand over instructions, a regulation for verifying disputed cases, and a procedure for reverting to the previous process.
Project artifacts: configurations, schemes, regulations
A project is considered delivered not when the model shows good metrics on a holdout set, but when the clinic works with it without our involvement.
Therefore, we hand over not only the model itself but also the entire surrounding framework: the path of an image to the treatment plan, access rights, update procedure, and rules for the doctor.
- Image processing scheme — from loading a study from a tomograph or scanner to finished markup: where data is de-identified, where intermediate results are stored, and in what form they go to the archive.
- Calculation configurations: run parameters, confidence thresholds for cephalometric points — at low confidence, markup goes to the doctor for confirmation rather than into the treatment plan.
- Policies for access to patient data: roles, separation by office and branch, export and printing rules, and a log of access to each study.
- Instructions for the doctor: how to read the model's output, which landmarks to check first, how to dispute a calculation — discrepancies become material for further training.
- Model update regulation: on what data we retrain, by what metrics we accept a new version, how we compare it with the current one, and how we roll back if quality has dropped.
- Support procedure: response window for a failure, who on the clinic side makes decisions, how often we check accuracy on fresh images.
Handover is done by an acceptance act and with acceptance testing on real studies: we run a flow of images together with the doctor, verify markup manually, and record under what conditions the result requires re-checking.
Case study: planning without repeat retakes and edits
The practice started with manual cephalometry: the doctor marked points on the cephalogram, separately superimposed CT and scans of the dental arches, and transferred measurements into the plan manually.
Preparation took about 40 minutes per patient, and discrepancies in markup between doctors reached 2–3 mm — the forecast “drifted,” and the plan was edited already during treatment, when the appliance was already on the patient.
We assembled a de-identified corpus of the practice's images with markup by its own orthodontists and configured a model: it finds cephalometric points and contours, calculates angular and linear parameters, compares them with a 3D model of the dental arch, and proposes 2–3 movement scenarios with timeline forecasts.
The draft opens in Dolphin Imaging: the doctor checks the points, changes the scenario, and approves the plan — the signature and responsibility remain with the doctor.
The effect was measured by three metrics on a patient flow.
| Metric | Before | After |
|---|---|---|
| Plan preparation per patient | ~40 min | ~11 min |
| Retakes during treatment | 21% of cases | 6% |
| Edits to the approved plan per course | 3–4 | 1 |
Retakes did not disappear by themselves: the model shows in advance where movement goes beyond the boundaries of bone support, and the doctor changes tactics before the start, not in the middle.
All versions of the plan are saved — when switching to another treatment scheme, we return to the previous markup without repeat measurements. When the methodology changes, the model is recalibrated on new cases from the practice.
What to do if orthodontics already has manual markup?
Manual markup is not a hindrance but a support: it becomes the reference against which the model is checked.
We do not rewrite the established order but embed automation on top of the current process — the doctor conducts appointments in a familiar environment, for example in Dolphin Imaging, while the calculation of cephalometric points, the forecast, and 3D model construction proceed in parallel and are placed into the patient's record as a separate layer.
Migration of the image archive is done in batches. Cephalograms and tomograms are parsed by metadata, lateral images are brought to a uniform scale and orientation, and old images undergo preprocessing.
Then the model calculates points on accumulated cases, and the result is compared line by line with manual protocols — the output is a map of discrepancies for each point and each image type.
The first weeks are parallel work of the doctor and the model in shadow mode: markup from the model is recorded, but the doctor's version goes into the treatment plan.
Discrepancies accumulate in a report, and thresholds are tuned based on them: where the calculation is stable and where it drifts on a specific morphology — pronounced asymmetry, crowding, condition after orthognathic surgery.
A gradual transition is built on the same data. Points with small discrepancy go into auto-acceptance with spot checks; disputed ones always remain for doctor verification.
Every edit is returned to the training loop, so the share of cases accepted without correction grows from flow to flow, and patient appointments do not stop for a single day.
Clinic concerns: data, accuracy, connection to the tomograph
Clinic concerns are almost always not about the quality of the report but about embedding the model into the current flow: what to do with an already captured archive, who is responsible for discrepancies in measurements, and whether equipment will have to be changed. These are the four questions asked before implementation starts.
Will our images work — or do we need to change equipment?
There is no need to change the clinic's tomograph or software — we work with current data. The series undergoes input control: coverage of the area, motion artifacts, correct calibration.
If quality is below the threshold, the system does not fill in what is missing but names the reason and requests a repeat scan.
Who is responsible for a model error?
The doctor: the model prepares calculations and plan options; the decision and signature remain with the specialist. Every measurement can be traced to the points on which it is based — they are visible on the image, can be corrected manually, and recalculated. The model version and input data are saved, so a discrepancy can be analyzed step by step.
Will you train the system for our methods?
Yes: your own cephalometric points, norms, and plan templates. We further train on the clinic's labeled archive, and with a small volume we start with a base model and refine it as doctors' decisions accumulate. On rare combinations of anomalies, accuracy is lower — we flag such cases directly in the report.
How will the result reach the doctors?
Into familiar tools: markup in the viewer, a report in the patient's record, and export to the clinical program. There is no need to learn a separate interface — the outputs are read like a regular conclusion, and disputed points are highlighted.
Let's discuss the task and come back with an implementation plan
For an assessment, send a de-identified set of images — lateral cephalograms or CBCT data — a list of points and measurements mandatory for your doctors, and a sample conclusion you want to receive as output.
Add a description of the clinical process: where the solution works, in the office or in the cloud, who confirms the calculation, and what happens with a questionable result. This determines both the scope of work and the timeline.
Together we will analyze which cephalometry points are critical and which are considered derived; how we measure suitability — by discrepancy with the markup of an expert doctor; which error class is unacceptable and what necessarily goes to manual review by a specialist.
In response, we will come back not with general considerations but with a plan:
- scope of work by stages: data preparation and markup, training, holdout validation, embedding into the clinical process;
- solution scheme and points of integration with what is already used in the clinic;
- acceptance metrics for each stage and how they are measured, rather than a “by eye” assessment;
- data requirements on your side: image format, volume, who annotates and by what protocol;
- role distribution: expert orthodontist, our engineers, your specialist for access and patient data protection;
- stage timelines and rollback procedure if quality on the control set is not confirmed.
Describe the task — we will come back with an assessment, an implementation plan, and timelines.







