AI Translation Quality Estimation Without Reference Translations

Problem: BLEU misses style and terminology errors A translator spent two hours post-editing a translation that the system rated 0.95 BLEU. The client rejected the project due to style guide violations and incorrect terminology. BLEU compares against a reference, but the reference often ignores co

AI Development Areas

Frequently Asked Questions

Latest works

  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1285
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1240
  • image_logo-advance_0.webp
    B2B Advance company logo design
    696
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    982
  • image_logo-aider_0.webp
    AIDER company logo development
    918
  • image_crm_chasseurs_493_0.webp
    CRM development for Chasseurs
    1032

Problem: BLEU misses style and terminology errors

A translator spent two hours post-editing a translation that the system rated 0.95 BLEU. The client rejected the project due to style guide violations and incorrect terminology. BLEU compares against a reference, but the reference often ignores context, tone, and domain-specific vocabulary. We solve this with Quality Estimation (QE) — an AI system that evaluates translations without a reference, like a human reviewer. Budget savings on review can reach 60%.

How Quality Estimation Without a Reference Translation Works

Quality Estimation analyzes the source text and translation, producing a score of 0–1 at the segment, word, and document levels. Segment-level scores indicate which sentences need review. Word-level QE tags each word as OK or BAD — the reviewer sees errors immediately. Document-level QE evaluates coherence, terminology consistency, and stylistic unity.

How QE Saves 40–60% of Reviewer Time

Assume you process 10,000 segments per day. Without QE, a reviewer checks every segment. With QE, only segments with score < 0.7 (typically 20–30%). At a threshold of 0.9 for auto-publishing, 10–15% of segments skip review entirely. We implemented such a pipeline in a fintech application localization project: the reviewer handled 3,000 segments instead of 10,000, and style errors dropped by 80%. This resulted in estimated annual savings of $120,000 in reviewer costs.

Why CometKiwi Is Better Than BLEU for QE

CometKiwi (Unbabel/wmt22-cometkiwi-da) is a transformer-based model trained on thousands of human judgments. It outperforms BLEU and traditional metrics in correlation with human evaluation. Here’s a comparison of key metrics:

Metric Requires Reference? Human Correlation Word-level Support Processing Time (1K segments)
BLEU Yes 0.3–0.4 No 1 sec
COMET Yes 0.6–0.7 No 10 sec
CometKiwi (QE) No 0.6–0.7 Yes (via MQM) 15 sec

CometKiwi needs no reference and provides word-level errors through MQM taxonomy.

Error Types Classified by QE (MQM Taxonomy)

We use the MQM taxonomy, dividing errors into four classes:

  • Accuracy — mistranslation, omissions, additions.
  • Fluency — grammar, spelling, punctuation.
  • Terminology — glossary violations, inconsistent term usage.
  • Style — tone of voice mismatches, stylistic inconsistencies.

Common Mistakes When Implementing QE

Typical errors include: using a model without fine-tuning for the language pair (lower precision), choosing the wrong score threshold (missing errors or overburdening the reviewer), ignoring word-level QE (losing context for individual words), and not integrating MQM (difficult to improve the process).

How We Integrate QE Into Your Pipeline

  1. Audit the current process: measure volume, latency, and existing metrics.
  2. Model selection: CometKiwi, OpenKiwi, or fine-tuning for your language pair.
  3. Integration: REST API or gRPC — wrapped in a microservice.
  4. MQM taxonomy setup: connect error type classification via an LLM (GPT-4o or LLaMA 3).
  5. Testing: measure precision/recall on your dataset.
  6. Deployment: Kubernetes + GPU (T4 or A10).
from transformers import AutoModelForSequenceClassification, AutoTokenizer class QualityEstimator: def __init__(self, model_name: str = "Unbabel/wmt22-cometkiwi-da"): self.model = load_comet_model(model_name) def estimate_segment(self, source: str, hypothesis: str) -> QEScore: score = self.model.predict( [{"src": source, "mt": hypothesis}], batch_size=8 ).scores[0] return QEScore( score=score, # 0-1, where 1 = excellent quality requires_review=score < 0.7, error_probability=1 - score ) def estimate_batch( self, segments: list[tuple[str, str]] ) -> list[QEScore]: data = [{"src": src, "mt": mt} for src, mt in segments] scores = self.model.predict(data, batch_size=32).scores return [QEScore(score=s, requires_review=s < 0.7) for s in scores] 

QE Model Comparison

Model Language Pairs Size Speed (1K segments) Word-level
CometKiwi Any 1.2B 15 sec Yes (via MQM)
OpenKiwi Limited 100M 5 sec Yes
Fine-tuned Your pair Task-dependent Depends on size Optional

What's Included in the Work

  • Audit of the current translation pipeline with metric measurement.
  • Selection and configuration of a QE model (CometKiwi, OpenKiwi, fine-tuning).
  • Integration via REST API or gRPC with documentation.
  • Error classification training for your MQM taxonomy.
  • Deployment on infrastructure (Kubernetes, GPU).
  • Team training and 1 month of support.

Timeline and Pricing

Timelines range from 2 to 6 weeks depending on complexity (volume, number of languages, need for fine-tuning). Pricing is calculated individually. For a typical project with 5 language pairs and 50,000 segments, pricing starts at $15,000. Get a consultation — send a description of your current pipeline, and we’ll evaluate your project within 2 days.

Why Work With Us

  • Over 5 years of experience in NLP and machine translation.
  • Completed 15+ quality evaluation projects for fintech, e-commerce, and software localization.
  • Certified engineers in PyTorch and MLOps.
  • We guarantee reviewer time savings of at least 40% or your money back.

Learn more about Quality Estimation on Wikipedia. Typical use cases include e-commerce product descriptions, financial reports, and software UI localization.

Contact us to get a consultation on implementing QE into your translation process. Order a project evaluation in 2 days — send a description of your current pipeline.