Problem: BLEU misses style and terminology errors
A translator spent two hours post-editing a translation that the system rated 0.95 BLEU. The client rejected the project due to style guide violations and incorrect terminology. BLEU compares against a reference, but the reference often ignores context, tone, and domain-specific vocabulary. We solve this with Quality Estimation (QE) — an AI system that evaluates translations without a reference, like a human reviewer. Budget savings on review can reach 60%.
How Quality Estimation Without a Reference Translation Works
Quality Estimation analyzes the source text and translation, producing a score of 0–1 at the segment, word, and document levels. Segment-level scores indicate which sentences need review. Word-level QE tags each word as OK or BAD — the reviewer sees errors immediately. Document-level QE evaluates coherence, terminology consistency, and stylistic unity.
How QE Saves 40–60% of Reviewer Time
Assume you process 10,000 segments per day. Without QE, a reviewer checks every segment. With QE, only segments with score < 0.7 (typically 20–30%). At a threshold of 0.9 for auto-publishing, 10–15% of segments skip review entirely. We implemented such a pipeline in a fintech application localization project: the reviewer handled 3,000 segments instead of 10,000, and style errors dropped by 80%. This resulted in estimated annual savings of $120,000 in reviewer costs.
Why CometKiwi Is Better Than BLEU for QE
CometKiwi (Unbabel/wmt22-cometkiwi-da) is a transformer-based model trained on thousands of human judgments. It outperforms BLEU and traditional metrics in correlation with human evaluation. Here’s a comparison of key metrics:
| Metric | Requires Reference? | Human Correlation | Word-level Support | Processing Time (1K segments) |
|---|---|---|---|---|
| BLEU | Yes | 0.3–0.4 | No | 1 sec |
| COMET | Yes | 0.6–0.7 | No | 10 sec |
| CometKiwi (QE) | No | 0.6–0.7 | Yes (via MQM) | 15 sec |
CometKiwi needs no reference and provides word-level errors through MQM taxonomy.
Error Types Classified by QE (MQM Taxonomy)
We use the MQM taxonomy, dividing errors into four classes:
- Accuracy — mistranslation, omissions, additions.
- Fluency — grammar, spelling, punctuation.
- Terminology — glossary violations, inconsistent term usage.
- Style — tone of voice mismatches, stylistic inconsistencies.
Common Mistakes When Implementing QE
Typical errors include: using a model without fine-tuning for the language pair (lower precision), choosing the wrong score threshold (missing errors or overburdening the reviewer), ignoring word-level QE (losing context for individual words), and not integrating MQM (difficult to improve the process).
How We Integrate QE Into Your Pipeline
- Audit the current process: measure volume, latency, and existing metrics.
- Model selection: CometKiwi, OpenKiwi, or fine-tuning for your language pair.
- Integration: REST API or gRPC — wrapped in a microservice.
- MQM taxonomy setup: connect error type classification via an LLM (GPT-4o or LLaMA 3).
- Testing: measure precision/recall on your dataset.
- Deployment: Kubernetes + GPU (T4 or A10).
from transformers import AutoModelForSequenceClassification, AutoTokenizer class QualityEstimator: def __init__(self, model_name: str = "Unbabel/wmt22-cometkiwi-da"): self.model = load_comet_model(model_name) def estimate_segment(self, source: str, hypothesis: str) -> QEScore: score = self.model.predict( [{"src": source, "mt": hypothesis}], batch_size=8 ).scores[0] return QEScore( score=score, # 0-1, where 1 = excellent quality requires_review=score < 0.7, error_probability=1 - score ) def estimate_batch( self, segments: list[tuple[str, str]] ) -> list[QEScore]: data = [{"src": src, "mt": mt} for src, mt in segments] scores = self.model.predict(data, batch_size=32).scores return [QEScore(score=s, requires_review=s < 0.7) for s in scores] QE Model Comparison
| Model | Language Pairs | Size | Speed (1K segments) | Word-level |
|---|---|---|---|---|
| CometKiwi | Any | 1.2B | 15 sec | Yes (via MQM) |
| OpenKiwi | Limited | 100M | 5 sec | Yes |
| Fine-tuned | Your pair | Task-dependent | Depends on size | Optional |
What's Included in the Work
- Audit of the current translation pipeline with metric measurement.
- Selection and configuration of a QE model (CometKiwi, OpenKiwi, fine-tuning).
- Integration via REST API or gRPC with documentation.
- Error classification training for your MQM taxonomy.
- Deployment on infrastructure (Kubernetes, GPU).
- Team training and 1 month of support.
Timeline and Pricing
Timelines range from 2 to 6 weeks depending on complexity (volume, number of languages, need for fine-tuning). Pricing is calculated individually. For a typical project with 5 language pairs and 50,000 segments, pricing starts at $15,000. Get a consultation — send a description of your current pipeline, and we’ll evaluate your project within 2 days.
Why Work With Us
- Over 5 years of experience in NLP and machine translation.
- Completed 15+ quality evaluation projects for fintech, e-commerce, and software localization.
- Certified engineers in PyTorch and MLOps.
- We guarantee reviewer time savings of at least 40% or your money back.
Learn more about Quality Estimation on Wikipedia. Typical use cases include e-commerce product descriptions, financial reports, and software UI localization.
Contact us to get a consultation on implementing QE into your translation process. Order a project evaluation in 2 days — send a description of your current pipeline.







