Industry Insights · January 25, 2022

How Should MT Quality Be Evaluated? From BLEU Scores to Human Evaluation

Illustrated machine translation quality evaluation

Every enterprise adopting machine translation asks: which engine is best? In the past, the answer was often a single number — the BLEU score. But the industry has long recognized the gap between automatic metrics and real usability: a high-BLEU translation may not read well, and two engines with similar BLEU scores can perform worlds apart on actual business content.

In 2022, evaluation systems centered on human assessment are becoming the standard practice. The mainstream approach has bilingual evaluators score MT output on dimensions like accuracy, fluency and terminology consistency, or directly rank two engines' outputs (pairwise comparison). Crucially, evaluation must run on the enterprise's own real corpus — an engine's performance on legal contracts, product manuals and marketing copy can differ dramatically.

More sophisticated practice also measures "edit distance": how much a linguist must actually change the MT output to reach delivery standard. This metric maps directly to post-editing cost and forms the technical basis for MTPE pricing.

The advice for enterprise buyers is simple: don't choose an engine on brochure numbers — run a small-scale evaluation on your own documents from the past year and let the data speak. For translation companies, MT evaluation capability itself is becoming a professional service that can be offered to clients.

Let's talk about your language needs

Tell us about your project — we'll reply with a quote and delivery plan within one business day.