How Translation Quality Is Measured: Error Rates, MQM, and Benchmarks

What the numbers on a translation quality report actually mean, how they are calculated, and how to set a threshold that survives an argument.
Diagram of how translation quality metrics turn logged errors into a scored quality benchmark

A vendor sends back a localization quality report. The score is 98.7. Nobody in the room can say whether that is good, bad, or meaningless, and the meeting moves on.

The problem is not the number. It arrived without the three things that would make it interpretable: how much content was reviewed, what separates a major error from a minor one, and what score would have failed.
Without those, a quality score cannot be argued with, compared to last quarter, or acted on. It gets filed.

This is the mechanics piece. How an error-based score is actually built, why the industry standard scoring model produces numbers that all look like 99, how much content you need to review before the score means anything, and what threshold to set for which content. The conceptual background on the metric families themselves sits in the complete guide to AI translation quality. This post assumes you have already picked a framework and now have to run it.

What a translation quality metric actually measures

A translation quality metric converts errors found in a reviewed sample into a single number, using severity weights and a fixed reference length as the denominator. It measures defect density. It does not measure correctness.

That distinction gets lost immediately in stakeholder conversations. A score of 98.7 does not mean 98.7 percent of the translation is right. It means the weighted penalties found in the sample, scaled to a thousand words, came to 13 points. Two reviewers working on the same file with different severity definitions will produce different scores, and both will be defensible. The number is a function of the scorecard, not a property of the text.

This is also why a high score does not rule out the worst failure mode in AI-assisted translation, which is fluent output that is quietly wrong. Fluency errors are cheap to spot and cheap to weight. A mistranslated dosage or a reversed liability clause reads perfectly and carries a critical penalty. The score only catches it if someone reviewed that segment.

The fix is procedural. Before comparing any two quality scores, confirm they were produced under the same error typology, severity scale and sample. If they were not, you are comparing two measurements that happen to share a scale.

How an error-based quality score is calculated

Every error is logged with a category and a severity, weighted, summed, and normalised to a fixed reference length of a thousand words. That normalised penalty total is then subtracted from 100 to produce the score.

The Multidimensional Quality Metrics (MQM) framework, maintained by the MQM Council and harmonised with ISO 5060:2024, is the model most providers now use. Its defining choice is that severity weights rise exponentially rather than linearly. A minor error costs one point. A major error costs five. A critical error costs twenty-five. One critical error outweighs twenty-five minor ones, which is the correct behaviour for commercial content and the opposite of what a simple error count would tell you.

The five steps behind any error-based quality score. Ask for step three and step four, not just step five.

The five steps behind any error-based quality score. Ask for step three and step four, not just step five.

Severity Weight What it looks like in a SaaS product Penalty
Neutral 0 A valid alternative wording the reviewer would not have chosen 0 each
Minor 1 Inconsistent capitalisation in a settings label 1 each
Major 5 A destructive action button labelled with a softer verb 5 each
Critical 25 A billing screen stating the wrong currency or tax treatment 25 each

MQM severity weights. The exponential gap is deliberate: it reflects risk, not effort.

The arithmetic is worth running once by hand. A reviewer checks 2,000 words of onboarding copy and logs six minor errors and two major ones. The Absolute Penalty Total is six plus ten, so sixteen. Normalised to a thousand words that is eight penalty points. The raw score is 100 minus 0.8, so 99.2. Under MQM’s raw passing threshold of 99.0, that file passes.

Any vendor running linguistic quality assurance (LQA) should be able to hand over the penalty total, the reviewed word count and the error log behind the score. If they can only give you the score, it is a claim rather than a measurement.

Why raw MQM scores all cluster around 99

The raw MQM passing threshold is 99.0, which leaves exactly one point of scale between passing and perfect. Every meaningful difference between two vendors gets compressed into that single point.

MQM’s own documented examples make the problem visible. A legal translation buyer sets a threshold of 99.5, which allows five penalty points per thousand words. A technical help content buyer accepts 97.2, which allows twenty-eight. Those are wildly different quality bars, and to anyone reading a dashboard they both look like near-perfect numbers. Ten minor errors and two major errors per thousand words both land on 99.0, despite being very different experiences for the user.

Raw scores compress every commercially relevant difference into the top one percent of the scale. Calibration is what makes them readable

Raw scores compress every commercially relevant difference into the top one percent of the scale. Calibration is what makes them readable.

MQM defines a calibrated score for exactly this reason. A scaling factor redistributes the narrow passing interval across a range people can actually read, using the calibrated score formula that takes the normalised penalty total, subtracts the acceptable penalty points, multiplies by the scaling factor, and adds the passing threshold. The output is a number where a 92 and a 97 are visibly different suppliers.

The practical rule: raw scores are for linguists doing rework. Calibrated scores are for anything that reaches a vendor scorecard, a quarterly review, or a procurement decision. Reporting raw scores upward guarantees that nobody upstream can tell good from adequate.

Error rate and quality score are not the same number

Error rate counts how many errors appear per unit of text. A quality score weights those errors by severity. A file can have a low error rate and still fail, and a file with a high error rate can comfortably pass.

Two files make the point. File A contains twenty minor errors in a thousand words: twenty penalty points, a raw score of 98.0. File B contains a single critical error in a thousand words: twenty-five penalty points, a raw score of 97.5. File B has one twentieth of the error count and the worse score. It is also the only one of the two that could trigger a refund, a regulatory finding, or a public correction.

Report both numbers, because they answer different questions. Error rate tells you how much rework the batch needs and therefore what it will cost to fix. Quality score tells you how much risk is sitting in the batch right now. A team that tracks only one of them will either over-invest in polishing low-risk content or ship a critical error inside a statistically clean file.

How much content you actually need to review

Above a few thousand words nobody reviews everything, which means every quality score you have ever seen is a sample estimate. If that sample is not stratified by content type and risk, the score describes the easiest content in the batch.

This is the step most published guidance skips. ISO 5060:2024 sets requirements across the pre-evaluation, evaluation and post-evaluation phases and includes recommendations on sampling methodology, but a standard cannot choose the sample. Teams improvise, and the improvisation has a predictable shape: the reviewer opens the largest file and reads the first thousand words.

In a typical SaaS release, the largest file is marketing copy or the longest help article. The billing strings, the destructive-action confirmations, the error messages and the legal footer are all short files. They are where critical errors live, and an unstratified sample never touches them. The score comes back at 99.4 and the checkout flow still says the wrong thing in German.

A workable default for sample depth. These are practical starting points rather than a published standard, and they assume stratification

A workable default for sample depth. These are practical starting points rather than a published standard, and they assume stratification.

Fix the composition of the sample before arguing about the number it produces. A stratified thousand-word sample that includes every content type in the release is worth more than an unstratified five-thousand-word one, costs less, and is the only version that can honestly be compared to last month’s score.

What score actually counts as good

There is no universal passing score. The threshold is set by the buyer, based on what a failure in that specific content would cost. Any provider quoting an industry-standard quality number without naming a content type is quoting nothing.

Content type Risk Raw pass score Penalty points per 1,000 words
Internal docs, release notes, knowledge base Low 97.0 30
Marketing pages, blog, in-app education Medium 98.5 15
Product UI, onboarding, billing flows High 99.0 10
Legal, medical, financial disclosures Critical 99.5 5

Workable starting thresholds by content risk. Set them before the first review, not after the first disappointing score.

Setting a single threshold for an entire release is the more common mistake, and it fails in both directions at once. A 99.5 bar applied to internal release notes burns reviewer hours on content nobody will sue over. The same 99.5 applied to a marketing landing page is usually fine. Applied to a payment confirmation screen it is arguably too low.

The NEX Translation Matrix™ exists to make that routing decision explicit. Five inputs get scored per piece of content: content type, business risk, customer impact, regulatory requirement, and quality expectation. Those inputs route the content to one of three outcomes: AI drafting is sufficient, human linguists refine the meaning, or an independent reviewer signs off before publication. The rule that matters is that the matrix routes content, not projects. A single product release usually produces all three outcomes at once, so routing runs per string or per document, never per job.

The NEX Translation Matrix. Five inputs, three review depths, applied per string rather than per project.

The NEX Translation Matrix. Five inputs, three review depths, applied per string rather than per project.

The same logic sits underneath NexTranslate’s transparent three-tier pricing, where the review layers included at each tier match the risk of the content the tier is built for, and human proofreading is included at every tier rather than sold as an add-on. Content routed to the top outcome gets independent LQA sign-off from a second linguist who did not produce the translation, which is the only arrangement where the quality score is genuinely independent of the person being scored.

Where automatic metrics belong in the process

Automatic metrics are a screen, not a verdict. They tell you which segments deserve a human reviewer’s attention. They do not tell you whether content is fit to publish.

The three families behave differently and fail differently. Surface metrics such as BLEU and chrF compare string overlap against a reference translation, which makes them fast, cheap, and prone to punishing perfectly good rewordings. Neural metrics such as COMET and MetricX encode source, output and reference into a shared semantic space and correlate far better with human judgment, but they return a number with no error category attached, so there is nothing for a linguist to act on. LLM-as-judge approaches such as GEMBA-MQM prompt a large model to identify error spans and assign MQM severities, which is closer to actionable, with the caveat that a model is grading work produced by models trained on similar data.

The productive arrangement is triage. Run a neural metric across the entire batch, rank segments by predicted quality, and route the bottom of that distribution into human review. That is how a thousand-word human sample can meaningfully cover a two-hundred-thousand-word release: the sample stops being random and starts being targeted at the segments most likely to contain errors. In a machine translation post-editing (MTPE) workflow this also tells the post-editor where to slow down, which is the difference between post-editing and re-reading.

Where the models themselves are the product being assessed rather than the tool, that is a separate discipline. Benchmarking multilingual model output across languages is what AI evaluation services are built for, and it uses the same MQM error typology applied to a very different question.

Frequently asked questions

What is a good translation quality score?

There is no single good score, because the threshold depends on the content. As a starting point, a raw MQM score of 99.5 suits legal and medical content, 99.0 suits product UI and billing flows, 98.5 suits marketing content, and 97.0 is reasonable for internal documentation. Set the threshold before the first review, and state it in the same document as the score.

What is the difference between MQM and BLEU?

MQM is a human error typology that produces a weighted, diagnosable score. BLEU is an automatic string-overlap metric that produces a number with no error categories behind it. MQM tells a linguist what to fix. BLEU tells an engineer whether a model changed. They answer different questions and are not interchangeable, which is why a rising BLEU score is not evidence of publishable quality.

How many words should be reviewed in a quality check?

For batches under 2,000 words, review all of it, because sampling saves nothing at that size. Between 2,000 and 20,000 words, a stratified sample of 1,000 to 2,000 words is a reasonable default. Above that, sample across every content type in the release rather than increasing the raw word count, since composition affects the score far more than sample size does.

Can AI score translation quality without a human?

It can produce a usable estimate for triage, but not a defensible sign-off. LLM-as-judge systems now identify error spans and assign severities with reasonable agreement against human annotators on general content, and they degrade on domain-specific terminology, regulatory language and brand voice, which is exactly where a critical error would occur. Use automatic scoring to decide what a human reviews, not to replace the review.

Does a higher quality score always mean a better translation?

No, and assuming it does is how teams get caught out. A score is a function of the sample, the error typology and the severity definitions used. A vendor reviewing easy content leniently will out-score a vendor reviewing billing strings against a strict typology every time. Compare the scorecards before comparing the scores.

Conclusion: a quality score is only as honest as the threshold behind it

Most translation quality reporting fails before the review starts. The threshold was never set, so no score can fail. The sample was never stratified, so the score describes the safest content in the batch. The number gets reported raw, so everything looks like 99. That is not a measurement problem. It is a decision problem wearing a measurement costume.

The corrective is small and mostly free. Decide the passing threshold per content type before the first file is reviewed. Stratify the sample so every content type in the release is represented. Report the penalty total and reviewed word count alongside the score. Calibrate anything that goes to a stakeholder. Those four habits turn a decorative number into one that can settle an argument.

If your current quality reporting cannot answer what was sampled and what would have failed, that is the place to start. We run scored LQA against your content types and thresholds rather than a generic scorecard, and we will show the error log behind every number. Get a quote with a sample of your content, or read the wider guide to benchmarking AI translation output for the framework this sits inside.

Written by Karuppusamy Arunachalam, NexTranslate
Published August 2026 . Filed under AI & LLM Evaluation

Picture of Karuppusamy Arunachalam

Karuppusamy Arunachalam

Karuppusamy Arunachalam is the founder of NexTranslate Private Limited, a language solutions company helping businesses communicate globally through AI-powered and human-refined translation services. With experience in SaaS solution consulting and enterprise communication systems, he is passionate about building technology-enabled solutions that bridge languages and cultures.

Table of Contents

Let’s Go Global Together

Ready to reach new markets and speak to your customers in their own language? Let’s make it happen –  faster, smarter, and more affordably.