Why two quality figures rarely compare

Picture three responses to the same tender; the figures are invented for illustration. One vendor reports that 98.6% of segments passed review. Another reports a quality score of 97 out of 100. A third reports fewer than two errors per thousand words. A scorecard with a single column headed "quality" will rank them, and the ranking will mean very little, because each figure is almost certainly counting something different.

The first may be the share of segments a reviewer left unchanged. The second may be a penalty-point score in which a style slip costs one point and a meaning error costs ten. The third may have dropped formatting and tag errors before anything was counted. Each is a legitimate way to measure. None can be set against the others until the buyer knows what was counted, what it was divided by, which material was reviewed, how severity was weighted and what record stands behind it.

This guide is about reading the figures vendors report. Designing a QA model of your own, with review depth, error typology, sampling and acceptance, is covered in choosing a translation QA model. The two meet at the scorecard: a buyer with a model can ask every vendor to report against it.

Pin down the unit

Ask each vendor to state, in one sentence, what a single count in the figure represents. An error a reviewer found, a segment with no errors, a penalty point, a reviewer edit and a machine-generated similarity score are five different units, and all of them travel under the word "quality".

The differences are not cosmetic. A segment pass rate treats a misplaced comma and a reversed instruction as the same failure. An edit count rises with reviewer preference, so a strict reviewer makes sound work look worse. An automatic score such as BLEU measures how closely machine output matches a set of reference translations; it is not a count of errors a qualified reviewer found. If the unit cannot be stated plainly, the figure is not ready for the scorecard.

Ask what the figure was divided by

A rate is only as meaningful as its denominator. Words, segments, pages, files or projects? Source or target? The reviewed sample only, or the whole delivery? Were any error classes, such as formatting, tags or preferential changes, removed before the division? Each answer can move the figure a long way without a single sentence of translation changing.

Two denominators deserve particular suspicion. A per-project pass rate counts a two-page leaflet and a year-long programme as one project each, so it says little about the work about to be awarded. A figure averaged across every language a vendor handles hides the weakest pair inside the strongest. Ask for the denominator for each language pair in scope, not for the vendor as a whole.

Find out who chose the sample

The same vendor can produce very different figures from different samples. Was the reviewed material drawn at random from production, aimed at high-risk passages, picked by the vendor, or supplied by the buyer as a test? How many segments were reviewed for each pair, over what period, and in what subject matter?

A test translation prepared for a tender is not production work, and a result from a familiar pair on marketing copy cannot stand in for a different script on regulated content. When a sample is small or does not match the work being awarded, record that limitation beside the figure rather than giving it a precision it has not earned.

Read the severity scale before the score

A single score blends errors of very different consequence. Ask for the severity definitions the figure was scored under: what separates minor, major and critical, who assigned each severity, and what weight each carried. Then ask whether one critical error fails a batch regardless of the overall score. A weighted score can clear its threshold while containing one error that changes the meaning of a dosage line or a contract clause.

Where vendors score on different scales, the scores cannot be reconciled reliably after the fact. The workable fix is to issue one severity scale in the request and ask every bidder to report against it.

Ask what stands behind the number

A figure without a record is a statement. Ask for the error log behind it, with segment references, categories and severities, so a portion of the findings can be checked by someone independent. Ask who translated, who reviewed and who made the final call, and whether the reviewer was separate from the translator. Ask which source version and which date the figure describes.

Certification is evidence of a different kind. An ISO 17100 certificate shows that a translation service process meets the standard's requirements; it is not a measurement of any particular delivery. ISO 17100 vendor qualification covers that step. Keep process evidence and delivery evidence in separate rows.

Put the definitions on the scorecard

For every reported figure, store six fields: the unit, the observed value, the denominator, the sample and how it was chosen, the severity scale, and a link to the evidence. An illustrative row might read "major meaning errors: 2 in 400 reviewed segments, drawn at random from the agreed test set, scored on the buyer's severity scale, log attached". That row shows the format only. It is not a MoniSa result and not a recommended pass mark; the threshold belongs to the buyer, set from the risk of the content before any vendor figure is read.

Where a vendor has not measured something, write "not measured". Turning missing evidence into a zero punishes by guesswork, and turning it into a pass rewards the same way. A "not measured" cell also shows which small test would settle the question. The supplier evidence matrix offers a working layout for keeping each claim beside its evidence.

Where MoniSa fits

The only quality figures that genuinely share a column are figures produced on the same content under the same definitions: the buyer's own material, reviewed against the buyer's unit, sample rule and severity scale. MoniSa runs a pilot batch first. A small first batch is produced and reviewed in full, and the buyer approves it before full production opens. Sending the scorecard definitions with the brief lets that pilot be scoped against them, so its result lands in the same column, on the same terms, as any other bidder measured the same way.

Design the model, qualify the vendor, record the evidence

These live pages cover the neighbouring steps of the same buying decision.

Settle these before comparing quality figures

Agree each point with every bidder before a single figure goes into the scorecard.

  • The unit each figure counts, stated in one sentence
  • The denominator for each language pair, and any error classes excluded before division
  • How the sample was chosen, its size per pair, its period and its subject matter
  • The severity definitions, who assigned severity, the weights and any single-error fail rule
  • The evidence behind the figure: error log, reviewer roles, source version and date
  • One severity scale and one sample rule issued to all bidders in the request

Where to stop and ask again

Each of these calls for a follow-up question, not a conclusion that the work is poor.

  • A quality percentage with no stated unit or denominator
  • One figure covering every language the vendor offers, with no per-pair breakdown
  • A sample the vendor selected, with no record of how
  • A score that blends severities and has no critical-error rule
  • A certificate offered in place of delivery evidence
  • An automatic similarity score presented as a human-reviewed error rate

What to send MoniSa

These inputs let a pilot be scoped against the definitions the scorecard will apply.

  • Content type, intended reader and what an error in that content would cost
  • The language pairs in scope and a representative source sample
  • The unit, denominator, sample rule and severity scale the scorecard uses
  • Acceptance criteria and who holds authority to accept or reject a batch
  • Any glossary, style guide or reference material the review should check against