When every language answers a different question
A preference task goes out in several languages under one translated instruction, and the results will not sit together. For example, one language group favours the longer response, a second splits almost evenly, a third agrees with itself unusually tightly. The tempting reading is that the languages differ. The less comfortable one is that the task did.
"Which response is better?" stops being one question once it is translated. Better can mean more complete, more polite, more direct or more correct, and the word chosen in each language can lean towards a different one. The judges answered honestly; they answered different questions. Five questions, asked before collection, separate the two readings and end in one decision: what to hold identical across languages, and what to let each language adapt.
What are the judges actually being asked to prefer?
Start with the construct, not the wording. Before anything is translated, write down what a preferred response has that the other lacks: it answers the question, it is correct, it follows the requested format, it suits the stated user. Separate preference from preconditions; a wrong or unsafe response may need ruling out before anyone is asked which of two they like.
An undefined "better" invites each judge to supply a personal definition, and within one language group those definitions can share habits. That is how an instruction can produce a language effect that is really a definition effect. The fix is a criterion stated as a concept, with worked examples showing it applied, so judges weigh the same property whatever word carries it. Say which criterion wins when two conflict, and settle the tie rule here: whether a judge may call two responses equal can change the data as much as the scale does.
Is the material equivalent, or translated from one source?
Prompts translated from a single source language carry its structure, idiom and assumptions into every judgement, so a judge comparing two responses to a translated prompt is partly judging the translation.
Two routes work, and the right one depends on what the data is for. Translate and adapt, with a native reviewer confirming that each prompt reads like something a user in that language would write. Or author items natively against a shared specification of task type, difficulty, length band and topic, so they are equivalent in design rather than identical in wording. A mixed design keeps a small translated set as a common anchor and authors the rest natively. Record the route per item: a split that appears only on translated items is a finding about the translation.
Who is judging, and are the panels comparable?
Domain specialists in one language set against general contributors in another can disagree for reasons that have nothing to do with language. Less visible differences do the same: how much each group reads AI-generated text, whether judges work in their dominant language, which region or dialect they come from, and how much time per item the task allows.
Write one panel specification with the same fields in every language, including recruitment route, and record them for every judge so a later difference can be tested within comparable subgroups. Match the dialect to the users the data will serve, not to a broad language name, because judges from one region can read register and naturalness differently from users elsewhere who nominally share the language.
Is a split real, or did the instruction land differently?
A split between language groups may be real. Expectations about directness, formality or length can differ, and preference data exists partly to capture that. It may equally mean the translated instruction asked something slightly different. Reporting the split as a cultural difference before checking turns a translation defect into a conclusion.
The checks are cheap if designed in before collection. Back-translate the instruction and scale labels, and have a bilingual reviewer compare intent rather than wording. Pilot in each language with judges giving a reason for each choice; if one group's reasons name a criterion the instruction never intended, the instruction is the cause. Include a few anchor pairs where the task definition makes the answer unambiguous, such as a response that ignores the requested format. These test the instrument, not the judges: if a whole language group splits on them, the instruction or the material is broken in that language.
Then look at where a split sits. A difference concentrated on items that turn on one translated term, or one that vanishes when judges from comparable regions are compared, is not a difference between languages. The annotation guideline QA checklist applies the same test to written rules, language by language, before anyone works from them.
What gets standardised, and what must be allowed to vary?
Standardise whatever defines the meaning of a judgement: the decision the data feeds, the criteria as concepts and their priority, the response format, the scale and tie rule, the reason codes, the anchor set, the panel fields and the sequence of pilot before volume. Changing any of these changes the question.
Let the surface vary wherever identical wording would make the task unequal: the instruction's phrasing, translated for meaning and tested; worked examples, written natively; prompts, realised against the shared specification; and each panel's dialect make-up, matched to the intended users. The test for every item is whether keeping it identical makes the task equally clear and equally hard in every language, or only equally worded.
One decision comes first. If judgements in each language are never pooled or compared, consistency within each language matters more than comparability between them, and more can be adapted locally. Settle that before deciding how much of the list is binding.
Comparable numbers, incomparable meaning
Identical instructions produce numbers that sit neatly side by side on a dashboard. They do not produce judgements that mean the same thing. Treating the first as the second is how a multilingual preference set can pass acceptance and still fail to support the comparison it was collected for.
MoniSa runs preference-style comparisons from a written rubric, with a calibration round on shared examples before scored batches begin. MoniSa's network of 110,000+ verified language specialists · Counted from our linguist database · verified June 2026 spans 300+ languages and 4,500+ dialects, so a panel specification can name a region and a domain rather than stop at a language name. The useful starting point is the task as it stands, reviewed for what to lock and what to adapt before collection.
The next question
These live pages take one part of this decision further.
- Annotation guideline QA checklist: Use to test instruction wording on real material.
- Why disagreement needs a taxonomy: Use when judges disagree within one language.
- Why usable data is not always ready to ship: Use when a delivered dataset may not support its intended use.
Audit the task before the first judgement
Run the task design against this list before collection starts in any language.
- A one-paragraph statement of what a preferred response has, written before translation
- Criteria stated as concepts with a priority order, plus native worked examples in each language
- One response format, scale and tie rule, identical in every language
- A per-item record of whether each prompt was translated, adapted or natively authored
- One panel specification, its fields recorded for every judge in every language
- A back-translated instruction and a pilot with reason codes and anchor pairs, reviewed before volume
Where the data will not compare
Each is a reason to pause and check the design, not proof the data is unusable.
- The instruction was translated once and never tried with judges in that language
- "Better", "natural" or "helpful" is left for judges to define
- Every prompt came from one source language, with no record of which items were translated
- Panels were recruited by availability rather than against one specification
- A cross-language difference is being reported before anyone checked how the instruction landed
What to send for a scope discussion
Bring the task design as it stands; the programme owner still decides what the data must support.
- The decision the preference data will feed, and whether results must compare across languages
- The current instruction, scale and tie rule in the source language, plus any translations already made
- A sample of the prompts and responses, with a note of how each was produced
- The target languages and, where known, the regions or dialects of the intended users
- Any pilot results or unexplained splits that prompted the question