Buyer tool

Annotation Guideline QA Checklist

Live work instructions often contain unresolved rules and contradictory clarifications that only surface when somebody applies them to real material. This checklist produces an approved clarification written back into the instruction, rather than left in a call.

About this resource

Use it: Before an annotation or linguistic-review guideline reaches the people who will apply it.

300+ languages

1,000+ organizations served since 2015 across AI data, translation, localization, media, and interpretation.

1,000+ organizations served · 300+ languages · 4,500+ dialects

When to use this

Before an annotation or linguistic-review guideline reaches the people who will apply it.

This checklist is for the questions that surface only when somebody tries to apply the guideline to real material. It does not replace a complete annotation QA standard.

What it produces

An approved clarification, written back into the guideline itself, with the section that changed and confirmation the working team received it.

The checklist

Annotation Guideline QA Checklist

Every question below is on the page as text. Use the link on any row to send a colleague straight to that one question.

Ask thisA useful answer is specific aboutWhat the answer actually tells you
What must be checked on every item, even when no correction is made? # The fixed checks, what can vary by item, and a valid no-change example alongside an example that needs action. Whether the guideline defines review, or only describes how to edit obvious errors.
When should a term be translated, transliterated, or left unchanged? # The rule for each language or variety, with an example where script choice changes the acceptable output. Whether two careful annotators can reach the same form.
When should a disfluency, self-correction, or recognition error be kept, edited, or flagged? # One example of each action and the reason the action changes. Whether the guideline protects meaning or invites silent rewriting.
What wins when the source form and the surrounding context disagree? # The deciding rule and a neutral example that does not contain real names or source text. Whether names and entities will be handled consistently without exposing source material.
When must an annotator choose unknown, ambiguous, or abstain? # The limit on inference, especially for subjective or sensitive attributes, plus an example of when not to force a label. Whether the task rewards unsupported guessing.
What happens when reviewer feedback conflicts? # Which instruction has priority, who resolves the conflict, and how the earlier advice is marked as replaced. Whether different reviewers can quietly create different rules.
How does an approved clarification reach everyone doing the work? # Where the change is written, who approves it, what section changed, and how the team is told. Whether the answer becomes part of the guideline or remains trapped in a call or message.

How to fill it in

Make each answer specific before you assess it.

Work through each question with the supplier. An answer should name the evidence behind it; a general yes leaves the question open. Write down unresolved questions to discuss together.

What counts as evidence

Something you can inspect, tied to the work you are buying.

A named owner, a document you can open, a sample produced under the instructions you will actually use, or a test whose result changes the decision. A capability statement, an organisation-wide headcount, or a claim about a different assignment does not answer these rows.

The document

Take it with you.

The PDF needs no sign-in, carries every row on this page, and is free to forward, copy into your own process, or quote in a document of your own. If it will not open for you, tell us and we will fix it the same day.

Where this fits

Buyer guides that use this tool.

The guide explains what to look for; this checklist is what you take into the conversation. Both are published in full.

AI Data Annotation Vendor

A buyer-side evaluation framework for annotation, review, security, and pilot-to-production discipline.

Read the guide

Content safety review

A procurement framework for policy ownership, reviewer calibration, IAA diagnostics, escalation control, security, and batch-level safety reporting.

Read the guide

Stuck on a row?

If one row here is hard to answer for your project, send us that row — you’ll get a written answer within one working day, from a person, with no meeting and no follow-up unless you ask.

Send us that row

Scope a project Call