Verifying what, exactly?

"Is this person who they say they are?" sounds like one question. In a contributor pool it is several, each needing different evidence.

Is this a distinct person, or one of several accounts run by the same person? Does the person match what the dataset will record: variety, region, an age band if the brief specifies one? Is the person doing the work the one who was approved? And, only where a contract, payment or legal rule requires it, who is this person on paper?

Capability is a separate question: can this person do this task in this language? A passport establishes legal identity and says nothing about whether its holder grew up speaking Bhojpuri. A clean sample task shows capability and says nothing about how many accounts its author holds. Merge the two and the result is heavy document collection that still waves through a standard-language speaker claiming the harder-to-fill regional variety.

For collected data, the contributor is the first link in the source history that the usable-versus-shippable pillar traces. Decide which of these claims the dataset depends on before choosing any evidence.

What would a determined bad actor gain?

Scope each check to the exposure in front of it, not to everything technically verifiable.

Per-item pay with a per-person cap rewards duplicate accounts, so the check that matters is uniqueness. A hard-to-fill variety quota rewards claiming the variety, which no identity document settles. A shareable login invites substitution, which only continuing review catches. Machine-translated text or recycled audio is fraud in the output, a job for output QA rather than onboarding.

Confidential work changes the calculation. Where contributors see an unreleased product, a pre-release model's outputs or a client's customer conversations, the gain is access and the harm is disclosure, so stronger identity assurance, contractual commitments and access limits are proportionate.

What evidence is reasonable to ask for?

Choose the lightest evidence that settles each claim.

Uniqueness signals often already sit in contact and payment records: a confirmed contact channel, a payout account no other contributor uses. Shared household accounts are ordinary in some places, so a match is a question, not a rejection. Variety and region need a structured self-declaration (where the person grew up, what is spoken at home) plus a short task reviewed by a speaker of that variety. Legal identity needs a document check only where a payment, contract or client rule requires one, and legal and finance owners decide that, not the data team's comfort level.

Formal credentials need the most care. Many capable contributors in regional and low-resource languages hold no certificate in the language, often because none exists. Requiring one filters for paperwork and tilts the pool toward speakers schooled in the standard variety, often not the one the dataset needs. Where a credential exists and the task requires it, ask for it; elsewhere it is one signal, never the gate. Building training data for low-resource languages covers finding native speakers when formal channels run out.

Biometric checks, such as face matching or voiceprints, are a decision for privacy and legal owners, not an onboarding default. Tell contributors what is checked, why and what is kept, before they begin. A disclosed check against a named risk is verification; undisclosed or open-ended collection is where surveillance starts.

How do you verify capability when nobody in the building reads the language?

This is where most verification quietly gives up. The screening call runs in English, the candidate sounds confident, the profile lists the language, and the account is approved. That is confidence in a different language, not verification of this one.

Four practices make capability checkable without a reader on the team.

A known-answer set. Two trusted speakers of the variety agree a small set of items with settled answers: a recording to transcribe, sentences to translate from a pivot language, labels with an agreed key. Anyone can then score a candidate, because the language knowledge was spent once, on the key. Reviewer calibration applies the same discipline to reviewers.

A reviewer for the variety, not only the language. A speaker of one Arabic variety can confirm that a candidate speaks Arabic; confirming natural Egyptian Arabic rather than a practised imitation needs someone who grew up with it. Name the variety the reviewer must hold.

Independence. If the recruiter is the only reviewer, the check is self-certification. Small language communities may not allow full independence, so use two reviewers, or one from another region who can still judge the variety, and record any relationship.

Checking after approval. The screening task is one data point; the first production batch, reviewed in depth, is a second; routine sampling supplies the rest. Sampling is also where substitution shows: a sudden shift in accent or error pattern from one account is worth a question.

What do you keep afterwards, and for how long?

Keep the result, not the evidence. A record of what was checked, when, by which role and with what outcome answers most later questions. A folder of passport scans answers none better and is a breach target nobody needed. Where a rule requires keeping a document, the rule sets the period, and the copy sits apart from the data under restricted access.

Identity should never travel with the dataset. Contributors appear under pseudonymous keys, and the link from key to person is held separately under narrow access. That link is how a withdrawn contributor's items are found, so minimisation means keeping the join, not keeping nothing. The consent and reviewer QA checklist covers reporting without exposing identities.

Set each item's retention period and deletion owner before the first check runs. Privacy owners set the periods, which differ by jurisdiction and task.

A check that excludes the people you need is not a control

The candidates most likely to fail a document-heavy, credential-first check are often speakers of the scarce varieties the dataset was commissioned to cover. The damage shows later, as a model weak in exactly those varieties, far from the decision that caused it.

So measure the check itself: how many candidates of each variety pass each step. A variety losing most candidates at the document step and few at the capability step is being filtered as a population, not screened for fraud.

The specialist network page sets out MoniSa's own qualification path, where checking duplicate records and reachability is a separate step from confirming credentials, language pairs and experience. For a pool being designed now, start from the task, the varieties and the risk, not a list of documents.

Source history and contributor questions

Each goes deeper on one part of the decision.

Scope a proportional verification step

Work through this before choosing a tool or writing an onboarding form.

  • Name each claim the dataset relies on: distinct person, variety and region, continuity, capability, and legal identity where a rule requires it
  • Write down what a bad actor would gain from each claim, and the harm if it is wrong
  • Match the lightest sufficient evidence to each claim
  • Assign capability review to a speaker of the variety who did not recruit the candidate
  • Tell contributors what is checked, why and what is kept
  • Set retention and a deletion owner for each item, and track drop-out at each step by variety

Where verification has drifted

Each deserves a second look at the process, not only at the contributor.

  • Public and confidential tasks run the same document-and-selfie check
  • Capability "verified" by a profile, a certificate requirement or a call in another language
  • A variety claim accepted on the strength of a test in the broader language
  • The recruiter is the only reviewer of the candidate's sample
  • Identity documents stored beside, or delivered with, the dataset

What to send MoniSa to scope the check

Send the task and the risk, not a document list, and the check can be scoped claim by claim and variety by variety.

  • The task, and whether contributors will see public prompts, confidential material or unreleased content
  • The languages, and the varieties or regions the metadata must record
  • Which attributes must be verified and which may be declared
  • Any identity or record-keeping requirement already set by legal, privacy or a client contract
  • Where candidates drop out today, if the pool already exists