A pilot that rejected 11 of 20.

An AI company needed a Bengali data collection pilot in the Bangladeshi dialect, recruited and delivered inside a fixed window. Twenty participants were recruited. Nine cleared screening.

178 hours - 20 - 9

178 hours Annotation delivered
20 Participants recruited
9 Participants cleared
Bengali pilot, screened visual: Two reviewers listening through a pilot submission queue.
Measured outcomes Bengali pilot, screened
178 hours Annotation delivered
20 Participants recruited
9 Participants cleared
55% Screening rejection rate
Paid production pilot, not a free trial Engagement type

The project

Bengali pilot, screened

Client
confidential AI company, via an LSP partner platform
Service
Data collection and annotation pilot
Language
Bengali (Bangladeshi)
Volume
178 hours
Resources
9 annotators, from 20 recruited

That 20-to-9 funnel is the case. A 55% rejection rate looks like a recruitment failure until you consider the alternative: eleven annotators who did not meet the standard, contributing to a dataset the client would train on.

The dialect requirement narrowed the pool before screening even started. Bengali spoken in Bangladesh is not interchangeable with Bengali spoken in West Bengal for annotation purposes, and a brief that says only "Bengali" gets a dataset that averages across both.

MoniSa handled the work under ISO 9001:2015 for process control and ISO 27001:2022 for information handling. ISO 17100:2015 is scoped to translation, so it is not claimed for this work.

This ran as a paid production pilot rather than a free trial, which changes the incentive structure on both sides: the client was buying real output, and the screening had to be defensible rather than generous.

The problem to solve

Pilots carry a structural temptation. The vendor wants the pilot to convert into a production contract, and the fastest way to make a pilot look successful is to pass everyone through screening and deliver the headcount the brief asked for.

That trade is invisible at pilot stage and expensive later. Annotation quality problems do not surface in the hour count; they surface when the client trains on the data and cannot explain why performance is uneven.

The dialect constraint made sourcing genuinely harder. Participants needed the Bangladeshi Bengali background specifically, plus annotation aptitude — which is a separate skill from speaking the language and does not correlate with it.

Annotation aptitude is the part buyers most often leave out of a brief. A fluent speaker who cannot apply a labelling guideline consistently produces data that looks fine line by line and is unusable in aggregate.

The fixed timeline compounded both. Recruiting for a narrow dialect and screening for a separate aptitude, inside a window that does not move, is where a vendor either holds the standard or quietly drops it.

For buyers, the check that matters is simple and rarely asked: what was the screening rejection rate, and what happened to the people who failed. A vendor that cannot answer either has not really screened.

What MoniSa changed

Twenty Bengali-speaking participants were recruited against the Bangladeshi dialect requirement, deliberately over-recruiting against the target headcount because a screening step only means something if it can afford to reject.

  • Over-recruit to allow rejection

    Twenty were recruited for a smaller target, because a screening step that cannot afford to fail anyone is not a screening step.

  • Aptitude separate from fluency

    Annotation consistency was tested as its own requirement — a fluent speaker is not automatically a reliable annotator.

  • Remove before production, not after

    The eleven who did not meet standard were removed before contributing data, rather than corrected downstream once the dataset existed.

  • Report the funnel, not the survivors

    The 20-to-9 ratio was reported as the result. A pilot reporting only its final headcount conceals whether screening happened at all.

Results

Measured outcomes from this engagement.

178 hours of annotation were completed and delivered as paid production work in the Bangladeshi Bengali dialect.

Annotation delivered178 hours
Participants recruited20
Participants cleared9
Screening rejection rate55%
Engagement typePaid production pilot, not a free trial

What supported the result

Why the fit was real

The work needed dialect-specific sourcing plus a screening standard that could reject more than half the pool without missing the delivery window.

What decided the result

Holding the annotation standard under pilot-conversion pressure is what protected the dataset — and it is the behaviour that converts pilots into production contracts.

What buyers can reuse

  • Ask every pilot vendor for the screening funnel: recruited, cleared, and what disqualified the rest. A project-scoped quality review pass rate is a finding, not a reassurance.
  • Specify the dialect, not just the language. Bengali (Bangladeshi) and Bengali (West Bengal) are different sourcing pools for annotation purposes.
  • Annotation aptitude is a separate requirement from language fluency and has to be screened separately. Fluency does not predict labelling consistency.
  • Over-recruitment is what makes screening real. A vendor recruiting exactly the target headcount has no capacity to reject anyone.
  • Prefer a paid production pilot to a free trial. It produces usable output and removes the incentive to build a demonstration rather than a sample of real work.
  • A useful pilot brief names the dialect, the annotation guideline, the aptitude threshold, the funnel reporting requirement, and who owns re-screening cost.
  • Quality control that happens after data enters the set is correction. Quality control that happens before is screening. Only the second protects the dataset.

Continue from this proof

Useful comparisons for the same problem.

Use these links to compare the case with the matching service, buyer guide, and language coverage.

Languages named

Examples referenced in the engagement.

  • Bengali (Bangladeshi)
  • Annotation screening
  • Dialect-specific sourcing
  • Paid production pilot

case evidence

Related projects.

These related cases keep the next click close to the same kind of work.

AI data services~300 hours of voice bot conversation transcribed with disfluencies preserved for training value.

Voice-bot transcription

The challenge. A partner needed human-side conversational transcription across six languages, where cleaning the transcript would destroy the training signal.

What we did. MoniSa held a verbatim convention across six per-language standards and kept the two Spanish variants operationally separate.

The result. ~50 hours per language delivered, with false starts and corrections intact.

Open full case

Recognise your own project in one of these?

Send the language list and volume
AI data services213.89 hours of Japanese short-form audio accepted under project-scoped review on partner review.

Japanese short-form audio

Problem. An AI data partner needed Japanese transcription of high-item-count short-form audio without convention drift between transcribers.

Action. MoniSa deployed a small stable four-person team working inside the partner's own production and review workflow.

Result. 213.89 recorded hours delivered, accepted under project-scoped review by the partner's review.

Open full case
Translation services607,000 words across 17 rare languages and 5 scripts, delivered in two phases with accuracy reported per phase.

Rare-language TEP, two phases

Problem. An LSP partner needed a 10-day rare-language surge followed by a four-month programme covering materially harder languages.

Action. MoniSa activated a pre-built bench, ran staggered parallel production, and applied QA per script system including dual-script Kashmiri.

Result. Phase 1 under project-scoped review in 10 days; Phase 2 under project-scoped review across 12 languages over four months.

Open full case

Buyer questions

Common questions.

What was delivered on this engagement?

Annotation delivered: 178 hours. Participants recruited: 20. Participants cleared: 9

What control kept the work stable?

Holding the annotation standard under pilot-conversion pressure is what protected the dataset — and it is the behaviour that converts pilots into production contracts.

Where should similar work go next?

Use AI data services for the delivery model, AI data annotation buyer guide for buyer-side evaluation, and the contact page for a scoped brief.

What happens if you cannot staff one of my language pairs?

For the proposed project, ask for pair-by-pair availability or a recruitment window in writing before agreeing a date. Define qualification and pilot approval for any new contributor before live work. A coverage claim should be checkable before the scope is signed.

Similar brief

Send the constraint behind the metric.

A useful follow-up to a case study names the language mix, review model, deadline, and what proof your buyer team needs before approval.

Send a brief

Do not paste raw outputs, source records, transcripts, third-party personal data or confidential files here. We will agree a transfer path after scoping.

Required. We will reply about your project.

Scope a project Call