Case study
A pilot that rejected 11 of 20.
An AI company needed a Bengali data collection pilot in the Bangladeshi dialect, recruited and delivered inside a fixed window. Twenty participants were recruited. Nine cleared screening.
178 hours - 20 - 9
Project overview
What landed, and what made it hard.
An AI company needed a Bengali data collection pilot in the Bangladeshi dialect, recruited and delivered inside a fixed window. Twenty participants were recruited. Nine cleared screening.
Delivery snapshot
Bengali pilot, screened
- Client
- confidential AI company, via an LSP partner platform
- Service
- Data collection and annotation pilot
- Language
- Bengali (Bangladeshi)
- Volume
- 178 hours
- Resources
- 9 annotators, from 20 recruited
Why this mattered
Outcome before process.
That 20-to-9 funnel is the case. A 55% rejection rate looks like a recruitment failure until you consider the alternative: eleven annotators who did not meet the standard, contributing to a dataset the client would train on.
The dialect requirement narrowed the pool before screening even started. Bengali spoken in Bangladesh is not interchangeable with Bengali spoken in West Bengal for annotation purposes, and a brief that says only "Bengali" gets a dataset that averages across both.
MoniSa handled the work under ISO 9001:2015 for process control and ISO 27001:2022 for information handling. ISO 17100:2015 is scoped to translation, so it is not claimed for this work.
This ran as a paid production pilot rather than a free trial, which changes the incentive structure on both sides: the client was buying real output, and the screening had to be defensible rather than generous.
The problem to solve
Why the work was difficult, and what MoniSa changed in-flight.
Pilots carry a structural temptation. The vendor wants the pilot to convert into a production contract, and the fastest way to make a pilot look successful is to pass everyone through screening and deliver the headcount the brief asked for.
The challenge
The problem to solve
That trade is invisible at pilot stage and expensive later. Annotation quality problems do not surface in the hour count; they surface when the client trains on the data and cannot explain why performance is uneven.
The dialect constraint made sourcing genuinely harder. Participants needed the Bangladeshi Bengali background specifically, plus annotation aptitude — which is a separate skill from speaking the language and does not correlate with it.
Annotation aptitude is the part buyers most often leave out of a brief. A fluent speaker who cannot apply a labelling guideline consistently produces data that looks fine line by line and is unusable in aggregate.
The fixed timeline compounded both. Recruiting for a narrow dialect and screening for a separate aptitude, inside a window that does not move, is where a vendor either holds the standard or quietly drops it.
For buyers, the check that matters is simple and rarely asked: what was the screening rejection rate, and what happened to the people who failed. A vendor that cannot answer either has not really screened.
Operating response
What MoniSa changed
Twenty Bengali-speaking participants were recruited against the Bangladeshi dialect requirement, deliberately over-recruiting against the target headcount because a screening step only means something if it can afford to reject.
- Over-recruit to allow rejection Twenty were recruited for a smaller target, because a screening step that cannot afford to fail anyone is not a screening step.
- Aptitude separate from fluency Annotation consistency was tested as its own requirement — a fluent speaker is not automatically a reliable annotator.
- Remove before production, not after The eleven who did not meet standard were removed before contributing data, rather than corrected downstream once the dataset existed.
- Report the funnel, not the survivors The 20-to-9 ratio was reported as the result. A pilot reporting only its final headcount conceals whether screening happened at all.
Results
Measured outcomes from this engagement.
178 hours of annotation were completed and delivered as paid production work in the Bangladeshi Bengali dialect.
| Annotation delivered | 178 hours |
|---|---|
| Participants recruited | 20 |
| Participants cleared | 9 |
| Screening rejection rate | 55% |
| Engagement type | Paid production pilot, not a free trial |
Selection logic
What protected the result.
The selection came down to whether MoniSa could source and review the work at standard, and whether that would hold across the full run.
Why the fit was real
Why the fit was real
The work needed dialect-specific sourcing plus a screening standard that could reject more than half the pool without missing the delivery window.
What decided the result
What decided the result
Holding the annotation standard under pilot-conversion pressure is what protected the dataset — and it is the behaviour that converts pilots into production contracts.
What buyers can reuse
What buyers can reuse
- Ask every pilot vendor for the screening funnel: recruited, cleared, and what disqualified the rest. A project-scoped quality review pass rate is a finding, not a reassurance.
- Specify the dialect, not just the language. Bengali (Bangladeshi) and Bengali (West Bengal) are different sourcing pools for annotation purposes.
- Annotation aptitude is a separate requirement from language fluency and has to be screened separately. Fluency does not predict labelling consistency.
- Over-recruitment is what makes screening real. A vendor recruiting exactly the target headcount has no capacity to reject anyone.
- Prefer a paid production pilot to a free trial. It produces usable output and removes the incentive to build a demonstration rather than a sample of real work.
- A useful pilot brief names the dialect, the annotation guideline, the aptitude threshold, the funnel reporting requirement, and who owns re-screening cost.
- Quality control that happens after data enters the set is correction. Quality control that happens before is screening. Only the second protects the dataset.
Continue from this proof
Useful comparisons for the same problem.
Use these links to compare the case with the matching service, buyer guide, and language coverage.
Mapped context
Service and buyer context
Languages named
Examples referenced in the engagement.
- Bengali (Bangladeshi)
- Annotation screening
- Dialect-specific sourcing
- Paid production pilot
More proof
Related proof
Compare this case with Low-resource speech data QA controls and Annotation manual version control to judge whether the operating pattern fits your brief.
case evidence
Nearest proof pattern.
These related cases keep the next click close to the same kind of work.
Rare-language TEP, two phases
The challenge. An LSP partner needed a 10-day rare-language surge followed by a four-month programme covering materially harder languages.
What we did. MoniSa activated a pre-built bench, ran staggered parallel production, and applied QA per script system including dual-script Kashmiri.
The result. Phase 1 at project-scoped quality review in 10 days; Phase 2 at project-scoped quality review across 12 languages over four months.
Voice-bot transcription
Problem. A partner needed human-side conversational transcription across six languages, where cleaning the transcript would destroy the training signal.
Action. MoniSa held a verbatim convention across six per-language standards and kept the two Spanish variants operationally separate.
Result. ~50 hours per language delivered, with false starts and corrections intact.
Japanese short-form audio
Problem. An AI data partner needed Japanese transcription of high-item-count short-form audio without convention drift between transcribers.
Action. MoniSa deployed a small stable four-person team working inside the partner's own production and review workflow.
Result. 213.89 recorded hours delivered, accepted at project-scoped quality review by the partner's review.
Buyer questions
Ask the questions weak vendors avoid.
Short answers for buyers checking fit, coverage, quality method, and next-step readiness.
What was delivered on this engagement?
Annotation delivered: 178 hours. Participants recruited: 20. Participants cleared: 9
What control kept the work stable?
Holding the annotation standard under pilot-conversion pressure is what protected the dataset — and it is the behaviour that converts pilots into production contracts.
Where should similar work go next?
Use AI data services for the delivery model, AI data annotation buyer guide for buyer-side evaluation, and the contact page for a scoped brief.
Similar brief
Send the constraint behind the metric.
A useful follow-up to a case study names the language mix, review model, deadline, and what proof your buyer team needs before approval.
Production-ready brief
01Closest matching challenge from this case02Language pair, dialect, and script coverage03Volume, cadence, or hours to deliver04Reviewer model and acceptance criteria05Security or platform constraints06Proof needed for stakeholder approvalCapability at a glance
The answers most briefs open by asking for.
Buyers rarely start with who we are. They start with a list of fields to fill. Here are ours, so the first email can be about the work instead.
- Languages and locales
- 300+ languages and 4,500+ dialects, quoted per locale rather than per language — because the dialect decides whether a dataset is usable, whether a market accepts a release, and which specialist the work goes to.
- Specialist network
- 110,000+ verified language specialists — linguists, annotators, and reviewers — plus voice talent and subtitlers, matched to the language, domain and task before assignment.
- Capacity and mobilisation
- Named availability confirmed per pair before scoping. Coverage is reported as staffed today or needing a recruitment window, in writing, before a launch date or release window is agreed.
- Sourcing constraints
- Specialists can be sourced against geographic, residency, locale and demographic requirements — including native-only, in-country, and speaker-diversity quotas where a data programme demands them.
- Deliverables and specs
- Work is delivered to the receiving specification: structured formats and schemas for data and annotation work, and timed-text, audio and platform conformance for media — subtitle reading speed, line limits, cue timing, channel and sample-rate requirements included.
- Comparable work
- 62 documented case studies stating the scope, the constraint that made it difficult, and the measured result — across AI data programmes, partner overflow, and media releases. 2,000+ AI projects delivered and 1,000+ brands served since 2015.
- Certifications
- ISO 9001:2015 quality management, ISO 27001:2022 information security, and ISO 17100 translation services — scoped to translation specifically, and stated that way rather than implied across every line.
- Commercial basis
- Quoted in the unit the work is measured in — per word, per audio hour, per approved hour, per finished minute, per batch, per item — with what the unit includes stated alongside it, whether the quote is for you or for a client you quote onward.
Need this against your own template? Convert your scope between units and check the deadline, then send the brief with your language list, content type, volume and deadline, and the acceptance criteria you will judge the output against — those four decide feasibility, and the reply addresses them directly.