Case study

A pilot that rejected 11 of 20.

An AI company needed a Bengali data collection pilot in the Bangladeshi dialect, recruited and delivered inside a fixed window. Twenty participants were recruited. Nine cleared screening.

178 hours - 20 - 9

178 hours Annotation delivered
20 Participants recruited
9 Participants cleared
Bengali pilot, screened visual: Speech transcription QA workflow with qualification, task tracking, and correction controls.
Measured outcomes Bengali pilot, screened
178 hours Annotation delivered
20 Participants recruited
9 Participants cleared
55% Screening rejection rate
Paid production pilot, not a free trial Engagement type

Project overview

What landed, and what made it hard.

An AI company needed a Bengali data collection pilot in the Bangladeshi dialect, recruited and delivered inside a fixed window. Twenty participants were recruited. Nine cleared screening.

Delivery snapshot

Bengali pilot, screened

Client
confidential AI company, via an LSP partner platform
Service
Data collection and annotation pilot
Language
Bengali (Bangladeshi)
Volume
178 hours
Resources
9 annotators, from 20 recruited

Why this mattered

Outcome before process.

That 20-to-9 funnel is the case. A 55% rejection rate looks like a recruitment failure until you consider the alternative: eleven annotators who did not meet the standard, contributing to a dataset the client would train on.

The dialect requirement narrowed the pool before screening even started. Bengali spoken in Bangladesh is not interchangeable with Bengali spoken in West Bengal for annotation purposes, and a brief that says only "Bengali" gets a dataset that averages across both.

MoniSa handled the work under ISO 9001:2015 for process control and ISO 27001:2022 for information handling. ISO 17100:2015 is scoped to translation, so it is not claimed for this work.

This ran as a paid production pilot rather than a free trial, which changes the incentive structure on both sides: the client was buying real output, and the screening had to be defensible rather than generous.

The problem to solve

Why the work was difficult, and what MoniSa changed in-flight.

Pilots carry a structural temptation. The vendor wants the pilot to convert into a production contract, and the fastest way to make a pilot look successful is to pass everyone through screening and deliver the headcount the brief asked for.

The challenge

The problem to solve

That trade is invisible at pilot stage and expensive later. Annotation quality problems do not surface in the hour count; they surface when the client trains on the data and cannot explain why performance is uneven.

The dialect constraint made sourcing genuinely harder. Participants needed the Bangladeshi Bengali background specifically, plus annotation aptitude — which is a separate skill from speaking the language and does not correlate with it.

Annotation aptitude is the part buyers most often leave out of a brief. A fluent speaker who cannot apply a labelling guideline consistently produces data that looks fine line by line and is unusable in aggregate.

The fixed timeline compounded both. Recruiting for a narrow dialect and screening for a separate aptitude, inside a window that does not move, is where a vendor either holds the standard or quietly drops it.

For buyers, the check that matters is simple and rarely asked: what was the screening rejection rate, and what happened to the people who failed. A vendor that cannot answer either has not really screened.

Operating response

What MoniSa changed

Twenty Bengali-speaking participants were recruited against the Bangladeshi dialect requirement, deliberately over-recruiting against the target headcount because a screening step only means something if it can afford to reject.

  • Over-recruit to allow rejection Twenty were recruited for a smaller target, because a screening step that cannot afford to fail anyone is not a screening step.
  • Aptitude separate from fluency Annotation consistency was tested as its own requirement — a fluent speaker is not automatically a reliable annotator.
  • Remove before production, not after The eleven who did not meet standard were removed before contributing data, rather than corrected downstream once the dataset existed.
  • Report the funnel, not the survivors The 20-to-9 ratio was reported as the result. A pilot reporting only its final headcount conceals whether screening happened at all.

Results

Measured outcomes from this engagement.

178 hours of annotation were completed and delivered as paid production work in the Bangladeshi Bengali dialect.

Annotation delivered178 hours
Participants recruited20
Participants cleared9
Screening rejection rate55%
Engagement typePaid production pilot, not a free trial

Selection logic

What protected the result.

The selection came down to whether MoniSa could source and review the work at standard, and whether that would hold across the full run.

Why the fit was real

Why the fit was real

The work needed dialect-specific sourcing plus a screening standard that could reject more than half the pool without missing the delivery window.

What decided the result

What decided the result

Holding the annotation standard under pilot-conversion pressure is what protected the dataset — and it is the behaviour that converts pilots into production contracts.

What buyers can reuse

What buyers can reuse

  • Ask every pilot vendor for the screening funnel: recruited, cleared, and what disqualified the rest. A project-scoped quality review pass rate is a finding, not a reassurance.
  • Specify the dialect, not just the language. Bengali (Bangladeshi) and Bengali (West Bengal) are different sourcing pools for annotation purposes.
  • Annotation aptitude is a separate requirement from language fluency and has to be screened separately. Fluency does not predict labelling consistency.
  • Over-recruitment is what makes screening real. A vendor recruiting exactly the target headcount has no capacity to reject anyone.
  • Prefer a paid production pilot to a free trial. It produces usable output and removes the incentive to build a demonstration rather than a sample of real work.
  • A useful pilot brief names the dialect, the annotation guideline, the aptitude threshold, the funnel reporting requirement, and who owns re-screening cost.
  • Quality control that happens after data enters the set is correction. Quality control that happens before is screening. Only the second protects the dataset.

Continue from this proof

Useful comparisons for the same problem.

Use these links to compare the case with the matching service, buyer guide, and language coverage.

Languages named

Examples referenced in the engagement.

  • Bengali (Bangladeshi)
  • Annotation screening
  • Dialect-specific sourcing
  • Paid production pilot

case evidence

Nearest proof pattern.

These related cases keep the next click close to the same kind of work.

Translation services607,000 words across 17 rare languages and 5 scripts, delivered in two phases with accuracy reported per phase.

Rare-language TEP, two phases

The challenge. An LSP partner needed a 10-day rare-language surge followed by a four-month programme covering materially harder languages.

What we did. MoniSa activated a pre-built bench, ran staggered parallel production, and applied QA per script system including dual-script Kashmiri.

The result. Phase 1 at project-scoped quality review in 10 days; Phase 2 at project-scoped quality review across 12 languages over four months.

Open full case
AI data services~300 hours of voice bot conversation transcribed with disfluencies preserved for training value.

Voice-bot transcription

Problem. A partner needed human-side conversational transcription across six languages, where cleaning the transcript would destroy the training signal.

Action. MoniSa held a verbatim convention across six per-language standards and kept the two Spanish variants operationally separate.

Result. ~50 hours per language delivered, with false starts and corrections intact.

Open full case
AI data services213.89 hours of Japanese short-form audio accepted at project-scoped quality review on partner review.

Japanese short-form audio

Problem. An AI data partner needed Japanese transcription of high-item-count short-form audio without convention drift between transcribers.

Action. MoniSa deployed a small stable four-person team working inside the partner's own production and review workflow.

Result. 213.89 recorded hours delivered, accepted at project-scoped quality review by the partner's review.

Open full case

Buyer questions

Ask the questions weak vendors avoid.

Short answers for buyers checking fit, coverage, quality method, and next-step readiness.

What was delivered on this engagement?

Annotation delivered: 178 hours. Participants recruited: 20. Participants cleared: 9

What control kept the work stable?

Holding the annotation standard under pilot-conversion pressure is what protected the dataset — and it is the behaviour that converts pilots into production contracts.

Where should similar work go next?

Use AI data services for the delivery model, AI data annotation buyer guide for buyer-side evaluation, and the contact page for a scoped brief.

Similar brief

Send the constraint behind the metric.

A useful follow-up to a case study names the language mix, review model, deadline, and what proof your buyer team needs before approval.

Production-ready brief

01Closest matching challenge from this case02Language pair, dialect, and script coverage03Volume, cadence, or hours to deliver04Reviewer model and acceptance criteria05Security or platform constraints06Proof needed for stakeholder approval

Capability at a glance

The answers most briefs open by asking for.

Buyers rarely start with who we are. They start with a list of fields to fill. Here are ours, so the first email can be about the work instead.

Languages and locales
300+ languages and 4,500+ dialects, quoted per locale rather than per language — because the dialect decides whether a dataset is usable, whether a market accepts a release, and which specialist the work goes to.
Specialist network
110,000+ verified language specialists — linguists, annotators, and reviewers — plus voice talent and subtitlers, matched to the language, domain and task before assignment.
Capacity and mobilisation
Named availability confirmed per pair before scoping. Coverage is reported as staffed today or needing a recruitment window, in writing, before a launch date or release window is agreed.
Sourcing constraints
Specialists can be sourced against geographic, residency, locale and demographic requirements — including native-only, in-country, and speaker-diversity quotas where a data programme demands them.
Deliverables and specs
Work is delivered to the receiving specification: structured formats and schemas for data and annotation work, and timed-text, audio and platform conformance for media — subtitle reading speed, line limits, cue timing, channel and sample-rate requirements included.
Comparable work
62 documented case studies stating the scope, the constraint that made it difficult, and the measured result — across AI data programmes, partner overflow, and media releases. 2,000+ AI projects delivered and 1,000+ brands served since 2015.
Certifications
ISO 9001:2015 quality management, ISO 27001:2022 information security, and ISO 17100 translation services — scoped to translation specifically, and stated that way rather than implied across every line.
Commercial basis
Quoted in the unit the work is measured in — per word, per audio hour, per approved hour, per finished minute, per batch, per item — with what the unit includes stated alongside it, whether the quote is for you or for a client you quote onward.

Need this against your own template? Convert your scope between units and check the deadline, then send the brief with your language list, content type, volume and deadline, and the acceptance criteria you will judge the output against — those four decide feasibility, and the reply addresses them directly.

Scope a project Call