Voice data that came back.

A client needed roughly 500 hours of US English voice data with specific requirements on sample diversity, voice characteristics, and security — inside a budget that did not flex.

~500 hours - English (US) - reviewed quality

110,000+ native linguists and AI data contributors · Founder-reported combined network · 4 Oct 2026
300+ languages across active service lines
4,500+ dialects and regional variants
110+ rare, indigenous and low-resource languages
1,000+ organizations served since 2015
Voice data, 500 hours visual: A voice contributor recording in a booth while an engineer runs the session from the console.
Measured outcomes Voice data, 500 hours
~500 hours Volume collected
English (US) Language
reviewed quality Client satisfaction
Extended into follow-on engagements Commercial outcome
Included rare-language work Follow-on scope

The project

Voice data, 500 hours

Client
confidential AI product company
Service
Voice data collection
Volume
~500 hours
Language
English (US)
Client rating
reviewed quality

That combination is the ordinary shape of speech data work and it is where most engagements quietly degrade. Diversity costs money. Security controls cost money. A fixed budget pushes a vendor toward the cheapest available speakers, and the dataset narrows without anyone deciding that it should.

Collection ran through the client's own mobile recording application, which meant the workflow had to fit their tooling rather than ours, with privacy compliance handled inside their environment.

MoniSa handled the work under ISO 9001:2015 for process control and ISO 27001:2022 for information handling. ISO 17100:2015 is scoped to translation, so it is not claimed for this work.

The outcome that mattered commercially was not the hour count. It was that the client returned with follow-on work, including rare-language engagements — converting a single project into a continuing relationship.

The problem to solve

Voice data has a diversity requirement that a raw hour count cannot express. 500 hours from a narrow speaker pool trains a model that works for that pool. The specification named sample diversity and voice characteristics precisely because the count alone would not have protected the dataset.

Budget was a stated constraint rather than a background pressure. The engagement required balancing collection quality against cost explicitly, which is a harder brief than either "cheapest" or "best" and one where the failure is invisible until model performance drops.

Privacy compliance shaped the workflow. Voice data is personal data. Collection through the client's mobile application meant security measures had to operate inside their app environment and their consent flow, not around it.

Working inside a client's own tooling removes a vendor's usual levers. There is no option to substitute a familiar recording pipeline or QC harness; the quality controls have to be built to fit what the application already does.

The QC burden in voice work is also different from text. A sample can be linguistically perfect and technically unusable — background noise, clipping, inconsistent microphone distance, wrong format. Both layers have to be checked, and only one of them is audible to a casual listener.

For buyers, the practical checks are: who defines speaker diversity and how it is evidenced, where consent and personal data are held, what the technical rejection criteria are, and who bears the cost of re-recording.

What MoniSa changed

Collection ran through the client's mobile recording application, keeping voice data and consent inside the environment the client already controlled rather than moving personal data into a second system.

  • Client-side tooling

    Collection ran inside the client's own mobile recording application, so voice data and consent never moved into a second environment.

  • Privacy by design

    Security controls for personal data were part of the collection design rather than a review applied after the recordings existed.

  • Two-layer QC

    Samples were checked against technical specification and linguistic quality, because a recording fails on format as completely as on content.

  • Budget held without narrowing the pool

    Commercial terms were negotiated so cost pressure did not quietly reduce speaker diversity — the usual hidden cost of a fixed-budget speech project.

Results

Measured outcomes from this engagement.

Roughly 500 hours were collected at a reviewed quality client satisfaction rating, inside the budget constraint that framed the engagement.

Volume collected~500 hours
LanguageEnglish (US)
Client satisfactionreviewed quality
Commercial outcomeExtended into follow-on engagements
Follow-on scopeIncluded rare-language work

What supported the result

Why the fit was real

The work required collection discipline inside the client's own tooling and privacy environment, with quality held against a budget that could not move.

What decided the result

Protecting speaker diversity under cost pressure is what made the dataset usable — and what earned the follow-on rare-language scope.

What buyers can reuse

  • Hour count is not dataset quality. Ask how speaker diversity is defined, evidenced, and protected when budget tightens.
  • Voice data is personal data. Confirm where recordings and consent are held, and who is accountable if the collection tooling belongs to the client.
  • Technical rejection criteria should be written before collection. Noise, clipping, microphone distance, and format fail a sample as completely as content does.
  • A fixed budget is a specification, not a background condition. Ask explicitly what a vendor will trade away to meet it.
  • Repeat scope is a better quality signal than a satisfaction rating. Ask which engagements were extended and whether the follow-on work was harder than the original.
  • A useful speech brief names the language variety, speaker diversity requirements, technical specification, consent and storage path, and the re-record cost owner.
  • Approximate volumes reported as approximate are a good sign. Round numbers in speech collection usually describe a target rather than a delivery.

Continue from this proof

Useful comparisons for the same problem.

Use these links to compare the case with the matching service, buyer guide, and language coverage.

Languages named

Examples referenced in the engagement.

  • English (US)
  • Voice data collection
  • Privacy-compliant recording
  • Rare language follow-on

case evidence

Related projects.

These related cases keep the next click close to the same kind of work.

AI data services178 annotated hours delivered with 11 of 20 recruits removed before they touched production data.

Bengali pilot, screened

The challenge. An AI company needed a Bangladeshi Bengali annotation pilot on a fixed timeline, where dialect and annotation aptitude are separate requirements.

What we did. MoniSa over-recruited, screened aptitude separately from fluency, and removed those who did not meet standard before production.

The result. 9 of 20 cleared screening; 178 hours delivered as paid production work with the funnel reported in full.

Open full case

Recognise your own project in one of these?

Send the language list and volume
AI data services~300 hours of voice bot conversation transcribed with disfluencies preserved for training value.

Voice-bot transcription

Problem. A partner needed human-side conversational transcription across six languages, where cleaning the transcript would destroy the training signal.

Action. MoniSa held a verbatim convention across six per-language standards and kept the two Spanish variants operationally separate.

Result. ~50 hours per language delivered, with false starts and corrections intact.

Open full case
AI data services213.89 hours of Japanese short-form audio accepted under project-scoped review on partner review.

Japanese short-form audio

Problem. An AI data partner needed Japanese transcription of high-item-count short-form audio without convention drift between transcribers.

Action. MoniSa deployed a small stable four-person team working inside the partner's own production and review workflow.

Result. 213.89 recorded hours delivered, accepted under project-scoped review by the partner's review.

Open full case

Buyer questions

Common questions.

What was delivered on this engagement?

Volume collected: ~500 hours. Language: English (US). Client satisfaction: reviewed quality

What control kept the work stable?

Protecting speaker diversity under cost pressure is what made the dataset usable — and what earned the follow-on rare-language scope.

Where should similar work go next?

Use AI data services for the delivery model, Speech data collection buyer guide for buyer-side evaluation, and the contact page for a scoped brief.

What happens if you cannot staff one of my language pairs?

For the proposed project, ask for pair-by-pair availability or a recruitment window in writing before agreeing a date. Define qualification and pilot approval for any new contributor before live work. A coverage claim should be checkable before the scope is signed.

Similar brief

Send the constraint behind the metric.

A useful follow-up to a case study names the language mix, review model, deadline, and what proof your buyer team needs before approval.

Send a brief

Do not paste raw outputs, source records, transcripts, third-party personal data or confidential files here. We will agree a transfer path after scoping.

Required. We will reply about your project.

Scope a project Call