Case study

Transcribing the human half.

A partner needed transcription of the human side of phone conversations with voice bots, across six languages — data that feeds directly back into voice bot improvement.

~300 hours - ~50 hours - 6, including Lao

~300 hours Total transcribed
~50 hours Per language
6, including Lao Languages
Voice-bot transcription visual: Multilingual audio and voice data pipeline on a compressed timeline.
Measured outcomes Voice-bot transcription
~300 hours Total transcribed
~50 hours Per language
6, including Lao Languages
Spain and Mexico handled separately Spanish variants
False starts and corrections preserved Speech features

Project overview

What landed, and what made it hard.

A partner needed transcription of the human side of phone conversations with voice bots, across six languages — data that feeds directly back into voice bot improvement.

Delivery snapshot

Voice-bot transcription

Client
confidential AI data partner
Service
Conversational transcription
Languages
German, Italian, French, Spanish (Spain), Spanish (Mexico), Lao
Volume
~300 hours, ~50 hours per language
Content type
Human-side voice bot conversations

Why this mattered

Outcome before process.

The language set is deliberately uneven: German, Italian, French, and two distinct Spanish variants, plus Lao. Five are well-resourced European languages; one is a Southeast Asian language with a far smaller professional transcription pool.

The two Spanish variants matter. Spain Spanish and Mexican Spanish are not one transcription pool, and treating them as one produces a dataset that misrepresents both.

MoniSa handled the work under ISO 9001:2015 for process control and ISO 27001:2022 for information handling. ISO 17100:2015 is scoped to translation, so it is not claimed for this work.

The defining decision was what to do with imperfect speech. People talking to voice bots restart sentences, correct themselves, and talk over the system. Those features are the signal, not the noise.

The problem to solve

Why the work was difficult, and what MoniSa changed in-flight.

Transcription for voice bot training inverts the usual quality instinct. Most transcription work rewards a clean, readable transcript. Here a clean transcript is a damaged one, because the disfluencies are exactly what the model needs to learn to handle.

The challenge

The problem to solve

That is a hard instruction to hold across a distributed team. Transcribers trained on conventional standards will tidy by reflex — dropping a false start, resolving a self-correction, smoothing an overlap — and each of those edits removes a training signal.

Natural conversational speech is also genuinely harder to transcribe than prepared speech. Overlapping turns, mid-word restarts, and background environments all reduce intelligibility in ways that prepared audio does not.

Six languages meant six distinct quality requirements rather than one standard applied six times. What counts as a self-correction, and how it is represented, differs by language.

The two Spanish variants created a specific risk: a transcriber comfortable in one variant will unconsciously normalize the other toward their own, which quietly corrupts the variant distinction the client was paying to preserve.

Lao carried the sourcing risk. A six-language project where five languages are easy to source and one is not will run at the speed of the hard one, and a plan that averages across all six will miss its dates.

Operating response

What MoniSa changed

Transcription ran across all six languages on the partner's own platform, keeping the workflow inside the environment the client already used rather than exporting and re-importing.

  • Preserve disfluency False starts, self-corrections, and overlaps were kept in the transcript, because they are the training signal a voice bot needs rather than noise to tidy away.
  • Per-language standards Six languages were treated as six quality requirements, since how a self-correction is represented is language-specific.
  • Spanish variants kept apart Spain and Mexico Spanish ran as separate streams so neither was normalized toward the other by transcriber habit.
  • Even per-language volume ~50 hours per language kept the streams comparable and stopped the hardest language hiding inside a combined total.

Results

Measured outcomes from this engagement.

Roughly 300 hours were transcribed across six languages, at approximately 50 hours per language.

Total transcribed~300 hours
Per language~50 hours
Languages6, including Lao
Spanish variantsSpain and Mexico handled separately
Speech featuresFalse starts and corrections preserved

Selection logic

What protected the result.

The selection came down to whether MoniSa could source and review the work at standard, and whether that would hold across the full run.

Why the fit was real

Why the fit was real

The work needed six-language coverage including Lao, plus the discipline to hold an anti-intuitive transcription convention across a distributed team.

What decided the result

What decided the result

Preserving disfluency and keeping the Spanish variants apart protected the exact properties that made the data worth collecting.

What buyers can reuse

What buyers can reuse

  • State the transcription convention explicitly. "Clean" and "verbatim" produce different datasets, and for voice AI training the difference decides whether the data is usable.
  • Disfluencies are signal in conversational AI data. A transcript that reads well may have removed the feature the model most needs to learn.
  • Treat language variants as separate pools. Spain and Mexico Spanish merge easily and cannot be separated after delivery.
  • Watch the hardest language in a multi-language brief. A combined hour total can hide a badly under-delivered stream.
  • Ask how per-language quality is evidenced. One project-level number across six languages is an average, not a quality report.
  • A useful conversational-transcription brief names the convention, the disfluency handling rules, per-language volume, and the variant boundaries.
  • Do not accept an accuracy figure that the engagement never measured. A vendor quoting one anyway is quoting a habit, not a result.

Continue from this proof

Useful comparisons for the same problem.

Use these links to compare the case with the matching service, buyer guide, and language coverage.

Languages named

Examples referenced in the engagement.

  • German
  • Italian
  • French
  • Spanish (Spain)
  • Spanish (Mexico)
  • Lao

case evidence

Nearest proof pattern.

These related cases keep the next click close to the same kind of work.

Translation services607,000 words across 17 rare languages and 5 scripts, delivered in two phases with accuracy reported per phase.

Rare-language TEP, two phases

The challenge. An LSP partner needed a 10-day rare-language surge followed by a four-month programme covering materially harder languages.

What we did. MoniSa activated a pre-built bench, ran staggered parallel production, and applied QA per script system including dual-script Kashmiri.

The result. Phase 1 at project-scoped quality review in 10 days; Phase 2 at project-scoped quality review across 12 languages over four months.

Open full case
AI data services213.89 hours of Japanese short-form audio accepted at project-scoped quality review on partner review.

Japanese short-form audio

Problem. An AI data partner needed Japanese transcription of high-item-count short-form audio without convention drift between transcribers.

Action. MoniSa deployed a small stable four-person team working inside the partner's own production and review workflow.

Result. 213.89 recorded hours delivered, accepted at project-scoped quality review by the partner's review.

Open full case
AI data services500 hours transcribed at project-scoped quality review on peer review, with accent-specific transcriber pools.

Regional-accent transcription

Problem. A partner needed French Canadian, Russian, and Persian transcription where regional variety and technical terminology both had to hold.

Action. MoniSa sourced by variety rather than language, deployed 15 transcribers, and ran peer review and spot checks during production.

Result. 500 hours with project-scoped quality review measured by spot-check peer review.

Open full case

Buyer questions

Ask the questions weak vendors avoid.

Short answers for buyers checking fit, coverage, quality method, and next-step readiness.

What was delivered on this engagement?

Total transcribed: ~300 hours. Per language: ~50 hours. Languages: 6, including Lao

What control kept the work stable?

Preserving disfluency and keeping the Spanish variants apart protected the exact properties that made the data worth collecting.

Where should similar work go next?

Use AI data services for the delivery model, Speech data collection buyer guide for buyer-side evaluation, and the contact page for a scoped brief.

Similar brief

Send the constraint behind the metric.

A useful follow-up to a case study names the language mix, review model, deadline, and what proof your buyer team needs before approval.

Production-ready brief

01Closest matching challenge from this case02Language pair, dialect, and script coverage03Volume, cadence, or hours to deliver04Reviewer model and acceptance criteria05Security or platform constraints06Proof needed for stakeholder approval

Capability at a glance

The answers most briefs open by asking for.

Buyers rarely start with who we are. They start with a list of fields to fill. Here are ours, so the first email can be about the work instead.

Languages and locales
300+ languages and 4,500+ dialects, quoted per locale rather than per language — because the dialect decides whether a dataset is usable, whether a market accepts a release, and which specialist the work goes to.
Specialist network
110,000+ verified language specialists — linguists, annotators, and reviewers — plus voice talent and subtitlers, matched to the language, domain and task before assignment.
Capacity and mobilisation
Named availability confirmed per pair before scoping. Coverage is reported as staffed today or needing a recruitment window, in writing, before a launch date or release window is agreed.
Sourcing constraints
Specialists can be sourced against geographic, residency, locale and demographic requirements — including native-only, in-country, and speaker-diversity quotas where a data programme demands them.
Deliverables and specs
Work is delivered to the receiving specification: structured formats and schemas for data and annotation work, and timed-text, audio and platform conformance for media — subtitle reading speed, line limits, cue timing, channel and sample-rate requirements included.
Comparable work
62 documented case studies stating the scope, the constraint that made it difficult, and the measured result — across AI data programmes, partner overflow, and media releases. 2,000+ AI projects delivered and 1,000+ brands served since 2015.
Certifications
ISO 9001:2015 quality management, ISO 27001:2022 information security, and ISO 17100 translation services — scoped to translation specifically, and stated that way rather than implied across every line.
Commercial basis
Quoted in the unit the work is measured in — per word, per audio hour, per approved hour, per finished minute, per batch, per item — with what the unit includes stated alongside it, whether the quote is for you or for a client you quote onward.

Need this against your own template? Convert your scope between units and check the deadline, then send the brief with your language list, content type, volume and deadline, and the acceptance criteria you will judge the output against — those four decide feasibility, and the reply addresses them directly.

Scope a project Call