Case study
Transcribing the human half.
A partner needed transcription of the human side of phone conversations with voice bots, across six languages — data that feeds directly back into voice bot improvement.
~300 hours - ~50 hours - 6, including Lao
Project overview
What landed, and what made it hard.
A partner needed transcription of the human side of phone conversations with voice bots, across six languages — data that feeds directly back into voice bot improvement.
Delivery snapshot
Voice-bot transcription
- Client
- confidential AI data partner
- Service
- Conversational transcription
- Languages
- German, Italian, French, Spanish (Spain), Spanish (Mexico), Lao
- Volume
- ~300 hours, ~50 hours per language
- Content type
- Human-side voice bot conversations
Why this mattered
Outcome before process.
The language set is deliberately uneven: German, Italian, French, and two distinct Spanish variants, plus Lao. Five are well-resourced European languages; one is a Southeast Asian language with a far smaller professional transcription pool.
The two Spanish variants matter. Spain Spanish and Mexican Spanish are not one transcription pool, and treating them as one produces a dataset that misrepresents both.
MoniSa handled the work under ISO 9001:2015 for process control and ISO 27001:2022 for information handling. ISO 17100:2015 is scoped to translation, so it is not claimed for this work.
The defining decision was what to do with imperfect speech. People talking to voice bots restart sentences, correct themselves, and talk over the system. Those features are the signal, not the noise.
The problem to solve
Why the work was difficult, and what MoniSa changed in-flight.
Transcription for voice bot training inverts the usual quality instinct. Most transcription work rewards a clean, readable transcript. Here a clean transcript is a damaged one, because the disfluencies are exactly what the model needs to learn to handle.
The challenge
The problem to solve
That is a hard instruction to hold across a distributed team. Transcribers trained on conventional standards will tidy by reflex — dropping a false start, resolving a self-correction, smoothing an overlap — and each of those edits removes a training signal.
Natural conversational speech is also genuinely harder to transcribe than prepared speech. Overlapping turns, mid-word restarts, and background environments all reduce intelligibility in ways that prepared audio does not.
Six languages meant six distinct quality requirements rather than one standard applied six times. What counts as a self-correction, and how it is represented, differs by language.
The two Spanish variants created a specific risk: a transcriber comfortable in one variant will unconsciously normalize the other toward their own, which quietly corrupts the variant distinction the client was paying to preserve.
Lao carried the sourcing risk. A six-language project where five languages are easy to source and one is not will run at the speed of the hard one, and a plan that averages across all six will miss its dates.
Operating response
What MoniSa changed
Transcription ran across all six languages on the partner's own platform, keeping the workflow inside the environment the client already used rather than exporting and re-importing.
- Preserve disfluency False starts, self-corrections, and overlaps were kept in the transcript, because they are the training signal a voice bot needs rather than noise to tidy away.
- Per-language standards Six languages were treated as six quality requirements, since how a self-correction is represented is language-specific.
- Spanish variants kept apart Spain and Mexico Spanish ran as separate streams so neither was normalized toward the other by transcriber habit.
- Even per-language volume ~50 hours per language kept the streams comparable and stopped the hardest language hiding inside a combined total.
Results
Measured outcomes from this engagement.
Roughly 300 hours were transcribed across six languages, at approximately 50 hours per language.
| Total transcribed | ~300 hours |
|---|---|
| Per language | ~50 hours |
| Languages | 6, including Lao |
| Spanish variants | Spain and Mexico handled separately |
| Speech features | False starts and corrections preserved |
Selection logic
What protected the result.
The selection came down to whether MoniSa could source and review the work at standard, and whether that would hold across the full run.
Why the fit was real
Why the fit was real
The work needed six-language coverage including Lao, plus the discipline to hold an anti-intuitive transcription convention across a distributed team.
What decided the result
What decided the result
Preserving disfluency and keeping the Spanish variants apart protected the exact properties that made the data worth collecting.
What buyers can reuse
What buyers can reuse
- State the transcription convention explicitly. "Clean" and "verbatim" produce different datasets, and for voice AI training the difference decides whether the data is usable.
- Disfluencies are signal in conversational AI data. A transcript that reads well may have removed the feature the model most needs to learn.
- Treat language variants as separate pools. Spain and Mexico Spanish merge easily and cannot be separated after delivery.
- Watch the hardest language in a multi-language brief. A combined hour total can hide a badly under-delivered stream.
- Ask how per-language quality is evidenced. One project-level number across six languages is an average, not a quality report.
- A useful conversational-transcription brief names the convention, the disfluency handling rules, per-language volume, and the variant boundaries.
- Do not accept an accuracy figure that the engagement never measured. A vendor quoting one anyway is quoting a habit, not a result.
Continue from this proof
Useful comparisons for the same problem.
Use these links to compare the case with the matching service, buyer guide, and language coverage.
Mapped context
Service and buyer context
Languages named
Examples referenced in the engagement.
- German
- Italian
- French
- Spanish (Spain)
- Spanish (Mexico)
- Lao
More proof
Related proof
Compare this case with Long-form transcription across 4 locales and Audio transcription at standing scale to judge whether the operating pattern fits your brief.
case evidence
Nearest proof pattern.
These related cases keep the next click close to the same kind of work.
Japanese short-form audio
The challenge. An AI data partner needed Japanese transcription of high-item-count short-form audio without convention drift between transcribers.
What we did. MoniSa deployed a small stable four-person team working inside the partner's own production and review workflow.
The result. 213.89 recorded hours delivered, accepted at project-scoped quality review by the partner's review.
Recognise your own project in one of these?
Send the language list and volumeRare-language TEP, two phases
Problem. An LSP partner needed a 10-day rare-language surge followed by a four-month programme covering materially harder languages.
Action. MoniSa activated a pre-built bench, ran staggered parallel production, and applied QA per script system including dual-script Kashmiri.
Result. Phase 1 at project-scoped quality review in 10 days; Phase 2 at project-scoped quality review across 12 languages over four months.
Regional-accent transcription
Problem. A partner needed French Canadian, Russian, and Persian transcription where regional variety and technical terminology both had to hold.
Action. MoniSa sourced by variety rather than language, deployed 15 transcribers, and ran peer review and spot checks during production.
Result. 500 hours with project-scoped quality review measured by spot-check peer review.
Buyer questions
Answers in writing, before you ask for a call.
The questions buyers send before a scope conversation, answered on the page rather than in a meeting. Take them to your team, then send us the one we did not answer.
What was delivered on this engagement?
Total transcribed: ~300 hours. Per language: ~50 hours. Languages: 6, including Lao
What control kept the work stable?
Preserving disfluency and keeping the Spanish variants apart protected the exact properties that made the data worth collecting.
Where should similar work go next?
Use AI data services for the delivery model, Speech data collection buyer guide for buyer-side evaluation, and the contact page for a scoped brief.
What happens if you cannot staff one of my language pairs?
You are told before a date is agreed, not after. Coverage is reported pair by pair as staffed today or needing a recruitment window, with the window stated — in writing, while the scope is still being agreed. Nobody new goes onto live work until a pilot batch has been reviewed and signed off. A coverage claim you cannot check before signing is not coverage.
Similar brief
Send the constraint behind the metric.
A useful follow-up to a case study names the language mix, review model, deadline, and what proof your buyer team needs before approval.
Production-ready brief
01Closest matching challenge from this case02Language pair, dialect, and script coverage03Volume, cadence, or hours to deliver04Reviewer model and acceptance criteria05Security or platform constraints06Proof needed for stakeholder approval