Case study
Transcribing the human half.
A partner needed transcription of the human side of phone conversations with voice bots, across six languages — data that feeds directly back into voice bot improvement.
~300 hours - ~50 hours - 6, including Lao
Project overview
What landed, and what made it hard.
A partner needed transcription of the human side of phone conversations with voice bots, across six languages — data that feeds directly back into voice bot improvement.
Delivery snapshot
Voice-bot transcription
- Client
- confidential AI data partner
- Service
- Conversational transcription
- Languages
- German, Italian, French, Spanish (Spain), Spanish (Mexico), Lao
- Volume
- ~300 hours, ~50 hours per language
- Content type
- Human-side voice bot conversations
Why this mattered
Outcome before process.
The language set is deliberately uneven: German, Italian, French, and two distinct Spanish variants, plus Lao. Five are well-resourced European languages; one is a Southeast Asian language with a far smaller professional transcription pool.
The two Spanish variants matter. Spain Spanish and Mexican Spanish are not one transcription pool, and treating them as one produces a dataset that misrepresents both.
MoniSa handled the work under ISO 9001:2015 for process control and ISO 27001:2022 for information handling. ISO 17100:2015 is scoped to translation, so it is not claimed for this work.
The defining decision was what to do with imperfect speech. People talking to voice bots restart sentences, correct themselves, and talk over the system. Those features are the signal, not the noise.
The problem to solve
Why the work was difficult, and what MoniSa changed in-flight.
Transcription for voice bot training inverts the usual quality instinct. Most transcription work rewards a clean, readable transcript. Here a clean transcript is a damaged one, because the disfluencies are exactly what the model needs to learn to handle.
The challenge
The problem to solve
That is a hard instruction to hold across a distributed team. Transcribers trained on conventional standards will tidy by reflex — dropping a false start, resolving a self-correction, smoothing an overlap — and each of those edits removes a training signal.
Natural conversational speech is also genuinely harder to transcribe than prepared speech. Overlapping turns, mid-word restarts, and background environments all reduce intelligibility in ways that prepared audio does not.
Six languages meant six distinct quality requirements rather than one standard applied six times. What counts as a self-correction, and how it is represented, differs by language.
The two Spanish variants created a specific risk: a transcriber comfortable in one variant will unconsciously normalize the other toward their own, which quietly corrupts the variant distinction the client was paying to preserve.
Lao carried the sourcing risk. A six-language project where five languages are easy to source and one is not will run at the speed of the hard one, and a plan that averages across all six will miss its dates.
Operating response
What MoniSa changed
Transcription ran across all six languages on the partner's own platform, keeping the workflow inside the environment the client already used rather than exporting and re-importing.
- Preserve disfluency False starts, self-corrections, and overlaps were kept in the transcript, because they are the training signal a voice bot needs rather than noise to tidy away.
- Per-language standards Six languages were treated as six quality requirements, since how a self-correction is represented is language-specific.
- Spanish variants kept apart Spain and Mexico Spanish ran as separate streams so neither was normalized toward the other by transcriber habit.
- Even per-language volume ~50 hours per language kept the streams comparable and stopped the hardest language hiding inside a combined total.
Results
Measured outcomes from this engagement.
Roughly 300 hours were transcribed across six languages, at approximately 50 hours per language.
| Total transcribed | ~300 hours |
|---|---|
| Per language | ~50 hours |
| Languages | 6, including Lao |
| Spanish variants | Spain and Mexico handled separately |
| Speech features | False starts and corrections preserved |
Selection logic
What protected the result.
The selection came down to whether MoniSa could source and review the work at standard, and whether that would hold across the full run.
Why the fit was real
Why the fit was real
The work needed six-language coverage including Lao, plus the discipline to hold an anti-intuitive transcription convention across a distributed team.
What decided the result
What decided the result
Preserving disfluency and keeping the Spanish variants apart protected the exact properties that made the data worth collecting.
What buyers can reuse
What buyers can reuse
- State the transcription convention explicitly. "Clean" and "verbatim" produce different datasets, and for voice AI training the difference decides whether the data is usable.
- Disfluencies are signal in conversational AI data. A transcript that reads well may have removed the feature the model most needs to learn.
- Treat language variants as separate pools. Spain and Mexico Spanish merge easily and cannot be separated after delivery.
- Watch the hardest language in a multi-language brief. A combined hour total can hide a badly under-delivered stream.
- Ask how per-language quality is evidenced. One project-level number across six languages is an average, not a quality report.
- A useful conversational-transcription brief names the convention, the disfluency handling rules, per-language volume, and the variant boundaries.
- Do not accept an accuracy figure that the engagement never measured. A vendor quoting one anyway is quoting a habit, not a result.
Continue from this proof
Useful comparisons for the same problem.
Use these links to compare the case with the matching service, buyer guide, and language coverage.
Mapped context
Service and buyer context
Languages named
Examples referenced in the engagement.
- German
- Italian
- French
- Spanish (Spain)
- Spanish (Mexico)
- Lao
More proof
Related proof
Compare this case with Long-form transcription across 4 locales and Audio transcription at standing scale to judge whether the operating pattern fits your brief.
case evidence
Nearest proof pattern.
These related cases keep the next click close to the same kind of work.
Rare-language TEP, two phases
The challenge. An LSP partner needed a 10-day rare-language surge followed by a four-month programme covering materially harder languages.
What we did. MoniSa activated a pre-built bench, ran staggered parallel production, and applied QA per script system including dual-script Kashmiri.
The result. Phase 1 at project-scoped quality review in 10 days; Phase 2 at project-scoped quality review across 12 languages over four months.
Japanese short-form audio
Problem. An AI data partner needed Japanese transcription of high-item-count short-form audio without convention drift between transcribers.
Action. MoniSa deployed a small stable four-person team working inside the partner's own production and review workflow.
Result. 213.89 recorded hours delivered, accepted at project-scoped quality review by the partner's review.
Regional-accent transcription
Problem. A partner needed French Canadian, Russian, and Persian transcription where regional variety and technical terminology both had to hold.
Action. MoniSa sourced by variety rather than language, deployed 15 transcribers, and ran peer review and spot checks during production.
Result. 500 hours with project-scoped quality review measured by spot-check peer review.
Buyer questions
Ask the questions weak vendors avoid.
Short answers for buyers checking fit, coverage, quality method, and next-step readiness.
What was delivered on this engagement?
Total transcribed: ~300 hours. Per language: ~50 hours. Languages: 6, including Lao
What control kept the work stable?
Preserving disfluency and keeping the Spanish variants apart protected the exact properties that made the data worth collecting.
Where should similar work go next?
Use AI data services for the delivery model, Speech data collection buyer guide for buyer-side evaluation, and the contact page for a scoped brief.
Similar brief
Send the constraint behind the metric.
A useful follow-up to a case study names the language mix, review model, deadline, and what proof your buyer team needs before approval.
Production-ready brief
01Closest matching challenge from this case02Language pair, dialect, and script coverage03Volume, cadence, or hours to deliver04Reviewer model and acceptance criteria05Security or platform constraints06Proof needed for stakeholder approvalCapability at a glance
The answers most briefs open by asking for.
Buyers rarely start with who we are. They start with a list of fields to fill. Here are ours, so the first email can be about the work instead.
- Languages and locales
- 300+ languages and 4,500+ dialects, quoted per locale rather than per language — because the dialect decides whether a dataset is usable, whether a market accepts a release, and which specialist the work goes to.
- Specialist network
- 110,000+ verified language specialists — linguists, annotators, and reviewers — plus voice talent and subtitlers, matched to the language, domain and task before assignment.
- Capacity and mobilisation
- Named availability confirmed per pair before scoping. Coverage is reported as staffed today or needing a recruitment window, in writing, before a launch date or release window is agreed.
- Sourcing constraints
- Specialists can be sourced against geographic, residency, locale and demographic requirements — including native-only, in-country, and speaker-diversity quotas where a data programme demands them.
- Deliverables and specs
- Work is delivered to the receiving specification: structured formats and schemas for data and annotation work, and timed-text, audio and platform conformance for media — subtitle reading speed, line limits, cue timing, channel and sample-rate requirements included.
- Comparable work
- 62 documented case studies stating the scope, the constraint that made it difficult, and the measured result — across AI data programmes, partner overflow, and media releases. 2,000+ AI projects delivered and 1,000+ brands served since 2015.
- Certifications
- ISO 9001:2015 quality management, ISO 27001:2022 information security, and ISO 17100 translation services — scoped to translation specifically, and stated that way rather than implied across every line.
- Commercial basis
- Quoted in the unit the work is measured in — per word, per audio hour, per approved hour, per finished minute, per batch, per item — with what the unit includes stated alongside it, whether the quote is for you or for a client you quote onward.
Need this against your own template? Convert your scope between units and check the deadline, then send the brief with your language list, content type, volume and deadline, and the acceptance criteria you will judge the output against — those four decide feasibility, and the reply addresses them directly.