Why is a number of hours not a specification?

Two suppliers can receive the same request, a set number of hours of speech in a named language, and each deliver exactly that. One collection might be mostly younger city speakers reading sentences in quiet rooms, transcribed to the written standard. The other might be speakers of mixed ages talking freely into their own phones at home and outdoors, with every hesitation and language switch kept. Both meet the brief. They represent different populations and support different uses, and the request never said which was wanted.

That choice belongs to the programme owner, before recruitment opens. It is a separate job from choosing who does the collecting, which the speech data collection buyer guide covers, and from checking whether a finished dataset may be used as planned, which why usable data is not always ready to ship covers. Both checks are measured against the specification, so it comes first.

Who is speaking, and how will they be found?

A language name describes a population that no collection reaches in full. The recruitment channel decides which part of it turns up. People who answer an online task listing, a university notice, a community organisation's request or a chain of referrals can differ sharply in age, region, schooling, familiarity with the written standard and ease with a recording app. The channel shapes the dataset's character before the first file exists, and no later step changes who spoke.

The specification should name the population the collection stands for: the language variants, regions and other attributes the intended use depends on, and which of them must be balanced, which only recorded, and which do not matter. It need not prescribe a recruitment channel, but it should require the channel to be recorded for every speaker, so any skew is visible in the data rather than discovered in a model review.

Speaking how: read, prompted or spontaneous?

A contributor can read a prepared sentence, answer a prompt in their own words (asking to check an account balance, for instance), or talk freely within a task. These are three different datasets. Read speech gives control over vocabulary and coverage of specific phrases, with the rhythm of someone reading aloud. Prompted speech keeps the intent controlled while the wording stays the speaker's own. Spontaneous speech carries the hesitations, restarts, overlaps and language switches that reading rarely produces.

A brief that names none of them leaves the choice to whatever is quickest to collect and check. The specification should state the mode, or the mix and the reason for it. Whether a mix helps a given model is for testing to show; the specification makes the mix a decision rather than an accident. Read speech cannot be turned into spontaneous speech afterwards.

In what conditions: the product's, or a quiet room?

A quiet room is the right condition for some collections and, for others, a room that exists nowhere the product will be used. The question is which conditions the collection must represent: the settings people will actually speak in, the background around them, their distance from the device, whether they are moving, and the kind of device they speak into. The answer can deliberately be clean audio, as long as it is a decision and not a default.

Each session's conditions should be recorded as metadata, so the model team can separate them later. Two related choices belong here. One is whether each speaker in a conversation is captured on a separate track, which decides how much of an overlap a transcriber can later recover. The other is whether noise is recorded in the room or added afterwards: people speak differently in a noisy place, and adding noise to a clean recording does not recreate that change.

Transcribed to what convention?

A transcript is a decision about what survives. Before anyone transcribes, the convention should say what happens to hesitations and fillers, false starts and self-corrections, overlapping speech, speaker turns, switches into another language, borrowed words and the script they are written in, numbers and names, dialect forms that differ from the written standard, and stretches nobody can make out. Each can be kept, tagged, normalised or dropped. Whatever is dropped is gone for every later user of the dataset.

Every rule needs a worked example, because two transcribers given the same one-line instruction can make different and equally reasonable choices, and the difference only shows when batches are compared. This is the one decision that can be revisited after recording, but only by paying for transcription again.

With what permission, for how long, and what happens on withdrawal?

A voice recording can identify the person speaking. Where consent is the route, what may be done with the recording is bounded by what the contributor was told and agreed to at the time. The specification should state the intended use, such as evaluation, model training, a customer-facing product or onward licensing; how long recordings and derived material are kept; and what the programme does if a contributor withdraws, including what happens to transcripts already made from their recordings.

Which rules apply, and which permission route fits, are decisions for the buyer's privacy and legal owners. The specification makes sure those answers exist before the first contributor agrees to anything, because a consent cannot be widened after capture without going back to every contributor, and some will not be reachable. The consent and reviewer QA checklist covers how that permission is written and checked.

The sentence to finish before anyone presses record

Every decision above can be settled on paper before recruitment. After recording, three of them can only be changed by recording again, the permission only by going back to every contributor, and the transcript only by transcribing again. The programme owner should be able to complete one sentence without using the word hours: this collection represents ___, found through ___, speaking ___, in ___ conditions, transcribed so that ___ survives, under permission that covers ___.

A blank that cannot yet be filled is not a failure; it is the next question to settle. The same sentence is where MoniSa's collection work begins: the collection specification is agreed with the buyer alongside the source and rights brief, and a pilot set is tested against both before collection scales, so the pilot tests the specification rather than a guess at it.

Where the specification leads next

Each of these picks up a question the specification hands on.

What the collection specification should settle

Settle each of these in writing before recruitment opens.

  • The population the collection must represent, the attributes to balance, and the recruitment channel recorded for every speaker
  • The speaking mode (read, prompted or spontaneous) and the reason for any mix
  • The capture conditions to match, recorded per session as metadata
  • The transcription convention, with a worked example for every keep, tag, normalise or drop rule
  • The intended use, retention period and withdrawal procedure, confirmed by the buyer's privacy and legal owners

The request still describes audio, not a collection

Any of these means a decision is being left to whoever finds it quickest to make.

  • The requirement is stated only in hours, languages and a deadline
  • The speaker profile is a language name and a headcount
  • Nobody has said whether the speech is read, prompted or spontaneous
  • Conditions are described by what to avoid, not by where the product will be used
  • The transcription instruction is one line, or will be written after the first batch
  • Consent wording is drafted before the intended use is agreed

What to send MoniSa

A draft is enough to start a scope discussion, and an unfilled blank is useful information.

  • The sentence above, with open blanks marked as open
  • The intended use and the settings the product will be used in
  • The speaker population, language variants and attributes to balance
  • Any existing transcription convention or sample transcripts
  • Permission, retention and withdrawal requirements from your privacy and legal owners, or the questions still open with them