Why is "everything available" the wrong requirement?

The draft statement of work has a line for source history, and the easiest word to write there is "complete". Nobody is criticised for asking for more evidence. But "complete" has no edge. It does not say whether the record attaches to a batch or to each item, which handoffs it follows, how long it is kept or who will read it. A supplier can meet it with a thin spreadsheet or with session logs nobody opens.

The two ways of getting this wrong fail differently. An over-specified record wastes collection time, participant effort and storage, all spent the moment the data is gathered. An under-specified record fails later, when someone needs it, and cannot be repaired then. Source history is decided at collection: once a session has ended and batches have been merged, a record written afterwards is a reconstruction, not evidence.

So the depth has to be set before signature. Why usable data is not always ready to ship covers why source history decides whether a usable dataset can be put to work; the five questions below decide where to stop.

What will this data be used for?

Use sets the floor, and three common cases give three different answers.

Research and internal evaluation, where the data informs a decision and never reaches a customer. A batch-level record is often enough: source class, permitted use, collection period and supplier handoff for each delivered batch. It has to answer one narrow question: were you entitled to use this material for this work?

A shipped product, where a model trained or tuned on the data goes to customers. The record now has to tie each delivered batch to its source route and to a permitted use covering commercial deployment and derivative models, stated explicitly rather than implied by a contract that describes the project in general terms.

A regulated deployment, in health, finance, public services or any setting where applicable rules impose documentation duties. The floor is not set by preference here. Your legal and privacy owners decide which rules apply, and the requirement is written to meet their decision. That often means item-level traceability, permission or licence scope tied to each item, and a record of what was done to the data before delivery.

Then ask the harder version: what is the most demanding use this data could plausibly reach? Research sets have a way of becoming training sets once a result looks promising. If that move is realistic, write the requirement for the tier the data may reach, because a batch record cannot become an item record after collection. If it is not realistic, say so in writing and do not pay for a tier you will never use.

Who will ask where it came from, and when?

Almost nobody asks at delivery. The questions arrive later, under pressure, from people who were not in the room: a customer's legal team during a deal, an internal auditor, a privacy lead who has inherited the dataset, a certification or regulatory enquiry, or a contributor asking to withdraw where the permission route allows it.

Each asks at a different grain. The customer wants to know whether commercial use is covered. The auditor wants the chain from source to delivered file. A withdrawal request means finding one person's items, which a batch-level record cannot do. Name your likely questioners, write down the question each would ask, and check that the record answers it without help from anyone who worked on the collection. A record that depends on memory expires with the people who hold it.

Timing counts too. If questions can arrive for as long as a model trained on the data stays in service, the record has to stay secured and findable for that long.

What can you reconstruct today?

Test rather than assume. Take a dataset your organisation already holds, ideally one bought a year or more ago, and try to answer your questioners' questions from the records alone. Which source route did this batch come from? What use was permitted? Who handled it before it reached your team? Note every point where you had to email someone or rely on recollection.

Where the old record answered comfortably, your current detail is probably enough for that tier of use. Where it failed, the failure names the exact field to add, which beats a general instruction to require more. Run the same test on a prospective supplier before signing: ask for a pilot manifest and try to answer the same questions from it.

What does the extra detail cost?

Source history is not free, and treating it as free is how requirements drift towards "everything".

Collection time. Per-item permission explanations, metadata entry and session logging add minutes to every session. As an example, two extra minutes is nothing once and weeks of collector time across a few thousand sessions.

Participant burden. Longer forms and more personal questions reduce willingness to take part. In smaller language communities, where each contributor is harder to replace, an over-specified record can narrow the very coverage the data was bought for.

Retention and review. Every personal field kept has to be secured, access-controlled, held for a defined period, deleted on time and checked for accuracy. Detail kept without a use is exposure, not assurance.

One rule keeps the trade-off honest: add a field only when you can name who will ask for it and what they will ask. A field without a questioner is cost without a buyer.

Where is the line for this programme?

Write the line down before signature, for this programme rather than as a company-wide policy. One paragraph is enough: the intended use and the most demanding use it may reach; the unit of record, batch or item; each required field and the questioner it serves; how long the record is kept and who owns it; and the detail deliberately left out, with the reason.

That paragraph belongs in the statement of work, not an email thread. The Data rights and provenance SOW checklist turns it into a schedule with a named owner against each row. Then test it on real items before volume. MoniSa agrees the source record in a source and rights brief and checks it against a pilot manifest, the point where moving the line is still cheap. Bring the paragraph you wrote, and the pilot shows whether it answers the questions it was written for.

The decisions on either side of this one

These pages cover the questions that come before and after setting the line.

Before the line goes into the SOW

Confirm each of these before the requirement is signed.

  • The intended use is written down, with the most demanding use the data could plausibly reach
  • The unit of record is chosen, batch or item, with the reason stated
  • Every required field names who will ask for it and what they will ask
  • An existing dataset has been tested against those questions from records alone
  • The cost of each field beyond the minimum, and the retention period, are accepted by a named owner

A requirement set by habit

These suggest the depth was chosen by reflex rather than by the programme.

  • The requirement says "complete" or "full" source history with no field list behind it
  • Item-level personal detail is demanded for internal research with no stated reason
  • A product or regulated use is planned, yet the record stops at batch level
  • Personal detail kept only for traceability has no retention end date

What to send for a pilot scope

Send these and the pilot can be built to test your line, not a generic one.

  • The intended use now, and the most demanding use the data may reach later
  • The line you wrote: unit of record, required fields, the questioner each serves and the retention period
  • The data type, languages and source classes in scope
  • Any source-history question a past dataset failed to answer, with private contributor details withheld