The comparison usually prices the wrong thing
The spreadsheet looks much the same wherever this decision is made: seats and setup on one side, a per-item or per-hour quote on the other, throughput estimates beside both. It is a reasonable way to compare two purchases. It is a poor way to compare two ways of running annotation, because the expensive moments sit in neither column.
Those moments are specific. Two reviewers read "partially visible" differently, and both readings are defensible under the guideline as written. A new language arrives in the next batch and nobody on the team can judge it. A batch fails the model team's acceptance test, and no one can say whether the guideline, the labeler or the test was wrong. Each has an owner under either model. The choice between a tool and a managed service is mostly a choice about who that owner is, and whether that person exists yet.
Three decisions neither option makes for you
A tool gives the buyer direct control over task design, access, queues, assignments and reporting. A managed service adds trained people and an accountable delivery process, but only for the responsibilities the contract actually names. Three decisions come with neither.
The first is task definition: what each label means, which edge cases fall inside it, and who may change the guideline once work has started. A partner cannot infer the policy behind a difficult category. It can only apply the policy it is given, and ask when that policy runs out.
The second is adjudication: who settles conflicting labels, how the ruling gets back into the guideline, and at what level of disagreement production stops rather than continuing on a guess.
The third is acceptance: who tests a finished batch against the way the model will actually use it, and with what rule. A batch that passes the annotation team's own review can still fail downstream, and the person who decides that must sit on the buyer's side.
Write a name against each. In a tool-led model every name is internal, and so are the names behind every control the tool exposes: someone recruits, trains and supervises labelers, someone reads the disagreement reports. In a managed model, recruitment, training, first-pass review, calibration and correction can move to the provider. Label meaning and acceptance do not move. If the internal names are blank, buying more seats makes the gap easier to see without closing it. If a provider cannot say who on its side does its share, "quality included" is a slogan; the procurement questions on annotation QA turn it into questions a provider has to answer.
Which names your team can actually fill
The internal model tends to hold when the taxonomy is stable, internal reviewers can settle disagreements with authority, the team has time to recruit and supervise labelers, and the product team needs to change tasks directly or keep data inside an existing environment. Control is immediate, and so is the obligation to staff it.
It strains when language or domain coverage changes from batch to batch, when recruiting qualified labelers is itself a project, or when review needs someone who did not write the task. These are not scores to add up. They are signals about which names on the map your team can fill this quarter, and which it cannot.
Design one pilot that either model can fail
Most tool-versus-service pilots are not comparisons. Each arm gets a slightly different sample, a different version of the guideline and a different reviewer, and the result reflects those differences more than the operating models. A fair pilot holds everything constant except the model.
Give both arms the same guideline version, the same language mix and the same representative sample, including the hard cases rather than only the clean ones. Before the pilot starts, run the shared guideline through the annotation guideline QA checklist so neither arm inherits a contradiction the other resolves by luck. Hold back a set of items with agreed answers that neither arm sees, and have the same buyer-side owner score both arms against the same acceptance rule.
Treat clarifications as part of the experiment. If one arm's question produces a ruling on the second day, for example, send the same ruling to the other arm that day and log it. The log is evidence: it shows how many decisions each model pushed back to your team, and how quickly they were closed.
Count your own hours in both arms. Guideline clarification, supervision, recruitment and rework are real costs whether they land in the internal arm or in answering a provider's questions. Compare accepted items, exception handling and total cost including that internal time, not items processed.
Keep the answer reversible
The pilot does not have to crown a winner. The result is often a split: internal task design with managed execution, or internal execution with specialist review for the languages or domains the team cannot staff. Whatever the split, keep the guideline and its decision log portable, confirm that labeled data and reviewer feedback return to your environment in your format, and write the division of ownership down before scale makes it expensive to change.
What the managed arm should show you
A managed arm is only useful in a comparison if its ownership is visible. MoniSa runs managed annotation from a written guideline and a gold set. Reviewers work independently, agreement is checked on a pilot batch, and ambiguous cases are escalated and folded back into the guideline rather than absorbed quietly by the labeler who met them. Throughput rises only after agreement holds, and the buyer approves the pilot before production volume opens. For multilingual work, native-speaker reviewer fit is confirmed for the specific language before a batch is scaled.
Those escalations and guideline versions are what a fair comparison needs from the managed side: which questions reached your team, how the guideline changed in response, and which decisions stayed yours throughout.
The rest of the annotation sourcing decision
These pages take one part of the decision further once the ownership map is drafted.
- Annotation guideline QA checklist: Test the guideline both arms will share before the pilot starts.
- AI annotation QA procurement questions: Calibration, agreement and adjudication questions for any managed provider.
- Pilot-to-production ramp planner: Record what made the pilot work before volume rises.
- Data annotation services: How a managed annotation workstream runs from guideline and gold set to an approved pilot.
The ownership map to complete before comparing quotes
Put a name against each line before either option is compared on price.
- Who writes the guideline, and who may change it once work has started
- Who settles conflicting labels, and the level of disagreement that pauses production
- Who scores a finished batch against downstream use, and with which acceptance rule
- Who recruits, trains and calibrates labelers for each language or domain in scope
- Where the labeled data, guideline versions and decision log live after the pilot ends
Signs the comparison will not hold
Any one of these means the pilot result describes the setup rather than the operating model.
- "Quality included" with no named sample, error classes or correction owner
- Each arm working from a different sample, guideline version or acceptance rule
- A clarification sent to one arm and not the other
- Internal supervision, clarification and rework hours left out of the cost
- The result reported as items processed rather than items accepted
- No one on the buyer side with the authority to change the rubric
What to send MoniSa to scope the managed arm
Send the same inputs the internal arm will use, so the two arms start level.
- The current annotation guideline and label taxonomy, with its version
- A representative sample that includes the hard cases, plus any gold examples with agreed answers
- The languages, scripts and domains in scope
- The acceptance rule and the name of the person who will apply it
- How the internal arm will run, so the managed arm can mirror it