What does the dashboard count today?

Most review dashboards begin with whatever the tooling records without being asked: items reviewed per hour, average time per item, backlog cleared and the share of items passed. Each is a timestamp or a status change. None required anyone to decide what good review looks like, which is why they arrive first and stay longest.

The trouble starts when a finance lead asks what the review layer is worth. A layer defended on items per hour has proved only that it is fast, and a fast layer invites the obvious follow-up: could it sample less, skip the easy task types, or go altogether? Throughput is an argument for a cheaper review layer, not for having one.

What does counting speed reward?

It rewards the reviewer who moves quickly through easy material and quietly penalises the one who stops. Take, as an example, a preference pair in which both responses are fluent and one contains a subtle factual error. The reviewer who catches it has to stop and check the claim, and on a volume view looks slower than a colleague who accepted ten clean pairs in the same time, although only one of them did the work the layer is paid for.

Speed metrics also put a price on disagreement. Accepting is always quicker than rejecting, because a rejection needs a reason, sometimes an escalation, often a second look. Every flag costs the reviewer time; every acceptance is free. The pass rate drifts upward and the dashboard reads the rise as improving quality. It cannot tell whether producers got better or reviewers stopped looking.

What is the review layer for?

Strip the proxies away and a review layer exists for two things. The first is errors caught before they reach the client or the training run. The second is disagreement surfaced rather than smoothed over: a reviewer who doubts a label, a guideline two people read differently, a producer whose work has drifted. A doubt settled silently, by one reviewer overruling their own hesitation, is a lesson the programme never receives.

It also decides what a sign-off is worth. Why usable data is not always ready to ship separates two sign-offs: whether a dataset meets its technical criteria, and whether it is cleared for the intended use. The first rests heavily on the review layer, and if that layer is measured in a way that cannot tell looking from not looking, "passed review" tells the person signing off very little.

Can it be measured without becoming a target?

This is the hard part, and it has no clean answer. Any number used to judge people will be optimised by people, and each better signal has its own way of going wrong.

Count errors flagged, and reviewers learn to flag borderline items: flags become cheap and adjudication fills up. Measure catch rate on seeded errors, known mistakes placed in a batch before review, and the risk moves to the seeds. If seeds are templated, obviously planted or drawn from the same few error types, reviewers learn to find seeds rather than errors, and the seeded catch rate climbs while real errors still pass. Seeds also test only the error types someone thought to seed, so a strong result is a floor on vigilance for known problems, not a verdict on quality.

Escalation rate can be gamed in both directions. Where escalating counts against a reviewer, escalations stop. Where it moves responsibility upward, everything borderline gets escalated. Agreement between reviewers fails more quietly: reviewers who can see each other's decisions, or who have learned which answer is safe, converge without judging.

A few habits keep these signals honest. Pair each with a counterweight, so caution in one direction shows up in another. Report at the level of the layer or task type, and keep individual figures for calibration conversations, never rankings. Replace targets with expected ranges and a rule for investigating outside them. Keep the mechanics hidden at item level while the policy stays open: reviewers can know that seeding and double review exist, and why, without knowing which items are involved. And treat a rate computed on a handful of seeds as an anecdote.

What would a fairer set of signals look like?

A workable starting set, offered as an example rather than a standard, has three parts: catch rate on seeded errors, split by error type and read beside false flags on seeded items known to be correct; escalation rate, read with escalation yield, meaning the share of escalations that changed a decision or clarified the guideline; and disagreement between two independent reviewers on a blind double-review sample, reported by item difficulty. None is reliable alone. Together they show whether the layer catches what it should, raises what it cannot settle, and still disagrees where the material is hard.

What to seed depends on what is expensive when it escapes. In preference data, a useful seed might be a pair where the more fluent response is factually wrong, or where the preferred response breaks the written policy.

The same set carries the cost case that throughput cannot. To take an illustrative example, "the layer caught most seeded errors of the type that costs most downstream, and its escalations changed the guideline twice this quarter" argues for the layer's existence, not for a cheaper version of it.

Blind double review is the hardest signal to run from inside, because a second reviewer usually shares the first one's training, tools and incentives. Independence is easier to keep when the second pass sits outside the team being measured. MoniSa organises review around that separation: production and review stay apart, reviewers are calibrated on shared examples before volume, and disputed items are adjudicated rather than averaged away.

Four easy metrics and what to put beside each

Items reviewed per hour rewards speed on easy material. Beside it: seeded catch rate by error type. Gaming risk: seeds that look planted, so build them from real errors found in earlier batches and refresh them.

Average time per item rewards skimming and penalises checking. Beside it: blind double-review disagreement by item difficulty. Gaming risk: convergence when reviewers can see each other's decisions, so keep the second pass blind.

Pass rate rewards agreeing with the producer. Beside it: flag precision, the share of flags upheld at adjudication, with false flags on seeded correct items. Gaming risk: reviewers stop flagging borderline items to protect precision, which the seeded catch rate will expose.

Backlog cleared rewards closing items, including by settling doubts silently. Beside it: escalation rate with escalation yield. Gaming risk: defensive escalation, or none at all, so never attach a target to the rate itself.

A review layer that never disagrees

Run one uncomfortable check before defending the layer's cost. If the blind sample shows near-total agreement on hard items, every seed is caught without a false flag, and escalations never change a decision, suspect measures that cannot fail before crediting an exceptional team. A review layer that never disagrees with itself is not being measured. It is being ignored.

Calibration, disagreement and QA questions

These live pages go deeper on the parts of review measurement that sit next to this decision.

Redesigning a review dashboard

Adapt these to the task and its costliest errors rather than adopting them as a fixed set.

  • Name the costliest downstream error types and seed each from real past errors
  • Seed known-correct items too, so false flags sit beside catches
  • Run a blind double-review sample and report disagreement by item difficulty
  • Report escalation rate with escalation yield
  • Report at layer or task level; keep individual figures for calibration
  • Swap any target for an expected range and an investigation rule
  • Refresh seeds so reviewers find errors, not seeds

Signs the measures are measuring themselves

Each of these suggests the dashboard describes the metric rather than the review.

  • Throughput is the only evidence offered for the layer's cost
  • Pass rate climbs and nobody can say why
  • Seeded catch rate is near perfect while real errors still reach the client
  • The second reviewer on a double-review sample can see the first decision
  • Escalations never change a decision or a guideline
  • Any of these signals is turned into a reviewer ranking

What to send MoniSa to scope an independent review sample

Sample-safe material is enough to start. The useful output of the discussion is an agreed set of signals, each with its counterweight.

  • The task type, languages and guideline version the review works against
  • The metrics currently on the dashboard, and who uses them to judge cost
  • The error types that cost most when they escape review
  • Any known-error or known-correct items already available for seeding
  • The current escalation route and who owns adjudication