Quality framework

How we measure annotation quality

Quality is a number you can act on or it is an opinion. This page explains exactly which numbers we produce, how they are calculated, what they can and cannot tell you, and what arrives in the QA report with every batch.

Every figure shown on this site as an example is labelled as an example. We do not publish accuracy guarantees, because a guarantee made before seeing your data is not a commitment — it is a marketing number.

Quality is designed in, not inspected in

A final inspection pass can catch mistakes. It cannot catch ambiguity, and ambiguity is what actually degrades a Japanese dataset. If the guideline permits two readings of a case, both readings pass inspection, and the inconsistency arrives in the training data with a clean bill of health.

So the work that determines quality happens before annotation: writing the guideline, breaking it on real edge cases during a pilot, measuring whether independent annotators converge, and fixing the guideline where they do not. Reviewing is the second line of defence, not the first.

What follows is the whole system, including the parts that are uncomfortable to publish — what each metric fails to capture, and where a number can look healthy while the dataset is not.

At a glance

Measured on
Inter-annotator agreement, measured gold-set accuracy, adjudication rate and label drift.
Reported per
Batch, with a per-label and per-criterion breakdown rather than one blended figure.
Review depth
Set per project. Deeper review where judgment is required, lighter where the task is mechanical.
Disagreements
Adjudicated in writing by a senior reviewer and written back into the guideline.
What we do not do
Publish accuracy guarantees, or quote a figure measured on a different project.

The system

Six stages, running continuously

This loop runs for the length of a project, not once at the start.

  1. 01

    Specify

    The guideline is written with worked examples and counter-examples for every decision the task requires. Anything left implicit here becomes inconsistency later.

  2. 02

    Calibrate

    Annotators independently label the same calibration set. Divergence identifies where the guideline is ambiguous — the guideline is fixed, not the annotator.

  3. 03

    Annotate

    Production work, with the guideline version recorded per item so a later question about a record can be answered precisely.

  4. 04

    Review

    A second annotator reviews at the depth the task warrants. Review is a separate role with its own criteria, not a faster repeat of annotation.

  5. 05

    Adjudicate

    Disagreements go to a senior reviewer, who records the resolution and the reason. Every resolution becomes a worked example in the guideline.

  6. 06

    Measure and feed back

    Agreement, gold-set accuracy and drift are computed per batch, compared against previous batches, and used to decide where the next guideline revision is needed.

Metrics

What each number actually means

Defined here so that a figure in a QA report is interpretable without asking us what it refers to.

Inter-annotator agreement
How often independent annotators produce the same label on the same item. Reported as raw agreement together with a chance-corrected coefficient, because raw agreement on an imbalanced taxonomy can look excellent while carrying almost no information.
Cohen’s κ / Krippendorff’s α
Chance-corrected agreement coefficients. Cohen’s κ for two annotators on categorical labels; Krippendorff’s α when there are more than two annotators, missing judgments or ordinal scales. Which one is used is stated in the report rather than left implicit.
Gold-set accuracy
Accuracy against a reference set annotated and adjudicated by senior reviewers. It measures correctness rather than consistency — two annotators can agree perfectly and both be wrong, which agreement statistics alone will never reveal.
Adjudication rate
The proportion of items that required a senior reviewer to resolve. A rising adjudication rate is usually the earliest signal that the data has shifted or the guideline has a gap, and it moves before accuracy does.
Label drift
Movement in the label distribution or in annotator behaviour over time. Long projects drift for ordinary reasons — the source data changes, annotators become more confident — and drift monitoring is what separates that from genuine degradation.
Boundary agreement
For span tasks, how closely annotators agree on where a span starts and ends, reported separately from whether they agree on its label. A Japanese span dataset can have excellent label agreement and unusable boundaries.

Quality overview

Sample QA Dashboard

Agreement

97.8%

2.1% vs last 7 days

Gold-set accuracy

98.6%

1.4% vs last 7 days

Items reviewed

24826

18.7% this week

Open disagreements

18

12 vs last 7 days

Label drift (30 days)

Observed drift Alert threshold
0%2.5%5% Day 1Day 8Day 15Day 22Day 30

Quality summary

  • Guideline adherence
  • Reviewer calibration
  • Outlier monitoring
  • Drift detection

Illustrative interface. The figures shown are sample data used to explain how we monitor quality, not reported results.

Review

How much of a batch gets reviewed

Review depth is a project setting, agreed with you and stated in the contract. Heavier is not always better — a mechanical task reviewed at full depth spends budget where no judgment is required.

Illustrative review configurations. The actual depth for a project is agreed during scoping.
Review depthTypically used forWhat happens
Sampled reviewHigh-volume mechanical tasks with stable agreementA defined share of items is reviewed; failures trigger a wider pass
Full second passJudgment-dependent labeling and most Japanese span workEvery item is reviewed by a second annotator
Blind double annotationEvaluation sets, benchmarks and gold setsTwo annotators work independently; all disagreements are adjudicated
Expert reviewRegulated domains and professional judgmentA domain specialist reviews, with written reasoning on contested items

Error taxonomy

Naming the failures

A defect that has a name can be counted, trended and fixed. A generic "error" count cannot tell you what to change.

Guideline gap

The case is not covered by the guideline. Fixed by revising the guideline, not by correcting the item — the same case will otherwise recur.

Guideline misapplication

The case is covered and the annotator applied the rule incorrectly. Fixed by feedback and, if it recurs, by a clearer worked example.

Boundary error

The label is right and the span is not. Specific to Japanese span work, where boundary rules and tokenization interact.

Register or nuance error

Linguistically correct, socially wrong: the wrong politeness level, or a refusal too soft to read as one.

Factual error

A verifiable claim, name reading or figure is wrong. Distinguished from judgment errors because the fix is verification, not calibration.

Source defect

The item itself is unusable — inaudible audio, an illegible scan, an ambiguous prompt. Flagged and returned, never filled in with a plausible guess.

Reporting

What arrives with every batch

The report is designed so that you could audit the batch without us in the room.

  • Agreement figures per label or criterion, with the coefficient used stated explicitly.
  • Gold-set accuracy for the batch, and the trend against previous batches.
  • Adjudication rate, with the adjudication log and the reasoning recorded for each resolution.
  • Error counts broken down by the taxonomy above, rather than as a single defect total.
  • The guideline version the batch was annotated against, and a change log against the previous version.
  • Items flagged as source defects, listed rather than silently dropped.

FAQ

About quality measurement

Questions about what these numbers do and do not prove.

We do not publish an accuracy guarantee, and we would treat one from any vendor with caution. Achievable accuracy depends on the task, the taxonomy, the quality of the source data and how much genuine ambiguity the domain contains — none of which are known before a pilot. What we commit to is the measurement method: which metrics, computed how, reported at what frequency, with what happens when a batch falls short. After a pilot on your data we can talk about realistic targets, because by then there is something real to talk about.

They are sample data. They exist to show what the interface reports and what the metrics look like in motion, and they sit behind a visible badge and disclaimer for exactly that reason. They are not results from a client project, and we will not present them as such.

Not on its own. Agreement measures consistency, not correctness — annotators who share the same misunderstanding agree perfectly. It is also inflated by imbalanced taxonomies, which is why we report a chance-corrected coefficient alongside raw agreement, and why gold sets adjudicated by senior reviewers exist as an independent check on correctness.

It does not ship. Depending on the failure the response is targeted re-annotation, a guideline revision followed by re-work of affected items, or escalation to you when the cause is upstream — an ambiguous taxonomy or source data that cannot support the task. We would rather deliver late with an explanation than on time with a number that will not survive contact with your model.

Yes, and it is a reasonable thing to ask for. You can review the guideline and its version history, take a sample for independent scoring, and compare your findings against our reported figures. Where a client runs their own QA in parallel we treat divergence between the two as a finding worth investigating rather than a dispute to be argued.

Quality

Ask us for the numbers on your own data.

A pilot produces real agreement figures, a real error breakdown and an honest view of which parts of your taxonomy will hold up at volume.

NDA before you share any data. Pilot scoping is free.