Guideline gap
The case is not covered by the guideline. Fixed by revising the guideline, not by correcting the item — the same case will otherwise recur.
Quality framework
Quality is a number you can act on or it is an opinion. This page explains exactly which numbers we produce, how they are calculated, what they can and cannot tell you, and what arrives in the QA report with every batch.
Every figure shown on this site as an example is labelled as an example. We do not publish accuracy guarantees, because a guarantee made before seeing your data is not a commitment — it is a marketing number.
A final inspection pass can catch mistakes. It cannot catch ambiguity, and ambiguity is what actually degrades a Japanese dataset. If the guideline permits two readings of a case, both readings pass inspection, and the inconsistency arrives in the training data with a clean bill of health.
So the work that determines quality happens before annotation: writing the guideline, breaking it on real edge cases during a pilot, measuring whether independent annotators converge, and fixing the guideline where they do not. Reviewing is the second line of defence, not the first.
What follows is the whole system, including the parts that are uncomfortable to publish — what each metric fails to capture, and where a number can look healthy while the dataset is not.
At a glance
The system
This loop runs for the length of a project, not once at the start.
The guideline is written with worked examples and counter-examples for every decision the task requires. Anything left implicit here becomes inconsistency later.
Annotators independently label the same calibration set. Divergence identifies where the guideline is ambiguous — the guideline is fixed, not the annotator.
Production work, with the guideline version recorded per item so a later question about a record can be answered precisely.
A second annotator reviews at the depth the task warrants. Review is a separate role with its own criteria, not a faster repeat of annotation.
Disagreements go to a senior reviewer, who records the resolution and the reason. Every resolution becomes a worked example in the guideline.
Agreement, gold-set accuracy and drift are computed per batch, compared against previous batches, and used to decide where the next guideline revision is needed.
Metrics
Defined here so that a figure in a QA report is interpretable without asking us what it refers to.
Agreement
97.8%
2.1% vs last 7 days
Gold-set accuracy
98.6%
1.4% vs last 7 days
Items reviewed
24826
18.7% this week
Open disagreements
18
12 vs last 7 days
Illustrative interface. The figures shown are sample data used to explain how we monitor quality, not reported results.
Review
Review depth is a project setting, agreed with you and stated in the contract. Heavier is not always better — a mechanical task reviewed at full depth spends budget where no judgment is required.
| Review depth | Typically used for | What happens |
|---|---|---|
| Sampled review | High-volume mechanical tasks with stable agreement | A defined share of items is reviewed; failures trigger a wider pass |
| Full second pass | Judgment-dependent labeling and most Japanese span work | Every item is reviewed by a second annotator |
| Blind double annotation | Evaluation sets, benchmarks and gold sets | Two annotators work independently; all disagreements are adjudicated |
| Expert review | Regulated domains and professional judgment | A domain specialist reviews, with written reasoning on contested items |
Error taxonomy
A defect that has a name can be counted, trended and fixed. A generic "error" count cannot tell you what to change.
The case is not covered by the guideline. Fixed by revising the guideline, not by correcting the item — the same case will otherwise recur.
The case is covered and the annotator applied the rule incorrectly. Fixed by feedback and, if it recurs, by a clearer worked example.
The label is right and the span is not. Specific to Japanese span work, where boundary rules and tokenization interact.
Linguistically correct, socially wrong: the wrong politeness level, or a refusal too soft to read as one.
A verifiable claim, name reading or figure is wrong. Distinguished from judgment errors because the fix is verification, not calibration.
The item itself is unusable — inaudible audio, an illegible scan, an ambiguous prompt. Flagged and returned, never filled in with a plausible guess.
Reporting
The report is designed so that you could audit the batch without us in the room.
FAQ
Questions about what these numbers do and do not prove.
We do not publish an accuracy guarantee, and we would treat one from any vendor with caution. Achievable accuracy depends on the task, the taxonomy, the quality of the source data and how much genuine ambiguity the domain contains — none of which are known before a pilot. What we commit to is the measurement method: which metrics, computed how, reported at what frequency, with what happens when a batch falls short. After a pilot on your data we can talk about realistic targets, because by then there is something real to talk about.
They are sample data. They exist to show what the interface reports and what the metrics look like in motion, and they sit behind a visible badge and disclaimer for exactly that reason. They are not results from a client project, and we will not present them as such.
Not on its own. Agreement measures consistency, not correctness — annotators who share the same misunderstanding agree perfectly. It is also inflated by imbalanced taxonomies, which is why we report a chance-corrected coefficient alongside raw agreement, and why gold sets adjudicated by senior reviewers exist as an independent check on correctness.
It does not ship. Depending on the failure the response is targeted re-annotation, a guideline revision followed by re-work of affected items, or escalation to you when the cause is upstream — an ambiguous taxonomy or source data that cannot support the task. We would rather deliver late with an explanation than on time with a number that will not survive contact with your model.
Yes, and it is a reasonable thing to ask for. You can review the guideline and its version history, take a sample for independent scoring, and compare your findings against our reported figures. Where a client runs their own QA in parallel we treat divergence between the two as a finding worth investigating rather than a dispute to be argued.
Quality
A pilot produces real agreement figures, a real error breakdown and an honest view of which parts of your taxonomy will hold up at volume.
NDA before you share any data. Pilot scoping is free.