FAQ
Frequently asked questions
The questions we are actually asked, answered at the length they deserve rather than in a sentence. Grouped so you can find yours without reading the rest.
If something here is not covered, send it with your request and we will answer it directly rather than pointing you back at this page.
01
Getting started
What a first engagement looks like.
With a scoped pilot. You send a representative data sample and a description of what the model should do; we draft the guideline decisions the task will require, annotate a small batch, and report measured agreement on it. The pilot exists to find out where two reasonable annotators disagree, because that is where your specification is incomplete. Production sizing and pricing follow from what the pilot measures rather than from an estimate.
A data sample, the task description, the output format you need, and any constraints — deadline, confidentiality requirements, tooling you want us to work inside. A quote produced from a task description alone is a guess; we would rather look at real data for an hour than produce a number that changes after the first batch.
Typically one to three weeks from receiving data, depending on how much specification work the task needs. Most of that time is guideline drafting and calibration, not annotation. A task with an existing, tested guideline moves considerably faster.
No fixed minimum. We run small pilots and scale to high-volume production programs. What sets a practical floor is setup cost — guideline work, tooling and calibration are largely fixed regardless of volume, so a very small one-off job can carry a disproportionate share of it. We will tell you when that is the case rather than quoting around it.
Per unit of work — per item, per audio hour, per image, per evaluation — with the rate driven by task complexity, review depth and the level of expertise required. Guideline development and calibration are scoped separately because they are real work that happens before any item is labelled. We do not charge for pilot scoping conversations.
02
Services and scope
What we do, and where we are not the right fit.
Japanese is what we are built for. We handle Japanese–English bilingual work — translation review, cross-lingual alignment, parallel evaluation sets — because that draws on the same Japanese judgment. For volume work in other languages we would rather tell you plainly that we are not the right vendor than take the project and subcontract it.
Yes. We work in client-provided platforms and virtual desktops, including where policy requires that data never leaves your infrastructure, and in our own secure environment when you would rather not provision accounts. The tooling changes; the guideline, the review passes and the QA reporting do not.
We collect and create data where it is Japanese-language work — writing instruction data, constructing evaluation prompts, recording scripted speech, sourcing document samples. We are not a general-purpose data collection operation, and for large-scale field collection you would be better served by a specialist.
Yes, and the first thing we do is read their guideline and re-annotate a sample of their delivered data blind. That tells you what you actually have — including whether the existing labels are internally consistent — before we add to it. Continuing a dataset without measuring what it currently contains is how inconsistency gets compounded.
Yes. A dedicated team arrangement is common for long-running AI programs, where the same annotators carry the guideline forward and their accumulated context is a real asset. It is also the arrangement where guideline drift matters most, so it comes with periodic re-calibration against a gold set.
03
Japanese language
What Japanese coverage means in practice.
Formal and informal registers across kanji, hiragana, katakana and Latin script, including honorific language, regional variation, slang, industry jargon and mixed Japanese–English technical writing. Normalization rules for script choice, numerals, punctuation, emoji and full-width characters are agreed at kickoff, because those choices determine whether the corpus is internally consistent.
Yes. This matters for judgments that depend on how a sentence lands rather than on what it denotes — whether a refusal reads as a refusal, whether a politeness level is appropriate to the relationship, whether a phrasing sounds translated. For mechanical tasks with a tight guideline, that requirement is less critical, and we staff accordingly rather than charging for expertise a task does not need.
Yes, with the caveat that dialect coverage is a staffing question and we will be specific about which varieties we can cover reliably. The more important decision is upstream: whether dialect is transcribed as spoken or normalised to standard Japanese. That choice changes the dataset entirely and needs to be made deliberately.
A written normalization convention, applied at annotation time and enforced by automated checks before delivery. It covers script choice for frequent words, long-vowel spelling in katakana, Arabic versus kanji numerals, full-width and half-width characters, spacing around Latin text and punctuation. Orthographic variation is the most common defect in unspecified Japanese datasets and the easiest to prevent.
Yes, and we would rather adopt yours than impose ours — a model that already writes in your house style should keep doing so. What we will do is test it against real edge cases during the pilot, because most style guides have two or three rules that turn out to conflict once they meet actual data.
04
Quality and delivery
How quality is measured and what arrives.
Inter-annotator agreement with a chance-corrected coefficient, gold-set accuracy, adjudication rate, and error counts broken down by a named taxonomy rather than as a single defect total. Everything is reported per label or criterion rather than blended, because a blended figure hides exactly the category that is failing. The quality page sets out how each number is calculated and what it does not capture.
JSONL, CSV, CoNLL-U, COCO, SRT, CTM, TextGrid, hOCR and custom schemas defined at kickoff. Span data always includes character offsets underneath whatever tag format you take, so a later change of tokenizer is a re-export rather than a re-annotation. Every batch ships with its guideline version and a change log against the previous one.
Revisions go through the same review and adjudication path as first-pass work — a fast correction lane is how datasets drift between batches. Where the cause is a guideline gap we revise the guideline and re-work the affected items, not just the ones you flagged, because the same gap will have produced the same error elsewhere.
Throughput depends on task complexity, review depth and how much genuine ambiguity the data contains, and it varies more between tasks than most estimates admit. We measure real throughput during the pilot on your data and commit to a plan from that. A delivery date quoted before anyone has annotated your data is a guess presented as a commitment.
We look at the specific items with you. Usually one of three things is true: the guideline was ambiguous, in which case we revise it and re-work; the guideline was clear and misapplied, in which case we correct and feed it back into review; or the guideline says something you no longer want it to say, in which case the change is a scope decision. Sorting a disagreement into one of those three is more productive than arguing about the individual label.
05
Security, legal and commercial
Confidentiality, contracts and how we work with procurement.
An NDA before anything is shared, a data processing agreement where personal information is involved, role-based access limited to the people staffed on the project, an isolated working environment, and retention and deletion terms written into the contract rather than left to policy. The security page covers each control and, equally importantly, what we do not claim.
We do not claim any certification that is not written into our contract with you. Ask about a specific scheme and you will get a direct answer, including when that answer is no. We would rather lose a procurement process for being accurate than pass it on a claim that will not survive an audit.
You do. Deliverables and the derived annotations are yours under the engagement contract. We do not retain client data for internal reuse, for model training or as portfolio material, and any sample material shown on this website is our own, created for illustration.
Yes, that is normal for enterprise engagements. Where a clause commits us to something we cannot actually deliver — a specific accuracy figure, a certification we do not hold — we will say so and propose alternative wording rather than signing and hoping. It is a slower conversation and a shorter one than the alternative.
Where subcontracting is involved it is disclosed, and it is covered by the same confidentiality obligations that bind our own staff. It is never assumed silently. If your policy prohibits subcontracting entirely, tell us and we will confirm whether we can staff the project on that basis before you commit to it.
06
Expert annotation
Domain specialists, and the limits of what we claim.
For projects that need it, we build teams around the required domain expertise, which can include people with relevant qualifications or practical experience in fields such as law, tax, accounting, labor and social security, or corporate registration. We do not maintain a standing roster of licensed professionals. Availability is assessed per project against its scope, volume, specialisation and timeline, and if a qualification cannot be sourced for your project we will tell you before it starts.
By what kind of judgment the task requires. Terminology and document conventions can be transferred to a trained annotator with a good lexicon and a calibration round, and for most work that is the better answer — faster, more consistent and less expensive. A specialist is warranted when the judgment itself is professional: whether an interpretation is defensible, whether an answer would mislead a reader about their obligations.
Yes. Domain specialists can score generated responses for factual accuracy, reasoning quality, professional appropriateness and compliance with your criteria, and can produce preference rankings for alignment work, benchmark sets, retrieval evaluation and red teaming. The reasoning behind each judgment is recorded, which is usually what makes an expert evaluation worth its cost.
No. They are annotating and evaluating data. The output is training or evaluation data for your AI system, not legal, tax or accounting advice to you or to your users, and nothing produced in an annotation project should be presented to an end user as professional advice. We are explicit about this because the distinction matters to both of us.
Still have a question?
Ask it directly.
Send it with your request and you will get a specific answer from someone who has run the kind of project you are describing — not a link back to this page.
We reply within one business day.