Services

Japanese AI training data, across every modality

Five services covering what an AI system needs at each stage — from the instruction data that teaches a model to the human evaluation that tells you whether it worked.

All five run on the same foundation: Japan-based annotators, guidelines that are versioned rather than remembered, and quality that is measured before it is claimed.

One team, the whole data lifecycle

Most annotation problems are not labeling problems. They are definition problems: two reasonable people read the same guideline and label the same Japanese sentence differently, and nobody notices until the model is trained on both. Splitting our work into five services is a way of making those definitions explicit — each service has its own taxonomy questions, its own failure modes and its own quality checks.

The services are not packages you have to pick between. A single project routinely uses three of them: NLP annotation to build a labelled corpus, LLM training data to turn it into instruction and preference pairs, and evaluation to measure what the resulting model actually does. What stays constant is the Japanese-language judgment underneath all of it.

At a glance

Modalities
Text, speech, image, video, document and multimodal Japanese data.
Model stages
Corpus filtering, supervised fine-tuning, preference and RLHF data, evaluation and benchmarking.
Delivery formats
JSONL, CSV, CoNLL-U, SRT, CTM, COCO — or a schema you define.
Tooling
We work inside your annotation platform, or in our own secure environment.
Starting point
A scoped pilot. Taxonomy and quality are proven before volume ramps.

What we do

Five services, plus the expert layer

Each has its own page with the taxonomy decisions, the Japanese-specific pitfalls and the deliverables spelled out.

Choosing

Which service do I need?

Start from what you are trying to change about the model. The service follows from that, not from the file type you happen to have.

Matching a goal to a service. Most programs end up combining two or three.
If your goal is…The service isTypical first deliverable
Teach the model to answer in natural, correctly-registered JapaneseLLM Training DataA pilot set of instruction–response pairs with a locked style guide
Find out where the model is wrong before you ship itLLM EvaluationA scored evaluation set with a rubric and inter-rater agreement
Make the model hear Japanese speech accuratelySpeech & AudioTranscribed audio with a written normalization convention
Read Japanese documents, forms or on-screen textImage & VideoOCR and layout annotation on a representative document sample
Extract structured facts from Japanese textText & NLPAn entity and relation schema tested against real edge cases
Have specialists judge answers in a regulated fieldExpert AnnotationAn expert-reviewed evaluation set with written adjudication notes

Lifecycle

Where each service fits

The same dataset rarely serves two stages. Data built to teach a model is a poor test of it, because the model has already seen it.

  1. 01

    Corpus preparation

    Filtering, deduplication rules, quality and safety flags on raw Japanese text, speech or documents. Decides what is worth annotating at all.

  2. 02

    Supervised fine-tuning

    Instruction–response pairs, multi-turn dialogue and task demonstrations written in the register the product actually uses.

  3. 03

    Preference and alignment

    Ranked or pairwise comparisons, safety and refusal behaviour, and the written criteria that make one response better than another.

  4. 04

    Evaluation

    Held-out sets and rubric scoring by people who were not involved in building the training data, so the measurement is independent.

  5. 05

    Production monitoring

    Sampling live model output, scoring it against the same rubric, and feeding recurring failures back into the next data cycle.

Delivery

Formats and delivery

Every batch ships with the taxonomy version it was labelled against and a change log against the previous batch, so datasets stay comparable across versions.

Common delivery formats. Custom schemas are defined at kickoff and versioned with the data.
Data typeTypical deliveryAlso available
Text and NLPJSONLCSV, CoNLL-U, brat standoff, Parquet
LLM instruction and preferenceJSONL (messages / chosen–rejected)CSV, Hugging Face dataset layout
SpeechJSON with timingsSRT, VTT, CTM, TextGrid, plain transcript
Image and videoCOCO JSONYOLO, Pascal VOC, CVAT XML, per-frame JSONL
Document and OCRJSON with page geometryhOCR, ALTO, CSV field extraction
Evaluation resultsJSONL scores plus a QA reportCSV, scoring notebook, adjudication log

FAQ

About the services

Scope and process questions that come up before a pilot is defined.

Yes, and most programs do. A typical sequence is NLP annotation to establish a labelled corpus, LLM training data built from it, then evaluation run by a separate group of annotators. Combining them under one engagement means the taxonomy, the style guide and the QA metrics stay consistent instead of being re-derived per vendor.

Japanese is what we are built for, and Japanese–English bilingual work sits inside that. For volume work in other languages we would rather tell you plainly that we are not the right vendor than take the project and subcontract it.

Pricing is per unit of work — per item, per audio hour, per image or per evaluation — and depends on task complexity, how much reviewer time each item needs and the level of expertise required. We quote after scoping a pilot on your real data, because a quote based on a task description is a guess and usually a wrong one.

We draft them, you review them, and both happen before production. The pilot exists to break the first draft on real edge cases; every disagreement resolved during it is written back into the document. The guideline is then versioned, and each delivered batch records which version it was labelled against.

Start here

Tell us what the model needs to learn.

Send the task, a data sample and the constraints you are working under. We come back with a scoped pilot plan, the taxonomy questions we would need answered, and an honest view of what is hard about it.

NDA before you share any data. Pilot scoping is free.