Data labeling
Quality process in data labeling
Consistency in a dataset doesn't come from individual effort, it comes from process. This page describes ours, stage by stage, including what we do when annotators disagree.
Send a small batch. We'll return proposed guidelines and an effort estimate.
The stages of the process
The same path on every project, from the first batch to versioned delivery.
-
Building the guidelines
We write the definition of each class, with positive examples, negative examples and what to do in case of doubt. A guideline without edge cases isn't a guideline.
-
First sample
A small batch is annotated first, deliberately. It exists to surface what the guideline didn't anticipate while fixing it is still cheap.
-
Criteria validation
You review the sample and confirm or correct the criteria. Volume only starts after that sign-off.
-
Annotation
Production against the approved guideline, by annotators trained on the specific criteria of the project.
-
Review
A second pass over the material, checking adherence to the guideline and not merely the absence of obvious errors.
-
Sampling audit
A slice of each batch is independently re-evaluated, to measure consistency rather than assume it.
-
Handling disagreement
Disagreement between annotators is recorded and resolved in the guideline, not case by case. Recurring disagreement points to an ambiguous criterion, not a weak annotator.
-
Export
Conversion to your pipeline's format, with an integrity check on the generated files.
-
Versioning
Every delivery is identified together with the guideline version used. Without that, there's no way to explain why two batches differ.
How we measure consistency
Quality that isn't measured is opinion.
-
Inter-annotator agreement
The same material is annotated by more than one person on a sample, and the divergence is measured rather than estimated.
-
Adherence to the guideline
Review checks whether the annotation follows the written criterion, including when the result “looks right”.
-
Divergence by class
When one class concentrates disagreement, the problem is its definition. That's where the guideline gets rewritten.
-
Stability across batches
Comparison over time, to catch criterion drift before it reaches the model.
Common mistakes the process prevents
Each stage above exists because one of these problems is expensive to fix later.
-
A class defined by one example
A class explained with a single image produces different readings as soon as the material varies.
-
Grey areas without a rule
If the guideline doesn't say what to do with the doubtful case, each annotator decides alone — and the model learns the inconsistency.
-
Scaling before validating
Volume on the wrong criterion multiplies rework. That's why the first sample isn't optional.
-
Delivery without a version
Without versioning, a problem found later can't be traced back to the batch and guideline that produced it.
Frequently asked questions about labeling quality
What gets asked before trusting a dataset to a vendor.
How do you handle disagreement between annotators?
We record it, resolve it in the guideline and re-annotate the affected material. Repeated disagreement signals an ambiguous criterion, and the criterion is what we fix.
Can you audit an existing dataset?
Yes. We assess a sample, measure consistency and point out where criteria diverge, before proposing correction or re-annotation.
Who writes the guidelines?
We write them together. We bring the structure and the edge cases that tend to appear; you bring the business definition of right and wrong.
What if the criteria change mid-project?
A new guideline version, affected material identified and re-annotated. What matters is that the change is recorded rather than implicit.
Talk to our team
Let's understand your challenge
Prefer to book the diagnosis conversation? Use the scheduling link. Or fill in the form — our technical team replies.