Data Quality & Labeling

Why Data Definition Is Hard in Real-World ML

The hidden trap of label inconsistency in unstructured datasets, resolving annotator disagreement, and establishing Human Level Performance baselines.

Jul 02, 20216 min read

Major types of data problems

A high-priority ML initiative rarely fails because the model architecture is weak. It fails because the ground truth is poorly defined.

In textbook settings, labels are usually clean, consistent, and easy to reason about. In real business applications, especially with unstructured or semi-structured data, annotation quality becomes the first real bottleneck.

Notebook example introducing the data-definition problem


The label inconsistency problem

When data comes from human judgment, boundary cases produce disagreement:

  • radiologists disagree on subtle medical findings
  • annotators disagree on nuanced sentiment and sarcasm
  • reviewers classify edge events differently across teams

This means the target itself is noisy before the model even trains.

If two annotators label the same sample differently, the model is not learning a single truth; it is learning a mixture of contradictory supervision signals.

Why label consistency matters


Why this is harder than it looks

The challenge is not just getting labels. It is defining what the label means.

A dataset may look clean on the surface but fail in practice because:

  • annotation rubrics are underspecified
  • edge cases are not documented
  • labelers interpret the same example differently
  • class definitions shift over time

This creates a hidden noise floor that no model architecture can fully fix.

How inconsistent labels distort the learned relationship


The statistical impact

If the label is inconsistent, then even a strong model cannot learn a stable signal. In practical terms, the system may have:

  • reduced precision on ambiguous examples
  • unstable validation trends
  • inconsistent human-level performance estimates
  • poor transfer from pilot to production

This is why data definition is often the first real engineering task in ML work.

Human-level performance and evaluation

Another human-level performance example


What good teams do

The best data teams do not just collect examples. They define:

  • annotation rules
  • edge-case policies
  • disagreement review workflows
  • inter-annotator agreement checks
  • human-level performance baselines

They treat labels as a product, not as an afterthought.

If your data definition is vague, your ML system will be brittle regardless of the model you choose.