Model Evaluation

Why Low Average Error Isn't Good Enough

A 99.2% aggregate test accuracy can disguise complete failure on critical query cohorts, catastrophic edge cases, and high-value customer segments.

Aug 20, 20217 min read

Performance on disproportionately important examples

A model can have a remarkably low average error and still be unacceptable for production.

This is the core lesson behind slice-based evaluation: aggregate metrics can hide catastrophic failures on important subgroups.


The trap of average performance

Suppose a system has 98.7% overall accuracy. On the surface, this looks excellent. But if the model fails on a small but critical subset of examples, then the apparently good number is misleading.

This is especially dangerous in systems where some failure modes are dramatically more costly than others.

Deployment example from the original notebook


Why slices matter

In many real-world systems, the data is heavily imbalanced:

  • the majority of calls are routine
  • a small fraction are safety-critical or business-critical
  • the minority group drives the most severe consequences

If the model performs poorly there, the average metric becomes irrelevant.

A product can look highly accurate while still being operationally unsafe.


Search engine example

Imagine a ranking system where 95% of queries are informational and only 2% are safety-critical or high-risk.

The model may do extremely well on the majority cohort while failing badly on the small but crucial subset.

That creates a dangerous illusion:

  • aggregate accuracy looks great
  • true product risk is hidden
  • user trust erodes in the most important cases

What production teams need

A better evaluation system includes:

  • cohort-based metrics
  • error slices by user segment or query class
  • calibration checks
  • risk-aware thresholds
  • cost-sensitive decision evaluation

The goal is not only to minimize mean error, but to make sure the system remains reliable where it matters most.

A low average error is useful, but it is not sufficient. In production, the dangerous errors are often the rare ones.