MLOps & Systems

ML Data Is A First Class Citizen in Production

Why ML code in production is just a drop in the ocean compared to dynamic data pipelines, schema skew, and covariate drift.

Sep 29, 20218 min read

Machine learning development lifecycle

In classical academia, machine learning seems like a neat optimization problem: clean data, a model, a benchmark, and a report. In production, the reality is different. Model code is only a small part of the system. The critical burden sits in the data layer: ingestion, validation, drift monitoring, feature quality, and feedback loops.

The real challenge is not training a better model. The real challenge is keeping the data contract healthy over time.

The production reality

Data in an academic or research setting is usually curated, static, and well-behaved. In production, the data pipeline is continuous and messy:

  • Schema drift changes the meaning of the same columns over time
  • Data quality issues create silent model degradation
  • Covariate shift changes the distribution seen at serving time
  • Label delay makes evaluation lag behind business reality

That is why production ML systems are more about observability, monitoring, and feedback loops than pure predictive performance.

Production machine learning systems


Why model code is not enough

When you deploy a model, your work does not stop. It actually begins.

  1. Scoping: define the business problem and the right success metric
  2. Data pipeline: ensure ingestion, feature engineering, and schema contracts are stable
  3. Modeling and error analysis: inspect slices, not just aggregate accuracy
  4. Deployment and monitoring: detect drift and trigger automated retraining

A model can be algorithmically correct and still fail in production because the incoming data no longer matches the assumptions of training.

Machine learning production pipeline


Distribution shift in practice

The most common types of drift are:

Data changes after deployment

1. Concept drift

Changes in the relationship between features and target:

Ptrain(y∣x)≠Pserve(y∣x)P_{train}(y \mid x) \neq P_{serve}(y \mid x)

2. Covariate shift

Changes in the input distribution itself:

Ptrain(x)≠Pserve(x)whileP(y∣x) remains stableP_{train}(x) \neq P_{serve}(x) \quad \text{while} \quad P(y \mid x) \text{ remains stable}

3. Schema skew

The feature columns arrive with new types, missing values, or unanticipated categories.

This is why MLOps teams care deeply about data validation frameworks, lineage tracking, and continuous monitoring dashboards.

Data and model monitoring workflow

Production data quality example

The original notebook also walks through a TensorFlow Extended production pipeline:

TensorFlow Extended pipeline

TFX pipeline components

Additional TensorFlow Extended component view


Operational checklist

A production pipeline should include:

  • data contract validation
  • baseline and business KPI tracking
  • drift detection thresholds
  • replayable training data
  • canary rollout and rollback gates
  • retraining triggered by evidence, not guesswork

Feedback and continuous learning loop

Monitoring and retraining example

Serving and feedback workflow

Final thought

The highest-leverage system in machine learning is not the model itself. It is the data health system around it. If the data pipeline is healthy, the model has a chance to remain useful. If not, the production system silently decays.