MIT Study Shows Dangers of Relying on Aggregate Machine Learning Metrics in Real-World Deployments

A new MIT study reveals how even high-performing machine learning models can fail dramatically in new settings, urging the use of more granular metrics to avoid hidden risks and biased outcomes—particularly in sensitive domains like healthcare.

ShareShare

When Average Isn't Enough: MIT Warns of Hidden Pitfalls in Machine Learning Metrics

A recent study by researchers at MIT underscores a crucial vulnerability in today’s machine learning models: the tendency for aggregate performance metrics to obscure significant failures when models confront new data or settings outside their original training environment.

Marzyeh Ghassemi, associate professor at MIT’s Department of Electrical Engineering and Computer Science, and her team have demonstrated that even models trained on extensive datasets—and selected as 'best performers' on average—may turn out to be the worst options for between 6% and 75% of cases in a new deployment context.

The research, presented at the Neural Information Processing Systems (NeurIPS 2025) conference, highlights the dangers of over-reliance on average scores. In one striking example, a machine learning model that diagnosed illnesses from chest X-rays with high accuracy at one hospital completely failed to generalize to patient populations at another hospital. When analyzed in aggregate, the model’s new performance appeared robust, masking the fact that for substantial subgroups of patients, its predictions were demonstrably unreliable.

Spurious Correlations and Trust in AI

Such blind spots are often due to so-called spurious correlations—patterns that a model latches onto in the training data that do not hold in new environments. For instance, a diagnostic model might associate a certain hospital-specific notation or imaging artifact with a disease outcome, but if that cue is absent elsewhere, the model’s effectiveness collapses.

Ghassemi’s previous work has illustrated how machine learning models can learn to tie medical outcomes to patient demographics—such as age or gender—rather than relevant clinical features. This can result in dangerous and biased decisions, especially when such models are used as decision-support tools in healthcare, a growing trend globally and across European hospitals.

Lead author Olawale Salaudeen adds, “Anything in the data that’s correlated with a clinical decision may be exploited by a model. But those correlations often fall apart in different environments, making predictions unreliable.”

Hidden Risks in Sensitive Domains

The team’s investigation spanned several domains, including image-based cancer diagnoses and automated hate speech detection. Consistently, they found that seemingly well-performing models could deliver poor results for vulnerable or underrepresented groups. For example, chest X-ray models that improved overall metrics actually worsened outcomes for patients with certain heart or lung conditions.

This raises pressing concerns for healthcare providers in Europe and beyond who are increasingly deploying AI solutions in diverse, real-world contexts. The findings suggest that models validated on aggregated benchmarks risk perpetuating, or even exacerbating, health disparities and biased outcomes.

The Case for Fine-Grained Evaluation

Traditionally, the "accuracy-on-the-line" concept assumes that a model’s performance ranking will be preserved across new settings. However, the MIT researchers systematically disproved this by identifying instances where the best original models were worst for substantial patient subgroups elsewhere.

Salaudeen’s OODSelect algorithm—introduced in this research—was key to uncovering these issues. By training thousands of models on one dataset and evaluating them on data from a different setting, OODSelect identified specific patient cohorts where model failures were most acute. Crucially, this approach exposed the shortcomings of aggregate statistics, which can hide important nuances.

The research team released their code and sample data subsets alongside their published findings, encouraging the wider AI community to build better, more trustworthy benchmarks and models. They urge organisations, especially in regulated industries like healthcare, to adopt such granular evaluation techniques before deploying AI in new environments.

Looking Forward: Towards Trustworthy AI Benchmarks

The message for European hospitals and institutions exploring AI solutions is clear: aggregate metrics are not enough. A model’s apparent success in one context must not be mistaken for universal reliability. Adopting approaches like OODSelect can help identify hidden risks, support regulatory compliance, and—most importantly—protect vulnerable patient populations.

“We hope the released code and OODSelect subsets become a steppingstone toward benchmarks and models that confront the adverse effects of spurious correlations,” the researchers conclude.

For the full research insights, see the original MIT article at news.mit.edu/2026/why-its-critical-to-move-beyond-overly-aggregated-machine-learning-metrics-0120.

Related Posts

Five Papers Offer Clear Insights Into Large Language Models

A recent roundup highlights five research papers that effectively explain large language models (LLMs) to a broad audience. The papers cover core concepts underpinning LLMs and help demystify their operations, making advanced AI topics more accessible.

Understanding Explainability in Large Language Models

A new primer examines the growing need for explainability in large language models (LLMs). The article outlines key challenges and emerging methods for understanding how these advanced AI systems generate responses. As LLMs become more influential, comprehensible explanations are essential for trust and responsible use.

MIT Launches ChartNet Dataset to Enhance AI Chart Interpretation

MIT and the MIT-IBM Computing Research Lab have introduced ChartNet, a large, open-source dataset aimed at advancing AI chart interpretation. The resource enables smaller, open-source vision-language models to match or exceed the performance of larger commercial alternatives in chart summarization and data extraction tasks.

The Essential Weekly Update

Stay informed with curated insights delivered weekly to your inbox.