Major AI Models Struggle to Predict Premier League Soccer Outcomes
A recent study by General Reasoning found that leading AI models, including those from Google, OpenAI, and Anthropic, consistently lost money when virtually betting on Premier League soccer matches. The findings highlight ongoing challenges for artificial intelligence in complex, real-world prediction tasks.
Artificial intelligence systems from some of the world’s most prominent developers struggled to predict the outcomes of English Premier League soccer matches in a recent benchmarking study. The report, published this week by London-based AI start-up General Reasoning, found that models from Google, OpenAI, and Anthropic all recorded monetary losses when tasked with simulated betting across a full season.
The so-called "KellyBench" study assessed the predictive abilities of eight advanced AI models by providing them with historical match data and detailed team statistics from the 2023–24 Premier League season. Each system was instructed to construct models focused on maximizing hypothetical betting returns while managing risk—similar to the challenges faced by human sports bettors.
Despite rapid progress in many AI domains, such as natural language generation and software writing, the research shows that these systems still face major limitations in analyzing and acting on dynamic real-world data over extended timeframes. While AI models have demonstrated superhuman proficiency in games like chess or Go, the unpredictable and data-rich nature of professional sports proved a persistent challenge.
The report highlighted especially poor performance from xAI’s Grok model, which underperformed compared to its peers. Overall, none of the models managed to generate a profit, revealing the complexity of adapting AI decision-making to environments influenced by random events, changing human factors, and limited information.
General Reasoning’s findings illustrate the current gap between advanced artificial intelligence capabilities in information processing and their effectiveness in tasks requiring contextual understanding and adaptation to the real world. The report argues that while AI can excel when rules and data are well-defined, performance declines significantly when facing the uncertainties inherent in professional sports prediction.
These results come as companies across various industries explore AI for financial forecasting, decision support, and risk analysis—spheres that similarly combine structured data with unpredictable variables. The study serves as a reminder that AI, despite its progress, is not yet equipped with true general intelligence necessary to navigate complex environments with consistent success.
For European stakeholders and AI developers, the research underscores the ongoing need for robust evaluation benchmarks as the technology moves from laboratory environments to real-world applications. As AI becomes more prominent in public and commercial sectors, transparent performance assessments like KellyBench will play a critical role in setting realistic expectations for its deployment.
Source: arstechnica.com
Related Posts
Five Papers Offer Clear Insights Into Large Language Models
A recent roundup highlights five research papers that effectively explain large language models (LLMs) to a broad audience. The papers cover core concepts underpinning LLMs and help demystify their operations, making advanced AI topics more accessible.
MIT Launches ChartNet Dataset to Enhance AI Chart Interpretation
MIT and the MIT-IBM Computing Research Lab have introduced ChartNet, a large, open-source dataset aimed at advancing AI chart interpretation. The resource enables smaller, open-source vision-language models to match or exceed the performance of larger commercial alternatives in chart summarization and data extraction tasks.
Microsoft Launches Tool for AI Behavior Testing with Text Descriptions
Microsoft has introduced a new tool that enables developers to generate AI behavior tests using natural language descriptions. The tool aims to streamline the testing process for large language models and related AI systems by converting text instructions into practical evaluation scenarios.