GPT-5.5 Matches Mythos Preview in Cybersecurity Benchmark Tests
OpenAI's GPT-5.5 has demonstrated cybersecurity capabilities similar to Anthropic's Mythos Preview model, according to new tests by the UK's AI Security Institute. Both models performed comparably in complex evaluation scenarios, suggesting that cybersecurity risks from advanced AI are not confined to any single provider.
UK Government Evaluates Anthropic’s Mythos AI for Cybersecurity Risks
The UK government's AI Security Institute has published an independent evaluation of Anthropic's Mythos Preview AI model, focusing on its cybersecurity capabilities. Early tests show Mythos excels at chaining together multiple steps in cyberattack simulations, surpassing previous models in complex, task-based challenges. The evaluation adds public verification to Anthropic's internal findings and highlights both progress and risks in advanced AI for security.
Major AI Models Struggle to Predict Premier League Soccer Outcomes
A recent study by General Reasoning found that leading AI models, including those from Google, OpenAI, and Anthropic, consistently lost money when virtually betting on Premier League soccer matches. The findings highlight ongoing challenges for artificial intelligence in complex, real-world prediction tasks.
Galtea Secures $3.2M Seed Funding to Improve AI Testing
Galtea, an AI evaluation infrastructure provider, raised $3.2 million in seed funding to develop its platform for generating test scenarios for generative AI agents. The platform aims to streamline and reduce the cost of AI testing for enterprises, helping to accelerate deployment. The new funding will support platform development and team expansion.
J-PAL Launches Initiative to Evaluate Impact of AI on Poverty
The Abdul Latif Jameel Poverty Action Lab (J-PAL) at MIT has launched Project AI Evidence to rigorously evaluate the effectiveness and risks of artificial intelligence innovations in poverty reduction. The initiative funds new research and connects policymakers, tech companies, and non-profits to evidence-based insights on AI's real-world social impact. Early projects target sectors such as education, health, and employment, aiming to scale up responsible and effective AI solutions.
MIT Study Reveals Unreliability in LLM Ranking Platforms
A new MIT study finds that popular platforms ranking large language models (LLMs) can be highly sensitive to small changes in user feedback, potentially compromising the reliability of their rankings. The researchers developed a method to test and identify data points most influential in shifting results, underscoring the need for more robust evaluation strategies.
Debate Grows Over Interpreting METR's Influential AI Progress Graph
A widely cited graph by the nonprofit METR has fueled debates about the pace of progress in frontier large language models. While the graph suggests exponential increases in AI capabilities, experts caution that its key metric—the 'time horizon'—is frequently misunderstood and not a universal measure of AI intelligence.