← Back to topics page

Articles about "AI evaluation"

arstechnica.com

UK Government Evaluates Anthropic’s Mythos AI for Cybersecurity Risks

The UK government's AI Security Institute has published an independent evaluation of Anthropic's Mythos Preview AI model, focusing on its cybersecurity capabilities. Early tests show Mythos excels at chaining together multiple steps in cyberattack simulations, surpassing previous models in complex, task-based challenges. The evaluation adds public verification to Anthropic's internal findings and highlights both progress and risks in advanced AI for security.

arstechnica.com

Major AI Models Struggle to Predict Premier League Soccer Outcomes

A recent study by General Reasoning found that leading AI models, including those from Google, OpenAI, and Anthropic, consistently lost money when virtually betting on Premier League soccer matches. The findings highlight ongoing challenges for artificial intelligence in complex, real-world prediction tasks.

tech.eu

Galtea Secures $3.2M Seed Funding to Improve AI Testing

Galtea, an AI evaluation infrastructure provider, raised $3.2 million in seed funding to develop its platform for generating test scenarios for generative AI agents. The platform aims to streamline and reduce the cost of AI testing for enterprises, helping to accelerate deployment. The new funding will support platform development and team expansion.

news.mit.edu

J-PAL Launches Initiative to Evaluate Impact of AI on Poverty

The Abdul Latif Jameel Poverty Action Lab (J-PAL) at MIT has launched Project AI Evidence to rigorously evaluate the effectiveness and risks of artificial intelligence innovations in poverty reduction. The initiative funds new research and connects policymakers, tech companies, and non-profits to evidence-based insights on AI's real-world social impact. Early projects target sectors such as education, health, and employment, aiming to scale up responsible and effective AI solutions.

news.mit.edu

MIT Study Reveals Unreliability in LLM Ranking Platforms

A new MIT study finds that popular platforms ranking large language models (LLMs) can be highly sensitive to small changes in user feedback, potentially compromising the reliability of their rankings. The researchers developed a method to test and identify data points most influential in shifting results, underscoring the need for more robust evaluation strategies.

technologyreview.com

Debate Grows Over Interpreting METR's Influential AI Progress Graph

A widely cited graph by the nonprofit METR has fueled debates about the pace of progress in frontier large language models. While the graph suggests exponential increases in AI capabilities, experts caution that its key metric—the 'time horizon'—is frequently misunderstood and not a universal measure of AI intelligence.

The Essential Weekly Update

Stay informed with curated insights delivered weekly to your inbox.