OpenAI Sets Ambitious Goal for Fully Automated AI Researcher
OpenAI is prioritizing the development of an AI system capable of conducting advanced research autonomously. The company aims to deliver a prototype research intern AI by September, with plans for a full multi-agent system by 2028. The project raises significant technical, ethical, and governance questions regarding oversight and concentration of power.
OpenAI has announced a major shift in its research agenda, focusing efforts and resources on building a fully automated AI researcher. The company’s new "North Star" is the creation of an agent-based system able to independently solve complex scientific and technical problems, with plans to unveil a prototype research intern by September 2026 and a multi-agent version by 2028.
This strategy builds on OpenAI’s advancements in large language models (LLMs) and autonomous coding agents, such as Codex. LLMs are AI models trained on vast amounts of text data, enabling them to process and generate human-like language. Codex, for example, can autonomously handle substantial programming tasks, and is seen by OpenAI as a conceptual template for future research agents capable of tackling problems in code, mathematics, life sciences, and policy.
Jakub Pachocki, OpenAI's chief scientist, has played a significant role in steering this initiative. He is convinced that recent leaps in AI—particularly advances in "reasoning models," which allow systems to work through tasks step by step—will enable sustained autonomous research. Pachocki emphasizes that while full human parity is not expected by 2028, even sub-human but highly capable systems could have transformative impacts.
Recent applications of OpenAI’s technology have included automated solutions to unsolved math problems and new findings in biology, chemistry, and physics. These successes underscore Pachocki's argument that continued scaling of LLMs and training on complex, long-horizon tasks will result in agents that can operate independently on extended research projects.
However, the pursuit raises a number of unresolved concerns. Pachocki acknowledges significant risks, such as the potential for misuse or system misalignment, especially if these agents operate with minimal human supervision. OpenAI is currently addressing these issues through "chain-of-thought monitoring," a method where AI agents document their reasoning process step by step, allowing for human or machine oversight.
The question of control and oversight remains pressing. Pachocki points out that as AI systems become more capable, power could become highly concentrated, with very few entities—or even individuals—able to direct the work of a fully autonomous research lab. He stresses that responsibility for governing these systems must extend beyond OpenAI and involve national policymakers.
The news comes amid heightened competition within the AI sector, with companies like Anthropic and Google DeepMind also pursuing long-term agent-based AI projects. The broader debate over regulation and the appropriate use of such powerful technologies continues, with recent disputes in the United States over military applications underscoring society’s uncertainty regarding boundaries for AI deployment.
While OpenAI’s stated mission remains to create artificial general intelligence (AGI) that benefits humanity, the focus for the next several years appears to be on "economically transformative" automated research—technology capable of reshaping scientific discovery, innovation, and potentially broader sectors of society.
Source: technologyreview.com.
Related Posts
Five Papers Offer Clear Insights Into Large Language Models
A recent roundup highlights five research papers that effectively explain large language models (LLMs) to a broad audience. The papers cover core concepts underpinning LLMs and help demystify their operations, making advanced AI topics more accessible.
MIT Launches ChartNet Dataset to Enhance AI Chart Interpretation
MIT and the MIT-IBM Computing Research Lab have introduced ChartNet, a large, open-source dataset aimed at advancing AI chart interpretation. The resource enables smaller, open-source vision-language models to match or exceed the performance of larger commercial alternatives in chart summarization and data extraction tasks.
Microsoft Launches Tool for AI Behavior Testing with Text Descriptions
Microsoft has introduced a new tool that enables developers to generate AI behavior tests using natural language descriptions. The tool aims to streamline the testing process for large language models and related AI systems by converting text instructions into practical evaluation scenarios.