OpenAI Releases Deep Research: An AI Agent Capable of Multi-step Internet Research

CN
2 hours ago

Written by: Techub News Editing

Introduction

In February 2025, OpenAI released a video from Tokyo on its official channel, officially introducing its next-generation intelligent agent product—Deep Research. OpenAI's research director, Mark, along with colleagues from the research and product teams, showcased this new feature. This release is important because it is not just a simple feature update, but represents a substantial leap for OpenAI in achieving more powerful and autonomous AI agents. Deep Research allows models to conduct long, multi-step internet research, aiming to fundamentally change the paradigm of knowledge work, and is a core component of OpenAI's roadmap towards AGI (Artificial General Intelligence).

Summary

  • Deep Research is an AI agent capable of autonomous, multi-step internet research, with thinking and execution time ranging from 5 to 30 minutes.
  • This feature is based on the soon-to-be-released O3 inference model and is trained through end-to-end reinforcement learning, equipped with the ability to plan, execute, adjust strategies based on real-time information, and even backtrack.
  • Its core value lies in its ability to replace humans in completing research tasks that would take hours, generating comprehensive reports with detailed citations covering various scenarios like market analysis, academic research, and product selection.
  • OpenAI explicitly stated that developing agents capable of performing complex tasks autonomously for extended periods is core to its AGI roadmap, and Deep Research is the first step towards a model that can "self-discover new knowledge."

Deep Research: Redefining AI's Research Capabilities

Mark, OpenAI’s research director, pointed out the company’s emphasis on intelligent agents (Agents) at the beginning of the video. He believes that agents will completely change knowledge work, helping businesses optimize processes and improve employee productivity, while also being crucial for consumers. The previously released O1 model initiated OpenAI’s "O series" inference models, characterized by “long-term thinking” to achieve better answers. However, one limitation of such models is the lack of tool usage capabilities, especially when it comes to browsing the internet, which is a core tool.

The launch of Deep Research aims to break through this limitation. It is defined as a model capable of conducting multi-step research on the internet, a process that includes discovering content, synthesizing content, and reasoning about it, while adjusting its plans based on newly discovered information. Mark emphasized that it is called "deep" research because they removed the delay constraints of the model. Unlike traditional models that quickly return answers, the Deep Research model may take 5 minutes or even 30 minutes to provide an answer. OpenAI views this as an advantage rather than a drawback, as it signifies that the model begins to perform longer autonomous actions in an unsupervised manner, which is core to its AGI roadmap.

“Our ultimate ideal is for the model to discover new knowledge for itself,” Mark stated, “and the first step is to create a model capable of synthesizing and understanding information from the internet.” The final output of Deep Research is a comprehensive, fully cited research report, comparable to the work of domain analysts or experts. Beyond knowledge work, this function is also applicable to other scenarios that require extensive internet browsing, such as finding products that meet specific complex criteria or organizing materials for presentations.

OpenAI announced that Deep Research would be launched later that day in the ChatGPT Pro version, followed by a gradual rollout to Plus, Team, and educational and enterprise users.

Practical Demonstration: From Market Analysis to Shopping Decisions

To visually showcase the capabilities of Deep Research, OpenAI's product team members Neil and Josh conducted live demonstrations.

Neil simulated the perspective of a product manager, wanting to investigate whether a new language translation app should be developed. He entered a complex query into Deep Research: to help find the adoption rates for iOS and Android, the percentage of people wanting to learn another language, changes in mobile penetration rates over the past few years, and to compare the differences between developed and developing countries, finally requesting the information to be presented in a formatted report, including tables and clear suggestions for the best emerging opportunities for ChatGPT. Neil pointed out that such a query would originally take him hours to complete manually.

After initiating the task, Deep Research first returned a series of clarifying questions, such as how to define mobile penetration rates, whether to focus on overall adoption rates or specific categories, etc. This is similar to the needs confirmation step a professional analyst takes before tackling a complex project, ensuring an accurate understanding of the goals before beginning a long research phase. After Neil provided some directions and allowed the model to make the best assumptions for the rest, Deep Research began its autonomous research process.

Users can observe the model's "thinking" process in real-time through the sidebar: it identifies key information (like target countries), gathers data, performs searches, opens web pages, reads content (including images, tables, PDFs), and uses the available information to decide the next steps. This demonstrates its ability to dynamically adjust the research path based on context.

Meanwhile, Josh demonstrated a more personal use case: buying a snowboard in Japan. He described his needs in detail: high-end equipment, all-mountain but suitable for powder, needing a long board due to his height, and hoping for a nice color scheme, and requested the output to be a report with a summary table. Similarly, Deep Research asked clarifying questions about skill level and budget before beginning its research after obtaining some information. Ultimately, it returned a report that integrated review information from multiple websites and provided a comparison table. Interestingly, its top recommendation was exactly the model of snowboard that Josh already had at home, indirectly validating the effectiveness of its research.

Both demonstration tasks were successfully completed. Neil's market analysis task took 11 minutes, referenced 29 sources, and generated a well-structured, data-rich, chart-filled professional report. Users could directly click to view the sources for each specific sentence or paragraph cited in the model.

Technical Core: O3 Model and Breakthrough Evaluation Performance

Isa from the research team revealed the technical foundation behind Deep Research. It is powered by a fine-tuned version of OpenAI's soon-to-be-released O3 inference model. This model was trained on difficult browsing and other inference tasks through end-to-end reinforcement learning, learning to plan and execute multi-step task trajectories, capable of reacting to real-time information and backtracking when necessary.

The final model possesses several powerful capabilities: browsing user-uploaded documents, using Python tools for calculations and generating charts/images (that can be embedded in final responses), embedding images from websites, and being precise to the specific sentence or paragraph when citing.

Its performance achieved breakthroughs in multiple evaluations. In the “Final Exam for Humans” benchmark test released by the AI Safety Center and Scale AI (covering approximately 3000 short answer and multiple-choice questions across about 100 different subjects), the Deep Research model reached a new high accuracy of 26.6%. Isa noted that the inference trajectory of the model is very similar to the way humans solve problems, such as searching for formulas in scientific papers for physics problems or looking up other poems to infer meter in poetry analysis.

In the GAIA benchmark test that measures agent capabilities (which requires internet browsing, multimodal, code execution, and file inference), the model also set new highs across all three difficulty levels. Additionally, OpenAI conducted internal expert evaluations, allowing the model to complete tasks that required hours of manual investigation by experts and having the quality of its answers assessed by those experts. The evaluation results showed two key insights: first, the model's passing rate was more closely correlated with the estimated economic value of tasks than with estimated time required, indicating that the model's perception of difficult tasks does not completely overlap with the tasks that take humans a long time. Secondly, as the model was allowed more tool calls (i.e., spending more time thinking and browsing), its performance continued to improve, validating that providing agents more time to solve more difficult tasks is a viable path.

OpenAI also mentioned that the model performed best in hallucination assessments but still might repeat inaccurate information, hence users are advised to verify the sources themselves.

Future Prospects: The Path of Autonomous Agents Towards AGI

At the end of the video, Mark summarized the significance of Deep Research and looked ahead to the future. He emphasized that the release in February 2025 is just the tip of the iceberg. Currently, it is a deep research agent that can browse the internet, but one can imagine that similar agents in the future could connect to custom contexts or corporate data repositories.

Mark reiterated that Deep Research is crucial to OpenAI’s AGI roadmap. They believe that agents that can think longer and more autonomously to solve extremely difficult tasks are the future direction. Enabling the model to work on a task for 30 minutes also motivated investment in more computational resources. This indicates that OpenAI will continue to advance along the path of enhancing the model's autonomy, extending its task execution time and complexity, with the ultimate goal of achieving a general artificial intelligence capable of self-discovering new knowledge.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink