OpenAI Chief Scientist Ilya Sutskever: Reflections and Prospects on the Road to AGI, from Benchmark Saturation to Automated Research

CN
1 hour ago

Written by: Techub News整理

Introduction

In the latest episode of the OpenAI official podcast, OpenAI co-founder and chief scientist Ilya Sutskever and researcher Simon Sidorenko joined host Andrew Maine for an in-depth conversation. They traced their common high school days, discussed the challenges of measuring AI progress today, shared an insider perspective on the groundbreaking performance achieved by the August 2025 model in the Math Olympiad, and looked ahead to the road to AGI and potential breakthroughs in the future. As a core architect of OpenAI's research roadmap, Ilya Sutskever's insights reflect the profound contemplation of this company, which is at the center of the AI wave, regarding the current state and future direction of technological development.

Summary

  • Traditional AI capability benchmarks (such as standardized tests) are facing "saturation," with models reaching or exceeding human levels in many areas, making snapshot measurements insufficient.
  • The focus of measuring AI progress should shift to its impact and utility in the real world, particularly its potential for automated discovery and creation of new technologies.
  • The groundbreaking performance in math and programming competitions in August 2025 demonstrates that AI has made substantial advancements in deep, creative reasoning, beyond mere memory and pattern matching.
  • The next major breakthrough may be in further expanding the model's "persistence" and long-range reasoning capabilities, allowing it to invest vast computational resources in focused exploration of important problems, similar to human researchers.

From Programming Competitions to AGI Reflections: The Growth Trajectories of Two Researchers

The conversation began with amusing personal histories. Ilya Sutskever and Simon Sidorenko are not only colleagues at OpenAI but also high school alumni. They both come from a Polish high school focused on computer science competitions and were influenced by an inspiring computer teacher who exposed them early to programming, graph theory, and other knowledge far beyond the high school curriculum. Simon Sidorenko believes that with tools like ChatGPT available today, it may become easier for more people to engage in similar depth of knowledge exploration, but he also emphasized that the emotional support and growth space provided by their teacher back then is something that AI currently struggles to replicate on its own. This gives rise to the point: AI will not replace education; rather, it will become a powerful assistive tool for teachers.

This growth experience also laid the groundwork for their perspectives on AI milestones. Whether it’s the International Mathematical Olympiad (IMO) or the International Olympiad in Informatics (IOI), these competitions represent not just "benchmarks," but concrete symbols of intellectual challenges at the peak of personal growth. This emotional connection enables them to evaluate AI progress from both an internal technical perspective and an understanding of potential cognitive differences in the external world.

The Evolution of the Definition of AGI: From Abstract Concepts to Concrete Impacts

When asked how to define AGI (Artificial General Intelligence), Ilya Sutskever candidly stated that the concept has undergone a substantial evolution over the past few years. A few years ago, although deep learning had a promising future, AGI still seemed abstract and distant. Whether it was "human-level intelligence," "natural conversation," or "solving mathematical problems," these goals appeared to fall into the same vague category.

Now, with technological advancements, these capabilities have proven to be quite distinct. AI has already engaged in natural conversations on a wide range of topics and can solve mathematical problems—such as winning a gold medal in the IMO, which has long been seen as a milestone on the road to AGI. However, Ilya Sutskever pointed out that this snapshot milestone measurement is becoming increasingly inadequate. His focus has shifted to the actual impact of AI on the world, especially its potential for the automated discovery and production of new technologies.

"We tend to associate fundamental technological advancements with human intelligence," he said, "but it's hard to internalize: most of this process can be automated." He envisions a powerful computing system capable of generating fundamentally new ideas that change our understanding of the world, and he believes that day is not far off. For him, an important hallmark of AGI is the ability to operate like an automated, super-efficient team of researchers or engineers, continuously producing code, designs, and even scientific discoveries, thereby dramatically accelerating the pace of technological progress.

The Dilemma of Benchmarks: Saturation, Limitations, and Seeking New Metrics

The discussion delved into the challenges currently facing the assessment of AI model capabilities. Ilya Sutskever and Simon Sidorenko both pointed out the benchmark "saturation problem". As models reach or exceed human top levels in many standardized tests (especially in elite high school competitions), it's become increasingly difficult to find more challenging and adequately constrained metrics.

Moreover, the previous "rising tide lifts all boats" model of broadly enhancing capabilities simply by expanding pre-training size (from GPT-1 to GPT-4) is changing. Today, researchers can train models more efficiently for specific capabilities (like mathematics), but this may lead to models that perform excellently on specific benchmarks yet do not represent their overall intelligence level. A model might win a math competition but perform mediocrely in creative writing.

Simon Sidorenko shared an amusing anecdote: when he excitedly told a multilingual colleague, Anna, about the progress in IMO, she responded with "What's IMO?" This made him realize that the research community sometimes lives in a "bubble," where milestones they consider important may be insignificant to others. This prompted them to seek more universal metrics.

An imperfect but practical metric is the ChatGPT itself and its wide variety of use cases. Hundreds of millions of users apply it to countless scenarios, providing a broad “stress test.” However, they also foresee that a key future measure of progress may lie in whether models can leverage computational resources far exceeding what individual users can afford to solve significant problems of value to many, such as medical research or the development of next-generation AI models themselves.

Breakthrough Moments: When Models Begin to "Surprise" Us

Reflecting on the fast pace of AI development, Simon Sidorenko described his "personal AGI moment." During the GPT-4 era, he felt for the first time that models could sometimes say things that surprised him. From early simple character predictions and flawed sentiment analysis (like misclassifying "this movie is good" as negative), to the shock of coherent paragraphs generated by GPT-2, and eventually the surprises from GPT-4, culminating in today's models competing against top participants in programming contests, the progress over this decade seems "incredibly astonishing" to him.

Ilya Sutskever mentioned the breakthroughs achieved in reasoning models (like o1) in August 2025. The team significantly enhanced the model's complex problem-solving abilities by allowing it to engage in longer "thought chains" or internal monologues. He recalled that when they first observed that this method was indeed effective and that increasing training data continuously improved performance, the entire team felt "shocked." He even mentioned a night when they had a serious conversation with Sam Altman and Mira Murati, discussing whether the organization was ready to embrace such rapid developments, adding, "Sometimes we really are taken aback by these results."

He also shared a dramatic competition story: during the AtCoder competition in Japan (a 10-hour single-problem heuristics optimization contest), their model competed against a friend of Simon Sidorenko's who was a top contender in that event, named Siho. The model ultimately secured second place, while Siho won first. Siho, exhausted after the competition, quipped in an interview that OpenAI's model was "very, very bad," as he just wanted to sleep. This little episode vividly illustrates the struggle between human endurance and AI intelligence.

The Road Ahead: Persistence, Scalability, and Trust

Looking ahead to future breakthroughs, Ilya Sutskever emphasized not to underestimate the enduring importance of scalability. The pre-training expansion paradigm has not disappeared and will compound with new research directions. A clear direction is to expand the "time horizon" or "persistence" of models for planning and reasoning.

He noted that from GPT-4 to GPT-4o or more advanced models, the computational resources consumed per answer might only increase tenfold or twentyfold but could yield significantly better answers. However, for truly important questions (like medical research), the computational resources that people are willing to invest will be uncomparably vast. Therefore, enabling models to handle single complex problems over long periods of time with focus is the next critical step.

What might AGI-level capabilities look like for average ChatGPT users in the coming years? Ilya Sutskever once again returned to the image of automated researchers. This isn't just a black box; it is an intelligent entity capable of interacting with humans, absorbing information, running experiments, and producing code and designs. At the same time, the interface of interaction between users and AI will become more humanized and persistent, allowing people to form deeper connections with AI, which will become an important social issue.

Finally, their advice to today's high school students is practical and sincere: learn programming. Simon Sidorenko refuted claims of "don't learn programming," stating that programming is an excellent way to exercise the structured thinking ability needed to break down complex problems into modules, a skill that will remain crucial in the future. Ilya Sutskever encouraged young people to break through self-imposed limits, dare to have ambitious goals, and believe they can make a meaningful positive impact on the world—that's the inspiring spirit he feels in the Silicon Valley community.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink