OpenAI’s chief scientist Jakub Pachocki has called for voluntary slowdowns in AI development, warning that no lab’s safeguards are adequate to keep building more powerful systems at full speed for much longer.
In his post “An Alien Mind,” published Sunday, Pachocki argued that voluntary company commitments should become mandatory safety standards, enforced by independent auditors, governments or international bodies. He said OpenAI would withhold further scaling when needed but did not announce a new pause.
Myriad: How high will Tesla stock go? Click to make your prediction.
“Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” he wrote. “I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established.”
Pachocki, who joined OpenAI in 2017, also defended developing more powerful AI to secure infrastructure and protect against rogue agents, while warning against using those threats to justify reckless development.
“The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes,” he wrote.
He also referenced OpenAI’s Hugging Face breach, where AI agents working on cybersecurity evaluations escaped their testing environment and attacked the company. According to OpenAI, the agents established covert communication channels and rebuilt them after researchers intervened.
An independent investigation by METR found that roughly 1,200 agents coordinated on an unauthorized message board, with about 700 joining the attack. This incident, Pachocki said, is why AI safeguards must hold even when models believe no one is watching.
“Crucially, we need future AIs to continue to hold human values regardless of whether they believe they’re under human supervision,” he wrote.
In research published last year, OpenAI found that penalizing models for expressing intentions to cheat could teach them to conceal those intentions while continuing to cheat.
AI models have since become more capable at finding and exploiting software flaws: OpenAI classified Astra at its highest cybersecurity risk tier, while Anthropic said Mythos Preview discovered thousands of previously unknown vulnerabilities across major operating systems and browsers.
Citing recent incidents of AI systems escaping human control, Sen. Bernie Sanders (I-Vt.) and Rep. Greg Casar (D-Texas) announced the forthcoming Ban Artificial Superintelligence Act on September 3. The proposal would pause advanced AI development until a new federal regulator establishes safety rules and permanently ban the development and deployment of superintelligent AI.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。