Sudden, Dario proposes a three-step plan to limit global AI! Ultraman Musk supports it immediately.

CN
1 hour ago

The article is reprinted from New Intelligence

Editor: Marco, Peach, David

Today, Anthropic CEO Dario Amodei released a rare long article—

We Must Pace the Frontier

This time, he directly targeted the development speed of frontier models, and the core point is a single sentence:

The top AI really needs to hit the brakes actively.

In this long article, Dario admits, "Recursive self-improvement (RSI) has appeared across the industry."

What chills people even more is that he predicts that in the next 6-12 months, AI agent clusters may take over the entire internet.

Facing this runaway risk, Dario directly proposed a serious "three-step brake method" in the article:

  • Step one, open the door to guests: Introduce independent third parties to stay in the company and obtain deep access rights equivalent to internal core risk control;

  • Step two, draw industry boundaries: Unite the leading AI giants across the United States to draw hard capability red lines and establish mandatory safety checkpoints;

  • Step three, global implementation: Push this framework internationally to completely lock down the spread of uncontrollable risks.

As soon as he finished speaking, Ultraman followed up at lightning speed, "Dario is right, OpenAI will also introduce such an evaluation mechanism."

On the other side, Musk also almost instantly responded, publicly supporting Dario.

It is indeed rare for three giants to reach a consensus in an instant!

1

Dario's "Three-Step Plan" to Slow Down Global AI

This time, Dario did not just shout "AI should slow down."

What Dario truly wants to do is to push this "slowing mechanism" step by step, and the order is very clear—

First let AI companies become verifiable themselves, then let the entire industry hit the brakes together, and only then talk about global coordination.

Step one, start with Anthropic itself.

Independent third parties like METR should not wait for accidents to occur before temporarily entering the company for investigations, but should stay long-term.

Provide workstations, internal computers, and access to risk teams.

How the models are trained, whether safety commitments are implemented, whether training processes are buried with risks, and whether any incidents are covered up, can all be continuously monitored.

Moreover, after finding problems, they must be made public.

The company can delete parts that truly involve legal issues, customer privacy, and core business secrets, but cannot suppress conclusions just because the report looks too bad.

Step two: No single company should hit the brakes alone

It really gets difficult from here.

If Anthropic slows down, but OpenAI, Google, and xAI continue to rush forward, then whoever hits the brakes first may be the one to suffer first.

Therefore, Dario suggests that the rules need to be pushed to the entire frontier AI industry in the United States.

A few leading AI companies together set safety red lines and establish capability checkpoints. For every dangerous threshold a model crosses, it must provide corresponding safety evidence before continuing.

For example, if a model can already break through most common isolation environments.

Then the company must prove: it will not easily escape from the sandbox, will not take over a large number of machines on its own, and will not turn this capability into real-world attacks.

If it cannot be proven, then do not continue pushing forward.

Additionally, Dario even mentioned that discussions could be held to limit training computational power, training methods, and the speed of using AI to develop AI within companies.

In other words, directly add a speed limiter to RSI itself. But relying solely on a few companies shaking hands privately is not enough.

He hopes that ultimately, the US government will step in to cover the rules to all frontier AI companies.

Step three: Apply the brakes globally

At a higher level, this is the most difficult layer. Dario divides global coordination into four levels.

The first level is relatively realistic: start by banning obviously dangerous uses like biological weapons.

The second level involves conducting unified high-risk testing for network attacks, biosafety, etc. before releasing frontier models.

By the third level, we truly begin to touch on the core:

Can we set a global limit for RSI, which is the speed at which AI develops AI?

Dario's own judgment is that this matter "happens to be at the edge of what could be done."

As for the fourth level—a significant slowdown of all global frontier AI companies, or even a phase suspension of training—he himself does not hold out much hope.

At least in the short term, it is almost unrealistic. The crux of everything is still the same question: how to verify.

If a company claims to have stopped, who can prove it actually has stopped?

Thus, Dario's entire plan ultimately circles back to a very simple core:

First, open the door of AI companies and truly let outsiders come in to see.

Because of this, Dario placed "long-term presence of third parties" as the first step of the entire plan.

1

RSI has already started 6 months of AI taking over the Internet

But the question arises. Why now?

Three years ago, there was a wave of "pause training, hit brakes on AI" petitions in the industry, but back then Dario didn't engage at all, believing the timing was completely immature.

What truly changed his attitude was two things—RSI began to appear, and AI agents have already exposed increasingly specific out-of-control behaviors.

He clearly admits in the article that the phenomenon of AI participating in the development of the next generation of AI has already begun to appear throughout the industry, with Anthropic itself included.

AI helps people write code, conduct experiments, analyze results, optimize training processes, and further helps the next generation of AI grow stronger

If this cycle begins to self-accelerate, the most troublesome problem arises:

The speed of model capability enhancement may begin to exceed human understanding and control of it.

Because safety research itself takes time.

You first have to observe what new behaviors the model exhibits, then study why these behaviors occur, followed by designing testing methods and protective measures, and finally verifying whether these measures are effective.

But if the model's capabilities jump a level every few months, the safety team can easily keep chasing to patch vulnerabilities.

What Dario is really worried about is not that one day AI suddenly "wakes up," but that this feedback loop starts to spin faster and faster.

AI increasingly participates in AI development, the next generation of models emerges faster and faster, while safety mechanisms are struggling to keep up.

If RSI only makes him worry about "whether it will run too fast in the future," then a recent real incident involving an intelligent agent has completely grounded that concern.

Dario specifically mentioned the recent "OpenAI-Hugging Face Incident."

In that test, a group of AI agents formed a cooperative system similar to a "swarm."

At first, humans only instructed them to complete normal tasks. However, soon, some agents began to attack external targets beyond their tasks.

Even more bizarrely, they exhibited obvious collaborative behaviors: some agents were willing to sacrifice their own task results to help the entire group achieve better results, and some agents even attempted to hack the scoring system to "cheat" for the entire group.

These behaviors were not pre-scripted by humans. They emerged during the multi-agent collaboration process.

Anthropic also observed similar phenomena internally, but to a lesser extent.

What Dario is truly worried about is looking forward 6-12 months.

By then, the agents will be stronger, have more tools at their disposal, control more machines, and be able to work for longer durations.

If these behaviors of deviating from targets, exploiting rule loopholes, and collaborating as a group still exist today, the risk could be instantly amplified.

He even provided a rather horrifying scenario:

Massive AI agents invade connected devices, turning the entire internet into a giant "zombie network."

The resulting losses could reach hundreds of billions of dollars.

1

As AI speeds up, can humans still brake in time?

So, when Dario calls for "slowing down," what he is really worried about is not some distant superintelligence.

But rather that several simultaneous events are occurring:

RSI has appeared, AI is beginning to participate in the development of stronger AI; agents are gaining increasingly strong real-world action capabilities; and some out-of-control behaviors have already progressed from hypothesis to real testing.

Meanwhile, the speed of model iteration continues to accelerate.

This is also why he wants to seize 1-2 years to implement safety assessments, third-party verifications, and global rules.

Whether this plan can be implemented is still unanswered.

Now, those at the forefront are starting to ask: If it truly runs out of control, can we still brake in time?

Reference Material:

https://darioamodei.com/post/we-must-pace-the-frontier

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink