After GPT-6, does the story of the surge in computing power make sense again?

CN
链捕手
Follow
6 hours ago

After the release of GPT-6 Astra on September 3, semiconductor and memory stocks rebounded first. On September 4, SOXX rose 3.5% in a single day, Micron increased by 6.1%, and SanDisk jumped 11.9%; on September 7, Korea's KOSPI rose 4.61%, Samsung Electronics was up 5.7%, and SK hynix rose by about 8%.

Prior to this, the market's concerns about overheated AI capital expenditure and peak demand for computing power had caused related sectors to fluctuate for quite some time. MarketWatch even described this rise as Astra “reigniting the memory-chip trade.”

Tae Kim, author of "The Nvidia Way" and former technology reporter for Barron’s, provided a more aggressive explanation: AI may be entering the fourth round of exponential demand growth for computing power in the past four years.

The first three rounds came from chatbots, reasoning, and coding agents, while the fourth round may be centered around the Computer Use showcased by Astra. Kim's logic is that chatbots helped hundreds of millions consume reasoning power, reasoning made each answer require more computation, and coding agents transformed models from "answering once" into working continuously for tens of minutes or even hours. Now, Computer Use extends the continuously running agents from programmers to ordinary knowledge workers in Excel, Blender, CAD, Power BI, and browsers.

Kim has long focused on Nvidia and semiconductor investments. His newsletter, Key Context, centers on judgments about technology investments and has leaned towards a bullish stance on AI infrastructure in recent articles.

Thus, “fourth round exponential growth” is primarily an investment hypothesis proposed by AI computing bulls and is not yet a validated industrial norm. However, its timing is quite delicate.

Before Astra's release, the agent paradigm had just undergone a phase of “capability hitting a wall”: short tasks became increasingly robust, but once in a long-term, multi-step, and real software environment, reliability still significantly declined, which also contributed to the earlier mentioned doubts about peak computing power. If the agent’s capability ceiling does not open, the narrative of continued growth in computing power falls apart.

Can Astra truly break this ceiling?

Every 1 human workday now corresponds to 3.1 Agent workdays

The first three rounds of AI "exponential growth" refer to the significant increase in computing capacity that a user can consume every time there is a change in the usage paradigm over the past few years.

ChatGPT introduced reasoning to the general market for the first time; reasoning models began to perform longer internal calculations for a single question; by the time of the coding agents, a simple instruction could trigger a continuous cycle of reading code, modifying it, testing, checking for errors, and modifying again.

Simultaneously, agents have begun to break through the physical limitation that “one person can only work 8 hours a day.”

On September 6, OpenAI just disclosed a set of internal data. By mid-August this year, among company researchers, those in the 50th percentile used computing power equivalent to over 600 dollars per day through coding agents; researchers in the 90th percentile exceeded 7000 dollars per day. According to Business Insider’s data compilation from OpenAI, the median value in July was only about 162 dollars, indicating an increase of nearly 3.7 times in just over a month.After GPT-6, does the story of power surge make sense again?Figure: By mid-August, every 1 human workday already corresponds to 3.1 Agent workdays; the median researcher uses Agent reasoning equivalent to over 600 dollars daily on an API price basis

This internal usage was calculated based on API pricing, making it more akin to “single user reasoning intensity.”

Prior to June, the total running time of all agents in the OpenAI research department did not exceed their own work hours; by mid-August, every 1 human workday now corresponds to 3.1 Agent workdays. More and more researchers are also running multiple agents simultaneously.

This means that after the agent paradigm, multiple “digital processes” capable of parallel operation are beginning to emerge behind one user, no longer constrained by the physical limitation of “human input speed.” Now, one person can keep three or four agents working continuously, or even have an agent create sub-agents.

The measurement unit for computing power demand may thus become: how many agents are running simultaneously behind one user and how many hours these agents work each day.

Computer Use expands the workspace

While coding agents have grown rapidly, they still have a clear boundary: most who run multiple agent tasks are still programmers, who are only a small part of knowledge workers globally.

Computer Use expands the space even further. After the release of Astra, Kim did a test himself. He previously did not know how to use Blender, so he let Astra research spacecraft, open Blender, and complete 3D modeling on a Mac. About 10 minutes later, a model that could be rotated and viewed was created.

Ordinary dialogue roughly follows the pattern of “input—reasoning—output”; coding agents became “read code—reason—modify—test—reason again”; computer use evolves further into “view screen—understand interface—decide actions—execute—wait for results—observe again—validate—correct mistakes.”

A task that takes a human ten minutes to complete may involve dozens of rounds of visual understanding, reasoning, and tool invocation behind the AI.

Once this capability transitions from browsers and code editors to Excel, Salesforce, SAP, Power BI, Photoshop, CAD, and various internal enterprise systems, the potential user base expands to nearly all knowledge workers sitting in front of computers.

This is the core of what Kim refers to as “fourth round demand”:

Coding agents seek to capture the computing time of programmers, while computer use aims to capture all white-collar workers' computing time.

But so far, this remains only a possibility.

OpenAI has not disclosed the amount of Computer Use tasks following the release of Astra, nor has it released daily curves for single-user token, GPU utilization, or reasoning throughput. Thus, it cannot be definitively proven that “Astra has brought about the fourth round of computing growth.”

Has computing power really begun to tighten?

After the release of Astra, Tencent Technology noted that some developer communities reported a feedback that is difficult to quantify but noteworthy: it seems that the user experience across different regions and accounts has begun to exhibit discrepancies.

Some developers in the Asia-Pacific region reported that ChatGPT and Codex have recently slowed down, and the stability of complex tasks is not as reliable as accounts from the U.S. region. The community has taken to referring to this experience directly as “dumb down.”

At least part of this is not entirely subjective.

On September 4, OpenAI's official status page separately reported a performance decline in the Asia-Pacific service, affecting several products including ChatGPT, Work, and Codex Cloud. Four days later, OpenAI announced a multi-year agreement with Firmus, a data center operator supported by Nvidia, to acquire dedicated computing power from two data centers in Malaysia.After GPT-6, does the story of power surge make sense again?

Currently, there is also no evidence to prove that OpenAI has reduced model reasoning budgets or switched to weaker models for Asia-Pacific users due to the increased load from Astra.

What is called “Asia-Pacific dumb down” could at least mix three distinct factors: regional infrastructure and routing issues, account-level risk controls and rate limits, and dynamic resource scheduling during peak times. All of these may manifest on the user side as slower response times, higher failure rates, tool invocation interruptions, or even worse final answers on complex tasks.

Yet, this feedback still provides another way to observe the pressure on AI computing power.

As models become increasingly capable of continuously consuming computation, supply tightness may first manifest as variations in latency, failure rates, and service quality among different regions and accounts.

The issue is that this judgment still lacks a key set of data:

For the same model, same plan, and same task, is there a consistent difference in first token latency, total reasoning time, tool success rates, and final task completion rates across different regions?

Until these data are obtained, “dumb down” can only serve as a clue, not as evidence of computing power tension.

After GPT-6, does the story of power surge make sense again?

Figure: On September 8, a user shared the prompt "capacity full" from GPT-6 Astra and claimed that all four of their accounts were unable to be used normally, with the system suggesting switching to other models. Tibo jokingly referred to this as “we are so back”, interpreting “model full” as a signal of renewed AI demand.

Capital markets have begun to place bets again

Although the actual computing power usage curve after the release of Astra has not been made public, the capital markets are already trading on the story that “AI demand has not peaked.”

MarketWatch reported that after the release of Astra, the iShares Semiconductor ETF briefly rose by approximately 3.5%, with memory-related companies such as Samsung, SK hynix, and Kioxia performing even better. Foreign media have directly described this round of market activity as Astra “reigniting the memory-chip trade.”

It is notable that in this round of growth, the most eye-catching is not Nvidia, but memory.

If Computer Use truly becomes a continuous agent workload, the increase required will not be limited to GPU computing. Longer contexts, KV Cache, concurrent agents, virtual machines, browsers, and software environments could continue to drive up demands for HBM, DRAM, CPU, networking, and storage.

Thus, what the capital markets have been trading recently is not entirely “GPT-6 has become more powerful”; the main logic is still: would a smarter model lead users to be willing to purchase more computation.

This also marks a fundamental difference between this round of AI infrastructure story and traditional software.

As traditional software becomes more optimized, the server resources needed by the same user may decrease.

Yet generative AI may exhibit the opposite “Jevons Paradox”: the more efficient the model and the higher the task success rate, the more willing people are to delegate more, longer, and more complex tasks to AI, resulting in overall computing power consumption continuing to increase.

However, it is important to remain cautious, as the capital markets also have their own “butts.”

AI infrastructure had previously experienced a round of significant increases and adjustments, and semiconductor stocks had seen marked withdrawals due to overheated capital expenditure and declining free cash flow for cloud vendors this year. Therefore, the few days of rebound brought by Astra can only indicate that investors are starting to bet again; it is not sufficient to prove that the fourth round of computing power demand has indeed emerged.

An economic account

For Kim's judgment to ultimately hold, it depends on three more realistic questions.

The first remains reliability.

Real enterprise work is not a demo once. Whether an agent can continuously operate ERP, financial models, or engineering software for several hours while maintaining correctness in the face of pop-ups, permission changes, and data updates is a completely different issue.

The second is cost.

What will it cost to complete one hour of equivalent work by a human computer in the future?

If AI can complete 50 dollars worth of work for an employee at the cost of 10 dollars, demand can easily explode; if it requires 100 dollars, the market space is entirely different.

The internal data of OpenAI showing 600 dollars/day suggests that agents can create huge demand, but may also imply that this work model remains very expensive today.

The third question: Will Computer Use ultimately eliminate some of its own demand?

Currently, models need to look at screens and find buttons because existing software is designed for humans. Once Salesforce, Microsoft, Adobe, or internal enterprise systems begin to open direct APIs, MCP, and agent interfaces, many tasks may not need to continue simulating mouse and keyboard operations.

The more likely final form is a fusion of Computer Use and Tool Use: directly calling where there are structured interfaces, while old systems and open webpages that lack interfaces or are difficult to transform continue to rely on visual operations.

Thus, the market truly opened up by Computer Use may not be about “making AI click on computers like a human forever,” but rather through the method of “my human master authorizes me to click your software, so you have no say,” enabling agents to forcibly break through the existing application ecosystem. Agents can finally enter the last large work environment that previous automation has not been able to cover.

This is also the part that Kim refers to as “the true verification of the fourth round of exponential growth,” questioning whether “Computer Use is a stepping stone or really another paradigm shift?”

After GPT-6, has the once difficult narrative of computing power due to agents hitting a wall become coherent again?

The capital markets have begun to place bets again.

After the explosion of generative AI, will the market continue to slap the face of the “computing power growth” shorts, or will this time be different?

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink