The Odysseus Moment of China's Large Models

CN
5 hours ago

Written by: Evening Team




The expedition of China's large models.


Text by丨Li Geng


Edited by丨Tang Jianbo


On June 17, 2026, Zhipu released GLM-5.2, ranking third globally in the Artificial Analysis intelligence index, with UBS referring to it as a “milestone in the development of AI in China.” Less than a month later, in the late night of July 16, Moonlight released Kimi K3. With 28 trillion parameters, it topped the Frontend Code Arena with a score of 1679, and achieved 57 points in the Artificial Analysis intelligence index—third in the world, setting a historic high for open source. Within 48 hours of its release, user requests approached the physical limits of the computing cluster, prompting Moonlight to announce a pause on new subscriptions for consumers; within 72 hours, the AI sector in the U.S. evaporated $470 billion, and the Philadelphia Semiconductor Index dropped 12.5% in a week. The market called it the “DeepSeek Moment 2.0.” In the first week of August, on OpenRouter, the world's largest AI model invocation platform, 6 out of the top 10 models were from China. The top model, DeepSeek V4 Flash, processed 82 trillion tokens in a week, an increase of over 999% month-over-month. The strongest U.S. model, GPT-5.6 Luna, ranked fifth. 47% of the platform's users are U.S. developers, while Chinese users account for only 6%. a16z partner Martin Casado revealed to The Economist that among Silicon Valley AI startups seeking funding, the proportion using Chinese open source models might be as high as 80%. a16z summarized a widely circulated statement based on Kimi K3's paid subscription data from Sensor Tower: “America debates Chinese AI while the rest of the world buys it.” — Americans are arguing, the world is placing orders.


Capital is also “placing orders.” It is understood that the F round of Moonlight began as early as June, with most investors having finalized their decisions before the K3 release, completing the deal by the end of July, raising over $3.5 billion, equivalent to about 25 billion yuan (at the time's exchange rate). This investment list includes newcomers: the National AI Industry Investment Fund (sole lead), China Life, ICBC, ABC, CCB, CITIC Group, and CICC. Previous investors, such as Tencent, Meituan, Sequoia, IDG, CATL, and KKR, all chose to increase their investments.


From entering the global top three with GLM-5.2, to K3 topping the programming charts, to DeepSeek leading in invocation volume, Chinese open source models are transitioning from the margins to the mainstream. Compared to giants with massive capital and computing power, Chinese model companies resemble Odysseus facing giants—lacking brute strength, they can only rely on “wisdom,” looking for paths in resource efficiency, architectural innovation, open source ecology, and organizational capital methods.


Chinese models, just 3 points short, and 70% cheaper


Being overlooked is the “starting point” for Chinese large models. By the end of 2024, Chinese models accounted for less than 2% of token share on OpenRouter, regarded as “cheap but not good” alternatives. The first turning point began in early 2025. The DeepSeek-R1 inference model was open-sourced, sparking multiple waves of open-sourcing among domestic large model companies, gradually breaking the industry’s unspoken rules of “low-spec open source.” Over the past 12 months, domestic models maintained the global open source model scale limit for 9 months. In July this year, the Ministry of Industry and Information Technology disclosed that cumulative downloads of Chinese open source large models globally surpassed 10 billion. Hugging Face’s spring report showed that downloads of self-developed open source models in China accounted for 41%, surpassing the U.S. for the first time. Out of every 10 downloads of large models globally, 6 came from China.


The second turning point was the release of Kimi K3 in July this year. K3 has a total of 28 trillion parameters, with about 104 billion activated per inference—896 experts waking 16 each time, a sparsity of 56:1; the self-developed KDA mixed linear attention expanded the context window to 1 million tokens. In the Artificial Analysis intelligence index (comprising nine evaluations), K3 scored 57 points, ranking third out of 189 measured models at launch, just behind Claude Fable 5 (60 points) and GPT-5.6 Sol (59 points), achieving the highest score in the history of open source models—previously the strongest open source model GLM-5.2 scored 51, and DeepSeek V4 Pro scored 44. Analysts abroad summarized: on a scale where the difference between the frontier and the middle exceeds 15 points, a 3-point gap “equals a tie.” Individual item evaluations also highlight the issues. In blind tests on the frontend coding platform Frontend Code Arena, K3 topped with an Elo of 1679, securing six first places out of seven sub-fields; it surpassed Fable 5 in the officially announced long-term software engineering assessment SWE Marathon. On the day of release, Musk commented “Impressive,” and the next day xAI announced the launch of Grok4.6, with a parameter target aimed directly at 28 trillion. UC Berkeley professor Ion Stoica judged that the gap between Chinese open source models and top-tier closed models globally has shrunk from 6 to 9 months to 2 to 3 months. Hugging Face CEO Clément Delangue indicated that China holds a leading position in the field of open-source weight AI models and is expected to catch up with American leading model companies by the end of 2026 or in 2027.


Moonlight officially acknowledges that K3's overall experience still lags behind Fable 5 and GPT-5.6 Sol, yet the industry's pricing rules have been permanently altered. The per-token price is just a tag; companies actually calculate the total bill for completing a task (including retries, long contexts, and inference tokens). Two independent measures provide consistent scales: according to weighted empirical tests from Artificial Analysis, K3's cost per single task is $0.86, while GPT-5.6 Sol costs $1.23 and Fable 5 costs $3.15—K3 is about 27% of Fable 5, and 65%-70% cheaper under different metrics. The significance of K3 lies not in being “cheap”—the real value is that it offers intelligence comparable to existing frontier models at a notably lower cost. Plotting ability scores against single task costs on the same graph connects the highest capability points of each pricing tier. After K3's release, the “ability premium” pricing power of flagship closed-source models was substantively compressed for the first time. Anthropic responded swiftly: on July 24, they launched Claude Opus 5 at half the previous price. Fable 5, which had been taken offline “for being too powerful,” was directly pushed to the exit zone after coming back online: it was flanked by the newly half-priced Opus 5 and K3, priced at $0.86.


K3's pricing strategy itself indicates issues—non-cache outputs are priced at $15 per million tokens, 70% lower than Fable 5's $50. Intelligence index score of 57. It is the first model in the open-source camp to charge in the “capability tier,” rather than just competitively pricing in the “sufficient tier.” Each level upwards significantly increases the difficulty of overturning opponents: at similar price points, competitors are no longer just open source peers but rather the frontline defenses of Fable 5 and GPT-5.6; weaknesses in hallucination rates and response speeds will be amplified into renewal rate issues in high-value tasks.


Goldman Sachs' research report pointed out that K3 set the mixed API pricing at $2.3 per million tokens, signifying that domestic AI companies are moving from “price wars” to competing for “pricing power.” The industry’s accounting unit is transitioning from “per million tokens” to “per qualified delivery.” Huang Renxun, in an interview with Axios, offered a calm judgment: “The market misunderstood the impact of DeepSeek for the first time, and now they misunderstand the impact of Kimi again.” In his view, cheaper and more powerful AI models will drive broader adoption, further increasing demand for chips and data center infrastructure.


After K3's release, 17 Wall Street investment banks lowered their target prices for AI chip companies. The more profound impact fell on the EDA sector—Synopsys fell 7.85%, and Cadence fell 9.47%. The market interpreted the trigger as K3 independently completing the design, optimization, and verification of a 4mm² chip in 48 hours, aided by an open-source EDA toolchain.


Computing Power is Scarce, Giants Locking Out


In recent years, the most common standard for determining which large model company is leading has been one: financing. OpenAI completed approximately $122 billion in financing, with a valuation of $852 billion, while Anthropic pushed its valuation to $965 billion with $65 billion in Series H funding, with both companies together absorbing 43% of the global total in risk investment in the first half of this year. The scale of financing has become a proxy for capability. The problem is that this standard measures “how much has been raised” rather than “how much intelligence has been produced with that money.”


The complicated cross-holdings among Silicon Valley giants make it even harder to discern the value behind the financing figures. Google, Microsoft, and Amazon are both competitors and appear on each other's shareholder lists for their model companies; Nvidia is the chip supplier for all factions and also an investor in them. The most complete capital internal circulation occurs between Microsoft and OpenAI: Microsoft holds about 27%, with OpenAI committing to a further $250 billion procurement of Azure services, enabling Microsoft to expand data centers and procure chips, which in turn funds OpenAI.


The infrastructure construction cycle for AI is measured in decades, and the giants use equity binding to lock in long-term supply and demand for computing power, which is a rational arrangement to share risks. But the issue with this standard is apparent: it is held in the hands of interested parties themselves. The consequence is that it becomes increasingly difficult for outsiders to discern how much of these numbers reflect “demand” and how much reflects “narratives.” The cost of “inbreeding” is reflected on the balance sheet. The global capital spending on computing power has reached up to $3 to $4 trillion, and to achieve returns, it needs to correspond with annual revenues of at least $8 to $10 trillion for the entire industry; yet the most optimistic projections—covering cloud, subscriptions, and programming—only amount to around $3 trillion in annual revenue. There is a gap of nearly three times between the investment and the most optimistic upper revenue limits. In April 2026, Microsoft and OpenAI terminated their exclusive cloud partnership, with IP licensing changing to non-exclusive. The exclusivity that once supported the deep bond between the two parties began to loosen, redefining the most important certainty factor in this AI supply chain.


On the other end are Chinese companies. Zhipu has raised a cumulative $7 billion (including refinancing after going public); DeepSeek completed a spotlight first round of financing this year, also about $7 billion; Moonlight has completed multiple rounds of financing this year, totaling about $9.5 billion, approaching $10 billion. Even so, the combined financing of these three still doesn’t exceed a fraction of the recent financing rounds of foreign giants.



On May 31, the U.S. Department of Commerce expanded chip licensing controls to “any Chinese parent company buyer,” blocking offshore computing power channels. After K3's open-source release, domestic computing power platforms initiated a Day-0 adaptation race—Huawei's Ascend 950 super node natively supports K3's quantized weights, Alibaba Cloud's Zhenwu M890 completed adaptation on the same day, reducing first token latency by about 35%, and Haiguang's DCU fully adapted, with performance losses controlled within 12%. Domestic GPUs are stable for the first time when running trillion-parameter models. However, all adaptations are limited to inference, with no public breakthroughs on the pre-training side; under HBM constraints, domestic high-end chip production in 2026 is estimated to be only 250,000 to 300,000 units.


The Eight Immortals Cross the Sea, Competing on the Same Boat


The tighter the supply of computing power, the more vigorous the demand. K3 saw a consumer-side circuit breaker within 48 hours of launch, leading to a month of reservation number releases; Zhipu raised API prices by 83%, yet paid token consumption still increased fourfold month-over-month; OpenAI plans to invest about $600 billion in computing power before 2030. These phenomena point to the same structure: the gap is not on the demand side, but on the supply side. In the 19th century, Jevons made an counterintuitive judgment in “The Coal Question”: the efficiency increase of the steam engine did not reduce coal consumption; instead, it expanded total consumption—efficiency reduces usage costs, which releases pent-up demand. JPMorgan's June research report offered a contemporary version: the token price dropped more than 87% in a year, yet the agents' consumption skyrocketed exponentially. Lower prices did not kill demand; instead, they created greater demand. The greater demand expands, the tighter supply gets, the more excited capital becomes.


The slope of growth provides support. Anthropic’s $2 trillion target corresponds to an approximate tenfold annual ARR growth over the past three years (approximately $1 billion by the end of 2024, $47 billion by May 2026), translating to a forward price-to-sales ratio of 17-20 times, far below the 55 times of concept stocks like Palantir. Zhipu's ARR grew 15 times in half a year, and in mid-July, it announced it would hit the $1 billion guidance ahead of schedule, with its market value stabilizing at $75 billion; DeepSeek continued to exchange price for volume with V4 Flash while restarting a first-round financing valued at $70 billion.


Moonlight completed a financing round exceeding $3.5 billion on July 29, with a post-investment valuation of $35 billion, being closed early due to subscription amounts exceeding the initial target by three times; Pre-IPO round’s pre-investment valuation anchored at $50 billion. This valuation has been validated by the market: on the day K3 was released, global single-day revenue jumped eightfold; according to public information, on the day of weight open-source release, five paid inference service providers became its certified “launch partners” and simultaneously launched sales in North America, and the next day, Microsoft also announced it would list K3.


In the global capital contest, Chinese model companies are gaining more and more supporters. The simplest structure comes from DeepSeek. According to public information, its equity structure is highly concentrated: founder Liang Wenfeng controls it through partnerships; external shareholders are primarily the National AI Industry Investment Fund, along with a few others from Tencent, CATL, NetEase, and JD through partnerships, along with a few institutions like Monolith, IDG, Zhengxin Valley, and Shixiang. This is a “small but certain” list—sources are concentrated, but each investment is substantial.


Moonlight presents another contrast. Its investors almost touch every aspect of the Chinese capital map: National team: National AI Industry Investment Fund, social security fund; Chinese mobile internet giants: Alibaba, Tencent, Meituan; U.S. dollar funds: Sequoia, IDG, Monolith; Industrial capital: CATL, BAIC;


Banks and brokerages: ICBC, ABC, CCB, CITIC Group, CICC; International capital: KKR (Europe), Granite Asia (Singapore), Orix (Japan), Future Assets (Korea); Long-term capital: HKIC (Hong Kong Investment Management Company), Chow Tai Fook family. Completely different capitals are willing to put their chips in the same boat. Beyond the pier, more international capital is queuing up. There is an important intersection in the lists of the two: the National AI Industry Investment Fund. It controls a scale of hundreds of billions and first participated in Moonlight’s C round last year, then joined the much-anticipated first round financing of DeepSeek this year. As early as July this year, before Moonlight had released Kimi K3, this national team fund had already planned to lead Moonlight's F round.


There is more than one common shareholder. IDG, Tencent, and some LP investors behind GPs such as Jiuyan Medical and By-Health. The simultaneous moves of the national team can still be read as route layout, while the dual betting by market-oriented institutions approaches pure return calculations—between cost and performance paths, they have already given up the option. Zhipu's trillion-dollar sprint once left peers lamenting; now, the capital queuing at the pier, both domestic and foreign, are unwilling to miss the boat.


The open-source strait, converging paths


The open-source navigation channels, in the beginning, had various measures. China's open-source fleet has explored two routes in the same strait. One route is the cost-performance path taken by DeepSeek. In January 2025, when DeepSeek R1 first became popular worldwide, it had nothing to do with being cheap. Firstly, R1's mathematical reasoning performance equaled that of OpenAI's o1, with the former scoring 79.8 on AIME 2024, and the latter scoring 79.2; meanwhile, R1's API output price is only about 3% of o1's. Powerful performance at a very low cost led the industry to label DeepSeek R1 as the “Pinduoduo of AI.” Before the model exhibited cost-effectiveness, it first needed to be useful. Even when DeepSeek later actively initiated a price war, it didn’t automatically translate into sales. It was not until August 2026 that DeepSeek V4 Pro was officially launched, with official data showing it exceeded the Claude rival model in the CyberGym and AutomationBench tests; it scored 87.9 on Terminal Bench2.1, surpassing Opus4.8. With performance enhancement, DeepSeek implemented peak-valley pricing, raising output prices from 6 yuan to 27 yuan during peak times. The other route is the performance path represented by Kimi. K3, with an intelligence index scoring 57, directly intersects the capability tier price band, with an output price of $15 per million tokens, 70% lower than Fable 5. It is the first model in the open-source camp to charge in the “capability tier,” rather than engaging in price wars. Each upward tier significantly increases the difficulty of overturning competitors: the opponents are no longer just open-source peers but the frontline defenses of Fable 5 and GPT-5.6; weaknesses in hallucination rates and response speeds will be amplified into renewal rate challenges in high-value tasks. However, for top-tier capabilities to achieve large-scale applications, costs must continue to decline. K3's architectural innovation has improved expansion efficiency by about 2.5 times, activating only 104 billion parameters per inference, leaving room for subsequent large-scale applications. Yet K3 does not seem eager to lower prices: since its release, over 70% of new users come from overseas, with overseas subscription revenue increasing 14 times, and API revenue increasing 10 times. Different paths, common premise: whether through cost-performance or price increases, sufficient capability is essential.


The endpoints of both paths are in the enterprise market. The reality of this sea's water level is marked by two sets of numbers: one set shows rapid increases in usage—Chinese model token shares have risen from about 1.2% in October 2024 to over 60% now; total global downloads of Chinese open source large models have surpassed 10 billion. The other set reveals that in terms of enterprise production flow metrics (Vercel AI Gateway data from June), open-source models handle 29% of tokens, yet account for less than 4% of expenditures; Anthropic, with 32% of tokens, captures 61% of spending, and the four leading American labs collectively account for 95%. Usage has surpassed half; revenue is just beginning.


But the slope is changing: the open-source token share surged from 11% to 29% within two months. Moonlight's own transformation is isomorphic: Kimi has shifted from a consumer product company to focusing on APIs and the enterprise version, with API revenue now making up over 70% of total revenue. BlackRock estimates global corporate AI spending could reach $5-8 trillion by 2030, and China's AI-related industry scale, according to the National Development and Reform Commission, could exceed 10 trillion yuan by 2030, all falling into “non-aggressive” ranges. According to the classic script of disruptive innovation, challengers enter from the low end and move up the performance curve until surpassing the mainstream customer’s “sufficient performance line.” Over the past two years, China’s open-source models have grown from just 1.2% corner traffic to a 60% share; K3 is currently the only open-source model to have crossed this line.


Once capability is in place, price adjustments become credible.


The Path to AGI: One Company, or One Ecosystem


The essence of AGI may not be reaching the other side, but rather the return journey—returning to a certain form of civilization where humans do not have to work for survival. Renowned tech investor Gavin Baker revealed in the All-In podcast that he heard from multiple trusted sources that Anthropic CEO Dario Amodei once mentioned that Anthropic could become “the only private company in the world”—“in this vision, there is Anthropic, then governments, and nothing else.” Baker's reaction was twofold: he acknowledged that “they have checkpoints more advanced than Fable up their sleeves, executed very well,” while also stating, “I would definitely advise Dario to never say that again” and, “I would bet that they can't do it.”


Even the most confident giants are now publicly contemplating a world “with only one private company,” indicating the convergence at the table, needing no further argument. If this idea comes true, the disappearance of corporate forms may merely be a return of human civilization—modern corporate structure has only a history of about 400 years. But history has provided another path: a whole ecosystem exists to satisfy needs. One path may remain, or both may coexist. Windows dominated the desktop as a “commercially closed integrated system,” while Linux won the server, cloud computing, and supercomputing markets with an “open-source kernel and distribution ecosystem.” The mobile war pushed the separation of “volume” and “profit” to extremes: open-source Android captured about 70% of the global market share, while closed-source iOS had significantly fewer users but commanded much higher profits. Ecosystems win in “volume,” while companies win in “profit.”


Over time, super companies will also embrace ecosystems. Google defeated Microsoft's IE using the open-source Chromium kernel and later seized the mobile market with open-source Android; Microsoft transitioned from Ballmer's “Linux is a cancer” to Nadella’s “Microsoft loves Linux,” learning the lesson of “join when you can’t win” with twenty years of stock price. The era of large models is replaying this script, but with two new variables. One is that the ecosystem’s “volume” is coming faster than ever in history; the inference cost has dropped fiftyfold in three years, and open-source models have become mainstream choices on OpenRouter; second, the super company camp is showing a multi-headed pattern, with the dual champions OpenAI and Anthropic racing ahead, leaving behind the “previous generation of seven giants” such as Google, Meta, Microsoft, and Nvidia. The variables of the ecosystem are still increasing, while the converging variables do not seem to stop.


The industry's flywheel is brutally clear: to build the strongest model, one must buy the most computing power; to buy the most computing power, one must secure the largest financing; to get the largest financing, one must prove with ARR that “strong capability = high revenue.” The industry's concentration is not a result of market choice; it's a result of cost structure. The list is shrinking further, and there is no longer a Cursor in the world. On August 14, SpaceX completed the acquisition of Cursor for $60 billion—the fastest software company to achieve $4 billion in ARR in history, the sharpest independent player in the AI programming lane, will now merge into SpaceXAI, with its brand being gradually phased out. This company's flagship model was previously found by developers to be based on Kimi K2.5.


The open-source camp remains weak. Agent Arena evaluates models using real user tasks assigned randomly by the platform, which cannot be selected or predicted, and then uses causal inference to calculate the “net improvement” from millions of real behavioral signals—no fixed question banks to overfit, no voting to inflate results, and no reward models to please, making it difficult to reward hack (cheat) mechanism-wise. On this list, only one open-source model appears in the top 10. Weakness is not just reflected in the rankings: export controls have capped computing power limits, and the full-chain implementation of domestic chips is still in progress; the competitive environment faced by Chinese large model startups is more challenging than that of their counterparts.



Images are from Arena. To this day, from resources to markets, closed-source giants maintain absolute dominance, resembling deities. The endgame of AGI could converge into a single company, whether it is Anthropic or not, or it may leave behind a flourishing open-source ecosystem, resembling more of a value choice than a prediction. A subject that may impact the form of civilization is being counted with every use and investment. Everyone is part of this. Odysseus's ship has already set sail. Ahead lies the surging waves.


Cover image source: “The Odyssey”


- FIN -


免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink